6.8 KiB
6.8 KiB
🔬 Hardware Architecture & Key Technical Learnings
This document details the practical hardware behaviors, quirks, and engineering solutions discovered while building and optimizing this 3-GPU Turing inference rig on an HPE ProLiant DL380p Gen8 server.
1. NVIDIA CMP 50HX Mining GPUs: VBIOS & Power States
Why CMP 50HX Mining Cards Idle at ~75W–85W Stock (The 100% Video Engine Pinning Bug)
- No Display / Headless Architecture: CMP 50HX mining cards lack physical display outputs and raster display engines.
- The 100% Video Engine Pinning Bug: The
.17(and newer.1A) CMP mining VBIOSes have a firmware bug where the GPU's hardware Video Engine (NVDEC) registers as 100% utilized at all times. This prevents the GPU power controller from automatically dropping into its lowest P8 sleep state, trapping stock idle power at roughly 75W–85W per card. - The Fix:
- On Windows: Use NVIDIA Inspector's Multi-Display Power Saver. Add AI backends (e.g.,
ollama_server.exeorpython.exe) to the high-performance threshold list to force the cards into true ~5W–8W idle states when inactive. - On Linux: Use our custom dynamic power governor (
power-governor.py) withnvidia-smiscript wrappers to manually clamp idle clocks to 300 MHz (-lgc 300,300) and apply a 150W power limit (-pl 150).
- On Windows: Use NVIDIA Inspector's Multi-Display Power Saver. Add AI backends (e.g.,
OEM Reference VBIOS vs. MSI VBIOS (90.02.60.00.17)
Comparing the hardware profiles and the cross-flashing outcome:
- GPU 2 (Original MSI Board
0x1462, VBIOS90.02.60.00.17):- Idles at
~31.8W – 33.7Wwhen core is locked to 300 MHz. - Minimum fan speed floor:
25%. - Aggressive P3 voltage gating tables.
- Idles at
- GPU 1 (Original NVIDIA Reference OEM
0x10DE, VBIOS90.02.60.00.01):- Idled at
~61.2W – 61.8Wstock (locked 40% fan curve floor, higher static VRM leakage).
- Idled at
- Cross-Flashing Result:
- Cross-flashed GPU 1 with the MSI VBIOS ROM (
vbios/msi_cmp50hx_90.02.60.00.17.rom) usingnvflash -6. - Verified Outcome: GPU 1 idle power immediately dropped from
61.8Wdown to31.89W(a ~30W direct reduction!), and fan curve synchronized to 25% minimum floor across both cards.
- Cross-flashed GPU 1 with the MSI VBIOS ROM (
2. Power Cap vs Inference Throughput
- 100W Power Cap (
-pl 100): Capping CMP 50HX cards to 100W throttles GPU 1's core clock down to 1,305 MHz during tensor-parallel splits, dropping generation speed from 18.0 tok/s\rightarrow~14.0 tok/s. - 150W Power Cap (
-pl 150): Provides sufficient headroom for full 1,800–1,900 MHz boost clocks, sustaining 18.03 tok/s while protecting against unnecessary 225W power spikes and heat.
3. Automated Dynamic Power Governor (power-governor.py)
- The Problem with GPU Utilization Polling: On Turing GPUs during batch=1 token generation,
utilization.gpufluctuates between 0% and 5% between token steps. Relying on GPU utilization causes false idle triggers. - The Solution: Direct slot polling against
llama-server(http://127.0.0.1:8080/slotsis_processing: True). - Governor Behavior:
- When
is_processing: True: Instantly sets unconstrained clocks (nvidia-smi -rgc)\rightarrow18.03 tok/s. - When idle for > 15 seconds: Sets low-power 300–600 MHz core clocks
\rightarrowdrops power to ~31W/card.
- When
4. HPE ProLiant DL380p Gen8 Platform Learnings
Sandy Bridge-EP (v1) vs. Ivy Bridge-EP (v2) PCIe 3.0 Jitter
- The Sandy Bridge-EP v1 Bug: Intel's first-generation PCIe 3.0 controller on
Xeon E5-2600 v1(2012) suffered from transmitter signal margin attenuation (Intel Errata BD78/BD105). Over riser cables, the 8.0 GHz eye diagram degrades, forcing links down to Gen 1 (2.5 GT/s). - The Ivy Bridge-EP v2 Fix:
Xeon E5-2600 v2(22nm Tri-Gate) completely redesigned the PCIe 3.0 PHY with CTLE equalization, reliably locking Gen 3 speeds across risers.
NUMA Socket Pinning
- Dual-socket Xeon servers have two distinct NUMA nodes.
- If all GPUs are plugged into Primary Riser 1, they are physically connected to CPU Socket 1 (NUMA Node 1).
- Binding
llama-serverto NUMA Node 1 (CPUAffinity&NUMAPolicy=bind) eliminates all host-to-device memory traffic over the cross-socket Intel QPI bus.
External PSU & Common Grounding
- When GPUs receive PCIe data signals from the server motherboard but 12V power from an external PSU:
- Ensure the server chassis and external PSU casing share a solid common ground to eliminate high-frequency ground loop noise on PCIe differential clock lines (
REFCLK).
6. Low-Latency NCCL Ring Buffer Tuning (2x Prefill Speedup)
- The Problem: Default NCCL ring allocations (4MB–8MB) introduce high PCIe bus synchronization latency on multi-GPU AllReduce barriers on Turing architecture (
sm_75). - The Solution:
export CUDA_DEVICE_MAX_CONNECTIONS=1 export NCCL_BUFFSIZE=2097152 # 2MB ring buffer (aligned to Turing L2 cache lines) export NCCL_NET_GDR_LEVEL=0 # Pure local PCIe bus routing export NCCL_P2P_DISABLE=0 # Direct P2P transfers export NCCL_ALGO=RING # Direct 1-stage ring AllReduce export NCCL_PROTO=SIMPLE # Eliminates LL128 packet framing overhead over PCIe - Measured Benchmark:
- Baseline Prefill Speed:
72.0 tok/s - Tuned NCCL Prefill Speed:
140.3 – 179.4 tok/s(+149% prompt processing throughput!)
- Baseline Prefill Speed:
7. Speculative Decoding & Multimodal (mmproj) Interaction
- The Conflict: In
llama.cpp, loading a secondary GGUF draft model (-md) simultaneously with a multimodal vision projector (--mmproj) causes an internal context initialization conflict insidelibmtmd.so. - The Resolution:
- Multimodal Mode: Use N-gram context-lookup speculation (
--spec-type ngram-simple --spec-ngram-simple-size-n 24 --spec-ngram-simple-size-m 4), which runs entirely in primary context with 0 additional VRAM and full vision support. - Pure Text Mode: Remove
--mmprojto enable secondary neural draft models (-md /opt/models/gguf/Qwen2.5-0.5B-Instruct-Q4_K_M.gguf).
- Multimodal Mode: Use N-gram context-lookup speculation (
8. 20GB VRAM Modding Feasibility (CMP 50HX TU102)
- Architecture: CMP 50HX uses the TU102 PCB layout (320-bit bus, 10 memory pads).
- Memory Swap:
- Desolder 10x 1GB (8Gbit) GDDR6 BGA-180 chips (e.g., Samsung
K4Z80325BC-HC14). - Solder 10x 2GB (16Gbit) GDDR6 BGA-180 chips (e.g., Samsung
K4ZAF325BM-HC14or MicronD9ZCL). - Modify memory strapping resistor dividers to signal 16Gbit density to the TU102 memory controller.
- Desolder 10x 1GB (8Gbit) GDDR6 BGA-180 chips (e.g., Samsung
- Payoff:
- 2x Modded Cards = 40 GB Total VRAM (capable of running full 70B/72B models like
Llama-3.3-70BorQwen2.5-72Bon just 2 dedicated x16 slots). - 4x Modded Cards = 80 GB Total VRAM (enterprise A100-tier capacity for under $1,000).
- 2x Modded Cards = 40 GB Total VRAM (capable of running full 70B/72B models like