8.3 KiB
8.3 KiB
🔬 Hardware Architecture & Key Technical Learnings
This document details the practical hardware behaviors, quirks, and engineering solutions discovered while building and optimizing this 3-GPU Turing inference rig on an HPE ProLiant DL380p Gen8 server.
1. NVIDIA CMP 50HX Mining GPUs: VBIOS & Power States
Why CMP 50HX Mining Cards Idle at ~75W–85W Stock (The 100% Video Engine Pinning Bug)
- No Display / Headless Architecture: CMP 50HX mining cards lack physical display outputs and raster display engines.
- The 100% Video Engine Pinning Bug: The
.17(and newer.1A) CMP mining VBIOSes have a firmware bug where the GPU's hardware Video Engine (NVDEC) registers as 100% utilized at all times. This prevents the GPU power controller from automatically dropping into its lowest P8 sleep state, trapping stock idle power at roughly 75W–85W per card. - The Fix:
- On Windows: Use NVIDIA Inspector's Multi-Display Power Saver. Add AI backends (e.g.,
ollama_server.exeorpython.exe) to the high-performance threshold list to force the cards into true ~5W–8W idle states when inactive. - On Linux: Use our custom dynamic power governor (
power-governor.py) withnvidia-smiscript wrappers to manually clamp idle clocks to 300 MHz (-lgc 300,300) and apply a 150W power limit (-pl 150).
- On Windows: Use NVIDIA Inspector's Multi-Display Power Saver. Add AI backends (e.g.,
OEM Reference VBIOS vs. MSI VBIOS (90.02.60.00.17)
Comparing the hardware profiles and the cross-flashing outcome:
- GPU 2 (Original MSI Board
0x1462, VBIOS90.02.60.00.17):- Idles at
~31.8W – 33.7Wwhen core is locked to 300 MHz. - Minimum fan speed floor:
25%. - Aggressive P3 voltage gating tables.
- Idles at
- GPU 1 (Original NVIDIA Reference OEM
0x10DE, VBIOS90.02.60.00.01):- Idled at
~61.2W – 61.8Wstock (locked 40% fan curve floor, higher static VRM leakage).
- Idled at
- Cross-Flashing Result:
- Cross-flashed GPU 1 with the MSI VBIOS ROM (
vbios/msi_cmp50hx_90.02.60.00.17.rom) usingnvflash -6. - Verified Outcome: GPU 1 idle power immediately dropped from
61.8Wdown to31.89W(a ~30W direct reduction!), and fan curve synchronized to 25% minimum floor across both cards.
- Cross-flashed GPU 1 with the MSI VBIOS ROM (
📊 Cluster Power Optimization Benchmark Table
| Optimization State | Per-Card Draw (CMP 50HX) | RTX 2060 Draw | 3x Cluster Idle Total | Idle VRAM Temp |
|---|---|---|---|---|
| Stock (Bugged VBIOS Idle) | 75W – 85W | 11.5W | ~165W – 185W (240W before flash) | ~55°C – 62°C |
| Linux Governor (P3 Clamped @ 300MHz) | 31.8W – 33.7W |
10.9W |
~76.5W ⚡ (−108W wall savings) |
27°C – 39°C |
| Ideal P-State Sleep (P8 Deep Sleep) | ~5W – 15W | ~10W | ~25W – 40W | ~25°C – 30°C |
2. Power Cap vs Inference Throughput
- 100W Power Cap (
-pl 100): Capping CMP 50HX cards to 100W throttles GPU 1's core clock down to 1,305 MHz during tensor-parallel splits, dropping generation speed from 18.0 tok/s\rightarrow~14.0 tok/s. - 150W Power Cap (
-pl 150): Provides sufficient headroom for full 1,800–1,900 MHz boost clocks, sustaining 18.03 tok/s while protecting against unnecessary 225W power spikes and heat.
3. Automated Dynamic Power Governor (power-governor.py)
- The Problem with GPU Utilization Polling: On Turing GPUs during batch=1 token generation,
utilization.gpufluctuates between 0% and 5% between token steps. Relying on GPU utilization causes false idle triggers. - The Solution: Direct slot polling against
llama-server(http://127.0.0.1:8080/slotsis_processing: True). - Governor Behavior:
- When
is_processing: True: Instantly sets unconstrained clocks (nvidia-smi -rgc)\rightarrow18.03 tok/s. - When idle for > 15 seconds: Sets low-power 300–600 MHz core clocks
\rightarrowdrops power to ~31W/card.
- When
4. HPE ProLiant DL380p Gen8 Platform Learnings
Sandy Bridge-EP (v1) vs. Ivy Bridge-EP (v2) PCIe 3.0 Jitter
- The Sandy Bridge-EP v1 Bug: Intel's first-generation PCIe 3.0 controller on
Xeon E5-2600 v1(2012) suffered from transmitter signal margin attenuation (Intel Errata BD78/BD105). Over riser cables, the 8.0 GHz eye diagram degrades, forcing links down to Gen 1 (2.5 GT/s). - The Ivy Bridge-EP v2 Fix:
Xeon E5-2600 v2(22nm Tri-Gate) completely redesigned the PCIe 3.0 PHY with CTLE equalization, reliably locking Gen 3 speeds across risers.
NUMA Socket Pinning
- Dual-socket Xeon servers have two distinct NUMA nodes.
- If all GPUs are plugged into Primary Riser 1, they are physically connected to CPU Socket 1 (NUMA Node 1).
- Binding
llama-serverto NUMA Node 1 (CPUAffinity&NUMAPolicy=bind) eliminates all host-to-device memory traffic over the cross-socket Intel QPI bus.
External PSU & Common Grounding
- When GPUs receive PCIe data signals from the server motherboard but 12V power from an external PSU:
- Ensure the server chassis and external PSU casing share a solid common ground to eliminate high-frequency ground loop noise on PCIe differential clock lines (
REFCLK).
6. Low-Latency NCCL Ring Buffer Tuning (2x Prefill Speedup)
- The Problem: Default NCCL ring allocations (4MB–8MB) introduce high PCIe bus synchronization latency on multi-GPU AllReduce barriers on Turing architecture (
sm_75). - The Solution:
export CUDA_DEVICE_MAX_CONNECTIONS=1 export NCCL_BUFFSIZE=2097152 # 2MB ring buffer (aligned to Turing L2 cache lines) export NCCL_NET_GDR_LEVEL=0 # Pure local PCIe bus routing export NCCL_P2P_DISABLE=0 # Direct P2P transfers export NCCL_ALGO=RING # Direct 1-stage ring AllReduce export NCCL_PROTO=SIMPLE # Eliminates LL128 packet framing overhead over PCIe - Measured Benchmark:
- Baseline Prefill Speed:
72.0 tok/s - Tuned NCCL Prefill Speed:
140.3 – 179.4 tok/s(+149% prompt processing throughput!)
- Baseline Prefill Speed:
7. Speculative Decoding & Multimodal (mmproj) Interaction
- The Conflict: In
llama.cpp, loading a secondary GGUF draft model (-md) simultaneously with a multimodal vision projector (--mmproj) causes an internal context initialization conflict insidelibmtmd.so. - The Resolution:
- Multimodal Mode: Use N-gram context-lookup speculation (
--spec-type ngram-simple --spec-ngram-simple-size-n 24 --spec-ngram-simple-size-m 4), which runs entirely in primary context with 0 additional VRAM and full vision support. - Pure Text Mode: Remove
--mmprojto enable secondary neural draft models (-md /opt/models/gguf/Qwen2.5-0.5B-Instruct-Q4_K_M.gguf).
- Multimodal Mode: Use N-gram context-lookup speculation (
8. 20GB VRAM Modding Feasibility (CMP 50HX TU102)
- Architecture: CMP 50HX uses the TU102 PCB layout (320-bit bus, 10 memory pads).
- Hardware Swap: Desolder 10x 1GB (8Gbit) GDDR6 BGA-180 chips and replace with 10x 2GB (16Gbit) chips (e.g., Samsung
K4ZAF325BM-HC14or MicronD9ZCL). Adjust memory strapping resistor dividers. - Driver Patching for 20GB (WPR2 Fix): See
docs/20GB_MOD_GUIDE.mdfor details on patching the NVIDIA open kernel driver to bypass the0x8dFWSEC error. - Capacity Scaling:
- 2x Modded Cards = 40 GB Total VRAM (runs full 70B/72B models on 2 slots).
- 4x Modded Cards = 80 GB Total VRAM (enterprise A100-tier capacity for under $1,000).
9. Private NVAPI NvAPI_GPU_SetForcePstate (0x025BFB10)
- The Discovery: While
nvidia-smicannot force P8 state on headless Linux, the private NVAPI entry pointNvAPI_GPU_SetForcePstate(0x025BFB10) insidelibnvidia-api.so.1can explicitly force the GPU into P8 state (645 MHz core / 405 MHz memory). - Implementation: We packaged
tools/cmp-pstate.pyto directly invoke this entry point per PCI Bus ID. - Restoration: Passing
--pstate 16returns power management control to unconstrained P0/P3 states before active inference.
10. Automated PCIe Link Retraining Service (cmp50hx-gen2)
- The Boot Lock: CMP 50HX cards feature an internal PCIe speed lock register that takes a short period after boot to become accessible.
- The Solution: A systemd boot service triggers a link retrain on the upstream PCIe port after boot, ensuring the bus reliably trains at full speed without manual re-initialization.