2.4 KiB
2.4 KiB
PCIe Layout & Performance Baseline (Pre-Hardware Change)
Date: 2026-09-02
Host: HPE ProLiant DL380p Gen8 (Dual Intel Xeon E5-2620 v2 @ 2.10GHz)
Driver: NVIDIA 610.43.03 (CUDA 13.3)
Server Binary: xrip/llama.cpp-avx1-numa-sm75 with native NCCL Ring AllReduce
Model: Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P.gguf (MTP Speculative n_max=2)
1. Baseline Hardware & PCIe Bus Layout
| GPU ID | Model | VRAM | Bus ID | Current Link | Max Link Capability | Current Power Limit |
|---|---|---|---|---|---|---|
| GPU 0 | NVIDIA GeForce RTX 2060 12GB | 12 GB | 0000:04:00.0 |
PCIe Gen 1 x16 | PCIe Gen 3 x16 | 184W (Max 215W) |
| GPU 1 | NVIDIA CMP 50HX 10GB | 10 GB | 0000:07:00.0 |
PCIe Gen 2 x4 | PCIe Gen 2 x16 | 150W (Max 225W) |
| GPU 2 | NVIDIA CMP 50HX 10GB | 10 GB | 0000:21:00.0 |
PCIe Gen 2 x4 | PCIe Gen 2 x16 | 150W (Max 225W) |
- Total VRAM: 32 GB GDDR6
- Idle Cluster Power: ~33W (11W / 9W / 13W)
- Model Cold-Start Load Time: 15.7 seconds
2. Baseline Inference Benchmarks
| Workload | Prompt Tokens | Completion Tokens | Prefill Speed | Decode Speed | Draft Acceptance Rate | Wall Clock Time |
|---|---|---|---|---|---|---|
| Short Prompt (Code Generation) | 69 tokens | 256 tokens | 14.8 tok/s (1825 ms) | 28.4 tok/s (8988 ms) | 61.4% (140/228) | 11.67 s |
| Medium Context (~1,000 tokens) | 1,074 tokens | 256 tokens | 276.6 tok/s (3731 ms) | 37.1 tok/s (6877 ms) | 63.4% (142/224) | 11.39 s |
| Long Context (~4,800 tokens) | 4,862 tokens | 256 tokens | 380.4 tok/s (12672 ms) | 38.8 tok/s (6571 ms) | 67.0% (146/218) | 20.23 s |
3. What to Watch For After PCIe Reseating / Slot Change
- RTX 2060 Link Generation:
- Current: Gen 1 x16 (2.5 GT/s)
- Target / Potential: Gen 3 x16 (8.0 GT/s) or Gen 3 x8 -> 3.2x to 6.4x bandwidth increase to GPU 0.
- CMP 50HX Link Width & Generation:
- Current: Gen 2 x4 (5.0 GT/s x4)
- Target / Potential: Gen 2 x8 or Gen 2 x16 (if plugged into x16 electrical slots or riser 2).
- NCCL Ring AllReduce Latency:
- Upgrading GPU 0 to Gen 3 and opening higher link widths on GPUs 1 & 2 will directly reduce AllReduce synchronization time between transformer layers, boosting both prefill tok/s and decode tok/s.