Files

3.8 KiB

Multi-GPU Quantization Benchmark: Qwen3.8-27B vs Gemma4-26B-A4B (Turing SM75)

Comprehensive performance benchmark, memory footprint comparison, context window scaling, and architecture analysis across 3x NVIDIA Turing GPUs (RTX 2060 12GB + 2x CMP 50HX 10GB).


1. Hardware & System Architecture

  • Host Processor: Intel Xeon E5-2697 v2 (12C/24T, Ivy Bridge @ 2.70 GHz base, AVX1 support)
  • Active GPUs: 3x NVIDIA Turing GPUs (32,768 MiB physical VRAM total)
    • CUDA0 (Bus 21:00.0): NVIDIA GeForce RTX 2060 12GB (11,841 MiB usable) @ PCIe 3.0 x16
    • CUDA1 (Bus 24:00.0): NVIDIA CMP 50HX 10GB (9,798 MiB usable) @ PCIe 3.0 x16
    • CUDA2 (Bus 27:00.0): NVIDIA CMP 50HX 10GB (9,798 MiB usable) @ PCIe 3.0 x16
  • Interconnect: NVIDIA NCCL / Layer Split pipeline across 3 PCIe buses

2. Head-to-Head Benchmark Matrix: Dense vs MoE

Metric / Benchmark Stage Gemma4-26B-A4B-QAT @ 256K Ctx (Active Now) Qwen3.8-27B (Q5_K_P @ 192K) Qwen3.8-27B (Q4_K_P @ 256K)
Architecture Type Sparse MoE (128 experts / 8 active) Dense (64 layers) Dense (64 layers)
Active Params / Total Params 4.0B active / 26B total 27.5B active / 27.5B total 27.5B active / 27.5B total
Model Disk Size 15.64 GB (Q4_K_M) 20.22 GB (Q5_K_P) 17.92 GB (Q4_K_P)
KV Cache Precision q5_0 (5.5-bit) q5_0 (5.5-bit) q4_0 (4.5-bit)
Max Context Window 256K (262,144 tokens) 192K (196,608 tokens) 256K (262,144 tokens)
Multimodal Support Text + Vision + Video (mmproj-BF16) Text + Vision (mmproj-f16) Text + Vision (mmproj-f16)
Split Mode Layer Split (--split-mode layer) NCCL Tensor Split NCCL Tensor Split
Active Memory Footprint 9.09 GB / 9.25 GB / 7.67 GB 11.60 GB / 9.07 GB / 9.05 GB 11.75 GB / 8.89 GB / 8.87 GB
Total VRAM Allocated 26.00 GB (83%) 29.72 GB (95%) 29.51 GB (94%)
Free Headroom / GPU +2.75 GB / +0.55 GB / +2.13 GB +240 MB / +720 MB / +740 MB +90 MB / +905 MB / +923 MB
1. Short Prompt (100 tok)
• Prefill Speed 95.87 tok/s 53.66 tok/s 52.46 tok/s
• Decode Speed 46.00 tok/s 37.14 tok/s 38.12 tok/s
• Total Wall Time 2.44 s 3.76 s 3.47 s
2. Medium Context (~5k tok)
• Prefill Speed 2,279.22 tok/s (5.8x faster) 388.85 tok/s 391.23 tok/s
• Decode Speed 45.42 tok/s 33.81 tok/s 41.42 tok/s
• Total Wall Time 4.40 s (3.8x faster) 16.79 s 16.33 s
3. Long Context (~6k tok)
• Prefill Speed 2,374.57 tok/s (6.0x faster) 396.52 tok/s 398.61 tok/s
• Decode Speed 43.67 tok/s 35.99 tok/s 41.00 tok/s
• Total Wall Time 4.97 s (3.9x faster) 19.60 s 19.47 s
4. Sustained Decode (256 tok)
• Decode Speed 47.50 tok/s (+58% faster) 30.10 tok/s 31.97 tok/s
• Total Wall Time 5.82 s 10.71 s 10.17 s

3. Context Scaling Matrix on Gemma4-26B-A4B

Context Window KV Cache Quant Type Total VRAM Across Rig RTX 2060 12GB CMP 50HX #1 10GB CMP 50HX #2 10GB Status
64K (65,536) q5_0 (5.5-bit) 19.5 GB 6.8 GB 7.1 GB 5.6 GB Rock Solid
128K (131,072) q5_0 (5.5-bit) 22.0 GB 7.7 GB 7.9 GB 6.5 GB Rock Solid
192K (196,608) q5_0 (5.5-bit) 24.0 GB 8.4 GB 8.6 GB 7.0 GB Stable
256K (262,144) q5_0 (5.5-bit) 26.0 GB 9.09 GB 9.25 GB 7.67 GB Full Native Ceiling (+550MB margin)
192K (196,608) q8_0 (8.5-bit) 27.8 GB 9.8 GB 9.6 GB 8.4 GB Max Precision Ceiling