Commit Graph

  • 7269a96313 fix(proxy): dynamically report exact 128K context window (131072) and Q5_K_P quantization across all API endpoints main wmantly 2026-09-18 02:10:00 +00:00
  • ff29e72442 fix(memory): tune cache-ram to 12288MB to prevent host RAM OOM-kills on 48GB servers wmantly 2026-09-18 02:07:18 +00:00
  • aebc82e403 fix(proxy): handle client aborts and broken pipe disconnections gracefully without 500 errors wmantly 2026-09-16 00:34:14 +00:00
  • f5f8e5bb79 feat(cache): enable radix prefix sharing, chunk cache reuse, and 32GB disk-backed slot persistence wmantly 2026-09-15 17:03:08 +00:00
  • ae7c1e24e3 build(server): update launcher for upstream b11015 numa syntax wmantly 2026-09-14 14:07:20 +00:00
  • 05720793ad fix(proxy): add safe_json_loads to prevent 500 errors on streaming tool call decoding wmantly 2026-09-13 02:18:52 +00:00
  • cd989b3e70 feat: update to Q5_K_P @ 128K ctx, non-blocking proxy sessions, and full multi-GPU benchmark suite wmantly 2026-09-13 02:14:25 +00:00
  • cfeac35ad6 Configure HauhauCS FastMTP 32K draft head, reasoning parameters, and non-thinking mode wmantly 2026-09-06 22:46:11 +00:00
  • 2af141b38e fix(proxy): expose /props with vision modality and architecture meta in /v1/models for Maki local discovery wmantly 2026-09-05 02:42:20 +00:00
  • 096fb421d1 fix(proxy): add input_modalities: [text, image] in /v1/models for Maki vision discovery wmantly 2026-09-04 20:55:12 +00:00
  • ac66e70c89 feat: advertise clip family and vision capabilities in Ollama metadata and handle Anthropic image blocks wmantly 2026-09-04 20:51:11 +00:00
  • 41265cf9db feat: re-enable multimodal vision ViT projector offloaded to 12GB RTX 2060 (CUDA2) wmantly 2026-09-04 20:29:40 +00:00
  • 9ab0583cab feat: implement session-managed persistent KV cache architecture with slot persistence and management API wmantly 2026-09-04 02:53:54 +00:00
  • 5daafe24e8 refactor: remove container-side power governor in favor of host-level cmp-tune wmantly 2026-09-02 21:01:52 +00:00
  • e037f7a122 tuning: add stable-fast profile for 24/7 FlashAttention inference wmantly 2026-09-02 21:01:06 +00:00
  • 9e543f1291 docs(lxc): add NUMA socket pinning recommendations and benchmark metrics wmantly 2026-09-02 20:14:49 +00:00
  • d26bcb34d3 docs(benchmarks): add PCIe layout comparison report wmantly 2026-09-02 20:08:28 +00:00
  • 09f10e0b54 docs(benchmarks): add pre-hardware change PCIe layout baseline wmantly 2026-09-02 16:08:47 +00:00
  • ba0aaed7bb feat: integrate xrip/llama.cpp-avx1-numa-sm75 fork with NCCL & NUMA optimizations wmantly 2026-09-02 16:02:55 +00:00
  • 38b308c6fd Docs: Update README with 38.7-44.6 tok/s decode, 502 tok/s prefill, and 33W cluster idle power benchmarks wmantly 2026-09-02 01:04:49 +00:00
  • 4a03c75710 Perf: Enable native MTP speculative decoding (draft-n-max 2), achieving 38.73 tok/s (+120% speedup) at 64.8% draft acceptance wmantly 2026-09-02 01:00:57 +00:00
  • d1353e9b0d Perf: Add cmp-tune insane profile, update Proxmox LXC privileged container docs, and update benchmarks to 28.8 tok/s with PCIe Gen2 and unlocked 610.43.03 driver wmantly 2026-09-02 00:56:37 +00:00
  • d726501d71 Docs: Add master CMP Reverse Engineering Ecosystem & Community Map (projects, techniques, hardware mods, and AI relevance) wmantly 2026-09-02 00:07:59 +00:00
  • 53e2b110a9 Enhancement: Integrate NVAPI P8 deep idle downclocking into power governor script wmantly 2026-09-02 00:05:23 +00:00
  • 3edf469c43 Hardware: Add 20GB modding guide, cmp-pstate NVAPI tool, and document private NvAPI_GPU_SetForcePstate and PCIe retrain mechanisms wmantly 2026-09-02 00:02:24 +00:00
  • d51de0df6a Docs: Add cluster power optimization benchmark table and Linux headless power management guidelines wmantly 2026-09-01 19:31:56 +00:00
  • 2a21eae435 Hardware: Add MSI CMP 50HX VBIOS ROM and document the 100% Video Engine Pinning bug and cross-flash idle power drop wmantly 2026-09-01 19:24:00 +00:00
  • 5f261c191f Docs: Update HARDWARE_LEARNINGS and README with low-latency NCCL ring buffer benchmarks and speculative decoding findings wmantly 2026-09-01 16:06:32 +00:00
  • 261631457c Update start-server.sh with validated ngram context-lookup speculation and NCCL pipeline flags wmantly 2026-09-01 15:56:43 +00:00
  • 2b4e826437 Add sglang-inspired low-latency NCCL ring buffer and CUDA connection optimization flags to start-server.sh wmantly 2026-09-01 14:55:56 +00:00
  • c30f4772f9 Fix power governor active retention during long prompt evaluations and remove redundant proxy governor thread wmantly 2026-09-01 14:40:19 +00:00
  • aeb72cecfa Initial commit: Complete deployment scripts, power governor, systemd units, and architecture documentation for Turing multi-GPU LLM rig wmantly 2026-09-01 01:55:19 +00:00