13 Commits
Author SHA1 Message Date
wmantly ff29e72442 fix(memory): tune cache-ram to 12288MB to prevent host RAM OOM-kills on 48GB servers 2026-09-18 02:07:18 +00:00
wmantly f5f8e5bb79 feat(cache): enable radix prefix sharing, chunk cache reuse, and 32GB disk-backed slot persistence 2026-09-15 17:03:08 +00:00
wmantly ae7c1e24e3 build(server): update launcher for upstream b11015 numa syntax 2026-09-14 14:07:20 +00:00
wmantly cd989b3e70 feat: update to Q5_K_P @ 128K ctx, non-blocking proxy sessions, and full multi-GPU benchmark suite 2026-09-13 02:14:25 +00:00
wmantly cfeac35ad6 Configure HauhauCS FastMTP 32K draft head, reasoning parameters, and non-thinking mode 2026-09-06 22:46:11 +00:00
wmantly 41265cf9db feat: re-enable multimodal vision ViT projector offloaded to 12GB RTX 2060 (CUDA2) 2026-09-04 20:29:40 +00:00
wmantly 9ab0583cab feat: implement session-managed persistent KV cache architecture with slot persistence and management API 2026-09-04 02:53:54 +00:00
wmantly ba0aaed7bb feat: integrate xrip/llama.cpp-avx1-numa-sm75 fork with NCCL & NUMA optimizations 2026-09-02 16:02:55 +00:00
wmantly 4a03c75710 Perf: Enable native MTP speculative decoding (draft-n-max 2), achieving 38.73 tok/s (+120% speedup) at 64.8% draft acceptance 2026-09-02 01:00:57 +00:00
wmantly d1353e9b0d Perf: Add cmp-tune insane profile, update Proxmox LXC privileged container docs, and update benchmarks to 28.8 tok/s with PCIe Gen2 and unlocked 610.43.03 driver 2026-09-02 00:56:37 +00:00
wmantly 261631457c Update start-server.sh with validated ngram context-lookup speculation and NCCL pipeline flags 2026-09-01 15:56:43 +00:00
wmantly 2b4e826437 Add sglang-inspired low-latency NCCL ring buffer and CUDA connection optimization flags to start-server.sh 2026-09-01 14:55:56 +00:00
wmantly aeb72cecfa Initial commit: Complete deployment scripts, power governor, systemd units, and architecture documentation for Turing multi-GPU LLM rig 2026-09-01 01:55:19 +00:00