wmantly
|
2af141b38e
|
fix(proxy): expose /props with vision modality and architecture meta in /v1/models for Maki local discovery
|
2026-09-05 02:42:20 +00:00 |
|
wmantly
|
096fb421d1
|
fix(proxy): add input_modalities: [text, image] in /v1/models for Maki vision discovery
|
2026-09-04 20:55:12 +00:00 |
|
wmantly
|
ac66e70c89
|
feat: advertise clip family and vision capabilities in Ollama metadata and handle Anthropic image blocks
|
2026-09-04 20:51:11 +00:00 |
|
wmantly
|
41265cf9db
|
feat: re-enable multimodal vision ViT projector offloaded to 12GB RTX 2060 (CUDA2)
|
2026-09-04 20:29:40 +00:00 |
|
wmantly
|
9ab0583cab
|
feat: implement session-managed persistent KV cache architecture with slot persistence and management API
|
2026-09-04 02:53:54 +00:00 |
|
wmantly
|
5daafe24e8
|
refactor: remove container-side power governor in favor of host-level cmp-tune
|
2026-09-02 21:01:52 +00:00 |
|
wmantly
|
e037f7a122
|
tuning: add stable-fast profile for 24/7 FlashAttention inference
|
2026-09-02 21:01:06 +00:00 |
|
wmantly
|
9e543f1291
|
docs(lxc): add NUMA socket pinning recommendations and benchmark metrics
|
2026-09-02 20:14:49 +00:00 |
|
wmantly
|
d26bcb34d3
|
docs(benchmarks): add PCIe layout comparison report
|
2026-09-02 20:08:28 +00:00 |
|
wmantly
|
09f10e0b54
|
docs(benchmarks): add pre-hardware change PCIe layout baseline
|
2026-09-02 16:08:47 +00:00 |
|
wmantly
|
ba0aaed7bb
|
feat: integrate xrip/llama.cpp-avx1-numa-sm75 fork with NCCL & NUMA optimizations
|
2026-09-02 16:02:55 +00:00 |
|
wmantly
|
38b308c6fd
|
Docs: Update README with 38.7-44.6 tok/s decode, 502 tok/s prefill, and 33W cluster idle power benchmarks
|
2026-09-02 01:04:49 +00:00 |
|
wmantly
|
4a03c75710
|
Perf: Enable native MTP speculative decoding (draft-n-max 2), achieving 38.73 tok/s (+120% speedup) at 64.8% draft acceptance
|
2026-09-02 01:00:57 +00:00 |
|
wmantly
|
d1353e9b0d
|
Perf: Add cmp-tune insane profile, update Proxmox LXC privileged container docs, and update benchmarks to 28.8 tok/s with PCIe Gen2 and unlocked 610.43.03 driver
|
2026-09-02 00:56:37 +00:00 |
|
wmantly
|
d726501d71
|
Docs: Add master CMP Reverse Engineering Ecosystem & Community Map (projects, techniques, hardware mods, and AI relevance)
|
2026-09-02 00:07:59 +00:00 |
|
wmantly
|
53e2b110a9
|
Enhancement: Integrate NVAPI P8 deep idle downclocking into power governor script
|
2026-09-02 00:05:23 +00:00 |
|
wmantly
|
3edf469c43
|
Hardware: Add 20GB modding guide, cmp-pstate NVAPI tool, and document private NvAPI_GPU_SetForcePstate and PCIe retrain mechanisms
|
2026-09-02 00:02:24 +00:00 |
|
wmantly
|
d51de0df6a
|
Docs: Add cluster power optimization benchmark table and Linux headless power management guidelines
|
2026-09-01 19:31:56 +00:00 |
|
wmantly
|
2a21eae435
|
Hardware: Add MSI CMP 50HX VBIOS ROM and document the 100% Video Engine Pinning bug and cross-flash idle power drop
|
2026-09-01 19:24:00 +00:00 |
|
wmantly
|
5f261c191f
|
Docs: Update HARDWARE_LEARNINGS and README with low-latency NCCL ring buffer benchmarks and speculative decoding findings
|
2026-09-01 16:06:32 +00:00 |
|
wmantly
|
261631457c
|
Update start-server.sh with validated ngram context-lookup speculation and NCCL pipeline flags
|
2026-09-01 15:56:43 +00:00 |
|
wmantly
|
2b4e826437
|
Add sglang-inspired low-latency NCCL ring buffer and CUDA connection optimization flags to start-server.sh
|
2026-09-01 14:55:56 +00:00 |
|
wmantly
|
c30f4772f9
|
Fix power governor active retention during long prompt evaluations and remove redundant proxy governor thread
|
2026-09-01 14:40:19 +00:00 |
|
wmantly
|
aeb72cecfa
|
Initial commit: Complete deployment scripts, power governor, systemd units, and architecture documentation for Turing multi-GPU LLM rig
|
2026-09-01 01:55:19 +00:00 |
|