This website requires JavaScript.
Explore
Help
Sign In
wmantly
/
turing-multi-gpu-llm-server
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
7269a963133267acfc621b1446b77e09a4eeff3a
turing-multi-gpu-llm-server
/
scripts
T
History
wmantly
7269a96313
fix(proxy): dynamically report exact 128K context window (131072) and Q5_K_P quantization across all API endpoints
2026-09-18 02:10:00 +00:00
..
build-nccl-llama.sh
Initial commit: Complete deployment scripts, power governor, systemd units, and architecture documentation for Turing multi-GPU LLM rig
2026-09-01 01:55:19 +00:00
ollama-proxy.py
fix(proxy): dynamically report exact 128K context window (131072) and Q5_K_P quantization across all API endpoints
2026-09-18 02:10:00 +00:00
start-server.sh
fix(memory): tune cache-ram to 12288MB to prevent host RAM OOM-kills on 48GB servers
2026-09-18 02:07:18 +00:00