This website requires JavaScript.
Explore
Help
Sign In
wmantly
/
turing-multi-gpu-llm-server
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
Packages
Projects
Releases
Wiki
Activity
Files
f5f8e5bb797e5cc6d6a62830afe68676ade39d6d
turing-multi-gpu-llm-server
/
scripts
T
History
wmantly
f5f8e5bb79
feat(cache): enable radix prefix sharing, chunk cache reuse, and 32GB disk-backed slot persistence
2026-09-15 17:03:08 +00:00
..
build-nccl-llama.sh
Initial commit: Complete deployment scripts, power governor, systemd units, and architecture documentation for Turing multi-GPU LLM rig
2026-09-01 01:55:19 +00:00
ollama-proxy.py
fix(proxy): add safe_json_loads to prevent 500 errors on streaming tool call decoding
2026-09-13 02:18:52 +00:00
start-server.sh
feat(cache): enable radix prefix sharing, chunk cache reuse, and 32GB disk-backed slot persistence
2026-09-15 17:03:08 +00:00