9.9 KiB
Connecting clients to the llama.cpp server
This guide covers wiring the Qwen3.8-27B llama.cpp server (OpenAI-compatible
API on :8080) into Open WebUI, opencode, Claude Code, and other tools.
Server: http://<HOST>:8080
Base URL: http://<HOST>:8080/v1 ← use this for OpenAI-style clients
Note: llama.cpp exposes an OpenAI-compatible API. It is not an Ollama-protocol server (no
GET /api/tags,/api/chat, etc.). Most tools accept OpenAI-style endpoints, so that's fine. If you need a true Ollama clone API, see §6.
0. Quick reference — endpoints
| Endpoint | Purpose |
|---|---|
GET /v1/models |
list models |
POST /v1/chat/completions |
chat (reasoning model: returns reasoning_content) |
POST /v1/completions |
raw completions |
POST /v1/embeddings |
embeddings |
GET /health |
liveness |
GET /props |
server params (context size, etc.) |
Auth is optional (no --api-key set). If you set one later, pass it as
Authorization: Bearer <key>.
1. Open WebUI
Open WebUI (openwebui.com) connects fine via its OpenAI API connection.
1.1 Docker (recommended)
docker run -d -p 3000:8080 \
-v open-webui:/app/backend/data \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:main
1.2 Point it at the llama.cpp server
- Open
http://<HOST>:3000and create an admin account. - Admin panel → Settings → Connections → OpenAI API.
- API Base URL:
http://<HOST>:8080/v1 - API Key: anything non-empty, e.g.
local - Enable if you want: "Enable Reasoning Content" (shows
thinkingblocks).
- API Base URL:
- Click refresh/save. The model
Qwen3.8-27B-Uncensored-...should appear in the model picker.
Notes:
- Open WebUI's Ollama connection type will not see this server (different protocol). Use the OpenAI API connection type.
- Optional: set
OPENAI_API_BASE_URL/OPENAI_API_KEYenv vars instead of the UI form.
2. opencode
opencode supports arbitrary OpenAI-compatible providers via the
@ai-sdk/openai-compatible driver.
2.1 opencode.json in your project
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"qwenlocal": {
"npm": "@ai-sdk/openai-compatible",
"name": "Qwen3.8-27B (local)",
"options": {
"baseURL": "http://<HOST>:8080/v1",
"apiKey": "local"
},
"models": {
"Qwen3.8-27B-Uncensored": {
"name": "Qwen3.8-27B Uncensored"
}
}
}
}
}
Then start with:
opencode
# model picker -> qwenlocal/Qwen3.8-27B-Uncensored
or force it per run:
opencode --model qwenlocal/Qwen3.8-27B-Uncensored
2.2 Global config (optional)
Put the same provider block in ~/.config/opencode/opencode.json to make it
available in every project.
Note: the model is a reasoning model — opencode will show the
reasoning_content stream as thinking output.
3. Claude Code
Claude Code speaks the Anthropic Messages protocol, so it needs a small translation layer to talk to llama.cpp's OpenAI API.
3.1 Use claude-code-router (CCR)
npm install -g @musistudio/claude-code-router
ccr --set-base-url http://<HOST>:8080/v1
ccr --set-provider openai
ccr
Point ANTHROPIC_BASE_URL at the router and run Claude Code as usual.
3.2 Alternative: a generic OpenAI→Anthropic proxy
Any tool that translates /v1/chat/completions (OpenAI) to the Anthropic
Messages shape, e.g. LiteLLM, works:
pip install litellm[proxy]
litellm --model openai/qwen3.8-27b --api_base http://<HOST>:8080/v1 --port 4000
# then: export ANTHROPIC_BASE_URL=http://localhost:4000
Caveat: the uncensored Q4_K_P model has no tools/functions baked in beyond basic chat — agentic tool-calling may be unreliable.
4. Other OpenAI-compatible clients
4.1 curl
curl http://<HOST>:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen",
"messages": [{"role":"user","content":"Hello"}],
"max_tokens": 200
}'
4.2 Python (openai SDK)
pip install openai
from openai import OpenAI
client = OpenAI(base_url="http://<HOST>:8080/v1", api_key="local")
resp = client.chat.completions.create(
model="qwen",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=200,
)
print(resp.choices[0].message.content)
# reasoning available as: resp.choices[0].message.reasoning_content
4.3 Node.js
npm i openai
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://<HOST>:8080/v1", apiKey: "local" });
const r = await client.chat.completions.create({
model: "qwen",
messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);
4.4 Local web UIs that accept an OpenAI endpoint
- AnythingLLM — Settings → LLM → "OpenAI" → custom base URL.
- LM Studio / Jan — treat the llama.cpp server as a remote OpenAI endpoint.
- SillyTavern — Chat Completion → Custom OpenAI → set base URL.
- Continue.dev (VS Code) —
models.yamlwith an openai provider + base URL.
5. Using the reasoning model properly
Qwen3.8-27B emits thinking internally before its response.
- OpenAI-compatible clients that surface
reasoning_content(Open WebUI, opencode) will show it automatically. - To disable/trim reasoning (lower latency, shorter answers), the chat template
supports
reasoning_effort:"low"|"medium"|"xhigh"(default). Most clients pass extrachat_template_kwargs; alternatively send it via thechat_template_kwargsfield or setenable_thinking: false:"chat_template_kwargs": {"reasoning_effort": "low"} - If a client shows raw
<thinking>/<response>tags, strip them, e.g. in Python:re.sub(r"</?thinking>", "", text).
6. "I want an Ollama clone API"
Done — a built-in proxy provides one. /opt/llama-server/ollama-proxy.py
fakes both the Ollama API (/api/chat, /api/generate, /api/tags,
/api/ps, /api/show, /api/embed, …) and the Anthropic Messages API
(/v1/messages, /v1/messages/count_tokens) on port 11434, translating
every request to the llama.cpp OpenAI backend on :8080. It runs as a systemd
service (ollama-proxy.service), pure Python stdlib, no third-party deps.
6.1 What the proxy exposes
| Endpoint | Purpose | Status |
|---|---|---|
GET /api/version |
version string | ✅ |
GET /api/tags |
list models (Qwen3.8-Uncensored) |
✅ |
GET /api/ps |
running models | ✅ |
GET /api/status |
cloud status (launcher) | ✅ |
GET /api/experimental/model-recommendations |
launcher hint | ✅ |
POST /api/show |
model details | ✅ |
POST /api/chat |
chat (stream + non-stream) | ✅ |
POST /api/generate |
single-prompt completion | ✅ |
POST /api/embed / /api/embeddings |
embeddings | ✅ |
POST /v1/messages |
Anthropic Messages API (Claude Code) | ✅ |
POST /v1/messages/count_tokens |
rough token estimate | ✅ |
GET /v1/models |
OpenAI-style model list | ✅ |
POST /v1/chat/completions |
OpenAI passthrough | ✅ |
Reasoning output from Qwen is exposed as reasoning_content on the Ollama
shape and thinking/text blocks on the Anthropic shape.
Context length: all four discovery endpoints (/api/tags, /api/ps,
/api/show, /v1/models) report context_length: 262144 so tools don't
down-scale to a default (e.g. 128k).
6.2 Usage
# any Ollama-native client, point it at this box
OLLAMA_HOST=http://192.168.1.198:11434 ollama run Qwen3.8-Uncensored
curl http://192.168.1.198:11434/api/chat -d '{
"model": "Qwen3.8-Uncensored",
"messages": [{"role": "user", "content": "hi"}]
}'
6.3 Claude Code via ollama launch claude
ollama launch claude makes Claude Code talk to the Anthropic /v1/messages
endpoint at OLLAMA_HOST. The proxy implements it, so:
OLLAMA_HOST="http://192.168.1.198:11434" \
CLAUDE_CODE_MAX_CONTEXT_TOKENS=65536 \
ollama launch claude --model Qwen3.8-Uncensored
Important:
- Use
192.168.1.198, this box's LAN IP — not a different address. Earlier guidance referenced.165, which is not this host. - The proxy presents
ANTHROPIC_BASE_URL=OLLAMA_HOST, so Claude Code talks directly to:11434/v1/messages. - If you don't use
ollama launch, the equivalent manual setup is:export ANTHROPIC_AUTH_TOKEN=ollama export ANTHROPIC_API_KEY= export ANTHROPIC_BASE_URL=http://192.168.1.198:11434 claude --model Qwen3.8-Uncensored - Tool calling: supported. Anthropic
tools,tool_use, andtool_resultblocks translate to/from llama.cpp OpenAI function calls (verified: the model returns propertool_useblocks, and multi-turn tool results are answered correctly). Streaming emitsinput_json_deltaevents for tool_use blocks.
6.4 Service management
systemctl status llama-server # backend (llama.cpp, :8080)
systemctl status ollama-proxy # API faker (:11434)
journalctl -u ollama-proxy -f # proxy logs
Both are enabled at boot. The proxy needs the backend up; the unit has
After=llama-server.service.
7. Troubleshooting
| Symptom | Fix |
|---|---|
| Connection refused | Server stopped (systemctl start llama-server); wrong host/port |
401 |
A --api-key was set; add Authorization: Bearer <key> |
| Model not in client list | Client cached models; hit refresh, or curl /v1/models to confirm |
| Slow first token | Reasoning model thinking; set reasoning_effort: "low" |
| OOM / VRAM errors | Reduce -c, or drop --cache-type-k/v q4_0 trade-offs (see README §11) |
| Client needs Ollama API | Use the built-in proxy on :11434 (§6) |
ollama launch <tool> fails "something went wrong" |
Check journalctl -u ollama-proxy -f; the launcher does HEAD / (heartbeat) + GET /api/tags first — both must return 200 |
| Tool calls return empty / no tool_use | Confirm the model supports tools via POST /api/chat with a tools array; the Q4_K_P GGUF does |