2.7 KiB
2.7 KiB
🤖 API, OpenWebUI & Thinking / Reasoning Control Guide
The API proxy (ollama-proxy.py) translates incoming Ollama, OpenAI, and Anthropic Messages API calls into optimized requests for the llama-server backend.
1. Controlling Reasoning & Thinking
By default, Qwen3.8 and DeepSeek reasoning models have their Jinja template default set to xhigh effort, which can cause excessive thinking on simple queries. The proxy provides full granular control over thinking modes.
Option A: Model Tag Aliases (Recommended for OpenWebUI Dropdown)
Select any of the registered alias tags directly in your client:
| Model Tag | Thinking Mode | Behavior |
|---|---|---|
qwen:fast / qwen:nothink |
Disabled (0 reasoning tokens) | Answers immediately with zero thinking delay! |
qwen / qwen:latest |
Low Effort (Default) | Concise, focused 1–3 sentence reasoning trace before answering. |
qwen:think / qwen:deep |
High Effort (xhigh) |
Full deep multi-step reasoning for complex math/coding. |
Option B: OpenWebUI UI Controls & Parameters
- Thinking Toggle: Toggle "Thinking" ON/OFF in the chat interface.
- Reasoning Effort Setting:
none/off\rightarrowThinking disabled.low\rightarrowBrief, focused reasoning.medium\rightarrowBalanced reasoning.high/xhigh\rightarrowDeep reasoning.
Option C: API Payload Parameters
1. Disabling Thinking (Ollama Format)
{
"model": "qwen",
"messages": [{"role": "user", "content": "What is 2+2?"}],
"options": {
"enable_thinking": false
}
}
2. Specifying Thinking Token Budget (Anthropic / OpenAI Format)
{
"model": "qwen",
"messages": [{"role": "user", "content": "Solve this equation: 3x + 12 = 45"}],
"thinking": {
"type": "enabled",
"budget_tokens": 512
}
}
2. Multimodal Vision Support
Send images directly via standard base64 strings in the images array (Ollama format) or image_url data URLs (OpenAI/Anthropic format).
The proxy features automatic magic-byte MIME detection supporting image/png, image/jpeg, image/webp, and image/gif.
3. Supported API Endpoints
- Ollama Endpoints:
POST /api/chat(Streaming & non-streaming)POST /api/generate(Streaming & non-streaming)GET /api/tagsPOST /api/showGET /api/psPOST /api/embed&POST /api/embeddings
- Anthropic Messages Endpoint:
POST /v1/messages(Claude Code, Continue.dev, Anthropic SDK)POST /v1/messages/count_tokens
- OpenAI Backend:
POST /v1/chat/completions(Forwarded directly tollama-server)