83 lines
2.7 KiB
Markdown
83 lines
2.7 KiB
Markdown
# 🤖 API, OpenWebUI & Thinking / Reasoning Control Guide
|
||
|
||
The API proxy (`ollama-proxy.py`) translates incoming **Ollama**, **OpenAI**, and **Anthropic Messages** API calls into optimized requests for the `llama-server` backend.
|
||
|
||
---
|
||
|
||
## 1. Controlling Reasoning & Thinking
|
||
|
||
By default, Qwen3.8 and DeepSeek reasoning models have their Jinja template default set to `xhigh` effort, which can cause excessive thinking on simple queries. The proxy provides full granular control over thinking modes.
|
||
|
||
### Option A: Model Tag Aliases (Recommended for OpenWebUI Dropdown)
|
||
|
||
Select any of the registered alias tags directly in your client:
|
||
|
||
| Model Tag | Thinking Mode | Behavior |
|
||
| :--- | :--- | :--- |
|
||
| **`qwen:fast`** / **`qwen:nothink`** | **Disabled** (0 reasoning tokens) | Answers immediately with zero thinking delay! |
|
||
| **`qwen`** / **`qwen:latest`** | **Low Effort** (Default) | Concise, focused 1–3 sentence reasoning trace before answering. |
|
||
| **`qwen:think`** / **`qwen:deep`** | **High Effort (`xhigh`)** | Full deep multi-step reasoning for complex math/coding. |
|
||
|
||
---
|
||
|
||
### Option B: OpenWebUI UI Controls & Parameters
|
||
|
||
* **Thinking Toggle**: Toggle "Thinking" ON/OFF in the chat interface.
|
||
* **Reasoning Effort Setting**:
|
||
* `none` / `off` $\rightarrow$ Thinking disabled.
|
||
* `low` $\rightarrow$ Brief, focused reasoning.
|
||
* `medium` $\rightarrow$ Balanced reasoning.
|
||
* `high` / `xhigh` $\rightarrow$ Deep reasoning.
|
||
|
||
---
|
||
|
||
### Option C: API Payload Parameters
|
||
|
||
#### 1. Disabling Thinking (Ollama Format)
|
||
```json
|
||
{
|
||
"model": "qwen",
|
||
"messages": [{"role": "user", "content": "What is 2+2?"}],
|
||
"options": {
|
||
"enable_thinking": false
|
||
}
|
||
}
|
||
```
|
||
|
||
#### 2. Specifying Thinking Token Budget (Anthropic / OpenAI Format)
|
||
```json
|
||
{
|
||
"model": "qwen",
|
||
"messages": [{"role": "user", "content": "Solve this equation: 3x + 12 = 45"}],
|
||
"thinking": {
|
||
"type": "enabled",
|
||
"budget_tokens": 512
|
||
}
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## 2. Multimodal Vision Support
|
||
|
||
Send images directly via standard base64 strings in the `images` array (Ollama format) or `image_url` data URLs (OpenAI/Anthropic format).
|
||
|
||
The proxy features automatic **magic-byte MIME detection** supporting `image/png`, `image/jpeg`, `image/webp`, and `image/gif`.
|
||
|
||
---
|
||
|
||
## 3. Supported API Endpoints
|
||
|
||
* **Ollama Endpoints**:
|
||
* `POST /api/chat` (Streaming & non-streaming)
|
||
* `POST /api/generate` (Streaming & non-streaming)
|
||
* `GET /api/tags`
|
||
* `POST /api/show`
|
||
* `GET /api/ps`
|
||
* `POST /api/embed` & `POST /api/embeddings`
|
||
* **Anthropic Messages Endpoint**:
|
||
* `POST /v1/messages` (Claude Code, Continue.dev, Anthropic SDK)
|
||
* `POST /v1/messages/count_tokens`
|
||
* **OpenAI Backend**:
|
||
* `POST /v1/chat/completions` (Forwarded directly to `llama-server`)
|