Configuration

Settings

Wire up endpoints and dial in your cost/quality trade-off.

Local Endpoint

Models detected on this PC via Ollama

Detected Local Models

Download a Model

One-click pull from Ollama — Gemma, Qwen, Llama, Phi, Mistral, DeepSeek

Downloaded models appear in Detected Local Models above and can be selected instantly. Larger models than your GPU fits will still run, but slower (CPU offload).

Remote Endpoint

Frontier model for hard prompts

Routing Threshold

Higher = more aggressive local routing (cheaper). Lower = safer cloud fallback.

Confidence Threshold
0.72
0.0 · always cloud0.5 · balanced1.0 · always local