Providers & models
Connect OAuth subscriptions, API-key providers, or a local model — and switch models any time.
A provider is the service hosting AI models; a model is the specific one you talk to. Tau ships with several built-in providers and lets you add your own OpenAI-compatible endpoints (including local models).
The fastest setup: /login
Start Tau and use /login to connect a provider. The provider picker includes
a search field, which is especially useful for the longer API-key provider list:
tau
/login # choose a login method
/login openai # save an OpenAI API key
/login openai-codex # authenticate a Codex/ChatGPT subscription via OAuth
/login anthropic-subscription # authenticate Claude Pro/Max via OAuth
/login anthropic-api # save an Anthropic API key
/login github-copilot # authenticate GitHub Copilot with a device code
/login opencode-go # save an OpenCode Go API key
/login nvidia # save an NVIDIA NIM API key
/login custom # add an OpenAI-compatible custom provider
Built-in providers include OpenAI, Anthropic, OpenAI Codex (subscription), GitHub Copilot, OpenCode Go, OpenCode Zen, Moonshot AI (Kimi), Kimi Code (subscription), OpenRouter, Hugging Face, and NVIDIA NIM.
OAuth subscriptions
Choose Subscription / OAuth in /login for:
| Tau provider | Login flow | Prerequisite |
|---|---|---|
openai-codex | Browser callback with pasted-code fallback | A supported ChatGPT/Codex subscription |
anthropic | Browser callback with PKCE and pasted-code fallback | Claude Pro/Max with Anthropic extra usage available |
github-copilot | GitHub device code | An active Copilot plan; organization policy must allow the selected model |
GitHub Copilot asks for a GitHub Enterprise Server URL/domain. Leave it blank
for github.com. Device login also works in SSH/headless sessions: open the
shown verification URL on any device and enter the displayed code.
Anthropic uses distinct direct-login aliases so the authentication method is
unambiguous: /login anthropic-subscription starts OAuth, while
/login anthropic-api saves an API key. The top-level /login picker still
lists Anthropic under both Subscription / OAuth and API key. OAuth
subscription requests use Anthropic’s required
Claude Code identity and may be billed as extra usage rather than consuming
ordinary Claude plan limits. Check Anthropic’s current account terms before
using it.
OAuth tokens refresh automatically. /logout removes Tau’s local credential,
but does not revoke the grant remotely; use the provider’s account settings for
remote revocation.
OpenAI prompt caching
For direct OpenAI API and Codex OAuth sessions, Tau sends a stable session-derived prompt-cache key with every model request, including continuations after tool calls. OpenAI can use that key to keep successive append-only requests on the same cache path, improving reuse of the system prompt, tool schemas, and conversation prefix. Resuming a Tau session reuses its key; starting or branching a session creates a new one.
Requests remain stateless: Tau keeps provider storage disabled and resends the complete transcript. The key improves cache affinity but cannot preserve a hit if the prefix changes or the provider cache expires. OpenAI-compatible gateways do not receive these fields unless their catalog compatibility settings explicitly opt in. The TUI sidebar’s latest-request cache rate shows whether the most recent request actually hit.
Anthropic prompt caching
Tau marks cache breakpoints on Anthropic requests so the system prompt, tool schemas, and conversation history are reused between turns instead of being reprocessed. Which retention Tau asks for depends on how you authenticated:
- Claude Pro/Max via OAuth requests the one-hour cache. Subscription auth is not billed per token, and the five-minute default is shorter than a build, a test run, or the time it takes to read a diff — any of which would otherwise expire the cache mid-session.
- An Anthropic API key uses the five-minute default, because one-hour cache writes cost more per token and that should be a deliberate choice.
Providers that speak the Anthropic protocol through a gateway rather than being
Anthropic itself — minimax, minimax-cn, fireworks, and vercel-ai-gateway —
send no cache breakpoints, since not every gateway accepts them. Watch the
sidebar’s cache hit rate to see caching working; see
The interactive session for how to read it.
Anthropic reports input_tokens as the fresh, uncached portion after the last
cache breakpoint. Tau preserves that value as fresh input, retains the separate
cache-read and cache-write counters, and reports total input as their sum. This
keeps usage and cost displays aligned with Anthropic’s response contract.
Codex subscription context limits
OpenAI’s public API and the ChatGPT/Codex subscription are separate serving surfaces. A model with the same ID can have a smaller, rollout-specific context window through Codex OAuth than through an API key. For example, the public GPT-5.6 Sol API advertises a 1.05M-token window, while Codex has advertised substantially smaller limits through its authenticated model catalog.
Tau loads the last successful account-specific snapshot from
~/.tau/codex-models-store.json when a session starts, so models discovered in a
previous session are immediately available without opening a picker. It refreshes
the authenticated catalog in the background when /model or /scoped-models
opens. Since OpenAI filters the response by official-client version, Tau resolves
the current stable @openai/codex release from the npm registry and caches that
safe, non-secret version for four hours at ~/.tau/codex-version-store.json.
Failed lookups use the stale cached version or Tau’s bundled fallback. A
successful refresh replaces the Codex model picker inventory, so newly enabled
and newly client-gated models appear without a Tau release. Models unavailable to
the account are not copied from the separate public API catalog. Tau also uses
reported context windows, compaction thresholds, input modalities, and reasoning
efforts when present.
If discovery is unavailable or invalid, Tau retains the last account-matched
snapshot, or the checked-in Codex model list when no snapshot is available; it
does not reuse the public API inventory or limits. /session reports whether the
active context value came from the live provider catalog or Tau’s configured
fallback. The model cache contains only parsed account/model metadata (no OAuth
tokens), is scoped to the saved Codex account ID, and is never written to
catalog.toml, providers.json, or the models.dev cache.
Starting or resuming a live-only Codex model discovers the account inventory
before validating the selection. This requires successful discovery; offline
startup remains available for static models. If a later refresh omits the active
model, Tau preserves its runtime metadata without adding it back to the picker.
Resume restores the transcript’s saved model selection. If an older session
restores a stale selection, resume it and explicitly choose the intended model
in /model once to record the correction.
Live limits can vary by account or rollout and may change independently of Tau.
A discovery failure is non-fatal: Tau reports it in /session and continues with
the fallback. Direct OpenAI API sessions retain the context limits documented on
the API model page. Vision-capable Codex models retain their image-input
metadata separately from these runtime context limits, allowing image files read
by Tau to reach the model. The gpt-5.6 alias, which routes to GPT-5.6 Sol, is only
available through the direct OpenAI API; Codex subscription users should select
the explicit gpt-5.6-sol model instead. Tau tombstones the API-only alias for
the Codex provider, so older user catalog overlays and saved preferences cannot
restore it after an upgrade.
OpenCode Go and Zen
OpenCode Go and OpenCode Zen are API-key providers, not OAuth providers. Sign in at the OpenCode console, subscribe to Go or fund Zen, copy the API key, and then run:
/login opencode-go # subscription limits; https://opencode.ai/zen/go/v1
/login opencode # Zen pay-as-you-go; https://opencode.ai/zen/v1
Both can also read OPENCODE_API_KEY. Tau stores their saved credentials under
separate opencode-go and opencode names, allowing different keys when
needed. Available models and plan limits change over time; consult the
OpenCode Go and
OpenCode Zen pages for the current list.
Z.AI
Log in with /login zai or set ZAI_API_KEY. Tau sends Z.AI’s provider-specific
thinking object rather than an OpenAI reasoning_effort field:
{"thinking": {"type": "enabled"}} for an enabled logical mode and
{"thinking": {"type": "disabled"}} for off. This keeps configured
thinking enabled for GLM models whose endpoint does not accept the raw effort
field.
Z.AI documents reasoning_effort only for GLM-5.2 and newer, with model-specific
values. Tau therefore emits that additional field only when the selected model’s
compatibility metadata explicitly supports it; unsupported providers and models
continue to omit it. See the Z.AI deep-thinking
reference and chat-completion
schema for the authoritative
wire contract.
Hugging Face Inference Providers
Log in with /login huggingface or set HF_TOKEN. Tau’s Hugging Face model
list is generated at build time from models.dev and routed
through https://router.huggingface.co/v1. It includes tool-capable models from
DeepSeek, Gemma, GLM, GPT OSS, Kimi, Llama, MiniMax, MiMo, Qwen, Step, and other
families. Use /model to search the generated list; availability and the backing
inference provider can vary over time and by account.
Like Pi, Tau’s released snapshot includes model names, limits, costs, modalities,
reasoning support, and verified effort values. Thinking levels are off,
minimal, low, medium, high, xhigh, and max; max is distinct from
xhigh. Empty or toggle-only reasoning options do not replace provider/manual
behavior. Hugging Face zai-org/GLM-5.2 currently uses a narrow verified
correction exposing off, high, and max, so Tau never sends its unsupported
medium value. A bundled snapshot keeps startup offline. Opening /model
refreshes catalogs in the background and caches them in
~/.tau/models-store.json; run tau update --models to force revalidation.
New models and capability changes can therefore arrive without upgrading Tau.
Set TAU_OFFLINE=1 to use cached/bundled data without catalog network access.
For a new session without an explicit preference, Hugging Face initially routes
the model automatically. After the first successful response, Tau reads Hugging
Face’s x-inference-provider response header and keeps that backing provider as
a sticky automatic route. If that route later exhausts its normal retries with a
retryable HTTP failure before producing output, Tau retries the interrupted turn
once through unsuffixed automatic routing and makes the successful replacement
the new sticky route. To choose a fixed provider instead, add a per-model
inference_providers preference to ~/.tau/providers.json:
{
"schema_version": 2,
"default_provider": "huggingface",
"provider_preferences": {
"huggingface": {
"default_model": "zai-org/GLM-5.2",
"inference_providers": { "zai-org/GLM-5.2": "deepinfra" }
}
},
"scoped_models": []
}
Use the exact provider suffix advertised for that model by Hugging Face. Tau
sends zai-org/GLM-5.2:deepinfra on the wire and continues to display and store
the logical zai-org/GLM-5.2 model. The pin survives resume; changing the
preference does not rewrite existing sessions. /session shows the active pin.
Route selection is available through the external
alejandro-ao/tau-huggingface
extension rather than a built-in command. It requires Tau 0.3.10 or newer. Clone
and load it explicitly:
git clone https://github.com/alejandro-ao/tau-huggingface.git
tau -e ./tau-huggingface
Then use /hf route to pick from the model’s currently live routes,
/hf route <provider> to select a fixed route, or /hf route automatic to
return to recoverable automatic routing. Switching models uses that model’s
configured fixed route or starts automatic resolution again. /session reports
automatic (currently <provider>) for a sticky automatic route and
<provider> (fixed) for an explicit route.
Transient failures first retry on the same wire model, and stream failures are
not retried or rerouted after model output has started. After those retries are
exhausted, only sticky routes selected in automatic mode fail over; routes chosen
through /hf route <provider> or the inference_providers preference remain
fixed so Tau never overrides explicit user intent. Automatic failover emits
visible retry progress and durable provider diagnostics. Pinning can reduce cold
prefix-cache misses caused by cross-provider routing, but cannot prevent eviction,
TTL expiry, or load balancing among workers within the chosen provider. A reroute
may require a cold prefix prefill, and account-wide rate limits may still fail on
the automatic retry. See
Configuration.
Moonshot AI API vs. Kimi Code
Both Kimi providers authenticate requests with Bearer API keys; neither uses OAuth. They are separate because the keys come from different consoles, use different endpoints, and charge against different billing plans:
| Tau provider | Access and billing | Model | Endpoint | Environment variable |
|---|---|---|---|---|
moonshotai | Pay-as-you-go key from the Kimi Open Platform | kimi-k2.7-code | https://api.moonshot.ai/v1 | MOONSHOT_API_KEY |
kimi-code | Subscription key from the Kimi Code console | k3 or rolling kimi-for-coding alias | https://api.kimi.com/coding/v1 | KIMI_CODE_API_KEY |
Kimi K3 uses the k3 model ID, accepts text and image input, and supports up to
a 1,048,576-token context window on eligible plans. It supports three
reasoning-effort levels via the reasoning_effort field: low, high, and
max (default). Tau exposes these as the distinct low, high, and max
thinking levels, and starts new K3 sessions at max unless a remembered
per-model choice exists. Start a new session when switching to K3 so the
previous model’s context cache is not re-prefilled. See
Kimi’s model documentation
for current plan availability and context limits.
A key for one service should not be treated as interchangeable with a key for
the other. Tau stores them independently under the moonshotai and kimi-code
credential names, so /login moonshotai and /login kimi-code can configure
both at once. The distinct environment variable names provide the same
separation when credentials are supplied through the shell.
Credentials saved through /login live in ~/.tau/credentials.json with
private 0600 permissions and atomic file replacement. The file is not
encrypted; protect your Tau home directory and do not share its contents. The custom-provider
flow asks for the provider name, display name, base URL, API-key environment
variable, default model, and API key; it writes the provider definition to
~/.tau/catalog.toml and runtime preferences to ~/.tau/providers.json.
Check what’s configured and how each provider will authenticate:
tau providers
Managing saved credentials
Use these slash commands inside Tau:
/login [provider] # add or refresh a saved credential
/logout [provider] # remove a saved credential
Saved credentials take precedence over environment variables. /logout only
edits saved credentials — it never touches your environment or providers.json.
OAuth troubleshooting
/login. If a Copilot model reports that it is unsupported, enable it in
Copilot Chat’s model selector or ask your organization administrator;
provider/model access varies by plan and policy.Choosing and switching models
/model— open the picker (lists models across configured providers; choosing one can switch the active provider too).tau -m <model>ortau --provider <name> -m <model>— choose at launch.- Ctrl+P / Shift+Ctrl+P — cycle your scoped (favorite) models
forward / backward without opening the picker.
Build the list with
/scoped-models, or pressSpaceon a model in the/modelpicker. In the/scoped-modelsmodal, pressTabto switch to a scoped-only view whereEnterremoves models from the list.
Tau validates the selected model against the active provider’s configured model
list before creating or refreshing a runtime provider. This prevents accidental
provider/model mismatches, such as trying to send an API-only OpenAI model to the
separate openai-codex subscription provider.
When a switch crosses provider APIs, Tau compiles existing tool history for the target provider. Provider-specific tool-call IDs are deterministically translated to a portable format, with the same translated ID used for each call and result. When compiling history for Anthropic, Tau also omits opaque reasoning signatures created by other APIs. This lets a session continue after tools have run without exposing users to provider validation errors or rewriting the saved JSONL history.
Claude Opus 5
Tau supports Anthropic’s claude-opus-5 through the direct anthropic
provider. The model has a 1M-token context window, accepts text and images,
generates up to 128k tokens, and costs $5 / $25 per million input/output tokens.
Anthropic enables adaptive thinking by default. Tau exposes models.dev’s
verified low, medium, high, xhigh, and max efforts as distinct modes.
Use /login anthropic-api or /login anthropic-subscription, then select
Claude Opus 5 in /model. See Anthropic’s
Claude Opus 5 guide
for current behavior and availability.
Dynamic extension providers
Tau’s extension API has a process-local DynamicProvider contract validated
by a permanent second fake backend and a test-only Ollama adapter. Unlike
providers created by /login custom or tau setup, these definitions are
source/generation-owned overlays and are never copied into catalog.toml,
providers.json, sessions, or generic disk storage. Source ownership is a stable
host identity derived from the canonical extension entry path—not the display
name. Tau freezes every discovered identity before importing extension code, so
symlink retargeting cannot change registration or cleanup ownership; separate
same-name extension files cannot remove one another’s providers. They support
dormant model sets, deeply immutable compatibility metadata, atomic
model-snapshot refresh, per-caller refresh deadlines, retry-safe coalescing, and
required/optional/no authentication without fake keys or exposed auth provenance.
Custom auth-resolution exceptions are reduced to a categorical host error during
runtime creation; Tau’s required-key guidance remains actionable.
Phase 1 established the contracts and registry mechanics. The /local host now
provides the generic TUI flow for registered dynamic local backends. Dynamic
providers still do not become durable catalog entries or automatic startup
fallbacks: configure a backend, then choose its provider/model explicitly. The
trusted built-in llama.cpp provider is the narrow scoped-model exception: Tau
may persist only its stable provider ID plus exact model ID. An unloaded/stale
reference remains visible as unavailable and cannot trigger load or download.
User and project dynamic providers cannot opt into durable references. See
the local backends guide and
Extensions.
Adding a custom / local provider
Any OpenAI-compatible endpoint works — including local servers like llama.cpp or Ollama. The easiest interactive path is:
/login custom
Tau prompts for the provider details, saves the API key, writes the provider
metadata to ~/.tau/catalog.toml, and makes the provider available immediately.
Built-in llama.cpp backend
Tau’s first-class llama.cpp integration is configured through the provider-neutral
/local command. For download/load/unload management, start llama.cpp
independently in router mode without a model argument:
llama-server --models-max 1 --parallel 1 --flash-attn auto
See the official router guide and Tau’s complete guide below before adding hardware- or model-specific flags.
Then open Tau and run:
/local
Choose and confirm the recommended llama.cpp backend. Tau automatically
checks the saved, environment, or default endpoint; use Configure for a
server elsewhere and an optional API key. Tau discovers exact loaded model IDs
through OpenAI-compatible or compatible router discovery; it does not use a
fake key or fake model. Compatible b9688–b10595 routers show arrow-key
navigable model states plus explicitly confirmed load, unload, Hugging Face GGUF
search, and server-side download actions. Load/download confirmations preselect
Cancel as an unlabelled safety default. Closing /local leaves an active
server-side download running; reopen it to cancel explicitly. No key means no
Authorization header. A saved key takes precedence over LLAMA_API_KEY.
Use the discovered ID explicitly from the TUI or print mode:
tau --provider llama.cpp --model <model-id>
tau --provider llama.cpp --model <model-id> --print "summarize this project"
The configured endpoint and safe model snapshot are stored under
~/.tau/state/extensions/llama.cpp.json; secrets stay in the credential store.
A cached snapshot allows explicit startup during temporary server downtime.
/local never scans ports or processes, stops the external server, or deletes
model files. See the complete llama.cpp guide
for endpoint precedence, Doctor, reset, and troubleshooting.
An older manually configured provider named llama-cpp remains separate and is
not migrated. Configure the built-in llama.cpp layer through /local, and use
the custom-provider flow below for Ollama and other OpenAI-compatible servers.
Tau does not ship an Ollama backend.
For scripted or one-off setup with another OpenAI-compatible server, use the
same tau setup flow. For example, Ollama’s OpenAI-compatible endpoint usually
runs at http://localhost:11434/v1:
tau --provider local \
--base-url http://localhost:11434/v1 \
--api-key-env LOCAL_API_KEY \
--model qwen \
setup
This writes the provider definition to ~/.tau/catalog.toml, writes runtime
preferences to ~/.tau/providers.json, and (by default) makes it the default
provider.
For reusable provider definitions, add a user-level catalog overlay at
~/.tau/catalog.toml:
schema_version = 1
[[providers]]
name = "local-gateway"
display_name = "Local Gateway"
kind = "openai-compatible"
base_url = "http://localhost:11434/v1"
api_key_env = "LOCAL_GATEWAY_API_KEY"
credential_name = "local-gateway"
models = ["qwen-coder"]
default_model = "qwen-coder"
docs_url = "https://example.test/local-gateway"
[providers.context_windows]
qwen-coder = 64000
Tau loads its bundled src/tau_coding/data/catalog.toml first, then overlays
~/.tau/catalog.toml. A user entry with the same name can extend or override a
built-in provider: scalar fields replace built-in values, models are merged
with your models first, and context_windows are merged.
There is intentionally no project-level .tau/catalog.toml. Only the
user-level ~/.tau/catalog.toml is loaded, so cloning a repository cannot
silently redirect a provider’s base_url or credentials to an unexpected
service.
Run the custom provider with:
tau --provider local-gateway
tau --provider local-gateway "summarize this project" # TUI with an initial prompt
tau --provider local-gateway -p "summarize this project" # one-shot print mode
Catalog TOML is for provider and model metadata. It does not accept runtime
request options such as custom HTTP headers, timeouts, or retry settings. Put
those in ~/.tau/providers.json instead. Saved providers.json entries support
headers, timeout_seconds, max_retries, and max_retry_delay_seconds. For
the full JSON shape, the catalog TOML shape, and thinking_levels for custom
models, see Configuration.
Hugging Face org billing
To send a Hugging Face billing header, keep the provider definition in the
catalog, then add the header to the matching provider preference in
~/.tau/providers.json:
{
"default_provider": "huggingface",
"provider_preferences": {
"huggingface": {
"default_model": "openai/gpt-oss-120b",
"headers": { "X-HF-Bill-To": "my-org" },
"thinking_defaults": { "openai/gpt-oss-120b": "low" },
"timeout_seconds": 60,
"max_retries": 2,
"max_retry_delay_seconds": 1
}
},
"scoped_models": []
}
How credentials are resolved
For a given provider, Tau uses, in order: a stored credential in
~/.tau/credentials.json, then the environment variable named by the provider’s
api_key_env. OAuth credentials are refreshed immediately before a request and
the replacement is saved atomically. Use /login for built-in providers or
/login custom for OpenAI-compatible custom providers.