Advanced 6 - Brain fallback for billing failover
The most common reason a production agent app stops working mid-conversation isn't a bug - it's the LLM provider returning a billing error. The user's credit ran out, the team's monthly cap was reached, the API key was suspended. The turn fails and the session dies.
Brain fallback is the answer. Each agent's brain can declare a
fallback: block - a complete second brain config - that the
runtime swaps to the moment the primary fails with a billing
error. The fallback is typically a different provider entirely (a
free local model, a backup API key, a cheaper plan), so the
session keeps running.
When the failover triggers
The runtime swaps to the fallback brain on billing-class errors
only - the primary's error message matching any of insufficient,
quota, balance, billing, payment, 402, exceeded your current quota, or budget. Everything else takes its normal path
instead of triggering a fallback:
- Rate limits (429, "too many requests") get retried against the primary with exponential backoff - a rate limit clears on its own, so there's nothing to fail over to.
- Auth errors (401, invalid API key), configuration errors (unknown model, bad request shape), and context-overflow errors fail immediately, with no retry and no fallback - these are all cases where the fallback brain would fail identically (a bad key stays bad, an unknown model stays unknown).
The intent is narrow: this failover is for soft-fail billing situations, not disaster recovery. For DR you want a different mechanism (a load balancer, a multi-region daemon, a circuit breaker). For "the account ran out of credit", brain fallback is exactly the right shape.
The YAML
agents:
- id: main
role: assistant
brain:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_main
scope: per_user
provider: deepseek
config:
api_key: "{{env.DEEPSEEK_API_KEY}}"
base_url: https://api.deepseek.com/v1
temperature: 0
max_tokens: 64
fallback: # ← the new piece
provider: anthropic
model: claude-haiku-4-5
backend: anthropic
config:
api_key: claude-code # uses the local OAuth file
temperature: 0
max_tokens: 64
fallback: takes the same shape as brain: itself - provider,
model, backend, config/credential, and its own sampling
settings. Whatever it declares is what runs once it's switched to.
The example above pairs DeepSeek (cheap, fast, sometimes runs dry) with Claude Haiku via the Claude Code OAuth alias (no extra key needed, billed against the local Claude Code subscription). A common alternative is the same provider with a different key - your team's shared backup key takes over when the per-user key empties.
What the runtime actually does
Inside the turn's retry loop, the moment a call to the primary brain returns a billing-classified error - and only then - the runtime:
- Logs the switch (
llm_billing_exhausted, naming the primary and the fallback). - Rebuilds the outgoing request around the fallback's provider, model, and backend (re-resolving its own API key for BYOK apps).
- Retries immediately - no backoff, since this isn't a transient failure to wait out.
This happens at most once per turn: if the fallback brain also fails, that failure takes its own normal path (retried if transient, surfaced as a turn error otherwise) rather than looping between brains. There's no cross-turn "stickiness" - the next turn always starts on the primary again and re-triggers the fallback fresh if the primary is still failing.
Verified
The switch itself is covered by a dedicated test at the exact
boundary where the runtime hands a request to the LLM client: a
fake client fails every call carrying the primary's model and
succeeds on the fallback's, so the test proves the request was
actually rebuilt around fallback: - not just retried against the
same, still-broken primary - and that the switch happens with zero
backoff delay. A second test confirms nothing changes for an agent
that never declared a fallback: - the same billing error surfaces
immediately, exactly as it did before this existed.
Picking the right fallback
Three patterns work in practice.
Same provider, backup key. Your team has a primary per-user key (which can drain) and a shared service-account key (which is throttled but never empty). When the user's key drains, the shared one takes over. Cheapest to set up, no provider integration needed.
fallback:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_team_backup # different credential
scope: per_app_shared
provider: deepseek
Cheaper model, same provider. The primary is a flagship model; the fallback is the cheapest model on the same provider. If billing fails (rare but possible), at least the session keeps running with reduced quality.
fallback:
provider: openai
model: gpt-4o-mini # cheaper than the primary's gpt-4o
...
Different provider entirely. Most resilient. If the primary provider has a regional outage or a billing system bug, the fallback dispatches to a completely separate vendor. Best for production-critical apps.
fallback:
provider: anthropic
model: claude-haiku-4-5
...
The fallback's quality and cost matter less than its availability. A response that's slightly worse than the primary is better than a session-killing error.
Composition with other defences
Brain fallback is one of several layers protecting the session against provider failures. The full stack:
| Layer | Catches | Lives in |
|---|---|---|
| Daemon retry loop | Rate limits and other transient failures - exponential backoff against the same brain | The turn's own retry loop |
| Brain fallback | Billing / quota / payment errors | brain.fallback in YAML |
| Capability deny | Action-level rejection (not LLM-related) | Capability gate 4 |
| Session-level abort | User-triggered or daemon-triggered cancellation | /sessions/{sid}/abort |
Each layer handles a distinct failure class. They don't substitute for each other - brain fallback won't help on a rate limit; the retry loop won't help on a drained billing account.
Going further
- The full agent brain reference, including the fallback schema: Agents.
- The
claude-codeAPI key alias (used in the fallback example to read from~/.claude/.credentials.jsonwithout declaring a credential): llm_provider module. - For multi-app call chaining as a different kind of resilience pattern (one app's output drives another): Composition.