Skip to main content

Advanced 6 - Brain fallback for billing failover

The most common reason a production agent app stops working mid-conversation isn't a bug - it's the LLM provider returning a billing error. The user's credit ran out, the team's monthly cap was reached, the API key was suspended. The turn fails and the session dies.

Brain fallback is the answer. Each agent's brain can declare a fallback: block - a complete second brain config - that the runtime swaps to the moment the primary fails with a billing error. The fallback is typically a different provider entirely (a free local model, a backup API key, a cheaper plan), so the session keeps running.

When the failover triggers​

The runtime swaps to the fallback brain on billing-class errors only - the primary's error message matching any of insufficient, quota, balance, billing, payment, 402, exceeded your current quota, or budget. Everything else takes its normal path instead of triggering a fallback:

  • Rate limits (429, "too many requests") get retried against the primary with exponential backoff - a rate limit clears on its own, so there's nothing to fail over to.
  • Auth errors (401, invalid API key), configuration errors (unknown model, bad request shape), and context-overflow errors fail immediately, with no retry and no fallback - these are all cases where the fallback brain would fail identically (a bad key stays bad, an unknown model stays unknown).

The intent is narrow: this failover is for soft-fail billing situations, not disaster recovery. For DR you want a different mechanism (a load balancer, a multi-region daemon, a circuit breaker). For "the account ran out of credit", brain fallback is exactly the right shape.

The YAML​

yaml
agents:
- id: main
role: assistant
brain:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_main
scope: per_user
provider: deepseek
config:
api_key: "{{env.DEEPSEEK_API_KEY}}"
base_url: https://api.deepseek.com/v1
temperature: 0
max_tokens: 64

fallback: # ← the new piece
provider: anthropic
model: claude-haiku-4-5
backend: anthropic
config:
api_key: claude-code # uses the local OAuth file
temperature: 0
max_tokens: 64

fallback: takes the same shape as brain: itself - provider, model, backend, config/credential, and its own sampling settings. Whatever it declares is what runs once it's switched to.

The example above pairs DeepSeek (cheap, fast, sometimes runs dry) with Claude Haiku via the Claude Code OAuth alias (no extra key needed, billed against the local Claude Code subscription). A common alternative is the same provider with a different key - your team's shared backup key takes over when the per-user key empties.

What the runtime actually does​

Inside the turn's retry loop, the moment a call to the primary brain returns a billing-classified error - and only then - the runtime:

  1. Logs the switch (llm_billing_exhausted, naming the primary and the fallback).
  2. Rebuilds the outgoing request around the fallback's provider, model, and backend (re-resolving its own API key for BYOK apps).
  3. Retries immediately - no backoff, since this isn't a transient failure to wait out.

This happens at most once per turn: if the fallback brain also fails, that failure takes its own normal path (retried if transient, surfaced as a turn error otherwise) rather than looping between brains. There's no cross-turn "stickiness" - the next turn always starts on the primary again and re-triggers the fallback fresh if the primary is still failing.

Verified​

The switch itself is covered by a dedicated test at the exact boundary where the runtime hands a request to the LLM client: a fake client fails every call carrying the primary's model and succeeds on the fallback's, so the test proves the request was actually rebuilt around fallback: - not just retried against the same, still-broken primary - and that the switch happens with zero backoff delay. A second test confirms nothing changes for an agent that never declared a fallback: - the same billing error surfaces immediately, exactly as it did before this existed.

Picking the right fallback​

Three patterns work in practice.

Same provider, backup key. Your team has a primary per-user key (which can drain) and a shared service-account key (which is throttled but never empty). When the user's key drains, the shared one takes over. Cheapest to set up, no provider integration needed.

yaml
fallback:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_team_backup # different credential
scope: per_app_shared
provider: deepseek

Cheaper model, same provider. The primary is a flagship model; the fallback is the cheapest model on the same provider. If billing fails (rare but possible), at least the session keeps running with reduced quality.

yaml
fallback:
provider: openai
model: gpt-4o-mini # cheaper than the primary's gpt-4o
...

Different provider entirely. Most resilient. If the primary provider has a regional outage or a billing system bug, the fallback dispatches to a completely separate vendor. Best for production-critical apps.

yaml
fallback:
provider: anthropic
model: claude-haiku-4-5
...

The fallback's quality and cost matter less than its availability. A response that's slightly worse than the primary is better than a session-killing error.

Composition with other defences​

Brain fallback is one of several layers protecting the session against provider failures. The full stack:

LayerCatchesLives in
Daemon retry loopRate limits and other transient failures - exponential backoff against the same brainThe turn's own retry loop
Brain fallbackBilling / quota / payment errorsbrain.fallback in YAML
Capability denyAction-level rejection (not LLM-related)Capability gate 4
Session-level abortUser-triggered or daemon-triggered cancellation/sessions/{sid}/abort

Each layer handles a distinct failure class. They don't substitute for each other - brain fallback won't help on a rate limit; the retry loop won't help on a drained billing account.

Going further​

  • The full agent brain reference, including the fallback schema: Agents.
  • The claude-code API key alias (used in the fallback example to read from ~/.claude/.credentials.json without declaring a credential): llm_provider module.
  • For multi-app call chaining as a different kind of resilience pattern (one app's output drives another): Composition.