Middleware
Pluggable middleware at two levels:
| Level | Wraps | Where |
|---|---|---|
| App-level | The LLM call inside the agent loop | runtime.middleware |
| Module-level | Any module's tool calls | tools.modules.<id>.middleware |
The two levels don't share one protocol. App-level middleware
implements a before / after pair: before runs top-to-bottom
and can short-circuit the LLM call by returning a response string
instead of empty; after runs bottom-to-top over the LLM's
response. Module-level middleware is a single wrapping step - it
receives the call plus a next function and decides whether (and
how many times) to invoke it; chaining several nests them, first
entry outermost.
App-level middleware
runtime:
middleware:
- mask_secrets:
patterns: [password, api_key, token]
- prompt_inject:
system: "Always respond in French."
- content_filter:
block_patterns: ["DROP TABLE", "rm -rf /"]
- rag_inject:
max_chunks: 5
max_chars: 2000
- response_filter:
max_length: 5000
mask_secrets: true
Built-in app middlewares
mask_secrets
Masks sensitive patterns in user messages and LLM responses.
- mask_secrets:
patterns: [password, api_key, secret_key] # additional patterns
replacement: "[MASKED]" # default
Built-in patterns include: password, api_key, secret_key,
token, bearer, sk-* (OpenAI-shaped), ghp_* (GitHub),
glpat-* (GitLab).
prompt_inject
Injects text into the system prompt on every turn.
- prompt_inject:
system: "Always respond in French."
position: append # append (default) | prepend
content_filter
Short-circuits the LLM call with a fixed rejection message when a user message matches a block pattern. There's no "warn and continue" mode.
- content_filter:
block_patterns: ["DROP TABLE", "rm -rf", "DELETE FROM"]
rejection_message: "This request has been blocked for safety."
rag_inject
Retrieves chunks for the latest user message and appends (or prepends) them to the system prompt as a "Relevant context" block, before the LLM call.
- rag_inject:
max_chunks: 5 # default 5
max_chars: 2000 # default 2000
position: append # append (default) | prepend
Requires a retriever to be wired for the app - a no-op otherwise.
response_filter
Truncates the LLM's response past a length, and can reuse
mask_secrets' pattern set on the way out.
- response_filter:
max_length: 5000 # 0 = no truncation (default)
mask_secrets: true # default false
Module-level middleware
tools:
modules:
filesystem:
middleware:
- audit:
log_params: true
log_result: false
- retry:
max_attempts: 3
base_delay: 1.0
backoff: exponential
- timeout:
seconds: 30.0
Built-in module middlewares
Nine ship in total. Three cover the everyday cases:
audit
Logs every call - module, tool, session, duration,
success/failure. Errors are always logged; log_params and
log_result add the raw request/response, unredacted.
- audit:
log_params: true # log input parameters (verbatim)
log_result: false # log the return value (may be huge)
retry
Retries a failed call on any error (no per-error-type filter), up
to max_attempts.
- retry:
max_attempts: 3 # default 3
base_delay: 1.0 # seconds
backoff: exponential # default; only the literal "fixed" turns off doubling
Exponential backoff doubles each attempt (1 s, 2 s, 4 s, ...),
capped at max_delay (default 30 s).
timeout
Per-call time ceiling; the call fails with a timeout error past it.
- timeout:
seconds: 30.0 # default 30
The remaining six - circuit_breaker, dedup, semantic_cache,
auto_heal, cross_context, budget - are real and
compile-time-validated the same way, but not detailed here.
semantic_cache needs an embedding backend the daemon doesn't
currently wire up, so it's a no-op in practice.
Custom app middleware
runtime.middleware also accepts a custom: entry that proxies
before / after calls out to a worker process over gRPC, instead
of running one of the built-ins:
runtime:
middleware:
- custom:
module: my-plugin
kind: my-worker-kind
timeout: 5
fail_open: true
module and kind are required - kind selects which worker
process handles the call. There's no way to point this at a plain
script or class file; it always dispatches to a worker. This
extension point does not exist at the module level - module-level
middleware is limited to the nine built-in names above.
Cross-references
- App-config block reference (
runtime.middleware+tools.modules.<id>.middleware): App Configuration → runtime - Field-by-field reference and worked examples: Middleware Pipeline
- Hooks (different mechanism - fires on agent-loop events, not LLM calls): hooks.md
- Behaviour engine (per-tool runtime checks): Behavior Engine