Middleware Pipeline
Middleware sits between the agent loop and the data flowing through it - intercepting, transforming, masking, or even short-circuiting LLM calls and tool calls. Two pipelines exist:
| Pipeline | Wraps |
|---|---|
| App-level | Every LLM call (before / after) |
| Module-level | Every tool call to a specific module |
App-level middleware
Declared under runtime.middleware, runs around every LLM
call in the agent loop. Order matters: middleware runs
top-to-bottom in before, bottom-to-top in after (standard
wrapping pattern).
runtime:
middleware:
- mask_secrets:
patterns: [api_key, password]
replacement: "[MASKED]"
- prompt_inject:
position: prepend
system: "Today is {{sys.date}}. Be concise."
- content_filter:
block_patterns: ["delete from .*"]
rejection_message: "That request was blocked."
- response_filter:
max_length: 2000
The before / after shape
Every app middleware implements two steps:
| Step | Receives | Returns |
|---|---|---|
before | The turn context (see below) | Empty to proceed, or a response string to short-circuit the LLM call (that string becomes the agent's reply, no LLM is invoked). |
after | Same context + the LLM's response | The (possibly modified) response string. |
The turn context carries: agent id, session id, user id, turn
number, system prompt, the message list, and a metadata map.
Middleware can mutate the system prompt or messages in place during
before.
Built-in app middleware
mask_secrets
Mask sensitive patterns in user messages before sending to the
LLM, and in the response after. Default regex catches
password=, api_key=, Bearer X, sk-... (OpenAI), ghp_...
(GitHub), glpat-... (GitLab).
- mask_secrets:
patterns: [internal_token, my_secret] # extra keywords beyond defaults
replacement: "[MASKED]" # default
prompt_inject
Inject extra text into the system prompt at every turn (useful for runtime context that should refresh every call - date, user identity, deployment info).
- prompt_inject:
position: append # default; or "prepend"
system: |
Current time: {{sys.timestamp}}
content_filter
Block user messages that match a regex, short-circuiting the LLM call with a fixed rejection message.
- content_filter:
block_patterns:
- "(?i)(drop|truncate)\\s+table"
- "(?i)delete\\s+from"
rejection_message: "I can't help with destructive SQL."
There is no "warn and continue" mode - a match always short-circuits.
response_filter
Truncate the LLM's response past a length, and optionally reuse
mask_secrets' pattern set on the way out.
- response_filter:
max_length: 4000 # 0 = no truncation (default)
mask_secrets: true # default false
The truncation suffix is fixed ([Response truncated]), not
configurable per entry.
rag_inject
Retrieve chunks for the latest user message and inject them into the system prompt before the LLM call, when the app has a RAG retriever wired (see RAG) - a no-op otherwise.
- rag_inject:
max_chunks: 5 # default
max_chars: 2000 # default
position: append # default; or "prepend"
Module-level middleware
Declared under tools.modules.<module_id>.middleware, runs around
every action call for that module. Order matters the same way
as app-level middleware: the first entry wraps outermost.
tools:
modules:
database:
middleware:
- audit:
log_params: true
log_result: false
- retry:
max_attempts: 3
backoff: exponential
- timeout:
seconds: 30
How a module middleware runs
Unlike app-level middleware's separate before/after steps, each
module middleware wraps the call in a single step: it receives the
call plus a next function, and decides whether (and how) to call
next - run it once, run it several times, time-box it, or skip it
and return a result of its own. Chaining several just nests these
wrappers, outermost first.
Built-in module middleware
audit
Log every action call - module, tool, session, duration, and
success/failure. Always logs errors; log_params and log_result
add the raw request/response, unredacted.
- audit:
log_params: true # log call parameters (verbatim, no redaction)
log_result: false # log the return value (may be huge)
retry
Retry a failed call with backoff, up to max_attempts (any error
triggers a retry - there's no per-error-type filter).
- retry:
max_attempts: 3
backoff: exponential # anything else behaves as "fixed"
base_delay: 1.0 # seconds
max_delay: 30.0 # seconds
timeout
Enforce a per-action time ceiling. When it elapses, the call fails
with a timeout error - which retry, if chained next, will retry
like any other error.
- timeout:
seconds: 30
Other module middleware
Five more ship alongside the three above, undocumented here in
detail: circuit_breaker (opens after repeated failures and stays
open for a recovery window), dedup (returns the cached result for
an identical call made recently in the same session instead of
re-running it), auto_heal (suggests an alternate tool name when a
call fails to resolve), cross_context (shares recent tool outputs
across a chain), and budget (caps calls or cost per hour). A
ninth, semantic_cache, is also compile-time valid but needs an
embedding backend the daemon doesn't currently wire up, so it's a
no-op in practice.
Pipeline ordering
App middlewares declared earlier in the YAML wrap outside later ones. Same for module middlewares. Concretely, with:
runtime:
middleware:
- mask_secrets: { ... }
- content_filter: { ... }
- prompt_inject: { ... }
Execution order:
mask_secrets.before→content_filter.before→prompt_inject.before→ LLM call →prompt_inject.after→content_filter.after→mask_secrets.after
This matches every standard middleware framework. If
content_filter.before short-circuits with a string,
prompt_inject.before and the LLM call are skipped, and only the
already-fired mask_secrets.before ran (its after won't fire
because there was no LLM response - short-circuit returns the
string directly).
Custom app middleware
runtime.middleware also accepts a custom: entry that proxies
before / after out to a worker process over gRPC instead of
running a built-in:
runtime:
middleware:
- custom:
module: my-plugin
kind: my-worker-kind
timeout: 5
fail_open: true
module and kind are required - kind selects which worker
process handles the call. There's no way to point this at a plain
script or class file; it always dispatches to a worker. This
extension point doesn't exist at the module level - module-level
middleware is limited to the nine built-in names above.
Choosing app-level vs module-level
| Goal | Pipeline |
|---|---|
| Mask secrets in user messages before any LLM sees them. | App (mask_secrets). |
| Inject runtime-fresh context every turn (date, user, deployment). | App (prompt_inject). |
| Block dangerous user requests before reaching the agent. | App (content_filter). |
| Audit every database query for compliance. | Module (database.middleware: [audit]). |
| Retry HTTP calls on transient failures. | Module (http.middleware: [retry]). |
| Cap the maximum time an MCP tool can take. | Module (mcp.middleware: [timeout]). |
| Per-module audit / retry / timeout. | Module middleware list on that module. |
| Cross-cutting agent-loop behaviour. | App runtime.middleware list. |
Compile-time validation
The compiler resolves every entry under runtime.middleware and
tools.modules.*.middleware against the registered middleware set.
A typo (mask_secret instead of mask_secrets) raises a clear
error pointing at the bad entry. Unknown middleware names fail closed at compile time.
Cross-references
- App-config block reference (
runtime.middleware,tools.modules.<id>.middleware): App Configuration - Hooks vs middleware (different timing, different scope): Tool Hooks