Skip to main content

Middleware

Pluggable middleware at two levels:

LevelWrapsWhere
App-levelThe LLM call inside the agent loopruntime.middleware
Module-levelAny module's tool callstools.modules.<id>.middleware

The two levels don't share one protocol. App-level middleware implements a before / after pair: before runs top-to-bottom and can short-circuit the LLM call by returning a response string instead of empty; after runs bottom-to-top over the LLM's response. Module-level middleware is a single wrapping step - it receives the call plus a next function and decides whether (and how many times) to invoke it; chaining several nests them, first entry outermost.

App-level middleware​

yaml
runtime:
middleware:
- mask_secrets:
patterns: [password, api_key, token]
- prompt_inject:
system: "Always respond in French."
- content_filter:
block_patterns: ["DROP TABLE", "rm -rf /"]
- rag_inject:
max_chunks: 5
max_chars: 2000
- response_filter:
max_length: 5000
mask_secrets: true

Built-in app middlewares​

mask_secrets​

Masks sensitive patterns in user messages and LLM responses.

yaml
- mask_secrets:
patterns: [password, api_key, secret_key] # additional patterns
replacement: "[MASKED]" # default

Built-in patterns include: password, api_key, secret_key, token, bearer, sk-* (OpenAI-shaped), ghp_* (GitHub), glpat-* (GitLab).

prompt_inject​

Injects text into the system prompt on every turn.

yaml
- prompt_inject:
system: "Always respond in French."
position: append # append (default) | prepend

content_filter​

Short-circuits the LLM call with a fixed rejection message when a user message matches a block pattern. There's no "warn and continue" mode.

yaml
- content_filter:
block_patterns: ["DROP TABLE", "rm -rf", "DELETE FROM"]
rejection_message: "This request has been blocked for safety."

rag_inject​

Retrieves chunks for the latest user message and appends (or prepends) them to the system prompt as a "Relevant context" block, before the LLM call.

yaml
- rag_inject:
max_chunks: 5 # default 5
max_chars: 2000 # default 2000
position: append # append (default) | prepend

Requires a retriever to be wired for the app - a no-op otherwise.

response_filter​

Truncates the LLM's response past a length, and can reuse mask_secrets' pattern set on the way out.

yaml
- response_filter:
max_length: 5000 # 0 = no truncation (default)
mask_secrets: true # default false

Module-level middleware​

yaml
tools:
modules:
filesystem:
middleware:
- audit:
log_params: true
log_result: false
- retry:
max_attempts: 3
base_delay: 1.0
backoff: exponential
- timeout:
seconds: 30.0

Built-in module middlewares​

Nine ship in total. Three cover the everyday cases:

audit​

Logs every call - module, tool, session, duration, success/failure. Errors are always logged; log_params and log_result add the raw request/response, unredacted.

yaml
- audit:
log_params: true # log input parameters (verbatim)
log_result: false # log the return value (may be huge)

retry​

Retries a failed call on any error (no per-error-type filter), up to max_attempts.

yaml
- retry:
max_attempts: 3 # default 3
base_delay: 1.0 # seconds
backoff: exponential # default; only the literal "fixed" turns off doubling

Exponential backoff doubles each attempt (1 s, 2 s, 4 s, ...), capped at max_delay (default 30 s).

timeout​

Per-call time ceiling; the call fails with a timeout error past it.

yaml
- timeout:
seconds: 30.0 # default 30

The remaining six - circuit_breaker, dedup, semantic_cache, auto_heal, cross_context, budget - are real and compile-time-validated the same way, but not detailed here. semantic_cache needs an embedding backend the daemon doesn't currently wire up, so it's a no-op in practice.

Custom app middleware​

runtime.middleware also accepts a custom: entry that proxies before / after calls out to a worker process over gRPC, instead of running one of the built-ins:

yaml
runtime:
middleware:
- custom:
module: my-plugin
kind: my-worker-kind
timeout: 5
fail_open: true

module and kind are required - kind selects which worker process handles the call. There's no way to point this at a plain script or class file; it always dispatches to a worker. This extension point does not exist at the module level - module-level middleware is limited to the nine built-in names above.

Cross-references​