Skip to main content

Middleware Pipeline

Middleware sits between the agent loop and the data flowing through it - intercepting, transforming, masking, or even short-circuiting LLM calls and tool calls. Two pipelines exist:

PipelineWraps
App-levelEvery LLM call (before / after)
Module-levelEvery tool call to a specific module

App-level middleware​

Declared under runtime.middleware, runs around every LLM call in the agent loop. Order matters: middleware runs top-to-bottom in before, bottom-to-top in after (standard wrapping pattern).

yaml
runtime:
middleware:
- mask_secrets:
patterns: [api_key, password]
replacement: "[MASKED]"
- prompt_inject:
position: prepend
system: "Today is {{sys.date}}. Be concise."
- content_filter:
block_patterns: ["delete from .*"]
rejection_message: "That request was blocked."
- response_filter:
max_length: 2000

The before / after shape​

Every app middleware implements two steps:

StepReceivesReturns
beforeThe turn context (see below)Empty to proceed, or a response string to short-circuit the LLM call (that string becomes the agent's reply, no LLM is invoked).
afterSame context + the LLM's responseThe (possibly modified) response string.

The turn context carries: agent id, session id, user id, turn number, system prompt, the message list, and a metadata map. Middleware can mutate the system prompt or messages in place during before.

Built-in app middleware​

mask_secrets​

Mask sensitive patterns in user messages before sending to the LLM, and in the response after. Default regex catches password=, api_key=, Bearer X, sk-... (OpenAI), ghp_... (GitHub), glpat-... (GitLab).

yaml
- mask_secrets:
patterns: [internal_token, my_secret] # extra keywords beyond defaults
replacement: "[MASKED]" # default

prompt_inject​

Inject extra text into the system prompt at every turn (useful for runtime context that should refresh every call - date, user identity, deployment info).

yaml
- prompt_inject:
position: append # default; or "prepend"
system: |
Current time: {{sys.timestamp}}

content_filter​

Block user messages that match a regex, short-circuiting the LLM call with a fixed rejection message.

yaml
- content_filter:
block_patterns:
- "(?i)(drop|truncate)\\s+table"
- "(?i)delete\\s+from"
rejection_message: "I can't help with destructive SQL."

There is no "warn and continue" mode - a match always short-circuits.

response_filter​

Truncate the LLM's response past a length, and optionally reuse mask_secrets' pattern set on the way out.

yaml
- response_filter:
max_length: 4000 # 0 = no truncation (default)
mask_secrets: true # default false

The truncation suffix is fixed ([Response truncated]), not configurable per entry.

rag_inject​

Retrieve chunks for the latest user message and inject them into the system prompt before the LLM call, when the app has a RAG retriever wired (see RAG) - a no-op otherwise.

yaml
- rag_inject:
max_chunks: 5 # default
max_chars: 2000 # default
position: append # default; or "prepend"

Module-level middleware​

Declared under tools.modules.<module_id>.middleware, runs around every action call for that module. Order matters the same way as app-level middleware: the first entry wraps outermost.

yaml
tools:
modules:
database:
middleware:
- audit:
log_params: true
log_result: false
- retry:
max_attempts: 3
backoff: exponential
- timeout:
seconds: 30

How a module middleware runs​

Unlike app-level middleware's separate before/after steps, each module middleware wraps the call in a single step: it receives the call plus a next function, and decides whether (and how) to call next - run it once, run it several times, time-box it, or skip it and return a result of its own. Chaining several just nests these wrappers, outermost first.

Built-in module middleware​

audit​

Log every action call - module, tool, session, duration, and success/failure. Always logs errors; log_params and log_result add the raw request/response, unredacted.

yaml
- audit:
log_params: true # log call parameters (verbatim, no redaction)
log_result: false # log the return value (may be huge)

retry​

Retry a failed call with backoff, up to max_attempts (any error triggers a retry - there's no per-error-type filter).

yaml
- retry:
max_attempts: 3
backoff: exponential # anything else behaves as "fixed"
base_delay: 1.0 # seconds
max_delay: 30.0 # seconds

timeout​

Enforce a per-action time ceiling. When it elapses, the call fails with a timeout error - which retry, if chained next, will retry like any other error.

yaml
- timeout:
seconds: 30

Other module middleware​

Five more ship alongside the three above, undocumented here in detail: circuit_breaker (opens after repeated failures and stays open for a recovery window), dedup (returns the cached result for an identical call made recently in the same session instead of re-running it), auto_heal (suggests an alternate tool name when a call fails to resolve), cross_context (shares recent tool outputs across a chain), and budget (caps calls or cost per hour). A ninth, semantic_cache, is also compile-time valid but needs an embedding backend the daemon doesn't currently wire up, so it's a no-op in practice.

Pipeline ordering​

App middlewares declared earlier in the YAML wrap outside later ones. Same for module middlewares. Concretely, with:

yaml
runtime:
middleware:
- mask_secrets: { ... }
- content_filter: { ... }
- prompt_inject: { ... }

Execution order:

  • mask_secrets.before → content_filter.before → prompt_inject.before → LLM call → prompt_inject.after → content_filter.after → mask_secrets.after

This matches every standard middleware framework. If content_filter.before short-circuits with a string, prompt_inject.before and the LLM call are skipped, and only the already-fired mask_secrets.before ran (its after won't fire because there was no LLM response - short-circuit returns the string directly).

Custom app middleware​

runtime.middleware also accepts a custom: entry that proxies before / after out to a worker process over gRPC instead of running a built-in:

yaml
runtime:
middleware:
- custom:
module: my-plugin
kind: my-worker-kind
timeout: 5
fail_open: true

module and kind are required - kind selects which worker process handles the call. There's no way to point this at a plain script or class file; it always dispatches to a worker. This extension point doesn't exist at the module level - module-level middleware is limited to the nine built-in names above.

Choosing app-level vs module-level​

GoalPipeline
Mask secrets in user messages before any LLM sees them.App (mask_secrets).
Inject runtime-fresh context every turn (date, user, deployment).App (prompt_inject).
Block dangerous user requests before reaching the agent.App (content_filter).
Audit every database query for compliance.Module (database.middleware: [audit]).
Retry HTTP calls on transient failures.Module (http.middleware: [retry]).
Cap the maximum time an MCP tool can take.Module (mcp.middleware: [timeout]).
Per-module audit / retry / timeout.Module middleware list on that module.
Cross-cutting agent-loop behaviour.App runtime.middleware list.

Compile-time validation​

The compiler resolves every entry under runtime.middleware and tools.modules.*.middleware against the registered middleware set. A typo (mask_secret instead of mask_secrets) raises a clear error pointing at the bad entry. Unknown middleware names fail closed at compile time.

Cross-references​

  • App-config block reference (runtime.middleware, tools.modules.<id>.middleware): App Configuration
  • Hooks vs middleware (different timing, different scope): Tool Hooks