Skip to main content

Security 2 - The gate chain

Every tool call an LLM makes passes through an in-order chain of security gates before it runs. Each gate answers one precise question; the first gate that says "no" stops the call - denied outright, or paused for human approval at gate 4 - and the rest of the chain never runs.

This page walks through the real gate sequence and demonstrates the most commonly used one - gate 2, the risk-level cap.

The sequence​

Higher-numbered gates only run if every lower-numbered gate passed. Gate 6 (rate limit) is checked as a separate step right after the chain, not technically part of it, but it behaves the same way from an app author's perspective: it can still deny the call.

What each gate checks​

Gate 0 - inactive. The app must be deployed and active (app.enabled: false fails this unconditionally - there's no admin bypass here). Useful for putting an app into a "soft-undeploy" state without deleting its bundle.

Gate 1a - module. The agent's profile lists which modules it can call. If the action's module isn't there, the call is blocked. This is what per-agent module restriction relies on.

Gate 1c - mcp_server. Only relevant to mcp_<server> modules: if the app declares an MCP server allowlist, a server not on it is blocked. A no-op for every app that doesn't declare one.

Gate 1d - connector association. Only relevant to the connector module: a connector's tools are callable only once the user has connected that account and associated it with the app in Settings → Connectors. A no-op for apps that use no connectors.

Gate 1b - hidden. tools.capabilities.hidden_actions lists actions the agent cannot see in its tool index. Different from deny: a hidden action can still be called from setup steps, hooks, or channel pipelines; it's only invisible to the LLM.

Gate 2 - risk level. Every tool declares a risk_level: low | medium | high. max_risk_level caps the ceiling. Actions above the ceiling are filtered out when the tool index is built, before the LLM ever sees them. Demonstrated below.

Gate 3 - permissions. A tool can declare required_permissions that the agent's granted actions must cover. In practice, today, only bash.run declares one - most tools don't use this mechanism, and an explicit tools.capabilities.grant entry for an action satisfies this gate directly regardless.

Gate 4 - policy. The big one. Resolves (module, action) against a four-step policy: explicit deny → explicit approve → explicit grant → app default_policy. First match wins. Approve pauses the call for human review (covered in Security 1).

Gate 5 - classification. Compares a static label on the tool (public / internal / confidential / restricted) against tools.capabilities.max_data_classification on the app. It's a rank comparison, not content scanning - no built-in module declares a classification today, so this gate is a no-op unless you set one yourself on a custom tool.

Gate 6 - rate limit. Configured via tools.capabilities.rate_limits: { "module.action": N } (a "*" key sets a default for anything not listed). A sliding one-minute window; bursting past N calls in that window denies the call without consuming further budget.

Live demo - gate 2 in action​

Save this as gates-bot.yaml. The interesting line is max_risk_level: medium: any action with risk_level: high is filtered out of the tool schema before the LLM ever sees it.

app.yaml
app:
app_id: gates-bot
name: Gates Bot
version: "1.0"

runtime:
mode: conversation
workdir_mode: auto
max_turns: 4
timeout: 60

agents:
- id: main
role: assistant
brain:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_main
scope: per_user
provider: deepseek
config:
api_key: "{{env.DEEPSEEK_API_KEY}}"
base_url: https://api.deepseek.com/v1
temperature: 0
max_tokens: 200
system_prompt: |
You can use Bash and filesystem tools. Be concise. If a
tool is rejected, report what was rejected and why in one
sentence.

tools:
modules:
bash: {}
filesystem: {}
capabilities:
default_policy: auto
max_risk_level: medium # bash.run declares risk=high → filtered
grant:
- module: filesystem
actions: [read, glob, grep]

bash.run declares risk_level: high; with max_risk_level: medium, gate 2 removes it from the tool index before the schema ever reaches the LLM. The model isn't refusing to call Bash - it never sees a Bash tool exists in the first place, since the filtering happens at schema-build time, not as a runtime rejection of an attempted call. That's a meaningfully stronger guarantee: a runtime rejection means the model saw an attractive tool and tried it; a schema filter never offers the choice.

Why each gate exists​

Gates 0-1 are the deployment surface: is this app running, can this agent reach this module or server, is this action hidden?

Gate 2 is the risk ceiling: a coarse cap that says "no high-risk actions in this tier", regardless of which specific actions you remembered to list. New high-risk actions added to the toolbox later are auto-blocked.

Gate 3 is a narrower permission escape hatch for the rare tool that declares one - most apps never interact with it directly since an explicit grant already satisfies it.

Gate 4 is the policy resolver: the explicit grant / approve / deny decisions, with default_policy as the catch-all. This is where the action actually clears, pauses (approve), or fails (deny).

Gate 5 is available for content-classification use cases if you declare one on a custom tool - out of the box, no built-in module uses it.

Gate 6 is flow control: stops a runaway loop from burning budget on the same action.

Composing the gates​

A typical production app uses two or three of these and ignores the rest:

  • default_policy: block plus explicit grants (gate 4)
  • max_risk_level: medium so future high-risk additions are auto-blocked (gate 2)
  • approve: on the actions where a human must always confirm (gate 4 with approval branch)

Reach for the others when:

  • You want to expose an action only to setup pipelines, not the LLM (move it to hidden_actions, gate 1b).
  • An app connects to multiple MCP servers or connectors and needs to restrict which ones an agent can reach (gates 1c/1d).
  • You're running a public-facing app and want per-action rate limits on expensive tools (gate 6).

Going further​