Skip to main content

Advanced 10 - Parallel tool execution with run_parallel

Tutorial 4 showed how to spawn parallel sub-agents. That's the right pattern when the work needs distinct system prompts or specialised brains. For the simpler case - fire N independent tool calls and wait for all of them - there's a lighter primitive: run_parallel.

The agent passes a list of {tool, args} tuples; the runtime fires every call concurrently, collects the results, and returns them as a single bundle. One tool call in the agent's view, N parallel calls internally.

When to use it​

  • Fan-out reads: search 5 different sources for the same query, merge the results.
  • Independent shell commands: run git status, git diff --stat, git log -3 in parallel and present a composite summary.
  • Bulk fetch: http.request against 10 URLs at once instead of serializing them.
  • Validation across files: lint 20 files concurrently, collect errors.

The constraint: the calls must be independent. If call B needs the result of call A, you want sequential chaining (a normal agent loop) or a hook pipe (Advanced 7), not parallel fan-out.

The YAML​

Save as parallel-bot.yaml. The agent gets bash.run plus the run_parallel primitive.

app.yaml
app:
app_id: parallel-bot
name: Parallel Bot
version: "1.0"

runtime:
mode: conversation
workdir_mode: auto
max_turns: 4
timeout: 60

agents:
- id: main
role: assistant
brain:
provider: deepseek
model: deepseek-chat
backend: openai_compat
credential:
ref: deepseek_main
scope: per_user
provider: deepseek
config:
api_key: "{{env.DEEPSEEK_API_KEY}}"
base_url: https://api.deepseek.com/v1
temperature: 0
max_tokens: 400
system_prompt: |
You can execute multiple tool calls in parallel using
run_parallel. Pass `tasks` as a list of {tool, args}
dicts. The runtime fires them concurrently and returns
every result. Use it when several independent calls can
run at once. Reply with one short summary line.

tools:
modules:
bash: {}
capabilities:
default_policy: auto
max_risk_level: high
grant:
- module: bash
actions: [run]
- module: context_builder
actions: [run_parallel]

run_parallel is a meta-tool exposed by context_builder (auto-loaded) and always available - like background_run, it bypasses the gate chain entirely (meta_tool_bypass, see Security architecture), so the grant row above is harmless but not what makes it callable.

What the call looks like​

The user asks the agent to fire three Bash echoes in parallel:

text
> Use run_parallel to fire THREE Bash calls at once:
echo ONE, echo TWO, echo THREE. Then list each output.

The agent makes a single tool call - tool_calls_count: 1 from its own perspective, even though three Bash processes actually run:

json
run_parallel(tasks=[
{"tool": "bash.run", "args": {"command": "echo ONE"}},
{"tool": "bash.run", "args": {"command": "echo TWO"}},
{"tool": "bash.run", "args": {"command": "echo THREE"}}
])
json
{
"results": [
{ "name": "bash.run", "status": "completed", "content": "ONE\n" },
{ "name": "bash.run", "status": "completed", "content": "TWO\n" },
{ "name": "bash.run", "status": "completed", "content": "THREE\n" }
]
}

All three ran concurrently - the runtime doesn't wait for echo ONE to finish before starting echo TWO. For three near-instant echoes the wall-clock difference against running them one at a time is negligible; the win compounds with slower calls (a handful of HTTP requests, several file reads) where the batch takes roughly as long as the slowest single call instead of the sum of all of them.

Anatomy of the result​

The call accepts a fair amount of alias flexibility: the top-level list can be called tasks, actions, calls, tools, invocations, steps, or items; each item's tool name can be under tool, name, action, or tool_name; and its arguments under args, params, arguments, input, or parameters. The canonical form is {"tasks": [{"tool": ..., "args": {...}}]} - the rest are accepted so the model doesn't have to remember an exact shape.

The result is a single {"results": [...]} object. Each entry mirrors the input list in order:

  • name - the resolved FQN (bash.run, not the short alias used to call it)
  • status - "completed" or "errored"
  • error - present only when status is "errored"
  • content - the tool's text output, present when there is any

There's no separate aggregate summary (no total/succeeded/failed counts, no elapsed-time field) - the agent checks status on each entry itself. A cap of 256 actions per call applies, and run_parallel cannot be nested inside another run_parallel call in the same turn (fire everything in one flat list instead).

Failure mode​

run_parallel does not abort on the first failure - all N calls are awaited. A single failure shows up as status: "errored" on that entry, with the rest of the batch continuing normally. This matches the typical fan-out mental model: the failed fetches are reported alongside the successful ones, the agent decides what to do with the partial result.

To abort siblings on first failure, fall back to spawning a sub-agent (Tutorial 4) - sub-agents have isolated lifecycles and can be cancelled cleanly. run_parallel is for "best-effort batch", sub-agent spawn is for "ordered operation with cleanup semantics".

Comparison with the other parallelism primitives​

PrimitiveConcurrencyLifecycleBest for
run_parallelTrue parallelOne turn, await allBatch fan-out of N tool calls, no per-call agent context
background_runTrue parallel; agent gets task_idSurvives turns, cancellableLong-running tools the agent should poll, not block on
Agent (spawn)True parallel; full agent loopIndependent context, isolated brainSpecialists with different prompts / brains
Sequential callsSerialOne turn, await eachWhen call B depends on call A's result

The first three all run concurrently; the difference is what each call carries. run_parallel carries just a tool call. background_run carries a tool call plus a handle so it can outlive the turn. Agent carries an entire agent loop with its own brain config.

Pick the smallest primitive that does the job. run_parallel is the smallest and the cheapest.

Going further​