Images and multimodal content
Images reach the agent through the same blob mechanism as any
other attachment - there's no separate image store, no image_id
reference system, and no resolution-downgrading over turns. See
Attach files to chat for
the shared upload/inline pipeline; this page covers what's
image-specific.
Two sources
- User -> Agent: the user attaches an image (paperclip, drag, or paste). It becomes a blob on the message, same as a document.
- Tool -> Agent: a tool (a screenshot action, a chart generator, ...) returns an image in its output. The daemon stores it as a blob the same way, so it reaches the agent through the exact same path a user upload would.
Either way, the image lands in the conversation as a native multimodal content part - no tool call needed to "see" it, and no special handling for whether it came from the user or a tool.
What actually gets sent to the provider
Every image on the latest message is converted straight into a provider-native vision content part and sent through the LLM gateway, which normalizes it to whatever shape the target provider expects. The daemon doesn't hand-roll per-provider image formats itself.
Three brain fields gate this, and none of them did anything
before they were wired up - an untouched app keeps sending every
image, which is why vision defaults to on:
| Field | Default | Effect |
|---|---|---|
vision | true | false strips every image out of the message before it reaches the LLM - useful for a model you know can't see. |
max_images_per_turn | unlimited (0) | Caps how many images are kept on one outgoing request; images past the cap are dropped. |
image_detail | provider default | Forwarded per-image as the OpenAI-style low / high / auto fidelity hint - providers that don't have this concept ignore it. |
agents:
- id: main
brain:
provider: anthropic
model: claude-sonnet-4-5
vision: true
max_images_per_turn: 4
image_detail: high
There's no "resize to 512px after N turns" or "downgrade to a text description" behavior - an image is either sent (subject to the cap above) or, once it falls off the front of the conversation the same way any older content does under normal context management, no longer in context at all.
Limits
The composer enforces 10 MB per file, 25 MB cumulative per message, and 10 files per message - see Attach files to chat for the full breakdown. There's no separate, image-specific cap beyond those.
Cross-references
- Attach files to chat - the shared attachment pipeline
- Agents -> Brain - the
vision/max_images_per_turn/image_detailfields