Skip to main content

Images and multimodal content

Images reach the agent through the same blob mechanism as any other attachment - there's no separate image store, no image_id reference system, and no resolution-downgrading over turns. See Attach files to chat for the shared upload/inline pipeline; this page covers what's image-specific.

Two sources​

  • User -> Agent: the user attaches an image (paperclip, drag, or paste). It becomes a blob on the message, same as a document.
  • Tool -> Agent: a tool (a screenshot action, a chart generator, ...) returns an image in its output. The daemon stores it as a blob the same way, so it reaches the agent through the exact same path a user upload would.

Either way, the image lands in the conversation as a native multimodal content part - no tool call needed to "see" it, and no special handling for whether it came from the user or a tool.

What actually gets sent to the provider​

Every image on the latest message is converted straight into a provider-native vision content part and sent through the LLM gateway, which normalizes it to whatever shape the target provider expects. The daemon doesn't hand-roll per-provider image formats itself.

Three brain fields gate this, and none of them did anything before they were wired up - an untouched app keeps sending every image, which is why vision defaults to on:

FieldDefaultEffect
visiontruefalse strips every image out of the message before it reaches the LLM - useful for a model you know can't see.
max_images_per_turnunlimited (0)Caps how many images are kept on one outgoing request; images past the cap are dropped.
image_detailprovider defaultForwarded per-image as the OpenAI-style low / high / auto fidelity hint - providers that don't have this concept ignore it.
yaml
agents:
- id: main
brain:
provider: anthropic
model: claude-sonnet-4-5
vision: true
max_images_per_turn: 4
image_detail: high

There's no "resize to 512px after N turns" or "downgrade to a text description" behavior - an image is either sent (subject to the cap above) or, once it falls off the front of the conversation the same way any older content does under normal context management, no longer in context at all.

Limits​

The composer enforces 10 MB per file, 25 MB cumulative per message, and 10 files per message - see Attach files to chat for the full breakdown. There's no separate, image-specific cap beyond those.

Cross-references​