Advanced 20 - Multimodal image input
The chat API accepts image attachments on every user message. When the brain points at a vision-capable model, Digitorn forwards the image to the LLM gateway, which converts it to the provider-native shape. The agent sees the picture alongside the text and can reason about it like any other turn - no tool call needed.
What you build
- Generate a deterministic PNG client-side (here with PIL).
- Upload the raw bytes to
POST /api/apps/{app_id}/blobs(Content-Typeset to the image's mime type, body = the raw bytes) and get back{hash, mime, size}. - Send the message with that blob referenced in the
attachments: [...]field of/sessions(first message) or/sessions/<sid>/messages(subsequent) -{hash, mime, size, name}. - The vision model reasons about the image and returns text.
The YAML
app:
app_id: tuto-multimodal-image
name: Tuto - Multimodal Image Input
version: "1.0"
attachments:
- image
runtime:
mode: conversation
workdir_mode: none
max_turns: 3
timeout: 60
tool_injection: direct
agents:
- id: main
role: assistant
brain:
provider: openai
backend: openai_compat
model: gpt-5-mini
config:
api_key: placeholder
base_url: https://api.openai.com/v1
temperature: 0.1
# Reasoning models consume internal reasoning tokens
# BEFORE visible text. 4096 leaves ~3000 tokens of
# reasoning headroom plus enough for the answer.
max_tokens: 4096
vision: true
image_detail: low
max_images_per_turn: 5
system_prompt: |
You analyse images. When the user sends a message with
an attached image, describe what you see in 2-4 short
bullets. Be specific (colours, shapes, text, count). No
preamble.
tools:
modules: {}
capabilities:
default_policy: auto
max_risk_level: low
Two knobs to know:
app.attachments: [image]declares which attachment types this app accepts. Without it, the client composer hides the upload button.brain.vision: truepins vision-on (this is also the default when the field is omitted - an app that never touched this knob keeps sending every image). Set it tofalseand images are stripped out entirely before the LLM call instead of being described in text - there's no automatic "convert to a text description" fallback.brain.max_tokens: 4096. Reasoning models such asgpt-5-mini,o3, ando4-miniconsume internal reasoning tokens before producing visible text. With a small budget the model spends the whole budget thinking and returns an empty reply. 4096 leaves room for both.
The attachments array accepts any number of blobs per message,
each shaped {hash: <sha-256>, mime: <"image/png"|"image/jpeg"|...>, size: <bytes>, name: <str>} - the hash comes back from the blob
upload in step 2. max_images_per_turn on the brain caps how many
of the request's images actually reach the model; the rest are
dropped, not queued for a later turn.
Sample flow
Upload (POST /api/apps/tuto-multimodal-image/blobs, body =
raw PNG bytes, Content-Type: image/png):
{ "hash": "a941e692e449...", "mime": "image/png", "size": 4108 }
Send message (POST .../sessions/<sid>/messages):
{
"message": "What text is in this image? What colour is the background? Reply in 2 bullets.",
"attachments": [
{ "hash": "a941e692e449...", "mime": "image/png", "size": 4108, "name": "digitorn-test.png" }
]
}
The image reaches the model the same turn it's attached - there's no separate placeholder-then-inflate step, and no metadata (width, height, turn number) beyond what's shown above is tracked.
Assistant reply (vision-capable model):
- Text: "DIGITORN" centered, uppercase, white sans-serif.
- Background colour: solid crimson red.
Going further
- Send multiple images per turn: upload each one and add every
{hash, mime, size, name}to the sameattachmentsarray. All of them reach the model in one LLM call, capped atmax_images_per_turn. - Add
attachments: [document, image]to accept PDFs, Word docs, etc. Documents are extracted to text before reaching the model. - Give the app a real workdir (
workdir_modeother thannone) and grantfilesystemif the agent should be able to re-read a specific attachment later by path, on top of what was inlined automatically the first turn. - Image generation (DALL-E, Stable Diffusion, ...): set
brain.kind: image(orvideo) on a dedicated media-generation agent rather than a boolean on the main chat brain.