Skip to main content

Advanced 20 - Multimodal image input

The chat API accepts image attachments on every user message. When the brain points at a vision-capable model, Digitorn forwards the image to the LLM gateway, which converts it to the provider-native shape. The agent sees the picture alongside the text and can reason about it like any other turn - no tool call needed.

What you build​

  1. Generate a deterministic PNG client-side (here with PIL).
  2. Upload the raw bytes to POST /api/apps/{app_id}/blobs (Content-Type set to the image's mime type, body = the raw bytes) and get back {hash, mime, size}.
  3. Send the message with that blob referenced in the attachments: [...] field of /sessions (first message) or /sessions/<sid>/messages (subsequent) - {hash, mime, size, name}.
  4. The vision model reasons about the image and returns text.

The YAML​

app.yaml
app:
app_id: tuto-multimodal-image
name: Tuto - Multimodal Image Input
version: "1.0"
attachments:
- image

runtime:
mode: conversation
workdir_mode: none
max_turns: 3
timeout: 60
tool_injection: direct

agents:
- id: main
role: assistant
brain:
provider: openai
backend: openai_compat
model: gpt-5-mini
config:
api_key: placeholder
base_url: https://api.openai.com/v1
temperature: 0.1
# Reasoning models consume internal reasoning tokens
# BEFORE visible text. 4096 leaves ~3000 tokens of
# reasoning headroom plus enough for the answer.
max_tokens: 4096
vision: true
image_detail: low
max_images_per_turn: 5
system_prompt: |
You analyse images. When the user sends a message with
an attached image, describe what you see in 2-4 short
bullets. Be specific (colours, shapes, text, count). No
preamble.

tools:
modules: {}
capabilities:
default_policy: auto
max_risk_level: low

Two knobs to know:

  • app.attachments: [image] declares which attachment types this app accepts. Without it, the client composer hides the upload button.
  • brain.vision: true pins vision-on (this is also the default when the field is omitted - an app that never touched this knob keeps sending every image). Set it to false and images are stripped out entirely before the LLM call instead of being described in text - there's no automatic "convert to a text description" fallback.
  • brain.max_tokens: 4096. Reasoning models such as gpt-5-mini, o3, and o4-mini consume internal reasoning tokens before producing visible text. With a small budget the model spends the whole budget thinking and returns an empty reply. 4096 leaves room for both.

The attachments array accepts any number of blobs per message, each shaped {hash: <sha-256>, mime: <"image/png"|"image/jpeg"|...>, size: <bytes>, name: <str>} - the hash comes back from the blob upload in step 2. max_images_per_turn on the brain caps how many of the request's images actually reach the model; the rest are dropped, not queued for a later turn.

Sample flow​

Upload (POST /api/apps/tuto-multimodal-image/blobs, body = raw PNG bytes, Content-Type: image/png):

json
{ "hash": "a941e692e449...", "mime": "image/png", "size": 4108 }

Send message (POST .../sessions/<sid>/messages):

json
{
"message": "What text is in this image? What colour is the background? Reply in 2 bullets.",
"attachments": [
{ "hash": "a941e692e449...", "mime": "image/png", "size": 4108, "name": "digitorn-test.png" }
]
}

The image reaches the model the same turn it's attached - there's no separate placeholder-then-inflate step, and no metadata (width, height, turn number) beyond what's shown above is tracked.

Assistant reply (vision-capable model):

text
- Text: "DIGITORN" centered, uppercase, white sans-serif.
- Background colour: solid crimson red.

Going further​

  • Send multiple images per turn: upload each one and add every {hash, mime, size, name} to the same attachments array. All of them reach the model in one LLM call, capped at max_images_per_turn.
  • Add attachments: [document, image] to accept PDFs, Word docs, etc. Documents are extracted to text before reaching the model.
  • Give the app a real workdir (workdir_mode other than none) and grant filesystem if the agent should be able to re-read a specific attachment later by path, on top of what was inlined automatically the first turn.
  • Image generation (DALL-E, Stable Diffusion, ...): set brain.kind: image (or video) on a dedicated media-generation agent rather than a boolean on the main chat brain.