Skip to main content

Advanced 21 - RAG knowledge base query

The rag module ships a complete retrieval-augmented generation pipeline: pluggable vector backend (Qdrant, pgvector, Elasticsearch), embeddings worker embeddings, hybrid search (semantic + BM25), chunking, cache, and citations. This tutorial walks the agent through the bootstrap → ingest → query → answer with citations loop.

On digitorn.ai (managed)? You don't wire this by hand.

On the hosted platform, add a knowledge base from the Knowledge manager in Studio (files or a website) and attach it to your agent — it searches automatically through the built-in knowledge.search tool, with no YAML and no vector store to run. This tutorial covers the self-hosted rag module, for when you run your own daemon and backend.

The YAML​

app.yaml
app:
app_id: tuto-rag-kb
name: Tuto - RAG Knowledge Base Query
version: "1.0"

runtime:
mode: conversation
workdir_mode: none
max_turns: 10
timeout: 180
tool_injection: direct
direct_modules: [rag]

agents:
- id: main
role: assistant
brain:
provider: openai
backend: openai_compat
model: gpt-5-mini
config:
api_key: placeholder
base_url: https://api.openai.com/v1
temperature: 0.1
max_tokens: 4096
system_prompt: |
You are a Digitorn documentation assistant.

First-turn bootstrap (only once per session):
1. Call rag.list_knowledge_bases. If "default" is not in the
result, call rag.ingest_directory with
knowledge_base="default", path="./rag-kb",
extensions=[".md"]. Wait for it to finish (it returns the
chunk count). Ingesting into a knowledge base that does
not exist yet creates it automatically - a separate
rag.create_knowledge_base call is not required first.

Then, on EVERY user question:
1. Call rag.query with knowledge_base="default",
query=<the user question rephrased as a search>,
top_k=3.
2. Answer based on what rag.query returned. Cite the
file path for each fact. If the result list is
empty, say so explicitly instead of guessing.

tools:
modules:
rag:
config:
embedding_model: minilm-l12
backend:
type: qdrant
# Empty path = in-memory (no disk persistence).
# Production sets a path under the workspace for
# durability.
path: ""
sources:
- type: file
path: "./rag-kb"
extensions: [".md"]
recursive: false
auto_index:
on_start: true
max_documents: 1000
capabilities:
default_policy: auto
max_risk_level: low
grant:
- module: rag
actions:
- query
- list_knowledge_bases
- knowledge_base_stats
- ingest_directory
- create_knowledge_base

Four pitfalls to know:

  • config: wrapper is mandatory. The rag module's schema is unknown keys forbidden on top-level fields; anything under rag: that is not config: ..., setup:, constraints:, or middleware: is silently dropped.
  • auto_index.on_start: true fires lazily, per app, the first time the rag engine is actually built for that app (its first rag tool call in a session) - not at daemon boot, and not for every app that merely declares sources:. If nothing has touched rag yet for an app this session, the configured sources have not been synced, so the tutorial's agent-driven ingest is still the reliable path rather than something to rely on sight-unseen.
  • Ingesting auto-creates the knowledge base. rag.ingest_directory (and rag.ingest/rag.ingest_file) create the target knowledge base if it doesn't exist yet - there's no ordering requirement against rag.create_knowledge_base.
  • Grant the right actions. Agents need ingest_directory in addition to query to run this tutorial's bootstrap. The default direct-tools build only exposes what capabilities.grant lists.

Sample knowledge base​

Three small Markdown files under ./rag-kb/:

  • hooks.md: how the Hooks V2 engine works, the list of events, condition types, action types.
  • sub-agents.md: the 7 invocation modes of the Agent tool, what role: coordinator and role: specialist do.
  • modules.md: the module system, shared vs per-app instances, direct vs discovery injection modes.

Deploy and run​

bash
digitorn install tuto-rag-kb.yaml
digitorn chat tuto-rag-kb

chat is interactive - type (or paste) this as your first message:

text
EXACT bootstrap then question. Follow this order, do not skip a step:
1. Call rag.ingest_directory(knowledge_base="default", path="./rag-kb", extensions=[".md"]).
2. Call rag.query(knowledge_base="default", query="hooks Digitorn", top_k=3).
3. Based on the rag.query results, answer: How do hooks work in Digitorn? Cite the file path for each fact.

The hyper-explicit user prompt is needed because smaller models often skip a tool when they think the data is available from a previous call. A larger model would follow the system_prompt's bootstrap order without hand-holding.

Sample flow​

Tool call 1: rag.ingest_directory

json
params:
{
"knowledge_base": "default",
"path": "./rag-kb",
"extensions": [".md"],
"recursive": true,
"max_files": 1000
}

result:
{
"knowledge_base": "default",
"files": 3,
"chunks": 10
}

3 files ingested, chunked into 10 passages, written to the in-memory Qdrant collection (created on the fly - there was no default collection before this call).

Tool call 2: rag.query

json
params:
{
"knowledge_base": "default",
"query": "hooks Digitorn",
"top_k": 3
}

result:
{
"knowledge_bases": ["default"],
"results": [
{
"text": "# Hooks in Digitorn\n\nThe Hooks V2 engine intercepts every tool call before and after execution. Each hook declares an event (tool_start, tool_end, turn_start, turn_end, session_start, session_end, pre_compact, error, approval_request, agent_spawn, agent_complete), a condition, and an action.",
"source": "hooks.md",
"chunk": 0,
"score": 0.0328,
"knowledge_base": "default"
},
...
],
"count": 3
}

results[] entries carry source (the citation label - the filename by default) and chunk (its index within that file), which is what the agent cites from.

Assistant's final reply:

text
Plan: I ingested the Markdown docs, ran a search for
"hooks Digitorn", and will answer based on the retrieved
results (citing sources).

What I did:
- Ingested ./rag-kb (Markdown).
Ingest result: 3 files, 10 chunks added to KB "default".
- Queried the KB for "hooks Digitorn" (top_k=3). The
query returned relevant passages from:
- hooks.md
- modules.md (less relevant)

Answer - How hooks work in Digitorn (facts with source):
- The Hooks V2 engine intercepts every tool call both
before and after execution.
[hooks.md]
- Each hook is defined by three parts: an event it
listens for, a condition that must hold, and an action
to run when triggered.
[hooks.md]
- The supported hook events include: tool_start,
tool_end, turn_start, turn_end, session_start,
session_end, pre_compact, error, approval_request,
agent_spawn, and agent_complete.
[hooks.md]
- Hooks fire during the agent loop.
[hooks.md]
- Hook state is isolated per session.
[hooks.md]

Every fact is grounded in the indexed document, with the file path inline.

When to reach for this​

  • Documentation Q&A: the agent answers from your project docs, code comments, or wiki, with citations the user can verify.
  • Compliance / audit: ground every answer in a primary reference so the auditor can re-read the cited passage.
  • Long-running projects: an agent that re-indexes changing files (set sources[0].watch: true) so its knowledge stays fresh without manual re-ingest.

For sessionless, one-shot semantic search (no agent loop), call rag.query via the HTTP / module execute API when your deploy exposes it. There is no rag.sql_query tool in this build; structure SQL access through database.query or RAG sources configured for DB ingestion.

Production deployment​

  • Use a persistent backend in production: backend: {type: qdrant, path: "<workspace>/qdrant_data"} for local, or {type: qdrant, url: "https://...", api_key: "..."} for Qdrant Cloud.
  • Pre-build the KB offline with the tool so the agent does not pay the ingest cost on first user turn.
  • Set cache.enabled: true (opt-in, off unless declared) to dedupe repeated near-identical queries via the semantic cache.
  • For multilingual content beyond English/French, swap the embedding model to bge-m3 (100+ languages).