Skip to main content

RAG Module

The rag module indexes sources into knowledge bases and retrieves for agents.

On digitorn.ai (managed)? Use Studio Knowledge, not this module.

On the hosted platform, give an agent a knowledge base from the Knowledge manager in Studio (upload files or add a website, attach it to the agent) — the agent then searches it with the built-in knowledge.search tool, no YAML and no vector store to run. This page is the self-hosted rag module, for when you run your own daemon and vector backend.

Authoritative tool table: reference/modules/rag.

Agent tools (11)​

create_knowledge_base, delete_knowledge_base, list_knowledge_bases, knowledge_base_stats, index_stats, ingest, ingest_file, ingest_directory, query, reindex, migrate_embeddings.

Zero-config​

yaml
tools:
modules:
rag: {}

Configuration (high level)​

Config is bound from YAML under tools.modules.rag (and related module config). Important groups from Config in Go:

AreaNotes
embedding_modelString shortcut or {id, dimensions, pooling}
backendVector store settings (Qdrant / pgvector / Elasticsearch appear in code)
pipeline / chunking / citations / cache / aclRetrieval and safety knobs
sourcesFile, DB, web, kafka-style source entries for the indexer
auto_indexon_start, schedule (cron string for indexer triggers)
default_knowledge_baseDefault KB name
max_knowledge_bases / max_documentsCaps

Source entries can carry their own triggers (type, every, cron).

Example​

yaml
tools:
modules:
rag:
config:
embedding_model: minilm-l12
default_knowledge_base: docs
auto_index:
on_start: true
schedule: "" # optional cron for the indexer
sources:
- name: handbook
type: file
path: "{{workdir}}/docs"
extensions: [.md, .txt]
recursive: true
capabilities:
grant:
- module: rag