Agent Build Tutor
Map / Full lesson

RAG and knowledge graph

Find the right passage, prove it, and add graph edges where vectors miss.

What it is and why it exists

What

Retrieval-augmented generation puts retrieved text in front of the model at answer time. This module builds hybrid retrieval (vector plus keyword) over Postgres, an authority-aware ranking step, two kinds of graph expansion, and a cheap claims store for definitions.

Why

The knowledge in a RAG system is in the store, and the model weights stay unchanged. An answer is grounded only if a tool call returned the passage in that conversation. Retrieval quality therefore sets the ceiling on answer quality, and it has to be measured separately from the chat model.

How it works

Where it sits in the build order

Needs first

  • Inference engineeringSearch needs an embedding model at query time and at index time. The vector column width is fixed by that model, so the model choice comes first.Build out of order Stub it with: A function that hashes text into a fixed-length vector. Plumbing works, results are meaningless.
  • Knowledge managementRetrieval can only return what was ingested, and it ranks by signals created at ingestion: chunk boundaries, source paths, authority tiers. Curation mistakes become retrieval mistakes that no ranking fix removes.Build out of order Stub it with: Three or four hand-written text files in one folder.
  • State and storageChunks, vectors, the file ledger, and graph edges need a store that survives concurrent readers and a nightly writer. The reference build lost weeks to single-writer stores before moving to Postgres.Build out of order Stub it with: An in-memory list with brute-force cosine similarity. Fine up to a few thousand chunks.

Unlocks

  • EvaluationThe retrieval eval calls the search tool and scores what it returns. Gold questions are generated from indexed chunks, so the index has to exist.
  • Self-evolving agentsThe learner reads indexed chunks and writes claims that point back to passages.

In the reference build

PathRole
apps/agent-server/pgschema.sqlSchemas: learned knowledge, the file ledger, and one chunk index per collection.
apps/agent-server/kb_ingest.pyExtract, chunk, embed, and record files by content hash.
apps/agent-server/tools/knowledge_base.pyThe search tool: hybrid SQL, fusion, authority weight, dedup, graph expansion, trace log.
apps/agent-server/tools/recall_claims.pyFull-text recall over verified claims, returned with the source passage.
apps/agent-server/nightly/defines.pyBuilds definition edges from set definitional forms.
apps/agent-server/nightly/citations.pyBuilds citation edges between passages.
knowledge-bank/source_authority.jsonPath rules that assign each source an authority tier.

Process map

One search, from question to returned passages

  1. 1
    QuestionPlus the collection to search. Collections are separate indexes on purpose. (input)
  2. 2
    Embed the querySame model that embedded the chunks. Different models are not comparable. (model call)
  3. 3
    Dense top NNearest vectors through the HNSW index. Finds paraphrases. (storage)
  4. 4
    Lexical top NFull-text match ranked by cover density. Finds exact terms, codes, and section numbers. (storage)
  5. 5
    Fuse with RRFScore is the sum of 1/(60 + rank) across the two lists. (decision)
  6. 6
    Authority and dedupSmall tier bonus. One copy of a paragraph that exists in several files. (decision)
  7. 7
    Graph expansionAdd cited passages and defining passages when the question type calls for it. (tool)
  8. 8
    Return with sourcesText, source file, chunk position, and a diagnostic when nothing matched. (output)

Build steps

Each step states why it sits at this point. Open any step on its own. The first is open.

01Validate documents before indexing anything
Why this step is hereIndexing is expensive and garbage is silent. A bad file produces confident chunks that rank well for the wrong reasons.

Check file type by content, not by extension or size. On the reference build, 15 files with a .pdf extension were HTML error pages that passed a size check.

Fix the folder taxonomy and file names first. Source path is a ranking signal later.

python
def is_real_pdf(path: str) -> bool:
    with open(path, "rb") as f:
        return f.read(5) == b"%PDF-"
You are done whenEvery file in the bank passes a magic-byte check for its type.
02Create the store with one index per collection
Why this step is hereThe schema fixes the vector width and the isolation boundaries. Changing either after loading millions of rows means a rebuild.

Use Postgres with the pgvector extension. It gives crash-safe tables, concurrent readers and writers, and full-text search in the same database as the vectors.

Give each collection its own chunk table and its own indexes. An ingest for one collection then cannot damage another.

Keep a file ledger keyed by content hash so re-ingestion skips unchanged files.

pgschema.sql (reduced)
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE kb.files (
  collection text NOT NULL,
  source     text NOT NULL,
  sha256     text NOT NULL,
  status     text NOT NULL DEFAULT 'indexed',
  PRIMARY KEY (collection, source)
);

CREATE TABLE kb_finance.chunks (
  chunk_id    text PRIMARY KEY,          -- "<source>::<n>"
  source      text NOT NULL,
  chunk_index int  NOT NULL,
  text        text NOT NULL,
  context     text,                      -- short lead-in describing the document
  embedding   halfvec(768),              -- width set by the embedding model
  retired     boolean NOT NULL DEFAULT false
);

CREATE INDEX chunks_fts_idx ON kb_finance.chunks
  USING gin (to_tsvector('english', coalesce(context,'') || ' ' || text));
You are done when`\d kb_finance.chunks` shows the vector column and the GIN index.
03Chunk with a context line, then embed
Why this step is hereChunking decides what a hit can contain. It comes before embedding because the embedding is of the chunk text, and you cannot re-chunk without re-embedding.

The reference build uses 800-character chunks with 100 characters of overlap. Start there and change it only with a retrieval eval in hand.

Store a short context string per chunk naming the document and section. It is included in the full-text index so a chunk can match on its document's title.

Embed in batches and record each file's hash in the ledger when its chunks are committed.

python
def chunk(text: str, size: int = 800, overlap: int = 100):
    step = size - overlap
    for i in range(0, max(len(text) - overlap, 1), step):
        yield text[i:i + size]

async def embed(text: str) -> list[float]:
    r = await http.post(f"{OLLAMA_URL}/api/embeddings",
                        json={"model": "nomic-embed-text", "prompt": text})
    return r.json()["embedding"]
You are done whenRow count in the chunk table matches your expectation for the corpus, and no embedding is null.
04Bulk load first, build the vector index after
Why this step is hereBuilding an HNSW graph row by row during a large load is hours slower than building it once at the end.

Load all rows, then create the index. Raise the search breadth parameter at query time if recall matters more than a few milliseconds.

sql
CREATE INDEX chunks_hnsw_idx ON kb_finance.chunks
  USING hnsw (embedding halfvec_cosine_ops);

SET hnsw.ef_search = 1000;   -- per session, recall over speed
You are done whenEXPLAIN on a nearest-neighbor query shows an index scan on the HNSW index.
05Write the hybrid query with rank fusion
Why this step is hereDense and lexical search fail on different questions. You need both lists before you can fuse them, and fusion by rank avoids calibrating two unrelated score scales.

Take the top N from each side. Give every chunk a score of 1/(k + rank) per list and sum. The constant k is 60 by convention and results are not sensitive to it.

Pass the query vector as a parameter in the ORDER BY. On the reference build the vector first went through a CTE, which hid it from the planner. Every search scanned 2.4 million rows and took 5.4 seconds. As a direct parameter it uses the index and runs in under 2 seconds.

knowledge_base.py (hybrid SQL, reduced)
WITH q AS (SELECT websearch_to_tsquery('english', %(text)s) AS t),
dense AS (
  SELECT chunk_id, row_number() OVER (ORDER BY d) AS r
  FROM (SELECT chunk_id, embedding <=> %(vec)s::halfvec(768) AS d
        FROM kb_finance.chunks
        WHERE NOT retired AND embedding IS NOT NULL
        ORDER BY embedding <=> %(vec)s::halfvec(768)
        LIMIT %(n)s) c
),
lexical AS (
  SELECT chunk_id, row_number() OVER (ORDER BY rank DESC) AS r
  FROM (SELECT chunk_id,
               ts_rank_cd(to_tsvector('english', coalesce(context,'') || ' ' || text), q.t) AS rank
        FROM kb_finance.chunks, q
        WHERE NOT retired
          AND to_tsvector('english', coalesce(context,'') || ' ' || text) @@ q.t
        ORDER BY rank DESC LIMIT %(n)s) l
),
fused AS (
  SELECT chunk_id, sum(1.0 / (60 + r)) AS rrf
  FROM (SELECT chunk_id, r FROM dense UNION ALL SELECT chunk_id, r FROM lexical) u
  GROUP BY chunk_id
)
SELECT c.chunk_id, f.rrf, c.source, c.chunk_index, c.text
FROM fused f JOIN kb_finance.chunks c USING (chunk_id)
ORDER BY f.rrf DESC
LIMIT %(n)s;
You are done whenA question that uses an exact section number and a paraphrased question both return the right passage in the top five.
06Add authority weighting and passage dedup
Why this step is hereThese operate on a fused candidate list, so they follow fusion. They fix two failures you only see once search works: a summary outranking the rule it summarizes, and ten results that are one paragraph copied across files.

Assign each source a tier from path rules. Add a small bonus per tier. On the reference build the weight is 0.08, enough to break near-ties and too small to override relevance.

Collapse identical passages to one result. Track how many distinct passages are in the top ten as an eval metric.

A cross-encoder reranker over the top 30 is optional. Add it behind a flag and keep it only if the eval moves.

You are done whenFor a known rule, the primary source ranks above documents that quote it.
07Add graph edges for the question types vectors miss
Why this step is hereGraph expansion needs a working baseline to expand from, and it needs evidence of which questions fail. Building a graph first is how projects spend a month on infrastructure for a problem they have not measured.

Citation edges: when a passage says "see section X", store an edge to that passage. At query time, follow edges from top hits and add a bounded number of cited passages.

Definition edges: detect set definitional forms ("X means", "the term X refers to") at index time and store an edge from the term to the defining passage. For definitional questions, add those passages.

Both are plain tables with a from and a to column. The reference build deferred a full entity-graph system until simpler tiers were shown to fail.

sql
CREATE TABLE brain.atom_links (
  from_id  text NOT NULL,
  to_id    text NOT NULL,
  relation text NOT NULL,          -- 'cites' | 'defines'
  PRIMARY KEY (from_id, to_id, relation)
);
You are done whenA definitional gold set scores higher with definition edges on than off. If it does not, turn them off.
08Add a claims store for cheap recall
Why this step is hereIt depends on having passages to cite. It exists because loading the embedder can evict the chat model, and many questions only need a definition.

Store short statements extracted from sources, each with a status and a pointer to its source passage. Index the statement text for full-text search.

Return the source passage beside each claim and tell the model that the passage is the authority. A claim is a pointer to evidence.

You are done whenA definition question is answered through the claims tool with no embedding model load in the engine log.
09Expose search as a tool that explains empty results
Why this step is hereThe tool wrapper is last because it wraps everything above. It is the interface the harness sees.

Return passages with source and position. When nothing matches, return a diagnostic naming the filter that emptied the result.

Write one trace line per search: query, candidate counts per side, final ranks. Offline diagnosis depends on it.

python
SCHEMA = {
  "type": "function",
  "function": {
    "name": "search_knowledge_base",
    "description": "Search a document collection. Returns passages with source and position.",
    "parameters": {
      "type": "object",
      "properties": {
        "collection": {"type": "string", "enum": ["FINANCE", "K12"]},
        "query": {"type": "string"},
        "top_k": {"type": "integer", "default": 6}
      },
      "required": ["collection", "query"]
    }
  }
}
You are done whenA query with no match returns an empty list and a diagnostic string, and a trace line is written.

How it works with the other parts

What went wrong in the real build

2026-08

A corrupt vector segment stopped sync for four weeks

What happened
The single-process vector store had one unreadable segment. Search on the main collection failed and ingestion could not proceed.
Fix
Move vectors into Postgres tables with write-ahead logging and a rebuildable index.
Lesson
Treat the vector index as derived data that can be rebuilt from rows you trust.
2026-09

Every search scanned the whole table

What happened
The query vector was passed through a CTE. The planner could not use the HNSW index and each search took 5.4 seconds.
Fix
Pass the vector as a query parameter in the ORDER BY clause.
Lesson
Run EXPLAIN on the retrieval query. An index that exists is not an index that is used.
2026-07

One collection leaked into another

What happened
Finance passages appeared in education-scoped chats because the knowledge attachment was set per model.
Fix
Separate indexes per collection and a deterministic scope filter.
Lesson
Isolation belongs in the data layout. A prompt instruction is not a boundary.
2026-10

The top ten was one paragraph ten times

What happened
The same paragraph existed in several files and filled the result list.
Fix
Collapse duplicate passages and track distinct passages in the top ten as a metric.
Lesson
Measure diversity of results as well as rank of the right one.

The same idea on other platforms

PlatformHow this module maps
DatabricksVector Search with a Delta Sync index replaces the chunk table and HNSW index, and it offers hybrid search. Chunks live in a Delta table governed by Unity Catalog. Graph edges are Delta tables you join. Lakebase gives you Postgres if you want to keep this SQL as written.
IBM watsonxwatsonx.data provides the lakehouse and a vector engine (Milvus). Orchestrate agents attach knowledge bases backed by it. You keep the same pipeline stages: validate, chunk, embed, hybrid query, rerank.
CodexCodex does not provide a retrieval store. Run this stack as a service and expose search as an MCP tool. Codex then calls it like any other tool.
CursorCursor indexes your code for its own use. Domain retrieval is yours to host. Register your search service as an MCP server in the project config.
Claude Code / Agent SDKExpose the search tool through an MCP server. The reference build does this so a coding agent and the chat agent share one retrieval implementation.
Another machinePostgres with pgvector runs anywhere, including Docker. The SQL in this module is unchanged. Only the embedding endpoint URL differs.

Explain it back

Answer aloud first. Then open the answer and compare.

Why hybrid search and not vectors alone?
A strong answerVectors match meaning and miss exact tokens such as section numbers, codes, and rare terms. Keyword search does the reverse. The two fail on different questions, so fusing both ranked lists covers more than either.
Why does rank fusion use ranks and not scores?
A strong answerCosine distance and text rank are on unrelated scales. Ranks are comparable across any two lists, so fusion needs no calibration.
Why must knowledge management come before retrieval?
A strong answerRanking uses signals made at ingestion: chunk boundaries, source paths, authority tiers, and which files exist. Bad inputs produce confident wrong hits, and no query-time step can recover text that was never ingested correctly.
Why are graph edges added after hybrid search works?
A strong answerEdges expand from an initial hit list, so they need one. You also need failing questions to know which edges to build. The reference build added citation and definition edges because specific gold questions failed without them.
Explain the claims store in one sentence to someone who knows RAG.
A strong answerIt is a full-text index over short verified statements that each point at a source passage, used to answer definition questions without loading the embedding model.

Terms

RRFReciprocal rank fusion. Score is the sum over lists of 1/(k + rank).
HNSWA graph index for approximate nearest-neighbor search over vectors.
hit@kShare of questions whose gold passage is in the top k results.
Authority tierA trust level assigned to a source by path rule, used as a small ranking bonus.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module