RAG and knowledge graph
Find the right passage, prove it, and add graph edges where vectors miss.
What it is and why it exists
What
Retrieval-augmented generation puts retrieved text in front of the model at answer time. This module builds hybrid retrieval (vector plus keyword) over Postgres, an authority-aware ranking step, two kinds of graph expansion, and a cheap claims store for definitions.
Why
The knowledge in a RAG system is in the store, and the model weights stay unchanged. An answer is grounded only if a tool call returned the passage in that conversation. Retrieval quality therefore sets the ceiling on answer quality, and it has to be measured separately from the chat model.
How it works
- Documents are split into chunks. Each chunk gets an embedding vector and a full-text index entry.
- A query runs twice: nearest vectors (dense) and keyword match (lexical). Reciprocal rank fusion merges the two ranked lists without needing comparable scores.
- A small authority weight nudges primary sources above summaries of them. Duplicate passages that appear in several files collapse to one.
- Graph edges add passages that similarity search cannot reach. Citation edges follow references from a hit to the passage it cites. Definition edges fetch the passage that defines a term the question uses.
- A separate claims table stores short verified statements with full-text search. It answers definition questions without loading the embedding model.
Where it sits in the build order
Needs first
- Inference engineeringSearch needs an embedding model at query time and at index time. The vector column width is fixed by that model, so the model choice comes first.Build out of order Stub it with: A function that hashes text into a fixed-length vector. Plumbing works, results are meaningless.
- Knowledge managementRetrieval can only return what was ingested, and it ranks by signals created at ingestion: chunk boundaries, source paths, authority tiers. Curation mistakes become retrieval mistakes that no ranking fix removes.Build out of order Stub it with: Three or four hand-written text files in one folder.
- State and storageChunks, vectors, the file ledger, and graph edges need a store that survives concurrent readers and a nightly writer. The reference build lost weeks to single-writer stores before moving to Postgres.Build out of order Stub it with: An in-memory list with brute-force cosine similarity. Fine up to a few thousand chunks.
Unlocks
- EvaluationThe retrieval eval calls the search tool and scores what it returns. Gold questions are generated from indexed chunks, so the index has to exist.
- Self-evolving agentsThe learner reads indexed chunks and writes claims that point back to passages.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/pgschema.sql | Schemas: learned knowledge, the file ledger, and one chunk index per collection. |
| apps/agent-server/kb_ingest.py | Extract, chunk, embed, and record files by content hash. |
| apps/agent-server/tools/knowledge_base.py | The search tool: hybrid SQL, fusion, authority weight, dedup, graph expansion, trace log. |
| apps/agent-server/tools/recall_claims.py | Full-text recall over verified claims, returned with the source passage. |
| apps/agent-server/nightly/defines.py | Builds definition edges from set definitional forms. |
| apps/agent-server/nightly/citations.py | Builds citation edges between passages. |
| knowledge-bank/source_authority.json | Path rules that assign each source an authority tier. |
Process map
One search, from question to returned passages
- 1QuestionPlus the collection to search. Collections are separate indexes on purpose. (input)
- 2Embed the querySame model that embedded the chunks. Different models are not comparable. (model call)
- 3Dense top NNearest vectors through the HNSW index. Finds paraphrases. (storage)
- 4Lexical top NFull-text match ranked by cover density. Finds exact terms, codes, and section numbers. (storage)
- 5Fuse with RRFScore is the sum of 1/(60 + rank) across the two lists. (decision)
- 6Authority and dedupSmall tier bonus. One copy of a paragraph that exists in several files. (decision)
- 7Graph expansionAdd cited passages and defining passages when the question type calls for it. (tool)
- 8Return with sourcesText, source file, chunk position, and a diagnostic when nothing matched. (output)
Build steps
Each step states why it sits at this point. Open any step on its own. The first is open.
01Validate documents before indexing anything
Check file type by content, not by extension or size. On the reference build, 15 files with a .pdf extension were HTML error pages that passed a size check.
Fix the folder taxonomy and file names first. Source path is a ranking signal later.
def is_real_pdf(path: str) -> bool:
with open(path, "rb") as f:
return f.read(5) == b"%PDF-"02Create the store with one index per collection
Use Postgres with the pgvector extension. It gives crash-safe tables, concurrent readers and writers, and full-text search in the same database as the vectors.
Give each collection its own chunk table and its own indexes. An ingest for one collection then cannot damage another.
Keep a file ledger keyed by content hash so re-ingestion skips unchanged files.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE kb.files (
collection text NOT NULL,
source text NOT NULL,
sha256 text NOT NULL,
status text NOT NULL DEFAULT 'indexed',
PRIMARY KEY (collection, source)
);
CREATE TABLE kb_finance.chunks (
chunk_id text PRIMARY KEY, -- "<source>::<n>"
source text NOT NULL,
chunk_index int NOT NULL,
text text NOT NULL,
context text, -- short lead-in describing the document
embedding halfvec(768), -- width set by the embedding model
retired boolean NOT NULL DEFAULT false
);
CREATE INDEX chunks_fts_idx ON kb_finance.chunks
USING gin (to_tsvector('english', coalesce(context,'') || ' ' || text));03Chunk with a context line, then embed
The reference build uses 800-character chunks with 100 characters of overlap. Start there and change it only with a retrieval eval in hand.
Store a short context string per chunk naming the document and section. It is included in the full-text index so a chunk can match on its document's title.
Embed in batches and record each file's hash in the ledger when its chunks are committed.
def chunk(text: str, size: int = 800, overlap: int = 100):
step = size - overlap
for i in range(0, max(len(text) - overlap, 1), step):
yield text[i:i + size]
async def embed(text: str) -> list[float]:
r = await http.post(f"{OLLAMA_URL}/api/embeddings",
json={"model": "nomic-embed-text", "prompt": text})
return r.json()["embedding"]04Bulk load first, build the vector index after
Load all rows, then create the index. Raise the search breadth parameter at query time if recall matters more than a few milliseconds.
CREATE INDEX chunks_hnsw_idx ON kb_finance.chunks
USING hnsw (embedding halfvec_cosine_ops);
SET hnsw.ef_search = 1000; -- per session, recall over speed05Write the hybrid query with rank fusion
Take the top N from each side. Give every chunk a score of 1/(k + rank) per list and sum. The constant k is 60 by convention and results are not sensitive to it.
Pass the query vector as a parameter in the ORDER BY. On the reference build the vector first went through a CTE, which hid it from the planner. Every search scanned 2.4 million rows and took 5.4 seconds. As a direct parameter it uses the index and runs in under 2 seconds.
WITH q AS (SELECT websearch_to_tsquery('english', %(text)s) AS t),
dense AS (
SELECT chunk_id, row_number() OVER (ORDER BY d) AS r
FROM (SELECT chunk_id, embedding <=> %(vec)s::halfvec(768) AS d
FROM kb_finance.chunks
WHERE NOT retired AND embedding IS NOT NULL
ORDER BY embedding <=> %(vec)s::halfvec(768)
LIMIT %(n)s) c
),
lexical AS (
SELECT chunk_id, row_number() OVER (ORDER BY rank DESC) AS r
FROM (SELECT chunk_id,
ts_rank_cd(to_tsvector('english', coalesce(context,'') || ' ' || text), q.t) AS rank
FROM kb_finance.chunks, q
WHERE NOT retired
AND to_tsvector('english', coalesce(context,'') || ' ' || text) @@ q.t
ORDER BY rank DESC LIMIT %(n)s) l
),
fused AS (
SELECT chunk_id, sum(1.0 / (60 + r)) AS rrf
FROM (SELECT chunk_id, r FROM dense UNION ALL SELECT chunk_id, r FROM lexical) u
GROUP BY chunk_id
)
SELECT c.chunk_id, f.rrf, c.source, c.chunk_index, c.text
FROM fused f JOIN kb_finance.chunks c USING (chunk_id)
ORDER BY f.rrf DESC
LIMIT %(n)s;06Add authority weighting and passage dedup
Assign each source a tier from path rules. Add a small bonus per tier. On the reference build the weight is 0.08, enough to break near-ties and too small to override relevance.
Collapse identical passages to one result. Track how many distinct passages are in the top ten as an eval metric.
A cross-encoder reranker over the top 30 is optional. Add it behind a flag and keep it only if the eval moves.
07Add graph edges for the question types vectors miss
Citation edges: when a passage says "see section X", store an edge to that passage. At query time, follow edges from top hits and add a bounded number of cited passages.
Definition edges: detect set definitional forms ("X means", "the term X refers to") at index time and store an edge from the term to the defining passage. For definitional questions, add those passages.
Both are plain tables with a from and a to column. The reference build deferred a full entity-graph system until simpler tiers were shown to fail.
CREATE TABLE brain.atom_links (
from_id text NOT NULL,
to_id text NOT NULL,
relation text NOT NULL, -- 'cites' | 'defines'
PRIMARY KEY (from_id, to_id, relation)
);08Add a claims store for cheap recall
Store short statements extracted from sources, each with a status and a pointer to its source passage. Index the statement text for full-text search.
Return the source passage beside each claim and tell the model that the passage is the authority. A claim is a pointer to evidence.
09Expose search as a tool that explains empty results
Return passages with source and position. When nothing matches, return a diagnostic naming the filter that emptied the result.
Write one trace line per search: query, candidate counts per side, final ranks. Offline diagnosis depends on it.
SCHEMA = {
"type": "function",
"function": {
"name": "search_knowledge_base",
"description": "Search a document collection. Returns passages with source and position.",
"parameters": {
"type": "object",
"properties": {
"collection": {"type": "string", "enum": ["FINANCE", "K12"]},
"query": {"type": "string"},
"top_k": {"type": "integer", "default": 6}
},
"required": ["collection", "query"]
}
}
}How it works with the other parts
- Harness engineeringThe planner decides when to search and may search again after reading results. Retrieval is a tool, so its output size is capped by the harness.
- Guardrails and verificationThe guard compares the answer with what retrieval returned in this request. Retrieval output is the evidence.
- EvaluationA sealed gold set scores retrieval alone, so a chunking or ranking change shows up as a number.
- Self-evolving agentsThe nightly learner reads chunks and writes claims. Held-out chunks are excluded so the learner cannot study the test.
What went wrong in the real build
A corrupt vector segment stopped sync for four weeks
- What happened
- The single-process vector store had one unreadable segment. Search on the main collection failed and ingestion could not proceed.
- Fix
- Move vectors into Postgres tables with write-ahead logging and a rebuildable index.
- Lesson
- Treat the vector index as derived data that can be rebuilt from rows you trust.
Every search scanned the whole table
- What happened
- The query vector was passed through a CTE. The planner could not use the HNSW index and each search took 5.4 seconds.
- Fix
- Pass the vector as a query parameter in the ORDER BY clause.
- Lesson
- Run EXPLAIN on the retrieval query. An index that exists is not an index that is used.
One collection leaked into another
- What happened
- Finance passages appeared in education-scoped chats because the knowledge attachment was set per model.
- Fix
- Separate indexes per collection and a deterministic scope filter.
- Lesson
- Isolation belongs in the data layout. A prompt instruction is not a boundary.
The top ten was one paragraph ten times
- What happened
- The same paragraph existed in several files and filled the result list.
- Fix
- Collapse duplicate passages and track distinct passages in the top ten as a metric.
- Lesson
- Measure diversity of results as well as rank of the right one.
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Vector Search with a Delta Sync index replaces the chunk table and HNSW index, and it offers hybrid search. Chunks live in a Delta table governed by Unity Catalog. Graph edges are Delta tables you join. Lakebase gives you Postgres if you want to keep this SQL as written. |
| IBM watsonx | watsonx.data provides the lakehouse and a vector engine (Milvus). Orchestrate agents attach knowledge bases backed by it. You keep the same pipeline stages: validate, chunk, embed, hybrid query, rerank. |
| Codex | Codex does not provide a retrieval store. Run this stack as a service and expose search as an MCP tool. Codex then calls it like any other tool. |
| Cursor | Cursor indexes your code for its own use. Domain retrieval is yours to host. Register your search service as an MCP server in the project config. |
| Claude Code / Agent SDK | Expose the search tool through an MCP server. The reference build does this so a coding agent and the chat agent share one retrieval implementation. |
| Another machine | Postgres with pgvector runs anywhere, including Docker. The SQL in this module is unchanged. Only the embedding endpoint URL differs. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why hybrid search and not vectors alone?
Why does rank fusion use ranks and not scores?
Why must knowledge management come before retrieval?
Why are graph edges added after hybrid search works?
Explain the claims store in one sentence to someone who knows RAG.
Terms
| RRF | Reciprocal rank fusion. Score is the sum over lists of 1/(k + rank). |
| HNSW | A graph index for approximate nearest-neighbor search over vectors. |
| hit@k | Share of questions whose gold passage is in the top k results. |
| Authority tier | A trust level assigned to a source by path rule, used as a small ranking bonus. |
From the live build
Recent changes and files the sync job filed under this module.
- Corroboration needs two independent texts: copies of one passage count once; duplicate the regulation copy retired from the learner queue
- recall_claims returns the source passage beside each claim and tells the model the passage is the authority
- Claim fidelity eval: a sample of served claims is checked against the cited passage by the 35B model; weekly, with a report line
- Knowledge graph: 'defines' edges from set definitional forms, used in retrieval for definitional questions; gold set and nightly eval
- Answers that used tools stream line by line as they clear the guard; new chunks get their Qwen3 embedding after ingest
- Reranker: cap and release the MLX buffer cache (process had grown to 71 GB)
- Tests: isolate conversations.db and the retrieval trace
- Skills: '_none' margin for embedding routing; default chosen by the eval
- FINANCE: name the actual authority, don't just summarize — and show it in practice, not just recite the rule
- CLAUDE.md — build-host Mac Studio
- Full-stack audit — 2026-09-07
- Whole-System Audit and Revised Update Plan
- Eval set v2 — built 2026-08-19
- K12: name the standard, match the grade, teach it like a tutor — not a textbook
- guards on the retrieval changes of 2026-09-30. Run: venv/bin/python3 test_retrieval.py (no database, no network) 1. SQL v2 renders for nomic (768) and Qwen3 (1024) with no stray fields, and passes the query vector as a parameter (the 5.4 s bug was a CTE) 2.
- does search_knowledge_base find the right passage? (2026-09-30) The weekly quiz measures whether a 9B model answers better with learned atoms. Nothing measured retrieval itself, so a chunker, embedding or ranking change could make search worse and no number would move.
- the search_knowledge_base tool. This is the single highest-leverage piece of Phase 7 step 2 (v2 guide §4).
- the same paragraph from two copies of a document is one result. The the regulation is in the bank three times: the per-volume PDFs, Download-the regulation-PDF and Full-the regulation-Policy-PDF. Appropriations acts and committee reports repeat each other.
- the "defines" edge of the knowledge graph. venv/bin/python3 nightly/defines.py --index # scan new chunks (incremental) venv/bin/python3 nightly/defines.py --stats venv/bin/python3 nightly/defines.py --build-gold # once; sealed The citation layer links a passage to what it CITES.
- which term does a passage DEFINE, and which term does a question ask about. The citation layer answers "which passages cite 31 U.S.C. 1341". This is the other typed edge a budget analyst needs: "which passage says what an 'unliquidated obligation' IS".
- Standalone reranker server for build-host agent-server's rerank_route.py. Added 2026-08-03.
- searchable index of past conversations (2026-10-01). The audit scored memory 2 of 5: conversations.db keeps every exchange for 90 days and nothing can search it, so "what did we decide about X last week" has no answer.