Agent Build Tutor
Map / Outline

Knowledge management

Collect, validate, name, tier, and retire documents before any index sees them.

What it is and why it exists

What

Knowledge management is the pipeline and the rules that decide what enters the knowledge bank: where documents come from, how they are validated and named, which are authoritative, and when a document is retired.

Why

Retrieval ranks whatever it is given. The reference build's own guide calls this work 90 percent librarianship and 10 percent technology. Folder taxonomy, file names, and pruning duplicates moved quality more than any algorithm change.

How it works

Where it sits in the build order

Needs first

Nothing. You can start here.

Unlocks

  • RAG and knowledge graphRetrieval can only return what was ingested, and it ranks by signals created at ingestion: chunk boundaries, source paths, authority tiers. Curation mistakes become retrieval mistakes that no ranking fix removes.
  • EvaluationHeld-out material is selected from the corpus. You cannot seal chunks that have not been ingested.

In the reference build

PathRole
KNOWLEDGE_MANAGEMENT_GUIDE.mdThe three-tier plan: vector search, curated wiki, graph.
knowledge-bank/_inbox/Drop zone for new files.
knowledge-bank/source_authority.jsonPath rules for authority tiers.
apps/llm-wiki/llm_wiki.pyDrafts one page per concept for human review.
COLLECTION-PLAN.mdWhat is collected, from where, on what schedule.

The same idea on other platforms

PlatformHow this module maps
DatabricksLand raw files in a Unity Catalog volume. Use a Lakeflow pipeline for validation and chunk tables. The ledger is a Delta table, and tiers are a column.
IBM watsonxStore documents in watsonx.data or object storage and register them as knowledge sources. Keep the ledger and tier rules as your own tables.
CodexNot provided. Keep the bank as a repo or bucket with the same ledger and let agents reach it through a search tool.
CursorNot provided. Same approach: your pipeline, your ledger, exposed through a tool.
Claude Code / Agent SDKNot provided by the coding agent. Project knowledge features hold small document sets. A large bank still needs this pipeline.
Another machineFolders, a hash ledger, and a rules file work on any filesystem.

Explain it back

Answer aloud first. Then open the answer and compare.

Why keep retired files in the ledger?
A strong answerSo the system can explain why a passage no longer appears, and so a re-download of the same file is recognized and skipped.
A 56 GB folder of CSV files sits in the document bank. Should it be chunked and embedded?
A strong answerNo. Tabular records are an analysis corpus. Load them into a query engine and expose query tools. Chunking rows into passages produces confident nonsense.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module