Map / Outline
Knowledge management
Collect, validate, name, tier, and retire documents before any index sees them.
What it is and why it exists
What
Knowledge management is the pipeline and the rules that decide what enters the knowledge bank: where documents come from, how they are validated and named, which are authoritative, and when a document is retired.
Why
Retrieval ranks whatever it is given. The reference build's own guide calls this work 90 percent librarianship and 10 percent technology. Folder taxonomy, file names, and pruning duplicates moved quality more than any algorithm change.
How it works
- New files land in an inbox. A validation step checks type by content, rejects error pages saved as documents, and files each item into a collection folder.
- A ledger records every file with its content hash, collection, processing stage, and status. Re-ingestion skips unchanged files. Retired files stay in the ledger with a reason.
- A rules file assigns each source an authority tier by path. Primary rules outrank guidance, which outranks commentary.
- Scheduled collectors pull feeds through a relevance filter so off-topic reports do not enter the queue.
- A second tier of curated wiki pages, drafted by a model and approved by a person, captures judgment that no source document contains.
Where it sits in the build order
Needs first
Nothing. You can start here.
Unlocks
- RAG and knowledge graphRetrieval can only return what was ingested, and it ranks by signals created at ingestion: chunk boundaries, source paths, authority tiers. Curation mistakes become retrieval mistakes that no ranking fix removes.
- EvaluationHeld-out material is selected from the corpus. You cannot seal chunks that have not been ingested.
In the reference build
| Path | Role |
|---|---|
| KNOWLEDGE_MANAGEMENT_GUIDE.md | The three-tier plan: vector search, curated wiki, graph. |
| knowledge-bank/_inbox/ | Drop zone for new files. |
| knowledge-bank/source_authority.json | Path rules for authority tiers. |
| apps/llm-wiki/llm_wiki.py | Drafts one page per concept for human review. |
| COLLECTION-PLAN.md | What is collected, from where, on what schedule. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Land raw files in a Unity Catalog volume. Use a Lakeflow pipeline for validation and chunk tables. The ledger is a Delta table, and tiers are a column. |
| IBM watsonx | Store documents in watsonx.data or object storage and register them as knowledge sources. Keep the ledger and tier rules as your own tables. |
| Codex | Not provided. Keep the bank as a repo or bucket with the same ledger and let agents reach it through a search tool. |
| Cursor | Not provided. Same approach: your pipeline, your ledger, exposed through a tool. |
| Claude Code / Agent SDK | Not provided by the coding agent. Project knowledge features hold small document sets. A large bank still needs this pipeline. |
| Another machine | Folders, a hash ledger, and a rules file work on any filesystem. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why keep retired files in the ledger?
A strong answerSo the system can explain why a passage no longer appears, and so a re-download of the same file is recognized and skipped.
A 56 GB folder of CSV files sits in the document bank. Should it be chunked and embedded?
A strong answerNo. Tabular records are an analysis corpus. Load them into a query engine and expose query tools. Chunking rows into passages produces confident nonsense.
From the live build
Recent changes and files the sync job filed under this module.
- Claim fidelity: sample by bank folder, not by chunk
- Learner queue: tier 3 sorts after core material; documents that get no vocabulary pass move to claims in one statement
- Collector: feed filter matches broad terms on the title only, domain organizations on title and summary
- Collector: per-feed relevance filter; the oversight body all-reports and legal feeds keep only domain and financial-management items
- the spending dataset collector: fiscal-year scope never starts later than FY2021
- Stream eval in the weekly cycle and the report; spending skill: a follow-up means another query; tests keep out of the production guard log; wiki watcher drops queued triggers; tuner skips manual-only sources
- Answers that used tools stream line by line as they clear the guard; new chunks get their Qwen3 embedding after ingest
- Guard: a requested made-up example is recorded, not rejected; wiki ingest runs 23:00-07:00
- Eval set v2 — built 2026-08-19
- 02 — The Build Process, Step by Step
- 00 — Project Timeline: How build-host Got Built
- Self-assessment: the learning system after ~100 hours of autonomous running
- AIDATA — Consolidation & Streamlining Assessment
- 05 — Gap Analysis
- the same paragraph from two copies of a document is one result. The the regulation is in the bank three times: the per-volume PDFs, Download-the regulation-PDF and Full-the regulation-Policy-PDF. Appropriations acts and committee reports repeat each other.
- Stage 1 of learning: KNOW WHAT YOU HAVE. This is the layer that was missing, and its absence is why the first attempt at Diet B was wrong.
- go over EVERY file, in stages, resumably. This replaces the first Diet B, which walked 81,636 chunks in database order and had no idea what any of them were.
- the "hook" that auto-triggers llm_wiki.py when the knowledge banks change. What it does: watches knowledge-bank/ (excluding Wiki/, _index/, _archive/, _inbox/) with fswatch, and on any change, runs `llm_wiki.py ingest` for BOTH banks.
- a local, Karpathy-pattern "LLM Wiki" for the AI_DATA knowledge banks. Three layers (see knowledge-bank/Wiki/WIKI-SCHEMA.md for the full spec): 1. RAW — knowledge-bank/<Bank>/** (immutable source documents — never touched) 2.
- turn a source file into clean, structure-aware chunks. Replaces the extraction and chunking half of kb_ingest.py (2026-09-22). WHAT WAS WRONG pypdf's extract_text() returns a flat string per page with no notion of headings, columns, tables or running headers.
- build the chunk index that tools/knowledge_base.py searches. 2026-09-22 — REWRITTEN FOR POSTGRES + PGVECTOR. What changed and why: STORE chroma -> kb_finance.chunks / kb_k12.chunks (pgstore.py).
- put the code and scripts on AI_DATA under git (2026-09-30). Idempotent: re-running commits whatever changed since, with the same checks. 1. whitelist .gitignore (only code + config; secrets and data never) 2.