Live build
What changed in the reference build
A job on the build machine scans the project, removes anything sensitive, sorts each change into a module, and pushes a snapshot here. The status line shows whether the machine is reachable right now.
Checking the local build
2026-10-08snapshot date (UTC)
58changes
44design notes
136source files
74redactions applied
Latest eval runs
| Eval | When | Metrics |
|---|---|---|
| retrieval K12 main dedup-on | 2026-10-07 | n 100 · hit@1 0.55 · hit@5 0.7 · hit@10 0.74 · near@10 0.79 · doc@10 0.84 · mrr 0.625 · p50_seconds 1.76 |
| retrieval FINANCE define dedup-on | 2026-10-07 | n 60 · hit@1 0.45 · hit@5 0.75 · hit@10 0.8 · near@10 0.85 · doc@10 0.933 · mrr 0.567 · p50_seconds 1.87 |
| retrieval FINANCE cite dedup-on | 2026-10-07 | n 60 · hit@1 0.517 · hit@5 0.783 · hit@10 0.85 · near@10 0.883 · doc@10 0.933 · mrr 0.642 · p50_seconds 1.89 |
| retrieval FINANCE main dedup-on | 2026-10-07 | n 100 · hit@1 0.37 · hit@5 0.66 · hit@10 0.7 · near@10 0.76 · doc@10 0.89 · mrr 0.492 · p50_seconds 1.9 |
| retrieval FINANCE define dedup-off | 2026-10-07 | n 60 · hit@1 0.45 · hit@5 0.817 · hit@10 0.917 · near@10 0.917 · doc@10 0.967 · mrr 0.601 · p50_seconds 2.14 |
| retrieval FINANCE cite dedup-off | 2026-10-07 | n 60 · hit@1 0.517 · hit@5 0.85 · hit@10 0.95 · near@10 0.95 · doc@10 0.983 · mrr 0.663 · p50_seconds 3.45 |
| retrieval FINANCE main dedup-off | 2026-10-07 | n 100 · hit@1 0.37 · hit@5 0.68 · hit@10 0.76 · near@10 0.81 · doc@10 0.91 · mrr 0.511 · p50_seconds 2.09 |
| retrieval FINANCE define tier-check | 2026-10-07 | n 60 · hit@1 0.45 · hit@5 0.817 · hit@10 0.917 · near@10 0.917 · doc@10 0.967 · mrr 0.601 · p50_seconds 1.86 |
| retrieval FINANCE cite tier-check | 2026-10-07 | n 60 · hit@1 0.517 · hit@5 0.85 · hit@10 0.95 · near@10 0.95 · doc@10 0.983 · mrr 0.663 · p50_seconds 2.93 |
| retrieval FINANCE main tier-check | 2026-10-07 | n 100 · hit@1 0.37 · hit@5 0.68 · hit@10 0.76 · near@10 0.81 · doc@10 0.91 · mrr 0.511 · p50_seconds 1.89 |
| retrieval FINANCE cite define-check | 2026-10-07 | n 60 · hit@1 0.517 · hit@5 0.85 · hit@10 0.95 · near@10 0.95 · doc@10 0.983 · mrr 0.663 · p50_seconds 2.82 |
| retrieval FINANCE main define-check | 2026-10-07 | n 100 · hit@1 0.37 · hit@5 0.69 · hit@10 0.77 · near@10 0.81 · doc@10 0.9 · mrr 0.512 · p50_seconds 2 |
Recent changes
Each line is a commit subject. They are written as lessons learned.
- Corroboration needs two independent texts: copies of one passage count once; duplicate the regulation copy retired from the learner queueSelf-evolving agentsRAG and knowledge graph
- Learner: 8 minutes of source checks in every window, so claim verification keeps pace with what is writtenSelf-evolving agents
- Search returns one copy of a paragraph that is in several files; same@10 and distinct@10 in the retrieval eval; dated scheduler log; notice when a no-query draft is replaced; skill: a status is half an answerEvaluationSkills
- Claim judges: the agency and the agency are one name; claims rejected for that are checked againSelf-evolving agents
- claim_verify: a simplification alone keeps a claim in service; --redecide re-applies the rule to recorded votesSelf-evolving agentsGuardrails and verification
- Claim verification live: nightly two-vote check, claims failing both leave service as 'unsupported'; omission alone keeps a claim; VERIFY line in the daily reportSelf-evolving agents
- claim_verify: two-vote check of served claims against their source (dry run and calibration); off-topic the oversight body reports retired from the learner queueSelf-evolving agentsGuardrails and verification
- Out of tool rounds: one final turn without tools answers from the evidence gathered; restate standing after promotion; ordered promotion queueHarness engineeringSelf-evolving agents
- BUSY sentinel is held until a streamed response finishes, so the learner pauses for the whole chat answerSelf-evolving agentsHarness engineering
- recall_claims returns the source passage beside each claim and tells the model the passage is the authoritySelf-evolving agentsRAG and knowledge graph
- Claim fidelity: sample by bank folder, not by chunkSelf-evolving agentsKnowledge management
- Claim fidelity eval: a sample of served claims is checked against the cited passage by the 35B model; weekly, with a report lineSelf-evolving agentsRAG and knowledge graph
- Learner: claims queue rotates through authority, newest and reports lanes; gate-blocked claims are retired, not left in quarantineSelf-evolving agents
- Claims: restate_standing re-rates every claim in one statement; corroboration counts documents; legacy trusted claims re-rated and gate-checkedSelf-evolving agents
- Authority: hearings, committee prints and mandated reports are tier 2 and typed other; their single-source claims are no longer trustedSelf-evolving agents
- Concept promotion: cap what is promoted, not what is read; the pass had been stuck behind 500 gate failuresSelf-evolving agents
- Trusted vocabulary: definition authorities only, placeholders and headings rejected, existing rows demoted; claims pass gets trusted terms onlySelf-evolving agents
- Knowledge graph: 'defines' edges from set definitional forms, used in retrieval for definitional questions; gold set and nightly evalRAG and knowledge graphEvaluation
- Learner queue: tier 3 sorts after core material; documents that get no vocabulary pass move to claims in one statementSelf-evolving agentsKnowledge management
- Collector: feed filter matches broad terms on the title only, domain organizations on title and summaryKnowledge management
- Collector: per-feed relevance filter; the oversight body all-reports and legal feeds keep only domain and financial-management itemsKnowledge management
- Conversation learner ignores eval threads; report treats findings from unchanged threads as no new trafficState and storage
- the spending dataset collector: fiscal-year scope never starts later than FY2021Knowledge managementAPI and gateway
- query_spending returns the total and names the period it covers; award warehouse builds itself when a newer archive lands, with a report alarmTools and MCP
- Guard: on a spending turn a draft sourced from a web search alone is rejected and sent to the spending tools; skill says the sameGuardrails and verificationSkills
- Stream eval in the weekly cycle and the report; spending skill: a follow-up means another query; tests keep out of the production guard log; wiki watcher drops queued triggers; tuner skips manual-only sourcesSkillsKnowledge management
- Answers that used tools stream line by line as they clear the guard; new chunks get their Qwen3 embedding after ingestKnowledge managementRAG and knowledge graph
- Guard: a requested made-up example is recorded, not rejected; wiki ingest runs 23:00-07:00Knowledge managementGuardrails and verification
- Collector: the agency IG and the oversight site through the the oversight site listing (two-step), executive orders query fixed, closed the oversight body archive marked manualKnowledge management
- Gateway answers the watchdog probe; guard: no-tool check on data turns only, strict first draft; wiki samples start, middle and end; collector tuner floor and restored cadencesOperations and deploymentKnowledge management
- Answer guard: vendor names, answers given without a tool call, streaming gate; FINANCE on Qwen3 embeddings; wiki watcher batch cap, lock and wait; downloads scriptAPI and gatewayTools and MCP
- Chat eval: the round-limit message counts as a fallback, not a bad answer deliveredGuardrails and verificationEvaluation
- Answer guard checks counts; chat eval in the Sunday cycle; CHAT line in the daily reportGuardrails and verificationEvaluation
- Reranker: cap and release the MLX buffer cache (process had grown to 71 GB)RAG and knowledge graphInference engineering
- Answer guard: plain message when the model returns nothing twice; chat eval allows 6000 output tokensGuardrails and verificationEvaluation
- Retry model calls that return HTTP 500 (malformed tool call) and empty final answersTools and MCP
- Answer guard: dollar figures must come from tool results (both answer paths); follow-ups inherit the skill; multi-turn chat evalTools and MCPSkills
- Memory store on Postgres, second attempt (qualified upsert columns); recall skips the current thread, memory briefing cannot fail a requestMemoryState and storage
- Memory store on Postgres (mem schema), recall skips the current thread, memory briefing cannot fail a requestState and storageMemory
- Recall similarity floor 0.35 (eval: lowest hit 0.55, highest unrelated 0.24)MemoryEvaluation
- Conversation recall: mem.turns index, recall_conversations tool, recall gate in the daily cycleMemoryState and storage
- Chain audit: files removed from disk are counted separately, not as awaiting ingestOperations and deploymentKnowledge management
- Daily report: unreadable learning store is one explicit problem, never zeros; backfill status merges per collection; collector pause expires before the same-hour runOperations and deploymentKnowledge management
- Weekly agent quiz in the daily cycle; AGENT line in the reportEvaluation
- Audit fixes: self-gap dedupe, SQLite stores in the nightly backup; agent quizOperations and deploymentState and storage
- Tests: isolate conversations.db and the retrieval traceState and storageRAG and knowledge graph
- Skills: route by natural-language examples only (descriptions as fallback)SkillsGuardrails and verification
- Skills: '_none' margin for embedding routing; default chosen by the evalSkillsRAG and knowledge graph
- Skills: embedding routing, up to two skills per request, routing evalSkills
- Citation expansion on by default; citation index refreshed daily
- Citation layer: kb_finance.citations, citation gold set, graph expansion (off)Evaluation
- Daily regression tests; Qwen3 switch readiness in the reportEvaluation
- kb_extract: OCR scanned PDFs with macOS Vision
- K12 on Qwen3 embeddings, search tracing, retrieval tests, ingest fixRAG and knowledge graphKnowledge management
- Qwen3 embeddings (off by default) and a daily retrieval checkRAG and knowledge graphEvaluation
- Retrieval SQL v2: use the HNSW index (search 5.4 s -> about 1 s before rerank)RAG and knowledge graph
- Retrieval: Qwen3 reranker on by default (gold set hit@10 0.34 -> 0.62, MRR 0.162 -> 0.414)RAG and knowledge graphEvaluation
- Baseline 2026-09-30: code and scripts after the Postgres cutover, pinned MCP servers, nightly pg_dumpTools and MCPState and storage
Design notes
- FINANCE: name the actual authority, don't just summarize — and show it in practice, not just recite the ruleCitation (unchanged from Step 3 — the "just the facts" requirement) · Depth (added 2026-07-28, from graded baseline feedback) · A status is half an answer (added 2026-10-07, from the agent quiz)RAG and knowledge graphEvaluation
- Spending orchestration0. The question behind every question · 1. Routing table · 2. Mandatory sequences · 3. The reconciliation rule · 3b. DEFC — which law the money came from · 4. Status discipline · 5. Summing discipline · 6. Two failure modes with equal weightHarness engineeringSelf-evolving agents
- CLAUDE.md — build-host Mac StudioThe one thing to know · Which tool · The skills carry the routing — use them · Where you are matters · Never assert absence · Money words are not interchangeable · Presenting results · Operating this machineRAG and knowledge graphSelf-evolving agents
- Fixes applied — 2026-09-071. Tool results are bounded — agentloop.py · 2. selfobserve.py can no longer fabricate corroboration — and 184 bad findings removed · 3. The the document service credential is redacted and scrubbed — collect.py · 4. The watchdog now actually checks health — Install-Autostart.command · 5. Twenty-one tools removed from every prompt — mcpservers.json · 6. CLAUDE.md no longer misinforms every request · One thing I got wrong, and reverted · Not touched, deliberatelyOperations and deploymentSelf-evolving agents
- Finish the Update Today — remaining stepsWhy it hung — four facts from the disk and the log · Step 1 — Maintenance window (~25 min) · Step 2 — Dynamic gateway (~20 min) · Step 3 — Rewire the router (~10 min) · Step 4 — Rotate the disclosed credentials (~5 min) · Step 5 — Retire stale models (~10 min) · Not today · Today's checklistInference engineering
- Full-stack audit — 2026-09-07The short version · 1. Confirmed live bugs · 2. What actually works — and the number that says otherwise · 3. Where the work is stuck on a human · 4. Outdated information · 5. What I would do, in order · 6. The thing that is genuinely goodRAG and knowledge graphTools and MCP
- Self-evolution: what was broken, what changed1. The experience half of the loop was switched off · 2. The findings file was 97% duplicates · 3. 58% of the server log was one misconfigured probe · 4. There was no conversation to learn from · Cleanup · Left alone, deliberately · Still open · To activateSelf-evolving agentsMemory
- Runbook — making the Claude Code CLI smart about the agencyStep 1 — make the files executable (30 seconds) · Step 2 — selftest against the real stores (2 minutes) · Step 3 — start Claude Code and approve the server (2 minutes) · Step 4 — confirm the skills were discovered ⚠️ THE UNCERTAIN ONE · Step 5 — the test that means something · Step 6 — measure, do not assume · Step 7 — decide three things I could not decide for you · Step 8 — the two things that make this fragileSkillsTools and MCP
- AI Local Server → Vercel Integration GuideMulti-Model Chain Architecture with Intelligent Routing · 🔴 2026-08-06 — TTS went from working to 501. Root cause, and why the error message mislead · 🏁 Finished this pass (2026-08-03, round 2) — everything I could do without your Mac's shel · 📚 Institutional knowledge source (added 2026-08-03) · 🧩 Root cause: why 4 downloaded models don't show in ollama list or "the venv" (confirmed 2 · 🔍 MiniMax-H3 evaluated as a vision replacement (2026-08-05) — not swapped in, here's why · ✅ Implementation Status (2026-08-03) — What's Actually Built · 🔄 "Legacy" agent-server vs "multi-model" agent-server — one service, not twoOperations and deploymentInference engineering
- SOP — Operating the Local LLM ServerPart 0 — Quick reference · Part 1 — What this system is · Part 2 — How to use it · Part 3 — Monitoring · Part 4 — Maintenance · Part 5 — Updating · Part 6 — Backup and restore · Part 7 — TroubleshootingOperations and deploymentAPI and gateway
- Claude Code → your own models0 · What this is, and what it does not touch · 1 · Before you start — five checks · 2 · Install LiteLLM · 3 · Create the key · 4 · Start it and prove it works · 5 · Point Claude Code at it (on the Mac Studio) · 6 · The dropdown · 7 · Make it survive a rebootAPI and gatewaySDKs and frameworks
- Bug Fixes — 2026-08-20, post-updateP0 — Live outage: your gateway currently allows zero working models · P1 — The 401 was a wrong key, not a broken gateway · P2 — The watchdog is down and paused · P3 — You have no coding model · P4 — Correction: you don't have video generation. I was wrong. · P5 — The MLX swap worked, but the benchmark is not what it looks like · P6 — Disk reclaimed correctly; your df reading was just taken too early · P7–P9 — Lower priority, worth a look this weekAPI and gateway
- Dynamic Gateway + Full-Capability Model LineupPart 1 — Make the allow-list dynamic · Part 2 — Full-capability lineup · OrderAPI and gateway
- Whole-System Audit and Revised Update Plan1. What you've actually built · 2. Three findings that change the update plan · 3. Exposure review — the public surface · 4. Is the learning loop actually learning? · 5. Revised update plan · 6. Fit check · 7. Suggested order · SourcesInference engineeringRAG and knowledge graph
- AI Stack Update Runbook — 2026-08-20Two corrections to what I told you first · TL;DR — what actually buys you speed · Baseline — recorded 2026-08-20 · Finding — your vision rollback fallback does not exist · Step 1 — Upgrade Ollama (low value, do it anyway) · ⚠️ Homebrew hazard list — do NOT run bare brew upgrade · Step 2 — Move your chat models onto the MLX engine · Step 3 — The two big models worth a decisionInference engineering
- Eval set v2 — built 2026-08-19What changed, and why each change was necessary · citecheck.py — the key no longer rests on me · Running it · RESULT — 2026-08-18 21:16, and my predictions scored · What I expected, written down before the runRAG and knowledge graphKnowledge management
- Issues 7, 8 and 9 — diagnosis and fix0. RESULT — the quiz ran, and it settled the question · 0. The eval is saturated, and my ranking of the retrievers was wrong · 0-ter. Current result, and one retraction · 0-bis. Why atoms sat at exactly 27/30, five runs running · 0a. STOP — every score above this line is suspect · 0b. Why 65.5%, and what actually gets you higher · 1. The headline · 2. The trap, and why the obvious order is backwardsEvaluationMemory
- The Role Dimension — Architecture v20. What changed from v1, and why · 1. The learning-science foundation · 2. Authority research: what the role dimension can be built from · 3. The graph model · 4. How the graph thinks — named traversals · 5. Measured feasibility — the locus and artifact layers are nearly free · 6. Build sequence, revised · 7. Failure modesMulti-agent patterns
- The Role Dimension — Architecture0. The one-paragraph version · 1. Why not a corpus per role · 2. The axis model · 3. What a capability actually contains · 4. The binding problem — and the prerequisite nobody will notice until it blocks them · 5. Roles must be ingested, not invented · 6. The corner roles · 7. Build sequenceMulti-agent patterns
- 01 — System ArchitectureThe one-paragraph version · Layer 1 — Hardware and the model engine · Layer 2 — Two front doors, on purpose · Layer 3 — Knowledge, in three tiers · Layer 4 — The agent-server request pipeline · Layer 5 — Networking and exposure · Layer 6 — Supporting apps · Layer 7 — BackupInference engineering
- 02 — The Build Process, Step by StepPhase 0 — Mac foundation · Phase 1 — The engine · Phase 2 — The rest of the model stack · Phase 3 — Making it usable: Open WebUI and knowledge banks · Phase 4 — Connecting everything else · Phase 4½ — The public demo bridge · Phase 5 — Opening it up to a small team · Phase 6 — BackupKnowledge managementInference engineering
- 00 — Project Timeline: How build-host Got BuiltThe shape of the week · Day 0 — July 24: Foundation · July 25 — The engine, the models, and the first public face · July 26 — Knowledge banks and the LLM Wiki · July 27 — Fixing retrieval, and the agent is born · July 28 — The big build day · July 29 — Streaming, latency, and turning it into an operable system · July 30 (today) — Diagnosing and fixing a real latency bugKnowledge managementAPI and gateway
- 07 — How to Change or Add LLM ModelsFirst, understand why this is easy · The VRAM math you must do first · Adding a model · Removing a model · Trying a model before committing · Changing the embedding model — different process, do this carefully · Setting a default, or forcing a specific model for specific traffic · Tuning how Ollama serves whatever models you haveInference engineering
- AIDATA Local AI Server — Master Guide (Phased Plan)Phased Roadmap (Multi-Purpose Server) · ⭐ OPERATIONS BOARD — Status, Usage & Maintenance (as of 2026-07-24) · Phase 7 — Building Agentic AI on This Foundation — ⚠️ MOVED to its own doc, v2, 2026-07-27 · Software Shopping List — Everything Is Free · 1. What This Machine Can Run · 2. Top 5 Candidates · Phase 0 — Mac Foundation (Day 0) · Phase 1 — Core AI Server (MAIN model first)Inference engineeringOperations and deployment
- Merged into spending-investigation.md on 2026-08-15.Skills
- BUILD-HOST — user guide1. What the agent can do that it could not a week ago · 2. How to ask · 3. Reading the answers · 4. What it cannot do · 5. Vercel front-end — two changes needed · 6. Troubleshooting · 7. A worked exampleSelf-evolving agentsEvaluation
- What the failures teach about making the agent actually get smarter1. The single pattern behind every failure · 2. Three structural rules that follow · 3. The learning loop is not learning from any of this · 4. Where the epistemic model must extend · 5. Honest ranking of what to build next · 6. What the detector found on its first run · 7. Auto-relaxation, and the failure it nearly introduced · 7b. The fix that broke production, four hours laterSelf-evolving agentsGuardrails and verification
- BUILD-HOST — Operations RunbookStatus · PHASE A — Seal the eval sets · PHASE B — Put services under supervision · PHASE C — Three production fixes · What each patch actually does · PHASE D — Log hygiene · PHASE F — The scheduler (replaces manual Phase E) · PHASE G — Review, then switch to applySelf-evolving agentsOperations and deployment
- Self-assessment: the learning system after ~100 hours of autonomous runningThe headline numbers, and why each one is misleading · A. What genuinely works · B. What is broken · C. What was designed and never built · D. Honest assessment of my own design · E. What to fix, in orderSelf-evolving agentsKnowledge management
- BUILD-HOST — Cognitive Architecture & Self-Evolving Learning Loop0. The premise, stated honestly · 1. Design principles · 2. The brain map · 3. Memory architecture · 4. The nightly loop · 5. The promotion gate · 6. What the agent must never do at night · 7. Attention — fixing the wake-up routerSelf-evolving agentsSkills
- Stage 0 patches — ready to apply, deliberately NOT appliedPATCH 1 — TTS: stop synthesizing non-speech, stop returning 200 on garbage · PATCH 2 — Log nresults and topscore (2 lines, unblocks Diet A) · PATCH 3 — Token cap → 413 before dispatch · PATCH 4 — Preemption sentinel (required by R1, §4.0.3) · PATCH 5 — LaunchAgents with KeepAlive (Stage 0a) · PATCH 6 — Log rotation (Stage 0b)Inference engineeringAPI and gateway
- Production findings — surfaced by the first Phase 1 replay runFINDING 1 — TTS returns HTTP 200 with ~4 seconds of garbage audio · FINDING 2 — 157 HTTP 502s, 6.9% error rate for BrainBank · FINDING 3 — the highest-value failure signal is not being logged · FINDING 4 — 97% of tool-using requests never iterate · What to do, in orderInference engineering
- AIDATA — Consolidation & Streamlining AssessmentVerdict in one paragraph · 1. Disk — what is actually earning its space · 2. "Two environments — mandatory or not?" · 3. Scripts — 11 launchers, overlapping responsibility · 4. Documentation — 84,000 words, heavily overlapping · 5. Effectiveness — what's built vs. what's wired · 6. Prioritized plan · 7. One thing to watchInference engineeringKnowledge management
- 08 — Extending the Harness AgentAdding a new local tool · Adding a new MCP server · Adding a new skill · Adding a new hook · Using or adding a new model · Structural extensions worth building next · A general principle for extending this systemTools and MCPSkills
- 06 — Improvement RoadmapDo now — cheap, high-value, low-risk · Do soon — real work, real payoff · Consider later — bigger bets · A note on sequencingAPI and gatewayOperations and deployment
- 05 — Gap AnalysisSecurity and secrets · Knowledge base completeness · Retrieval and RAG quality · The agent harness itself (from reading the code directly) · Fine-tuning · Documentation and housekeeping · How to read this listKnowledge managementOperations and deployment
- 04 — Operations Cheat SheetThe five buttons (double-click at the root of AIDATA) · Emergency: something's broken · The two servers — don't confuse them · Talking to the server from a terminal · Managing API keys · Adding or removing a model · Changing agent behavior · Adding documents to a knowledge bankAPI and gatewayOperations and deployment
- 03 — The Agent Harness, Deep DiveThe request, start to finish · Tools: two sources, one registry · Memory: narrow by design · Retrieval: ChromaDB plus a trust layer · The self-referential routing gate, concretely · Guardrails, as actual code · Why this design holds upAPI and gatewayHarness engineering
- Local Agent Server — Build Guide v2Quick Start — what the keys are for, and how to wire one into a Vercel app · Master executive summary — read this first (new session / new context window) · Daily log · ACTION ITEMS — what YOU need to do (manual, on the real Mac) · 0. Mission, in one paragraph · 1. What changed from the v1 decisions (2026-07-27 → this rewrite) · 2. Target architecture · 3. The commercial-grade API key system (the MUST item)Tools and MCPMemory
- Agent-Server Latency Investigation — Summary (2026-07-30)Step 1 — Noticed the symptom · Step 2 — Went to the logs instead of guessing · Step 3 — Found the actual root cause · Step 4 — Compared the architecture against frontier agent design · Step 5 — Fixed the prompt (AGENT.md) · Step 6 — Added a deterministic backstop (Step 5.9: regex triage) · Step 7 — Real-world test exposed the gate's limit · Step 8 — Added a second, smarter tier (Step 5.10: embedding routing)Harness engineeringInference engineering
- MCP Server Catalog — candidates for mcpservers.jsonOfficial reference servers (modelcontextprotocol/servers repo) · Web Search / Web Fetch · PDF · Planning / Orchestration · Financial data (market data / filings) · Financial planning / accounting (books, invoicing) · Time series · How to actually enable any of theseTools and MCP
- K12: name the standard, match the grade, teach it like a tutor — not a textbookCitation (why this stays a "just the facts" requirement first) · Tone and depth (added 2026-07-28, from graded baseline feedback)RAG and knowledge graphKnowledge management
- AIDATA Server — High-Level Overview & Maintenance MapThe System in One Paragraph · The Pieces · How a Question Flows Through the System · What Needs Updating Later (Maintenance Map) · Current Status Snapshot (July 25, 2026 — verified on disk)Knowledge managementInference engineering
- Knowledge Management Guide — Making Your Local AI Domain-SmartThe Big Picture: Three Tiers · Tier 1 — Vector RAG Done Right (your foundation) · Tier 2 — The LLM Wiki (curated knowledge layer) · Tier 3 — Graph Knowledge (only when relationships are the question) · What Goes Where — Decision Table · Rollout PlanKnowledge managementRAG and knowledge graph
Source files
- guards on the retrieval changes of 2026-09-30. Run: venv/bin/python3 test_retrieval.py (no database, no network) 1. SQL v2 renders for nomic (768) and Qwen3 (1024) with no stray fields, and passes the query vector as a parameter (the 5.4 s bug was a CTE) 2.RAG and knowledge graphEvaluation
- the quarantined learned-knowledge store (BRAIN-ARCHITECTURE.md §3.1). Deliberately modelled on memory_store.py, right down to the conventions: contextmanager connection, schema-if-not-exists, 0600 permissions, PRAGMA journal_mode=TRUNCATE, and idempotent additive migrations via ...MemoryState and storage
- two copies of one text are one voice. venv/bin/python3 nightly/corroboration_audit.py # incremental venv/bin/python3 nightly/corroboration_audit.py --all # every claim with 2+ documents venv/bin/python3 nightly/corroboration_audit.py --stats "Corroborated" means a claim is ...Self-evolving agents
- the circadian clock (BRAIN-ARCHITECTURE.md §4.0). This is the piece that turns "a script you remember to run" into a system that runs itself. It owns two things and nothing else: R1 When the Vercel app calls the LLM server, learning PAUSES. External requests always win.Self-evolving agents
- check served claims against their source, two votes, and take the ones that fail both out of service.Self-evolving agentsGuardrails and verification
- builds the specific planner/tool graph on top of graph.py's generic engine, and exposes the same run() contract main.py already calls.Harness engineering
- a final answer may not state dollar figures the tools did not return (2026-10-02). Why: on 2026-10-02 the agent ran the right a commodity and a commodity queries, got the right rows back, and then answered with a table of an unrelated item purchases ("$18.96 Billion", "Nuccio shipyard").Guardrails and verification
- does search_knowledge_base find the right passage? (2026-09-30) The weekly quiz measures whether a 9B model answers better with learned atoms. Nothing measured retrieval itself, so a chunker, embedding or ranking change could make search worse and no number would move.EvaluationRAG and knowledge graph
- the search_knowledge_base tool. This is the single highest-leverage piece of Phase 7 step 2 (v2 guide §4).RAG and knowledge graph
- the same paragraph from two copies of a document is one result. The the regulation is in the bank three times: the per-volume PDFs, Download-the regulation-PDF and Full-the regulation-Policy-PDF. Appropriations acts and committee reports repeat each other.RAG and knowledge graphKnowledge management
- do the claims the agent serves say what their source says? venv/bin/python3 nightly/claim_fidelity.py --n 200 # print only venv/bin/python3 nightly/claim_fidelity.py --n 200 --write # + history "Authoritative" means the source document is tier 1.Self-evolving agents
- sandbox test suite for Step 5.7 (streaming + latency). Run: python3 test_streaming.py Tests against a REAL stub Ollama HTTP server (a live uvicorn process on a real port, speaking real SSE), not mocks — same discipline every prior step in this codebase used.API and gateway
- Agent Server — apps/agent-server/main.py ========================================= STEP 2 of the "mini Claude Code" build (see AGENT_SERVER_GUIDE_v2.md). Step 1 — bare passthrough, plaintext-key auth. Built/tested 2026-07-27.Tools and MCPAPI and gateway
- the agent-facing wrapper around claim_recall. REGISTERED 2026-08-18, AFTER IT EARNED IT. ------------------------------------------ This file sat unregistered until the T5 discrimination drill said it was worth registering.MemoryEvaluation
- lets the agent answer questions ABOUT ITS OWN LEARNING with facts instead of assumptions. WHY THIS EXISTS --------------- Asked "what new knowledge did you gain in the last 100 hours?", the agent answered: "Honest answer: nothing. And that's not a gap — it's by design.Self-evolving agents
- Stage 1 of learning: KNOW WHAT YOU HAVE. This is the layer that was missing, and its absence is why the first attempt at Diet B was wrong.Knowledge management
- go over EVERY file, in stages, resumably. This replaces the first Diet B, which walked 81,636 chunks in database order and had no idea what any of them were.Knowledge managementSelf-evolving agents
- the "defines" edge of the knowledge graph. venv/bin/python3 nightly/defines.py --index # scan new chunks (incremental) venv/bin/python3 nightly/defines.py --stats venv/bin/python3 nightly/defines.py --build-gold # once; sealed The citation layer links a passage to what it CITES.RAG and knowledge graphEvaluation
- which term does a passage DEFINE, and which term does a question ask about. The citation layer answers "which passages cite 31 U.S.C. 1341". This is the other typed edge a budget analyst needs: "which passage says what an 'unliquidated obligation' IS".RAG and knowledge graph
- learn from what people actually said. python3 nightly/conversation_learn.py # report, write nothing python3 nightly/conversation_learn.py --apply # record python3 nightly/conversation_learn.py --all # ignore the cursor WHY THIS EXISTS =============== This system had three ways ...Self-evolving agentsState and storage
- the federal spending tools, exposed to live traffic. Three tools, matching the three shapes a real question takes: query_spending aggregate: how much, by whom, on what, when search_requirements free text: what was ACTUALLY bought screen_purchase_request pre-award: should this be ...Tools and MCP
- the 2026-10-02 an unrelated item answer must fail the guard and the correct a commodity answer must pass. Exits non-zero on failure.Guardrails and verification
- the "hook" that auto-triggers llm_wiki.py when the knowledge banks change. What it does: watches knowledge-bank/ (excluding Wiki/, _index/, _archive/, _inbox/) with fswatch, and on any change, runs `llm_wiki.py ingest` for BOTH banks.Knowledge management
- end-to-end test through main.py's REAL FastAPI app. test_streaming.py proves agent_loop's streaming logic.API and gateway
- what the chat box actually does (2026-10-06). chat_eval.py and agent_quiz.py call agent_loop.run(), the non-streaming path. The chat box streams.API and gatewayEvaluation
- a local, Karpathy-pattern "LLM Wiki" for the AI_DATA knowledge banks. Three layers (see knowledge-bank/Wiki/WIKI-SCHEMA.md for the full spec): 1. RAW — knowledge-bank/<Bank>/** (immutable source documents — never touched) 2.Knowledge management
- how does the agent answer real multi-turn questions? (2026-10-02) The quiz and the gold sets are single questions with one right letter or one right passage. They never measured the thing a person does in the chat box: ask something, then follow up in five words.Evaluation
- the agent can look up what was said before (2026-10-01). Thin wrapper over conversation_index.search(). Registered after nightly/recall_eval.py passed its gate (see evals/memory/history.jsonl).State and storageMemory
- Standalone reranker server for build-host agent-server's rerank_route.py. Added 2026-08-03.Inference engineeringRAG and knowledge graph
- run the agent-server regression scripts (2026-09-30). Each is a plain script that exits non-zero on failure. Writes logs/tests.status.json. Safe: the brain test uses a throwaway <db>_test database and the endpoint tests use /tmp and stub model servers.Operations and deploymentAPI and gateway
- searchable index of past conversations (2026-10-01). The audit scored memory 2 of 5: conversations.db keeps every exchange for 90 days and nothing can search it, so "what did we decide about X last week" has no answer.State and storageRAG and knowledge graph
- Phase 1 of the nightly loop (BRAIN-ARCHITECTURE.md §4.1). "Hippocampal replay": read the day's traffic, find where the agent failed, turn each failure into a gap the agent can later study.Self-evolving agents
- Step 3.6 (AGENT_SERVER_GUIDE_v2.md §4/§7.1): the durable episodic + user-profile memory layer.Memory
- memory_store on Postgres behaves like it did on SQLite (2026-10-01). Runs the same calls against both backends under a throwaway partition, compares what comes back, removes its rows. Exits non-zero on a difference.Memory
- copy memory.db into Postgres schema mem (2026-10-01). Idempotent. agent:self observations are collapsed to the newest row per distinct finding (10,998 rows were 26 findings re-logged hourly). venv/bin/python3 ../ops/migrate_memory_pg.py # migrate + verifyMemoryState and storage
- collects every tool into the two things agent_loop.py needs: the OpenAI-style schema list to hand Ollama (so the model knows what it can call), and a name -> async-callable map to actually execute a call.Self-evolving agentsState and storage
- does conversation recall find the right exchange, stay quiet when nothing fits, and keep users apart? (2026-10-01) Seeds two synthetic users (partitions eval:recall-a / eval:recall-b) straight into mem.turns, asks paraphrased questions, removes the seed rows.MemoryEvaluation
- fill chunks.embedding_q3 with Qwen3-Embedding-0.6B (2026-09-30). Why: on the tier-1 gold test (evals/retrieval, embed-experiment) the Qwen3 embedder put the answering passage in the top 10 for 91% of questions against 77% for nomic-embed-text, and 96% against 89% with the ...RAG and knowledge graph
- the discrimination quiz through the REAL agent path (2026-09-30). The audit's top evaluation gap: quiz_runner.py measures the 9B model with and without atoms, never what a user gets.Evaluation
- nightly dump of the knowledge store (2026-09-30). WHAT: pg_dump --format=custom of the whole build-host database (brain.* learned atoms, kb.* file ledger, kb_finance / kb_k12 chunk vectors).Operations and deploymentState and storage
- Step 3: skills router (AGENT_SERVER_GUIDE_v2.md §4). Each file in this directory is a plain Markdown skill: a `---`-delimited frontmatter block with a one-line `description:`, then the skill body below the closing `---`.Skills
- does the router load the right skills? (2026-09-30) Gold: evals/skills/gold_v1.json, 58 hand-labelled requests (15 FINANCE, 15 spending, 12 K12, 4 that need two skills, 12 that need none).SkillsEvaluation
- the citation layer of the FINANCE knowledge graph (2026-09-30). The audit scored the knowledge graph 1 of 5: concept co-mention only, and retrieval never walks it.RAG and knowledge graphEvaluation
- find legal and financial-management citations in text and reduce each to one normal key (2026-09-30). Shared by nightly/citations.py (which indexes every FINANCE chunk into kb_finance.citations) and tools/knowledge_base.py (which looks up the keys a question mentions).Self-evolving agentsRAG and knowledge graph
- turn a source file into clean, structure-aware chunks. Replaces the extraction and chunking half of kb_ingest.py (2026-09-22). WHAT WAS WRONG pypdf's extract_text() returns a flat string per page with no notion of headings, columns, tables or running headers.Knowledge management
- build the chunk index that tools/knowledge_base.py searches. 2026-09-22 — REWRITTEN FOR POSTGRES + PGVECTOR. What changed and why: STORE chroma -> kb_finance.chunks / kb_k12.chunks (pgstore.py).Knowledge management
- does a newer embedding model find the right FINANCE passage more often than nomic-embed-text? (2026-09-30) The gold set (evals/retrieval/gold_v1.json) was built from tier-1 chunks (regulations, statute, the agency guidance).RAG and knowledge graphEvaluation
- where do the 5.5 seconds of one search go? (2026-09-30) Times each stage of search_knowledge_base for 10 gold-set questions: embedding (Ollama), dense-only SQL, lexical-only SQL, the full hybrid SQL.RAG and knowledge graph
- Backup-AI-Offsite.command [destination] [full] Writes ONE encrypted file containing everything on this Mac that cannot be recreated, to somewhere that is not this Mac. WHY THIS EXISTS AND Backup-AI.command DOES NOT REPLACE IT.Operations and deployment
- pin every MCP server in mcp_servers.json to an exact version (2026-09-30). Why: every server was launched as `npx -y <pkg>` or `uvx <pkg>`, which resolves the newest release on each cold start.Tools and MCP
- pg_restore_test.sh <dumpfile> — prove a dump restores (2026-09-30). Restores brain.*, kb.* and kb_k12.* from the dump into a throwaway database (build-host_restoretest), compares row counts with the live database, checks a K12 vector search works on the restored copy, then ...Operations and deployment
- let Backup-AI-Offsite.command run unattended (2026-09-30). Targeted edits, each asserted to match exactly once: * pause(): `read -k 1` becomes a no-op when NONINTERACTIVE is set * passphrase from the login Keychain when OFFSITE_KEYCHAIN_SERVICE is set (item created only by ...Operations and deployment
- put the code and scripts on AI_DATA under git (2026-09-30). Idempotent: re-running commits whatever changed since, with the same checks. 1. whitelist .gitignore (only code + config; secrets and data never) 2.Knowledge managementTools and MCP
- Double-click button: stops and restarts the NEW agent-server harness (apps/agent-server/main.py, FastAPI + agent_loop.py, port 8788).Operations and deploymentInference engineering
- Install-Autostart.command GOAL: after a power cut, a reboot, a sleep/wake, or a crash, every LLM service comes back BY ITSELF. No double-clicking anything. WHAT IS ALREADY TRUE (since the 2026-08-09 migration): six services are supervised by launchd with KeepAlive.Operations and deployment
- REWRITTEN by Install-Autostart.command. The services this script used to nohup are now launchd jobs, so starting them means LOADING them, not spawning a second copy.Operations and deploymentInference engineering
- move brain.db, claims_index.db and the chroma index into PostgreSQL + pgvector. One-way, resumable, verifiable.RAG and knowledge graphKnowledge management
- brainbank MCP server — exposes agent-server's tool registry to Claude Code. WHY THIS EXISTS --------------- Nothing the self-evolution pipeline learns is in the model's weights.Tools and MCP
- the missing wire between 144,386 learned atoms and a live answer, plus the instrumentation that proves whether it helps.MemorySelf-evolving agents
- creates the measurement substrate the whole learning loop depends on (BRAIN-ARCHITECTURE.md §4.5.1). RUN THIS BEFORE ANY LEARNING.Evaluation
- T5 Discrimination drill. The first thing this system has ever run that measures whether learning improved ANSWERS.EvaluationMemory
- fix what the learner studies first. venv/bin/python nightly/prioritise_corpus.py # report only venv/bin/python nightly/prioritise_corpus.py --apply # re-tier + persist WHAT WAS WRONG ============== `documents_at_stage` orders by type, then authority tier, then **size ...Knowledge management
- does the corpus actually support each eval item's key? WHY THIS EXISTS --------------- v1's answer key was mine. I wrote 18 discriminations from appropriations-law fundamentals and graded the agent against my own recollection.MemorySelf-evolving agents
- regenerates brain.html (BRAIN-ARCHITECTURE.md §8). Same pattern as the existing Eval-Dashboard.html and Knowledge-Bank-Pulls-Dashboard.html: a single self-contained HTML file with no server and no CDN, openable from Finder.Self-evolving agentsEvaluation
- regression tests for the learned-knowledge store. Run: python3 test_brain_store.py These are not exhaustive unit tests; they are guards on the SIX properties that make brain_store.py safe to run unattended.State and storageSelf-evolving agents
- the ONLY sanctioned way to reach the knowledge database. Replaces three stores, 2026-09-22: brain.db (SQLite) learned atoms, concepts, the document ledger claims_index.db (SQLite) an FTS5 copy of brain.db's serveable atoms chroma (HNSW) the chunk vectors behind ...State and storageMemory
- the PostgreSQL + pgvector store that replaced brain.db, claims_index.db and the chroma index on 2026-09-22. Idempotent. Applied by pgstore.ensure_schema() and by tools/migrate_to_pg.py. Safe to run against a live database. WHY POSTGRES brain.db SQLite, one writer at a time.State and storageRAG and knowledge graph
- what is every local model doing RIGHT NOW. python3 llm_activity.py one snapshot, in the terminal python3 llm_activity.py --watch refresh every 5s until Ctrl-C python3 llm_activity.py --html also write $AI_DATA/llm-activity.html python3 llm_activity.py --html --watch keep ...Inference engineering
- the "scan for changes" automation: ingest new/changed files in each knowledge bank, then take anything that has left the active tree (deleted outright, or moved into an `_archive/` subfolder) out of the index.Knowledge management
- embeddings for the knowledge base, batched. Same model and the same unprefixed input as every vector already stored (nomic-embed-text through Ollama), so old and new vectors share one space.RAG and knowledge graph
- Double-click button: gives Claude Code access to the knowledge bank. Built 2026-08-23, after a Claude Code session on local-qwen38 was asked about the agency a commodity spending and answered "I don't have that in my knowledge" — while brain.db held 144,781 atoms and the spending tables ...Tools and MCPSDKs and frameworks
- Check-Learning-Progress.command ONE BUTTON: is the agent actually learning, and is any of it usable yet? Built 2026-08-15. brain.html (the dashboard) shows totals.Self-evolving agentsOperations and deployment
- Check-All-LLM-Status.command ONE BUTTON: is every local LLM, every API surface, and every backend actually working right now? Built 2026-08-15.Operations and deployment
- Double-click button: backs up the irreplaceable state of the local AI stack to backups/<timestamp>/. Built 2026-07-29.Operations and deploymentState and storage
- REWRITTEN by Install-Autostart.command. Ollama, gateway.py and Open WebUI are now launchd jobs. `kill`ing them starts a fight launchd always wins and you always lose: KeepAlive restarts the process within ThrottleInterval, so the stop appears not to work.Operations and deployment
- Measure what parallel mining actually buys — 2026-09-10. The claim being tested is narrow and worth testing rather than assuming: mlx_lm.server answers one request at a time, so N servers should give close to N times the throughput UNTIL memory bandwidth saturates.Self-evolving agentsInference engineering
- re-checks the 2026-09-07 fixes, on demand. $AI_DATA/apps/agent-server/venv/bin/python3 $AI_DATA/apps/agent-server/verify_fixes.py Exits 0 if every check passes, 1 otherwise, so it can go in a cycle step.Guardrails and verificationTools and MCP
- learn from the agent's own failed tool calls. python3 nightly/self_observe.py # report, write nothing python3 nightly/self_observe.py --apply # record self-observation atoms python3 nightly/self_observe.py --all # ignore the cursor, re-read everything WHY THIS EXISTS ...Self-evolving agents
- the transcript the system never kept. WHY THIS EXISTS =============== Before 2026-09-07 this server learned from exactly two things: documents it read at night, and its own failed tool calls. It did not learn from the people using it, and the reason was not policy.State and storage
- what did the agent learn recently? The question this answers is the one you actually want to ask a system that studies while you sleep: not "is it running" (the scheduler log says that) and not "how far through the corpus is it" (corpus_pipeline --status says that), but WHAT DID ...Self-evolving agents
- measures the streaming win instead of asserting it. Replays THIS server's own real production numbers (from logs/agents/requests.jsonl: p50 26.3s, p90 88.7s, 72 tok/s effective) against a stub that generates at that same measured rate, and reports time-to-first- token vs total ...API and gateway
- Step 3.6: the ONLY code path allowed to write to memory_store.py from live traffic.MemoryState and storage
- copy-truncate rotation for AI_DATA logs. WHY COPY-TRUNCATE AND NOT MOVE: Ollama, agent-server and the MLX backends hold their log files open with O_APPEND.Operations and deploymentInference engineering
- Double-click button: adds a real WEB SEARCH tool to the agent. Built 2026-07-29, after the chatbot searched for a web-search tool and correctly reported it didn't have one. WHY IT WAS NEVER BUILT: it isn't a bug or an oversight.Tools and MCP
- Double-click button: installs the `fetch` MCP server (read a web page, convert to markdown) into its own isolated venv, verifies it actually works, and enables it in mcp_servers.json. Built 2026-07-29.Tools and MCP
- LocalAPI-Claude.command rev 2026-08-22b The one command for the Claude Code local API. Start, restart, or re-run after pulling a model. Idempotent.SDKs and frameworksOperations and deployment
- LocalAPI-Claude.command rev 2026-08-22b The one command for the Claude Code local API. Start, restart, or re-run after pulling a model. Idempotent.SDKs and frameworksAPI and gateway
- U7/U8: the account layer, and the reconciliation. Two tools: budget_execution the appropriation's own story: total resources -> obligated -> outlaid -> unobligated, by Treasury Account, program activity or object class. THE ONLY PLACE OUTLAYS LIVE.Harness engineering
- supplemental_appropriations: emergency and disaster money, by the law that created it.Harness engineering
- Disaster Emergency Fund Codes: the shared vocabulary. DEFC is how a federal dollar says which emergency, disaster or supplemental law it came from.SDKs and frameworksKnowledge management
- QUICKSTART-10-MIN.command Double-click this. It does everything: 1. installs litellm (if missing) 2. generates a key 3. writes a minimal config pointed at whichever model you actually have 4. starts the proxy in the background 5. tests it end to end 6.API and gateway
- verify-claude-code-gateway.sh Proves the proxy works BEFORE you open Claude Code, so a failure points at one specific thing instead of at "Claude Code is broken".API and gateway
- Restart-Claude-Code-Gateway.command The one to double-click after editing config.yaml.SDKs and frameworksOperations and deployment
- Stop-Claude-Code-Gateway.command Stops the LiteLLM proxy. Handles both the manual launcher and the launchd job — if the launchd job is loaded, killing the PID alone just makes launchd start it again, which looks like the stop failed.API and gatewayOperations and deployment
- Start-Claude-Code-Gateway.command Starts the LiteLLM Anthropic-format proxy that Claude Code talks to. Double-clickable from Finder, or run from a terminal. This is the MANUAL launcher.API and gateway
- document -> audio, English or Chinese. Double-click it, and it asks for a file or URL, then a voice.Operations and deploymentKnowledge management
- document -> verified audiobook (English + Chinese) $AI_DATA/apps/agent-server/venv/bin/python3 make_audiobook.py FILE All synthesis, normalization and quality checking lives in tts_engine.py.Inference engineering
- production text-to-speech for long documents (EN + ZH) ===================================================================== Replaces the naive "post a big chunk, hope it comes back" approach with the loop commercial audiobook pipelines actually use: normalize -> small units -> ...Inference engineeringKnowledge management
- verify the gateway without ever typing or pasting a key. WHY THIS EXISTS Every 401 on 2026-08-20 came from the same typo: TWO spaces after "Bearer".API and gateway
- double-click button, re-runnable Puts the DAILY CYCLE under launchd: com.example.collect-daily 03:15 every day daily-cycle.sh daily com.example.collect-weekly Sun 05:10 daily-cycle.sh full The cycle is collect -> retune plan from log -> INGEST -> audit learning chain ...Knowledge management
- double-click button Acquires authoritative Federal budget / financial-management sources into knowledge-bank/FINANCE-Knowledge-Bank/. Everything about WHERE material comes from lives in _collector/sources.json — adding a source is a JSON edit.Knowledge management
- Check-Public-Path.command Traces the FULL path an outside app takes to reach this Mac, hop by hop, and names the first hop that is broken. WHY A SEPARATE BUTTON. Check-All-LLM-Status.command answers "is the service healthy on this Mac".Operations and deployment
- Ollama ONLY, in the FOREGROUND, for debugging. start-ai-server.command — Ollama alone, foreground, watch the output start-ai-stack.command — the whole stack, via launchd The names differ by one word and both start Ollama.Inference engineeringOperations and deployment
- Enable-Unattended-Recovery.command Makes this Mac come back from a power cut, a crash, or a macOS update reboot entirely on its own — no password typed, no button pressed.State and storageOperations and deployment
- Check-Boot-Recovery.command ONE QUESTION: if the power died right now, would this Mac come back serving, entirely by itself? Built 2026-08-15. Every other piece of supervision on this Mac assumes a logged-in session exists.Operations and deploymentState and storage
- toggle the health watchdog off/on. WHY THIS EXISTS. Restart-Agent-Server.command deliberately never kills a backend you started by hand in a Terminal you are watching for tracebacks.Operations and deployment
- task-type detection + model selection for chat-family requests (default / fast / vision). Added 2026-08-03 on top of the existing Step 5.7 agent-server, per AI_Model_Integration_Guide.md.Inference engineeringAPI and gateway
- U3, U5 and U6: the layers that turn a number into a finding. The four tools in spending.py answer "how much". These answer the three questions that have to follow, and that the agent has so far had to guess at: U3 lookup_code what is this thing CALLED in the data?State and storageEvaluation
- the point where learned knowledge finally reaches a live answer. Until this file existed, the learning loop was an accumulator: it read documents, extracted a vocabulary, validated it, and stored it somewhere the question-answering path could not see.RAG and knowledge graphSelf-evolving agents
- dump the RAW response shape from a model backend. Exists because diet_kb.py's first real run failed with KeyError: <redacted> five times in a row, ~8s per call.Self-evolving agentsInference engineering
- /v1/audio/speech (TTS) for build-host agent-server (FastAPI). Added 2026-08-03, scaffolded — see the "NOT WIRED" status in model_catalog.json's _not_ollama.speech_tts entry.Inference engineering
- request-time authentication/authorization for the agent server. Runs BEFORE the request body is ever touched (same ordering discipline as step 1's main.py) so a bad key always gets a clean, fast rejection regardless of what — if anything — valid the body contains.API and gateway
- manage commercial-grade API keys for the agent server. Replaces apps/build-host-site/local-server/manage_keys.py for THIS server (that script still administers the old plaintext api_keys.json used by the legacy gateway.py — untouched, separate system, on purpose, until Phase ...API and gateway
- commercial-grade key storage for the agent server. Replaces the old plaintext apps/build-host-site/local-server/api_keys.json approach (still used by the legacy gateway.py, untouched, not this file's concern).API and gateway
- EDITED 2026-08-07 — the OLLAMA_* block that used to be duplicated here now lives in ollama-env.sh, the single source of truth.Inference engineering
- EDITED 2026-08-07 — this file used to be four lines, the last of which was: nohup $AI_DATA/start-ai-stack.command >/dev/null 2>&1 & That redirect discarded every message the start script produced, including fatal ones.Inference engineeringOperations and deployment
- SINGLE SOURCE OF TRUTH for Ollama's environment Created 2026-08-07 after the outage described below.Inference engineering
- Standalone speech-to-text server for build-host agent-server's new /v1/audio/transcriptions route (added to speech_route.py 2026-08-03).Inference engineering
- Standalone image-generation server for build-host agent-server's media_route.py. Added 2026-08-03.Operations and deploymentInference engineering
- /v1/images/generations for build-host agent-server (FastAPI). Added 2026-08-03, scaffolded — see the "NOT WIRED" status in model_catalog.json's _not_ollama.media_image_video entry.API and gateway
- /v1/rerank for build-host agent-server (FastAPI). Added 2026-08-03, scaffolded — see the "NOT WIRED" status in model_catalog.json's _not_ollama.reranker entry.RAG and knowledge graph
- /v1/embeddings for build-host agent-server (FastAPI). Adds the one route brainbank's knowledge engine needs. Everything else on agent-server is untouched.RAG and knowledge graphAPI and gateway
- Step 4 (AGENT_SERVER_GUIDE_v2.md §4): connects MCP (Model Context Protocol) servers configured in mcp_servers.json and registers every tool they expose into tools/registry.py — the exact same registry search_knowledge_base lives in.Tools and MCP
- Step 7 (AGENT_SERVER_GUIDE_v2.md §4 Step 7). Scope, stated honestly up front: this is the ONE piece of Step 7 that's genuinely buildable and testable without the real Mac's hardware — turning this project's own already-graded eval data into an MLX-LM-shaped training file.Inference engineeringSelf-evolving agents
- sends every question in knowledge-bank/test-questions.md through the REAL agent server (with tools enabled) and saves each answer + what the tool actually retrieved, so you can grade pass/fail by hand.Evaluation
- Step 5.5 (AGENT_SERVER_GUIDE_v2.md §7.3a). Reads the audit-log family this project already writes on every request (logs/agents/requests.jsonl, tool_calls.jsonl, feedback.jsonl) plus every graded knowledge-bank/test-questions-results-*.md run, and renders a single static local ...Knowledge managementEvaluation
- Step 3.6: the get_study_plan tool. This is the concrete deliverable behind the MUST-have requirement (2026-07-28): a customized K12 study plan grounded in this specific student's actual observed knowledge gaps, not a generic one the model would otherwise have to guess at.MemorySelf-evolving agents
- Step 5 (AGENT_SERVER_GUIDE_v2.md §4): a small, pluggable pre/post tool-call event registry. A pre-tool hook is `fn(tool_name: str, arguments: dict) -> dict | None`.Guardrails and verification
- hooks/ — Step 5 (AGENT_SERVER_GUIDE_v2.md §4): a small, pluggable pre/post tool-call event registry, plus the two first real hooks.Guardrails and verification
- Step 5's second real hook: after a filesystem write/edit/move tool call actually SUCCEEDS against a path under one of the two knowledge banks, automatically triggers kb_sync.py for that bank — the same reconciliation Sync-Knowledge-Base.command already runs by hand (§4 Step ...Guardrails and verification
- Step 5's first real hook: blocks any tool call that looks like a filesystem write/edit/create/move operation whose target path(s) fall outside an allow-listed set of directories.Tools and MCPGuardrails and verification
- one-shot diagnostic for the 5 "budget-exhausted" FAILs in the 2026-07-28 test-questions baseline (test-questions-results- 20260728-165833.md). All 5 hit the 6-round tool-call budget without ever finding a good passage.Knowledge managementRAG and knowledge graph
- Double-click button: scans both knowledge banks for changes — ingests new or edited files, and removes index entries for anything deleted or moved into an _archive/ subfolder.Knowledge management
- minimal state-graph execution engine. Why this exists (v2 refinement, 2026-07-27): step 2's agent_loop.py was a flat sequential loop — call the model, run whatever tools it asked for in one batch, repeat.Harness engineering
- Tool registry package. Each tool module exports SCHEMA (an OpenAI function-calling schema dict) and a callable matching that schema's parameters.State and storageTools and MCP
- title: Knowledge Scope Filter author: build-host setup version: 1.0.0 description: > Anti-leakage + noise-stripping inlet filter for Open WebUI RAG.Knowledge management