Evaluation
Sealed test sets, layered metrics, and enough repeats to trust a difference.
What it is and why it exists
What
Evaluation is the set of fixed questions, expected results, and scoring scripts that tell you whether a change made the system better. This module builds a sealed baseline, a retrieval gold set, routing and chat evals, a claim-fidelity check, and a history file that every run appends to.
Why
Agent systems change daily and fail quietly. Without numbers, each change is judged by the last conversation someone remembers. With a learning loop the risk is higher: a loop that can influence its own test will raise its score without learning anything.
How it works
- Each layer gets its own eval so a failure can be located. Retrieval is scored without the chat model. Skill routing is scored without retrieval. The full agent is scored end to end.
- Test material is sealed before any learning job runs. Held-out chunks are excluded from study. Human-written questions are copied to a read-only folder.
- Each run appends one line to a history file with the configuration that produced it. Variants are selected by environment flags, so the same script compares old and new behavior.
- Results are read with their sample size. A difference smaller than the run-to-run spread is reported as no difference.
Where it sits in the build order
Needs first
- RAG and knowledge graphThe retrieval eval calls the search tool and scores what it returns. Gold questions are generated from indexed chunks, so the index has to exist.Build out of order Stub it with: Score a hard-coded list of results. You can build and test the metric code with no search at all.
- Harness engineeringEnd-to-end evals send requests through the agent and read its logs. The log formats are defined by the harness.Build out of order Stub it with: Evaluate a bare model endpoint. The scripts are the same and the target URL differs.
- Knowledge managementHeld-out material is selected from the corpus. You cannot seal chunks that have not been ingested.Build out of order Stub it with: Twenty hand-written questions with expected answers.
Unlocks
- Multi-agent patternsEvery added agent adds cost and latency. Only an eval can show that a split improved results.
- Operations and deploymentScheduled evals and the daily report are how you learn that a deploy made things worse.
- Self-evolving agentsA loop optimizes what it can measure. Sealed evals have to exist before the loop runs, or it grades itself.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/nightly/seal_evals.py | Creates the frozen question set and the held-out chunk manifest. |
| apps/agent-server/nightly/retrieval_eval.py | Builds the retrieval gold set and scores hit@k, near, doc, and MRR. |
| apps/agent-server/nightly/skill_eval.py | Scores skill routing against a gold set kept separate from routing exemplars. |
| apps/agent-server/nightly/chat_eval.py | Multi-turn conversations through the full agent. |
| apps/agent-server/nightly/stream_eval.py | The same questions through the streaming path, repeated. |
| apps/agent-server/nightly/claim_fidelity.py | Samples served claims and checks each against its cited passage. |
| apps/agent-server/nightly/quiz_runner.py | Baseline versus with-knowledge arms on one question set. |
| evals/ | Gold sets, sealed material, run outputs, and history files. |
Process map
From a proposed change to a decision
- 1SealDone once, before learning. Frozen questions and held-out chunks become read-only. (storage)
- 2Baseline runScore the current system. Repeat enough to know the spread. (decision)
- 3Change behind a flagNew behavior is selectable by an environment variable. (input)
- 4Variant runSame script, same gold set, flag on. (decision)
- 5Compare with noiseIs the difference larger than the spread between repeats? (decision)
- 6Append historyOne JSON line per run with metrics and configuration. (storage)
- 7DecideMake it the default, keep it off, or collect more data. (output)
Build steps
Each step states why it sits at this point. Open any step on its own. The first is open.
01Seal the test material before anything learns
Copy every human-written question and every historical result into a frozen folder. Results captured before any learning existed are your only true pre-learning baseline.
Sample chunks from the index into a held-out manifest. The learner must treat the manifest as an exclusion list.
Make the files read-only. The defense is structural. A prompt asking the loop not to look is not a defense.
python3 nightly/seal_evals.py
chmod -R a-w evals/frozen
ls -l evals/frozen # every file -r--r--r--02Run a human-written baseline
Write ten to twenty real questions per collection with the answer and the source you expect.
Run them and grade by hand once. On the reference build the first run was 13 pass and 7 fail out of 20.
Add a question each time the system fumbles a real request.
03Build a retrieval gold set
Sample chunks: held-out first, a few per document at most, primary sources only, prose of moderate length.
Have a small model write one question that the chunk answers, in its own words. Drop any question that copies a six-word run from the chunk, since that tests string matching.
Keep 100 and make the file read-only. Build separate sets for question types you care about, such as citation and definition questions.
def copies_passage(question: str, passage: str, n: int = 6) -> bool:
q = question.lower().split()
p = " ".join(passage.lower().split())
return any(" ".join(q[i:i + n]) in p for i in range(len(q) - n + 1))04Score retrieval with rank metrics
For each question take the top ten results and compute the metrics below. Average across the set.
Report several. Exact-chunk hit rate is strict. Document-level hit rate shows whether you are in the right place and chunked badly.
def score(gold_id: str, results: list[str]) -> dict:
doc = gold_id.rsplit("::", 1)[0]
idx = int(gold_id.rsplit("::", 1)[1])
near = {f"{doc}::{idx + d}" for d in (-1, 0, 1)}
rank = results.index(gold_id) + 1 if gold_id in results else 0
return {
"hit@1": int(rank == 1),
"hit@5": int(0 < rank <= 5),
"hit@10": int(0 < rank <= 10),
"near@10": int(any(r in near for r in results[:10])),
"doc@10": int(any(r.startswith(doc + "::") for r in results[:10])),
"mrr": 1 / rank if rank else 0.0,
}05Compare variants with flags and a history file
Give each behavior change an environment flag. Run the eval with the flag off and on.
Append one line per run with timestamp, variant name, sample size, metrics, and the flags in effect.
Keep a change only if the metric moves. The reranker, dedup, and definition edges on the reference build were each decided this way.
python3 nightly/retrieval_eval.py --run
python3 nightly/retrieval_eval.py --run --variant rerank
tail -n 2 evals/retrieval/history.jsonl06Evaluate routing with examples the router never saw
Keep two files: exemplars the router matches against, and a gold set used only for scoring.
When a real request is routed wrong, add a line to the exemplars. Never copy gold items into exemplars.
07Evaluate the full agent, more than once
Send multi-turn conversations through the real endpoint, both streaming and non-streaming. The two paths have different code and have failed differently.
Run each question several times. One failure in six runs is a real defect that a single run usually misses.
Check mechanical properties with code: was a data tool called, do the figures match the tool output. Use a model judge only for what code cannot check.
08Use model judges with two votes and calibration
For claim fidelity, sample served claims and ask a stronger model whether the cited passage supports each one.
Use two independent votes. Remove a claim from service only when it fails both. An omission alone keeps the claim.
Review a sample of judge decisions by hand. On the reference build, judges rejected correct claims because an organization had two official names. The judge prompt was corrected and the rejected claims were rechecked.
09Read results with their sample size
The reference build once measured 93.3 percent with learned knowledge against 70.0 percent without, on 30 items, in one run.
A 75-item set then ran seven times in both arms. The means were 61.9 and 61.6 out of 75. The difference was 0.3 items, inside run-to-run noise, and negative twice.
Both results are recorded. Neither is quoted as the system's accuracy until the two instruments are reconciled.
10Schedule the evals and print one line each in the daily report
Run retrieval and routing evals nightly, claim fidelity weekly. Pause them while a user request is in flight.
Mine the logs for mechanically checkable failures. An empty result that returns rows when one filter is relaxed is a labeled example produced with no human.
How it works with the other parts
- Self-evolving agentsThe learner is allowed to change knowledge. Sealed evals are what stop it from grading itself.
- Guardrails and verificationEach guard rule started as an eval failure, and each has a regression test.
- Operations and deploymentEvals run from the scheduler, yield to live traffic, and write into the daily report.
- RAG and knowledge graphRanking changes ship behind flags and are kept or dropped on the retrieval history.
What went wrong in the real build
Two instruments disagreed about the main result
- What happened
- A 30-item single run showed a 23-point gain from learned knowledge. A 75-item set over seven runs showed 0.3 items.
- Fix
- Stop quoting either number. Record both with sample sizes and investigate the difference.
- Lesson
- A single run on a small set is an anecdote with a decimal point.
A failure that appeared once in six runs
- What happened
- A spending question was answered from a web search one time in six, with the data tools never called. Every figure matched the web result, so the guard passed.
- Fix
- Repeat stream evals. Add a wrong-source rule for data turns.
- Lesson
- Intermittent failures need repeats to see and a rule about sources to stop.
Judges rejected correct claims
- What happened
- Claim judges treated two names for one organization as different entities and voted against true claims.
- Fix
- State the equivalence in the judge prompt and recheck every claim rejected for that reason.
- Lesson
- Audit the judge. Its errors are systematic and silent.
No number moved when search broke
- What happened
- The only eval measured answers from a small model. A chunking or ranking change could degrade search with no visible signal.
- Fix
- A dedicated retrieval eval with a sealed gold set.
- Lesson
- Give each layer its own measurement.
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | MLflow evaluation runs scorers and judges over an evaluation dataset stored in Unity Catalog, and traces link each score to the run that produced it. Your gold sets become tables. Sealing becomes table permissions. The metric code in this module can be registered as custom scorers. |
| IBM watsonx | watsonx.governance evaluates and monitors deployed models and agents, and the Orchestrate ADK includes an evaluation framework for agent trajectories. Keep your own gold sets and feed them in. The sample-size discipline is unchanged. |
| Codex | Use the provider's evals tooling or plain scripts. Codex can run your eval scripts headless as part of a change, which makes eval a gate in the coding loop. |
| Cursor | Evals are scripts in your repo. A rule or hook can require the retrieval eval to pass before a change to ranking code is accepted. |
| Claude Code / Agent SDK | Run eval scripts through the coding agent, or score plugin and skill behavior with its eval commands. Keep gold files out of the agent's writable paths. |
| Another machine | The eval scripts are Python and JSON files. They move with the repo. Rebuild the gold set only if the corpus changed. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why must sealing happen before the learning loop runs, and why is a prompt not enough?
Your end-to-end score dropped. In what order do you check, and why?
A change improves a 30-item set by 7 points in one run. What do you say?
Why keep routing exemplars and routing gold in separate files?
What could you evaluate if you had only a model endpoint and no agent?
Terms
| Held-out | Material excluded from learning so it can test generalization. |
| MRR | Mean reciprocal rank. The average of 1/rank of the correct item, zero if absent. |
| Variant | A named configuration of flags evaluated against the same gold set. |
| Calibration | Measuring how often a judge agrees with trusted labels. |
From the live build
Recent changes and files the sync job filed under this module.
- Search returns one copy of a paragraph that is in several files; same@10 and distinct@10 in the retrieval eval; dated scheduler log; notice when a no-query draft is replaced; skill: a status is half an answer
- Knowledge graph: 'defines' edges from set definitional forms, used in retrieval for definitional questions; gold set and nightly eval
- Chat eval: the round-limit message counts as a fallback, not a bad answer delivered
- Answer guard checks counts; chat eval in the Sunday cycle; CHAT line in the daily report
- Answer guard: plain message when the model returns nothing twice; chat eval allows 6000 output tokens
- Recall similarity floor 0.35 (eval: lowest hit 0.55, highest unrelated 0.24)
- Weekly agent quiz in the daily cycle; AGENT line in the report
- Citation layer: kb_finance.citations, citation gold set, graph expansion (off)
- FINANCE: name the actual authority, don't just summarize — and show it in practice, not just recite the rule
- Issues 7, 8 and 9 — diagnosis and fix
- BUILD-HOST — user guide
- guards on the retrieval changes of 2026-09-30. Run: venv/bin/python3 test_retrieval.py (no database, no network) 1. SQL v2 renders for nomic (768) and Qwen3 (1024) with no stray fields, and passes the query vector as a parameter (the 5.4 s bug was a CTE) 2.
- does search_knowledge_base find the right passage? (2026-09-30) The weekly quiz measures whether a 9B model answers better with learned atoms. Nothing measured retrieval itself, so a chunker, embedding or ranking change could make search worse and no number would move.
- the agent-facing wrapper around claim_recall. REGISTERED 2026-08-18, AFTER IT EARNED IT. ------------------------------------------ This file sat unregistered until the T5 discrimination drill said it was worth registering.
- the "defines" edge of the knowledge graph. venv/bin/python3 nightly/defines.py --index # scan new chunks (incremental) venv/bin/python3 nightly/defines.py --stats venv/bin/python3 nightly/defines.py --build-gold # once; sealed The citation layer links a passage to what it CITES.
- what the chat box actually does (2026-10-06). chat_eval.py and agent_quiz.py call agent_loop.run(), the non-streaming path. The chat box streams.
- how does the agent answer real multi-turn questions? (2026-10-02) The quiz and the gold sets are single questions with one right letter or one right passage. They never measured the thing a person does in the chat box: ask something, then follow up in five words.
- does conversation recall find the right exchange, stay quiet when nothing fits, and keep users apart? (2026-10-01) Seeds two synthetic users (partitions eval:recall-a / eval:recall-b) straight into mem.turns, asks paraphrased questions, removes the seed rows.
- the discrimination quiz through the REAL agent path (2026-09-30). The audit's top evaluation gap: quiz_runner.py measures the 9B model with and without atoms, never what a user gets.