Agent Build Tutor
Map / Full lesson

Evaluation

Sealed test sets, layered metrics, and enough repeats to trust a difference.

What it is and why it exists

What

Evaluation is the set of fixed questions, expected results, and scoring scripts that tell you whether a change made the system better. This module builds a sealed baseline, a retrieval gold set, routing and chat evals, a claim-fidelity check, and a history file that every run appends to.

Why

Agent systems change daily and fail quietly. Without numbers, each change is judged by the last conversation someone remembers. With a learning loop the risk is higher: a loop that can influence its own test will raise its score without learning anything.

How it works

Where it sits in the build order

Needs first

  • RAG and knowledge graphThe retrieval eval calls the search tool and scores what it returns. Gold questions are generated from indexed chunks, so the index has to exist.Build out of order Stub it with: Score a hard-coded list of results. You can build and test the metric code with no search at all.
  • Harness engineeringEnd-to-end evals send requests through the agent and read its logs. The log formats are defined by the harness.Build out of order Stub it with: Evaluate a bare model endpoint. The scripts are the same and the target URL differs.
  • Knowledge managementHeld-out material is selected from the corpus. You cannot seal chunks that have not been ingested.Build out of order Stub it with: Twenty hand-written questions with expected answers.

Unlocks

  • Multi-agent patternsEvery added agent adds cost and latency. Only an eval can show that a split improved results.
  • Operations and deploymentScheduled evals and the daily report are how you learn that a deploy made things worse.
  • Self-evolving agentsA loop optimizes what it can measure. Sealed evals have to exist before the loop runs, or it grades itself.

In the reference build

PathRole
apps/agent-server/nightly/seal_evals.pyCreates the frozen question set and the held-out chunk manifest.
apps/agent-server/nightly/retrieval_eval.pyBuilds the retrieval gold set and scores hit@k, near, doc, and MRR.
apps/agent-server/nightly/skill_eval.pyScores skill routing against a gold set kept separate from routing exemplars.
apps/agent-server/nightly/chat_eval.pyMulti-turn conversations through the full agent.
apps/agent-server/nightly/stream_eval.pyThe same questions through the streaming path, repeated.
apps/agent-server/nightly/claim_fidelity.pySamples served claims and checks each against its cited passage.
apps/agent-server/nightly/quiz_runner.pyBaseline versus with-knowledge arms on one question set.
evals/Gold sets, sealed material, run outputs, and history files.

Process map

From a proposed change to a decision

  1. 1
    SealDone once, before learning. Frozen questions and held-out chunks become read-only. (storage)
  2. 2
    Baseline runScore the current system. Repeat enough to know the spread. (decision)
  3. 3
    Change behind a flagNew behavior is selectable by an environment variable. (input)
  4. 4
    Variant runSame script, same gold set, flag on. (decision)
  5. 5
    Compare with noiseIs the difference larger than the spread between repeats? (decision)
  6. 6
    Append historyOne JSON line per run with metrics and configuration. (storage)
  7. 7
    DecideMake it the default, keep it off, or collect more data. (output)

Build steps

Each step states why it sits at this point. Open any step on its own. The first is open.

01Seal the test material before anything learns
Why this step is hereThis is first because it cannot be done afterward. Once a learning job has read a chunk or seen a question, that item no longer measures generalization.

Copy every human-written question and every historical result into a frozen folder. Results captured before any learning existed are your only true pre-learning baseline.

Sample chunks from the index into a held-out manifest. The learner must treat the manifest as an exclusion list.

Make the files read-only. The defense is structural. A prompt asking the loop not to look is not a defense.

bash
python3 nightly/seal_evals.py
chmod -R a-w evals/frozen
ls -l evals/frozen        # every file -r--r--r--
You are done whenA write to the frozen folder fails, and the learner's queue contains no held-out chunk id.
02Run a human-written baseline
Why this step is hereYou need one number that predates your own tooling. It anchors everything measured later.

Write ten to twenty real questions per collection with the answer and the source you expect.

Run them and grade by hand once. On the reference build the first run was 13 pass and 7 fail out of 20.

Add a question each time the system fumbles a real request.

You are done whenYou have a dated results file and can name which questions fail.
03Build a retrieval gold set
Why this step is hereRetrieval is scored before end-to-end quality because it is deterministic, fast, and the most common root cause. It also needs no judge.

Sample chunks: held-out first, a few per document at most, primary sources only, prose of moderate length.

Have a small model write one question that the chunk answers, in its own words. Drop any question that copies a six-word run from the chunk, since that tests string matching.

Keep 100 and make the file read-only. Build separate sets for question types you care about, such as citation and definition questions.

python
def copies_passage(question: str, passage: str, n: int = 6) -> bool:
    q = question.lower().split()
    p = " ".join(passage.lower().split())
    return any(" ".join(q[i:i + n]) in p for i in range(len(q) - n + 1))
You are done whenThe gold file has 100 items, each with a question and the exact chunk id that answers it.
04Score retrieval with rank metrics
Why this step is hereMetric code follows the gold set because the metric definitions depend on what the gold item identifies: a chunk, its neighbors, and its document.

For each question take the top ten results and compute the metrics below. Average across the set.

Report several. Exact-chunk hit rate is strict. Document-level hit rate shows whether you are in the right place and chunked badly.

python
def score(gold_id: str, results: list[str]) -> dict:
    doc = gold_id.rsplit("::", 1)[0]
    idx = int(gold_id.rsplit("::", 1)[1])
    near = {f"{doc}::{idx + d}" for d in (-1, 0, 1)}
    rank = results.index(gold_id) + 1 if gold_id in results else 0
    return {
        "hit@1":  int(rank == 1),
        "hit@5":  int(0 < rank <= 5),
        "hit@10": int(0 < rank <= 10),
        "near@10": int(any(r in near for r in results[:10])),
        "doc@10":  int(any(r.startswith(doc + "::") for r in results[:10])),
        "mrr": 1 / rank if rank else 0.0,
    }
You are done whenA recent reference run on a citation gold set of 60: hit@1 0.52, hit@5 0.78, hit@10 0.85, doc@10 0.93, MRR 0.64, median 1.9 seconds.
05Compare variants with flags and a history file
Why this step is hereThis needs a stable baseline metric to compare against. It is the mechanism that turns every later change into an experiment.

Give each behavior change an environment flag. Run the eval with the flag off and on.

Append one line per run with timestamp, variant name, sample size, metrics, and the flags in effect.

Keep a change only if the metric moves. The reranker, dedup, and definition edges on the reference build were each decided this way.

bash
python3 nightly/retrieval_eval.py --run
python3 nightly/retrieval_eval.py --run --variant rerank
tail -n 2 evals/retrieval/history.jsonl
You are done whenTwo adjacent history lines differ only in variant and metrics.
06Evaluate routing with examples the router never saw
Why this step is hereRouting is a separate layer between request and retrieval. Scoring it alone tells you whether a bad answer started with the wrong skill.

Keep two files: exemplars the router matches against, and a gold set used only for scoring.

When a real request is routed wrong, add a line to the exemplars. Never copy gold items into exemplars.

You are done whenRouting accuracy is reported on the gold set, with a list of the misrouted requests.
07Evaluate the full agent, more than once
Why this step is hereEnd-to-end evals come after the layer evals so that when the full score drops you already know which layer to check.

Send multi-turn conversations through the real endpoint, both streaming and non-streaming. The two paths have different code and have failed differently.

Run each question several times. One failure in six runs is a real defect that a single run usually misses.

Check mechanical properties with code: was a data tool called, do the figures match the tool output. Use a model judge only for what code cannot check.

You are done whenEach question has a pass rate across repeats, and failures link to request ids in the tool log.
08Use model judges with two votes and calibration
Why this step is hereJudges are needed for free-text fidelity, and they are also models with their own errors. They come late because you calibrate them against the deterministic checks you already have.

For claim fidelity, sample served claims and ask a stronger model whether the cited passage supports each one.

Use two independent votes. Remove a claim from service only when it fails both. An omission alone keeps the claim.

Review a sample of judge decisions by hand. On the reference build, judges rejected correct claims because an organization had two official names. The judge prompt was corrected and the rejected claims were rechecked.

You are done whenYou can state the judge's agreement rate with your own labels on a sample of at least 50.
09Read results with their sample size
Why this step is hereThis is the discipline that makes the previous steps worth doing. It comes last because you need several instruments before you can see them disagree.

The reference build once measured 93.3 percent with learned knowledge against 70.0 percent without, on 30 items, in one run.

A 75-item set then ran seven times in both arms. The means were 61.9 and 61.6 out of 75. The difference was 0.3 items, inside run-to-run noise, and negative twice.

Both results are recorded. Neither is quoted as the system's accuracy until the two instruments are reconciled.

You are done whenEvery reported improvement names the set, its size, the number of runs, and the spread.
10Schedule the evals and print one line each in the daily report
Why this step is hereAutomation comes after the evals are trusted. Scheduling an eval you do not believe produces a dashboard nobody reads.

Run retrieval and routing evals nightly, claim fidelity weekly. Pause them while a user request is in flight.

Mine the logs for mechanically checkable failures. An empty result that returns rows when one filter is relaxed is a labeled example produced with no human.

You are done whenThe daily report shows each eval's latest value beside its previous value.

How it works with the other parts

What went wrong in the real build

2026-09

Two instruments disagreed about the main result

What happened
A 30-item single run showed a 23-point gain from learned knowledge. A 75-item set over seven runs showed 0.3 items.
Fix
Stop quoting either number. Record both with sample sizes and investigate the difference.
Lesson
A single run on a small set is an anecdote with a decimal point.
2026-10

A failure that appeared once in six runs

What happened
A spending question was answered from a web search one time in six, with the data tools never called. Every figure matched the web result, so the guard passed.
Fix
Repeat stream evals. Add a wrong-source rule for data turns.
Lesson
Intermittent failures need repeats to see and a rule about sources to stop.
2026-10

Judges rejected correct claims

What happened
Claim judges treated two names for one organization as different entities and voted against true claims.
Fix
State the equivalence in the judge prompt and recheck every claim rejected for that reason.
Lesson
Audit the judge. Its errors are systematic and silent.
2026-08

No number moved when search broke

What happened
The only eval measured answers from a small model. A chunking or ranking change could degrade search with no visible signal.
Fix
A dedicated retrieval eval with a sealed gold set.
Lesson
Give each layer its own measurement.

The same idea on other platforms

PlatformHow this module maps
DatabricksMLflow evaluation runs scorers and judges over an evaluation dataset stored in Unity Catalog, and traces link each score to the run that produced it. Your gold sets become tables. Sealing becomes table permissions. The metric code in this module can be registered as custom scorers.
IBM watsonxwatsonx.governance evaluates and monitors deployed models and agents, and the Orchestrate ADK includes an evaluation framework for agent trajectories. Keep your own gold sets and feed them in. The sample-size discipline is unchanged.
CodexUse the provider's evals tooling or plain scripts. Codex can run your eval scripts headless as part of a change, which makes eval a gate in the coding loop.
CursorEvals are scripts in your repo. A rule or hook can require the retrieval eval to pass before a change to ranking code is accepted.
Claude Code / Agent SDKRun eval scripts through the coding agent, or score plugin and skill behavior with its eval commands. Keep gold files out of the agent's writable paths.
Another machineThe eval scripts are Python and JSON files. They move with the repo. Rebuild the gold set only if the corpus changed.

Explain it back

Answer aloud first. Then open the answer and compare.

Why must sealing happen before the learning loop runs, and why is a prompt not enough?
A strong answerA loop optimizes whatever is measurable. If it can select or phrase its own test, the score climbs while nothing is learned. An instruction can be ignored. A file the loop cannot read or write cannot be.
Your end-to-end score dropped. In what order do you check, and why?
A strong answerRetrieval eval, then routing eval, then the guard logs, then the model. That order goes from deterministic and cheap to stochastic and expensive, and it follows the direction data flows.
A change improves a 30-item set by 7 points in one run. What do you say?
A strong answerThat it is not yet known. Seven points on 30 items is about two questions. Run it several times, report the spread, and use a larger set if the spread is that wide.
Why keep routing exemplars and routing gold in separate files?
A strong answerThe router matches requests against exemplars. If the gold items are in that file, the eval measures lookup of known strings and stops measuring routing.
What could you evaluate if you had only a model endpoint and no agent?
A strong answerEverything in the scoring scripts. Point them at the endpoint. You would learn the unaided baseline, which is the number the agent has to beat.

Terms

Held-outMaterial excluded from learning so it can test generalization.
MRRMean reciprocal rank. The average of 1/rank of the correct item, zero if absent.
VariantA named configuration of flags evaluated against the same gold set.
CalibrationMeasuring how often a judge agrees with trusted labels.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module