Map / Outline
Guardrails and verification
Deterministic checks on tool calls and on answers, each one written from a real failure.
What it is and why it exists
What
Guardrails are code that inspects what the agent is about to do or say. Hooks check tool calls. An answer guard checks the final draft against the evidence gathered in this request.
Why
Models produce fluent text whether or not it is supported. The reference agent once ran the right queries, got the right rows, and answered with an invented table. Only a check outside the model catches that.
How it works
- Extract checkable items from the draft: dollar figures, counts, organization names.
- Collect evidence: tool results from this request, plus what the user and earlier answers already said.
- An item is grounded when it appears in the evidence at some unit scale, or is the sum or difference of two grounded items.
- The first draft is judged strictly. A failed draft gets one corrective message. The retry is judged by majority. If it still fails, a fallback states what was found.
- Each rule has an enforce, observe, or off setting, and a test for the legitimate cases it must allow.
Where it sits in the build order
Needs first
- Harness engineeringThe guard is a stage in the loop: after the draft, before the response. Hooks wrap tool execution. Both need the loop.Build out of order Stub it with: Run the guard offline over logged answers.
- Tools and MCPThe guard compares the answer with tool results. Tools must return structured values the guard can parse.Build out of order Stub it with: A fixed evidence string.
Unlocks
- Self-evolving agentsPromotion gates and claim verification are guard logic applied to stored knowledge.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/answer_guard.py | Figure, count, and name checks. Corrective and fallback messages. Stream gate. |
| apps/agent-server/test_answer_guard.py | Regression tests, one per rule. |
| apps/agent-server/hooks/write_path_guard.py | Blocks writes outside an allowed folder. |
| apps/agent-server/nightly/claim_verify.py | Two-vote check of stored claims against their source. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | AI Gateway guardrails cover safety and PII at the endpoint. Groundedness checks like this one are custom code in your agent, or scorers run on traces. |
| IBM watsonx | watsonx.governance provides guardrails and monitors. Domain checks go in a tool or a post-processing step you write. |
| Codex | Sandbox and approval modes limit actions. Hooks and your own test commands act as checks on changes. |
| Cursor | Hooks can block or modify agent actions. Rules can require checks before completion. |
| Claude Code / Agent SDK | Hooks run before and after tool use and can block. Permission modes limit actions. Output checks are code you add in hooks or in the SDK loop. |
| Another machine | The guard is a pure Python module that takes an answer and evidence. Portable as is. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why judge the first draft strictly and the retry by majority?
A strong answerA strict first pass catches a single wrong figure among right ones. A lenient retry keeps one stubborn figure from turning a usable answer into a refusal.
Why does every guard rule need an observe mode?
A strong answerSo you can measure how often it would fire on real traffic, and what it would wrongly block, before it changes any answer.
From the live build
Recent changes and files the sync job filed under this module.
- claim_verify: a simplification alone keeps a claim in service; --redecide re-applies the rule to recorded votes
- claim_verify: two-vote check of served claims against their source (dry run and calibration); off-topic the oversight body reports retired from the learner queue
- Guard: on a spending turn a draft sourced from a web search alone is rejected and sent to the spending tools; skill says the same
- Guard: a requested made-up example is recorded, not rejected; wiki ingest runs 23:00-07:00
- Chat eval: the round-limit message counts as a fallback, not a bad answer delivered
- Answer guard checks counts; chat eval in the Sunday cycle; CHAT line in the daily report
- Answer guard: plain message when the model returns nothing twice; chat eval allows 6000 output tokens
- Skills: route by natural-language examples only (descriptions as fallback)
- What the failures teach about making the agent actually get smarter
- check served claims against their source, two votes, and take the ones that fail both out of service.
- a final answer may not state dollar figures the tools did not return (2026-10-02). Why: on 2026-10-02 the agent ran the right a commodity and a commodity queries, got the right rows back, and then answered with a table of an unrelated item purchases ("$18.96 Billion", "Nuccio shipyard").
- the 2026-10-02 an unrelated item answer must fail the guard and the correct a commodity answer must pass. Exits non-zero on failure.
- re-checks the 2026-09-07 fixes, on demand. $AI_DATA/apps/agent-server/venv/bin/python3 $AI_DATA/apps/agent-server/verify_fixes.py Exits 0 if every check passes, 1 otherwise, so it can go in a cycle step.
- Step 5 (AGENT_SERVER_GUIDE_v2.md §4): a small, pluggable pre/post tool-call event registry. A pre-tool hook is `fn(tool_name: str, arguments: dict) -> dict | None`.
- hooks/ — Step 5 (AGENT_SERVER_GUIDE_v2.md §4): a small, pluggable pre/post tool-call event registry, plus the two first real hooks.
- Step 5's second real hook: after a filesystem write/edit/move tool call actually SUCCEEDS against a path under one of the two knowledge banks, automatically triggers kb_sync.py for that bank — the same reconciliation Sync-Knowledge-Base.command already runs by hand (§4 Step ...
- Step 5's first real hook: blocks any tool call that looks like a filesystem write/edit/create/move operation whose target path(s) fall outside an allow-listed set of directories.