Agent Build Tutor
Map / Outline

Guardrails and verification

Deterministic checks on tool calls and on answers, each one written from a real failure.

What it is and why it exists

What

Guardrails are code that inspects what the agent is about to do or say. Hooks check tool calls. An answer guard checks the final draft against the evidence gathered in this request.

Why

Models produce fluent text whether or not it is supported. The reference agent once ran the right queries, got the right rows, and answered with an invented table. Only a check outside the model catches that.

How it works

Where it sits in the build order

Needs first

  • Harness engineeringThe guard is a stage in the loop: after the draft, before the response. Hooks wrap tool execution. Both need the loop.Build out of order Stub it with: Run the guard offline over logged answers.
  • Tools and MCPThe guard compares the answer with tool results. Tools must return structured values the guard can parse.Build out of order Stub it with: A fixed evidence string.

Unlocks

  • Self-evolving agentsPromotion gates and claim verification are guard logic applied to stored knowledge.

In the reference build

PathRole
apps/agent-server/answer_guard.pyFigure, count, and name checks. Corrective and fallback messages. Stream gate.
apps/agent-server/test_answer_guard.pyRegression tests, one per rule.
apps/agent-server/hooks/write_path_guard.pyBlocks writes outside an allowed folder.
apps/agent-server/nightly/claim_verify.pyTwo-vote check of stored claims against their source.

The same idea on other platforms

PlatformHow this module maps
DatabricksAI Gateway guardrails cover safety and PII at the endpoint. Groundedness checks like this one are custom code in your agent, or scorers run on traces.
IBM watsonxwatsonx.governance provides guardrails and monitors. Domain checks go in a tool or a post-processing step you write.
CodexSandbox and approval modes limit actions. Hooks and your own test commands act as checks on changes.
CursorHooks can block or modify agent actions. Rules can require checks before completion.
Claude Code / Agent SDKHooks run before and after tool use and can block. Permission modes limit actions. Output checks are code you add in hooks or in the SDK loop.
Another machineThe guard is a pure Python module that takes an answer and evidence. Portable as is.

Explain it back

Answer aloud first. Then open the answer and compare.

Why judge the first draft strictly and the retry by majority?
A strong answerA strict first pass catches a single wrong figure among right ones. A lenient retry keeps one stubborn figure from turning a usable answer into a refusal.
Why does every guard rule need an observe mode?
A strong answerSo you can measure how often it would fire on real traffic, and what it would wrongly block, before it changes any answer.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module