Harness engineering
The loop around the model: prompt assembly, tool execution, budgets, checks, logs.
What it is and why it exists
What
The harness is the program that runs the model. It assembles the prompt, offers tools, executes the tool calls the model asks for, feeds results back, decides when to stop, checks the answer, and records what happened.
Why
A model call is stateless and cannot act. Everything that makes an agent behave reliably lives in the harness: which tools exist, how many rounds are allowed, what gets truncated, what is checked before the user sees it. Two systems with the same model and different harnesses behave like different products.
How it works
- The harness exposes the same chat-completions contract it consumes. Callers cannot tell whether they are talking to a bare model or to the agent.
- A planner node calls the model. If the model answers, the run ends. If it requests tool calls, each call becomes a node and all of them run concurrently.
- Tool nodes append their results to the conversation and route back to the planner. Several tool nodes naming the same successor collapse into one planner run.
- A wave budget bounds the cycle. When the budget runs out, one final turn without tools answers from the evidence already gathered.
- Before the answer leaves, a deterministic guard compares it with the evidence from this request. A failed draft gets one corrective retry.
Where it sits in the build order
Needs first
- API and gatewayThe harness sits behind the gateway contract. Keys carry scopes, and the harness reads the scope to decide whether a caller gets tools at all.Build out of order Stub it with: A hard-coded flag: tools allowed, one anonymous caller.
- Inference engineeringThe planner is a model call. Round budgets, output caps, and prompt size limits are set from measured inference numbers.Build out of order Stub it with: A fake model function that returns a scripted tool call on the first turn and text on the second.
- State and storageThe loop is a state machine. Concurrent tool results have to merge into one conversation without overwriting each other, so the merge rules are defined before the loop that relies on them.Build out of order Stub it with: A plain dict and sequential tool execution.
- Tools and MCPA loop with no tools is a chat proxy. You need at least one registered tool with a schema to exercise the tool path.Build out of order Stub it with: One local function such as a clock or calculator.
Unlocks
- EvaluationEnd-to-end evals send requests through the agent and read its logs. The log formats are defined by the harness.
- Guardrails and verificationThe guard is a stage in the loop: after the draft, before the response. Hooks wrap tool execution. Both need the loop.
- MemoryMemory is read at prompt composition and written after the response. Both are harness stages.
- SkillsA skill is text injected during prompt composition. Without the composition step there is nowhere to load it.
- Multi-agent patternsA sub-agent is a harness loop run as a node. You need one working loop before you can run several.
- Operations and deploymentThe agent server is the main service. Its health endpoint and busy marker are what the supervisor and scheduler read.
- SDKs and frameworksYou evaluate an SDK by comparing it with a loop you understand. Without that, every framework's defaults look like requirements.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/main.py | HTTP surface: chat completions, models, feedback, health. |
| apps/agent-server/graph.py | A 110-line wave executor: fan-out, fan-in, merge rules, wave budget, trace. |
| apps/agent-server/agent_loop.py | The planner and tool graph, prompt composition, streaming, logging. |
| apps/agent-server/AGENT.md | The base system prompt. |
| apps/agent-server/tools/registry.py | One registry for hand-written and MCP tools. |
| apps/agent-server/hooks/ | Pre-tool and post-tool hooks. |
| apps/agent-server/answer_guard.py | Checks figures, counts, and names in the answer against tool results. |
Process map
One request through the harness
- 1AuthenticateKey, scope, rate limit, quota. Scope decides whether tools are offered. (decision)
- 2TriageCheap gates first: regex, then embedding similarity. Some questions skip tools. (decision)
- 3Compose promptBase instructions, at most one matched skill, the caller's memory section. (input)
- 4PlannerModel call with tool schemas. Returns an answer or a set of tool calls. (model call)
- 5Tool waveAll requested calls run at once. Hooks run before and after each. Results are capped. (tool)loops to step 4: back to planner, up to 6 rounds
- 6GuardThe draft is checked against this request's evidence. One corrective retry on failure. (decision)
- 7RespondStreamed or complete, in the same format the caller sent. (output)
- 8RecordLogs and memory writes happen after the response and never block it. (storage)
Build steps
Each step states why it sits at this point. Open any step on its own. The first is open.
01Fix the public contract first
Accept and return the OpenAI chat-completions format. Existing SDKs, editors, and UIs then work unchanged.
Keep one internal entry point. The reference build kept the same function signature when the loop was rewritten as a graph, so the HTTP layer needed no change.
# The whole public surface of the loop.
async def run(payload: dict, engine_url: str, allow_tools: bool, key_id: str):
"""-> (openai_response_dict, prompt_tokens, completion_tokens, model_calls)"""02Build the flat loop and bound it
Call the model with tool schemas. If it returns tool calls, run them, append the results as tool messages, and call again.
Stop when the model answers in plain text or when a round limit is reached. Six rounds is the reference default.
async def flat_loop(messages, tools, max_rounds=6):
for _ in range(max_rounds):
reply = await chat(messages, tools=tools)
msg = reply["choices"][0]["message"]
messages.append(msg)
calls = msg.get("tool_calls") or []
if not calls:
return msg["content"]
for call in calls:
result = await call_tool(call["function"]["name"],
json.loads(call["function"]["arguments"]))
messages.append({"role": "tool",
"tool_call_id": call["id"],
"content": json.dumps(result)[:8000]})
return None # budget exhausted, handled in a later step03Generalize to a wave graph
A node is an async function from state to a pair: state updates and the next nodes. A wave is every node scheduled at the same time.
Run a wave with gather. Merge every update. Then build the next wave as a dict keyed by node name. Two tool nodes that both name the planner produce one planner entry, which is fan-in by construction.
Accumulator keys extend or sum. All other keys overwrite. That rule is what lets concurrent nodes contribute without clobbering each other.
_EXTEND_KEYS = {"messages_append", "trace_notes"}
_SUM_KEYS = {"usage_prompt_tokens", "usage_completion_tokens", "model_calls"}
def merge_state(state: dict, update: dict) -> None:
for key, value in update.items():
if key in _EXTEND_KEYS:
target = key.removesuffix("_append")
state.setdefault(target, []).extend(value)
elif key in _SUM_KEYS:
state[key] = state.get(key, 0) + value
else:
state[key] = value
async def run_graph(entry_name, entry_fn, initial_state, max_waves=6):
state = dict(initial_state)
wave = {entry_name: entry_fn}
trace = []
for wave_idx in range(1, max_waves + 1):
if not wave:
break
trace.append({"wave": wave_idx, "nodes": list(wave)})
results = await asyncio.gather(*[fn(state) for fn in wave.values()])
for update, _ in results: # merge everything first
if update:
merge_state(state, update)
next_wave = {}
for _, edges in results: # then route
next_wave.update(edges) # same key twice = fan-in
wave = next_wave
else:
state["_budget_exhausted"] = True
state["_trace"] = trace
return state04Put every tool behind one registry and cap results
Register hand-written functions and MCP tools in the same table: name, schema, callable.
Cap each result (8,000 characters on the reference build) and cap the total per request (40,000). Tell the model when a result was truncated.
Log every call with arguments, duration, and result size.
05Compose the system prompt from parts
Start with a short base prompt. Add at most one matched skill. Add a memory section for the caller.
Keep the base short. It is charged on every turn. On the reference build a rule is added only after it has been broken once.
def compose_system_prompt(messages, partition_key):
parts = [read_agent_md()]
skill = match_skill(last_user_text(messages)) # zero or one
if skill:
parts.append(skill.body)
memory = compose_memory_section(partition_key)
if memory:
parts.append(memory)
return "\n\n".join(parts)06Put cheap deterministic gates in front of expensive ones
Some questions are about the agent itself and need no tools. Detect them with a regex first, then with embedding similarity to a few canonical examples, and only then let the model decide.
Log each triage decision with the reason, so a wrong route can be traced.
07Handle budget exhaustion as a designed outcome
When rounds run out, make one final model call with tools removed and an instruction to answer from the evidence gathered so far.
Count how often this happens. A question shape that regularly exhausts the budget needs a purpose-built batch tool.
08Add hooks around tool execution
A pre-tool hook can block a call. The reference build restricts file writes to one folder.
A post-tool hook can trigger follow-up work, such as re-indexing after a write.
def write_path_guard(name: str, args: dict) -> str | None:
"""Return a reason to block, or None to allow."""
if name in WRITE_TOOLS:
p = Path(args.get("path", "")).resolve()
if not p.is_relative_to(ALLOWED_ROOT):
return f"writes are limited to {ALLOWED_ROOT.name}/"
return None09Check the answer against the evidence
Extract precise figures, counts, and named entities from the draft. Look for each in the tool results of this request.
A figure is grounded if it equals an evidence number at some unit scale, or is the sum or difference of two grounded figures.
If the draft fails, send one corrective message naming what was unsupported and let the model retry. If the retry fails, return a plain statement of what was found, without the unsupported figures.
For streaming, hold tokens until the check passes.
10Stream, and record after responding
Emit server-sent events in the standard chunk format. Accumulate tool-call deltas until a call is complete, then execute it.
Write the exchange to memory as a fire-and-forget task with its own error handling.
Hold a busy marker for the whole stream so background jobs pause until the answer finishes.
How it works with the other parts
- SkillsThe harness matches one skill per request and injects its body. Skills change behavior without changing harness code.
- MemoryMemory is read during prompt composition and written after the response.
- Guardrails and verificationHooks act on tool calls. The answer guard acts on the final draft. Both are called by the loop.
- EvaluationEvery request, tool call, skill match, and triage decision is logged as JSON lines. Evals and self-observation read those logs.
- Multi-agent patternsThe same wave executor runs sub-agents as nodes. A sub-agent is a planner with its own tools and budget.
What went wrong in the real build
Right query, right rows, wrong answer
- What happened
- The agent ran correct queries, received correct rows, then answered with a table about an unrelated subject and invented totals. Nothing compared the answer with the tool results.
- Fix
- A deterministic guard that requires precise figures in the answer to appear in the evidence.
- Lesson
- Correct retrieval does not guarantee a grounded answer. Check the last step.
An answer with no tool call passed unchecked
- What happened
- A follow-up question was answered with three invented totals and no query. The guard only ran when tools had been called.
- Fix
- On data turns, check tool-free answers against the conversation so far.
- Lesson
- Decide what the default is when a check has nothing to compare against.
A guard that blocked a requested example
- What happened
- A user asked for a made-up numeric example. The guard rejected the invented amounts and the user got an apology.
- Fix
- Detect explicit requests for examples and skip the figure check on that turn.
- Lesson
- Every guard needs a list of legitimate cases it must let through, with tests.
Sixteen thousand log lines from one probe
- What happened
- The health endpoint required a key. The watchdog sent none, so 58 percent of the server log was 401 responses, and the supervisor could not tell healthy from sick.
- Fix
- An unauthenticated liveness response with no configuration detail. Details still require a key.
- Lesson
- Separate liveness from detail, and keep noise out of logs that other jobs read.
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Write the loop as an MLflow ResponsesAgent or with a framework such as LangGraph, log it to Unity Catalog, and deploy it to Model Serving. Tracing replaces your JSON-lines logs. Agent Bricks offers managed agents when you do not need a custom loop. |
| IBM watsonx | watsonx Orchestrate is the harness. You declare agents, tools, and collaborators with the Agent Development Kit, and the platform runs the loop. Custom loop logic goes into tools or a LangGraph agent you import. |
| Codex | Codex is a finished harness for coding work. You shape it with AGENTS.md, skills, MCP servers, and approval and sandbox settings. To build your own loop, use the OpenAI Agents SDK, which supplies the planner loop, handoffs, guardrails, and sessions. |
| Cursor | Cursor is a finished harness inside an editor. Rules, AGENTS.md, MCP, hooks, and subagents are your control points. You cannot change its loop, so policy belongs in hooks and tools. |
| Claude Code / Agent SDK | Claude Code is a harness with the same parts: a project instruction file, skills, hooks, MCP tools, subagents. The Agent SDK exposes that loop as a library so you can run it in your own service. |
| Another machine | The reference loop is plain Python with an HTTP client. Copy the two files and change the engine URL. Nothing in it is specific to the machine. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why build the flat loop before the wave graph?
How does the executor avoid running the planner three times after three tool calls?
Why does the harness depend on the state module?
Where would you enforce a rule that the agent may never write outside one folder, and why there?
You are moving this agent to a managed platform that owns the loop. What do you still own?
Terms
| Wave | The set of nodes that run concurrently in one step of the executor. |
| Fan-in | Several nodes converging on one successor, which then runs once. |
| Triage | A cheap decision made before the planner about how to handle a request. |
| Hook | Code that runs before or after a tool call and can block or extend it. |
From the live build
Recent changes and files the sync job filed under this module.
- Out of tool rounds: one final turn without tools answers from the evidence gathered; restate standing after promotion; ordered promotion queue
- BUSY sentinel is held until a streamed response finishes, so the learner pauses for the whole chat answer
- Spending orchestration
- 03 — The Agent Harness, Deep Dive
- Agent-Server Latency Investigation — Summary (2026-07-30)
- builds the specific planner/tool graph on top of graph.py's generic engine, and exposes the same run() contract main.py already calls.
- U7/U8: the account layer, and the reconciliation. Two tools: budget_execution the appropriation's own story: total resources -> obligated -> outlaid -> unobligated, by Treasury Account, program activity or object class. THE ONLY PLACE OUTLAYS LIVE.
- supplemental_appropriations: emergency and disaster money, by the law that created it.
- minimal state-graph execution engine. Why this exists (v2 refinement, 2026-07-27): step 2's agent_loop.py was a flat sequential loop — call the model, run whatever tools it asked for in one batch, repeat.