Agent Build Tutor
Map / Full lesson

Harness engineering

The loop around the model: prompt assembly, tool execution, budgets, checks, logs.

What it is and why it exists

What

The harness is the program that runs the model. It assembles the prompt, offers tools, executes the tool calls the model asks for, feeds results back, decides when to stop, checks the answer, and records what happened.

Why

A model call is stateless and cannot act. Everything that makes an agent behave reliably lives in the harness: which tools exist, how many rounds are allowed, what gets truncated, what is checked before the user sees it. Two systems with the same model and different harnesses behave like different products.

How it works

Where it sits in the build order

Needs first

  • API and gatewayThe harness sits behind the gateway contract. Keys carry scopes, and the harness reads the scope to decide whether a caller gets tools at all.Build out of order Stub it with: A hard-coded flag: tools allowed, one anonymous caller.
  • Inference engineeringThe planner is a model call. Round budgets, output caps, and prompt size limits are set from measured inference numbers.Build out of order Stub it with: A fake model function that returns a scripted tool call on the first turn and text on the second.
  • State and storageThe loop is a state machine. Concurrent tool results have to merge into one conversation without overwriting each other, so the merge rules are defined before the loop that relies on them.Build out of order Stub it with: A plain dict and sequential tool execution.
  • Tools and MCPA loop with no tools is a chat proxy. You need at least one registered tool with a schema to exercise the tool path.Build out of order Stub it with: One local function such as a clock or calculator.

Unlocks

  • EvaluationEnd-to-end evals send requests through the agent and read its logs. The log formats are defined by the harness.
  • Guardrails and verificationThe guard is a stage in the loop: after the draft, before the response. Hooks wrap tool execution. Both need the loop.
  • MemoryMemory is read at prompt composition and written after the response. Both are harness stages.
  • SkillsA skill is text injected during prompt composition. Without the composition step there is nowhere to load it.
  • Multi-agent patternsA sub-agent is a harness loop run as a node. You need one working loop before you can run several.
  • Operations and deploymentThe agent server is the main service. Its health endpoint and busy marker are what the supervisor and scheduler read.
  • SDKs and frameworksYou evaluate an SDK by comparing it with a loop you understand. Without that, every framework's defaults look like requirements.

In the reference build

PathRole
apps/agent-server/main.pyHTTP surface: chat completions, models, feedback, health.
apps/agent-server/graph.pyA 110-line wave executor: fan-out, fan-in, merge rules, wave budget, trace.
apps/agent-server/agent_loop.pyThe planner and tool graph, prompt composition, streaming, logging.
apps/agent-server/AGENT.mdThe base system prompt.
apps/agent-server/tools/registry.pyOne registry for hand-written and MCP tools.
apps/agent-server/hooks/Pre-tool and post-tool hooks.
apps/agent-server/answer_guard.pyChecks figures, counts, and names in the answer against tool results.

Process map

One request through the harness

  1. 1
    AuthenticateKey, scope, rate limit, quota. Scope decides whether tools are offered. (decision)
  2. 2
    TriageCheap gates first: regex, then embedding similarity. Some questions skip tools. (decision)
  3. 3
    Compose promptBase instructions, at most one matched skill, the caller's memory section. (input)
  4. 4
    PlannerModel call with tool schemas. Returns an answer or a set of tool calls. (model call)
  5. 5
    Tool waveAll requested calls run at once. Hooks run before and after each. Results are capped. (tool)
    loops to step 4: back to planner, up to 6 rounds
  6. 6
    GuardThe draft is checked against this request's evidence. One corrective retry on failure. (decision)
  7. 7
    RespondStreamed or complete, in the same format the caller sent. (output)
  8. 8
    RecordLogs and memory writes happen after the response and never block it. (storage)

Build steps

Each step states why it sits at this point. Open any step on its own. The first is open.

01Fix the public contract first
Why this step is hereThe contract is what callers depend on. If it stays constant, you can rebuild everything behind it without breaking a client.

Accept and return the OpenAI chat-completions format. Existing SDKs, editors, and UIs then work unchanged.

Keep one internal entry point. The reference build kept the same function signature when the loop was rewritten as a graph, so the HTTP layer needed no change.

python
# The whole public surface of the loop.
async def run(payload: dict, engine_url: str, allow_tools: bool, key_id: str):
    """-> (openai_response_dict, prompt_tokens, completion_tokens, model_calls)"""
You are done whenAn unmodified OpenAI client pointed at your server gets a valid completion.
02Build the flat loop and bound it
Why this step is hereThe flat loop is the smallest thing that is an agent. Get it correct and bounded before adding concurrency, because every later feature is a refinement of it.

Call the model with tool schemas. If it returns tool calls, run them, append the results as tool messages, and call again.

Stop when the model answers in plain text or when a round limit is reached. Six rounds is the reference default.

python
async def flat_loop(messages, tools, max_rounds=6):
    for _ in range(max_rounds):
        reply = await chat(messages, tools=tools)
        msg = reply["choices"][0]["message"]
        messages.append(msg)
        calls = msg.get("tool_calls") or []
        if not calls:
            return msg["content"]
        for call in calls:
            result = await call_tool(call["function"]["name"],
                                     json.loads(call["function"]["arguments"]))
            messages.append({"role": "tool",
                             "tool_call_id": call["id"],
                             "content": json.dumps(result)[:8000]})
    return None  # budget exhausted, handled in a later step
You are done whenA question that needs one tool call produces two model calls and one tool message.
03Generalize to a wave graph
Why this step is hereThis follows the flat loop because it is the same behavior with two additions: tool calls in one turn run concurrently, and their results merge safely. You need the sequential version as the reference for correctness.

A node is an async function from state to a pair: state updates and the next nodes. A wave is every node scheduled at the same time.

Run a wave with gather. Merge every update. Then build the next wave as a dict keyed by node name. Two tool nodes that both name the planner produce one planner entry, which is fan-in by construction.

Accumulator keys extend or sum. All other keys overwrite. That rule is what lets concurrent nodes contribute without clobbering each other.

graph.py (core)
_EXTEND_KEYS = {"messages_append", "trace_notes"}
_SUM_KEYS = {"usage_prompt_tokens", "usage_completion_tokens", "model_calls"}

def merge_state(state: dict, update: dict) -> None:
    for key, value in update.items():
        if key in _EXTEND_KEYS:
            target = key.removesuffix("_append")
            state.setdefault(target, []).extend(value)
        elif key in _SUM_KEYS:
            state[key] = state.get(key, 0) + value
        else:
            state[key] = value

async def run_graph(entry_name, entry_fn, initial_state, max_waves=6):
    state = dict(initial_state)
    wave = {entry_name: entry_fn}
    trace = []
    for wave_idx in range(1, max_waves + 1):
        if not wave:
            break
        trace.append({"wave": wave_idx, "nodes": list(wave)})
        results = await asyncio.gather(*[fn(state) for fn in wave.values()])
        for update, _ in results:           # merge everything first
            if update:
                merge_state(state, update)
        next_wave = {}
        for _, edges in results:            # then route
            next_wave.update(edges)         # same key twice = fan-in
        wave = next_wave
    else:
        state["_budget_exhausted"] = True
    state["_trace"] = trace
    return state
You are done whenThree tool calls requested in one turn show as one wave with three nodes in the trace, followed by one planner node.
04Put every tool behind one registry and cap results
Why this step is hereThe planner should not know where a tool lives. A single registry also gives you one place to enforce size limits, which protects the context budget from the inference module.

Register hand-written functions and MCP tools in the same table: name, schema, callable.

Cap each result (8,000 characters on the reference build) and cap the total per request (40,000). Tell the model when a result was truncated.

Log every call with arguments, duration, and result size.

You are done whenA tool that returns a megabyte of text reaches the model as a bounded string with a truncation note.
05Compose the system prompt from parts
Why this step is hereComposition comes after tools because the prompt has to describe when to use them. It comes before guards because guards reference which skill was loaded.

Start with a short base prompt. Add at most one matched skill. Add a memory section for the caller.

Keep the base short. It is charged on every turn. On the reference build a rule is added only after it has been broken once.

python
def compose_system_prompt(messages, partition_key):
    parts = [read_agent_md()]
    skill = match_skill(last_user_text(messages))   # zero or one
    if skill:
        parts.append(skill.body)
    memory = compose_memory_section(partition_key)
    if memory:
        parts.append(memory)
    return "\n\n".join(parts)
You are done whenThe skill-match log shows most requests matching no skill. That is the expected case.
06Put cheap deterministic gates in front of expensive ones
Why this step is hereTriage runs before the planner, so it has to exist before you tune the planner. It also removes whole classes of slow requests.

Some questions are about the agent itself and need no tools. Detect them with a regex first, then with embedding similarity to a few canonical examples, and only then let the model decide.

Log each triage decision with the reason, so a wrong route can be traced.

You are done whenAsking the agent what it can do returns quickly with zero tool calls in the log.
07Handle budget exhaustion as a designed outcome
Why this step is hereIt depends on the wave budget from step 3. Without it, hitting the limit returns an error after the user has waited the longest.

When rounds run out, make one final model call with tools removed and an instruction to answer from the evidence gathered so far.

Count how often this happens. A question shape that regularly exhausts the budget needs a purpose-built batch tool.

You are done whenWith the budget set to 1, a multi-step question still returns a grounded partial answer.
08Add hooks around tool execution
Why this step is hereHooks need the registry and the loop to exist. They are where policy is enforced in code, which is stronger than asking the model to behave.

A pre-tool hook can block a call. The reference build restricts file writes to one folder.

A post-tool hook can trigger follow-up work, such as re-indexing after a write.

python
def write_path_guard(name: str, args: dict) -> str | None:
    """Return a reason to block, or None to allow."""
    if name in WRITE_TOOLS:
        p = Path(args.get("path", "")).resolve()
        if not p.is_relative_to(ALLOWED_ROOT):
            return f"writes are limited to {ALLOWED_ROOT.name}/"
    return None
You are done whenA write outside the allowed folder is refused and the refusal appears in the tool log.
09Check the answer against the evidence
Why this step is hereThe guard is last in the run because it needs the complete evidence set and the draft. It is built after logging exists, since each rule was written from a logged failure.

Extract precise figures, counts, and named entities from the draft. Look for each in the tool results of this request.

A figure is grounded if it equals an evidence number at some unit scale, or is the sum or difference of two grounded figures.

If the draft fails, send one corrective message naming what was unsupported and let the model retry. If the retry fails, return a plain statement of what was found, without the unsupported figures.

For streaming, hold tokens until the check passes.

You are done whenForce a wrong figure into a draft in a test. The guard rejects it and names the figure.
10Stream, and record after responding
Why this step is hereStreaming changes how the loop yields output, so add it once the loop and guard are stable. Recording goes after the response so memory and logging never add latency.

Emit server-sent events in the standard chunk format. Accumulate tool-call deltas until a call is complete, then execute it.

Write the exchange to memory as a fire-and-forget task with its own error handling.

Hold a busy marker for the whole stream so background jobs pause until the answer finishes.

You are done whenFirst token arrives within a few seconds on a warm model, and a memory-store failure does not change the response.

How it works with the other parts

What went wrong in the real build

2026-10

Right query, right rows, wrong answer

What happened
The agent ran correct queries, received correct rows, then answered with a table about an unrelated subject and invented totals. Nothing compared the answer with the tool results.
Fix
A deterministic guard that requires precise figures in the answer to appear in the evidence.
Lesson
Correct retrieval does not guarantee a grounded answer. Check the last step.
2026-10

An answer with no tool call passed unchecked

What happened
A follow-up question was answered with three invented totals and no query. The guard only ran when tools had been called.
Fix
On data turns, check tool-free answers against the conversation so far.
Lesson
Decide what the default is when a check has nothing to compare against.
2026-10

A guard that blocked a requested example

What happened
A user asked for a made-up numeric example. The guard rejected the invented amounts and the user got an apology.
Fix
Detect explicit requests for examples and skip the figure check on that turn.
Lesson
Every guard needs a list of legitimate cases it must let through, with tests.
2026-09

Sixteen thousand log lines from one probe

What happened
The health endpoint required a key. The watchdog sent none, so 58 percent of the server log was 401 responses, and the supervisor could not tell healthy from sick.
Fix
An unauthenticated liveness response with no configuration detail. Details still require a key.
Lesson
Separate liveness from detail, and keep noise out of logs that other jobs read.

The same idea on other platforms

PlatformHow this module maps
DatabricksWrite the loop as an MLflow ResponsesAgent or with a framework such as LangGraph, log it to Unity Catalog, and deploy it to Model Serving. Tracing replaces your JSON-lines logs. Agent Bricks offers managed agents when you do not need a custom loop.
IBM watsonxwatsonx Orchestrate is the harness. You declare agents, tools, and collaborators with the Agent Development Kit, and the platform runs the loop. Custom loop logic goes into tools or a LangGraph agent you import.
CodexCodex is a finished harness for coding work. You shape it with AGENTS.md, skills, MCP servers, and approval and sandbox settings. To build your own loop, use the OpenAI Agents SDK, which supplies the planner loop, handoffs, guardrails, and sessions.
CursorCursor is a finished harness inside an editor. Rules, AGENTS.md, MCP, hooks, and subagents are your control points. You cannot change its loop, so policy belongs in hooks and tools.
Claude Code / Agent SDKClaude Code is a harness with the same parts: a project instruction file, skills, hooks, MCP tools, subagents. The Agent SDK exposes that loop as a library so you can run it in your own service.
Another machineThe reference loop is plain Python with an HTTP client. Copy the two files and change the engine URL. Nothing in it is specific to the machine.

Explain it back

Answer aloud first. Then open the answer and compare.

Why build the flat loop before the wave graph?
A strong answerThe graph is the flat loop plus concurrency and safe merging. The flat loop is the correctness reference and is easy to test. If you start with the graph you debug orchestration and agent behavior at the same time.
How does the executor avoid running the planner three times after three tool calls?
A strong answerThe next wave is a dict keyed by node name. Each tool node names the planner as its successor, and writing the same key three times leaves one entry.
Why does the harness depend on the state module?
A strong answerConcurrent tool nodes each produce updates. The merge rules decide which keys accumulate and which overwrite. Those rules have to be fixed before nodes run in parallel, or results are silently lost.
Where would you enforce a rule that the agent may never write outside one folder, and why there?
A strong answerIn a pre-tool hook. It runs in code on every call regardless of what the model was told, and it produces a logged refusal.
You are moving this agent to a managed platform that owns the loop. What do you still own?
A strong answerTool definitions and their result limits, the instructions and skills, the checks on the final answer, the evaluation sets, and the logs or traces you review. The loop itself is the most replaceable part.

Terms

WaveThe set of nodes that run concurrently in one step of the executor.
Fan-inSeveral nodes converging on one successor, which then runs once.
TriageA cheap decision made before the planner about how to handle a request.
HookCode that runs before or after a tool call and can block or extend it.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module