A playbook built from a working system
Build an agent stack, and know why each piece comes when it does. Fifteen modules, from serving a model to a loop that improves itself. Every dependency on the map carries its reason, and every module says what to stub if you want to build it out of order. Each one also shows the same idea on Databricks, watsonx, Codex, Cursor, Claude Code, and a second machine.
Full lesson Must exist first Unlocked nextSelect a module. Scroll sideways on a small screen.
Selected
Harness engineering The loop around the model: prompt assembly, tool execution, budgets, checks, logs.
Open lesson Why these come first API and gateway The harness sits behind the gateway contract. Keys carry scopes, and the harness reads the scope to decide whether a caller gets tools at all. Build out of order Stub it with: A hard-coded flag: tools allowed, one anonymous caller.Inference engineering The planner is a model call. Round budgets, output caps, and prompt size limits are set from measured inference numbers. Build out of order Stub it with: A fake model function that returns a scripted tool call on the first turn and text on the second.State and storage The loop is a state machine. Concurrent tool results have to merge into one conversation without overwriting each other, so the merge rules are defined before the loop that relies on them. Build out of order Stub it with: A plain dict and sequential tool execution.Tools and MCP A loop with no tools is a chat proxy. You need at least one registered tool with a schema to exercise the tool path. Build out of order Stub it with: One local function such as a clock or calculator.What this unlocks Evaluation End-to-end evals send requests through the agent and read its logs. The log formats are defined by the harness. Guardrails and verification The guard is a stage in the loop: after the draft, before the response. Hooks wrap tool execution. Both need the loop. Memory Memory is read at prompt composition and written after the response. Both are harness stages. Skills A skill is text injected during prompt composition. Without the composition step there is nowhere to load it. Multi-agent patterns A sub-agent is a harness loop run as a node. You need one working loop before you can run several. Operations and deployment The agent server is the main service. Its health endpoint and busy marker are what the supervisor and scheduler read. SDKs and frameworks You evaluate an SDK by comparing it with a loop you understand. Without that, every framework's defaults look like requirements. Why does A come before B? Pick any two modules. If one depends on the other you get the chain of reasons. If neither does, you can build them in either order.
Module A Inference engineering Knowledge management API and gateway State and storage RAG and knowledge graph Tools and MCP Harness engineering Evaluation Guardrails and verification Memory Skills Multi-agent patterns Operations and deployment SDKs and frameworks Self-evolving agents Module B Inference engineering Knowledge management API and gateway State and storage RAG and knowledge graph Tools and MCP Harness engineering Evaluation Guardrails and verification Memory Skills Multi-agent patterns Operations and deployment SDKs and frameworks Self-evolving agents
Knowledge management comes before Evaluation .
Evaluation needs Knowledge managementHeld-out material is selected from the corpus. You cannot seal chunks that have not been ingested. Start with a full lesson From the live build Last snapshot 2026-10-08: 58 changes, 44 design notes, 136 source files, sorted into modules. See what changed