Map / Outline
Tools and MCP
Give the model typed actions, and serve them over a protocol any harness can call.
What it is and why it exists
What
A tool is a function with a name, a description, and a JSON schema for its arguments. The Model Context Protocol is a standard way to serve tools from a separate process so any compatible client can list and call them.
Why
Tools are how an agent reads data and acts. Serving them over MCP means the chat agent, a coding agent, and an editor share one implementation. Tool design also decides how often the model picks the wrong instrument.
How it works
- Write each tool as a narrow function that returns structured data and, when empty, a diagnostic naming the filter that emptied the result.
- Describe when to use it in the schema description. Keep the number of tools small. Twelve tools without routing made the reference agent slower and more confidently wrong.
- Register tools in one table. The harness calls by name and does not care whether a tool is local or remote.
- Wrap the same functions in an MCP server process. Clients start it from a config entry with a command and environment.
- Measure schema cost. Tool schemas are sent on every turn. The reference server has a flag that prints its token cost.
Where it sits in the build order
Needs first
- API and gatewayTool calling rides on the chat contract: tool schemas go in the request and tool calls come back in the response. Key scope decides who gets tools.Build out of order Stub it with: Call tool functions directly from a test script.
Unlocks
- Harness engineeringA loop with no tools is a chat proxy. You need at least one registered tool with a schema to exercise the tool path.
- Guardrails and verificationThe guard compares the answer with tool results. Tools must return structured values the guard can parse.
- SDKs and frameworksTools are the part you carry between SDKs. Having them as MCP servers makes the comparison a configuration change.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/tools/registry.py | Schemas and the call_tool dispatcher. |
| apps/agent-server/mcp_client.py | Connects to external MCP servers and registers their tools. |
| apps/mcp-servers/brainbank/server.py | Serves the knowledge and data tools to any MCP client. |
| .mcp.json | Client config: command, arguments, environment, timeout. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Unity Catalog functions are governed tools. Managed MCP servers expose Vector Search, Genie, and functions. You can also host a custom MCP server as a Databricks App. |
| IBM watsonx | Orchestrate tools are Python functions, OpenAPI specs, or MCP servers imported with the ADK. A tool written as an MCP server here can be registered there. |
| Codex | Add MCP servers in the config file. Tools then appear to the agent. This is the direct route for reusing your own tools. |
| Cursor | Add servers to the project MCP config. Same server, no code change. |
| Claude Code / Agent SDK | Add servers to the project MCP file. The reference build shares one server between its chat agent and its coding agent this way. |
| Another machine | An MCP server is a process speaking JSON-RPC over stdio or HTTP. It runs anywhere Python or Node runs. |
Explain it back
Answer aloud first. Then open the answer and compare.
A tool returns an empty list. What should the agent be allowed to conclude?
A strong answerOnly that this filter matched nothing. It may report what it searched. It may not explain why the data cannot exist. The tool's diagnostic should name the filter to relax.
Why serve tools over MCP when the harness could import them?
A strong answerOne implementation then serves every harness. Moving to another platform means registering the server, with no rewrite.
From the live build
Recent changes and files the sync job filed under this module.
- query_spending returns the total and names the period it covers; award warehouse builds itself when a newer archive lands, with a report alarm
- Answer guard: vendor names, answers given without a tool call, streaming gate; FINANCE on Qwen3 embeddings; wiki watcher batch cap, lock and wait; downloads script
- Retry model calls that return HTTP 500 (malformed tool call) and empty final answers
- Answer guard: dollar figures must come from tool results (both answer paths); follow-ups inherit the skill; multi-turn chat eval
- Baseline 2026-09-30: code and scripts after the Postgres cutover, pinned MCP servers, nightly pg_dump
- Full-stack audit — 2026-09-07
- Runbook — making the Claude Code CLI smart about the agency
- 08 — Extending the Harness Agent
- Local Agent Server — Build Guide v2
- MCP Server Catalog — candidates for mcpservers.json
- Agent Server — apps/agent-server/main.py ========================================= STEP 2 of the "mini Claude Code" build (see AGENT_SERVER_GUIDE_v2.md). Step 1 — bare passthrough, plaintext-key auth. Built/tested 2026-07-27.
- the federal spending tools, exposed to live traffic. Three tools, matching the three shapes a real question takes: query_spending aggregate: how much, by whom, on what, when search_requirements free text: what was ACTUALLY bought screen_purchase_request pre-award: should this be ...
- pin every MCP server in mcp_servers.json to an exact version (2026-09-30). Why: every server was launched as `npx -y <pkg>` or `uvx <pkg>`, which resolves the newest release on each cold start.
- put the code and scripts on AI_DATA under git (2026-09-30). Idempotent: re-running commits whatever changed since, with the same checks. 1. whitelist .gitignore (only code + config; secrets and data never) 2.
- brainbank MCP server — exposes agent-server's tool registry to Claude Code. WHY THIS EXISTS --------------- Nothing the self-evolution pipeline learns is in the model's weights.
- Double-click button: gives Claude Code access to the knowledge bank. Built 2026-08-23, after a Claude Code session on local-qwen38 was asked about the agency a commodity spending and answered "I don't have that in my knowledge" — while brain.db held 144,781 atoms and the spending tables ...
- re-checks the 2026-09-07 fixes, on demand. $AI_DATA/apps/agent-server/venv/bin/python3 $AI_DATA/apps/agent-server/verify_fixes.py Exits 0 if every check passes, 1 otherwise, so it can go in a cycle step.
- Double-click button: adds a real WEB SEARCH tool to the agent. Built 2026-07-29, after the chatbot searched for a web-search tool and correctly reported it didn't have one. WHY IT WAS NEVER BUILT: it isn't a bug or an oversight.