Agent Build Tutor
Carry it elsewhere

Every module on every platform

The parts stay the same across platforms. What changes is who runs each part: you, or the platform. Read a row to see one idea in six places. Read a column to plan a build on one platform.

What you always own

Tool definitions, instructions and skills, answer checks, evaluation sets, and the curation of your knowledge. Keep these outside any one vendor's format where you can.

What a platform usually takes over

Model serving, the agent loop, state storage, tracing, scheduling, and access control. You configure them and you still have to measure them.

ModuleDatabricksIBM watsonxCodexCursorClaude Code / Agent SDKAnother machine
Inference engineeringModel Serving replaces the engine. Foundation Model APIs give pay-per-token endpoints, and provisioned throughput reserves capacity. Your memory budget becomes a throughput and cost budget. The endpoint is OpenAI-compatible, so the contract from step 4 carries over.watsonx.ai hosts foundation models behind an inference API, with on-demand deployments for dedicated capacity. Model choice is a catalog decision. You still measure tokens per second and time to first token per model.Codex is a client of an inference endpoint. It uses hosted models by default and can be pointed at another provider in its config, including a local OpenAI-compatible server. The serving work stays on whatever machine hosts the model.Cursor consumes inference. It can override the OpenAI base URL to reach your own endpoint for some features. Serving decisions are made wherever that endpoint runs.Claude Code and the Agent SDK call a hosted model, or a compatible gateway you configure. On the reference build a translation gateway lets the coding client talk to local models. Context budgets still apply and are charged on every turn.On an NVIDIA machine use vLLM or llama.cpp in place of the Mac runtimes. The budget is VRAM per card. Use systemd in place of launchd. Everything from step 3 onward is identical.
Knowledge managementLand raw files in a Unity Catalog volume. Use a Lakeflow pipeline for validation and chunk tables. The ledger is a Delta table, and tiers are a column.Store documents in watsonx.data or object storage and register them as knowledge sources. Keep the ledger and tier rules as your own tables.Not provided. Keep the bank as a repo or bucket with the same ledger and let agents reach it through a search tool.Not provided. Same approach: your pipeline, your ledger, exposed through a tool.Not provided by the coding agent. Project knowledge features hold small document sets. A large bank still needs this pipeline.Folders, a hash ledger, and a rules file work on any filesystem.
API and gatewayModel Serving endpoints are the API. AI Gateway adds rate limits, usage tracking, guardrails, and routing across providers. Identity is the workspace's tokens and service principals.watsonx.ai exposes inference endpoints behind IAM keys. A model gateway routes to third-party providers. Orchestrate exposes agent endpoints.Codex is a client. It needs a provider base URL and key. If you point it at your gateway, issue it its own key.A client. Give it a key of its own and a base URL override where supported.A client. A base URL setting can route it through your gateway. Issue a separate key per machine.A small FastAPI app or an off-the-shelf LLM proxy does the same job on any host.
State and storageDelta tables for durable state. Lakebase for Postgres-style transactional state such as threads and checkpoints. Run state lives in your agent code.Orchestrate manages thread state for its agents. Durable domain state goes in watsonx.data or a database you connect.Session history is managed by the tool and can be resumed. Durable project state is files in the repo.Chat history is managed by the editor. Durable state is files, plus whatever your MCP servers store.Sessions can be resumed. The SDK exposes session ids. Durable state is files and your own stores.Postgres in a container gives the same guarantees. SQLite is fine for one writer.
RAG and knowledge graphVector Search with a Delta Sync index replaces the chunk table and HNSW index, and it offers hybrid search. Chunks live in a Delta table governed by Unity Catalog. Graph edges are Delta tables you join. Lakebase gives you Postgres if you want to keep this SQL as written.watsonx.data provides the lakehouse and a vector engine (Milvus). Orchestrate agents attach knowledge bases backed by it. You keep the same pipeline stages: validate, chunk, embed, hybrid query, rerank.Codex does not provide a retrieval store. Run this stack as a service and expose search as an MCP tool. Codex then calls it like any other tool.Cursor indexes your code for its own use. Domain retrieval is yours to host. Register your search service as an MCP server in the project config.Expose the search tool through an MCP server. The reference build does this so a coding agent and the chat agent share one retrieval implementation.Postgres with pgvector runs anywhere, including Docker. The SQL in this module is unchanged. Only the embedding endpoint URL differs.
Tools and MCPUnity Catalog functions are governed tools. Managed MCP servers expose Vector Search, Genie, and functions. You can also host a custom MCP server as a Databricks App.Orchestrate tools are Python functions, OpenAPI specs, or MCP servers imported with the ADK. A tool written as an MCP server here can be registered there.Add MCP servers in the config file. Tools then appear to the agent. This is the direct route for reusing your own tools.Add servers to the project MCP config. Same server, no code change.Add servers to the project MCP file. The reference build shares one server between its chat agent and its coding agent this way.An MCP server is a process speaking JSON-RPC over stdio or HTTP. It runs anywhere Python or Node runs.
Harness engineeringWrite the loop as an MLflow ResponsesAgent or with a framework such as LangGraph, log it to Unity Catalog, and deploy it to Model Serving. Tracing replaces your JSON-lines logs. Agent Bricks offers managed agents when you do not need a custom loop.watsonx Orchestrate is the harness. You declare agents, tools, and collaborators with the Agent Development Kit, and the platform runs the loop. Custom loop logic goes into tools or a LangGraph agent you import.Codex is a finished harness for coding work. You shape it with AGENTS.md, skills, MCP servers, and approval and sandbox settings. To build your own loop, use the OpenAI Agents SDK, which supplies the planner loop, handoffs, guardrails, and sessions.Cursor is a finished harness inside an editor. Rules, AGENTS.md, MCP, hooks, and subagents are your control points. You cannot change its loop, so policy belongs in hooks and tools.Claude Code is a harness with the same parts: a project instruction file, skills, hooks, MCP tools, subagents. The Agent SDK exposes that loop as a library so you can run it in your own service.The reference loop is plain Python with an HTTP client. Copy the two files and change the engine URL. Nothing in it is specific to the machine.
EvaluationMLflow evaluation runs scorers and judges over an evaluation dataset stored in Unity Catalog, and traces link each score to the run that produced it. Your gold sets become tables. Sealing becomes table permissions. The metric code in this module can be registered as custom scorers.watsonx.governance evaluates and monitors deployed models and agents, and the Orchestrate ADK includes an evaluation framework for agent trajectories. Keep your own gold sets and feed them in. The sample-size discipline is unchanged.Use the provider's evals tooling or plain scripts. Codex can run your eval scripts headless as part of a change, which makes eval a gate in the coding loop.Evals are scripts in your repo. A rule or hook can require the retrieval eval to pass before a change to ranking code is accepted.Run eval scripts through the coding agent, or score plugin and skill behavior with its eval commands. Keep gold files out of the agent's writable paths.The eval scripts are Python and JSON files. They move with the repo. Rebuild the gold set only if the corpus changed.
Guardrails and verificationAI Gateway guardrails cover safety and PII at the endpoint. Groundedness checks like this one are custom code in your agent, or scorers run on traces.watsonx.governance provides guardrails and monitors. Domain checks go in a tool or a post-processing step you write.Sandbox and approval modes limit actions. Hooks and your own test commands act as checks on changes.Hooks can block or modify agent actions. Rules can require checks before completion.Hooks run before and after tool use and can block. Permission modes limit actions. Output checks are code you add in hooks or in the SDK loop.The guard is a pure Python module that takes an answer and evidence. Portable as is.
MemoryUse Lakebase or Delta tables keyed by user, read in your agent code. Long-term memory patterns ship with the agent framework and LangGraph stores.Orchestrate keeps conversation context. Long-term memory is a store you connect and read through a tool.Persistent instructions live in AGENTS.md. Memories are files the agent is told to read and update.Rules and memory features hold persistent preferences. Project facts belong in repo files.A project instruction file plus memory files and tools. The pattern is the same: write after, read at start, scope by project.A table with a partition column and a summarizer. Portable.
SkillsNo native skill format. Store skill files in a volume and load them in your agent code, or register prompts in the prompt registry.Agent instructions and per-agent guidelines play this role. Narrow agents as collaborators are the platform's way to scope instructions.Skills are supported as folders with a SKILL.md. The reference skill bodies can be moved with small edits.Rules files scoped by path or description, plus skills. The description line does the same matching job.Native. A folder per skill with a SKILL.md whose description decides when it loads.Markdown files and a 100-line loader. Fully portable.
Multi-agent patternsMulti-agent supervisors are available as a managed pattern, and LangGraph or the OpenAI Agents SDK run inside a deployed agent. Each sub-agent can be its own serving endpoint or a function.Orchestrate agents list collaborators and route work among them. This is the platform's core pattern.Subagents run tasks in separate contexts. The Agents SDK supports handoffs between agents.Subagents and background agents run tasks in parallel with their own context.Subagents have their own context window and tool set, defined in files. The SDK can spawn them from code.The wave executor is plain Python. Sub-agents are functions.
Operations and deploymentServing endpoints, Jobs, and Apps are supervised by the platform. You own versions, permissions, cost alerts, and promotion between workspaces with asset bundles.The platform runs the services. You own environments, deployment spaces, and promotion between them.Operations here means CI: headless runs, sandbox settings, and secrets handling for the agent.Same: team rules in the repo, background agent settings, and secrets.Headless runs in CI, managed settings, and permission policies. For a service built on the SDK, operate it like any other service.systemd units and timers in place of launchd. The backup and status scripts are shell and Python.
SDKs and frameworksThe agent framework accepts agents written with MLflow's agent interfaces, LangGraph, or the OpenAI Agents SDK, and adds deployment, tracing, and evaluation around them.The Orchestrate ADK defines agents and tools in files and a CLI. It also imports agents built with LangGraph and similar frameworks.The OpenAI Agents SDK provides the loop, handoffs, guardrails, sessions, and tracing. Codex itself can be run headless and scripted.Cursor provides a command-line agent and background agents. For your own service use a general SDK.The Claude Agent SDK exposes the coding agent's loop, tools, hooks, subagents, and MCP support as a library in Python and TypeScript.Any SDK that speaks the OpenAI-compatible API can target a local engine by base URL.
Self-evolving agentsLakeflow Jobs schedule the loop. Traces and labeled feedback are the raw material. Evaluation datasets grow from production traces, and judges can be aligned with human labels.Schedule jobs outside Orchestrate and write results to a governed store. Monitors in watsonx.governance provide the drift signals.A scheduled headless run can review logs and propose changes to AGENTS.md or skills as a pull request for review.Background agents can do the same review and open a pull request.A scheduled headless session can read logs and propose skill or instruction edits. Keep a human merge step.A scheduler such as launchd, systemd timers, or cron, plus the same scripts.

Platform feature names were checked in October 2026 and change often. Confirm against current vendor documentation before you commit to a design.