Agent Build Tutor
Map / Full lesson

Inference engineering

Serve models on fixed hardware and know what each second costs.

What it is and why it exists

What

Inference engineering is the work of turning model weights into a dependable endpoint. You choose which models run, how they share memory, how long they stay loaded, how many requests run at once, and how much context each request may use.

Why

Every other layer calls this one. Retrieval calls it to embed. The harness calls it to plan and to answer. The evaluator calls it to judge. If a single call takes 70 seconds because a model was evicted, no prompt change above it will fix that. You set the speed limit of the whole system here.

How it works

Where it sits in the build order

Needs first

Nothing. You can start here.

Unlocks

  • API and gatewayA gateway forwards to an engine. Its timeouts and limits are set from measured load times and generation rates.
  • RAG and knowledge graphSearch needs an embedding model at query time and at index time. The vector column width is fixed by that model, so the model choice comes first.
  • Harness engineeringThe planner is a model call. Round budgets, output caps, and prompt size limits are set from measured inference numbers.
  • Operations and deploymentThe engine is the first service to supervise, and its memory limits decide what can run at the same time.

In the reference build

PathRole
ollama-env.shServing settings: loaded-model cap, keep-alive, parallelism, context length.
apps/agent-server/model_router.pyPicks a model per request by task type. Two mechanisms: tag swap and base-URL swap.
apps/agent-server/model_catalog.jsonThe lineup: which model serves default, fast, and vision work.
apps/agent-server/bench_latency.pyMeasures time to first token and tokens per second from real requests.
apps/mlx/Second runtime. One process serves exactly one model.

Process map

What happens to one request at the engine

  1. 1
    Request arrivesOpenAI-style JSON with a model name, messages, and max tokens. (input)
  2. 2
    Is the model resident?If yes, skip to prompt evaluation. If no, the engine must load it and may evict another. (decision)
  3. 3
    Cold loadTens of seconds for a 20 to 60 GB model. This is the largest avoidable cost. (storage)
  4. 4
    Prompt evaluationCost grows with prompt size. Tool schemas and system prompts count on every turn. (model call)
  5. 5
    GenerationOutput tokens at a steady rate. Time is tokens divided by rate. (model call)
  6. 6
    Stream outTokens leave as they are produced, so the user sees the first one early. (output)

Build steps

Each step states why it sits at this point. Open any step on its own. The first is open.

01Write the memory budget before you choose models
Why this step is hereThe budget decides the lineup. Choosing models first leads to a lineup that cannot be resident together, and you find out in production as reload delays.

List usable model memory. On Apple silicon that is unified memory minus what the OS and your services need. On a GPU server it is VRAM per card.

For each candidate model record the weight size at your chosen quantization, then add headroom for context. Longer context and more parallel slots both raise memory use.

Decide which models must be warm at the same time. On the reference machine the workhorse (22 GB) and the vision model (29 GB) stay warm together. The 61 GB reasoner evicts both, so it is reserved for hard problems where a multi-minute swap is acceptable.

budget.txt
usable model memory        ~78 GB
workhorse   35B MoE  q4     22 GB   default, ~72 tok/s
vision      27B      q8     29 GB   image input only
reasoner    120B     q4     61 GB   runs alone
embedder    0.3 GB          loads on demand, can evict a chat model
rule: workhorse + vision fit together (51 GB). Anything + reasoner does not.
You are done whenYou can state, for any two models, whether they fit together and what a swap costs in seconds.
02Install the engine and put weights on a dedicated volume
Why this step is hereWeights are large and reproducible. Keeping them off the boot disk makes the whole stack one portable folder and keeps backups small.

Install the engine with your package manager. Point its model directory at the data volume before the first pull.

Run the engine under the OS supervisor so it restarts after a crash or reboot. That is covered in the operations module, and you can start it by hand until then.

bash
brew install ollama
export OLLAMA_MODELS="$AI_DATA/models/ollama"
ollama serve &
ollama pull <workhorse-model>
ollama pull nomic-embed-text
ollama list
You are done when`ollama list` shows the models and the files sit under your data volume.
03Set the serving parameters on purpose
Why this step is hereDefaults assume one user and one model. An agent makes many calls per question, so the defaults produce evictions and queueing that look like a slow model.

Cap loaded models at what your budget allows. Set keep-alive long enough that normal gaps between requests do not unload the model.

Set parallel slots to the concurrency you expect. Each slot reserves context memory.

Set context length explicitly. Then confirm it took effect by reading the engine log, because a setting that is present is not proof that it is in effect.

ollama-env.sh
export OLLAMA_MAX_LOADED_MODELS=2   # at most 2 models resident
export OLLAMA_KEEP_ALIVE=30m        # unload after 30 idle minutes
export OLLAMA_NUM_PARALLEL=4        # concurrent requests per model
export OLLAMA_CONTEXT_LENGTH=49152  # tokens of context per slot
export OLLAMA_FLASH_ATTENTION=1     # less memory for long contexts
You are done whenSend a request, then read the engine log. The reported prompt size is the ground truth for how much context you really sent.
04Call the endpoint the way every later layer will
Why this step is hereThe OpenAI-compatible chat endpoint is the contract. If you standardize on it now, you can swap engines, machines, or cloud providers later by changing a base URL.

Test with plain curl before you add any library. You want to know the engine works independent of your code.

Record tokens per second and time to first token for each model. These two numbers explain most latency complaints later.

bash
curl -s http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<workhorse-model>",
    "messages": [{"role": "user", "content": "Reply with the word ready."}],
    "max_tokens": 16
  }' | python3 -m json.tool
You are done whenYou get a JSON completion with a usage block. Divide completion tokens by elapsed seconds to get your rate.
05Add a second runtime only where it earns its place
Why this step is hereA small fast model is useful for cheap jobs such as triage, question writing, and judging. It should not compete with the workhorse for the two resident slots.

The reference build serves a 9B model through MLX on its own port. That runtime serves exactly one model per process, chosen at startup.

This creates two routing mechanisms. Models in the main engine are selected by the model field (tag swap). The MLX model is selected by calling a different base URL.

bash
# one process, one model, its own port
mlx_lm.server --model mlx-community/<small-model>-4bit --port 8082
You are done whenThe same curl works against the second port with no model field needed.
06Route by task type, and fail safe
Why this step is hereRouting comes after the endpoints exist and are measured. A router that guesses about models you have not benchmarked sends work to the wrong place with confidence.

Classify the request: default, fast, or vision. Map each class to a catalog entry.

If anything is missing or unknown, return nothing and let the existing behavior run. The router should never be the reason a request fails.

model_router.py (shape)
def resolve_for_request(payload: dict) -> dict | None:
    if payload.get("model"):            # caller was explicit: do nothing
        return None
    task = detect_task(payload)         # "default" | "fast" | "vision"
    entry = CATALOG.get(task)
    if not entry:
        return None                     # unknown task: do nothing
    if entry.get("base_url"):           # separate runtime: swap the URL
        return {"base_url": entry["base_url"]}
    if entry.get("tag"):                # same engine: swap the tag
        return {"model": entry["tag"]}
    return None                         # incomplete entry: do nothing
You are done whenDelete the catalog file and send a request. It should still succeed on the default model.
07Cut latency where it is actually spent
Why this step is hereDo this after you have logs from real requests. Tuning before measuring usually targets model speed, which is rarely the problem.

Output volume dominates. Set a default max-tokens value and ask for shorter answers in the system prompt.

Stream responses so the first token arrives early. Reuse one pooled HTTP client across requests.

Keep prompts small. Tool schemas, system prompts, and memory are charged on every turn.

Avoid evictions. A retrieval call that loads the embedder can push out the chat model. Prefer the cheap path that loads no model when a definition will do.

You are done whenMedian latency in your own request log moves, and the log shows fewer cold loads.
08Keep the engine private
Why this step is hereThe engine has no authentication. Anything that needs outside access goes through the gateway module, which adds keys, scopes, and limits.

Bind the engine to localhost or a private network. Expose only a proxy that authenticates callers.

You are done whenFrom a device outside your private network, the engine port does not answer.

How it works with the other parts

What went wrong in the real build

2026-08

A context setting that was present but not in effect

What happened
A context cap was configured for a coding client. Requests still sent far larger prompts and timed out after 15 minutes.
Fix
Read the engine log for the real prompt size on each request before touching anything else.
Lesson
A configured value is a claim. The log line is the evidence.
2026-09

The embedder evicted the chat model

What happened
With a cap of two resident models, a knowledge search loaded the embedding model and pushed the chat model out. The next answer paid a 35 to 70 second reload.
Fix
Route definition questions to a full-text lookup that loads no model. Use vector search when a passage and citation are needed.
Lesson
Count every model your request path can load, including the small ones.
2026-07

Latency blamed on model speed

What happened
Median latency was 26 seconds and p90 was 89 seconds. The generation rate was a healthy 72 tokens per second.
Fix
Cap default output tokens, stream, and extend keep-alive.
Lesson
Latency is tokens divided by rate. Check the numerator first.

The same idea on other platforms

PlatformHow this module maps
DatabricksModel Serving replaces the engine. Foundation Model APIs give pay-per-token endpoints, and provisioned throughput reserves capacity. Your memory budget becomes a throughput and cost budget. The endpoint is OpenAI-compatible, so the contract from step 4 carries over.
IBM watsonxwatsonx.ai hosts foundation models behind an inference API, with on-demand deployments for dedicated capacity. Model choice is a catalog decision. You still measure tokens per second and time to first token per model.
CodexCodex is a client of an inference endpoint. It uses hosted models by default and can be pointed at another provider in its config, including a local OpenAI-compatible server. The serving work stays on whatever machine hosts the model.
CursorCursor consumes inference. It can override the OpenAI base URL to reach your own endpoint for some features. Serving decisions are made wherever that endpoint runs.
Claude Code / Agent SDKClaude Code and the Agent SDK call a hosted model, or a compatible gateway you configure. On the reference build a translation gateway lets the coding client talk to local models. Context budgets still apply and are charged on every turn.
Another machineOn an NVIDIA machine use vLLM or llama.cpp in place of the Mac runtimes. The budget is VRAM per card. Use systemd in place of launchd. Everything from step 3 onward is identical.

Explain it back

Answer aloud first. Then open the answer and compare.

Why does the memory budget come before model selection?
A strong answerBecause residency is what makes responses fast, and residency is limited by memory. A lineup chosen without the budget can require a model swap on a common path, and each swap costs tens of seconds to minutes.
A request is slow. What do you look at first, and why?
A strong answerThe engine log for that request: whether a model was loaded, the real prompt size, and the output token count. Those three explain almost all latency, and each has a different fix.
Why are there two routing mechanisms on the reference build?
A strong answerThe main engine serves many models on one port and selects by the model field. The second runtime serves one model per process and has no model selection, so reaching it means calling a different base URL.
You could build retrieval before this module. What would you stub?
A strong answerThe embedding call. Any function that returns a fixed-length vector for a string lets you build and test ingestion and search plumbing. You replace it with the real embedder later.

Terms

ResidentLoaded in memory and ready to generate without a load delay.
Keep-aliveHow long an idle model stays resident before the engine unloads it.
TTFTTime to first token. What the user feels as responsiveness.
QuantizationStoring weights at lower precision to cut memory, with some quality cost.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module