Inference engineering
Serve models on fixed hardware and know what each second costs.
What it is and why it exists
What
Inference engineering is the work of turning model weights into a dependable endpoint. You choose which models run, how they share memory, how long they stay loaded, how many requests run at once, and how much context each request may use.
Why
Every other layer calls this one. Retrieval calls it to embed. The harness calls it to plan and to answer. The evaluator calls it to judge. If a single call takes 70 seconds because a model was evicted, no prompt change above it will fix that. You set the speed limit of the whole system here.
How it works
- A serving engine loads weights into memory and exposes an HTTP API. On the reference machine the engine is Ollama, with a second runtime (MLX) for one small fast model.
- Memory is the binding constraint. The reference machine has 96 GB of unified memory and about 78 GB usable for models. Two mid-size models fit together. The largest model fits only alone.
- The engine keeps a model resident for a keep-alive window. A request for a model that is not resident pays a cold load. Loading a third model evicts one of the two that were resident.
- Latency after load is mostly output tokens divided by tokens per second. On the reference build, measured p50 latency was 26 seconds at 72 tokens per second, which matches a median answer of about 1,700 tokens almost exactly.
Where it sits in the build order
Needs first
Nothing. You can start here.
Unlocks
- API and gatewayA gateway forwards to an engine. Its timeouts and limits are set from measured load times and generation rates.
- RAG and knowledge graphSearch needs an embedding model at query time and at index time. The vector column width is fixed by that model, so the model choice comes first.
- Harness engineeringThe planner is a model call. Round budgets, output caps, and prompt size limits are set from measured inference numbers.
- Operations and deploymentThe engine is the first service to supervise, and its memory limits decide what can run at the same time.
In the reference build
| Path | Role |
|---|---|
| ollama-env.sh | Serving settings: loaded-model cap, keep-alive, parallelism, context length. |
| apps/agent-server/model_router.py | Picks a model per request by task type. Two mechanisms: tag swap and base-URL swap. |
| apps/agent-server/model_catalog.json | The lineup: which model serves default, fast, and vision work. |
| apps/agent-server/bench_latency.py | Measures time to first token and tokens per second from real requests. |
| apps/mlx/ | Second runtime. One process serves exactly one model. |
Process map
What happens to one request at the engine
- 1Request arrivesOpenAI-style JSON with a model name, messages, and max tokens. (input)
- 2Is the model resident?If yes, skip to prompt evaluation. If no, the engine must load it and may evict another. (decision)
- 3Cold loadTens of seconds for a 20 to 60 GB model. This is the largest avoidable cost. (storage)
- 4Prompt evaluationCost grows with prompt size. Tool schemas and system prompts count on every turn. (model call)
- 5GenerationOutput tokens at a steady rate. Time is tokens divided by rate. (model call)
- 6Stream outTokens leave as they are produced, so the user sees the first one early. (output)
Build steps
Each step states why it sits at this point. Open any step on its own. The first is open.
01Write the memory budget before you choose models
List usable model memory. On Apple silicon that is unified memory minus what the OS and your services need. On a GPU server it is VRAM per card.
For each candidate model record the weight size at your chosen quantization, then add headroom for context. Longer context and more parallel slots both raise memory use.
Decide which models must be warm at the same time. On the reference machine the workhorse (22 GB) and the vision model (29 GB) stay warm together. The 61 GB reasoner evicts both, so it is reserved for hard problems where a multi-minute swap is acceptable.
usable model memory ~78 GB
workhorse 35B MoE q4 22 GB default, ~72 tok/s
vision 27B q8 29 GB image input only
reasoner 120B q4 61 GB runs alone
embedder 0.3 GB loads on demand, can evict a chat model
rule: workhorse + vision fit together (51 GB). Anything + reasoner does not.02Install the engine and put weights on a dedicated volume
Install the engine with your package manager. Point its model directory at the data volume before the first pull.
Run the engine under the OS supervisor so it restarts after a crash or reboot. That is covered in the operations module, and you can start it by hand until then.
brew install ollama
export OLLAMA_MODELS="$AI_DATA/models/ollama"
ollama serve &
ollama pull <workhorse-model>
ollama pull nomic-embed-text
ollama list03Set the serving parameters on purpose
Cap loaded models at what your budget allows. Set keep-alive long enough that normal gaps between requests do not unload the model.
Set parallel slots to the concurrency you expect. Each slot reserves context memory.
Set context length explicitly. Then confirm it took effect by reading the engine log, because a setting that is present is not proof that it is in effect.
export OLLAMA_MAX_LOADED_MODELS=2 # at most 2 models resident
export OLLAMA_KEEP_ALIVE=30m # unload after 30 idle minutes
export OLLAMA_NUM_PARALLEL=4 # concurrent requests per model
export OLLAMA_CONTEXT_LENGTH=49152 # tokens of context per slot
export OLLAMA_FLASH_ATTENTION=1 # less memory for long contexts04Call the endpoint the way every later layer will
Test with plain curl before you add any library. You want to know the engine works independent of your code.
Record tokens per second and time to first token for each model. These two numbers explain most latency complaints later.
curl -s http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<workhorse-model>",
"messages": [{"role": "user", "content": "Reply with the word ready."}],
"max_tokens": 16
}' | python3 -m json.tool05Add a second runtime only where it earns its place
The reference build serves a 9B model through MLX on its own port. That runtime serves exactly one model per process, chosen at startup.
This creates two routing mechanisms. Models in the main engine are selected by the model field (tag swap). The MLX model is selected by calling a different base URL.
# one process, one model, its own port
mlx_lm.server --model mlx-community/<small-model>-4bit --port 808206Route by task type, and fail safe
Classify the request: default, fast, or vision. Map each class to a catalog entry.
If anything is missing or unknown, return nothing and let the existing behavior run. The router should never be the reason a request fails.
def resolve_for_request(payload: dict) -> dict | None:
if payload.get("model"): # caller was explicit: do nothing
return None
task = detect_task(payload) # "default" | "fast" | "vision"
entry = CATALOG.get(task)
if not entry:
return None # unknown task: do nothing
if entry.get("base_url"): # separate runtime: swap the URL
return {"base_url": entry["base_url"]}
if entry.get("tag"): # same engine: swap the tag
return {"model": entry["tag"]}
return None # incomplete entry: do nothing07Cut latency where it is actually spent
Output volume dominates. Set a default max-tokens value and ask for shorter answers in the system prompt.
Stream responses so the first token arrives early. Reuse one pooled HTTP client across requests.
Keep prompts small. Tool schemas, system prompts, and memory are charged on every turn.
Avoid evictions. A retrieval call that loads the embedder can push out the chat model. Prefer the cheap path that loads no model when a definition will do.
08Keep the engine private
Bind the engine to localhost or a private network. Expose only a proxy that authenticates callers.
How it works with the other parts
- API and gatewayThe gateway is the only thing allowed to forward outside traffic to the engine.
- RAG and knowledge graphRetrieval uses the embedding model. Loading it can evict a chat model, so retrieval design affects inference cost.
- Harness engineeringEach tool round is another model call. Round budgets and prompt size are inference decisions made in the harness.
- EvaluationJudges and question writers are model calls too. A small second runtime lets evals run without evicting the serving model.
What went wrong in the real build
A context setting that was present but not in effect
- What happened
- A context cap was configured for a coding client. Requests still sent far larger prompts and timed out after 15 minutes.
- Fix
- Read the engine log for the real prompt size on each request before touching anything else.
- Lesson
- A configured value is a claim. The log line is the evidence.
The embedder evicted the chat model
- What happened
- With a cap of two resident models, a knowledge search loaded the embedding model and pushed the chat model out. The next answer paid a 35 to 70 second reload.
- Fix
- Route definition questions to a full-text lookup that loads no model. Use vector search when a passage and citation are needed.
- Lesson
- Count every model your request path can load, including the small ones.
Latency blamed on model speed
- What happened
- Median latency was 26 seconds and p90 was 89 seconds. The generation rate was a healthy 72 tokens per second.
- Fix
- Cap default output tokens, stream, and extend keep-alive.
- Lesson
- Latency is tokens divided by rate. Check the numerator first.
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Model Serving replaces the engine. Foundation Model APIs give pay-per-token endpoints, and provisioned throughput reserves capacity. Your memory budget becomes a throughput and cost budget. The endpoint is OpenAI-compatible, so the contract from step 4 carries over. |
| IBM watsonx | watsonx.ai hosts foundation models behind an inference API, with on-demand deployments for dedicated capacity. Model choice is a catalog decision. You still measure tokens per second and time to first token per model. |
| Codex | Codex is a client of an inference endpoint. It uses hosted models by default and can be pointed at another provider in its config, including a local OpenAI-compatible server. The serving work stays on whatever machine hosts the model. |
| Cursor | Cursor consumes inference. It can override the OpenAI base URL to reach your own endpoint for some features. Serving decisions are made wherever that endpoint runs. |
| Claude Code / Agent SDK | Claude Code and the Agent SDK call a hosted model, or a compatible gateway you configure. On the reference build a translation gateway lets the coding client talk to local models. Context budgets still apply and are charged on every turn. |
| Another machine | On an NVIDIA machine use vLLM or llama.cpp in place of the Mac runtimes. The budget is VRAM per card. Use systemd in place of launchd. Everything from step 3 onward is identical. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why does the memory budget come before model selection?
A request is slow. What do you look at first, and why?
Why are there two routing mechanisms on the reference build?
You could build retrieval before this module. What would you stub?
Terms
| Resident | Loaded in memory and ready to generate without a load delay. |
| Keep-alive | How long an idle model stays resident before the engine unloads it. |
| TTFT | Time to first token. What the user feels as responsiveness. |
| Quantization | Storing weights at lower precision to cut memory, with some quality cost. |
From the live build
Recent changes and files the sync job filed under this module.
- Reranker: cap and release the MLX buffer cache (process had grown to 71 GB)
- Finish the Update Today — remaining steps
- AI Local Server → Vercel Integration Guide
- Whole-System Audit and Revised Update Plan
- AI Stack Update Runbook — 2026-08-20
- 01 — System Architecture
- 02 — The Build Process, Step by Step
- Standalone reranker server for build-host agent-server's rerank_route.py. Added 2026-08-03.
- Double-click button: stops and restarts the NEW agent-server harness (apps/agent-server/main.py, FastAPI + agent_loop.py, port 8788).
- REWRITTEN by Install-Autostart.command. The services this script used to nohup are now launchd jobs, so starting them means LOADING them, not spawning a second copy.
- what is every local model doing RIGHT NOW. python3 llm_activity.py one snapshot, in the terminal python3 llm_activity.py --watch refresh every 5s until Ctrl-C python3 llm_activity.py --html also write $AI_DATA/llm-activity.html python3 llm_activity.py --html --watch keep ...
- Measure what parallel mining actually buys — 2026-09-10. The claim being tested is narrow and worth testing rather than assuming: mlx_lm.server answers one request at a time, so N servers should give close to N times the throughput UNTIL memory bandwidth saturates.
- copy-truncate rotation for AI_DATA logs. WHY COPY-TRUNCATE AND NOT MOVE: Ollama, agent-server and the MLX backends hold their log files open with O_APPEND.
- document -> verified audiobook (English + Chinese) $AI_DATA/apps/agent-server/venv/bin/python3 make_audiobook.py FILE All synthesis, normalization and quality checking lives in tts_engine.py.
- production text-to-speech for long documents (EN + ZH) ===================================================================== Replaces the naive "post a big chunk, hope it comes back" approach with the loop commercial audiobook pipelines actually use: normalize -> small units -> ...