Map / Outline
API and gateway
One contract for every caller, with keys, scopes, limits, and a single public door.
What it is and why it exists
What
The API is the contract callers use: the OpenAI-compatible chat, embeddings, and models endpoints. The gateway is the process in front that authenticates each caller and decides what that caller may do.
Why
The model engine has no authentication. Every outside caller needs an identity that can be limited and revoked without affecting the others. A stable contract also lets you swap what is behind it.
How it works
- Keys are stored hashed. Each key has an app name, scopes such as chat and tools, a rate limit, and a quota.
- The gateway validates the key, checks scope and limits, then forwards. The engine stays on localhost.
- Public access goes through an encrypted tunnel to the gateway port only. No router port is forwarded.
- A translation gateway maps one provider's API shape to another, which lets a coding client built for a hosted API talk to local models.
- Health has two levels. An unauthenticated probe answers only whether the process is serving. Details require a key.
Where it sits in the build order
Needs first
- Inference engineeringA gateway forwards to an engine. Its timeouts and limits are set from measured load times and generation rates.Build out of order Stub it with: An echo handler that returns a fixed completion.
Unlocks
- Tools and MCPTool calling rides on the chat contract: tool schemas go in the request and tool calls come back in the response. Key scope decides who gets tools.
- Harness engineeringThe harness sits behind the gateway contract. Keys carry scopes, and the harness reads the scope to decide whether a caller gets tools at all.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/auth.py | Key validation, scopes, rate limits, quotas. |
| apps/agent-server/keys_admin.py | Issue, list, and revoke keys. |
| apps/agent-server/main.py | The endpoints. |
| apps/claude-code-gateway/ | Translation proxy so a coding client can use local models. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Model Serving endpoints are the API. AI Gateway adds rate limits, usage tracking, guardrails, and routing across providers. Identity is the workspace's tokens and service principals. |
| IBM watsonx | watsonx.ai exposes inference endpoints behind IAM keys. A model gateway routes to third-party providers. Orchestrate exposes agent endpoints. |
| Codex | Codex is a client. It needs a provider base URL and key. If you point it at your gateway, issue it its own key. |
| Cursor | A client. Give it a key of its own and a base URL override where supported. |
| Claude Code / Agent SDK | A client. A base URL setting can route it through your gateway. Issue a separate key per machine. |
| Another machine | A small FastAPI app or an off-the-shelf LLM proxy does the same job on any host. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why one key per app?
A strong answerSo you can limit and revoke one integration without touching the others, and so logs attribute every request.
Why is liveness unauthenticated while details are not?
A strong answerA supervisor has to tell healthy from sick without holding a key. Details reveal configuration, so they stay behind one.
From the live build
Recent changes and files the sync job filed under this module.
- the spending dataset collector: fiscal-year scope never starts later than FY2021
- Answer guard: vendor names, answers given without a tool call, streaming gate; FINANCE on Qwen3 embeddings; wiki watcher batch cap, lock and wait; downloads script
- SOP — Operating the Local LLM Server
- Claude Code → your own models
- Bug Fixes — 2026-08-20, post-update
- Dynamic Gateway + Full-Capability Model Lineup
- 00 — Project Timeline: How build-host Got Built
- Stage 0 patches — ready to apply, deliberately NOT applied
- sandbox test suite for Step 5.7 (streaming + latency). Run: python3 test_streaming.py Tests against a REAL stub Ollama HTTP server (a live uvicorn process on a real port, speaking real SSE), not mocks — same discipline every prior step in this codebase used.
- Agent Server — apps/agent-server/main.py ========================================= STEP 2 of the "mini Claude Code" build (see AGENT_SERVER_GUIDE_v2.md). Step 1 — bare passthrough, plaintext-key auth. Built/tested 2026-07-27.
- end-to-end test through main.py's REAL FastAPI app. test_streaming.py proves agent_loop's streaming logic.
- what the chat box actually does (2026-10-06). chat_eval.py and agent_quiz.py call agent_loop.run(), the non-streaming path. The chat box streams.
- run the agent-server regression scripts (2026-09-30). Each is a plain script that exits non-zero on failure. Writes logs/tests.status.json. Safe: the brain test uses a throwaway <db>_test database and the endpoint tests use /tmp and stub model servers.
- measures the streaming win instead of asserting it. Replays THIS server's own real production numbers (from logs/agents/requests.jsonl: p50 26.3s, p90 88.7s, 72 tok/s effective) against a stub that generates at that same measured rate, and reports time-to-first- token vs total ...
- LocalAPI-Claude.command rev 2026-08-22b The one command for the Claude Code local API. Start, restart, or re-run after pulling a model. Idempotent.
- QUICKSTART-10-MIN.command Double-click this. It does everything: 1. installs litellm (if missing) 2. generates a key 3. writes a minimal config pointed at whichever model you actually have 4. starts the proxy in the background 5. tests it end to end 6.