Map / Outline
Operations and deployment
Supervision, scheduling, logs, backups, and recovery, so the system runs when nobody is watching.
What it is and why it exists
What
Operations covers how services start and restart, how background work is scheduled around live traffic, how logs are kept readable, how data is backed up and restored, and how changes are rolled out and rolled back.
Why
An agent that works in a terminal session and dies on reboot is a demo. Most of the reference build's lost time came from operational causes: a stripped service environment, a noisy probe, single-writer stores, backups on the same volume they protected.
How it works
- Every long-running service is under the OS supervisor with keep-alive and absolute paths. Each service uses its own virtual environment.
- A watchdog probes liveness. A status command reports every service, port, and model in one screen.
- A scheduler gives background work a duty cycle and preempts it when a user request arrives.
- Logs rotate on a schedule. Structured request and tool logs are separate from server logs.
- Backups copy what cannot be reproduced: databases, keys, prompts, skills, scripts. Weights and source documents are re-downloadable. One copy goes off the machine, encrypted, and restores are tested.
- Before editing a file, a dated copy is saved beside it. Each change is one commit whose message states what was learned.
Where it sits in the build order
Needs first
- Inference engineeringThe engine is the first service to supervise, and its memory limits decide what can run at the same time.Build out of order Stub it with: Start services by hand in terminal tabs.
- Harness engineeringThe agent server is the main service. Its health endpoint and busy marker are what the supervisor and scheduler read.Build out of order Stub it with: Supervise the engine alone.
- EvaluationScheduled evals and the daily report are how you learn that a deploy made things worse.Build out of order Stub it with: A manual smoke test after each change.
Unlocks
Nothing depends on this. It is an end point of the map.
In the reference build
| Path | Role |
|---|---|
| Install-Autostart.command | Installs supervised services. |
| Check-All-LLM-Status.command | One-screen status of every service and model. |
| apps/agent-server/nightly/scheduler.py | Duty cycle and preemption. |
| apps/ops/pg_backup.sh | Database backup, with a restore test beside it. |
| SOP-LLM-OPERATIONS.md | Operating procedures. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Serving endpoints, Jobs, and Apps are supervised by the platform. You own versions, permissions, cost alerts, and promotion between workspaces with asset bundles. |
| IBM watsonx | The platform runs the services. You own environments, deployment spaces, and promotion between them. |
| Codex | Operations here means CI: headless runs, sandbox settings, and secrets handling for the agent. |
| Cursor | Same: team rules in the repo, background agent settings, and secrets. |
| Claude Code / Agent SDK | Headless runs in CI, managed settings, and permission policies. For a service built on the SDK, operate it like any other service. |
| Another machine | systemd units and timers in place of launchd. The backup and status scripts are shell and Python. |
Explain it back
Answer aloud first. Then open the answer and compare.
Which files do you back up, and which do you not?
A strong answerBack up what cannot be reproduced: databases, keys, prompts, skills, eval sets, scripts. Skip model weights and source documents that can be downloaded again.
A service runs fine by hand and fails under the supervisor. What is the usual cause?
A strong answerEnvironment. The supervisor starts with a minimal PATH and no shell profile, so a binary or virtual environment is not found. Use absolute paths in the service definition.
From the live build
Recent changes and files the sync job filed under this module.
- Gateway answers the watchdog probe; guard: no-tool check on data turns only, strict first draft; wiki samples start, middle and end; collector tuner floor and restored cadences
- Chain audit: files removed from disk are counted separately, not as awaiting ingest
- Daily report: unreadable learning store is one explicit problem, never zeros; backfill status merges per collection; collector pause expires before the same-hour run
- Audit fixes: self-gap dedupe, SQLite stores in the nightly backup; agent quiz
- Fixes applied — 2026-09-07
- AI Local Server → Vercel Integration Guide
- SOP — Operating the Local LLM Server
- AIDATA Local AI Server — Master Guide (Phased Plan)
- BUILD-HOST — Operations Runbook
- 06 — Improvement Roadmap
- run the agent-server regression scripts (2026-09-30). Each is a plain script that exits non-zero on failure. Writes logs/tests.status.json. Safe: the brain test uses a throwaway <db>_test database and the endpoint tests use /tmp and stub model servers.
- nightly dump of the knowledge store (2026-09-30). WHAT: pg_dump --format=custom of the whole build-host database (brain.* learned atoms, kb.* file ledger, kb_finance / kb_k12 chunk vectors).
- Backup-AI-Offsite.command [destination] [full] Writes ONE encrypted file containing everything on this Mac that cannot be recreated, to somewhere that is not this Mac. WHY THIS EXISTS AND Backup-AI.command DOES NOT REPLACE IT.
- pg_restore_test.sh <dumpfile> — prove a dump restores (2026-09-30). Restores brain.*, kb.* and kb_k12.* from the dump into a throwaway database (build-host_restoretest), compares row counts with the live database, checks a K12 vector search works on the restored copy, then ...
- let Backup-AI-Offsite.command run unattended (2026-09-30). Targeted edits, each asserted to match exactly once: * pause(): `read -k 1` becomes a no-op when NONINTERACTIVE is set * passphrase from the login Keychain when OFFSITE_KEYCHAIN_SERVICE is set (item created only by ...
- Double-click button: stops and restarts the NEW agent-server harness (apps/agent-server/main.py, FastAPI + agent_loop.py, port 8788).
- Install-Autostart.command GOAL: after a power cut, a reboot, a sleep/wake, or a crash, every LLM service comes back BY ITSELF. No double-clicking anything. WHAT IS ALREADY TRUE (since the 2026-08-09 migration): six services are supervised by launchd with KeepAlive.
- REWRITTEN by Install-Autostart.command. The services this script used to nohup are now launchd jobs, so starting them means LOADING them, not spawning a second copy.