Map / Outline
Memory
What the agent keeps about a user and about its own past work, and how it gets back into the prompt.
What it is and why it exists
What
Memory is durable information written during or after conversations and read back in later ones. It includes per-user facts, summaries of past threads, and rules the agent learned from its own mistakes.
Why
Without memory every conversation starts cold. With careless memory the agent leaks one user's context into another's, or carries forward a wrong belief forever.
How it works
- Each caller has a partition key. All reads and writes are scoped to it.
- After a response, a background task extracts what is worth keeping and writes it. A failure there never affects the response.
- At prompt composition, a short memory section is built for the partition and added to the system prompt.
- A recall tool searches past conversations on demand, so history does not have to sit in every prompt.
- Memory is tested: a memory eval asks questions whose answers were stated in earlier threads.
Where it sits in the build order
Needs first
- Harness engineeringMemory is read at prompt composition and written after the response. Both are harness stages.Build out of order Stub it with: A JSON file per user, read at start and appended at end.
- State and storageMemory is durable state with concurrent writers: live chats and background jobs. It needs the storage guarantees from the state module.Build out of order Stub it with: A single-writer file store, for one user only.
Unlocks
- Self-evolving agentsLearned claims and rules are a form of memory. They use the same store, status fields, and recall path.
In the reference build
| Path | Role |
|---|---|
| apps/agent-server/memory_store.py | Per-partition reads and writes. |
| apps/agent-server/memory_pipeline.py | Post-response extraction, run in the background. |
| apps/agent-server/tools/conversation_recall.py | Search past threads on demand. |
| evals/memory/ | Memory eval sets. |
The same idea on other platforms
| Platform | How this module maps |
|---|---|
| Databricks | Use Lakebase or Delta tables keyed by user, read in your agent code. Long-term memory patterns ship with the agent framework and LangGraph stores. |
| IBM watsonx | Orchestrate keeps conversation context. Long-term memory is a store you connect and read through a tool. |
| Codex | Persistent instructions live in AGENTS.md. Memories are files the agent is told to read and update. |
| Cursor | Rules and memory features hold persistent preferences. Project facts belong in repo files. |
| Claude Code / Agent SDK | A project instruction file plus memory files and tools. The pattern is the same: write after, read at start, scope by project. |
| Another machine | A table with a partition column and a summarizer. Portable. |
Explain it back
Answer aloud first. Then open the answer and compare.
Why is the memory write fire-and-forget?
A strong answerThe user is waiting on the answer, not on bookkeeping. A slow or failed write should cost nothing in latency or correctness.
What stops one user's memory from reaching another?
A strong answerA partition key applied in the store on every read and write. A prompt instruction alone would not be a boundary.
From the live build
Recent changes and files the sync job filed under this module.
- Memory store on Postgres, second attempt (qualified upsert columns); recall skips the current thread, memory briefing cannot fail a request
- Memory store on Postgres (mem schema), recall skips the current thread, memory briefing cannot fail a request
- Recall similarity floor 0.35 (eval: lowest hit 0.55, highest unrelated 0.24)
- Conversation recall: mem.turns index, recall_conversations tool, recall gate in the daily cycle
- Self-evolution: what was broken, what changed
- Issues 7, 8 and 9 — diagnosis and fix
- Local Agent Server — Build Guide v2
- the quarantined learned-knowledge store (BRAIN-ARCHITECTURE.md §3.1). Deliberately modelled on memory_store.py, right down to the conventions: contextmanager connection, schema-if-not-exists, 0600 permissions, PRAGMA journal_mode=TRUNCATE, and idempotent additive migrations via ...
- the agent-facing wrapper around claim_recall. REGISTERED 2026-08-18, AFTER IT EARNED IT. ------------------------------------------ This file sat unregistered until the T5 discrimination drill said it was worth registering.
- the agent can look up what was said before (2026-10-01). Thin wrapper over conversation_index.search(). Registered after nightly/recall_eval.py passed its gate (see evals/memory/history.jsonl).
- Step 3.6 (AGENT_SERVER_GUIDE_v2.md §4/§7.1): the durable episodic + user-profile memory layer.
- memory_store on Postgres behaves like it did on SQLite (2026-10-01). Runs the same calls against both backends under a throwaway partition, compares what comes back, removes its rows. Exits non-zero on a difference.
- copy memory.db into Postgres schema mem (2026-10-01). Idempotent. agent:self observations are collapsed to the newest row per distinct finding (10,998 rows were 26 findings re-logged hourly). venv/bin/python3 ../ops/migrate_memory_pg.py # migrate + verify
- does conversation recall find the right exchange, stay quiet when nothing fits, and keep users apart? (2026-10-01) Seeds two synthetic users (partitions eval:recall-a / eval:recall-b) straight into mem.turns, asks paraphrased questions, removes the seed rows.
- the missing wire between 144,386 learned atoms and a live answer, plus the instrumentation that proves whether it helps.