Agent Build Tutor
Map / Outline

Self-evolving agents

A scheduled loop that studies sources and its own logs, promotes what is corroborated, and is graded on sealed tests.

What it is and why it exists

What

A self-evolving agent improves between conversations. On the reference build a nightly learner reads documents and logs, writes candidate claims, verifies and corroborates them, and promotes the survivors into what the agent can recall.

Why

Most improvement comes from noticing the same failure repeatedly and encoding the fix. Automating that is powerful and dangerous. The loop changes no model weights. It changes stored knowledge, routing exemplars, and rules, and all of it reaches an answer only through a tool call.

How it works

Where it sits in the build order

Needs first

  • EvaluationA loop optimizes what it can measure. Sealed evals have to exist before the loop runs, or it grades itself.Build out of order Stub it with: None that is safe. You can run the loop in report-only mode without evals, and you should not let it apply changes.
  • MemoryLearned claims and rules are a form of memory. They use the same store, status fields, and recall path.Build out of order Stub it with: Append findings to a Markdown file for human review.
  • Guardrails and verificationPromotion gates and claim verification are guard logic applied to stored knowledge.Build out of order Stub it with: Human approval of every promotion.
  • RAG and knowledge graphThe learner reads indexed chunks and writes claims that point back to passages.Build out of order Stub it with: Feed it plain text files.

Unlocks

Nothing depends on this. It is an end point of the map.

In the reference build

PathRole
apps/agent-server/nightly/scheduler.pyDuty cycle and preemption for background work.
apps/agent-server/nightly/self_observe.pyFinds confident wrong answers in the logs.
apps/agent-server/nightly/conversation_learn.pyLearns from real conversations, ignoring eval threads.
apps/agent-server/nightly/corroboration_audit.pyChecks that support is independent.
SELF-EVOLUTION-LESSONS.mdWhat went wrong and what was changed.

The same idea on other platforms

PlatformHow this module maps
DatabricksLakeflow Jobs schedule the loop. Traces and labeled feedback are the raw material. Evaluation datasets grow from production traces, and judges can be aligned with human labels.
IBM watsonxSchedule jobs outside Orchestrate and write results to a governed store. Monitors in watsonx.governance provide the drift signals.
CodexA scheduled headless run can review logs and propose changes to AGENTS.md or skills as a pull request for review.
CursorBackground agents can do the same review and open a pull request.
Claude Code / Agent SDKA scheduled headless session can read logs and propose skill or instruction edits. Keep a human merge step.
Another machineA scheduler such as launchd, systemd timers, or cron, plus the same scripts.

Explain it back

Answer aloud first. Then open the answer and compare.

What does the learner change, and what does it never change?
A strong answerIt changes rows in the knowledge store, routing exemplars, and rule files. It never changes model weights. Learned knowledge reaches an answer only when a tool returns it.
Why is corroboration counted by independent document and by day?
A strong answerTen copies of one paragraph are one source. Two failures in one session are one event. Counting them as many would let repetition pass for evidence.

From the live build

Recent changes and files the sync job filed under this module.

Ask the tutor about this module