← all systems
AGENT INFRAAI AGENTEVALS

Learning-Loop — Self-Improving Agent Memory

Agents that log lessons, retrieve them at the right moment, and measure whether the memory actually changed behavior.

Python · Claude Code Hooks · 144-test suite · Weekly cron · IDF-ranked retrieval

The Problem

Most "agent memory" counts retrievals and calls it learning. Retrieving a lesson is not the same as using it. I built Learning-Loop to close that gap: a drop-on engine for any Claude Code project or managed agent that injects memory at startup, retrieves at targeted stage checkpoints with IDF ranking, and — the part that matters — records confirmed-use telemetry: did the retrieved memory change what the agent did? Lessons that get retrieved forever without a confirmed use decay out. Lessons that prove out graduate from an episodic log into a pattern library, into the project contract, and finally into enforcement code that cannot be skipped.

Stack

🐍
Python
22 domain-neutral scripts; idempotent install.py; cross-project host registry
🪝
Claude Code Hooks
PostToolUse marks the session; Stop blocks until closeout runs
🧪
144-test suite
Run by the weekly audit gate
⏰
Weekly cron
Transcript mining + curation; risky changes staged for human review
🔍
IDF-ranked retrieval
Targeted stage checkpoints, not a firehose at startup

How It Works

The execution path

01Inject at Startup
→
02Retrieve at Checkpoint
→
03Confirm Use
→
04Graduate or Decay
Execution flow
Session start: memory injected— project contract + top patterns
Stage checkpoint: IDF-ranked retrieval of relevant lessons
Confirmed-use telemetry: did the retrieval change behavior?— not "was it retrieved"
Hook-enforced closeout: learning decision + use decision required
Weekly: mine transcripts for unlogged lessons; stage risky changes for one review file
Graduate: episodic log → pattern library → CLAUDE.md contract → enforcement code
Decay: retrieved-forever, never-used lessons fall out
A closeout gate that actually gates. A PostToolUse hook marks the session on first edit; a Stop hook blocks until closeout runs; closeout requires both a learning decision and a use decision. A weekly cron mines transcripts for lessons nobody logged and stages anything risky for one human review file. A fresh install starts bare on purpose — the loop's own instrumentation should tell you what its categories should be. Memory that adopts its own structure unreviewed is how loops poison themselves.

Key Design Decisions

📏
Measure Use, Not Retrieval
Confirmed-use telemetry records whether a retrieved memory changed the agent's action. Counting retrievals is vanity.
🚪
A Gate That Gates
Stop is blocked until closeout runs, and closeout requires decisions, not acknowledgements. No session ends unrecorded.
🎓
Lessons Graduate
Proven lessons move from log to pattern to contract to code that cannot be skipped. Unused ones decay.
🌱
Start Bare on Purpose
A fresh install has no categories. The loop's own data decides them, and a human reviews before anything sticks.

By The Numbers

22
Domain-neutral scripts
144
Tests in the audit gate
1
Human review file per week
0
Sessions that end without closeout
← back to all systemsmatthew batterson · gtm engineer