The Self-Improving Agent Test Kit — Test Any AI Agent's Memory Claims Before You Trust Them
The companion to the video. 'Self-improving' is the pitch on a wave of new AI agent harnesses — so we ran a preregistered, controlled experiment on one (Prime Agent v0.7.1, gpt-5.2 in both arms): a persistent LEARNER session that kept its stored lessons across five rounds versus a fresh-session RESET control, graded by hidden deterministic tests the agent never saw, with a blind transfer round at the end. The fresh-session control scored 38/38. The persistent arm scored 12/38 — it answered rounds 2-5 from memory in 3-5 seconds without touching a file, once claimed it 're-checked the README' with zero tool calls on record, and its own stored lessons correctly diagnosed the failure three times without ever changing its next answer. THIS KIT is the full protocol so you can run the same audit on any agent: (1) THE LEARNER-VS-RESET DESIGN, diagrammed — two lanes, five rounds, hidden graders, blind final round. (2) THE FIVE-ROUND BENCHMARK TEMPLATE — visible happy-path tests, 6-10 hidden spec tests, validated answer keys, the no-leak vocabulary grep, and the five transferable trap classes. (3) THE COPYABLE GRADING RUBRIC — hidden-tests-passed as the primary metric, with the exact preregistered definitions of repeated errors and regressions, verbatim. (4) THE FAILURE-REPORT TEMPLATE we fed the agent, verbatim — honest failure data with zero coaching. (5) THE SAFE REFINEMENT LOOP — fresh session per task, external grader, review refinements like code review, canary re-runs with rollback. (6) FOUR BAD-LESSON WARNING SIGNS — including sub-10-second DONEs and the playbook that grows while scores fall. (7) INSTRUCTION-BLOAT CHECKS with our token receipts (66,452 vs 53,705 input tokens; state growing 121KB to 181KB while performance sat at 4/30). (8) THE WHEN-MEMORY-HELPS-VS-HURTS DECISION TABLE. (9) THE ISOLATION AND ROLLBACK CHECKLIST — these harnesses run model-generated code with your user permissions, not in a sandbox. (10) OUR FULL MEASURED RESULTS — the complete scoreboard plus the agent's stored lessons, quoted verbatim, and the honest limitations. Independent test kit from Hyperautomation Labs — not affiliated with Prime Intellect.
Free. No spam. Unsubscribe anytime.