Across the cited comparisons summarized in a critical review of on-policy self-distillation for mathematical reasoning, benchmark scores, functional diversity - the range of useful alternatives a model produces - and training cost do not move together. In one comparison, pass@1, the one-attempt measure, rose while pass@16, the 16-attempt measure, fell; in another, fewer generated tokens came with a longer optimization step. The review presents these as trade-offs in reported comparisons, not as causal findings from a new experiment.
The paper is a structural, non-experimental critical review. It synthesizes terminology and separates claims treated as settled from those that remain disputed. The review is non-exhaustive and centered on mathematical reasoning; it does not cover the multimodal and tool-using branches.
Three places the review looks for collapse
In the OPSD loop described by the review, the model samples an on-policy rollout and then uses the same model as a teacher after giving it privileged information - information unavailable to the student. The teacher scores the student's tokens with forward KL, a token-level comparison of their predicted distributions; only the student is updated, while the teacher stays fixed.
That loop gives the review three places to look for collapse: how the token-level signal is placed or weighted, what privileged information the teacher receives, and when the teacher or its guidance changes. This three-lever framework is the review's organizing device, not a new empirical test.
The choice of divergence also split the outcomes in one cited comparison. Reverse KL had the best performance but the lowest diversity; forward KL showed the opposite pattern, and JSD was intermediate. These rankings belong to that comparison, and the review does not present them as a universal order.
The numbers pull apart
The performance figures come with a checkpoint qualification. A forward-KL OPSD comparison reported an AIME25 rise from 36.7 to 43.9. For Qwen3-1.7B, the review lists best-over-checkpoint scores of 57.2% on AIME24, 43.9% on AIME25 and 29.2% on HMMT25, while the end-of-training AIME25 score was 41.1%. The review reports no confidence interval.
On Qwen3-8B, another cited comparison showed pass@1, the one-attempt measure, rising from 71.9 to 73.4 while pass@16, the 16-attempt measure, fell from 83.6 to 78.5. The review interprets that pairing as higher mean success alongside lower functional diversity. It also says token entropy was not a valid stand-in for functional diversity.
Efficiency pointed in a different direction. GRPO sampled eight rollouts of up to 16,000 tokens per problem, whereas OPSD used one generation capped at 1,024 tokens. At comparable performance, OPSD used fewer generated tokens, but its equalized-budget step took 20.6 seconds, compared with 11.2 seconds for GRPO. The comparison separates token efficiency from optimization compute.
What the teacher sees matters
In a cited thinking-model comparison, privileged information was associated with a relative decrease of up to 17% in avg@16, a measure based on 16 attempts. The review says the direction depended on the model type and on how much information was shown.
Yet a separate strict self-distillation comparison reported an error-aligned critique scoring 5.27 points above reference-solution OPSD and 16.11 points above GRPO on avg@12, a measure based on 12 attempts. In a different distillation regime, a five-form comparison reported final-answer information at 59.5 versus 63.0 with no privileged information, while step-wise hints without execution scored 71.3 on C-Eval. Taken together, the cited comparisons report different results for different information designs.
Signal and timing stay open
The question is not only what the teacher sees, but where the learning signal is applied. A cited oracle analysis estimated that distilling only positively aligned tokens could improve the signal by a factor of 10 to 15. The estimate was made outside training and was not converted into an effective training gain.
An entropy-based comparison offered a more specific warning. Over training, the student's most-probable token changed 84 times on high-entropy tokens, compared with seven changes in the low-entropy regime. High-entropy tokens represented 6.8% of the student's tokens versus 18.5% of the teacher's. The review says entropy was not a valid proxy for functional diversity.
Timing is the third lever. Adaptive teacher exposure varied how much of the reasoning trace the teacher saw according to student progress. Reported gains ranged from 0.95 to 2.33 on avg@12 across Qwen3-1.7B, 4B and 8B. The review leaves long-run stability and the best schedule unresolved.
What remains unsettled
Capability retention remains unsettled. One cited study found less forgetting with self-distillation, while another found that dense self-distillation forgot more than sparse reinforcement learning and could collapse. The review therefore identifies conflicting reports, not a settled direction for forgetting.
That caution is central to the preprint's message. The document is an arXiv preprint and a structural review, not a new experiment. Its comparisons do not establish one best design. The review separates fewer generated tokens from lower optimization compute and treats a higher pass@1 result as compatible with lower functional diversity. The three design questions it highlights are signal placement, privileged-information content and teacher timing.
Paper data and sources
Original title: One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
Authors: Justin Robert, Raheel Qader
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text