Pith. sign in

hub Canonical reference

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Canonical reference. 80% of citing Pith papers cite this work as background.

34 Pith papers citing it
Background 80% of classified citations
abstract

Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we find that it can reduce response length while degrading performance. We trace this degradation to the suppression of epistemic verbalization - the model's expression of uncertainty during reasoning. Through controlled experiments varying conditioning context richness and task coverage, we show that conditioning the teacher on rich information suppresses uncertainty expression, enabling rapid in-domain optimization with limited task coverage but harming OOD performance, where unseen problems benefit from expressing uncertainty and adjusting accordingly. Across Qwen3-1.7B/8B, DeepSeek-Distill-Qwen-7B, and Olmo3-7B-Instruct, we observe performance drops of up to 40%. Our findings highlight that exposing appropriate levels of uncertainty is crucial for robust reasoning and underscore the importance of optimizing reasoning behavior beyond merely reinforcing correct answer traces.

hub tools

citation-role summary

background 4 other 1

citation-polarity summary

years

2026 34

polarities

background 4 unclear 1

representative citing papers

Rethinking On-Policy Self-Distillation for Thinking Models

cs.AI · 2026-07-06 · conditional · novelty 7.0

Privileged-context on-policy self-distillation degrades thinking models' long-budget accuracy by suppressing forking and self-correction behaviors, while helping instruction-tuned models.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

cs.CL · 2026-06-16 · unverdicted · novelty 7.0

ZPPO improves distillation to small vision-language models by using binary and negative candidate prompts plus a replay buffer for hard questions, outperforming standard distillation and GRPO on a 31-benchmark suite with largest gains at the 0.8B scale.

OPRD: On-Policy Representation Distillation

cs.LG · 2026-06-04 · unverdicted · novelty 7.0

OPRD performs distillation in hidden-state space on on-policy data for deterministic gradients and better math benchmark performance, plus OPRD-Bridge for cross-architecture transfer via low-rank projectors.

Learning from Language Feedback via Variational Policy Distillation

cs.LG · 2026-05-14 · unverdicted · novelty 7.0

VPD frames language feedback learning as variational EM so the teacher policy refines itself via trust-region updates on outcomes while the student learns dense token distributions on its own rollouts, outperforming fixed-teacher baselines on reasoning and code tasks.

KL for a KL: On-Policy Distillation with Control Variate Baseline

cs.LG · 2026-05-08 · unverdicted · novelty 7.0

vOPD stabilizes on-policy distillation gradients by subtracting a closed-form per-token negative reverse KL baseline as a detached control variate, preserving unbiasedness while lowering variance and matching expensive full-vocabulary methods.

DemoPSD: Disagreement-Modulated Policy Self-Distillation

cs.LG · 2026-07-02 · conditional · novelty 6.5

Disagreement-modulated reverse-KL barycenter targets let on-policy self-distillation attenuate privileged leakage while preserving exploration, beating SDPO and GRPO on SciKnowEval and GPQA.

Geometric Self-Distillation for Reasoning Generalization

cs.LG · 2026-07-07 · accept · novelty 6.0

A geometric self-distillation objective using Hellinger loss and Fisher-Rao proximal regularization prevents predictive drift in LLM post-training, improving out-of-distribution reasoning by 5.7-8.6 points.

Weak-to-Strong Generalization via Direct On-Policy Distillation

cs.LG · 2026-07-06 · conditional · novelty 6.0

Transferring a weak model’s RL-induced log-ratio policy shift on a strong student’s own rollouts raises AIME accuracy more cheaply than imitating the weak teacher or running matched-step RL on the student.

citing papers explorer

Showing 34 of 34 citing papers.