Pith. sign in

REVIEW 4 cited by

Revisiting Self-Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.08491 v1 pith:GEOPWP6E submitted 2022-06-17 cs.LG

classification cs.LG
keywords self-distillationteacherstudentdistillationexplanationsknowledgemodelprocedure
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowledge distillation is the procedure of transferring "knowledge" from a large model (the teacher) to a more compact one (the student), often being used in the context of model compression. When both models have the same architecture, this procedure is called self-distillation. Several works have anecdotally shown that a self-distilled student can outperform the teacher on held-out data. In this work, we systematically study self-distillation in a number of settings. We first show that even with a highly accurate teacher, self-distillation allows a student to surpass the teacher in all cases. Secondly, we revisit existing theoretical explanations of (self) distillation and identify contradicting examples, revealing possible drawbacks of these explanations. Finally, we provide an alternative explanation for the dynamics of self-distillation through the lens of loss landscape geometry. We conduct extensive experiments to show that self-distillation leads to flatter minima, thereby resulting in better generalization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    CurioSFT preserves exploration during supervised fine-tuning by distilling toward the model's own temperature-scaled distribution and adaptively increasing entropy at high-entropy tokens, improving SFT accuracy by ~2....

  2. Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Mask Pooling, which pools foreground and background separately using ground-truth masks, improves object-detection robustness on domain-shift benchmarks but requires masks at inference and lacks a rigorous causal derivation.

  3. Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Checkpoint fusion guided by a 'forget' metric, followed by distillation, lets a single model recover test points that were learned and then forgotten during training, improving accuracy, especially under label noise.

  4. Smooth-Distill: A Self-distillation Framework for Multitask Learning with Wearable Sensor Data

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Smooth-Distill applies EMA parameter averaging as a self-distillation teacher for multitask HAR and placement detection, and reports consistent but modest gains over multitask baselines.

Pith tools