Pith. sign in

REVIEW 1 cited by

Why distillation helps: a statistical perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.10419 v1 pith:OBFWZHGW submitted 2020-05-21 cs.LG stat.ML

classification cs.LGstat.ML
keywords distillationextremeknowledgelabelsmodelmulticlassobjectiveperspective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowledge distillation is a technique for improving the performance of a simple "student" model by replacing its one-hot training labels with a distribution over labels obtained from a complex "teacher" model. While this simple approach has proven widely effective, a basic question remains unresolved: why does distillation help? In this paper, we present a statistical perspective on distillation which addresses this question, and provides a novel connection to extreme multiclass retrieval techniques. Our core observation is that the teacher seeks to estimate the underlying (Bayes) class-probability function. Building on this, we establish a fundamental bias-variance tradeoff in the student's objective: this quantifies how approximate knowledge of these class-probabilities can significantly aid learning. Finally, we show how distillation complements existing negative mining techniques for extreme multiclass retrieval, and propose a unified objective which combines these ideas.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A compressed HuBERT can be trained with the original masked-prediction objective using k-means labels from the teacher, beating feature-distillation methods on four SUPERB tasks.

Pith tools