Pith. sign in

REVIEW 1 cited by

Regularizing Class-wise Predictions via Self-knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.13964 v2 pith:L7KE2CCK submitted 2020-03-31 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords predictionsclass-wisedistillationdistributiongeneralizationknowledgemethodnetworks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep neural networks with millions of parameters may suffer from poor generalization due to overfitting. To mitigate the issue, we propose a new regularization method that penalizes the predictive distribution between similar samples. In particular, we distill the predictive distribution between different samples of the same label during training. This results in regularizing the dark knowledge (i.e., the knowledge on wrong predictions) of a single network (i.e., a self-knowledge distillation) by forcing it to produce more meaningful and consistent predictions in a class-wise manner. Consequently, it mitigates overconfident predictions and reduces intra-class variations. Our experimental results on various image classification tasks demonstrate that the simple yet powerful method can significantly improve not only the generalization ability but also the calibration performance of modern convolutional neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self Distillation via Iterative Constructive Perturbations

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A self-distillation framework that aligns features of original and gradient-perturbed inputs reports an accuracy jump from 22.93% to 41.99% on CIFAR-100 and improved SSIM/FID on CUB.

Pith tools