Pith. sign in

REVIEW 1 cited by

BD-KD: Balancing the Divergences for Online Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.12965 v2 pith:WZ7V6DJN submitted 2022-12-25 cs.CV

classification cs.CV
keywords distillationonlineaccuracybd-kdcalibrationknowledgeperformancecompact
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the challenge of producing trustworthy and accurate compact models for edge devices. While Knowledge Distillation (KD) has improved model compression in terms of achieving high accuracy performance, calibration of these compact models has been overlooked. We introduce BD-KD (Balanced Divergence Knowledge Distillation), a framework for logit-based online KD. BD-KD enhances both accuracy and model calibration simultaneously, eliminating the need for post-hoc recalibration techniques, which add computational overhead to the overall training pipeline and degrade performance. Our method encourages student-centered training by adjusting the conventional online distillation loss on both the student and teacher losses, employing sample-wise weighting of forward and reverse Kullback-Leibler divergence. This strategy balances student network confidence and boosts performance. Experiments across CIFAR10, CIFAR100, TinyImageNet, and ImageNet datasets, and various architectures demonstrate improved calibration and accuracy compared to recent online KD methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Response-to-vision attention similarity predicts MLLM student performance far better than prompt-to-vision, and entropy-adaptive token-wise KL distillation (TRAG) transfers that signal effectively.

Pith tools