Pith. sign in

REVIEW 4 major objections 5 minor 87 references

Cross-View Consistency Regularisation for Knowledge Distillation

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Cross-view consistency regularisation sets new logit-distillation records on CIFAR-100, Tiny-ImageNet, and ImageNet, adding no parameters.

desk verdict Solid empirical KD paper with a simple consistency-regularisation idea; the universal-boosting claim overreaches and the lack of error bars blunts the SOTA claim. read the letter →

arxiv 2412.16493 v1 pith:TL73GNYS submitted 2024-12-21 cs.CV

classification cs.CV
keywords knowledgedistillationlogitconsistencyregularisationcross-viewlearningdataaugmentationconfirmationbiasimageclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a simple extra training signal—forcing the student's predictions to agree with the teacher's across both the same and differently augmented views of an image—pushes logit-based knowledge distillation past current methods on CIFAR-100, Tiny-ImageNet, and ImageNet. If the claim holds, it matters because the method adds no network parameters and no inference cost, and it can be grafted onto existing logit-based distillation methods to improve their accuracy. The paper also claims the method keeps working when ground-truth labels are unavailable, and that it transfers from a vision-transformer teacher to a convolutional student. A sympathetic reading of the evidence is that the key contribution is showing that cross-view consistency, previously thought to be inferior to same-view matching, becomes beneficial when combined with same-view consistency and confidence-filtered soft labels.

What carries the argument

The load-bearing object is a four-way prediction-matching loss defined on a weakly augmented view $T_w(x)$ and a strongly augmented view $T_s(x)$ of each input. Within-view terms use $\mathrm{KLD}(p_S^w, p_T^w)$ and $\mathrm{KLD}(p_S^s, p_T^s)$; cross-view terms use $\mathrm{KLD}(p_S^w, p_T^s)$ and $\mathrm{KLD}(p_S^s, p_T^w)$, where $p$ denotes the teacher's or student's softmax predictions. Each term is multiplied by a binary mask from confidence-based soft-label mining, keeping only teacher predictions whose top class probability exceeds thresholds $\tau_w$ and $\tau_s$. The full objective sums these gated KL terms with cross-entropy on both views to the ground-truth label. The mechanism is claimed to work because the strong view amplifies dark knowledge, the confidence masks suppress confirmation bias, and the cross-view terms regularise the student against overconfident teacher outputs.

What would settle it

Re-run the paper's Table 5 ablation on Tiny-ImageNet (within-view terms alone versus within-view plus cross-view terms, keeping the strong augmentation and thresholds fixed): the central claim predicts a clear accuracy gain from the two cross-view terms. If the gain is absent or negative on this second dataset, the claim that cross-view consistency is the key ingredient would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that consistency regularisation can be moved from semi-supervised learning into logit-based knowledge distillation with strong results. CRLD takes a weakly and a strongly augmented view of each image, passes both through teacher and student, and enforces four matching terms: the student's weak prediction should match the teacher's weak prediction, the student's strong should match the teacher's strong, and the two cross-view pairs should also match. A confidence-based soft-label mining step discards teacher predictions whose maximum class probability falls below a threshold, so that unreliable supervision does not reinforce confirmation bias. On CIFAR-100, Tiny-ImageNet, and ImageNet, across homogeneous and heterogeneous teacher-student architecture pairs, the paper reports new state-of-the-art accuracies, and it reports that adding CRLD to NormKD, DKD, NKD, MLLD, and vanilla KD improves each of them. The paper also positions CRLD as a refinement of the earlier finding that teacher and student must see identical views: cross-view matching is harmful alone but helps in combination with within-view matching.

Load-bearing premise

The load-bearing premise is that matching the student's prediction on one augmentation to the teacher's prediction on a different augmentation transfers useful knowledge when paired with same-view matching; the paper's own ablations show cross-view terms alone degrade accuracy and offer no mechanism to explain why the combination works.

Editorial extensions

If this is right

  • Grafting CRLD onto existing logit-based methods such as NormKD, DKD, NKD, MLLD, and vanilla KD improves their accuracy on CIFAR-100 by about one to three points, with no extra parameters or inference overhead.
  • CRLD outperforms both logit-based and feature-based distillation baselines on CIFAR-100, Tiny-ImageNet, and ImageNet across many teacher-student architecture pairs.
  • The method remains effective in a label-free setting where ground-truth labels are unavailable during distillation, a practical scenario for proprietary teachers.
  • CRLD transfers knowledge from a ViT-L teacher to a ResNet-18 student, suggesting it generalises across architectures, not just within a family.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the cross-view terms may be functioning mostly as a strong regulariser rather than as a richer knowledge channel, since they are harmful alone but beneficial on top of within-view terms; that would explain the ablation pattern and predict that the gain shrinks as the student already generalises well.
  • Beyond the paper: the two confidence thresholds make the method depend on teacher calibration; a natural extension is to make the thresholds class-conditional or adapt them during training, which the paper does not explore.
  • Beyond the paper: the strong augmentation strength is a critical knob—the paper shows performance drops when too many RandAugment operations are applied—so the method may be sensitive to the augmentation policy on datasets unlike CIFAR-100 or ImageNet.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes CRLD, a logit-based knowledge distillation method that augments standard KD losses with within-view and cross-view consistency regularisation between teacher and student predictions on weak and strongly augmented views, plus a confidence-based soft-label mining (SLS) step that discards low-confidence teacher predictions. The authors evaluate CRLD on CIFAR-100, Tiny-ImageNet, and ImageNet across a range of homogeneous and heterogeneous teacher-student pairs, and also graft CRLD onto existing logit-based methods (KD, NKD, MLLD, DKD, NormKD). The central claims are that CRLD sets new state-of-the-art results, introduces no extra network parameters, and boosts the performance of various existing methods by considerable margins.

Significance. If the central claims are correct, CRLD is an attractive plug-in regulariser for logit-based distillation: it is simple, adds no parameters, applies to multiple base methods, and is evaluated on several standard benchmarks. The paper has genuine strengths: the empirical sweep is broad (multiple datasets, many teacher-student pairs, ablations in Table 5, sensitivity analyses in Figures 6-7, a ViT-teacher experiment in Table 9, and a label-free KD study in Table 8), the method is clearly specified in Algorithm 1, and the authors state that code and models will be released. However, the headline conclusions are empirical claims, and the evidence as reported is not yet strong enough to support the 'state-of-the-art' and 'boosts various existing approaches' statements: there are no error bars or significance tests, and one of the paper's own tables contains counterexamples to the universal-boosting claim.

major comments (4)
  1. [Abstract; Section 4.3, Table 6] The abstract claims that CRLD 'boosts the performance of various existing approaches by considerable margins', but Table 6 directly contradicts this for DKD: on ResNet56->ResNet20, DKD+CRLD scores 70.70 versus DKD's 71.49, and on ResNet110->ResNet32, DKD+CRLD scores 73.45 versus DKD's 73.95. These are not marginal differences (0.79 and 0.50 accuracy points) and they appear in the same table used to support the generalisation claim. The universal-boosting statement must be qualified, or the two counterexamples must be explained and shown to be attributable to a specific identifiable mechanism rather than left as exceptions.
  2. [Section 4.2; Tables 1-4, 6] The paper reports only means over 3 runs with no standard deviations, confidence intervals, or significance tests, yet the central claims rest on small margins. For example, in Table 1 the CRLD versus DKD gap on WRN-40-2->WRN-16-2 is 76.45 versus 76.24 (0.21 points), and in Table 4 the CRLD versus LSKD gap on ResNet34->ResNet18 is 72.37 versus 72.08 (0.29 points). With three seeds, these differences are well within plausible run-to-run noise. At minimum, the authors should report per-seed results or error bars for the headline comparisons and state whether the reported margins are stable across seeds; otherwise the 'new state-of-the-art' claim is not statistically supported.
  3. [Section 3.2, Eq. (2); Section 4.4, Table 5] The core novelty, cross-view consistency (Eq. (2)), is shown in Table 5 to be harmful when used alone (Expt C: 74.10, Expt D: 75.38) relative to using only within-view terms (Expt A: 76.26, Expt B: 76.75), yet beneficial when combined with within-view terms (Expts E-J). The paper offers no principled account of why cross-view matching is only useful in combination, beyond the assertion in Section 4.4 that 'cross-view consistencies are harmful when used individually, but are rather beneficial when applied in concert with within-view consistencies'. This is load-bearing because the method's claimed superiority depends on the interaction, not on the cross-view term per se. The authors should either provide a mechanistic explanation or add targeted experiments that isolate the conditions under which the interaction helps (e.g., varying the relative strengths of the two views, the confidence thresholds, or the teacher's accuracy), so that the result is not simply an unexplained hyperparameter-dependent phenomenon.
  4. [Section 4.4; Figure 6; Supplementary B] Several decisive hyperparameters are tuned on validation data (tau_w, tau_s, lambda_WV_KD, lambda_CV_KD, n, and p_s), and Figure 6 shows that performance varies across the threshold space even if the method is 'not highly sensitive'. Since the SOTA comparisons in Tables 1-4 presumably use tuned thresholds, the paper should state clearly which hyperparameter values were used for each dataset/architecture pair and whether the same values transfer across pairs. Without this information, the reader cannot judge whether the reported gains are robust to reasonable hyperparameter choices or partly reflect per-pair tuning.
minor comments (5)
  1. [Section 4.5, first paragraph] The text says 'in Table 3 we plot the evolution of training and test accuracies', but the corresponding visualisation is Figure 3, not Table 3.
  2. [Section 4.4, 'Sensitivity to tau_w and tau_s'] The phrase 'over the entire hyperaprameter space' contains a typo; it should be 'hyperparameter space'.
  3. [Section 4.1, 'ImageNet'] The description says ImageNet images are 'annotated in 100 classes', but ImageNet is a 1000-class dataset; the text appears to be a typo.
  4. [Section 3.2, Eq. (1); Algorithm 1] Equation (1) defines the within-view loss without the SLS mask, while Algorithm 1 applies masks M_w and M_s to the same within-view terms. Since Section 3.3 states that 'other objectives are defined like-wise', the notation should be made consistent so that the reader knows whether the masks apply to Eq. (1) as well.
  5. [Supplementary Materials, 'Strengths of view transformations'] The supplementary text says 'From the results in Table 9', but the relevant table in the main text is Table 7; the cross-reference should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: CRLD is an empirical loss-combination method evaluated on external benchmarks; its reported gains do not reduce to its own definitions, fitted parameters, or load-bearing self-citations.

full rationale

The paper's derivation chain is empirical rather than circular. The proposed objectives in Equations 1-3 are new loss terms (within-view KLD, cross-view KLD, and soft-label masking), and the reported predictions are test-set accuracies on CIFAR-100, Tiny-ImageNet, and ImageNet, which are external benchmarks not constructed from the method's own loss values. Hyperparameters such as tau_w, tau_s, and loss weights are tuned on validation data, which is standard model selection and not a fitted-input-called-prediction pattern. The ablations in Table 5 explicitly show that cross-view losses alone degrade performance relative to within-view losses (Expt. C: 74.10 and Expt. D: 75.38 vs. Expt. B: 76.75), so the central claim is not forced by construction; it is a contingent empirical result about how the terms interact. The one self-citation, reference [72] in the Related Work list of downstream KD applications, is peripheral and not load-bearing for any claimed result. There is no uniqueness theorem imported from the authors' prior work, and the consistency-regularisation idea is borrowed from external SSL literature with appropriate citations. The absence of error bars and the Table 6 counterexamples (DKD+CRLD dropping on two pairs) are legitimate empirical-correctness concerns, but they are not circularity: the paper does not define or fit its way to those numbers. Overall, the central claim is self-contained against external benchmarks and exhibits no equation-level or citation-level circularity.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The central results rest on standard math (KLD, softmax), domain assumptions about augmentation and teacher confidence, and several hyperparameters tuned on validation. No invented entities are introduced.

free parameters (7)
  • lambda_KD = not stated in paper
    Balancing weight for the distillation loss in Eq. 3; tuned on the validation set.
  • lambda_WV_KD = not stated in paper
    Weight for the within-view consistency loss, part of L_KD in Eq. 3; tuned on validation.
  • lambda_CV_KD = not stated in paper
    Weight for the cross-view consistency loss, part of L_KD in Eq. 3; tuned on validation.
  • tau_w = not stated; tuned on validation (Figure 6)
    Confidence threshold for teacher predictions on the weak view; controls the soft label mining trade-off.
  • tau_s = not stated; tuned on validation (Figure 6)
    Confidence threshold for teacher predictions on the strong view; controls the soft label mining trade-off.
  • n = 2 (default)
    Number of RandAugment operations sampled for the strong view; varied in Figure 7.
  • p_s = not stated in paper; tuned via Figure 7 and supplementary Eq. 3-4
    Probability multiplier modulating the strength of strong view transformations.
assumptions (5)
  • standard math Kullback-Leibler divergence is an appropriate distance measure for soft probability distributions in distillation.
    Used to define L_WV and L_CV in Eq. 1-3; standard in the KD literature.
  • domain assumption Weak and strong augmentations preserve the semantic label of the input image.
    Required for consistency regularisation between view pairs; the strong view is assumed to share the weak view's label. Section 3.2.
  • domain assumption Teacher maximum softmax probability is a reliable indicator of prediction correctness.
    Basis of the confidence-based soft label mining in Section 3.3; not separately validated against teacher calibration.
  • domain assumption Non-target class probabilities (dark knowledge) are valuable for distillation and should be preserved as soft labels.
    Motivates keeping soft labels instead of converting to hard pseudo-labels; Section 3.3, inherited from Hinton et al.
  • domain assumption Strong augmentation amplifies dark information in teacher predictions.
    Cited from [62] and used to justify the strong view in cross-view matching; Section 1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cross-View Consistency Regularisation for Knowledge Distillation." pith.science (2026). https://pith.science/paper/TL73GNYS

@misc{pith2026241216493,
  author       = {Pith},
  title        = {Pith review of: Cross-View Consistency Regularisation for Knowledge Distillation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TL73GNYS}},
  note         = {Machine review of arXiv:2412.16493}
}
read the original abstract

Knowledge distillation (KD) is an established paradigm for transferring privileged knowledge from a cumbersome model to a lightweight and efficient one. In recent years, logit-based KD methods are quickly catching up in performance with their feature-based counterparts. However, previous research has pointed out that logit-based methods are still fundamentally limited by two major issues in their training process, namely overconfident teacher and confirmation bias. Inspired by the success of cross-view learning in fields such as semi-supervised learning, in this work we introduce within-view and cross-view regularisations to standard logit-based distillation frameworks to combat the above cruxes. We also perform confidence-based soft label mining to improve the quality of distilling signals from the teacher, which further mitigates the confirmation bias problem. Despite its apparent simplicity, the proposed Consistency-Regularisation-based Logit Distillation (CRLD) significantly boosts student learning, setting new state-of-the-art results on the standard CIFAR-100, Tiny-ImageNet, and ImageNet datasets across a diversity of teacher and student architectures, whilst introducing no extra network parameters. Orthogonal to on-going logit-based distillation research, our method enjoys excellent generalisation properties and, without bells and whistles, boosts the performance of various existing approaches by considerable margins.

Figures

Figures reproduced from arXiv: 2412.16493 by the authors.

Figure 1
Figure 1. A schematic comparison of logit-based distillation [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. In summary, the contributions of this paper include: (1) We introduce extensive within-view and cross-view consis￾tency regularisations to combat the overconfident teacher and over-fitting problems common in KD. (2) We design a reliable pseudo-label mining module to sidestep the negative impact of unreliable and erroneous supervisory signals from the teacher, thereby mitigating confirmation bias in KD. (3) We presen… view at source ↗
Figure 2
Figure 2. The CRLD framework. An input image is transformed into a weakly-transformed view and a strongly-transformed [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figures from the paper (3 more)
Figure 3
Figure 3. Figure 3: Evolution of training (top) and test (bottom) set [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: t-SNE visualisation of teacher’s and distilled stu [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 7
Figure 7. Figure 7: Sensitivity of CRLD against varying strengths [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 74 canonical work pages

  1. [1]

    Lawrence, and Zhenwen Dai

    Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D. Lawrence, and Zhenwen Dai. 2019. Variational Information Distillation for Knowledge Transfer. InCVPR

  2. [2]

    O’Connor, and Kevin McGuinness

    Eric Arazo, Diego Ortego, Paul Albert, Noel E. O’Connor, and Kevin McGuinness

  3. [3]

    Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel

    David Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. 2020. ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring. In ICLR

  4. [4]

    David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A. Raffel. 2019. MixMatch: A Holistic Approach to Semi-Supervised Learning. In NIPS

  5. [5]

    Lucas Beyer, Xiaohua Zhai, Amelie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. 2022. Knowledge Distillation: A Good Teacher is Patient and Consistent. In CVPR

  6. [6]

    Defang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang, Yan Feng, and Chun Chen

  7. [7]

    Defang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang, Zhe Wang, Yan Feng, and Chun Chen. 2021. Cross-Layer Distillation with Semantic Calibration. In AAAI

  8. [8]

    Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. 2021. Distilling Knowledge via Knowledge Review. In CVPR

Show all 87 references
  1. [9]

    Yudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu, Frank de Hoog, and Zi Huang

  2. [10]

    Zhihao Chi, Tu Zheng, Hengjia Li, Zheng Yang, Boxi Wu, Binbin Lin, and Deng Cai. 2023. NormKD: Normalized Logits for Knowledge Distillation. In arXiv:2308.00520

  3. [11]

    Jang Hyun Cho and Bharath Hariharan. 2019. On the Efficacy of Knowledge Distillation. In ICCV

  4. [12]

    In NeurIPS

    Improved Feature Distillation via Projector Ensemble. In NeurIPS

  5. [13]

    Ekin Dogus Cubuk, Barret Zoph, Jon Shlens, and Quoc Le. 2020. RandAugment: Practical Automated Data Augmentation with a Reduced Search Space. In NIPS

  6. [14]

    Kaiwen Cui, Yingchen Yu, Fangneng Zhan, Shengcai Liao, Shijian Lu, and Eric P Xing. 2023. KD-DLGAN: Data limited image generation via knowledge distillation. In CVPR

  7. [15]

    Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V

    Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le

  8. [16]

    Deepan Das, Haley Massa, Abhimanyu Kulkarni, and Theodoros Rekatsinas

  9. [17]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Ima- geNet: A Large-Scale Hierarchical Image Database. In CVPR

  10. [18]

    Terrance DeVries and Graham W. Taylor. 2017. Improved Regularization of Convolutional Neural Networks with Cutout. In arXiv:1708.04552

  11. [19]

    Wanyun Cui and Sen Yan. 2021. Isotonic Data Augmentation for Knowledge Distillation. In IJCAI

  12. [20]

    Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge distillation of large language models. In ICLR

  13. [21]

    In ICML Workshop

    An Empirical Analysis of the Impact of Data Augmentation on Knowledge Distillation. In ICML Workshop

  14. [22]

    Xiaoyang Guo, Shaoshuai Shi, Xiaogang Wang, and Hongsheng Li. 2021. LIGA- Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D Detector. In ICCV

  15. [23]

    Ziyao Guo, Haonan Yan, Hui Li, and Xiaodong Lin. 2023. Class Attention Transfer Based Knowledge Distillation. In CVPR

  16. [24]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...

  17. [25]

    Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan

    Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. 2020. AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty. In ICLR

  18. [26]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In ICML

  19. [27]

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowlegde in a Neural Network. In arXiv:1503.02531

  20. [28]

    Yu Hong, Hang Dai, and Yong Ding. 2022. Cross-modality Knowledge Distillation Network for Monocular 3D Object Detection. In ECCV

  21. [29]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR

  22. [30]

    Tao Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. 2022. Knowledge Distillation from a Stronger Teacher. In NeurIPS

  23. [31]

    Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi. 2019. A Comprehensive Overhaul of Feature Distillation. In ICCV

  24. [32]

    Ying Jin, Jiaqi Wang, and Dahua Lin. 2023. Multi-level Logit Distillation. InCVPR

  25. [33]

    Minha Kim, Shahroz Tariq, and Simon S Woo. 2021. Fretal: Generalizing deepfake detection using knowledge distillation and representation learning. In CVPR

  26. [34]

    Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam

    Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. In arXiv:1704.04861

  27. [35]

    Samuli Laine and Timo Aila. 2017. Temporal Ensembling for Semi-Supervised Learning. In ICLR

  28. [36]

    Yuge Huang, Jiaxiang Wu, Xingkun Xu, and Shouhong Ding. 2022. Evaluation- oriented knowledge distillation for deep face recognition. In CVPR

  29. [37]

    Sihao Lin, Hongwei Xie, Bing Wang, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang, and Gang Wang. 2022. Knowledge Distillation via the Target-aware Transformer. In CVPR

  30. [38]

    Dongyang Liu, Meina Kan, Shiguang Shan, and Xilin Chen. 2023. Function- Consistent Feature Distillation. In ICLR

  31. [39]

    Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images

  32. [40]

    Yifan Liu, Ke Chen, Chris Liu, Zengchang Qin, Zhenbo Luo, and Jingdong Wang

  33. [41]

    Zheng Li, Xiang Li, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo, Jun Li, and Jian Yang. 2023. Curriculum Temperature for Knowledge Distillation. In AAAI

  34. [42]

    Rafael Müller, Simon Kornblith, and Geoffrey Hinton. 2019. When Does Label Smoothing Help?. In NIPS

  35. [43]

    Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. 2019. Relational Knowledge Distillation. In CVPR

  36. [44]

    Li Liu, Qingle Huang, Sihao Lin, Hongwei Xie, Bing Wang, Xiaojun Chang, and Xiaodan Liang. 2021. Exploring Inter-Channel Correlation for Diversity- preserved Knowledge Distillation. In ICCV

  37. [45]

    Baoyun Peng, Xiao Jin, Jiaheng Liu, Shunfeng Zhou, Yichao Wu, Yu Liu, Dong- sheng Li, and Zhaoning Zhang. 2019. Correlation Congruence for Knowledge Distillation. In CVPR

  38. [46]

    Structured Knowledge Distillation for Semantic Segmentation. In CVPR

  39. [47]

    Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Mat- sukawa, and Hassan Ghasemzadeh. 2020. Improved Knowledge Distillation via Teacher Assistant. In AAAI

  40. [48]

    Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang- Chieh Chen. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In CVPR

  41. [49]

    Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Net- works for Large-Scale Image Recognition. In ICLR

  42. [50]

    Nikolaos Passalis and Anastasios Tefas. 2018. Learning Deep Representations with Probabilistic Knowledge Transfer. In ECCV

  43. [51]

    Wonchul Son, Jaemin Na, Junyong Choi, and Wonjun Hwang. 2021. Densely Guided Knowledge Distillation using Multiple Teacher Assistants. In ICCV

  44. [52]

    Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. In ICLR

  45. [53]

    Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. 2016. Regularization with Stochastic Transformations and Perturbations for Deep Semi-Supervised Learning. In NIPS

  46. [54]

    Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Representation Distillation. In ICLR

  47. [55]

    Frederick Tung and Greg Mori. 2019. Similarity-Preserving Knowledge Distilla- tion. In ICCV

  48. [56]

    Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel

    Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. 2020. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence. In NIPS

  49. [57]

    Tao Wang, Li Yuan, Xiaopeng Zhang, and Jiashi Feng. 2019. Distilling object detectors with fine-grained feature imitation. In CVPR

  50. [58]

    Shangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang, Rui Wang, and Xiaochun Cao

  51. [59]

    Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy. 2020. Knowledge Distillation Meets Self-Supervision. In ECCV

  52. [60]

    Antti Tarvainen and Harri Valpola. 2017. Mean Teachers Are Better Role Models: Weight-Averaged Consistency Targets Improve Semi-Supervised Deep Learning Results. In NIPS

  53. [61]

    Jing Yang, Brais Martinez, Adrian Bulat, and Georgios Tzimiropoulos. 2021. Knowledge Distillation via Softmax Regression Representation Learning. In ICLR. MM ’24, October 28-November 1, 2024, Melbourne, VIC, Australia Weijia Zhang, Dongnan Liu, Weidong Cai, and Chao Ma

  54. [62]

    Lihe Yang, Lei Qi, litong Feng, Wayne Zhang, and Yinghuan Shi. 2023. Revisiting Weak-to-Strong Consistency in Semi-Supervised Semantic Segmentation. In CVPR

  55. [63]

    Huan Wang, Suhas Lohit, Mike Jones, and Yun Fu. 2022. What Makes A "Good" Data Augmentation in Knowledge Distillation - A Statistical Perspective. InNIPS

  56. [64]

    Zhendong Yang, Zhe Li, Mingqi Shao, Dachuan Shi, Zehuan Yuan, and Chun Yuan. 2022. Masked Generative Distillation. In ECCV

  57. [65]

    Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le. 2020. Unsupervised Data Augmentation for Consistency Training. In NIPS

  58. [66]

    Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Km. 2017. A Gift from Knowl- edge Distillation: Fast Optimization, Network Minimization and Transfer Learn- ing. In CVPR

  59. [67]

    Chuanguang Yang, Zhulin An, Linhang Cai, and Yongjun Xu. 2021. Hierarchical Self-Supervised Augmented Knowledge Distillation. In IJCAI

  60. [68]

    Sergey Zagorukyo and Nikos Komodakis. 2016. Wide Residual Networks. In BMVC

  61. [69]

    Sergey Zagoruyko and Nikos Komodakis. 2017. Paying More Attention to Atten- tion: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR

  62. [70]

    Zhendong Yang, Zhe Li, Xiaohu Jiang, Yuan Gong, Zehuan Yuan, Danpei Zhao, and Chun Yuan. 2022. Focal and Global Knowledge Distillation for Detectors. In CVPR

  63. [71]

    Dauphin, and David Lopez-Paz

    Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond Empirical Risk Minimization. In ICLR

  64. [72]

    Zhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang, Chun Yuan, and Yu Li. 2023. From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft Labels. In ICCV

  65. [73]

    Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2018. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. In CVPR

  66. [74]

    Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng. 2020. Revisiting Knowledge Distillation via Label Smoothing Regularization. In CVPR

  67. [75]

    Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang. 2022. Decoupled Knowledge Distillation. In CVPR

  68. [76]

    Kaixiang Zheng and En-Hui Yang. 2024. Knowledge Distillation based on Trans- formed Teaching Matching. In ICLR

  69. [77]

    Bo Zhang, Jiacheng Sui, and Li Niu. 2023. Foreground Object Search by Distilling Composite Image Feature. In ICCV

  70. [78]

    Shengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou, and Chao Ma. 2023. UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird’s-Eye View. In CVPR. Cross-View Consistency Regularisation for Knowledge Distillation - Supplementary...

  71. [79]

    Weijia Zhang, Dongnan Liu, Chao Ma, and Weidong Cai. 2024. Alleviating Foreground Sparsity for Semi-Supervised Monocular 3D Object Detection. In W ACV

  72. [81]

    Hospedales, and Huchuan Lu

    Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu. 2018. Deep Mutual Learning. In CVPR

  73. [84]

    Zhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren, Wangmeng Zuo, Qibin Hou, and Ming-Ming Cheng. 2022. Localization Distillation for Dense Object Detection. In CVPR

  74. [86]

    Weak-Weak

    using two independently composed weakly-transformed views ∗Corresponding author. Figure A: Examples of ImageNet [? ] images transformed by the proposed weak and strong transformations and predic- tions made by a ResNet32×4 teacher. (denoted as “Weak-Weak”); 2) using two indepe...

  75. [87]

    Nature of regularisation: CS-KD uses self-KD whereas we use cross-agent matching; 3) Choice of matched pairs: CS-KD uses class- wise samples whereas we use transformed views instance-wise. A fundamental flaw of CS-KD is that the non-target dark knowledge, proven crucial in KD,...

  76. [2019]

    AutoAugment: Learning Augmentation Strategies from Data. In CVPR

  77. [2020]

    Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning

  78. [2022]

    Knowledge Distillation with the Reused Teacher Classifier. In CVPR

  79. [2024]

    Logit Standardization in Knowledge Distillation. In CVPR

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.