REVIEW 1 major objections 6 minor 68 references
Neural Collapse Inspired Knowledge Distillation
T0 review · 1 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that knowledge distillation works better when the student imitates the teacher's neural-collapse geometry, not just its logits or features, and that the proposed NCKD loss achieves state-of-the-art accuracy.
desk verdict NCKD is a promising distillation idea with a real under-specification: cross-architecture losses are undefined without a stated projection layer, so it needs a major revision but deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the simplex equiangular tight frame (ETF), the neural-collapse geometry in which each normalized class mean has inner product $-\frac{1}{K-1}$ with every other class mean and aligns with its classifier. The paper operationalizes three neural-collapse properties as loss terms: $L_{\mathrm{NC1}}$ pulls student features toward the teacher's class centroids through a contrastive cosine-similarity term; $L_{\mathrm{NC2}}$ pushes the matrix of similarities between student and teacher normalized centroids toward the teacher's ETF structure; and the NC3-inspired classifier reuses normalized centroids as classifier weights to remove the separate linear layer. The load-bearing mechanism is the $L_{\mathrm{NC2}}$ term: the ablation shows removing it costs the most accuracy, and the appendix proves that reaching the target matrix is equivalent to placing each student centroid in the teacher's equiangular frame.
What would settle it
Run the exact recipe on a cross-architecture pair with different last-layer widths, such as ResNet-32x4 to ShuffleNet-V1, without adding any projection layer: if the loss raises a dimension-mismatch error, the method as written cannot run on that reported setup. On a same-width pair, ablate the NC2 term while measuring student NC2: if accuracy does not drop even when the student's centroids stop tracking the teacher's ETF, then the reported gains are not caused by the transferred structure.
Extended reading notes
Core claim
The paper's central claim is that the teacher's neural-collapse structure is transferable knowledge: teaching the student to reproduce that structure improves distillation beyond matching logits or instance-level features. The authors support the claim with an empirical correlation—methods that distill better also move the student's NC metrics closer to the teacher's—and then turn the correlation into a method. Their total loss is $L_{\mathrm{total}} = L_{\mathrm{cls}} + \lambda_1 L_{\mathrm{NC1}} + \lambda_2 L_{\mathrm{NC2}}$, where $L_{\mathrm{NC1}}$ is a prototype-contrastive loss on cosine similarity between student features and teacher class centroids, and $L_{\mathrm{NC2}}$ minimizes the gap between the student-teacher centroid inner-product matrix and the teacher's simplex ETF matrix. They also replace the student's linear classifier with normalized centroids, invoking the NC3 self-duality property, and report this cuts training time without hurting accuracy. On their experiments, the combined loss improves every student architecture they test and outperforms prior distillation losses on classification and detection benchmarks.
Load-bearing premise
The losses compare student features with teacher class centroids as if both live in one shared vector space, but many teacher-student pairs in the experiments have different feature dimensions and the paper never says how those dimensions are matched before the inner products are computed.
Editorial extensions
If this is right
- NCKD works as a plug-in: adding it to CRD and SimKD improves their accuracy on CIFAR-100, not just when used alone.
- Students distilled with NCKD keep transferable features: frozen CIFAR-100 features give higher linear-classification accuracy on STL-10 and Tiny-ImageNet than baselines.
- In self-distillation, where no separate teacher exists, the same loss outperforms teacher-free baselines on ImageNet.
- Under NCKD, bigger teachers reliably produce better students, whereas existing methods plateau or degrade as the teacher grows.
- Replacing the standard classifier with the NC3-inspired normalized-centroid classifier cuts per-epoch training time by 8-13% without reducing accuracy.
Reading between the lines
- If the causal story holds, the NC metrics themselves could serve as a training signal or diagnostic: monitoring the gap between student and teacher NC2 would tell practitioners whether distillation is working, without needing a validation set.
- A natural extension the paper leaves untested is teacher selection by NC quality: choose the teacher whose class centroids are closest to a perfect ETF, since that is the structure being transferred.
- Because the loss formulas require a common feature space and the paper never specifies how cross-architecture pairs with different last-layer widths are aligned, the practical recipe must include an implicit or unstated projection; whether that projection is learned or fixed could change the measured gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Neural-Collapse-inspired Knowledge Distillation (NCKD), which adds two losses to the standard classification loss: a contrastive prototype-alignment loss L_NC1 (Eq. 7) that pulls student penultimate features toward the teacher's per-class centroids, and an ETF-structure loss L_NC2 (Eq. 8) that aligns the student's normalized class-mean matrix with the teacher's, so that the cross-class similarity matrix matches the simplex ETF Gram matrix. A third component replaces the student's linear classifier with normalized class centroids, motivated by the NC3 self-duality property. The total loss is L_total = L_cls + λ1 L_NC1 + λ2 L_NC2 (Eq. 9). The authors report state-of-the-art results on CIFAR-100, ImageNet, and MS-COCO for both homogeneous and cross-architecture teacher/student pairs, plus plug-in gains when combined with CRD and SimKD, and additional self-distillation and feature-transfer experiments.
Significance. If the method works as described, it is a simple and general plug-in distillation loss with broad applicability, and the paper supplies a plausible geometric interpretation connecting neural collapse to knowledge distillation. The experimental scope is wide: multiple CIFAR-100 architectures, ImageNet, COCO detection, self-distillation, and feature transfer, and the authors include an appendix proof of the ETF alignment property. However, the central technical gap described below prevents the reader from implementing or verifying the method on the very cross-architecture pairs that are prominent in the experiments, so the significance cannot be assessed as claimed without a fix.
major comments (1)
- [Appendix B, hyperparameters] The sensitivity study in Appendix D.1 and the statement in Appendix B that λ1 and λ2 are selected by grid search over [0, 2] do not report the chosen values for each experiment, nor the teacher/student pairs for which the grid search was conducted. Since the main results in Tables 1-3 presumably rely on per-pair or per-dataset hyperparameters, omitting them makes the results hard to reproduce even after the dimension mismatch is resolved. Please list the λ1, λ2 settings used for each reported configuration.
minor comments (6)
- [Eq. (7)] There is a mismatched parenthesis in the exponential term: the similarity expression should be written as sim(g_S(x_k^{(n)}), h_k^T)/τ.
- [Appendix B and Eq. (2)] The temperature τ for the NCKD losses is set to 0.1 'following the practice of CRD', but the KD loss in Eq. (2) uses τ = 4; please clarify which τ appears in Eq. (7) and whether the KD temperature is used only in the baseline comparisons.
- [Table 5 and Figure 2] The notation 'N C1', 'N C2', 'N C3' is inconsistent with the 'NC' abbreviation used elsewhere; please use 'NC1', 'NC2', 'NC3' uniformly.
- [Ablation Study, 'Distillation from Bigger Models'] The phrase 'GREAT TEACHERS PRODUCING OUTSTANDING STUDENTS' in all caps is informal and should be rephrased in normal prose.
- [References] The reference 'Kim, H.; and Kim, K. ???? Fixed Non-negative Orthogonal Classifier...' is missing the publication year and venue; it appears to be an ICLR submission and should be cited completely.
- [Figure 2 caption] The caption says 'The ideal NC results are characterized by NC1,2 approaching 0, and NC3 approaching 1', but the figure plots bars for accuracy and NC values without error bars or significance tests; please state that these are single-run or mean results and clarify the number of seeds.
Circularity Check
No circular derivation: teacher NC targets are fixed external quantities; accuracy is independent; the one self-citation is contextual only.
full rationale
The paper's central derivation is self-contained with respect to its empirical claims. The teacher's class centroids h^T_k and the ETF target matrix in Eq. (8) are fixed quantities computed from a pretrained teacher; the student losses L_NC1 (Eq. 7) and L_NC2 (Eq. 8) minimize alignment to those fixed external targets, and student accuracy is then evaluated on held-out test sets. The motivation in Figure 2 is an observational correlation between NC metrics and distillation accuracy, not an equation used to define the loss or to generate the reported accuracy numbers. The NC3 classifier sets the student classifier to its own normalized centroids, which makes the NC3 metric equal to 1 by construction, but the accuracy comparison with a standard classifier (Figure 5) is an independent empirical result. The only author self-citation (Zhang, Liu, and He 2024) appears in a related-work sentence about graph-level KD and is not load-bearing. The cross-architecture feature-dimension gap (e.g., ResNet-32x4 vs ShuffleNet-V1 in Tables 1-3) is an implementation under-specification that affects reproducibility, not circularity: nothing in the paper derives the reported gains from that undefined alignment. No fitted parameter is renamed as a prediction; lambda_1 and lambda_2 are grid-searched on CIFAR-100 and transferred to ImageNet, which is standard model selection rather than a forced statistical prediction. Because the paper contains one minor, non-load-bearing self-citation in the related work, I assign 2 on the rubric despite the absence of any circular reduction.
Assumptions & free parameters
free parameters (3)
- lambda_1 =
grid-searched in [0, 2]
- lambda_2 =
grid-searched in [0, 2]
- temperature_tau_NC1 =
0.1
assumptions (3)
- domain assumption Well-trained teachers exhibit neural collapse in their last-layer features.
- domain assumption Student and teacher features can be compared via cosine similarity in a common space.
- domain assumption Optimizing L_NC1 and L_NC2 in combination with the classification loss does not harm the student's class separability.
Cite this review
Pith. "Pith review of Neural Collapse Inspired Knowledge Distillation." pith.science (2026). https://pith.science/paper/HM42ON2K
@misc{pith2026241211788,
author = {Pith},
title = {Pith review of: Neural Collapse Inspired Knowledge Distillation},
year = {2026},
howpublished = {\url{https://pith.science/paper/HM42ON2K}},
note = {Machine review of arXiv:2412.11788}
}
read the original abstract
Existing knowledge distillation (KD) methods have demonstrated their ability in achieving student network performance on par with their teachers. However, the knowledge gap between the teacher and student remains significant and may hinder the effectiveness of the distillation process. In this work, we introduce the structure of Neural Collapse (NC) into the KD framework. NC typically occurs in the final phase of training, resulting in a graceful geometric structure where the last-layer features form a simplex equiangular tight frame. Such phenomenon has improved the generalization of deep network training. We hypothesize that NC can also alleviate the knowledge gap in distillation, thereby enhancing student performance. This paper begins with an empirical analysis to bridge the connection between knowledge distillation and neural collapse. Through this analysis, we establish that transferring the teacher's NC structure to the student benefits the distillation process. Therefore, instead of merely transferring instance-level logits or features, as done by existing distillation methods, we encourage students to learn the teacher's NC structure. Thereby, we propose a new distillation paradigm termed Neural Collapse-inspired Knowledge Distillation (NCKD). Comprehensive experiments demonstrate that NCKD is simple yet effective, improving the generalization of all distilled student models and achieving state-of-the-art accuracy performance.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
B.; Belkhir, N.; Popescu, S.; Manzanera, A.; and Franchi, G
Ammar, M. B.; Belkhir, N.; Popescu, S.; Manzanera, A.; and Franchi, G. 2023. NECO: NEural Collapse Based Out-of-distribution detection. arXiv preprint arXiv:2310.06823
arXiv 2023
-
[2]
Chen, D.; Mei, J.-P.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022. Knowledge Distillation with the Reused Teacher Classifier. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11933--11942
work page 2022
-
[3]
Chen, D.; Mei, J.-P.; Zhang, Y.; Wang, C.; Wang, Z.; Feng, Y.; and Chen, C. 2021 a . Cross-layer distillation with semantic calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, 7028--7036
work page 2021
-
[4]
Chen, K.; Wang, J.; Pang, J.; Cao, Y.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Xu, J.; et al. 2019. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155
arXiv 2019
-
[5]
Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021 b . Distilling knowledge via knowledge review. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5008--5017
work page 2021
-
[6]
Coates, A.; Ng, A.; and Lee, H. 2011. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth International Conference on Artificial Intelligence and Statistics, 215--223. JMLR Workshop and Conference Proceedings
work page 2011
-
[7]
Dang, H.; Tran, T.; Osher, S.; Tran-The, H.; Ho, N.; and Nguyen, T. 2023. Neural collapse in deep linear networks: from balanced to imbalanced data. arXiv preprint arXiv:2301.00437
arXiv 2023
-
[8]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929
arXiv 2020
Show all 68 references
-
[9]
L.; Zhang, X.; Zhao, P.; Zhu, F.; Zhao, R.; and Li, H
Ge, Y.; Choi, C. L.; Zhang, X.; Zhao, P.; Zhu, F.; Zhao, R.; and Li, H. 2021. Self-distillation with batch knowledge ensembling improves imagenet classification. arXiv:2104.13298
2021 arXiv
-
[10]
Girshick, R. 2015. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, 1440--1448
2015
-
[11]
Guo, Z.; Yan, H.; Li, H.; and Lin, X. 2023. Class attention transfer based knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11868--11877
2023
-
[12]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 a . Deep residual learning for image recognition. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, 770--778
2016
-
[13]
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 b . Identity mappings in deep residual networks. In Computer Vision--ECCV 2016: 14th European Conference
2016
-
[14]
Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the knowledge in a neural network. arXiv:1503.02531
2015 arXiv
-
[15]
Hu, J.; Shen, L.; and Sun, G. 2018. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7132--7141
2018
-
[16]
Huang, T.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2022. Knowledge distillation from a stronger teacher. Advances in Neural Information Processing Systems, 35: 33716--33727
2022
-
[17]
Huang, Z.; and Wang, N. 2017. Like what you like: Knowledge distill via neuron selectivity transfer. arXiv:1707.01219
2017 arXiv
-
[18]
Ji, M.; Shin, S.; Hwang, S.; Park, G.; and Moon, I.-C. 2021. Refine myself by teaching myself: Feature refinement via self-knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10664--10673
2021
-
[19]
Jin, Y.; Wang, J.; and Lin, D. 2023. Multi-level logit distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24276--24285
2023
-
[20]
Joyce, J. M. 2011. Kullback-leibler divergence. In International Encyclopedia of Statistical Science, 720--722. Springer
2011
-
[21]
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020. Supervised contrastive learning. Advances in neural information processing systems, 33: 18661--18673
2020
-
[22]
???? Fixed Non-negative Orthogonal Classifier: Inducing Zero-mean Neural Collapse with Feature Dimension Separation
Kim, H.; and Kim, K. ???? Fixed Non-negative Orthogonal Classifier: Inducing Zero-mean Neural Collapse with Feature Dimension Separation. In The Twelfth International Conference on Learning Representations
-
[23]
Kim, J.; Park, S.; and Kwak, N. 2018. Paraphrasing complex network: Network compression via factor transfer. In Advances in Neural Information Processing Systems
2018
-
[24]
Y.; and Jung, J.-H
Kim, J.; You, J.; Lee, D.; Kim, H. Y.; and Jung, J.-H. 2024. Do Topological Characteristics Help in Knowledge Distillation? In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Proceedings of the 41st International Confe...
2024
-
[25]
Kim, T.; Oh, J.; Kim, N.; Cho, S.; and Yun, S. 2021. Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, 2628--2635
2021
-
[26]
R.; Vakilian, V.; Behnia, T.; Gill, J.; and Thrampoulidis, C
Kini, G. R.; Vakilian, V.; Behnia, T.; Gill, J.; and Thrampoulidis, C. 2023. Supervised-contrastive loss learns orthogonal frames and batching matters. arXiv preprint arXiv:2306.07960
2023 arXiv
-
[27]
Kolesnikov, A.; Beyer, L.; Zhai, X.; Puigcerver, J.; Yung, J.; Gelly, S.; and Houlsby, N. 2020. Big transfer (bit): General visual representation learning. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part V 16, 491--5...
2020
-
[28]
Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images
2009
-
[29]
Le, Y.; and Yang, X. S. 2015. Tiny ImageNet Visual Recognition Challenge
2015
-
[30]
J.; and Shin, J
Lee, H.; Hwang, S. J.; and Shin, J. 2020. Self-supervised label augmentation via input transformations. In International Conference on Machine Learning, 5714--5724
2020
-
[31]
Lin, T.-Y.; Doll \'a r, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2117--2125
2017
-
[32]
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Doll \'a r, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 1...
2014
-
[33]
Liu, X.; Li, L.; Li, C.; and Yao, A. 2023. Norm: Knowledge distillation via n-to-one representation matching. arXiv preprint arXiv:2305.13803
2023 arXiv
-
[34]
Liu, Y.; Cao, J.; Li, B.; Yuan, C.; Hu, W.; Li, Y.; and Duan, Y. 2019. Knowledge distillation via instance relationship graph. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7096--7104
2019
-
[35]
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, 10012--10022
2021
-
[36]
Lu, J.; and Steinerberger, S. 2022. Neural collapse under cross-entropy loss. Applied and Computational Harmonic Analysis, 59: 224--241
2022
-
[37]
G.; Parshall, H.; and Pi, J
Mixon, D. G.; Parshall, H.; and Pi, J. 2022. Neural collapse with unconstrained features. Sampling Theory, Signal Processing, and Data Analysis, 20(2): 11
2022
-
[38]
Papyan, V.; Han, X.; and Donoho, D. L. 2020. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40): 24652--24663
2020
-
[39]
Park, W.; Kim, D.; Lu, Y.; and Cho, M. 2019. Relational knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 3967--3976
2019
-
[40]
Passalis, N.; and Tefas, A. 2018. Learning deep representations with probabilistic knowledge transfer. In Proceedings of the European Conference on Computer Vision (ECCV), 268--284
2018
-
[41]
Peng, B.; Jin, X.; Liu, J.; Li, D.; Wu, Y.; Liu, Y.; Zhou, S.; and Zhang, Z. 2019. Correlation congruence for knowledge distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 5007--5016
2019
-
[42]
P.; Liwicki, S.; and Cipolla, R
Poudel, R. P.; Liwicki, S.; and Cipolla, R. 2019. Fast-scnn: Fast semantic segmentation network. arXiv preprint arXiv:1902.04502
2019 arXiv
-
[43]
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28
2015
-
[44]
E.; Chassang, A.; Gatta, C.; and Bengio, Y
Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2014. Fitnets: Hints for thin deep nets. arXiv:1412.6550
2014 arXiv
-
[45]
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 211--252
2015
-
[46]
Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4510--4520
2018
-
[47]
Seo, M.; Koh, H.; Jeung, W.; Lee, M.; Kim, S.; Lee, H.; Cho, S.; Choi, S.; Kim, H.; and Choi, J. 2024. Learning Equi-angular Representations for Online Continual Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 23933--23942
2024
-
[48]
Tian, Y.; Krishnan, D.; and Isola, P. 2019. Contrastive representation distillation. arXiv:1910.10699
2019 arXiv
-
[49]
Wang, T.; Yuan, L.; Zhang, X.; and Feng, J. 2019. Distilling object detectors with fine-grained feature imitation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4933--4942
2019
-
[50]
Wang, Y.; Cheng, L.; Duan, M.; Wang, Y.; Feng, Z.; and Kong, S. 2023. Improving knowledge distillation via regularizing feature norm and direction. arXiv preprint arXiv:2305.17007
2023 arXiv
-
[51]
Xie, S.; Girshick, R.; Doll \'a r, P.; Tu, Z.; and He, K. 2017. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, 1492--1500
2017
-
[52]
Xu, T.-B.; and Liu, C.-L. 2019. Data-distortion guided self-distillation for deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence
2019
-
[53]
Xue, Y.; Joshi, S.; Gan, E.; Chen, P.-Y.; and Mirzasoleiman, B. 2023. Which features are learnt by contrastive learning? On the role of simplicity bias in class collapse and feature suppression. In International Conference on Machine Learning, 38938--38970. PMLR
2023
-
[54]
Yang, Y.; Yuan, H.; Li, X.; Lin, Z.; Torr, P.; and Tao, D. 2023. Neural collapse inspired feature-classifier alignment for few-shot class incremental learning. arXiv preprint arXiv:2302.03004
2023 arXiv
-
[55]
Yun, S.; Park, J.; Lee, K.; and Shin, J. 2020. Regularizing class-wise predictions via self-knowledge distillation. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, 13876--13885
2020
-
[56]
Zagoruyko, S.; and Komodakis, N. 2016 a . Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. arXiv:1612.03928
2016 arXiv
-
[57]
Zagoruyko, S.; and Komodakis, N. 2016 b . Wide Residual Networks. arXiv:1605.07146
2016 arXiv
-
[58]
Zhang, H.; Wu, C.; Zhang, Z.; Zhu, Y.; Lin, H.; Zhang, Z.; Sun, Y.; He, T.; Mueller, J.; Manmatha, R.; et al. 2022. Resnest: Split-attention networks. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, 2736--2746
2022
-
[59]
Zhang, L.; Song, J.; Gao, A.; Chen, J.; Bao, C.; and Ma, K. 2019. Be your own teacher: Improve the performance of convolutional neural networks via self distillation. In International Conference on Computer Vision
2019
-
[60]
Zhang, S.; Liu, H.; and He, K. 2024. Knowledge Distillation via Token-Level Relationship Graph Based on the Big Data Technologies. Big Data Research, 36: 100438
2024
-
[61]
Zhang, X.; Zhou, X.; Lin, M.; and Sun, J. 2018. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In CVPR, 6848--6856
2018
-
[62]
Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022. Decoupled knowledge distillation. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, 11953--11962
2022
-
[63]
Zhao, H.; Shi, J.; Qi, X.; Wang, X.; and Jia, J. 2017. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2881--2890
2017
-
[64]
Zheng, K.; and Yang, E.-H. 2024. Knowledge distillation based on transformed teacher matching. arXiv preprint arXiv:2402.11148
2024 arXiv
-
[65]
Zhou, J.; Li, X.; Ding, T.; You, C.; Qu, Q.; and Zhu, Z. 2022. On the optimization landscape of neural collapse under mse loss: Global optimality with unconstrained features. In International Conference on Machine Learning, 27179--27202. PMLR
2022
-
[66]
Zhu, Z.; Ding, T.; Zhou, J.; Li, X.; You, C.; Sulam, J.; and Qu, Q. 2021. A geometric analysis of neural collapse with unconstrained features. Advances in Neural Information Processing Systems, 34: 29820--29834
2021
-
[67]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[68]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.