REVIEW 3 major objections 5 minor 70 references
Role of Mixup in Topological Persistence Based Knowledge Distillation for Wearable Sensor Data
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Mixup and knowledge distillation share a smoothing mechanism, and per-teacher mixup strengths let a small wearable-sensor model absorb topological knowledge.
desk verdict A useful empirical sweep of mixup for multi-teacher KD on wearable sensor data, but the headline claim is undermined by test-set hyperparameter selection and small effect sizes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is Equation (7), a two-teacher distillation objective in which each teacher contributes its own knowledge-distillation loss and its own mixup-augmented loss, with mixup strengths sampled from $\mathrm{Beta}(\alpha_1,\alpha_1)$ and $\mathrm{Beta}(\alpha_2,\alpha_2)$ respectively. Around it sit three control knobs: the annealing strategy, which initializes the student from a model trained from scratch to reduce the knowledge gap; the temperature $T$, which smooths each teacher's logits; and partial mixup, which restricts the number of mixed pairs per batch to avoid excessive smoothing. Persistence images supply the topological teacher's input as stable 2D grid representations of persistent homology.
What would settle it
Compare the recommended recipe against a strict protocol where $\alpha_1$, $\alpha_2$, temperature, and mixup-pair ratio are selected on a held-out validation set (or on held-out subjects) and the untouched test set is scored once. If accuracy no longer exceeds the equal-$\alpha$ mixup baseline, the per-teacher mixup recommendation reflects test-set selection rather than a general property.
Extended reading notes
Core claim
The central claim is that smoothness is the connecting link between mixup and knowledge distillation: KD softens the teacher's output distribution through temperature, while mixup softens labels by blending inputs and targets. Because both inject smoothness, applying mixup to the student in KD can create a synergetic effect, but too much smoothness degrades performance. The paper shows this on wearable sensor data by comparing single-teacher distillation from persistence images, multi-teacher distillation from time series plus persistence images, and the annealed multi-teacher variant (Ann.), both with and without mixup. It finds that Ann. with mixup is consistently best, and that the two teachers transfer different statistical knowledge, so using different mixup strengths for each teacher yields the highest accuracy.
Load-bearing premise
The paper's headline results depend on choosing temperature, per-teacher mixup strengths, and mixup-pair ratios on the same GENEActiv and PAMAP2 test sets that later appear as reported accuracies, so the load-bearing premise is that those test-set-selected choices are a reliable guide to unseen subjects and datasets.
Editorial extensions
If this is right
- A time-series-only student can carry topological knowledge at inference time, avoiding the cost of computing persistence images on a wearable device.
- Mixup is best applied to the student rather than to the teachers in this multi-teacher setting.
- Different teachers need different mixup strengths; a single shared $\alpha$ leaves accuracy behind.
- Temperature and mixup both add smoothness, so too much smoothness can hurt, and partial mixup provides a control knob.
- The annealed multi-teacher student's solution space stays close to the from-scratch solution, which the parametric plots connect to flatter, less overfit behavior.
Reading between the lines
- Beyond the paper: a held-out validation protocol on new subjects would directly test whether per-teacher mixup strengths generalize, since the reported hyperparameters were selected on the same test sets that are scored.
- Beyond the paper: because the mechanism is smoothness, the per-teacher mixup recipe should transfer to other paired representations, such as spectrograms or wavelet features as a second teacher, provided the second teacher softens different information than the first.
- Beyond the paper: the same idea could apply to single-modality multi-teacher distillation, where teachers with different capacities or training schedules naturally produce differently smoothed outputs and could each receive its own mixup strength.
- Beyond the paper: the paper does not isolate whether the gain comes from topological content or from having a second, differently smoothed teacher; training a second teacher on a non-topological auxiliary representation would separate these explanations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates the role of mixup augmentation in knowledge distillation when topological persistence images are used as a second teacher modality for wearable-sensor activity recognition. It compares single- and multi-teacher distillation strategies (standard KD, Base, and the annealed multi-teacher variant labeled Ann.), with and without mixup applied to teachers and/or student, while varying temperature, partial mixup ratios, and per-teacher mixup strengths, on GENEActiv and PAMAP2. The paper reports that the annealed multi-teacher strategy with per-teacher mixup hyperparameters gives the best accuracy (e.g., 71.22% on GENEActiv and 88.13% on PAMAP2) and concludes that the smoothness injected by mixup improves knowledge distillation. It also presents parametric plots of interpolation between trained solutions and a sensitivity analysis of the mixup strength alpha, and it argues that topological features complement time-series features in the distillation process.
Significance. If the empirical claims hold, the paper would provide a useful practical recipe for injecting topological persistence knowledge into a lightweight time-series-only student and would extend the image-domain understanding of mixup and knowledge distillation to multimodal time-series data. The study is wide in coverage: it includes single-teacher and multi-teacher distillation, several baselines (AT, SP, DIST, SimKD, AVER, EBKD, CA-MKD), two public datasets, multiple teacher-student architectures, standard deviations over three runs, and a rough efficiency comparison. These are genuine strengths. However, the central quantitative claims rest on very small accuracy differences, many within one standard deviation of three runs, and the headline configurations are selected on the same test sets later used for reporting. The contribution is therefore best understood as an exploratory empirical study whose conclusions require a more rigorous evaluation protocol before they can be accepted as general recommendations.
major comments (3)
- [Sections 4.5-4.6, Tables 8-16, Figures 9-10] Key hyperparameters appear to be selected on the same test sets used for the final reported results. The temperature T, the partial-mixup proportions (PMU 0.1, PMU 0.5, FMU) in Tables 8-9, and the per-teacher alpha pairs in Tables 10-15 are all chosen by comparing test-set accuracies on the GENEActiv held-out subjects and PAMAP2 leave-one-subject-out folds that are later reported as the headline outcomes. No validation split, nested cross-validation, or multiple-comparison control is described. Because the grids include many configurations (e.g., seven alpha pairs times two teacher families in Tables 10-11), the best observed pair (0.2, 0.15) on GENEActiv and (0.1, 0.15) on PAMAP2 may be a selection artifact, and the claim that per-teacher mixup strengths yield the best student is not supported by the current protocol. I ask the authors to either introduce a validation split for hyperparameter selection, use nested cross-validation, or report an independent evaluation of the selected configuration on truly held-out data.
- [Abstract and Section 4.3, Figure 5, Table 8] The abstract's unqualified statement that "applying mixup to training a student in KD improves performance" is contradicted by several of the paper's own results. For example, Table 8 shows that on GENEActiv with WRN16-3 teachers, TS+KD with FMU reaches 68.94% versus 69.50% without mixup, and Figure 5 shows mixup degrading PI-alone KD and Base KD in several configurations. The claim should be qualified to the specific strategies where improvement is observed (e.g., the annealed multi-teacher setup) and should acknowledge the documented degradation cases. As written, the abstract overstates the findings relative to the evidence in the manuscript.
- [Section 4.1.2 and Tables 8-15] The reported improvements are generally smaller than the run-to-run variability. On PAMAP2, the standard deviations are around 2.2 percentage points across all reported cells, while the headline differences between configurations are between 0.1 and 0.5 percentage points (e.g., 88.13 vs. 87.98 in Table 15 and 87.98 vs. 87.12 in Table 9). On GENEActiv, the best gains are about 0.5 points (71.22 vs. 70.72 in Table 10) with standard deviations of 0.1-0.2. With three runs per configuration, these differences are not statistically distinguishable, and no significance test, confidence interval, or paired analysis across subjects or folds is provided. Without such analysis, the central conclusion that mixup, and especially per-teacher mixup strengths, improves the annealed student is not established.
minor comments (5)
- [Section 4.1.1] The phrase "writ-worn tri-axial accelerometer" contains a typo; it should be "wrist-worn."
- [Table 1] The GFLOPs and processing-time columns are misaligned in the rendered table, which makes the efficiency comparison difficult to read.
- [Section 4.5.1 and Figures 9-10] The sentence beginning "For both KD with time-series and Ann..." contains a duplicated teacher description ("WRN16-3 teacher and T is 12 for WRN16-3 teacher"); please clarify which teacher/student configuration each temperature statement refers to.
- [Introduction, Section 1] The sentence "In section 6, we discuss our findings and conclusions" is inconsistent with the actual structure, where Section 5 is Discussion and Section 6 is Conclusion; please update the section references.
- [Section 3.2] The training objective in Eq. (7) is introduced without a step-by-step training recipe; adding pseudocode or a short algorithm box would improve reproducibility, since the per-teacher mixup pairs and sampling order are load-bearing for the method.
Circularity Check
No significant circularity: the paper's central claims are empirical measurements on public benchmark datasets, and the cited self-works are prior experiments, not definitions of the result.
full rationale
This is an empirical study rather than a formal derivation, so there is no equation-level circularity in which an output is identified with an input by construction. The multi-teacher KD loss (Eqs. 5-7) and the mixup loss (Eqs. 1-2) are standard definitions taken from the literature; the paper's claims about which strategy works best are supported by Tables 2-16 and Figures 4-11, which are measurements on the public GENEActiv and PAMAP2 datasets with comparisons against external baselines. The 'smoothness connects mixup and KD' premise is imported from the authors' prior WACV 2023 study [15], but it is used as a testable hypothesis and is not invoked as a uniqueness theorem or as a formal proof step; the present paper's contribution is the new empirical evaluation on time-series and topological persistence, which stands or falls on the data. Self-citations to [3] and [15] are extensions of the authors' own earlier methods, but the measured accuracies are not defined by those citations. The main methodological weakness is that hyperparameters such as temperature, partial-mixup proportions, and per-teacher alpha pairs appear to be selected using the same test sets that are later reported as outcomes (Sec. 4.5-4.6, Tables 8-16), and gains are small with overlapping standard deviations; this is a generalization and selection-bias concern, not a circularity concern, because the reported numbers are empirical results rather than predictions forced by construction. No fitted parameter is renamed as a prediction, and no load-bearing premise is justified solely by a self-citation chain. Accordingly the circularity score is 1.
Assumptions & free parameters
free parameters (7)
- alpha (mixup strength) =
0.1 default; varied 0.05 to 0.4
- temperature T =
4 default; 12 often best
- tau (KD loss weight) =
0.7 (GENEActiv), 0.99 (PAMAP2)
- eta (teacher weighting) =
0.7 (GENEActiv), 0.3 (PAMAP2)
- alpha1, alpha2 =
(0.15, 0.2) for GENEActiv, (0.1, 0.15) for PAMAP2
- PMU proportion =
0.1, 0.5, or 1.0 depending on configuration
- PI parameters =
Gaussian std 0.25 (GENEActiv) / 0.015 (PAMAP2); range [-10,10] / [-1,1]; image 64x64
assumptions (4)
- domain assumption Persistence images computed from time series capture complementary shape information that improves activity classification.
- domain assumption Logits from teachers with different architectures and input modalities can be combined into a single student via a weighted KD loss.
- ad hoc to paper The smoothing mechanism linking mixup and KD in image models transfers to time-series and topological persistence data.
- domain assumption Annealing, i.e., initializing the student from a model learned from scratch, reduces the knowledge gap between heterogeneous teachers.
Cite this review
Pith. "Pith review of Role of Mixup in Topological Persistence Based Knowledge Distillation for Wearable Sensor Data." pith.science (2026). https://pith.science/paper/FZAN6DSS
@misc{pith2026250200779,
author = {Pith},
title = {Pith review of: Role of Mixup in Topological Persistence Based Knowledge Distillation for Wearable Sensor Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/FZAN6DSS}},
note = {Machine review of arXiv:2502.00779}
}
read the original abstract
The analysis of wearable sensor data has enabled many successes in several applications. To represent the high-sampling rate time-series with sufficient detail, the use of topological data analysis (TDA) has been considered, and it is found that TDA can complement other time-series features. Nonetheless, due to the large time consumption and high computational resource requirements of extracting topological features through TDA, it is difficult to deploy topological knowledge in various applications. To tackle this problem, knowledge distillation (KD) can be adopted, which is a technique facilitating model compression and transfer learning to generate a smaller model by transferring knowledge from a larger network. By leveraging multiple teachers in KD, both time-series and topological features can be transferred, and finally, a superior student using only time-series data is distilled. On the other hand, mixup has been popularly used as a robust data augmentation technique to enhance model performance during training. Mixup and KD employ similar learning strategies. In KD, the student model learns from the smoothed distribution generated by the teacher model, while mixup creates smoothed labels by blending two labels. Hence, this common smoothness serves as the connecting link that establishes a connection between these two methods. In this paper, we analyze the role of mixup in KD with time-series as well as topological persistence, employing multiple teachers. We present a comprehensive analysis of various methods in KD and mixup on wearable sensor data.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
A. Som, H. Choi, K. N. Ramamurthy, M. P. Buman, P. Turaga, Pi-net: A deep learning approach to extract topological persistence images, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 834–835
work page 2020
- [3]
- [4]
-
[5]
L. M. Seversky, S. Davis, M. Berger, On time-series topological data analysis: New data and opportunities, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016, pp. 59–67
work page 2016
-
[6]
Munch, A user’s guide to topological data analysis, Journal of Learning Analytics 4 (2) (2017)
E. Munch, A user’s guide to topological data analysis, Journal of Learning Analytics 4 (2) (2017)
work page 2017
- [7]
-
[8]
H. Edelsbrunner, J. L. Harer, Computational topology: an introduction, American Mathematical Society, 2022
work page 2022
Show all 70 references
-
[9]
E. S. Jeon, H. Choi, A. Shukla, Y . Wang, M. P. Buman, P. Turaga, Constrained adaptive distillation based on topological persistence for wearable sensor data, IEEE Transactions on Instrumentation and Measurement 72 (2023) 1–14. doi:10.1109/TIM.2023.3329818
2023
-
[10]
Rieck, T
B. Rieck, T. Yates, C. Bock, K. Borgwardt, G. Wolf, N. Turk-Browne, S. Krishnaswamy, Uncovering the topology of time-varying fmri data using cubical persistence, Advances in Neural Information Processing Systems 33 (2020) 6900–6912. 19
2020
-
[11]
Jiang, B
F. Jiang, B. Xu, Z. Zhu, B. Zhang, Topological data analysis approach to extract the persistent homology features of ballistocardiogram signal in unobstructive atrial fibrillation detection, IEEE Sensors Journal 22 (7) (2022) 6920–6930
2022
-
[12]
Yan, Y .-S
Y . Yan, Y .-S. Liu, C.-D. Li, J.-H. Wang, L. Ma, J. Xiong, X.-X. Zhao, L. Wang, Topological descriptors of gait nonlinear dynamics toward freezing-of-gait episodes recognition in parkinson’s disease, IEEE Sensors Journal 22 (5) (2022) 4294–4304
2022
-
[13]
Chazal, B
F. Chazal, B. Michel, An introduction to topological data analysis: Fundamental and practical aspects for data scientists, Frontiers in Artificial Intelligence 4 (2021)
2021
-
[14]
H. Wang, S. Lohit, M. N. Jones, Y . Fu, What makes a ”good” data augmentation in knowledge distillation - a statistical perspective, Advances in Neural Information Processing Systems 35 (2022) 13456–13469
2022
-
[15]
H. Choi, E. S. Jeon, A. Shukla, P. Turaga, Understanding the role of mixup in knowledge distillation: An empirical study, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2319–2328
2023
-
[16]
X. Li, H. Xiong, C. Xu, D. Dou, Smile: Self-distilled mixup for efficient transfer learning, arXiv preprint arXiv:2103.13941 (2021)
2021 arXiv
-
[17]
C. Yang, Z. An, H. Zhou, L. Cai, X. Zhi, J. Wu, Y . Xu, Q. Zhang, Mixskd: Self-knowledge distillation from mixup for image recognition, in: European Conference on Computer Vision, Springer, 2022, pp. 534–551
2022
-
[18]
G. Xu, Z. Liu, C. C. Loy, Computation-efficient knowledge distillation via uncertainty-aware mixup, Pattern Recognition 138 (2023) 109338
2023
-
[19]
C. M. Bishop, Training with noise is equivalent to tikhonov regularization, Neural computation 7 (1) (1995) 108–116
1995
-
[20]
S. Chen, E. Dobriban, J. H. Lee, A group-theoretic framework for data augmentation, Journal of Machine Learning Research 21 (245) (2020) 1–71
2020
-
[21]
R. Shen, S. Bubeck, S. Gunasekar, Data augmentation as feature manipulation, in: International conference on machine learning, PMLR, 2022, pp. 19773–19808
2022
-
[22]
Allen-Zhu, Y
Z. Allen-Zhu, Y . Li, Towards understanding ensemble, knowledge distillation and self-distillation in deep learning, in: The Eleventh Interna- tional Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=Uuf2q9TfXGA
2023
-
[23]
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, Y . Yoo, Cutmix: Regularization strategy to train strong classifiers with localizable features, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6023–6032
2019
-
[24]
Allen-Zhu, Y
Z. Allen-Zhu, Y . Li, Feature purification: How adversarial training performs robust deep learning, in: 2021 IEEE 62nd Annual Symposium on Foundations of Computer Science (FOCS), IEEE, 2022, pp. 977–988
2021
-
[25]
T. Zhao, Y . Liu, L. Neves, O. Woodford, M. Jiang, N. Shah, Data augmentation for graph neural networks, in: Proceedings of the aaai conference on artificial intelligence, V ol. 35, 2021, pp. 11015–11023
2021
-
[26]
D. Zou, Y . Cao, Y . Li, Q. Gu, The benefits of mixup for feature learning, in: International Conference on Machine Learning, PMLR, 2023, pp. 43423–43479
2023
-
[27]
Beyer, X
L. Beyer, X. Zhai, A. Royer, L. Markeeva, R. Anil, A. Kolesnikov, Knowledge distillation: A good teacher is patient and consistent, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10925–10934
2022
-
[28]
Zhang, M
H. Zhang, M. Cisse, Y . N. Dauphin, D. Lopez-Paz, mixup: Beyond empirical risk minimization, in: Proceedings of the International Confer- ence on Learning Representations, 2018. URL https://openreview.net/forum?id=r1Ddp1-Rb
2018
-
[29]
Verma, A
V . Verma, A. Lamb, C. Beckham, A. Najafi, I. Mitliagkas, D. Lopez-Paz, Y . Bengio, Manifold mixup: Better representations by interpolating hidden states, in: Proceedings of the International Conference on Machine Learning, 2019, pp. 6438–6447
2019
-
[30]
J.-H. Kim, W. Choo, H. Jeong, H. O. Song, Co-mixup: Saliency guided joint mixup with supermodular diversity, arXiv preprint arXiv:2102.03065 (2021)
2021 arXiv
-
[31]
L. N. Darlow, A. Joosen, M. Asenov, Q. Deng, J. Wang, A. Barker, Tsmix: time series data augmentation by mixing sources, in: Proceedings of the 3rd Workshop on Machine Learning and Systems, 2023, pp. 109–114
2023
-
[32]
Aggarwal, J
K. Aggarwal, J. Srivastava, Embarrassingly simple mixup for time-series, arXiv preprint arXiv:2304.04271 (2023)
2023 arXiv
-
[33]
Y . Zhou, L. You, W. Zhu, P. Xu, Improving time series forecasting with mixup data augmentation, in: ECML PKDD 2023 International Workshop on Machine Learning for Irregular Time Series, 2023. URL https://www.amazon.science/publications/improving-time-series-forecasting-with-mi...
2023
-
[34]
Y . Wang, R. Behroozmand, L. P. Johnson, L. Bonilha, J. Fridriksson, Topological signal processing and inference of event-related potential response, Journal of Neuroscience Methods 363 (2021) 109324. doi:https://doi.org/10.1016/j.jneumeth.2021.109324
2021
-
[35]
Gholizadeh, W
S. Gholizadeh, W. Zadrozny, A short survey of topological data analysis in time series and systems analysis, arXiv preprint arXiv:1809.10745 (2018)
2018 arXiv
-
[36]
S. Zeng, F. Graf, C. Hofer, R. Kwitt, Topological attention for time series forecasting, Advances in Neural Information Processing Systems 34 (2021) 24871–24882
2021
-
[37]
B. J. Stolz, Outlier-robust subsampling techniques for persistent homology, Journal of Machine Learning Research 24 (2023)
2023
-
[38]
Bucilu ˇa, R
C. Bucilu ˇa, R. Caruana, A. Niculescu-Mizil, Model compression, in: Proceedings of the ACM International Conference on Knowledge Discovery and Data Mining (KDD), 2006, pp. 535–541
2006
-
[39]
Hinton, O
G. Hinton, O. Vinyals, J. Dean, Distilling the knowledge in a neural network, in: Proceedings of the NeurIPS Deep Learning and Represen- tation Learning Workshop, V ol. 2, 2015
2015
-
[40]
J. H. Cho, B. Hariharan, On the efficacy of knowledge distillation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 4794–4802
2019
-
[41]
J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge distillation: A survey, International Journal of Computer Vision 129 (6) (2021) 1789–1819
2021
-
[42]
Zagoruyko, N
S. Zagoruyko, N. Komodakis, Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer, in: Proceedings of the International Conference on Learning and Representations (ICLR), 2017, pp. 1–13
2017
-
[43]
F. Tung, G. Mori, Similarity-preserving knowledge distillation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 1365–1374
2019
-
[44]
Y . Liu, W. Zhang, J. Wang, Adaptive multi-teacher multi-level knowledge distillation, Neurocomputing 415 (2020) 106–113
2020
-
[45]
Zhang, D
H. Zhang, D. Chen, C. Wang, Confidence-aware multi-teacher knowledge distillation, in: Proceedings of the IEEE International Conference 20 on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 4498–4502
2022
-
[46]
S. You, C. Xu, C. Xu, D. Tao, Learning from multiple teacher networks, in: Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, pp. 1285–1294
2017
-
[47]
Q. Wang, S. Lohit, M. J. Toledo, M. P. Buman, P. Turaga, A statistical estimation framework for energy expenditure of physical activities from a wrist-worn accelerometer, in: Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Soc...
2016
-
[48]
E. S. Jeon, A. Som, A. Shukla, K. Hasanaj, M. P. Buman, P. Turaga, Role of data augmentation strategies in knowledge distillation for wearable sensor data, IEEE Internet of Things Journal 9 (14) (2022) 12848–12860
2022
-
[49]
Reiss, D
A. Reiss, D. Stricker, Introducing a new benchmarked dataset for activity monitoring, in: Proceedings of the International Symposium on Wearable Computers, 2012, pp. 108–109
2012
-
[50]
Jordao, A
A. Jordao, A. C. Nazare Jr, J. Sena, W. R. Schwartz, Human activity recognition based on wearable sensor data: A standardization of the state-of-the-art, arXiv preprint arXiv:1806.05226 (2018)
2018 arXiv
-
[51]
N. Saul, C. Tralie, Scikit-tda: Topological data analysis for python (2019). doi:10.5281/zenodo.2533369. URL https://doi.org/10.5281/zenodo.2533369
2019 doi
-
[52]
Zagoruyko, N
S. Zagoruyko, N. Komodakis, Wide residual networks, in: Proceedings of the British Machine Vision Conference, 2016
2016
-
[53]
Komodakis, S
N. Komodakis, S. Zagoruyko, Paying more attention to attention: improving the performance of convolutional neural networks via attention transfer, in: Proceedings of the International Conference on Learning and Representations (ICLR), 2017, pp. 1–13
2017
-
[54]
Chen, J.-P
D. Chen, J.-P. Mei, H. Zhang, C. Wang, Y . Feng, C. Chen, Knowledge distillation with the reused teacher classifier, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 11933–11942
2022
-
[55]
Huang, S
T. Huang, S. You, F. Wang, C. Qian, C. Xu, Knowledge distillation from a stronger teacher, Advances in Neural Information Processing Systems 35 (2022) 33716–33727
2022
-
[56]
K. Kwon, H. Na, H. Lee, N. S. Kim, Adaptive knowledge distillation based on entropy, in: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 7409–7413
2020
-
[57]
Cortes, V
C. Cortes, V . Vapnik, Support-vector networks, Machine learning 20 (3) (1995) 273–297
1995
-
[58]
H. Choi, Q. Wang, M. Toledo, P. Turaga, M. Buman, A. Srivastava, Temporal alignment improves feature quality: an experiment on activity recognition with accelerometer data, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 349–357
2018
-
[59]
Y . Chen, Y . Xue, A deep learning approach to human activity recognition based on single accelerometer, in: Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics, 2015, pp. 1488–1492
2015
-
[60]
Ha, J.-M
S. Ha, J.-M. Yun, S. Choi, Multi-modal convolutional neural networks for activity recognition, in: Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics, 2015, pp. 3017–3022
2015
-
[61]
S. Ha, S. Choi, Convolutional neural networks for human activity recognition using multiple accelerometer and gyroscope sensors, in: Proceedings of the International Joint Conference on Neural Networks, 2016, pp. 381–388
2016
-
[62]
Catal, S
C. Catal, S. Tufekci, E. Pirmit, G. Kocabag, On the use of ensemble of classifiers for accelerometer-based activity recognition, Applied Soft Computing 37 (2015) 1018–1022
2015
-
[63]
H.-J. Kim, M. Kim, S.-J. Lee, Y . S. Choi, An analysis of eating activities for automatic food type recognition, in: Proceedings of the Asia Pacific Signal and Information Processing Association Annual Summit and Conference, 2012, pp. 1–5
2012
-
[64]
Rosenberg, J
A. Rosenberg, J. Hirschberg, V-measure: A conditional entropy-based external cluster evaluation measure, in: Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, 2007, pp. 410–420
2007
-
[65]
DeVries, G
T. DeVries, G. W. Taylor, Improved regularization of convolutional neural networks with cutout, arXiv preprint arXiv:1708.04552 (2017)
2017 arXiv
-
[66]
I. J. Goodfellow, O. Vinyals, A. M. Saxe, Qualitatively characterizing neural network optimization problems, arXiv preprint arXiv:1412.6544 (2014)
2014 arXiv
-
[67]
F. Zhu, Z. Cheng, X.-Y . Zhang, C.-L. Liu, Rethinking confidence calibration for failure prediction, in: Proceedings of the European Confer- ence on Computer Vision (ECCV), 2022, pp. 518–536
2022
-
[68]
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, P. T. P. Tang, On large-batch training for deep learning: Generalization gap and sharp minima, in: Proceedings of the International Conference on Learning and Representations (ICLR), 2017
2017
-
[69]
W. Han, X. Dong, Y . Zhang, D. Crandall, C.-Z. Xu, J. Shen, Asymmetric convolution: An efficient and generalized method to fuse feature maps in multiple vision tasks, IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[70]
X. Dong, J. Shen, F. Porikli, J. Luo, L. Shao, Adaptive siamese tracking with a compact latent network, IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (7) (2022) 8049–8062. 21
2022
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.