REVIEW 3 major objections 6 minor 60 references
Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A teacher trained on a full seven-week course can distill its knowledge into a student that sees only the first three or six weeks, yielding higher average recall and F1 for at-risk students than standard RNN, GRU, and LSTM baselines.
desk verdict Sensible incremental RNN-FitNets extension, but the evaluation never isolates KD from attention and the numbers are within plausible noise; fixable major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a pair of mean-squared-error distillation losses over internal representations. The hint loss $\mathcal{L}_{\mathrm{HD}}$ forces the student's hidden state at the early cut $n$ to equal the teacher's hidden state at the final week $m$, so the early network must encode the whole course's accumulated information in a single vector. The context-vector loss $\mathcal{L}_{\mathrm{CV}}$ forces the student's attention-weighted context vector to match the teacher's, which in turn pressures the student's attention weights $\alpha_i$ to concentrate on the time steps the full-sequence teacher found salient, countering the vanishing-gradient tendency of RNNs to forget early weeks. A third distillation term on the teacher's soft logits is included in the full objective but the ablation shows it hurts; the two representation-matching losses carry the gain.
What would settle it
Train the RNN-Attention student on weeks 1--3 or 1--6 with distillation from a teacher that saw all seven weeks, and compare it to the same student trained with a teacher whose later-week inputs (weeks 4--7) were shuffled or replaced by noise. If the distilled student still outperforms the early-only baseline, the gain does not come from genuine future information transferred through the hidden state and context vector. A second check: replace the teacher targets in $\mathcal{L}_{\mathrm{HD}}$ and $\mathcal{L}_{\mathrm{CV}}$ with random vectors of the same dimension; if recall and F1 stay unchanged, representation matching is not the active mechanism.
Extended reading notes
Core claim
The paper's central claim is that knowledge distillation can compress the time axis of a course rather than the model. A teacher RNN with attention trained on all $M$ weeks guides a student network that sees only weeks $1$ through $N$, by matching two representations: the teacher's final hidden state $h^t_m$ to the student's early hidden state $h^s_n$ through $\mathcal{L}_{\mathrm{HD}} = \mathrm{MSE}(h^t_m, h^s_n)$, and the teacher's attention context vector $c^t_m$ to the student's early context vector $c^s_n$ through $\mathcal{L}_{\mathrm{CV}} = \mathrm{MSE}(c^t_m, c^s_n)$. Because both models share the same hidden dimension, the student is trained to produce the representation the full sequence would have produced from only the early sequence. In six train/test splits built from four years of a seven-week programming course, the resulting RNN-Attention-KD model reports the highest average recall and F1-measure across datasets, with recall and F1 of $0.49$ and $0.51$ for weeks 1--3 and $0.51$ and $0.61$ for weeks 1--6, and beats the conventional baselines in four of the six splits.
Load-bearing premise
The claim depends on the assumption that making a student's early hidden state and attention context vector equal the teacher's full-sequence final hidden state and context vector is a valid way to transfer information about future weeks into the early model; if matching these representations does not actually carry usable future knowledge, the reported improvement would disappear.
Editorial extensions
If this is right
- At-risk flags are available by week 3: across all datasets the method reaches average recall 0.49 and F1 0.51 on weeks 1--3, and 0.51 and 0.61 on weeks 1--6.
- Because students at this institution could withdraw only after five weeks, a week-3 or week-6 flag arrives before the withdrawal deadline, enabling interventions while the student is still in the course.
- The ablation implies the framework can be simplified: hint loss plus context-vector loss is sufficient, and the teacher-logit soft cross-entropy term should be dropped or down-weighted.
- The results degrade when training data come from an on-site offering and test data from pandemic-era online offerings, so modality shifts between years are a practical limit of the method.
Reading between the lines
- The same time-compression reading of distillation could apply to other early-warning domains, such as medical monitoring or equipment failure, where a model with the full horizon supervises a model that must act after a few observations.
- Because teacher and student share architecture and hidden size, the method does not compress the model; it compresses the input horizon, so the practical saving is in how early a decision can be made, not in parameter count.
- A direct test of the mechanism would vary the early cut $N$ from 1 to 7 and check whether the student's performance approaches the teacher's monotonically as $N$ grows; abrupt jumps would suggest the losses are not transferring smoothly.
- The finding that soft logit distillation hurts in this small imbalanced setting suggests feature-level distillation may be more reliable than output-level dark knowledge for education data, a hypothesis that could be checked by sweeping the temperature and mixing weight $\lambda$.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RNN-Attention-KD, a knowledge distillation framework for early prediction of at-risk students in a university course. A teacher model, an RNN with an attention mechanism, is trained on full-course (7-week) data; a student model of the same architecture is trained on only the first 3–6 weeks and is guided by three distillation losses: a hidden-state hint loss (Eq. 6), a context-vector loss (Eq. 7), and a soft-target cross-entropy loss combined with hard labels (Eq. 8). The paper evaluates the model on six dataset splits derived from four years of a programming course, comparing against MLP, RNN, GRU, LSTM, Bi-GRU, and Bi-LSTM baselines (Table 5), and conducts an ablation study over the distillation objectives (Table 6). The reported results claim the highest average recall and F1 across datasets and that the hint loss and context-vector loss are the effective components.
Significance. If the central claim were well supported, the idea of using knowledge distillation for temporal compression in educational early-warning systems would be a useful and transferable contribution. The paper has several strengths: the research question is practically motivated, the authors provide a public code repository, and the limitations section is honest about the single-course setting and lack of deployment. However, the current empirical design does not isolate the effect of knowledge distillation from the effect of the attention mechanism, and the statistical evidence for the reported advantages is weak. The significance is therefore conditional: the proposed framework may be useful, but the evidence presented does not yet establish that the distillation component is what drives the improvements.
major comments (3)
- [§5.1, Table 5] The central claim that knowledge distillation improves early prediction is confounded by the absence of a same-architecture no-KD control. The baselines in Table 5 (MLP, RNN, GRU, LSTM, Bi-GRU, Bi-LSTM) differ from RNN-Attention-KD in two simultaneous ways: they lack the attention module and they lack the teacher-loss objectives in Eqs. (6)–(8). Therefore, the reported gains could be due entirely to the attention mechanism rather than to distillation. The ablation study in Table 6 is not a substitute, because every row keeps the teacher model and includes at least one distillation term; there is no row corresponding to an RNN-Attention student trained on the early weeks with only the hard-label loss. This missing control is load-bearing for the paper's mechanism claim and must be added.
- [§5.1, Tables 5 and 6] All results are reported as point estimates (means over 30 runs) with no standard deviations, confidence intervals, or significance tests. The test sets contain only 50–62 students (Table 1, Table 4), so F1 differences of 0.01–0.05 are within plausible sampling noise. For example, in Table 5, T20P21 weeks 1–3 shows RNN-Attention-KD F1 = 0.50 versus Bi-LSTM F1 = 0.49, and T21P22 weeks 1–6 shows F1 = 0.56 versus GRU and Bi-GRU at 0.55. Without a measure of variance, the claims of 'outperforms traditional neural network models' and 'the highest average recall and F1-measure' are not statistically established. The paper should report the full distributions of the 30 runs and apply appropriate pairwise significance tests or bootstrap intervals.
- [§5.2, Table 6] The ablation study does not support the paper's conclusion that the hint loss (L_HD) and context-vector loss (L_CV) 'can enhance the model's prediction performance'. The F1 differences between the full model and the single-loss or two-loss variants are mostly within 0.01–0.04, and some single-loss variants outperform the full model. For instance, in T19P20 weeks 1–6, Only L_HD achieves F1 = 0.74 versus 0.72 for the full model; in T192021P22 weeks 1–5, Only L_HD achieves 0.66 versus 0.65. Additionally, Only L_KD(Soft) is often within 0.02–0.03 of the full model (e.g., T19P20 weeks 1–6: 0.72 versus 0.72). Given the lack of significance testing, the ablation should be interpreted as exploratory, and the stated conclusion about the necessity of the two losses is not supported.
minor comments (6)
- [§4.2.2] The hyperparameter search is described for RNN-Attention-KD, but it is unclear whether the baseline models (MLP, RNN, Bi-GRU, Bi-LSTM) received the same grid search procedure; please specify their hyperparameters or state that they were tuned identically to ensure a fair comparison.
- [§4.1] The features in Table 2 are described as capturing 'student activities for each lecture', but the model input is weekly aggregated data; please clarify the temporal aggregation window and whether the SRP scores are computed per week or per lecture.
- [§5.2, Eq. (8)] The soft-target weight λ is stated to be 0.1 in the text of Section 5.2, but Section 4.2.2 does not describe how λ was set in the hyperparameter search; please state the value, whether it was tuned, and how sensitive the results are to it.
- [§5.1] The explanation that the PT2019 on-site versus online modality caused the performance drop in T19P20 and T1920P21 is speculative; please soften the wording or provide supporting evidence (e.g., a feature-distribution comparison or a targeted experiment).
- [§2.2] The phrase 'Hinton et al. [20]'s study' is awkward; consider rephrasing to 'the study by Hinton et al. [20]'.
- [Abstract] The abstract states specific numeric recall and F1 values (0.49/0.51 for weeks 1–3 and 0.51/0.61 for weeks 1–6) without noting that these are averages over the six datasets; please make this explicit in both the abstract and Section 5.1.
Circularity Check
No circular reduction: the KD objectives are training losses evaluated against external baselines, and the only self-citation is not load-bearing.
full rationale
The paper's derivation chain is self-contained. Eq. 6 (L_HD = MSE(h^t_m, h^s_n)) and Eq. 7 (L_CV = MSE(c^t_m, c^s_n)) are training objectives that transfer teacher representations to the student; they are not fitted constants or renamed outputs. Eq. 8 combines hard and soft cross-entropy so the student still learns ground-truth labels. The teacher is trained on full-course data from prior years and the student on early weeks from the same training split, with test-year courses held out (Table 4); thus no test label or test feature enters the training objective. The central claim (Table 5) compares RNN-Attention-KD against MLP/RNN/GRU/LSTM/Bi-GRU/Bi-LSTM, and the ablation (Table 6) varies the distillation objectives; neither table's quantity is defined in terms of the quantity being predicted. The only self-citation of note is Murata et al. [31], which shares an author and introduced RNN-FitNets time-series KD; the paper cites it as prior art and does not use it to justify the reported gains. The missing no-KD RNN-Attention control is a legitimate experimental-attribution weakness, but it is a confounding control problem, not a circular reduction of an equation to its own input. Therefore no circularity.
Assumptions & free parameters
free parameters (5)
- KD soft-loss weight lambda =
0.1
- GRU hidden units and layers =
1 layer, 4 units
- SRP quantile thresholds =
decile ranks 10 to 1, zero activity = 0
- Decision threshold for at-risk classification =
not reported (presumed 0.5)
- Optimizer hyperparameters =
learning rate 0.01 or 0.001, batch size 8, epochs 150, weight decay 1e-5
assumptions (4)
- domain assumption Final course grade (C, D, or F) is a valid binary at-risk label, and each student's label is available for all weekly predictions.
- domain assumption The teacher's final hidden state and context vector computed over all M weeks are suitable targets for early student models.
- domain assumption Knowledge distilled from a teacher trained on prior-year courses transfers to the next year's students despite modality shifts (on-site vs online).
- ad hoc to paper MSE between hidden states and context vectors is a valid similarity measure for knowledge transfer.
Cite this review
Pith. "Pith review of Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance." pith.science (2026). https://pith.science/paper/W3TQ5FQS
@misc{pith2026241214526,
author = {Pith},
title = {Pith review of: Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance},
year = {2026},
howpublished = {\url{https://pith.science/paper/W3TQ5FQS}},
note = {Machine review of arXiv:2412.14526}
}
read the original abstract
Educational data mining (EDM) is a part of applied computing that focuses on automatically analyzing data from learning contexts. Early prediction for identifying at-risk students is a crucial and widely researched topic in EDM research. It enables instructors to support at-risk students to stay on track, preventing student dropout or failure. Previous studies have predicted students' learning performance to identify at-risk students by using machine learning on data collected from e-learning platforms. However, most studies aimed to identify at-risk students utilizing the entire course data after the course finished. This does not correspond to the real-world scenario that at-risk students may drop out before the course ends. To address this problem, we introduce an RNN-Attention-KD (knowledge distillation) framework to predict at-risk students early throughout a course. It leverages the strengths of Recurrent Neural Networks (RNNs) in handling time-sequence data to predict students' performance at each time step and employs an attention mechanism to focus on relevant time steps for improved predictive accuracy. At the same time, KD is applied to compress the time steps to facilitate early prediction. In an empirical evaluation, RNN-Attention-KD outperforms traditional neural network models in terms of recall and F1-measure. For example, it obtained recall and F1-measure of 0.49 and 0.51 for Weeks 1--3 and 0.51 and 0.61 for Weeks 1--6 across all datasets from four years of a university course. Then, an ablation study investigated the contributions of different knowledge transfer methods (distillation objectives). We found that hint loss from the hidden layer of RNN and context vector loss from the attention module on RNN could enhance the model's prediction performance for identifying at-risk students. These results are relevant for EDM researchers employing deep learning models.
Figures
Reference graph
Works this paper leans on
-
[1]
Malak Abdullah, Mahmoud Al-Ayyoub, Farah Shatnawi, Saif Rawashdeh, and Rob Abbott. 2023. Predicting students’ academic performance using e-learning logs. IAES International Journal of Artificial Intelligence (IJ-AI) 12 (06 2023), 831. https://doi.org/10.11591/ijai.v12.i2.pp831-839
-
[2]
Fatima Ahmed Al-azazi and Mossa Ghurab. 2023. ANN-LSTM: A deep learning model for early student performance prediction in MOOC. Heliyon 9, 4 (4 2023), e15382. https://doi.org/10.1016/j.heliyon.2023.e15382
-
[3]
Balqis Albreiki, Tetiana Habuza, and Nazar Zaki. 2023. Extracting topological features to identify at-risk students using machine learning and graph convolu- tional network models. International Journal of Educational Technology in Higher Education 20, 1 (04 2023), 1–22. https://doi.org/10.1186/s41239-023-00389-3
-
[4]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In 3rd International Conference on Learning Representations, ICLR 2015, May 7-9, 2015, Conference Track Proceedings . Cornell University, San Diego, CA, USA, 15 pages. http: //arxiv.org/abs/1409.0473
arXiv 2015
-
[5]
Brijesh Kumar Baradwaj and Saurabh Pal. 2011. Mining Educational Data to Analyze Students Performance. International Journal of Advanced Computer Science and Applications 2, 6 (2011), 7 pages. https://doi.org/10.14569/IJACSA. 2011.020609
work page Pith review arXiv 2011
-
[6]
Yoshua Bengio, Patrice Y. Simard, and Paolo Frasconi. 1994. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks 5, 2 (3 1994), 157–166. https://doi.org/10.1109/72.279181
-
[7]
Gabriella Casalino, Giovanna Castellano, Andrea Mannavola, and Gennaro Ves- sio. 2020. Educational Stream Data Analysis: A Case Study. In 2020 IEEE 20th Mediterranean Electrotechnical Conference ( MELECON) . IEEE, New York, NY, USA, 232–237. https://doi.org/10.1109/MELECON48756.2020.9140510
-
[8]
Cheng-Huan Chen, Stephen J. H. Yang, Jian-Xuan Weng, Hiroaki Ogata, and Chien-Yuan Su. 2021. Predicting at-risk university students based on their e-book reading behaviours by using machine learning classifiers. Australasian Journal of Educational Technology 37, 4 (6 2021), 130–144. https://doi.org/10.14742/ajet.6116
Show all 60 references
-
[9]
Fu Chen and Ying Cui. 2020. Utilizing Student Time Series Behaviour in Learning Management Systems for Early Prediction of Course Performance. Journal of Learning Analytics 7, 2 (Sep. 2020), 1–17. https://doi.org/10.18608/jla.2020.72.1
2020 doi
-
[10]
Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker
-
[11]
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Rep- resentations using RNN Encoder–Decoder for Statistical Machine Translation. In 2014 Conference on Empirical Methods in Natural ...
2014 doi
-
[12]
Ouafae El Aissaoui, Yasser El Alami El Madani, Lahcen Oughdir, Ahmed Dakkak, and Youssouf El Allioui. 2020. A Multiple Linear Regression-Based Approach to Predict Student Performance. In Advanced Intelligent Systems for Sustainable Development (AI2SD’2019). Springer Internatio...
2020 doi
-
[13]
Jo Shan Fu. 2013. ICT in education: A critical literature review and its impli- cations. International Journal of Education and Development Using Informa- tion and Communication Technology (IJEDICT) 9, 1 (01 2013), 112–125. https: //files.eric.ed.gov/fulltext/EJ1182651.pdf
2013
-
[14]
Filippos Giannakas, Christos Troussas, Ioannis Voyiatzis, and Cleo Sgouropoulou
-
[15]
Yann Ling Goh, Yeh Huann Goh, Chun-Chieh Yip, Chen Hunt Ting, Kah Pin Chen, and Raymond Ling Leh Bin. 2020. Prediction of Students’ Academic Performance by K-Means Clustering. International Journal of Advanced Science and Technology 29, 10s (4 2020), 1–6. https://doi.org/10.31...
2020 doi
-
[16]
Maybank, and Dacheng Tao
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowl- edge Distillation: A Survey.International Journal of Computer Vision 129, 6 (March 2021), 1789–1819. https://doi.org/10.1007/s11263-021-01453-z
2021 doi
-
[17]
Alamri, Ranim S
Leena H. Alamri, Ranim S. Almuslim, Mona S. Alotibi, Dana K. Alkadi, Irfan Ullah Khan, and Nida Aslam. 2021. Predicting Student Academic Performance using Support Vector Machine and Random Forest. In Proceedings of the 2020 3rd International Conference on Education Technology ...
2021
-
[18]
Yanbai He, Rui Chen, Xinya Li, Chuanyan Hao, Sijiang Liu, Gangyao Zhang, and Bo Jiang. 2020. Online At-Risk Student Identification using RNN-GRU Joint Neural Networks. Information 11, 10 (10 2020), 474. https://doi.org/10.3390/ info11100474
2020
-
[19]
Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao
Arto Hellas, Petri Ihantola, Andrew Petersen, Vangel V. Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao. 2018. Predicting academic performance: a systematic literature review. In Proceedings Companion of the 23rd Annual ...
2018
- [20]
-
[21]
Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-Term Memory. Neural Computation 9, 8 (11 1997), 1735–1780. https://doi.org/10.1162/neco.1997. 9.8.1735 SAC ’25, March 31-April 4, 2025, Catania, Italy Sukrit Leelaluk, Cheng Tang, Valdemar Švábenský, and Atsushi Shimada
1997 doi
-
[22]
Ya-Han Hu, Chia-Lun Lo, and Sheng-Pao Shih. 2014. Developing early warning systems to predict students’ online learning performance. Computers in Human Behavior 36 (7 2014), 469–478. https://doi.org/10.1016/j.chb.2014.04.002
2014 doi
-
[23]
Mushtaq Hussain, Wenhao Zhu, Wu Zhang, and Syed Muhammad Raza Abidi
-
[24]
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for Natural Language Understand- ing. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu ...
2020 doi
- [25]
-
[26]
Charles Koutcheme, Sami Sarsa, Arto Hellas, Lassi Haaranen, and Juho Leinonen
-
[27]
Sukrit Leelaluk, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada. 2022. Predicting student performance based on Lecture Materials data using Neural Network Models. CEUR Workshop Proceedings 3120 (2022), 11–20. https://ceur-ws.org/Vol-3120/paper2.pdf
2022
-
[28]
Lopez Z, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada
Erwin D. Lopez Z, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada. 2022. Assessment of At-Risk Students’ Predictions From e- Book Activities Representations in Practical Applications. In 30th International Conference on Computers in Education, ICCE 2022 - ...
2022
-
[29]
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective Ap- proaches to Attention-based Neural Machine Translation, In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lluís Màrquez, Chris Callison-Burch, Jian Su, Daniele Pigh...
2015 arXiv
-
[30]
Nur Izzati Mohd Talib, Nazatul Aini Abd Majid, and Shahnorbanun Sahran. 2023. Identification of Student Behavioral Patterns in Higher Education Using K-Means Clustering and Support Vector Machine. Applied Sciences 13, 5 (3 2023), 3267. https://doi.org/10.3390/app13053267
2023 doi
-
[31]
Ryusuke Murata, Fumiya Okubo, Tsubasa Minematsu, Yuta Taniguchi, and At- sushi Shimada. 2023. Recurrent Neural Network-FitNets: Improving Early Pre- diction of Student Performanceby Time-Series Knowledge Distillation. Journal of Educational Computing Research 61, 3 (6 2023), 6...
2023
-
[32]
Hiroaki Ogata, Chengjiu Yin, Misato Oi, Fumiya Okubo, Atsushi Shimada, Ken- taro Kojima, and Masanori Yamada. 2015. E-book-based learning analytics in University education. In Proceedings of the 23rd International Conference on Com- puters in Education, ICCE 2015 . Asia-Pacifi...
2015
-
[33]
Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, and Shin’ichi Konomi
-
[34]
Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, and Hiroaki Ogata. 2017. A neural network approach for students’ performance prediction. In Proceedings of the Seventh International Learning Analytics & Knowledge Conference , Marek Hatala, Alyssa Wis, Phil Winne, Grace Lync...
2017
-
[35]
Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, Yuta Taniguchi, and Konomi Shin’ichi. 2018. On the prediction of students’ quiz score by recurrent neural network. CEUR Workshop Proceedings 2163 (2018), 6 pages. https://ceur- ws.org/Vol-2163/paper3.pdf
2018
-
[36]
Feiyue Qiu, Guodao Zhang, Xin Sheng, Lei Jiang, Lijia Zhu, Qifeng Xiang, Bo Jiang, and Ping-Kuo Chen. 2022. Predicting students’ performance in e-learning using learning process and behaviour data. Scientific Reports 12, 1 (01 2022). https://doi.org/10.1038/s41598-021-03867-8
2022 doi
-
[37]
Moises Riestra-González, Maria del Puerto Paule-Ruíz, and Francisco Ortin. 2021. Massive LMS log data analysis for the early prediction of course-agnostic student performance. Computers & Education 163 (4 2021), 104108. https://doi.org/10. 1016/j.compedu.2020.104108
2021
-
[38]
In Proceedings of the 25th International Con- ference on Computers in Education, ICCE 2017 - Main Conference Pro- ceedings
Students’ performance prediction using data of multiple courses by recurrent neural network. In Proceedings of the 25th International Con- ference on Computers in Education, ICCE 2017 - Main Conference Pro- ceedings. Asia-Pacific Society for Computers in Education, Taiwan, 439–
2017
-
[39]
T M Noviyanti Sagala, Syarifah Diana Permai, Alexander Agung Santoso Gu- nawan, Rehnianty Octora Barus, and Cito Meriko. 2022. Predicting Computer Science Student’s Performance using Logistic Regression. In2022 5th International Seminar on Research of Information Technology an...
2022
- [40]
- [41]
-
[42]
Francisco Segura Altamirano
Luis Vives, Ivan Cabezas, Juan Carlos Vives, Nilton German Reyes, Janet Aquino, Jose Bautista Cóndor, and S. Francisco Segura Altamirano. 2024. Prediction of Students’ Academic Performance in the Programming Fundamentals Course Using Long Short-Term Memory Neural Networks. IEE...
2024
-
[43]
Baker, Pavel Čeleda, Jan Vykopal, Jens Mache, and Ankur Chattopadhyay
Valdemar Švábenský, Kristián Tkáčik, Aubrey Birdwell, Richard Weiss, Ryan S. Baker, Pavel Čeleda, Jan Vykopal, Jens Mache, and Ankur Chattopadhyay. 2024. Detecting Unsuccessful Students in Cybersecurity Exercises in Two Different Learning Environments. In Proceedings of the 54...
2024 arXiv
-
[44]
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. arXiv:1412.6550 [cs.LG] https://arxiv.org/abs/1412.6550
2015 arXiv
- [45]
-
[46]
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021. MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers. In Findings of the Association for Computational Lin- guistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, We...
2021 doi
-
[47]
Xing Wei, Yuqing Liu, Jiajia Li, Huiyong Chu, Zichen Zhang, Feng Tan, and Pengwei Hu. 2022. B-AT-KD: Binary attention map knowledge distillation. Neu- rocomputing 511 (9 2022), 299–307. https://doi.org/10.1016/j.neucom.2022.09.064
2022 doi
-
[48]
Jacob Whitehill, Kiran Mohan, Daniel Seaton, Yigal Rosen, and Dustin Tingley
-
[49]
Mariana Windarti and Putri Taqwa Prasetyaninrum. 2020. Prediction Analysis Student Graduate Using Multilayer Perceptron. InProceedings of the International Conference on Online and Blended Learning 2019 (ICOBL 2019) . Atlantis Press, Paris, France, 5 pages. https://doi.org/10....
2020 doi
-
[50]
Yue Xie, Ruiyu Liang, Zhenlin Liang, Chengwei Huang, Cairong Zou, and Björn Schuller. 2019. Speech Emotion Classification Using Attention-Based LSTM. IEEE/ACM Transactions on Audio, Speech, and Language Processing 27, 11 (7 2019), 1675–1685. https://doi.org/10.1109/TASLP.2019.2925934
2019
-
[51]
Aljohani, Guanliang Chen, and Dragan Gasevic
Hajra Waheed, Saeed-Ul Hassan, Raheel Nawaz, Naif R. Aljohani, Guanliang Chen, and Dragan Gasevic. 2023. Early prediction of learners at risk in self-paced education: A neural network approach. Expert Systems with Applications 213 (3 2023), 118868. https://doi.org/10.1016/j.es...
2023
-
[52]
Yupei Zhang, Yue Yun, Rui An, Jiaqi Cui, Huan Dai, and Xuequn Shang. 2021. Educational Data Mining Techniques for Student Performance Prediction: Method Review and Comparison Analysis. Frontiers in Psychology 12 (12 2021), 19 pages. https://doi.org/10.3389/fpsyg.2021.698490
2021
- [56]
-
[59]
Sergey Zagoruyko and Nikos Komodakis. 2017. Paying More Attention to Atten- tion: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR. OpenReview.net, Online, 13 pages. https://arxiv.org/abs/1612. 03928
2017
-
[444]
https://kyushu-u.pure.elsevier.com/en/publications/students-performance- prediction-using-data-of-multiple-courses-by
-
[2017]
Wallach, Rob Fergus, S
Learning Efficient Object Detection Models with Knowledge Distilla- tion, In Advances in Neural Information Processing Systems, Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). Neural Information Pr...
2017
-
[2018]
Computational Intelligence and Neuro- science 2018, 1 (10 2018), 6347186
Student Engagement Predictions in an e-Learning System and Their Impact on Student Course Assessment Scores. Computational Intelligence and Neuro- science 2018, 1 (10 2018), 6347186. https://doi.org/10.1155/2018/6347186
2018 doi
-
[2021]
Applied Soft Computing 106 (7 2021), 107355
A deep learning classification framework for early prediction of team- based academic performance. Applied Soft Computing 106 (7 2021), 107355. https://doi.org/10.1016/j.asoc.2021.107355
2021
-
[2022]
In Proceed- ings of the 24th Australasian Computing Education Conference (ACE ’22) , Judy Sheard and Paul Denny (Eds.)
Methodological Considerations for Predicting At-risk Students. In Proceed- ings of the 24th Australasian Computing Education Conference (ACE ’22) , Judy Sheard and Paul Denny (Eds.). Association for Computing Machinery, New York, NY, USA, 105–113. https://doi.org/10.1145/35118...
-
[5898]
https://doi.org/10.1109/ACCESS.2024.3350169
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.