Pith. sign in

REVIEW 3 major objections 6 minor 60 references

Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A teacher trained on a full seven-week course can distill its knowledge into a student that sees only the first three or six weeks, yielding higher average recall and F1 for at-risk students than standard RNN, GRU, and LSTM baselines.

desk verdict Sensible incremental RNN-FitNets extension, but the evaluation never isolates KD from attention and the numbers are within plausible noise; fixable major revision. read the letter →

arxiv 2412.14526 v1 pith:W3TQ5FQS submitted 2024-12-19 cs.LG cs.CY

classification cs.LGcs.CY
keywords studentperformancepredictionat-riskstudentsknowledgedistillationrecurrentneuralnetworksattentionmechanismearlyeducationaldataminingtime-seriescompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a student model seeing only the first three or six weeks of a course can identify at-risk students more reliably than standard models trained on the same early data, provided it is trained by knowledge distillation from a teacher that saw all seven weeks. The proposed RNN-Attention-KD framework matches the teacher's final hidden state and attention context vector to the student's early representations, effectively compressing the time axis rather than the model. In four of six year-to-year datasets it beats MLP, RNN, GRU, LSTM, and bidirectional variants on recall and F1, and it holds the highest average recall and F1 across all datasets. An ablation attributes the gain to the hidden-state and context-vector losses, while soft cross-entropy against the teacher's logits hurts.

What carries the argument

The load-bearing mechanism is a pair of mean-squared-error distillation losses over internal representations. The hint loss $\mathcal{L}_{\mathrm{HD}}$ forces the student's hidden state at the early cut $n$ to equal the teacher's hidden state at the final week $m$, so the early network must encode the whole course's accumulated information in a single vector. The context-vector loss $\mathcal{L}_{\mathrm{CV}}$ forces the student's attention-weighted context vector to match the teacher's, which in turn pressures the student's attention weights $\alpha_i$ to concentrate on the time steps the full-sequence teacher found salient, countering the vanishing-gradient tendency of RNNs to forget early weeks. A third distillation term on the teacher's soft logits is included in the full objective but the ablation shows it hurts; the two representation-matching losses carry the gain.

What would settle it

Train the RNN-Attention student on weeks 1--3 or 1--6 with distillation from a teacher that saw all seven weeks, and compare it to the same student trained with a teacher whose later-week inputs (weeks 4--7) were shuffled or replaced by noise. If the distilled student still outperforms the early-only baseline, the gain does not come from genuine future information transferred through the hidden state and context vector. A second check: replace the teacher targets in $\mathcal{L}_{\mathrm{HD}}$ and $\mathcal{L}_{\mathrm{CV}}$ with random vectors of the same dimension; if recall and F1 stay unchanged, representation matching is not the active mechanism.

Watch

Extended reading notes

Core claim

The paper's central claim is that knowledge distillation can compress the time axis of a course rather than the model. A teacher RNN with attention trained on all $M$ weeks guides a student network that sees only weeks $1$ through $N$, by matching two representations: the teacher's final hidden state $h^t_m$ to the student's early hidden state $h^s_n$ through $\mathcal{L}_{\mathrm{HD}} = \mathrm{MSE}(h^t_m, h^s_n)$, and the teacher's attention context vector $c^t_m$ to the student's early context vector $c^s_n$ through $\mathcal{L}_{\mathrm{CV}} = \mathrm{MSE}(c^t_m, c^s_n)$. Because both models share the same hidden dimension, the student is trained to produce the representation the full sequence would have produced from only the early sequence. In six train/test splits built from four years of a seven-week programming course, the resulting RNN-Attention-KD model reports the highest average recall and F1-measure across datasets, with recall and F1 of $0.49$ and $0.51$ for weeks 1--3 and $0.51$ and $0.61$ for weeks 1--6, and beats the conventional baselines in four of the six splits.

Load-bearing premise

The claim depends on the assumption that making a student's early hidden state and attention context vector equal the teacher's full-sequence final hidden state and context vector is a valid way to transfer information about future weeks into the early model; if matching these representations does not actually carry usable future knowledge, the reported improvement would disappear.

Editorial extensions

If this is right

  • At-risk flags are available by week 3: across all datasets the method reaches average recall 0.49 and F1 0.51 on weeks 1--3, and 0.51 and 0.61 on weeks 1--6.
  • Because students at this institution could withdraw only after five weeks, a week-3 or week-6 flag arrives before the withdrawal deadline, enabling interventions while the student is still in the course.
  • The ablation implies the framework can be simplified: hint loss plus context-vector loss is sufficient, and the teacher-logit soft cross-entropy term should be dropped or down-weighted.
  • The results degrade when training data come from an on-site offering and test data from pandemic-era online offerings, so modality shifts between years are a practical limit of the method.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same time-compression reading of distillation could apply to other early-warning domains, such as medical monitoring or equipment failure, where a model with the full horizon supervises a model that must act after a few observations.
  • Because teacher and student share architecture and hidden size, the method does not compress the model; it compresses the input horizon, so the practical saving is in how early a decision can be made, not in parameter count.
  • A direct test of the mechanism would vary the early cut $N$ from 1 to 7 and check whether the student's performance approaches the teacher's monotonically as $N$ grows; abrupt jumps would suggest the losses are not transferring smoothly.
  • The finding that soft logit distillation hurts in this small imbalanced setting suggests feature-level distillation may be more reliable than output-level dark knowledge for education data, a hypothesis that could be checked by sweeping the temperature and mixing weight $\lambda$.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes RNN-Attention-KD, a knowledge distillation framework for early prediction of at-risk students in a university course. A teacher model, an RNN with an attention mechanism, is trained on full-course (7-week) data; a student model of the same architecture is trained on only the first 3–6 weeks and is guided by three distillation losses: a hidden-state hint loss (Eq. 6), a context-vector loss (Eq. 7), and a soft-target cross-entropy loss combined with hard labels (Eq. 8). The paper evaluates the model on six dataset splits derived from four years of a programming course, comparing against MLP, RNN, GRU, LSTM, Bi-GRU, and Bi-LSTM baselines (Table 5), and conducts an ablation study over the distillation objectives (Table 6). The reported results claim the highest average recall and F1 across datasets and that the hint loss and context-vector loss are the effective components.

Significance. If the central claim were well supported, the idea of using knowledge distillation for temporal compression in educational early-warning systems would be a useful and transferable contribution. The paper has several strengths: the research question is practically motivated, the authors provide a public code repository, and the limitations section is honest about the single-course setting and lack of deployment. However, the current empirical design does not isolate the effect of knowledge distillation from the effect of the attention mechanism, and the statistical evidence for the reported advantages is weak. The significance is therefore conditional: the proposed framework may be useful, but the evidence presented does not yet establish that the distillation component is what drives the improvements.

major comments (3)
  1. [§5.1, Table 5] The central claim that knowledge distillation improves early prediction is confounded by the absence of a same-architecture no-KD control. The baselines in Table 5 (MLP, RNN, GRU, LSTM, Bi-GRU, Bi-LSTM) differ from RNN-Attention-KD in two simultaneous ways: they lack the attention module and they lack the teacher-loss objectives in Eqs. (6)–(8). Therefore, the reported gains could be due entirely to the attention mechanism rather than to distillation. The ablation study in Table 6 is not a substitute, because every row keeps the teacher model and includes at least one distillation term; there is no row corresponding to an RNN-Attention student trained on the early weeks with only the hard-label loss. This missing control is load-bearing for the paper's mechanism claim and must be added.
  2. [§5.1, Tables 5 and 6] All results are reported as point estimates (means over 30 runs) with no standard deviations, confidence intervals, or significance tests. The test sets contain only 50–62 students (Table 1, Table 4), so F1 differences of 0.01–0.05 are within plausible sampling noise. For example, in Table 5, T20P21 weeks 1–3 shows RNN-Attention-KD F1 = 0.50 versus Bi-LSTM F1 = 0.49, and T21P22 weeks 1–6 shows F1 = 0.56 versus GRU and Bi-GRU at 0.55. Without a measure of variance, the claims of 'outperforms traditional neural network models' and 'the highest average recall and F1-measure' are not statistically established. The paper should report the full distributions of the 30 runs and apply appropriate pairwise significance tests or bootstrap intervals.
  3. [§5.2, Table 6] The ablation study does not support the paper's conclusion that the hint loss (L_HD) and context-vector loss (L_CV) 'can enhance the model's prediction performance'. The F1 differences between the full model and the single-loss or two-loss variants are mostly within 0.01–0.04, and some single-loss variants outperform the full model. For instance, in T19P20 weeks 1–6, Only L_HD achieves F1 = 0.74 versus 0.72 for the full model; in T192021P22 weeks 1–5, Only L_HD achieves 0.66 versus 0.65. Additionally, Only L_KD(Soft) is often within 0.02–0.03 of the full model (e.g., T19P20 weeks 1–6: 0.72 versus 0.72). Given the lack of significance testing, the ablation should be interpreted as exploratory, and the stated conclusion about the necessity of the two losses is not supported.
minor comments (6)
  1. [§4.2.2] The hyperparameter search is described for RNN-Attention-KD, but it is unclear whether the baseline models (MLP, RNN, Bi-GRU, Bi-LSTM) received the same grid search procedure; please specify their hyperparameters or state that they were tuned identically to ensure a fair comparison.
  2. [§4.1] The features in Table 2 are described as capturing 'student activities for each lecture', but the model input is weekly aggregated data; please clarify the temporal aggregation window and whether the SRP scores are computed per week or per lecture.
  3. [§5.2, Eq. (8)] The soft-target weight λ is stated to be 0.1 in the text of Section 5.2, but Section 4.2.2 does not describe how λ was set in the hyperparameter search; please state the value, whether it was tuned, and how sensitive the results are to it.
  4. [§5.1] The explanation that the PT2019 on-site versus online modality caused the performance drop in T19P20 and T1920P21 is speculative; please soften the wording or provide supporting evidence (e.g., a feature-distribution comparison or a targeted experiment).
  5. [§2.2] The phrase 'Hinton et al. [20]'s study' is awkward; consider rephrasing to 'the study by Hinton et al. [20]'.
  6. [Abstract] The abstract states specific numeric recall and F1 values (0.49/0.51 for weeks 1–3 and 0.51/0.61 for weeks 1–6) without noting that these are averages over the six datasets; please make this explicit in both the abstract and Section 5.1.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reduction: the KD objectives are training losses evaluated against external baselines, and the only self-citation is not load-bearing.

full rationale

The paper's derivation chain is self-contained. Eq. 6 (L_HD = MSE(h^t_m, h^s_n)) and Eq. 7 (L_CV = MSE(c^t_m, c^s_n)) are training objectives that transfer teacher representations to the student; they are not fitted constants or renamed outputs. Eq. 8 combines hard and soft cross-entropy so the student still learns ground-truth labels. The teacher is trained on full-course data from prior years and the student on early weeks from the same training split, with test-year courses held out (Table 4); thus no test label or test feature enters the training objective. The central claim (Table 5) compares RNN-Attention-KD against MLP/RNN/GRU/LSTM/Bi-GRU/Bi-LSTM, and the ablation (Table 6) varies the distillation objectives; neither table's quantity is defined in terms of the quantity being predicted. The only self-citation of note is Murata et al. [31], which shares an author and introduced RNN-FitNets time-series KD; the paper cites it as prior art and does not use it to justify the reported gains. The missing no-KD RNN-Attention control is a legitimate experimental-attribution weakness, but it is a confounding control problem, not a circular reduction of an equation to its own input. Therefore no circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim depends on hand-set hyperparameters, a hand-designed SRP feature encoding, and the assumption that final-grade labels can be predicted from early weekly logs. No new entities are introduced.

free parameters (5)
  • KD soft-loss weight lambda = 0.1
    Balances soft and hard cross-entropy in Eq. 8; chosen by the authors without a reported sensitivity analysis.
  • GRU hidden units and layers = 1 layer, 4 units
    Selected by 5-fold cross-validation on the training sets; architecture choice affects all reported results.
  • SRP quantile thresholds = decile ranks 10 to 1, zero activity = 0
    Hand-designed feature encoding in Table 3 converts raw activity counts into 12 inputs; these cutoffs are not learned.
  • Decision threshold for at-risk classification = not reported (presumed 0.5)
    Precision, recall, and F1 require a probability cutoff, but the paper never states the threshold used after model output.
  • Optimizer hyperparameters = learning rate 0.01 or 0.001, batch size 8, epochs 150, weight decay 1e-5
    Fixed grid-search range; final per-dataset values are not fully disclosed.
assumptions (4)
  • domain assumption Final course grade (C, D, or F) is a valid binary at-risk label, and each student's label is available for all weekly predictions.
    Section 4.1 defines at-risk on final grade, but the weekly notation (x_i, y_i) in Section 3.1 is ambiguous.
  • domain assumption The teacher's final hidden state and context vector computed over all M weeks are suitable targets for early student models.
    Equations 6 and 7 define MSE losses that assume matching these full-course representations helps early prediction.
  • domain assumption Knowledge distilled from a teacher trained on prior-year courses transfers to the next year's students despite modality shifts (on-site vs online).
    Section 5.1 notes performance drops when training on on-site and testing on online data, so transfer is not guaranteed.
  • ad hoc to paper MSE between hidden states and context vectors is a valid similarity measure for knowledge transfer.
    Proposed in Section 3.2.2 without independent theoretical justification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance." pith.science (2026). https://pith.science/paper/W3TQ5FQS

@misc{pith2026241214526,
  author       = {Pith},
  title        = {Pith review of: Knowledge Distillation in RNN-Attention Models for Early Prediction of Student Performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W3TQ5FQS}},
  note         = {Machine review of arXiv:2412.14526}
}
read the original abstract

Educational data mining (EDM) is a part of applied computing that focuses on automatically analyzing data from learning contexts. Early prediction for identifying at-risk students is a crucial and widely researched topic in EDM research. It enables instructors to support at-risk students to stay on track, preventing student dropout or failure. Previous studies have predicted students' learning performance to identify at-risk students by using machine learning on data collected from e-learning platforms. However, most studies aimed to identify at-risk students utilizing the entire course data after the course finished. This does not correspond to the real-world scenario that at-risk students may drop out before the course ends. To address this problem, we introduce an RNN-Attention-KD (knowledge distillation) framework to predict at-risk students early throughout a course. It leverages the strengths of Recurrent Neural Networks (RNNs) in handling time-sequence data to predict students' performance at each time step and employs an attention mechanism to focus on relevant time steps for improved predictive accuracy. At the same time, KD is applied to compress the time steps to facilitate early prediction. In an empirical evaluation, RNN-Attention-KD outperforms traditional neural network models in terms of recall and F1-measure. For example, it obtained recall and F1-measure of 0.49 and 0.51 for Weeks 1--3 and 0.51 and 0.61 for Weeks 1--6 across all datasets from four years of a university course. Then, an ablation study investigated the contributions of different knowledge transfer methods (distillation objectives). We found that hint loss from the hidden layer of RNN and context vector loss from the attention module on RNN could enhance the model's prediction performance for identifying at-risk students. These results are relevant for EDM researchers employing deep learning models.

Figures

Figures reproduced from arXiv: 2412.14526 by the authors.

Figure 1
Figure 1. RNN with an attention mechanism structure (RNN [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Knowledge Distillation Framework of the RNN with an attention mechanism structure (RNN-Attention-KD). [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 40 canonical work pages

  1. [1]

    Malak Abdullah, Mahmoud Al-Ayyoub, Farah Shatnawi, Saif Rawashdeh, and Rob Abbott. 2023. Predicting students’ academic performance using e-learning logs. IAES International Journal of Artificial Intelligence (IJ-AI) 12 (06 2023), 831. https://doi.org/10.11591/ijai.v12.i2.pp831-839

  2. [2]

    Fatima Ahmed Al-azazi and Mossa Ghurab. 2023. ANN-LSTM: A deep learning model for early student performance prediction in MOOC. Heliyon 9, 4 (4 2023), e15382. https://doi.org/10.1016/j.heliyon.2023.e15382

  3. [3]

    Balqis Albreiki, Tetiana Habuza, and Nazar Zaki. 2023. Extracting topological features to identify at-risk students using machine learning and graph convolu- tional network models. International Journal of Educational Technology in Higher Education 20, 1 (04 2023), 1–22. https://doi.org/10.1186/s41239-023-00389-3

  4. [4]

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In 3rd International Conference on Learning Representations, ICLR 2015, May 7-9, 2015, Conference Track Proceedings . Cornell University, San Diego, CA, USA, 15 pages. http: //arxiv.org/abs/1409.0473

  5. [5]

    Brijesh Kumar Baradwaj and Saurabh Pal. 2011. Mining Educational Data to Analyze Students Performance. International Journal of Advanced Computer Science and Applications 2, 6 (2011), 7 pages. https://doi.org/10.14569/IJACSA. 2011.020609

  6. [6]

    Simard, and Paolo Frasconi

    Yoshua Bengio, Patrice Y. Simard, and Paolo Frasconi. 1994. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks 5, 2 (3 1994), 157–166. https://doi.org/10.1109/72.279181

  7. [7]

    Gabriella Casalino, Giovanna Castellano, Andrea Mannavola, and Gennaro Ves- sio. 2020. Educational Stream Data Analysis: A Case Study. In 2020 IEEE 20th Mediterranean Electrotechnical Conference ( MELECON) . IEEE, New York, NY, USA, 232–237. https://doi.org/10.1109/MELECON48756.2020.9140510

  8. [8]

    Cheng-Huan Chen, Stephen J. H. Yang, Jian-Xuan Weng, Hiroaki Ogata, and Chien-Yuan Su. 2021. Predicting at-risk university students based on their e-book reading behaviours by using machine learning classifiers. Australasian Journal of Educational Technology 37, 4 (6 2021), 130–144. https://doi.org/10.14742/ajet.6116

Show all 60 references
  1. [9]

    Fu Chen and Ying Cui. 2020. Utilizing Student Time Series Behaviour in Learning Management Systems for Early Prediction of Course Performance. Journal of Learning Analytics 7, 2 (Sep. 2020), 1–17. https://doi.org/10.18608/jla.2020.72.1

  2. [10]

    Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker

  3. [11]

    Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Rep- resentations using RNN Encoder–Decoder for Statistical Machine Translation. In 2014 Conference on Empirical Methods in Natural ...

  4. [12]

    Ouafae El Aissaoui, Yasser El Alami El Madani, Lahcen Oughdir, Ahmed Dakkak, and Youssouf El Allioui. 2020. A Multiple Linear Regression-Based Approach to Predict Student Performance. In Advanced Intelligent Systems for Sustainable Development (AI2SD’2019). Springer Internatio...

  5. [13]

    Jo Shan Fu. 2013. ICT in education: A critical literature review and its impli- cations. International Journal of Education and Development Using Informa- tion and Communication Technology (IJEDICT) 9, 1 (01 2013), 112–125. https: //files.eric.ed.gov/fulltext/EJ1182651.pdf

  6. [14]

    Filippos Giannakas, Christos Troussas, Ioannis Voyiatzis, and Cleo Sgouropoulou

  7. [15]

    Yann Ling Goh, Yeh Huann Goh, Chun-Chieh Yip, Chen Hunt Ting, Kah Pin Chen, and Raymond Ling Leh Bin. 2020. Prediction of Students’ Academic Performance by K-Means Clustering. International Journal of Advanced Science and Technology 29, 10s (4 2020), 1–6. https://doi.org/10.31...

  8. [16]

    Maybank, and Dacheng Tao

    Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowl- edge Distillation: A Survey.International Journal of Computer Vision 129, 6 (March 2021), 1789–1819. https://doi.org/10.1007/s11263-021-01453-z

  9. [17]

    Alamri, Ranim S

    Leena H. Alamri, Ranim S. Almuslim, Mona S. Alotibi, Dana K. Alkadi, Irfan Ullah Khan, and Nida Aslam. 2021. Predicting Student Academic Performance using Support Vector Machine and Random Forest. In Proceedings of the 2020 3rd International Conference on Education Technology ...

  10. [18]

    Yanbai He, Rui Chen, Xinya Li, Chuanyan Hao, Sijiang Liu, Gangyao Zhang, and Bo Jiang. 2020. Online At-Risk Student Identification using RNN-GRU Joint Neural Networks. Information 11, 10 (10 2020), 474. https://doi.org/10.3390/ info11100474

  11. [19]

    Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao

    Arto Hellas, Petri Ihantola, Andrew Petersen, Vangel V. Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen, Chris Messom, and Soohyun Nam Liao. 2018. Predicting academic performance: a systematic literature review. In Proceedings Companion of the 23rd Annual ...

  12. [20]

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. https://doi.org/10.48550/arxiv.1503.02531

  13. [21]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-Term Memory. Neural Computation 9, 8 (11 1997), 1735–1780. https://doi.org/10.1162/neco.1997. 9.8.1735 SAC ’25, March 31-April 4, 2025, Catania, Italy Sukrit Leelaluk, Cheng Tang, Valdemar Švábenský, and Atsushi Shimada

  14. [22]

    Ya-Han Hu, Chia-Lun Lo, and Sheng-Pao Shih. 2014. Developing early warning systems to predict students’ online learning performance. Computers in Human Behavior 36 (7 2014), 469–478. https://doi.org/10.1016/j.chb.2014.04.002

  15. [23]

    Mushtaq Hussain, Wenhao Zhu, Wu Zhang, and Syed Muhammad Raza Abidi

  16. [24]

    Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for Natural Language Understand- ing. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu ...

  17. [25]

    Byung-Hak Kim, Ethan Vizitei, and Varun Ganapathi. 2018. GritNet: Student Performance Prediction with Deep Learning.. InEDM, Kristy Elizabeth Boyer and Michael Yudelson (Eds.). International Educational Data Mining Society (IEDMS), USA, 5 pages. https://doi.org/10.48550/arxiv....

  18. [26]

    Charles Koutcheme, Sami Sarsa, Arto Hellas, Lassi Haaranen, and Juho Leinonen

  19. [27]

    Sukrit Leelaluk, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada. 2022. Predicting student performance based on Lecture Materials data using Neural Network Models. CEUR Workshop Proceedings 3120 (2022), 11–20. https://ceur-ws.org/Vol-3120/paper2.pdf

  20. [28]

    Lopez Z, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada

    Erwin D. Lopez Z, Tsubasa Minematsu, Yuta Taniguchi, Fumiya Okubo, and Atsushi Shimada. 2022. Assessment of At-Risk Students’ Predictions From e- Book Activities Representations in Practical Applications. In 30th International Conference on Computers in Education, ICCE 2022 - ...

  21. [29]

    Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective Ap- proaches to Attention-based Neural Machine Translation, In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lluís Màrquez, Chris Callison-Burch, Jian Su, Daniele Pigh...

  22. [30]

    Nur Izzati Mohd Talib, Nazatul Aini Abd Majid, and Shahnorbanun Sahran. 2023. Identification of Student Behavioral Patterns in Higher Education Using K-Means Clustering and Support Vector Machine. Applied Sciences 13, 5 (3 2023), 3267. https://doi.org/10.3390/app13053267

  23. [31]

    Ryusuke Murata, Fumiya Okubo, Tsubasa Minematsu, Yuta Taniguchi, and At- sushi Shimada. 2023. Recurrent Neural Network-FitNets: Improving Early Pre- diction of Student Performanceby Time-Series Knowledge Distillation. Journal of Educational Computing Research 61, 3 (6 2023), 6...

  24. [32]

    Hiroaki Ogata, Chengjiu Yin, Misato Oi, Fumiya Okubo, Atsushi Shimada, Ken- taro Kojima, and Masanori Yamada. 2015. E-book-based learning analytics in University education. In Proceedings of the 23rd International Conference on Com- puters in Education, ICCE 2015 . Asia-Pacifi...

  25. [33]

    Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, and Shin’ichi Konomi

  26. [34]

    Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, and Hiroaki Ogata. 2017. A neural network approach for students’ performance prediction. In Proceedings of the Seventh International Learning Analytics & Knowledge Conference , Marek Hatala, Alyssa Wis, Phil Winne, Grace Lync...

  27. [35]

    Fumiya Okubo, Takayoshi Yamashita, Atsushi Shimada, Yuta Taniguchi, and Konomi Shin’ichi. 2018. On the prediction of students’ quiz score by recurrent neural network. CEUR Workshop Proceedings 2163 (2018), 6 pages. https://ceur- ws.org/Vol-2163/paper3.pdf

  28. [36]

    Feiyue Qiu, Guodao Zhang, Xin Sheng, Lei Jiang, Lijia Zhu, Qifeng Xiang, Bo Jiang, and Ping-Kuo Chen. 2022. Predicting students’ performance in e-learning using learning process and behaviour data. Scientific Reports 12, 1 (01 2022). https://doi.org/10.1038/s41598-021-03867-8

  29. [37]

    Moises Riestra-González, Maria del Puerto Paule-Ruíz, and Francisco Ortin. 2021. Massive LMS log data analysis for the early prediction of course-agnostic student performance. Computers & Education 163 (4 2021), 104108. https://doi.org/10. 1016/j.compedu.2020.104108

  30. [38]

    In Proceedings of the 25th International Con- ference on Computers in Education, ICCE 2017 - Main Conference Pro- ceedings

    Students’ performance prediction using data of multiple courses by recurrent neural network. In Proceedings of the 25th International Con- ference on Computers in Education, ICCE 2017 - Main Conference Pro- ceedings. Asia-Pacific Society for Computers in Education, Taiwan, 439–

  31. [39]

    T M Noviyanti Sagala, Syarifah Diana Permai, Alexander Agung Santoso Gu- nawan, Rehnianty Octora Barus, and Cito Meriko. 2022. Predicting Computer Science Student’s Performance using Logistic Regression. In2022 5th International Seminar on Research of Information Technology an...

  32. [40]

    Sungho Shin, Joosoon Lee, Junseok Lee, Yeonguk Yu, and Kyoobin Lee. 2022. Teaching Where to Look: Attention Similarity Knowledge Distillation for Low Resolution Face Recognition, In Computer Vision – ECCV 2022, Shai Avidan, Gabriel J. Brostow, Moustapha Cissé, Giovanni Maria F...

  33. [41]

    Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS’14, Vol. abs/1409.3215), Zoubin Ghahramani,...

  34. [42]

    Francisco Segura Altamirano

    Luis Vives, Ivan Cabezas, Juan Carlos Vives, Nilton German Reyes, Janet Aquino, Jose Bautista Cóndor, and S. Francisco Segura Altamirano. 2024. Prediction of Students’ Academic Performance in the Programming Fundamentals Course Using Long Short-Term Memory Neural Networks. IEE...

  35. [43]

    Baker, Pavel Čeleda, Jan Vykopal, Jens Mache, and Ankur Chattopadhyay

    Valdemar Švábenský, Kristián Tkáčik, Aubrey Birdwell, Richard Weiss, Ryan S. Baker, Pavel Čeleda, Jan Vykopal, Jens Mache, and Ankur Chattopadhyay. 2024. Detecting Unsuccessful Students in Cybersecurity Exercises in Two Different Learning Environments. In Proceedings of the 54...

  36. [44]

    Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. arXiv:1412.6550 [cs.LG] https://arxiv.org/abs/1412.6550

  37. [45]

    Kai Wang, Fei Yang, and Joost van de Weijer. 2022. Attention Distillation: self- supervised vision transformer students need more guidance.. In BMVC. BMVA Press, UK, 26 pages. https://doi.org/10.48550/arxiv.2210.00944

  38. [46]

    Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2021. MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers. In Findings of the Association for Computational Lin- guistics: ACL-IJCNLP 2021 , Chengqing Zong, Fei Xia, We...

  39. [47]

    Xing Wei, Yuqing Liu, Jiajia Li, Huiyong Chu, Zichen Zhang, Feng Tan, and Pengwei Hu. 2022. B-AT-KD: Binary attention map knowledge distillation. Neu- rocomputing 511 (9 2022), 299–307. https://doi.org/10.1016/j.neucom.2022.09.064

  40. [48]

    Jacob Whitehill, Kiran Mohan, Daniel Seaton, Yigal Rosen, and Dustin Tingley

  41. [49]

    Mariana Windarti and Putri Taqwa Prasetyaninrum. 2020. Prediction Analysis Student Graduate Using Multilayer Perceptron. InProceedings of the International Conference on Online and Blended Learning 2019 (ICOBL 2019) . Atlantis Press, Paris, France, 5 pages. https://doi.org/10....

  42. [50]

    Yue Xie, Ruiyu Liang, Zhenlin Liang, Chengwei Huang, Cairong Zou, and Björn Schuller. 2019. Speech Emotion Classification Using Attention-Based LSTM. IEEE/ACM Transactions on Audio, Speech, and Language Processing 27, 11 (7 2019), 1675–1685. https://doi.org/10.1109/TASLP.2019.2925934

  43. [51]

    Aljohani, Guanliang Chen, and Dragan Gasevic

    Hajra Waheed, Saeed-Ul Hassan, Raheel Nawaz, Naif R. Aljohani, Guanliang Chen, and Dragan Gasevic. 2023. Early prediction of learners at risk in self-paced education: A neural network approach. Expert Systems with Applications 213 (3 2023), 118868. https://doi.org/10.1016/j.es...

  44. [52]

    Yupei Zhang, Yue Yun, Rui An, Jiaqi Cui, Huan Dai, and Xuequn Shang. 2021. Educational Data Mining Techniques for Student Performance Prediction: Method Review and Comparison Analysis. Frontiers in Psychology 12 (12 2021), 19 pages. https://doi.org/10.3389/fpsyg.2021.698490

  45. [56]

    https://doi.org/ 10.48550/arxiv.1702.06404 arXiv:1702.06404 [cs.AI]

    Delving Deeper into MOOC Student Dropout Prediction. https://doi.org/ 10.48550/arxiv.1702.06404 arXiv:1702.06404 [cs.AI]

  46. [59]

    Sergey Zagoruyko and Nikos Komodakis. 2017. Paying More Attention to Atten- tion: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR. OpenReview.net, Online, 13 pages. https://arxiv.org/abs/1612. 03928

  47. [444]

    https://kyushu-u.pure.elsevier.com/en/publications/students-performance- prediction-using-data-of-multiple-courses-by

  48. [2017]

    Wallach, Rob Fergus, S

    Learning Efficient Object Detection Models with Knowledge Distilla- tion, In Advances in Neural Information Processing Systems, Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). Neural Information Pr...

  49. [2018]

    Computational Intelligence and Neuro- science 2018, 1 (10 2018), 6347186

    Student Engagement Predictions in an e-Learning System and Their Impact on Student Course Assessment Scores. Computational Intelligence and Neuro- science 2018, 1 (10 2018), 6347186. https://doi.org/10.1155/2018/6347186

  50. [2021]

    Applied Soft Computing 106 (7 2021), 107355

    A deep learning classification framework for early prediction of team- based academic performance. Applied Soft Computing 106 (7 2021), 107355. https://doi.org/10.1016/j.asoc.2021.107355

  51. [2022]

    In Proceed- ings of the 24th Australasian Computing Education Conference (ACE ’22) , Judy Sheard and Paul Denny (Eds.)

    Methodological Considerations for Predicting At-risk Students. In Proceed- ings of the 24th Australasian Computing Education Conference (ACE ’22) , Judy Sheard and Paul Denny (Eds.). Association for Computing Machinery, New York, NY, USA, 105–113. https://doi.org/10.1145/35118...

  52. [5898]

    https://doi.org/10.1109/ACCESS.2024.3350169

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.