Pith. sign in

REVIEW 3 major objections 4 minor 185 references

Towards Human-Centered Early Prediction Models for Academic Performance in Real-World Contexts

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Using the first week of passively sensed phone and wearable data plus self-reports, the paper claims end-of-term GPA category can be predicted within the same term with 0.81–0.92 balanced accuracy, though cross-year prediction remains…

desk verdict Promising within-term first-week prediction, but the abstract's week-one early-warning claim only works with same-term labels; the honest cross-year test puts LR and 1D-CNN at chance. read the letter →

arxiv 2504.12236 v2 pith:PBT4KJV5 submitted 2025-04-16 cs.HC

classification cs.HC
keywords academicperformancepredictionearlypassivesensinghuman-centeredmachinelearningexplainabilityfairnessgeneralizabilitymulti-task
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether end-of-term academic risk can be detected from a single week of passively sensed phone and wearable data combined with self-reports, and whether the resulting models are explainable, fair, and generalizable enough for real use. Its central claim is that logistic regression and a one-dimensional convolutional network classify students as low or high performers from first-week data with balanced accuracy around 0.81 to 0.92, matching earlier studies that required four or more weeks of data and therefore shifting early prediction three weeks earlier. The paper reports that this holds when models are trained and tested within the same term; when trained on 2018 and applied to 2019, these two models drop to majority-class baseline level (balanced accuracy 0.500 and 0.559), and only a multi-task variant reaches 0.712. The paper's main finding is a trade-off map: the transparent logistic model is reasonably fair but does not transfer across years, the 1D-CNN is the most accurate within a term but opaque, and the multi-task model transfers best but shows the largest fairness gaps for protected groups. It also flags that about 21% of students participated in both years, a caveat for cross-year comparisons.

What carries the argument

The load-bearing object is the first-week feature set built from continuous phone and Fitbit streams: phone unlock counts, Bluetooth scans, location bouts, sleep, steps, physical activity, and self-reported stress/health scales, organized into five daily epochs (morning, afternoon, evening, night, full day) and augmented with behavioral-change slopes and breakpoints. Three pipelines sit on top: logistic regression with correlation-based feature selection and SMOTE oversampling; a 1D-CNN over daily time series; and a multi-task 1D-CNN whose secondary task predicts prior-term Winter GPA, letting it train on 2018 and 2019 data without using the test year's end-of-term labels. The features supply the week-one signal; the multi-task auxiliary task is what gives the only demonstrated cross-year transfer.

What would settle it

Run a prospective deployment in a new Spring term: freeze the LR and 1D-CNN pipelines on data from previous terms before the term starts, collect first-week sensing, and compare week-one predictions against end-of-term GPA. The paper's Table 5 already supplies the predicted outcome for this test—balanced accuracy 0.500 and 0.559, equal to the majority-class baseline—so a prospective run would settle whether the within-term numbers transfer to real early warning.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that end-of-term GPA category can be predicted from the first week of a term: in Spring 2018 and Spring 2019, logistic regression and a 1D-CNN achieve balanced accuracy between 0.814 and 0.918 in classifying students as low or high performers, comparable to the prior four-week benchmark and three weeks earlier. The paper argues this demonstrates that at-risk students can be identified at week one. It is equally explicit, in Section 5.4, that the same two approaches do not carry that accuracy across terms: trained on 2018 and tested on 2019, they perform at the level of a majority-class baseline, while the multi-task 1D-CNN, which adds prior-term Winter GPA prediction as a secondary task, reaches 0.712 balanced accuracy. Across explainability, fairness, and generalizability, no single approach dominates: LR is transparent and fair in several comparisons but weak cross-year, 1D-CNN is strongest within a term but a black box, and MTL-1D-CNN generalizes best while showing the largest fairness gaps. The paper also notes the presence of returning students (about 21.3% year-to-year retention) as a reliability caveat.

Load-bearing premise

The load-bearing premise is that a model trained and tested on students from the same term can stand in for real early-warning use; in deployment, the model would have to be trained before the term's grades exist, and the paper's own cross-year test shows two of its three approaches then perform no better than predicting everyone as a high performer.

Editorial extensions

If this is right

  • If the within-term results hold, student support teams could begin outreach by the end of the first week of a term, roughly three weeks earlier than the previous state of the art, widening the window for intervention.
  • The logistic-regression feature rankings suggest concrete early warning signs—class attendance, timing of phone use, restless sleep, and self-reported stressors—that advisors could act on even without the model.
  • The Friday breakpoint seen in both years implies that many students' weekend behavior starts on Friday, so end-of-week check-ins may be a better-timed intervention point than midweek messages.
  • Because LR and 1D-CNN fail on the cross-year test, a deployed early-warning system cannot reuse last year's model; the MTL variant is the paper's only candidate for that role, and it still requires further validation.
  • Fairness varies by protected group and metric across all approaches, so an institution adopting any of these models should audit it on its own student population before acting on its outputs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's week-one claim is a same-term result; in a realistic deployment where the model must be trained before the term's grades exist, the Table 5 cross-year numbers (0.500 and 0.559 balanced accuracy for LR and 1D-CNN) are the more relevant estimate.
  • Prior-term GPA alone (the 1R-SVM baseline) performs at 0.672 balanced accuracy cross-year and perfectly identifies students who stay low performers, so the marginal value of week-one sensing over transcripts deserves a direct held-out comparison.
  • The Thursday/Friday behavioral breakpoint is testable as an intervention design: randomize students to a Friday-focused support message versus a Monday-focused one and compare end-of-term GPA.
  • The association between phone service provider and GPA, which the paper reads as a proxy for unmeasured socioeconomic context, could be resolved by collecting direct income or family-background measures and checking whether the association disappears.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents three modeling approaches (logistic regression, 1D-CNN, and MTL-1D-CNN) to classify college students as low or high academic performers (GPA below/above 3.2) using passive sensing and self-report data collected within the first week of two Spring terms (2018 and 2019). The authors report high within-year balanced accuracy (0.81-0.92) and evaluate the models on explainability, fairness, and generalizability, concluding that trade-offs exist across these human-centered principles. The central claim is that at-risk students can be identified as early as week one, which is stated in the abstract and Section 1; however, the paper's own cross-year evaluation (Table 5) shows the LR and 1D-CNN models trained on 2018 and tested on 2019 perform at or near the majority-class baseline (balanced accuracy 0.500 and 0.559, respectively).

Significance. The work addresses an important gap in the learning analytics and HCI literature by combining passive behavioral sensing with explicit consideration of explainability, fairness, and generalizability, and by evaluating on real longitudinal data. The within-year results, if interpreted as a retrospective feasibility study, are valuable: they show that first-week behavioral features carry signal about end-of-term GPA. The paper also honestly reports the cross-year failure of the two simpler models and includes a multi-task approach that achieves more encouraging cross-year balanced accuracy (0.712). These strengths are, however, undermined by the mismatch between the abstract's 'week one' claim and the deployment-relevant protocol, as well as by acknowledged test-data leakage in the LR feature-selection loop. The empirical contribution is real, but the claims need re-scoping before the paper can be accepted.

major comments (3)
  1. [Abstract and Section 1 vs. Section 5.4 and Table 5] The abstract's claim that 'these models can identify at-risk students as early as week one' is not supported by the deployment-relevant protocol. The balanced accuracies in Table 2 (LR 0.814/0.858, 1D-CNN 0.918/0.866) come from within-year training and testing, meaning the model uses end-of-term labels of the same term it is predicting. In a real-world week-one early-warning deployment, those labels do not exist at prediction time; the appropriate protocol is training on past terms and testing on a new term. Under that protocol (Table 5), LR achieves 0.500 and 1D-CNN achieves 0.559 balanced accuracy, essentially matching the majority-class baseline of 0.500. Section 5.4.1 explicitly concedes that neither approach outperforms the baseline on unseen data. The abstract and Section 1 should therefore re-scope the claim to distinguish retrospective within-year feasibility from cross-year deployment performance, or make the MTL model, which reaches 0.712, the focus of the early-prediction claim with its additional assumptions explicitly stated.
  2. [Appendix C.2.4] The LR feature-selection step selects the correlation threshold r by maximizing a_test - a_train on the held-out test subject within each LOSO-CV fold. This is a test-data leak in the model-selection loop and can inflate the 2018 within-year LR results reported in Table 2. The authors acknowledge the leakage but dismiss it as having 'no leakage' for the 2019 data. Since the within-year LR numbers are a primary support for the week-one claim, the paper must either re-run the 2018 analysis with r chosen only from training folds, or explicitly report the 2018 LR result as an optimistic upper bound and base the claim on the frozen-pipeline 2019 result. As written, the evidence for the headline claim is overstated.
  3. [Section 5.1.1] The statement that the robust performance 'across the 2018 and 2019 datasets provides strong evidence that early predictions of student performance ... are possible' conflates pipeline-level generalization with model-level generalization. In Table 2, the models are re-fit in each year using that year's own end-of-term labels, so the across-year consistency does not demonstrate that a model pre-trained on one term can make week-one predictions in a new term. The paper introduces the pipeline-level versus model-level distinction only later in Section 5.4; Section 5.1.1 should either be qualified at the point of the claim or the generalizability section should be moved earlier so readers can correctly interpret the early-prediction results.
minor comments (4)
  1. [Section 5] The word 'approachs' appears in the final paragraph of Section 5.3.2 and the word 'reffered' appears in Section 5.4; these typos should be corrected.
  2. [Figure 5 caption] The caption repeats '(a) 2018 spring term GPA' twice and is missing a description for one panel; this should be fixed for clarity.
  3. [Table 2] The 'Earliest predictable time' column lists 'before Spring term' for 1R-SVM, which uses prior-term GPA; since this baseline does not use first-week data, the column header should be clarified (e.g., 'earliest data time') to avoid implying a direct comparison on prediction timing.
  4. [Section 5.3] The fairness thresholds (difference between -0.1 and 0.1, ratio between 0.8 and 1.2) are taken from demographic-parity conventions and extended to equalized odds and equal opportunity without justification; the authors should add a sentence explaining why these thresholds are appropriate for the other two metrics.

Circularity Check

1 steps flagged · score 2.0 of 10

No load-bearing circularity; one acknowledged test-data leakage in LR feature-threshold selection affects only the 2018 within-term row, while 2019 and cross-year results remain independent.

  1. fitted input called prediction [Appendix C.2.4 (Feature Selection in the LR preprocessing pipeline)]
    "For each round of LOSO-CV, we performed a grid search to determine an optimal correlation threshold 𝑟, selecting the 𝑟 value that maximized the performance advantage (𝑎𝑑𝑖𝑓𝑓 = 𝑎𝑡𝑒𝑠𝑡 − 𝑎𝑡𝑟𝑎𝑖𝑛). We note that while the use of test data in determining 𝑟 introduces some leakage, this was only during feature inclusion, and no leakage occurred when applied to the 2019 data."

    Within every LOSO-CV fold, the held-out test subject's labels are used to select the feature-selection threshold r by maximizing atest - atrain, and Table 2 reports the resulting test-fold balanced accuracy for LR in 2018. The reported performance is therefore the same objective used to select r, making the 2018 LR accuracy partly a model-selection result rather than an unbiased prediction. The paper itself concedes that 'the use of test data in determining r introduces some leakage.' This is a real but partial instance of fitted-input-called-prediction: it affects the 2018 within-year LR row, while the 2019 numbers and the cross-year Table 5 provide independent checks, so the central week-one claim does not reduce to this leakage.

full rationale

This is an empirical study rather than a mathematical derivation, so most circularity patterns do not apply. The within-term training-and-test design (Figure 2a) uses same-term end-of-term GPA labels; this is legitimate retrospective cross-validation but not a deployment-protocol test, and the paper itself reports the deployment-like cross-year experiment in Table 5, where LR and 1D-CNN fall to 0.500 and 0.559 balanced accuracy. That mismatch is an external-validity and correctness concern, not circularity: no equation or definition makes the week-one claim equivalent to the within-term evaluation. The only concrete circularity-like step is the acknowledged use of test-fold labels to choose the LR correlation threshold r in Appendix C.2.4, which inflates the 2018 LR row of Table 2. This is disclosed, localized, and does not taint the 2019 hold-out or the cross-year comparison, so the central HCML evaluation retains independent content. Self-citations to GLOBEM [160], Sefidgar et al. [134], and Xu et al. [157] supply data and terminology, but no load-bearing uniqueness theorem or derivation is imported from them. Score 2 reflects one minor, partially self-referential evaluation step amid an otherwise self-contained benchmark study.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the assumption that first-week sensor and self-report data, after imputation and feature selection, contain enough signal to predict end-of-term GPA. The LR pipeline introduces several hand-chosen thresholds (GPA 3.2, collinearity 0.7, class attendance 50%, step threshold 12/min) and one test-tuned parameter (CFS threshold r). The CNN adds hyperparameters tuned on the 2018 training set. No new entities are postulated. The main axiomatic load is the independence of training and test data, which is violated in the LR feature selection step.

free parameters (5)
  • CFS correlation threshold r = chosen per LOSO fold via grid search maximizing a_test - a_train
    Feature selection in the LR pipeline selects r using test data (Appendix C.2.4), introducing leakage.
  • 1D-CNN hyperparameters = lr=0.0001, dropout=85%, epochs=150, batch size=6
    Selected via grid search on the 2018 training set (Appendix C.3.4); central to model performance.
  • GPA cutoff 3.2 = 3.2
    Chosen by authors to match university average GPA (Section 3.4); a modeling choice rather than a fitted value.
  • Class attendance threshold 50% = 50%
    Used to define class attendance from location data (Appendix B.2); an arbitrary threshold.
  • Fairness difference/ratio thresholds = difference 0.1, ratio 0.8 to 1.2
    Adopted from literature for demographic parity and extended to equalized odds and equal opportunity (Section 5.3).
assumptions (5)
  • domain assumption Train and test sets are independent, with no data leakage in model selection.
    Violated by CFS feature selection on the 2018 dataset (Appendix C.2.4), which uses test data to pick r.
  • domain assumption End-of-term GPA is a valid and reliable outcome measure of academic performance.
    Used as ground truth throughout (Section 3.4); the paper does not validate GPA against other outcomes.
  • domain assumption First-week sensor and self-report data are sufficiently complete and representative after preprocessing.
    Missing data rates up to 33% in 2018 (Table 7); imputation and feature removal assume ignorable missingness.
  • domain assumption Protected group definitions based on self-report are accurate and stable.
    Fairness analysis relies on these labels (Table 1); misclassification would bias fairness metrics.
  • domain assumption Fairness metrics computed on small subgroups are statistically reliable.
    No confidence intervals are reported; groups such as sexual minorities have n=21, making difference and ratio estimates noisy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Human-Centered Early Prediction Models for Academic Performance in Real-World Contexts." pith.science (2026). https://pith.science/paper/PBT4KJV5

@misc{pith2026250412236,
  author       = {Pith},
  title        = {Pith review of: Towards Human-Centered Early Prediction Models for Academic Performance in Real-World Contexts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PBT4KJV5}},
  note         = {Machine review of arXiv:2504.12236}
}
read the original abstract

Supporting student success requires collaboration among multiple stakeholders. Researchers have explored machine learning models for academic performance prediction; yet key challenges remain in ensuring these models are interpretable, equitable, and actionable within real-world educational support systems. First, many models prioritize predictive accuracy but overlook human-centered machine learning principles, limiting trust among students and reducing their usefulness for educators and institutional decision-makers. Second, most models require at least a month of data before making reliable predictions, delaying opportunities for early intervention. Third, current models primarily rely on sporadically collected, classroom-derived data, missing broader behavioral patterns that could provide more continuous and actionable insights. To address these gaps, we present three modeling approaches-LR, 1D-CNN, and MTL-1D-CNN-to classify students as low or high academic performers. We evaluate them based on explainability, fairness, and generalizability to assess their alignment with key social values. Using behavioral and self-reported data collected within the first week of two Spring terms, we demonstrate that these models can identify at-risk students as early as week one. However, trade-offs across human-centered machine learning principles highlight the complexity of designing predictive models that effectively support multi-stakeholder decision-making and intervention strategies. We discuss these trade-offs and their implications for different stakeholders, outlining how predictive models can be integrated into student support systems. Finally, we examine broader socio-technical challenges in deploying these models and propose future directions for advancing human-centered, collaborative academic prediction systems.

Figures

Figures reproduced from arXiv: 2504.12236 by the authors.

Figure 1
Figure 1. Overview of the whole modeling pipeline for the three approaches. All three approaches utilize the same data sources [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Overview of the training (highlighted in light gray) and testing process for the three approaches. (a) shows the training [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Radar charts comparing the fairness performance of three approaches (LR, 1D-CNN, MTL-1D-CNN) across four [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Accuracy of three approaches as well as the baselines in predicting academic performance consistency and transitions. [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Distributions of spring term GPA among all students and distribution of spring term GPA of high and low performers [PITH_FULL_IMAGE:figures/full_fig_p032_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

185 extracted references · 66 canonical work pages

  1. [1]

    Daniel A Adler, Emily Tseng, Khatiya C Moon, John Q Young, John M Kane, Emanuel Moss, David C Mohr, and Tanzeem Choudhury

  2. [2]

    Daniel A Adler, Fei Wang, David C Mohr, and Tanzeem Choudhury. 2022. Machine learning for passive mental health symptom prediction: Generalization across different longitudinal mobile sensing studies. Plos one 17, 4 (2022), e0266516

  3. [3]

    David W Aha, Dennis Kibler, and Marc K Albert. 1991. Instance-based learning algorithms. Machine learning 6, 1 (1991), 37–66

  4. [4]

    Samantha J Ahern. 2024. The potential and pitfalls of learning analytics as a tool for supporting student wellbeing. (2024)

  5. [5]

    Muhammad Aurangzeb Ahmad, Arpit Patel, Carly Eckert, Vikas Kumar, and Ankur Teredesai. 2020. Fairness in machine learning for healthcare. In Proceedings of the 26th acm sigkdd international conference on knowledge discovery & data mining . 3529–3530

  6. [6]

    Amanda Aird, Paresha Farastu, Joshua Sun, Elena Stefancová, Cassidy All, Amy Voida, Nicholas Mattei, and Robin Burke. 2024. Dynamic fairness-aware recommendation through multi-agent social choice. ACM Transactions on Recommender Systems 3, 2 (2024), 1–35

  7. [7]

    Abdulmajeed Al-Drees, Hamza Abdulghani, Mohammad Irshad, Abdulsalam Ali Baqays, Abdulaziz Ali Al-Zhrani, Sulaiman Abdullah Alshammari, and Norah Ibrahim Alturki. 2016. Physical activity and academic achievement among the medical students: A cross- sectional study. Medical teacher 38, sup1 (2016), S66–S72

  8. [8]

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 chi conference on human factors in 23 , , Zhang et al. computing systems. 1–13

Show all 185 references
  1. [9]

    Tariq Osman Andersen, Francisco Nunes, Lauren Wilcox, Enrico Coiera, and Yvonne Rogers. 2023. Introduction to the special issue on human-centred AI in healthcare: Challenges appearing in the wild. , 12 pages

  2. [10]

    Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al . 2020. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities a...

  3. [11]

    Shervin Assari, Ehsan Moazen-Zadeh, Cleopatra Howard Caldwell, and Marc A Zimmerman. 2017. Racial discrimination during adolescence predicts mental health deterioration in adulthood: Gender differences among Blacks. Frontiers in public health 5 (2017), 104

  4. [12]

    Christoph Augner and Gerhard W Hacker. 2012. Associations between problematic mobile phone use and psychological parameters in young adults. International journal of public health 57, 2 (2012), 437–441

  5. [13]

    Pranjal Awasthi, Alex Beutel, Matthäus Kleindessner, Jamie Morgenstern, and Xuezhi Wang. 2021. Evaluating fairness of machine learning models under uncertain and incomplete information. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. 206–214

  6. [14]

    Nikola Banovic, Zhuoran Yang, Aditya Ramesh, and Alice Liu. 2023. Being trustworthy is not enough: How untrustworthy artificial intelligence (AI) can deceive the end-users and gain their trust. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023), 1–17

  7. [15]

    Solon Barocas and Andrew D Selbst. 2016. Big data’s disparate impact. Calif. L. Rev. 104 (2016), 671

  8. [16]

    Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020. Explainable Artificial Intelligence (XAI): Concepts...

  9. [17]

    Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José MF Moura, and Peter Eckersley. 2020. Explainable machine learning in deployment. In Proceedings of the 2020 conference on fairness, accountability, and transparenc...

  10. [18]

    Sarah Bird, Miro Dudík, Richard Edgar, Brandon Horn, Roman Lutz, Vanessa Milan, Mehrnoosh Sameki, Hanna Wallach, and Kathleen Walker. 2020. Fairlearn: A toolkit for assessing and improving fairness in AI. Microsoft, Tech. Rep. MSR-TR-2020-32 (2020)

  11. [19]

    Christopher M Bishop and Nasser M Nasrabadi. 2006. Pattern recognition and machine learning . Vol. 4. Springer

  12. [20]

    Shahab Boumi and Adan Ernesto Vela. 2021. Quantifying the Impact of Students’ Semester Course Load on Their Academic Performance. In 2021 ASEE Virtual Annual Conference Content Access

  13. [21]

    Javier Bravo-Agapito, Sonia J Romero, and Sonia Pamplona. 2021. Early prediction of undergraduate Student’s academic performance in completely online learning: A five-year study. Computers in Human Behavior 115 (2021), 106595

  14. [22]

    Leo Breiman. 2001. Random forests. Machine learning 45, 1 (2001), 5–32

  15. [23]

    Kay Henning Brodersen, Cheng Soon Ong, Klaas Enno Stephan, and Joachim M Buhmann. 2010. The balanced accuracy and its posterior distribution. In 2010 20th international conference on pattern recognition . IEEE, 3121–3124

  16. [24]

    Maarten Buyl and Tijl De Bie. 2024. Inherent limitations of AI fairness. Commun. ACM 67, 2 (2024), 48–55

  17. [25]

    R Caruana. 1993. Multitask learning: A knowledge-based source of inductive bias1. In Proceedings of the Tenth International Conference on Machine Learning. Citeseer, 41–48

  18. [26]

    Rich Caruana. 1997. Multitask learning. Machine learning 28 (1997), 41–75

  19. [27]

    Diogo V Carvalho, Eduardo M Pereira, and Jaime S Cardoso. 2019. Machine learning interpretability: A survey on methods and metrics. Electronics 8, 8 (2019), 832

  20. [28]

    Laetitia Cassells. 2018. The effectiveness of early identification of ‘at risk’students in higher education institutions. Assessment & Evaluation in Higher Education 43, 4 (2018), 515–526

  21. [29]

    Stevie Chancellor. 2023. Toward practices for human-centered machine learning. Commun. ACM 66, 3 (2023), 78–85

  22. [30]

    Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002. SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002), 321–357

  23. [31]

    Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. 2018. Recurrent neural networks for multivariate time series with missing values. Scientific reports 8, 1 (2018), 1–12

  24. [32]

    Fu Chen and Ying Cui. 2020. Utilizing Student Time Series Behaviour in Learning Management Systems for Early Prediction of Course Performance. Journal of Learning Analytics 7, 2 (2020), 1–17

  25. [33]

    Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining . 785–794

  26. [34]

    Sheldon Cohen and Harry M Hoberman. 1983. Positive events and social supports as buffers of life change stress 1. Journal of applied social psychology 13, 2 (1983), 99–125. 24 Towards Human-Centered Early Academic Performance Prediction Models , ,

  27. [35]

    Sheldon Cohen, Tom Kamarck, and Robin Mermelstein. 1983. A global measure of perceived stress.Journal of health and social behavior (1983), 385–396

  28. [36]

    Toshka Coleman, Sarina Till, Jaydon Farao, Londiwe Shandu, Nonkululeko Khuzwayo, Livhuwani Muthelo, Masenyani Mbombi, Mamare Bopape, Alastair van Heerden, Tebogo Mothiba, et al. 2023. Reconsidering priorities for digital maternal and child health: community-centered perspectiv...

  29. [37]

    Olivier Corneille and Bertram Gawronski. 2024. Self-reports are better measurement instruments than implicit measures. Nature Reviews Psychology (2024), 1–12

  30. [38]

    Evandro B Costa, Baldoino Fonseca, Marcelo Almeida Santana, Fabrísia Ferreira de Araújo, and Joilson Rego. 2017. Evaluating the effectiveness of educational data mining techniques for early prediction of students’ academic failure in introductory programming courses. Computers...

  31. [39]

    Reagan G Cox, Lei Zhang, William D Johnson, and Daniel R Bender. 2007. Academic performance and substance use: findings from a state survey of public high school students. Journal of school health 77, 3 (2007), 109–115

  32. [40]

    Marcus Credé, Sylvia G Roch, and Urszula M Kieszczynka. 2010. Class attendance in college: A meta-analytic review of the relationship of class attendance with grades and student characteristics. Review of Educational Research 80, 2 (2010), 272–295

  33. [41]

    Vedant Das Swain, Victor Chen, Shrija Mishra, Stephen M Mattingly, Gregory D Abowd, and Munmun De Choudhury. 2022. Semantic gap in predicting mental wellbeing through passive sensing. In Proceedings of the 2022 CHI conference on human factors in computing systems. 1–16

  34. [42]

    Vedant Das Swain, Lan Gao, Abhirup Mondal, Gregory D Abowd, and Munmun De Choudhury. 2024. Sensible and Sensitive AI for Worker Wellbeing: Factors that Inform Adoption and Resistance for Information Workers. InProceedings of the CHI Conference on Human Factors in Computing Sys...

  35. [43]

    Vedant Das Swain, Lan Gao, William A Wood, Srikruthi C Matli, Gregory D Abowd, and Munmun De Choudhury. 2023. Algorithmic power or punishment: Information worker perspectives on passive sensing enabled ai phenotyping of performance and wellbeing. In Proceedings of the 2023 CHI...

  36. [44]

    Ali Daud, Naif Radi Aljohani, Rabeeh Ayaz Abbasi, Miltiadis D Lytras, Farhat Abbas, and Jalal S Alowibdi. 2017. Predicting student performance using advanced learning analytics. In Proceedings of the 26th international conference on world wide web companion . 415–421

  37. [45]

    Susan M De Luca, Cynthia Franklin, Yan Yueqi, Shannon Johnson, and Chris Brownson. 2016. The relationship between suicide ideation, behavioral health, and college academic performance. Community mental health journal 52, 5 (2016), 534–540

  38. [46]

    What are you doing, TikTok?

    Daniel Delmonaco, Samuel Mayworm, Hibby Thach, Josh Guberman, Aurelia Augusta, and Oliver L Haimson. 2024. " What are you doing, TikTok?": How Marginalized Social Media Users Perceive, Theorize, and" Prove" Shadowbanning. Proceedings of the ACM on Human-Computer Interaction 8,...

  39. [47]

    Pieter Delobelle, Paul Temple, Gilles Perrouin, Benoît Frénay, Patrick Heymans, and Bettina Berendt. 2021. Ethical adversaries: Towards mitigating unfairness with adversarial machine learning. ACM SIGKDD Explorations Newsletter 23, 1 (2021), 32–41

  40. [48]

    Carsten F Dormann, Jane Elith, Sven Bacher, Carsten Buchmann, Gudrun Carl, Gabriel Carré, Jaime R García Marquéz, Bernd Gruber, Bruno Lafourcade, Pedro J Leitão, et al. 2013. Collinearity: a review of methods to deal with it and a simulation study evaluating their performance....

  41. [49]

    Afsaneh Doryab, Prerna Chikarsel, Xinwen Liu, and Anind K Dey. 2018. Extraction of behavioral features from smartphone and wearable data. arXiv preprint arXiv:1812.10394 (2018)

  42. [50]

    Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608 (2017)

  43. [51]

    Mengnan Du, Ninghao Liu, and Xia Hu. 2019. Techniques for interpretable machine learning. Commun. ACM 63, 1 (2019), 68–77

  44. [52]

    Upol Ehsan, Samir Passi, Q Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl. 2024. The Who in XAI: How AI Background Shapes Perceptions of AI Explanations. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–32

  45. [53]

    Upol Ehsan and Mark O Riedl. 2020. Human-centered explainable ai: Towards a reflective sociotechnical approach. In HCI International 2020-Late Breaking Papers: Multimodality and Intelligence: 22nd HCI International Conference, HCII 2020, Copenhagen, Denmark, July 19–24, 2020, ...

  46. [54]

    Upol Ehsan, Koustuv Saha, Munmun De Choudhury, and Mark O Riedl. 2023. Charting the sociotechnical gap in explainable ai: A framework to address the gap in xai. Proceedings of the ACM on human-computer interaction 7, CSCW1 (2023), 1–32

  47. [55]

    Martin Ester, Hans-Peter Kriegel, Jörg Sander, Xiaowei Xu, et al. 1996. A density-based algorithm for discovering clusters in large spatial databases with noise.. In kdd, Vol. 96. 226–231

  48. [56]

    FairLearn Contributors. 2022. Fairlearn Metrics Package. https://fairlearn.org/v0.7.0/api_reference/fairlearn.metrics.html

  49. [57]

    Michael Feldman, Sorelle A Friedler, John Moeller, Carlos Scheidegger, and Suresh Venkatasubramanian. 2015. Certifying and removing disparate impact. In proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining . 259–268

  50. [58]

    Mireia Felez-Nobrega, Charles H Hillman, Kieran P Dowd, Eva Cirera, and Anna Puig-Ribera. 2018. ActivPAL™ determined sedentary behaviour, physical activity and academic achievement in college students. Journal of sports sciences 36, 20 (2018), 2311–2316. 25 , , Zhang et al

  51. [59]

    Daniel Darghan Felisoni and Alexandra Strommer Godoi. 2018. Cell phone usage and academic performance: An experiment.Computers & Education 117 (2018), 175–187

  52. [60]

    Fitbit Team. 2023. Fitbit development: Sleep logs. https://dev.fitbit.com/build/reference/web-api/sleep/

  53. [61]

    Barbara L Fredrickson. 2000. Extracting meaning from past affective experiences: The importance of peaks, ends, and specific emotions. Cognition & Emotion 14, 4 (2000), 577–606

  54. [62]

    Yoav Freund and Robert E Schapire. 1997. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of computer and system sciences 55, 1 (1997), 119–139

  55. [63]

    Sorelle A Friedler, Carlos Scheidegger, and Suresh Venkatasubramanian. 2016. On the (im) possibility of fairness. arXiv preprint arXiv:1609.07236 (2016)

  56. [64]

    Jerome H Friedman. 2001. Greedy function approximation: a gradient boosting machine. Annals of statistics (2001), 1189–1232

  57. [65]

    Yarin Gal and Zoubin Ghahramani. 2016. A theoretically grounded application of dropout in recurrent neural networks. Advances in neural information processing systems 29 (2016)

  58. [66]

    Robert P Gallagher. 2006. National survey of counseling center directors 2005. (2006)

  59. [67]

    Saul Geiser and Maria Veronica Santelices. 2007. Validity of high-school grades in predicting student success beyond the freshman year: High-school record vs. standardized tests as indicators of four-year college outcomes. (2007)

  60. [68]

    O’Reilly Media, Inc

    Aurélien Géron. 2022. Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow . " O’Reilly Media, Inc. "

  61. [69]

    Fausto Giunchiglia, Mattia Zeni, Elisa Gobbi, Enrico Bignotti, and Ivano Bison. 2018. Mobile social media usage and academic performance. Computers in Human Behavior 82 (2018), 177–185

  62. [70]

    Ana Allen Gomes, José Tavares, and Maria Helena P de Azevedo. 2011. Sleep and academic performance in undergraduates: a multi-measure, multi-predictor approach. Chronobiology international 28, 9 (2011), 786–801

  63. [71]

    Farley Grubb. 2006. Does going Greek impair undergraduate academic performance? A case study. American Journal of Economics and Sociology 65, 5 (2006), 1085–1110

  64. [72]

    Mark Andrew Hall. 1999. Correlation-based feature selection for machine learning. (1999)

  65. [73]

    Shaher H Hamaideh. 2011. Stressors and reactions to stressors among university students. International journal of social psychiatry 57, 1 (2011), 69–80

  66. [74]

    Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems 29 (2016)

  67. [75]

    Martin Hlosta, Zdenek Zdrahal, and Jaroslav Zendulka. 2017. Ouroboros: early identification of at-risk students without models based on legacy data. In Proceedings of the seventh international learning analytics & knowledge conference . 6–15

  68. [76]

    Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–16

  69. [77]

    Melissa K Holt, David Finkelhor, and Glenda Kaufman Kantor. 2007. Multiple victimization experiences of urban elementary school students: Associations with psychosocial functioning and academic performance. Child abuse & neglect 31, 5 (2007), 503–515

  70. [78]

    Sungsoo Ray Hong, Jessica Hullman, and Enrico Bertini. 2020. Human factors in model interpretability: Industry practices, challenges, and needs. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–26

  71. [79]

    algorithmic bias is a data problem

    Sara Hooker. 2021. Moving beyond “algorithmic bias is a data problem”. Patterns 2, 4 (2021), 100241

  72. [80]

    Justin Hunt and Daniel Eisenberg. 2010. Mental health problems and help-seeking behavior among college students. Journal of adolescent health 46, 1 (2010), 3–10

  73. [81]

    Virginia W Huynh, Que-Lam Huynh, and Mary-Patricia Stein. 2017. Not just sticks and stones: Indirect ethnic discrimination leads to greater physiological reactivity. Cultural Diversity and Ethnic Minority Psychology 23, 3 (2017), 425

  74. [82]

    Apple Inc. 2024. If an app asks to track your activity. https://support.apple.com/en-us/102420

  75. [83]

    Apple Inc. 2025. User privacy and data use. https://developer.apple.com/app-store/user-privacy-and-data-use/

  76. [84]

    Maral Jamalova and M Constantinovits. 2019. The comparative study of the relationship between smartphone choice and socio-economic indicators. Int. J. Mark. Stud 11, 11 (2019), 10–5539

  77. [85]

    Sandeep M Jayaprakash, Erik W Moody, Eitel JM Lauría, James R Regan, and Joshua D Baron. 2014. Early alert of academically at-risk students: An open source analytics initiative. Journal of Learning Analytics 1, 1 (2014), 6–47

  78. [86]

    Robert I Kabacoff, Daniel L Segal, Michel Hersen, and Vincent B Van Hasselt. 1997. Psychometric properties and diagnostic utility of the Beck Anxiety Inventory and the State-Trait Anxiety Inventory with older adult psychiatric outpatients. Journal of anxiety disorders 11, 1 (1...

  79. [87]

    Manjula G Kadapatti and AHM Vijayalaxmi. 2012. Stressors of academic stress-a study on pre-university students. Indian Journal of Scientific Research 3, 1 (2012), 171–175

  80. [88]

    Faisal Kamiran and Toon Calders. 2012. Data preprocessing techniques for classification without discrimination. Knowledge and information systems 33, 1 (2012), 1–33. 26 Towards Human-Centered Early Academic Performance Prediction Models , ,

  81. [89]

    Faisal Kamiran, Asim Karim, and Xiangliang Zhang. 2012. Decision theory for discrimination-aware classification. In 2012 IEEE 12th international conference on data mining . IEEE, 924–929

  82. [90]

    Jacob Merew Katamei and Gedion A Omwono. 2015. Intervention strategies to improve students’ academic performance in public secondary schools in arid and semi-arid lands in Kenya. Int’l J. Soc. Sci. Stud. 3 (2015), 107

  83. [91]

    Anna Kawakami, Shreya Chowdhary, Shamsi T Iqbal, Q Vera Liao, Alexandra Olteanu, Jina Suh, and Koustuv Saha. 2023. Sensing wellbeing in the workplace, why and for whom? envisioning impacts with organizational stakeholders. Proceedings of the ACM on Human-Computer Interaction 7...

  84. [92]

    Anupam Khan and Soumya K Ghosh. 2021. Student performance analysis and prediction in classroom learning: A review of educational data mining studies. Education and information technologies 26, 1 (2021), 205–240

  85. [93]

    Mariam Khan, Misja Ilcisin, and Katherine Saxton. 2017. Multifactorial discrimination as a fundamental cause of mental health inequities. International Journal for Equity in Health 16 (2017), 1–12

  86. [94]

    Seunghyun Kim, Afsaneh Razi, Gianluca Stringhini, Pamela J Wisniewski, and Munmun De Choudhury. 2021. A human-centered systematic literature review of cyberbullying detection algorithms. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (2021), 1–34

  87. [95]

    Doreen H Kinkel and Scott E Henke. 2006. Impact of undergraduate research on academic performance, educational planning, and career development. Journal of Natural Resources and Life Sciences Education 35, 1 (2006), 194–201

  88. [96]

    Serkan Kiranyaz, Onur Avci, Osama Abdeljaber, Turker Ince, Moncef Gabbouj, and Daniel J Inman. 2021. 1D convolutional neural networks and applications: A survey. Mechanical systems and signal processing 151 (2021), 107398

  89. [97]

    Serkan Kiranyaz, Turker Ince, Osama Abdeljaber, Onur Avci, and Moncef Gabbouj. 2019. 1-D convolutional neural networks for signal processing applications. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8360–8364

  90. [98]

    Kenji Kobayashi and Yuri Nakao. 2021. One-vs.-One Mitigation of Intersectional Bias: A General Method for Extending Fairness-Aware Binary Classification. In International Conference on Disruptive Technologies, Tech Ethics and Artificial Intelligence . Springer, 43–54

  91. [99]

    Nicholas D Lane, Mashfiqui Mohammod, Mu Lin, Xiaochao Yang, Hong Lu, Shahid Ali, Afsaneh Doryab, Ethan Berke, Tanzeem Choudhury, and Andrew Campbell. 2011. Bewell: A smartphone application to monitor, model and promote wellbeing. In 5th international ICST conference on pervasi...

  92. [100]

    Juan A Lara, David Lizcano, María A Martínez, Juan Pazos, and Teresa Riera. 2014. A system for knowledge discovery in e-learning environments within the European Higher Education Area–Application to student data from Open University of Madrid, UDIMA. Computers & Education 72 (...

  93. [101]

    Anders Larrabee Sønderlund, Emily Hughes, and Joanne Smith. 2019. The efficacy of learning analytics interventions in higher education: A systematic review. British Journal of Educational Technology 50, 5 (2019), 2594–2618

  94. [102]

    Andrew Lepp, Jacob E Barkley, and Aryn C Karpinski. 2015. The relationship between cell phone use and academic performance in a sample of US college students. Sage Open 5, 1 (2015), 2158244015573169

  95. [103]

    Javier López Zambrano, Juan Alfonso Lara Torralbo, Cristóbal Romero Morales, et al . 2021. Early prediction of student learning performance through data mining: A systematic review. Psicothema (2021)

  96. [104]

    Hong Lu, Jun Yang, Zhigang Liu, Nicholas D Lane, Tanzeem Choudhury, and Andrew T Campbell. 2010. The jigsaw continuous sensing engine for mobile phone applications. In Proceedings of the 8th ACM conference on embedded networked sensor systems . 71–84

  97. [105]

    Owen HT Lu, Anna YQ Huang, Jeff CH Huang, Albert JQ Lin, Hiroaki Ogata, and Stephen JH Yang. 2018. Applying learning analytics for the early prediction of Students’ academic performance in blended learning. Journal of Educational Technology & Society 21, 2 (2018), 220–232

  98. [106]

    Scott Lundberg. 2017. A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874 (2017)

  99. [107]

    Adilson Marques, Diana A Santos, Charles H Hillman, and Luís B Sardinha. 2018. How does academic achievement relate to cardiorespiratory fitness, self-reported physical activity and objectively reported physical activity: a systematic review in children and adolescents aged 6–...

  100. [108]

    Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276–282

  101. [109]

    Lakmal Meegahapola, Dimitris Spathis, Marios Constantinides, Han Zhang, Sofia Yfantidou, Niels van Berkel, and Anind K Dey. 2024. FairComp: 2nd International Workshop on Fairness and Robustness in Machine Learning for Ubiquitous Computing. In Companion of the 2024 on ACM Inter...

  102. [110]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR) 54, 6 (2021), 1–35

  103. [111]

    Gonzalo Mendez, Luis Galárraga, and Katherine Chiluiza. 2021. Showing academic performance predictions during term planning: effects on students’ decisions, behaviors, and preferences. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–17

  104. [112]

    Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences. Artificial intelligence 267 (2019), 1–38

  105. [113]

    Christoph Molnar. 2020. Interpretable machine learning. Lulu. com. 27 , , Zhang et al

  106. [114]

    Mehrab Bin Morshed, Koustuv Saha, Richard Li, Sidney K D’Mello, Munmun De Choudhury, Gregory D Abowd, and Thomas Plötz

  107. [115]

    Imani Mwalumbwe and Joel S Mtebe. 2017. Using learning analytics to predict students’ performance in Moodle learning management system: A case of Mbeya University of Science and Technology. The Electronic Journal of Information Systems in Developing Countries 79, 1 (2017), 1–13

  108. [116]

    Kevin L Nadal, Katie E Griffin, Yinglee Wong, Kristin C Davidoff, and Lindsey S Davis. 2020. The injurious relationship between racial microaggressions and physical health: Implications for social work. In Microaggressions and Social Work Research, Practice and Education. Rout...

  109. [117]

    Abdallah Namoun and Abdullah Alshanqiti. 2020. Predicting student performance using data mining and learning analytics techniques: A systematic literature review. Applied Sciences 11, 1 (2020), 237

  110. [118]

    Subigya Nepal, Weichen Wang, Vlado Vojdanovski, Jeremy F Huckins, Alex Dasilva, Meghan Meyer, and Andrew Campbell. 2022. COVID student study: A year in the life of college students during the COVID-19 pandemic through the lens of mobile phone sensing. In Proceedings of the 202...

  111. [119]

    Nguyen Thai Nghe, Paul Janecek, and Peter Haddawy. 2007. A comparative analysis of techniques for predicting academic performance. In 2007 37th Annual Frontiers In Education Conference - Global Engineering: Knowledge Without Borders, Opportunities Without Passports . T2G–7–T2G...

  112. [120]

    Opeyemi Ojajuni, Foluso Ayeni, Olagunju Akodu, Femi Ekanoye, Samson Adewole, Timothy Ayo, Sanjay Misra, and Victor Mbarika

  113. [121]

    Kana Okano, Jakub R Kaczmarzyk, Neha Dave, John DE Gabrieli, and Jeffrey C Grossman. 2019. Sleep quality, duration, and consistency are associated with better academic performance in college students. NPJ science of learning 4, 1 (2019), 1–5

  114. [122]

    Alexandra Olteanu, Carlos Castillo, Fernando Diaz, and Emre Kıcıman. 2019. Social data: Biases, methodological pitfalls, and ethical boundaries. Frontiers in Big Data 2 (2019), 13

  115. [123]

    Mallie J Paschall and Bridget Freisthler. 2003. Does heavy drinking affect academic performance in college? Findings from a prospective study of high achievers. Journal of Studies on Alcohol 64, 4 (2003), 515–519

  116. [124]

    Juliana L Pereira, Gisela Maria Guedes-Carneiro, Liana R Netto, Patrícia Cavalcanti-Ribeiro, Sidnei Lira, José F Nogueira, Carlos A Teles, Karestan C Koenen, Aline S Sampaio, Lucas C Quarantini, et al. 2018. Types of trauma, posttraumatic stress disorder, and academic performa...

  117. [125]

    Dana Pessach and Erez Shmueli. 2022. A Review on Fairness in Machine Learning. ACM Computing Surveys (CSUR) 55, 3 (2022), 1–44

  118. [126]

    Helen Pluut, Petru Lucian Curşeu, and Remus Ilies. 2015. Social and study related stressors and resources among university entrants: Effects on well-being and academic performance. Learning and Individual Differences 37 (2015), 262–268

  119. [127]

    Cassidy Pyle, Nicole B Ellison, and Nazanin Andalibi. 2023. Social Media and College-Related Social Support Exchange for First- Generation, Low-Income Students: The Role of Identity Disclosures. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–36

  120. [128]

    Shaojie Qu, Kan Li, Bo Wu, Xuri Zhang, and Kaihao Zhu. 2019. Predicting student performance and deficiency in mastering knowledge points in MOOCs using multi-task learning. Entropy 21, 12 (2019), 1216

  121. [129]

    Lenore Sawyer Radloff. 1977. The CES-D scale: A self-report depression scale for research in the general population. Applied psychological measurement 1, 3 (1977), 385–401

  122. [130]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144

  123. [131]

    Avi Rosenfeld and Ariella Richardson. 2019. Explainability in human–agent systems. Autonomous agents and multi-agent systems 33 (2019), 673–705

  124. [132]

    Sebastian Ruder. 2017. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098 (2017)

  125. [133]

    Akane Sano, Andrew J Phillips, Z Yu Amy, Andrew W McHill, Sara Taylor, Natasha Jaques, Charles A Czeisler, Elizabeth B Klerman, and Rosalind W Picard. 2015. Recognizing academic performance, sleep quality, stress level, and mental health using personality traits, wearable sens...

  126. [134]

    Yasaman S Sefidgar, Woosuk Seo, Kevin S Kuehn, Tim Althoff, Anne Browning, Eve Riskin, Paula S Nurius, Anind K Dey, and Jennifer Mankoff. 2019. Passively-sensed Behavioral Correlates of Discrimination Events in College Students. Proceedings of the ACM on Human-Computer Interac...

  127. [135]

    Andrew D Selbst, Danah Boyd, Sorelle A Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and abstraction in sociotechnical systems. In Proceedings of the conference on fairness, accountability, and transparency . 59–68

  128. [136]

    Anni Silvola, Piia Näykki, Anceli Kaveri, and Hanni Muukkonen. 2021. Expectations for supporting student engagement with learning analytics: An academic path perspective. Computers & Education 168 (2021), 104192. 28 Towards Human-Centered Early Academic Performance Prediction ...

  129. [137]

    Michelle J Sternthal, Natalie Slopen, and David R Williams. 2011. Racial disparities in health: how much does stress really matter? 1. Du Bois review: social science research on race 8, 1 (2011), 95–113

  130. [138]

    Arthur A Stone, Stefan Schneider, and James K Harter. 2012. Day-of-week mood patterns in the United States: On the existence of ‘Blue Monday’, ‘Thank God it’s Friday’and weekend effects.The Journal of Positive Psychology 7, 4 (2012), 306–314

  131. [139]

    Esther Y Strahan. 2003. The effects of social anxiety and social skills on academic performance. Personality and individual differences 34, 2 (2003), 347–366

  132. [140]

    Otgontsetseg Sukhbaatar, Tsuyoshi Usagawa, and Lodoiravsal Choimaa. 2019. An artificial neural network based early prediction of failure-prone students in blended learning course. International Journal of Emerging Technologies in Learning (iJET) 14, 19 (2019), 77–92

  133. [141]

    Evren Sumuer. 2021. The effect of mobile phone usage policy on college students’ learning. Journal of Computing in Higher Education 33, 2 (2021), 281–295

  134. [142]

    Harini Suresh and John V Guttag. 2019. A framework for understanding unintended consequences of machine learning. arXiv preprint arXiv:1901.10002 2 (2019)

  135. [143]

    Andrew J Thayer, Clayton R Cook, Aria E Fiat, Meghanne N Bartlett-Chase, and Jessie M Kember. 2018. Wise feedback as a timely intervention for at-risk students transitioning into high school. School Psychology Review 47, 3 (2018), 275–290

  136. [144]

    Christopher A Thurber and Edward A Walton. 2012. Homesickness and adjustment in university students. Journal of American college health 60, 5 (2012), 415–419

  137. [145]

    Mickey T Trockel, Michael D Barnes, and Dennis L Egget. 2000. Health-related variables and academic performance among first-year college students: Implications for sleep and other behaviors. Journal of American college health 49, 3 (2000), 125–131

  138. [146]

    Catrine Tudor-Locke, Ho Han, Elroy J Aguiar, Tiago V Barreira, John M Schuna Jr, Minsoo Kang, and David A Rowe. 2018. How fast is fast enough? Walking cadence (steps/min) as a practical estimate of intensity in adults: a narrative review. British Journal of Sports Medicine 52,...

  139. [147]

    Civil Service Commission Department o f Labor US Equal Employment Opportunity Commission, Department o f Justice, et al. 1978. Uniform guidelines on employee selection procedures. Federal Register 43, 166 (1978), 38295–38309

  140. [148]

    Petrie JAC Van der Zanden, Eddie Denessen, Antonius HN Cillessen, and Paulien C Meijer. 2018. Domains and predictors of first-year student success: A systematic review. Educational Research Review 23 (2018), 57–77

  141. [149]

    Sahil Verma and Julia Rubin. 2018. Fairness definitions explained. In2018 ieee/acm international workshop on software fairness (fairware) . IEEE, 1–7

  142. [150]

    Aleksandar Višnjić, Vladica Veličković, Dušan Sokolović, Miodrag Stanković, Kristijan Mijatović, Miodrag Stojanović, Zoran Milošević, and Olivera Radulović. 2018. Relationship between the manner of mobile phone use and depression, anxiety, and stress in university students. In...

  143. [151]

    Hajra Waheed, Saeed-Ul Hassan, Raheel Nawaz, Naif R Aljohani, Guanliang Chen, and Dragan Gasevic. 2023. Early prediction of learners at risk in self-paced education: A neural network approach. Expert Systems with Applications 213 (2023), 118868

  144. [152]

    R. Wang, F. Chenand Z. Chen, T. Li, G. Harari, S. Tignor, X. Zhou, D. Ben-Zeev, and A. T. Campbell. 2014. Studentlife: Assessing mental health, academic performance and behavioral trends of college students using smartphones.. In Proceedings of the 2014 ACM International Joint...

  145. [153]

    Rui Wang, Gabriella Harari, Peilin Hao, Xia Zhou, and Andrew T Campbell. 2015. SmartGPA: how smartphones can assess and predict academic performance of college students. In Proceedings of the 2015 ACM international joint conference on pervasive and ubiquitous computing. 295–306

  146. [154]

    Zhiguang Wang, Weizhong Yan, and Tim Oates. 2017. Time series classification from scratch with deep neural networks: A strong baseline. In 2017 International joint conference on neural networks (IJCNN) . IEEE, 1578–1585

  147. [155]

    Tammy Wyatt and Sara B Oswalt. 2013. Comparing mental health issues among undergraduate and graduate students. American journal of health education 44, 2 (2013), 96–107

  148. [156]

    Jie Xu, Yunyu Xiao, Wendy Hui Wang, Yue Ning, Elizabeth A Shenkman, Jiang Bian, and Fei Wang. 2022. Algorithmic fairness in computational medicine. EBioMedicine 84 (2022)

  149. [157]

    Xuhai Xu, Prerna Chikersal, Afsaneh Doryab, Daniella K Villalba, Janine M Dutcher, Michael J Tumminia, Tim Althoff, Sheldon Cohen, Kasey G Creswell, J David Creswell, et al. 2019. Leveraging routine behavior and contextually-filtered features for depression detection among col...

  150. [158]

    Xuhai Xu, Xin Liu, Han Zhang, Weichen Wang, Subigya Nepal, Yasaman Sefidgar, Woosuk Seo, Kevin S Kuehn, Jeremy F Huckins, Margaret E Morris, et al. 2023. GLOBEM: cross-dataset generalization of longitudinal human behavior modeling. Proceedings of the ACM on Interactive, Mobile...

  151. [159]

    Xing Xu, Jianzhong Wang, Hao Peng, and Ruilin Wu. 2019. Prediction of academic performance associated with internet usage behaviors using machine learning algorithms. Computers in Human Behavior 98 (2019), 166–173

  152. [160]

    Xuhai Xu, Han Zhang, Yasaman Sefidgar, Yiyi Ren, Xin Liu, Woosuk Seo, Jennifer Brown, Kevin Kuehn, Mike Merrill, Paula Nurius, et al. 2022. GLOBEM dataset: multi-year datasets for longitudinal human behavior modeling generalization. Advances in Neural 29 , , Zhang et al. Infor...

  153. [161]

    Xuhai Xu, Han Zhang, Yasaman S Sefidgar, Yiyi Ren, Xin Liu, Woosuk Seo, Jennifer Brown, Kevin Scott Kuehn, Mike A Merrill, Paula S Nurius, et al. 2022. GLOBEM: Multi-Year Datasets for Longitudinal Human Behavior Modeling Generalization. Thirty-sixth Conference on Neural Inform...

  154. [162]

    Mustafa Yağcı. 2022. Educational data mining: prediction of students’ academic performance using machine learning algorithms. Smart Learning Environments 9, 1 (2022), 11

  155. [163]

    Huaxiu Yao, Defu Lian, Yi Cao, Yifan Wu, and Tao Zhou. 2019. Predicting academic performance for college students: a campus behavior perspective. ACM Transactions on Intelligent Systems and Technology (TIST) 10, 3 (2019), 1–21

  156. [164]

    Johnson Yeboah and George Dominic Ewur. 2014. The impact of WhatsApp messenger usage on students performance in Tertiary Institutions in Ghana. Journal of Education and practice 5, 6 (2014), 157–164

  157. [165]

    Dong Whi Yoo, Hayoung Woo, Sachin R Pendse, Nathaniel Young Lu, Michael L Birnbaum, Gregory D Abowd, and Munmun De Choudhury. 2024. Missed Opportunities for Human-Centered AI Research: Understanding Stakeholder Collaboration in Mental Health AI Research. Proceedings of the ACM...

  158. [166]

    Liang-Chih Yu, Cheng-Wei Lee, HI Pan, Chih-Yueh Chou, Po-Yao Chao, ZH Chen, SF Tseng, CL Chan, and K Robert Lai. 2018. Improving early prediction of academic failure using sentiment analysis on self-evaluated comments. Journal of Computer Assisted Learning 34, 4 (2018), 358–365

  159. [167]

    Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, and Krishna P Gummadi. 2017. Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment. In Proceedings of the 26th international conference on world wide web. 1171–1180

  160. [168]

    Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. 2018. Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society . 335–340

  161. [169]

    Han Zhang, Margaret E Morris, Paula S Nurius, Kelly Mack, Jennifer Brown, Kevin S Kuehn, Yasaman S Sefidgar, Xuhai Xu, Eve A Riskin, Anind K Dey, et al. 2022. Impact of Online Learning in the Context of COVID-19 on Undergraduates with Disabilities and Mental Health Concerns. A...

  162. [170]

    Han Zhang, Vedant Das Swain, Leijie Wang, Nan Gao, Yilun Sheng, Xuhai Xu, Flora D Salim, Koustuv Saha, Anind K Dey, and Jennifer Mankoff. 2024. Illuminating the Unseen: A Framework for Designing and Mitigating Context-induced Harms in Behavioral Sensing. arXiv preprint arXiv:2...

  163. [171]

    still” to “walking

    Han Zhang, Leijie Wang, Yilun Sheng, Xuhai Xu, Jennifer Mankoff, and Anind K Dey. 2023. A framework for designing fair ubiquitous computing systems. In Adjunct Proceedings of the 2023 ACM International Joint Conference on Pervasive and Ubiquitous Computing & the 2023 ACM Inter...

  164. [175]

    the direction of behavioral change ( i.e., increases or decreases in sleep duration) and 2) magnitude of the behavioral change (i.e., steep or gradual changes in sleep duration) within the first week, as well as the first half (Monday to Wednesday) and second half (Thursday to...

  165. [176]

    A year End-of-year GPA (2-class, 3-class & 4-class) Academic records; admissions information Accuracy≈ 72% (4-class); 80% (3-class); 93% (2-class) × × × [21] A term (14-17 wks) End-of-Year GPA (continuous) Learning Management System Log data; academic records; demographics 𝑅= ...

  166. [177]

    14 weeks Term GPA (2-class) Learning Management System Log Data Accuracy=93%, Recall =0.95 × × ×

  167. [178]

    12 wks Term GPA (2-class) Learning Management System Log data; academic records Avg accuracy≈92%, Avg sensitivity=65%, Avg precision≈75%, Avg F1=66% × × ×

  168. [179]

    10 wks Cumulative GPA (continuous) Behavioral data from sensor; self-reports MAE=0.18, 𝑟=0.81, 𝑅2=0.56 ✓ × ×

  169. [180]

    10 wks Term GPA (2-class) Online Learning Log data Accuracy=94%, Precision=0.82, Recall=0.90, Specificity=0.95 × × ×

  170. [181]

    6 wks Term GPA (continuous) Learning Management System Log data; academic records PMSE=159.71, 𝑅2=0.56 ✓ × ×

  171. [182]

    5 wks Single-class GPA (2-class) Online Learning Log data, demographics, assessment-relatd data Accuracy=69% Precision=0.70 Recall=0.70 AUC=0.71 ✓ × ×

  172. [183]

    5 wks Term GPA (2-class) Academic records; self-evaluation comments Accuracy=71%, F1=0.71 × × ×

  173. [184]

    4 wks Term GPA (2-class) Behavioral data from sensor; self-reports Accuracy=92% ✓ × ×

  174. [185]

    Data information

    4 wks Single-class GPA (2-class) Learning Management System Log data AUC=0.75 (original) AUC=0.63 (unseen data)× × ✓ 38 Towards Human-Centered Early Academic Performance Prediction Models , , Table 7. Data information. Statistics, dropout rate and data missingness for 2018 and...

  175. [2019]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1–21

    Prediction of mood instability with passive sensing. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1–21

  176. [2021]

    In Computational Science and Its Applications–ICCSA 2021: 21st International Conference, Cagliari, Italy, September 13–16, 2021, Proceedings, Part IX 21

    Predicting student academic performance using machine learning. In Computational Science and Its Applications–ICCSA 2021: 21st International Conference, Cagliari, Italy, September 13–16, 2021, Proceedings, Part IX 21 . Springer, 481–491

  177. [2022]

    Proceedings of the ACM on Human-computer Interaction 6, CSCW2 (2022), 1–48

    Burnout and the quantified workplace: Tensions around personal sensing interventions for stress in resident physicians. Proceedings of the ACM on Human-computer Interaction 6, CSCW2 (2022), 1–48

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.