Pith. sign in

REVIEW 3 major objections 4 minor 39 references

LLM difficulty ratings systematically underestimate how hard misconception-driven fraction items are for real students.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 10:26 UTC pith:ZXFBMOAC

load-bearing objection A credible and well-written paper identifying a new bias—LLMs underestimating misconception-driven fraction items—but the missing scale alignment between 1–100 ratings and IRT β makes the quantitative claim non-reproducible; worth sending to review with a required calibration fix. the 3 major comments →

arxiv 2607.26067 v1 pith:ZXFBMOAC submitted 2026-06-22 cs.CY cs.AI

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

classification cs.CY cs.AI
keywords LLM difficulty estimationitem response theoryconceptual misconceptionfraction arithmeticcognitive difficultycurricular difficultyadaptive assessmenteducational data mining
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tests whether LLM-generated difficulty ratings match actual student performance on basic arithmetic items. Using 32 items, responses from 770 undergraduates, and repeated queries to four cutting-edge LLMs, it finds that models rank items in roughly the right order but consistently rate fraction problems as far easier than they truly are. The authors argue that LLMs approximate curricular difficulty—what instruction says should be easy—rather than cognitive difficulty caused by persistent misconceptions. They call this systematic bias the 'Easy Trap' and warn that adaptive assessment systems relying on LLM difficulty estimates without empirical grounding will mis-sequence content and miscalibrate student ability.

Core claim

The paper's central claim is that LLM difficulty estimates are systematically biased for items whose difficulty stems from conceptual misconceptions rather than procedural complexity. In the 32-item arithmetic assessment, fraction items such as '100 ÷ 1/2' were consistently rated near the easy end of the LLM 1–100 scale (mean 12.3) yet only 34.2% of students answered correctly. The five most underestimated items were all fraction items, and error analysis showed they involve well-known misconceptions like 'division makes smaller' and whole-number bias. The authors interpret this as evidence that LLMs encode curricular sequencing instead of learner cognition, producing a measurable and predic

What carries the argument

The central mechanism is the contrast between curricular difficulty and cognitive difficulty, operationalized by comparing LLM ratings on a 1–100 scale with empirical difficulty from Classical Test Theory (proportion correct) and a 2-parameter Item Response Theory model (difficulty parameter β). The paper uses Spearman rank correlation, RMSE/MAE, and item-level error analysis to show that fraction items deviate systematically from the expected trend while non-fraction items do not. A run-to-run flip-rate analysis further shows that fraction items produce less stable LLM predictions, supporting the claim of a domain-specific bias rather than random noise.

Load-bearing premise

The load-bearing premise is that LLM ratings on a 1–100 scale and empirical IRT difficulty β values (roughly −3 to +3) are placed on a common scale before computing discrepancies; if the normalization is chosen to align non-fraction items, the apparent fraction-specific underestimation could be partly a calibration artifact.

What would settle it

Re-analyze the 32-item data after calibrating LLM ratings to empirical β values using only non-fraction items, then check whether the eleven fraction items still show significant systematic underestimation; if the bias disappears, the 'Easy Trap' is an artifact of scale alignment rather than a genuine cognitive prediction failure. Alternatively, collect a larger bank of fraction items from multiple institutions and test whether the underestimation replicates beyond this single sample.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Adaptive testing systems that use ungrounded LLM difficulty estimates will select items that are too hard for learners in misconception-heavy domains, miscalibrating ability estimates.
  • Correlation-only validation of LLM difficulty estimates is misleading; two systems with similar rank correlations can differ substantially in absolute error on the very items that matter.
  • LLM-generated difficulty labels should be treated as provisional hypotheses requiring empirical calibration, not as ground truth for assessment design.
  • The 'Easy Trap' predicts that the most underestimated items will be those with strong conceptual-error patterns, which can be identified from distractor analysis rather than text alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the curricular-vs-cognitive explanation would be to prompt LLMs with the same items plus information about common misconceptions (e.g., typical wrong answers) and see whether underestimation shrinks; if it does, part of the bias is a prompt-formatting effect, not an inherent limitation.
  • The 'Easy Trap' likely generalizes to other 'bottleneck' topics with known misconceptions (e.g., algebra word problems, negative-number operations) — a prediction the paper does not test but that follows from its mechanism.
  • The direction of the bias should reverse for items that are procedurally complex but conceptually trivial (e.g., multi-digit arithmetic without carry errors): LLMs would overestimate difficulty. This inversion would provide a strong confirmation of the curricular-difficulty hypothesis.
  • If the bias is confirmed in larger item banks, a practical fix could be to calibrate LLM outputs against empirical p-values for a small anchor set of items before using the model at scale, making the LLM a useful prior rather than the final estimate.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper compares LLM-generated difficulty ratings (four models, five runs, 32 arithmetic items) with empirical difficulty estimated from 770 Indonesian undergraduates' responses using CTT p-values and a 2PL IRT model. It reports moderate Spearman rank correlations (0.52–0.77) between LLM ratings and empirical difficulty, but claims that LLMs systematically underestimate the difficulty of fraction items, terming this the 'Easy Trap.' The strongest example is P32 (100 ÷ 1/2), with an LLM mean rating of 12.30/100 but only 34.2% correct student responses. The paper interprets the pattern as evidence that LLMs encode curricular rather than cognitive difficulty, and cautions against using LLM difficulty estimates without empirical calibration.

Significance. If the claim holds, the result is practically important: LLM-based difficulty estimates are increasingly used in item generation and adaptive assessment, and a systematic bias on misconception-driven items would have direct consequences for test construction and learner modeling. The study has clear strengths: it uses real student response data rather than simulated learners, includes multiple frontier LLMs with repeated runs, reports rank-based agreement alongside absolute error, and provides item-level error and misconception analyses. The authors also state a commitment to an open repository. However, the central quantitative claim of 'underestimation' depends on an unstated alignment between the LLM 1–100 scale and the IRT/CTT difficulty scales. Until that alignment is specified and tested, the paper's headline magnitude is not reproducible, and the post hoc selection of five items from 11 fraction items is not sufficient to establish a systematic domain-specific bias. The limitations section is candid about sample size and single-institution data, but the main inferential gap remains.

major comments (3)
  1. [Section 3.3, Table 1, Table 2] The paper never specifies how LLM ratings on a 1–100 scale are put on the same scale as 2PL IRT β (roughly −3 to +3) or CTT (1−p) before computing RMSE, MAE, and residual patterns. Section 3.3 says β 'can then be compared directly,' but that is not a scaling argument. The reported RMSE values (e.g., Claude 0.809, ChatGPT 0.974) are numerically inconsistent with raw 1–100 ratings against β in [−3,3], which implies that some transformation or standardization was applied without being described. This is load-bearing: 'underestimation' has no coordinate-independent meaning on arbitrary scales. The problem is visible in Table 2 itself: P28 has IRT β = −0.319 (below-average difficulty) yet is listed as one of the five most underestimated items; on the raw β scale, an LLM rating of 13.25 is not obviously an underestimate. The authors must state the exact rescaling/equating procedure, justify it
  2. [Section 4, Table 2] The five 'most underestimated' fraction items are selected post hoc from the same data used to claim systematic underestimation, and no inferential test is reported. The qualitative pattern in Figure 2 is suggestive, but the claim that fraction items are systematically underestimated relative to non-fraction items requires a formal test of an item-type × empirical-difficulty interaction on residuals, or a permutation/block bootstrap analysis. With only 11 fraction items and 21 non-fraction items, a small number of extreme items could drive the apparent pattern. The authors acknowledge the small item sample in the limitations, but the central 'systematic bias' claim is precisely what needs an inferential test rather than a list of five selected cases.
  3. [Figure 3 and Section 4] The flip-rate analysis is used to support the claim that fraction items are not only misestimated but also less stable across runs. The manuscript reports high flip rates for fraction items, e.g., Gemini showing about 91% inconsistency across 5 runs for 10 of 11 fraction items, but no statistical comparison or confidence intervals are given. With n=11 fraction items and only 5 runs per model, sampling variability is large. The definition of flip rate for more than two runs also needs a precise formula. This is a secondary claim, but it should be quantified or softened.
minor comments (4)
  1. [Section 3.2.2] The text says 'four repetitions' but the formula says '5 repetitions' and N=640 = 4×5×32. Please correct the inconsistency and state the exact prompt template in an appendix or repository listing.
  2. [Table 1] The table is difficult to read. The row 'CTT 0.912* 0.366 0.257 Empirical baseline' is unexplained; if this is a baseline correlation or a comparison with CTT itself, it should be labeled clearly. Also, 'Significance in a = 0.05' is a typo for 'α = 0.05.'
  3. [Section 3.3] Typo: 'motives the use' should be 'motivates the use.' Also, the statement that β 'can then be compared directly' is not supported by citation [12], which is an IRT textbook and does not address LLM-rating alignment.
  4. [Various] The manuscript refers to Figure 1 and Figure 2 but does not include them in the provided text; if the figures are not available, the prompt template and residual plots need to be included in the final version. The fractional notation in Table 2/3 is also garbled (e.g., '100 ÷ !"', '2!" + 1#!'); use proper math typesetting.

Circularity Check

0 steps flagged

Empirical comparison against external student-response benchmarks; no fitted-parameter circularity; the 'Easy Trap' is a post-hoc interpretation rather than an input to the analysis.

full rationale

This paper is an empirical comparison, not a derivation. LLM difficulty ratings (1–100) are generated by four independent model systems from item text, while empirical difficulty is estimated from 770 student responses using CTT and 2PL IRT. No parameter is fitted to the target quantity, and no equation is constructed such that the predicted difficulty equals the empirical difficulty by definition. The central 'underestimation' claim rests on comparing LLM ratings with empirical measures; the absence of an explicit normalization between the 1–100 rating scale and IRT beta/logit scale is a measurement-validity concern, not a circularity, because the residuals are not forced by construction. The 'Easy Trap' is introduced after observing the residual pattern, and the curricular-vs-cognitive interpretation is an explanatory label applied to those residuals, not a mechanistic assumption used to produce them. There is a minor self-citation ([20] La Hadi & Dedyerianto, used for item validation), but it is not load-bearing for the central comparison and does not determine the results. The authors also acknowledge limitations—32 items, 11 fraction items, single institution—which weaken generalizability but do not indicate circular reasoning. Overall, the derivation chain is self-contained empirical benchmarking with no significant circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 1 invented entities

The central comparison rests on standard psychometric estimation (2PL IRT and CTT) and on an unstated scale alignment between LLM ratings and IRT β. No independent free parameters are introduced beyond the fitted IRT item parameters; the 'Easy Trap' is a post-hoc label for the observed residual pattern, not an independently evidenced mechanism.

free parameters (2)
  • 2PL item difficulty parameters β_i (i=1..32) = Estimated via MML; examples: P32 0.667, P3 0.843, P11 0.328, P20 0.185, P28 -0.319
    Fitted to 770 students' responses in mirt; these values define the empirical difficulty benchmark against which LLM ratings are compared.
  • 2PL item discrimination parameters a_i (i=1..32) = Range 0.28–1.83, M=1.20, SD=0.37
    Fitted in the same 2PL model; used to justify 2PL over 1PL, not directly in the LLM comparison.
axioms (5)
  • domain assumption The 2PL IRT model is appropriate: unidimensionality, local independence, and monotone item characteristic curves.
    Section 3.3 adopts the 2PL model without reporting unidimensionality or local-independence checks; discrimination heterogeneity is used to justify 2PL over 1PL.
  • ad hoc to paper LLM 1–100 ratings and IRT β can be compared directly after some normalization.
    Section 3.3 says the IRT difficulty parameter 'can then be compared directly' with LLM estimates, but the transformation is never stated, yet RMSE/MAE and residual plots depend on it.
  • domain assumption Fraction items are the relevant misconception-driven subset and non-fraction items serve as a control.
    Section 3.2.1 defines the item set; the classification of fraction items as misconception-prone is based on prior literature rather than measured item properties.
  • domain assumption Student response distributions reflect cognitive misconceptions rather than test artifacts.
    Section 4, Table 3 interprets common wrong answers as evidence of specific misconceptions; no distractor analysis or cognitive interview validates this link.
  • domain assumption The standardized prompt mirrors realistic educator use without misconception information.
    Section 5 acknowledges that richer prompts might improve alignment; the study intentionally withholds misconception context to reflect current practice.
invented entities (1)
  • Easy Trap (named bias construct) no independent evidence
    purpose: Label for the systematic underestimation of misconception-driven item difficulty by LLMs.
    The construct is defined from the same observed residual pattern that is used to evidence it; no independent prediction or measure of 'curricular difficulty' is provided.

pith-pipeline@v1.3.0-alltime-deepseek · 8019 in / 13392 out tokens · 127410 ms · 2026-08-02T10:26:22.291668+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.

Figures

Figures reproduced from arXiv: 2607.26067 by Amanda La Hadi, A. Taufiq Asyhari, Guanliang Chen, Muhammad Johan Alibasa.

Figure 1
Figure 1. Figure 1: Standardized Prompting Template That Used in All LLMs 3.2.2 LLM Difficulty Estimation Prompt Difficulty predictions were obtained from four frontier LLM based AI systems: 1) Claude Sonnet 4.5 (Anthropic, September 2025) 2) Gemini 3 (Google DeepMind, November 2025) 3) ChatGPT 5.2 (OpenAI, GPT-5 series, December 2025) 4) Microsoft Copilot (Powered by GPT-5, December 2025) [PITH_FULL_IMAGE:figures/full_fig_p… view at source ↗
Figure 2
Figure 2. Figure 2: compares LLM-predicted difficulty with empirical IRT difficulty. For CTT, we are using (1 – p) to align the magnitude with other measurements. Non-fraction items follow the expected positive trend, indicating general alignment between model predictions and empirical measures. However, fraction items systematically deviate from this relationship. As empirical difficulty increases (β > 0), fraction items bec… view at source ↗
Figure 3
Figure 3. Figure 3: Across most models, fraction items produced higher rates [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 20 canonical work pages

  1. [1]

    T., Carpuat, M., & Rudinger, R

    Acquaye, C., Huang, Y. T., Carpuat, M., & Rudinger, R. . (2026). Take Out Your Calculators- Estimating the Real Difficulty of Question Items with LLM Student Simulations. arXiv preprint arXiv:2601.09953

  2. [2]

    B., & Kim, S

    Baker, F. B., & Kim, S. H. (2004). Item response theory: Parameter estimation techniques. CRC press

  3. [3]

    All students had completed Indonesia-K13 mathematics curricula through Grade 9 – 12 (before the national shift to the Merdeka Curriculum in 2021)

    METHOD 3.1 Participant Participants were 770 second-year undergraduate students enrolled during 2020 to 2025 in mandatory mathematics courses at a large public university in Indonesia (79.74% female, 20.26% male; mean age = 19.2 years). All students had completed Indonesia-K13 mathematics curricula through Grade 9 – 12 (before the national shift to the Me...

  4. [4]

    S., & Hawn, A

    Baker, R. S., & Hawn, A. (2021). Algorithmic Bias in Education. International Journal of Artificial Intelligence in Education, 32(4), 1052–1092. https://doi.org/10.1007/s40593-021-00285-9

  5. [5]

    Biancini, G., Ferrato, A., & Limongelli, C. (2024). Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights Adjunct Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization,

  6. [6]

    J., & Bentley, B

    Bossé, M. J., & Bentley, B. (2018). College Students’ Understanding of Fraction Operations. International Electronic Journal of Mathematics Education, 13(3). https://doi.org/10.12973/iejme/3881

  7. [7]

    Castleman, J., Nadeem, N., Namjoshi, T., & Liu, L. T. . (2025). Rethinking Math Benchmarks for LLMs using IRT. Proceedings of Machine Learning Research, 66–82

  8. [8]

    Chalmers, R. P. (2012). Mirt_A Multidimensional Item Response Theory Package for the R Environment. ournal of statistical Software, 48, 1–29

  9. [9]

    Chu, Y., Peng He, Hang Li, Haoyu Han, Kaiqi Yang, Yu Xue, Tingting Li, Joseph Krajcik, and Jiliang Tang. . (2025). Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation. arXiv preprint arXiv:2504.05276. https://doi.org/10.5281/zenodo.15870304

  10. [10]

    Crocker, L., & Algina, J. (1986). Introduction to classical and modern test theory. Holt, Rinehart and Winston

  11. [11]

    Cui, J., Wang, L., Li, D., & Zhou, X. (2024). Verbalized arithmetic principles correlate with mathematics achievement. Br J Educ Psychol, 94(1), 41–57. https://doi.org/10.1111/bjep.12632

  12. [12]

    Deeken, C., Neumann, I., & Heinze, A. (2019). Mathematical Prerequisites for STEM Programs: What do University Instructors Expect from New STEM Undergraduates? International Journal of Research in Undergraduate Mathematics Education, 6(1), 23–41. https://doi.org/10.1007/s40753-019-00098-1

  13. [13]

    E., & Reise, S

    Embretson, S. E., & Reise, S. P. (2013). Item response theory for psychologists. Psychology Press

  14. [14]

    Gonzalez, M. A. A., Hernandez, M. B., Perez, M. A. P., Orozco, B. L., Soto, J. T. C., & Malagon, S. . (2025). Do Repetitions Matter? Strengthening Reliability in LLM Evaluations. arXiv preprint arXiv:2509.24086

  15. [15]

    S., Loud, B

    Gray, S. S., Loud, B. J., & Sokolowski, C. P. (2009). Calculus Students' Use and Interpretation of Variables: Algebraic vs. Arithmetic Thinking. Canadian Journal of Science, Mathematics and Technology Education, 9(2), 59–72. https://doi.org/10.1080/14926150902873434

  16. [16]

    K., & Swaminathan, H

    Hambleton, R. K., & Swaminathan, H. . (2013). Item response theory: Principles and applications. Springer Science & Business Media

  17. [17]

    Hijji, B. M. (2025). An Indispensable Requirement for Medical Dosage Calculation: Basic Mathematical Skills of Baccalaureate Nursing Students. Nurs Rep, 15(5). https://doi.org/10.3390/nursrep15050150

  18. [18]

    Hwang, J., & Riccomini, P. J. (2019). A Descriptive Analysis of the Error Patterns Observed in the Fraction-Computation Solution Pathways of Students With and Without Learning Disabilities. Assessment for Effective Intervention, 46(2), 132–142. https://doi.org/10.1177/1534508419872256

  19. [19]

    L., Sani, A., & Jafar

    Kadir, K., Kodirun, Cahyono, E., La Hadi, A. L., Sani, A., & Jafar. (2020). The ability of prospective teachers to pose contextual word problem about fractions addition. ournal of Physics: Conference Series, 1581(1)

  20. [20]

    P., Nayak, A., K, M

    Kumar, A. P., Nayak, A., K, M. S., Chaitanya, & Ghosh, K. (2023). A Novel Framework for the Generation of Multiple Choice Question Stems Using Semantic and Machine-Learning Techniques. International Journal of Artificial Intelligence in Education, 34(2), 332–375. https://doi.org/10.1007/s40593-023-00333-6

  21. [21]

    La Hadi, A., & Dedyerianto, D. . (2020). Analisis data miskonsepsi siswa sekolah menengah pertama dalam menyelesaikan operasi aritmatika dasar. l-Ta'dib: Jurnal Kajian Ilmu Kependidikan, 18–33

  22. [22]

    Lee, H.-J., & Boyadzhiev, I. (2020). Underprepared College Students’ Understanding of and Misconceptions with Fractions. International Electronic Journal of Mathematics Education, 15(3). https://doi.org/10.29333/iejme/7835

  23. [23]

    Lestari, M., Johar, R., Mailizar, M., & Ridho, A. (2023). Measuring Learning Loss Due to Disruptions from COVID-19: Perspectives from the Concept of Fractions. Jurnal Didaktik Matematika, 10(1), 131–151. https://doi.org/10.24815/jdm.v10i1.28580

  24. [24]

    Li, M., Jiao, H., Zhou, T., Zhang, N., Peters, S., & Lissitz, R. W. (2025). Item Difficulty Modeling Using Fine-tuned Small and Large Language Models. Educ Psychol Meas, 00131644251344973. https://doi.org/10.1177/00131644251344973

  25. [25]

    Li, Y., & Kulm, G. (2008). Knowledge and confidence of pre-service mathematics teachers: the case of fraction division. Zdm, 40(5), 833–843. https://doi.org/10.1007/s11858-008-0148-2

  26. [26]

    Nasution, N. E. A. (2023). Using artificial intelligence to create biology multiple choice questions for higher education. Agricultural and Environmental Education, 2(1). https://doi.org/10.29333/agrenvedu/13071

  27. [27]

    Nita, S., Sussolaikah, K., & Aldida, J. D. . (2023). The Role of Artificial Intelligence-Based Technology with ChatGPT as an Educational Learning Media Innovation in Indonesia. International Journal of Multidisciplinary Sciences and Arts, 2(4), 235–241

  28. [28]

    E., Patel, M., van Wamelen, P., Kodeswaran, B., Woolf, S., & Young, M

    Ormerod, C., Lottridge, S., Harris, A. E., Patel, M., van Wamelen, P., Kodeswaran, B., Woolf, S., & Young, M. (2022). Automated Short Answer Scoring Using an Ensemble of Neural Networks and Latent Semantic Analysis Classifiers. International Journal of Artificial Intelligence in Education, 33(3), 467–496. https://doi.org/10.1007/s40593-022-00294-2

  29. [29]

    Razavi, P., & Powers, S. J. (2025). Estimating Item Difficulty Using Large Language Models and Tree-Based Machine Learning Algorithms. arXiv preprint arXiv:2504.08804

  30. [30]

    Rudolph, E., Seer, H., Mothes, C., & Albrecht, J. (2024). Automated feedback generation in an intelligent tutoring system for counselor education Proceedings of the 19th Conference on Computer Science and Intelligence Systems (FedCSIS),

  31. [31]

    E., & Burks, L

    Ryals, M., Hill-Lindsay, S., Pilgrim, M. E., & Burks, L. C. (2025). 'Simple Mistakes' in College Algebra: an Analysis of Students' Perceptions of Their Errors Using Attribution Theory. International Journal of Research in Undergraduate Mathematics Education. https://doi.org/10.1007/s40753-025-00269-3

  32. [32]

    S., Duncan, G

    Siegler, R. S., Duncan, G. J., Davis-Kean, P. E., Duckworth, K., Claessens, A., Engel, M., Susperreguy, M. I., & Chen, M. (2012). Early predictors of high school mathematics achievement. Psychol Sci, 23(7), 691–697. https://doi.org/10.1177/0956797612440101

  33. [33]

    Spitzer, M. W. H., & Moeller, K. (2022). Predicting fraction and algebra achievements online: A large‐scale longitudinal study using data from an online learning environment. Journal of Computer Assisted Learning, 38(6), 1797–1806. https://doi.org/10.1111/jcal.12721

  34. [34]

    Stewart, S., & Reeder, S. (2017). Algebra Underperformances at College Level: What Are the Consequences? In And the Rest is Just Algebra (pp. 3–18). https://doi.org/10.1007/978-3-319-45053-7_1

  35. [35]

    A., Priyolistiyanto, A., Pinandhita, F., KA, A

    Susanto, D. A., Priyolistiyanto, A., Pinandhita, F., KA, A. P., & Bimo, D. S. (2024). Utilizing ChatGPT on designing English language teaching (ELT) materials in Indonesia: Opportunities and challenges. Celt: A Journal of Culture, English Language Teaching & Literature, 24(1), 157–171

  36. [36]

    Susnjak, T. (2023). Beyond Predictive Learning Analytics Modelling and onto Explainable Artificial Intelligence with Prescriptive Analytics and ChatGPT. International Journal of Artificial Intelligence in Education, 34(2), 452–482. https://doi.org/10.1007/s40593-023-00336-3

  37. [37]

    Tariq, V. (2003). Diagnosis of Mathematical Skills Among Bioscience Entrants

  38. [38]

    VanLehn, K., Milner, F., Banerjee, C., & Wetzel, J. (2023). A Step-Based Tutoring System to Teach Underachieving Students How to Construct Algebraic Models. International Journal of Artificial Intelligence in Education, 34(2), 224–246. https://doi.org/10.1007/s40593-023-00328-3

  39. [39]

    Weegar, R., & Idestam-Almquist, P. (2023). Reducing Workload in Short Answer Grading Using Machine Learning. International Journal of Artificial Intelligence in Education, 34(2), 247–273. https://doi.org/10.1007/s40593-022-00322-1