Pith. sign in

REVIEW 4 major objections 8 minor 2 cited by

Do Tutors Learn from Equity Training and Can Generative AI Assess It?

T0 review · 4 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A short scenario-based lesson improves tutors' equity-responsive skills, and few-shot GPT-4o can grade their open responses at 88–89% accuracy.

desk verdict A transparent, useful dataset and a promising LLM-grading pipeline, but the learning-gain claim rests on a confounded pretest/posttest comparison that the paper's own adjustments do not rescue. read the letter →

arxiv 2412.11255 v1 pith:LLE6BBY6 submitted 2024-12-15 cs.HC cs.AI

classification cs.HCcs.AI
keywords tutortrainingequityscenario-basedlearninglargelanguagemodelsautomaticassessmentopen-endedresponsesfew-shotpromptinganalytics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a short online lesson can teach tutors to respond well to middle school students experiencing possible inequities, and whether a large language model can grade tutors' open-ended answers reliably. On 81 remote tutors, the authors find a marginally significant improvement from pretest to posttest, with significant gains only when the Jeremiah scenario came first; tutors also reported sharply higher confidence. For assessment, GPT-4o with few-shot prompting matched human coders on 89% of predict responses and 88% of explain responses, at a fraction of the cost and time of human grading. The authors read this as evidence that equity-focused tutor training can be delivered and automatically assessed at scale.

What carries the argument

The load-bearing mechanism is the modified predict-observe-explain (POE) cycle: tutors predict how to respond to a student in an inequitable situation, justify their choice, observe a research-based recommendation, and then transfer to a second scenario. Two scenario batteries (Jeremiah, lacking home internet; Alexis, seated where she cannot hear) serve as counterbalanced pretest and posttest, with binary human coding of open responses. On the AI side, the key machinery is few-shot prompting with chain-of-thought and contextual priming, asking GPT-4o to return a JSON score and rationale.

What would settle it

Give the same lesson to two new cohorts but reverse which scenario is pretest; if the Jeremiah-first advantage does not follow the scenario order, the learning-gain claim is an artifact of battery difficulty. Separately, run GPT-4o few-shot on a third, unseen inequity scenario and compare to fresh human labels; if agreement falls below the 0.88–0.89 range, the scalability claim fails.

Watch

Extended reading notes

Core claim

The central claim is that a scenario-based predict-observe-explain lesson increases tutors' skill at helping students recognize inequity and advocate for themselves, and that generative AI can be a practical substitute for human coders in scoring the lesson's open-ended responses. The learning evidence is a main effect of time at $F(1,79)=3.20$, $p=.078$, qualified by an interaction: only the Jeremiah-to-Alexis order showed a significant gain ($M=0.12$, $p=.001$). The assessment evidence is stronger: GPT-4o with few-shot prompting reaches 0.89 accuracy on predict and 0.88 on explain against binary human labels, and few-shot consistently beats zero-shot. The authors conclude that GPT-4o few-shot is the preferred model for large-scale grading, balancing accuracy, speed, and cost.

Load-bearing premise

The pretest and posttest scenario batteries (Jeremiah and Alexis) measure the same equity skill, so that a pretest-to-posttest difference counts as learning; if the batteries are not exchangeable in difficulty or content, the learning-gain conclusion collapses.

Editorial extensions

If this is right

  • If the learning gain is real, a one-session scenario lesson is enough to shift tutors toward recognizing inequity and encouraging student self-advocacy.
  • GPT-4o few-shot can grade this lesson's open responses at near-human agreement, making automated feedback and large-scale deployment feasible.
  • The cost comparison implies that grading 1,000 lesson completions with GPT-4o few-shot costs about $8.85 and 3.5 hours, versus roughly $500 and 16.7 hours for human graders.
  • Few-shot prompting consistently outperformed zero-shot, so future automated assessment should include worked examples in the prompt.
  • Released datasets, rubrics, and prompts allow other researchers to replicate and extend the equity-assessment pipeline.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial: The learning-gain conclusion rests on the exchangeability of the two scenario batteries; because gains appeared only in one order and the Alexis battery was easier at pretest, the true effect size may be smaller or order-dependent.
  • Editorial: The 0.89/0.88 agreement is with binary labels on a narrow rubric; on a new, harder scenario the same few-shot prompt may need retuning, so the practical claim should be tested on out-of-sample situations.
  • Editorial: The large confidence gain (3.44 to 4.51) with no correlation to measured learning suggests confidence may reflect perceived relevance rather than skill acquisition; future work could tie both to real tutoring transcripts.
  • Editorial: A direct test of the assessment claim would be to run GPT-4o few-shot on a third scenario battery and compare its scores to a fresh set of human labels; if agreement holds, the method generalizes beyond the two scenarios.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper reports an online, scenario-based equity training lesson for 81 undergraduate remote tutors and evaluates two outcomes: whether tutors learn equity-responsive skills (RQ1, RQ2) and whether GPT-4o and GPT-4-turbo can assess tutors' open-ended responses accurately enough for scalable grading (RQ3, RQ4). The lesson uses two counterbalanced scenarios (Jeremiah and Alexis) in a predict-observe-explain format. The authors find a marginally significant main effect of time (F(1,79)=3.20, p=.078), a significant time-by-scenario interaction driven by gains in only the Jeremiah-to-Alexis order (M=0.12, p=.001), and significantly increased self-reported confidence among the 35 tutors who completed the post-survey. For LLM grading, GPT-4o few-shot achieves 0.89 accuracy on predict responses and 0.88 on explain responses, with GPT-4-turbo similar; the authors recommend GPT-4o few-shot on cost and speed grounds. The paper releases the lesson log data, human annotation rubrics, and LLM prompts.

Significance. If the learning-gain and LLM-assessment results were solid, this would be a useful contribution to learning analytics: equity-focused tutor training is under-resourced, open-ended response assessment is costly, and the authors provide a rare public dataset, coding rubrics, and exact prompts. The human inter-rater reliability (Cohen's kappa 0.75 and 0.73) is a genuine strength, and the cost/throughput comparison for human versus LLM grading is practically informative. However, the significance is substantially weakened by the marginal and order-dependent learning-gain evidence and by the absence of inferential statistics in the LLM evaluation. The paper is best read as an exploratory demonstration plus a reproducibility-oriented dataset contribution rather than as a definitive demonstration that the lesson produces learning or that the LLM assessments are statistically equivalent to human grading.

major comments (4)
  1. [§4.1 and §5.1] The evidence for RQ1 does not establish a learning gain. The main effect of time is marginal (F(1,79)=3.20, p=.078) and the interaction is driven entirely by one scenario order: the Jeremiah-to-Alexis order shows a gain of M=0.12 (p=.001), while the reverse order shows M=-0.01 (p=.846). Because the Alexis battery was easier at pretest (79.9% vs. 72.7%), this is exactly the pattern a difficulty confound would produce. The z-score adjustment still leaves only a marginal time effect (F(1,79)=2.82, p=.097), and the Rasch-adjusted time effect is also marginal (β=0.42, p=.066). The non-significant t-test on pretest battery difficulty (t(12.87)=0.88, p=.393) does not rule out the confound given the small number of items and low reliability. The authors should either provide a stronger equivalence argument for the two scenario batteries or explicitly downgrade the RQ1 conclusion to exploratory and hedge the abstract accordingly.
  2. [§3.4 and §4.1] The eight-item test battery has a split-half reliability of only 0.489. Low reliability attenuates pretest-posttest difference scores and makes order-specific gains difficult to interpret; it also weakens the construct-validity link between the instrument and the equity skill being measured. The authors mention this limitation but still use the same scores for the central learning-gain claim and as the human labels for LLM evaluation. Please report reliability separately for each counterbalancing order, discuss the maximum detectable effect size at this reliability, and state how much the learning-gain conclusion could change under a correction for measurement error.
  3. [§4.3 and Table 5] The LLM evaluation is reported as point estimates without confidence intervals or any statistical comparison between models, prompting methods, or response types. For example, GPT-4o few-shot and GPT-4-turbo few-shot have identical predict accuracy (0.89) and explain accuracies of 0.88 and 0.89; without intervals or paired tests, the claim that few-shot outperforms zero-shot and the practical equivalence of the two models is not quantified. Add bootstrapped confidence intervals or McNemar-type tests for the accuracy/F1 differences, and show per-item agreement rather than only aggregate accuracy.
  4. [§3.5 and Future Work] The few-shot prompts were selected after iterative tuning, and only the best-performing prompt iteration is reported. Because the evaluation data are the same data used to select the prompt, the reported accuracy is likely optimistically biased. The Future Work section acknowledges this, but the current RQ3/RQ4 claims are nonetheless presented as the performance of 'GPT-4o few-shot' rather than of a prompt-selection procedure. Please evaluate on a held-out set, report results across multiple prompt variants, or clearly label the reported numbers as in-sample prompt-tuning results.
minor comments (8)
  1. [Abstract] The abstract contains a typo: 'abilities topredict' should be 'abilities to predict'.
  2. [§3.4] The description of the mixed-effects ANOVA says 'test time as a random effect'; time is a within-subjects factor, with subjects as the random effect. Please clarify the model specification.
  3. [§4.1] 'Shining light on the significant interaction' should be 'Shedding light on the significant interaction'.
  4. [§5.1] The text uses 'Jeremy' in one place ('Tutors who had the Jeremy scenario followed by the Alexis scenario') while the rest of the manuscript uses 'Jeremiah'. Please make the naming consistent.
  5. [§6] The final two paragraphs of the Limitations section are duplicated verbatim (from 'Only 35 out of 81 tutors completed the post-lesson survey' through 'capturing common misconceptions'). Remove the duplicate.
  6. [Table 4 and §3.5] The scoring prompt in Table 4 describes the scenario as 'a middle school student struggling to understand a math problem', but the actual assessment scenarios concern homework access and classroom seating. Align the prompt context with the real scenario content.
  7. [§5.4] 'What would it cost for humans to perform this same task?' is an incomplete sentence; please rephrase as part of a full sentence.
  8. [Figures 5 and 6] Figure 5 would be more informative with individual pretest/posttest data points or error bars, and Figure 6 should state how processing time estimates were derived (e.g., tokens/sec measurements) in the caption.

Circularity Check

1 steps flagged · score 2.0 of 10

LLM assessment accuracy is partly an in-sample fit; the learning-gain analysis is confounded but not circular.

  1. fitted input called prediction [Section 3.5, Section 4.3 (Table 5), Section 7]
    "The creation of these prompts followed an iterative process, with several rounds of adjustments informed by feedback from initial model outputs. ... Future work ... research could investigate the performance and variability of all combinations of few-shot prompts, rather than only reporting the best iteration, to better understand the impact of prompt design."

    The few-shot prompt is the fitted parameter: it was iteratively adjusted against the same human labels that define the reported accuracy, and only the best-performing iteration is reported. The paper then presents GPT-4o few-shot accuracy (0.89/0.88) as evidence of assessment proficiency, but this is an in-sample fit statistic rather than an out-of-sample prediction. The prompt embeds the same rubric and learner-sourced examples used to create the human labels, so high agreement partly reflects the model being tuned to reproduce those labels. This is a mild validation circularity, not a closed-form identity, because the model could still disagree; however, the 'prediction' that GPT-4o can assess tutor responses is not independently tested on held-out data.

full rationale

The learning-gain derivation (RQ1) is not circular: pretest and posttest scores are independently coded by human raters, and the paper's own adjusted analyses (z-score transformation and Rasch model) are additional robustness checks, not restatements of the input. The scenario-difficulty imbalance is a validity threat to the learning-gain claim, but it is not a definitional equivalence and therefore does not constitute circularity. The LLM evaluation (RQ3/RQ4) is also not circular by construction, since the model could disagree with the human labels; however, the few-shot prompts were iteratively tuned against the same human labels and the best iteration was reported, so the reported accuracy is partly a fit statistic rather than an independent prediction. Self-citations to prior same-author work [7,34] supply the lesson framework and prior learning-gain expectations, but they do not carry the central empirical claims of this paper. Overall, the central claims retain independent content, with one mild validation-circularity concern in the LLM assessment component.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper does not introduce new entities or fitted constants. The main burden falls on the validity of the human rubric and the exchangeability of the two scenario batteries, both of which the paper partly addresses with reliability checks and difficulty adjustments.

assumptions (3)
  • domain assumption Human binary coding of open-ended tutor responses is a valid ground truth for equity-focused responding.
    Sections 3.2 and 3.3 define correct/incorrect based on lesson objectives and report inter-rater reliability (kappa 0.75 and 0.73). Both the learning-gain analysis and the LLM evaluation rely on this coding as the standard.
  • domain assumption The pretest and posttest scenario batteries measure the same construct and are exchangeable for computing learning gains.
    Section 3.4 computes learning gains as posttest minus pretest; Section 5.1 acknowledges the Alexis battery was easier at pretest (79.9% vs 72.7%), so the exchangeability assumption is load-bearing and only partially addressed by z-score and Rasch adjustments.
  • standard math Standard mixed-effects ANOVA and Rasch model assumptions hold for the small sample of 81 tutors with binary item scores.
    Section 3.4 describes the mixed-effects ANOVA and Rasch-based split-half reliability (0.489). Small sample size, item imbalance, and possible ceiling effects are acknowledged in Section 6.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Tutors Learn from Equity Training and Can Generative AI Assess It?." pith.science (2026). https://pith.science/paper/LLE6BBY6

@misc{pith2026241211255,
  author       = {Pith},
  title        = {Pith review of: Do Tutors Learn from Equity Training and Can Generative AI Assess It?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LLE6BBY6}},
  note         = {Machine review of arXiv:2412.11255}
}
read the original abstract

Equity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment.

Figures

Figures reproduced from arXiv: 2412.11255 by the authors.

Figure 1
Figure 1. The modified predict-observe-explain cycle for the pretest and posttest scenarios. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The scenario involving student Jeremiah with the open-ended question prompting a tutor to predict the best approach. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The scenario involving student Alexis with the open-ended question prompting a tutor to predict the best approach. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Mean pretest and posttest scores between scenario order conditions and measurement points. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Average open response scores (2 pts total) at pretest and posttest by scenario for each LLM model and human graders. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Average performance and cost of GPT models for 1,000 lesson completions. Bubble size is proportional to processing time. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detecting LLM-Generated Short Answers and Effects on Learner Performance

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A fine-tuned GPT-4o detects human-annotated LLM-generated short answers at 80% accuracy, outperforming GPTZero, and flagged LLM use is associated with higher posttest MCQ scores.

  2. Comparing Few-Shot Prompting of GPT-4 LLMs with BERT Classifiers for Open-Response Assessment in Tutor Equity Training

    cs.HC 2025-01 conditional novelty 4.0 of 10

    Fine-tuned BERT outperformed few-shot GPT-4o and GPT-4 Turbo on all four open-response equity training assessment tasks in a small cross-validated study.

Reference graph

Works this paper leans on

40 extracted references · 34 canonical work pages · cited by 2 Pith papers

  1. [1]

    Reem H Alattar. 2019. The effectiveness of using scenario-based learning strategy in developing EFL eleventh graders’ speaking and prospective thinking skills. The Islamic University of Gaza, Palestine (2019)

  2. [2]

    Lisa Bardach, Robert M Klassen, Tracy L Durksen, Jade V Rushby, Keiko CP Bostwick, and Lynn Sheridan. 2021. The power of feedback and reflection: Testing an online scenario-based learning intervention for student teachers. Computers & Education 169 (2021), 104194

  3. [3]

    Paul Black and Dylan Wiliam. 2009. Developing the theory of formative assessment.Educational Assessment, Evaluation and Accountability (formerly: Journal of personnel evaluation in education) 21 (2009), 5–31

  4. [4]

    T Brown, B Mann, N Ryder, M Subbiah, JD Kaplan, P Dhariwal, A Neelakantan, P Shyam, G Sastry, A Askell, et al. 2020. Language models are few-shot learners Advances in neural information processing systems 33. (2020). Manuscript submitted to ACM 16 Thomas et al

  5. [5]

    Andrew C Butler. 2018. Multiple-choice testing in education: Are the best practices for assessment also good for learning? Journal of Applied Research in Memory and Cognition 7, 3 (2018), 323–331

  6. [6]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  7. [7]

    Danielle R Chine, Pallavi Chhabra, Adetunji Adeniran, Shivang Gupta, and Kenneth R Koedinger. 2022. Development of scenario-based mentor lessons: an iterative design process for training at scale. In Proceedings of the Ninth ACM Conference on Learning@ Scale . 469–471

  8. [8]

    Chine, Pallavi Chhabra, Adetunji Adeniran, Joseph Kopko, Cindy Tipper, Shivang Gupta, and Kenneth R

    Danielle R. Chine, Pallavi Chhabra, Adetunji Adeniran, Joseph Kopko, Cindy Tipper, Shivang Gupta, and Kenneth R. Koedinger. 2022. Scenario-based training and on-the-job support for equitable mentoring. InProceedings of The Learning Ideas Conference 2022 . Springer, Cham, Switzerland, 581–592

Show all 40 references
  1. [9]

    Aubrey Condor, Max Litster, and Zachary Pardos. 2021. Automatic Short Answer Grading with SBERT on Out-of-Sample Questions. International Educational Data Mining Society (2021)

  2. [10]

    Jeffrey Duncan-Andrade. 2009. Note to educators: Hope required when growing roses in concrete. Harvard educational review 79, 2 (2009), 181–194

  3. [11]

    Casey Foundation

    The Annie E. Casey Foundation. 2015. Race Equity and Inclusion Action Guide: Embracing Equity: 7 Steps to Advance and Embed Race Equity and Inclusion Within Your Organization. https://assets.aecf.org/m/resourcedoc/AECF_EmbracingEquity7Steps-2014.pdf Accessed: 2024-09-05

  4. [12]

    Graham Gibbs. 1988. Learning by doing: A guide to teaching and learning methods. Further Education Unit (1988)

  5. [13]

    Taucia González, Kate M McCabe, and Carolina Lobo De Castro. 2017. An Equity Toolkit for Inclusive Schools: Centering Youth Voice in School Change. Equity Assistance Center Region III, Midwest and Plains Equity Assistance Center (2017)

  6. [14]

    Jonathan Guryan, Jens Ludwig, Monica P Bhatt, Philip J Cook, Jonathan MV Davis, Kenneth Dodge, George Farkas, Roland G Fryer Jr, Susan Mayer, Harold Pollack, et al. 2023. Not too late: Improving academic outcomes among adolescents. American Economic Review 113, 3 (2023), 738–765

  7. [15]

    Zaretta Hammond. 2021. Liberatory education: Integrating the science of learning and culturally responsive practice. American Educator 45, 2 (2021), 4

  8. [16]

    Zifei FeiFei Han, Jionghao Lin, Ashish Gurung, Danielle Thomas, Eason Chen, Conrad Borchers, Shivang Gupta, and Ken Koedinger. 2024. Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation. InAI for Education: Bridging Innovation and Responsibility at t...

  9. [17]

    Owen Henkel, Libby Hills, Adam Boxer, Bill Roberts, and Zach Levonian. 2024. Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability To Mark Short Answer Questions in K-12 Education. In Proceedings of the Eleventh ACM Conference on Learning@ Scale ...

  10. [18]

    Wen-Juan Hou and Jia-Hao Tsao. 2011. AUTOMATIC ASSESSMENT OF STUDENTS’FREE-TEXT ANSWERS WITH DIFFERENT LEVELS. International Journal on Artificial Intelligence Tools 20, 02 (2011), 327–347

  11. [19]

    Cigdem Hursen and Funda Gezer Fasli. 2017. Investigating the efficiency of scenario based learning and reflective learning approaches in teacher education. European Journal of Contemporary Education 6, 2 (2017), 264–279

  12. [20]

    Sanjit Kakarla, Danielle Thomas, Jionghao Lin, Shivang Gupta, and Kenneth R Koedinger. 2024. Using large language models to assess tutors’ performance in reacting to students making math errors. arXiv preprint arXiv:2401.03238 (2024)

  13. [21]

    Mohammad Khalil, Paul Prinsloo, and Sharon Slade. 2023. Fairness, trust, transparency, equity, and responsibility in learning analytics. Journal of Learning Analytics 10, 1 (2023), 1–7

  14. [22]

    Seonghoon Kim and Leonard S Feldt. 2010. The estimation of the IRT reliability coefficient and its lower and upper bounds, with comparisons to CTT reliability statistics. Asia Pacific Education Review 11 (2010), 179–188

  15. [23]

    Kenneth R Koedinger, Jihee Kim, Julianna Zhuxin Jia, Elizabeth A McLaughlin, and Norman L Bier. 2015. Learning is not a spectator sport: Doing is better than watching for learning from a MOOC. In Proceedings of the second (2015) ACM conference on learning@ scale . 111–120

  16. [24]

    Matthew A Kraft and Grace T Falken. 2021. A blueprint for scaling tutoring and mentoring across public schools. Aera Open 7 (2021), 23328584211042858

  17. [25]

    Thomas, Wei Tan, Ngoc Dang Nguyen, and Kenneth R

    Jionghao Lin, Eason Chen, Zifei Han, Ashish Gurung, Danielle R. Thomas, Wei Tan, Ngoc Dang Nguyen, and Kenneth R. Koedinger. 2024. How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses. In Proceedings of the 17th International Confer...

  18. [26]

    Jionghao Lin, Zifei Han, Danielle R Thomas, Ashish Gurung, Shivang Gupta, Vincent Aleven, and Kenneth R Koedinger. 2024. How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses. International Journal of Artificial Intelligence in Education (2024), 1–27

  19. [27]

    Nitin Madnani, Jill Burstein, John Sabatini, and Tenaha O’Reilly. 2013. Automated Scoring of Summary-Writing Tasks Designed to Measure Reading Comprehension. Grantee Submission (2013)

  20. [28]

    Susan F McLean. 2016. Case-based learning and its application in medical and health-care fields: a review of worldwide literature. Journal of medical education and curricular development 3 (2016), JMECD–S20377

  21. [29]

    OpenAI. 2024. OpenAI API Pricing. https://openai.com/api/pricing/ Accessed: 2024-09-01

  22. [30]

    PLUS- Personalized Learning Squared. 2024. PLUS- Personalized Learning Squared. https://tutors.plus/ Accessed: 2024-09-05

  23. [31]

    Roger C Schank, Tamara R Berman, and Kimberli A Macpherson. 2013. Learning by doing. In Instructional-design theories and models . Routledge, 161–181

  24. [32]

    Shenbanjo, A

    T. Shenbanjo, A. Buonaspina, A. Bhagwat, and S. Baumgartner. 2024. Championing Change: A Practitioner Guide for Leading Inclusive and Equity-Infused Rapid-Cycle Learning. Mathematica. https://www.mathematica.org/ Accessed: 2024-09-05. Manuscript submitted to ACM Do Tutors Lear...

  25. [33]

    Hongye Tan, Chong Wang, Qinglong Duan, Yu Lu, Hu Zhang, and Ru Li. 2023. Automatic short answer grading by encoding student responses via a graph convolutional network. Interactive Learning Environments 31, 3 (2023), 1636–1650

  26. [34]

    Danielle Thomas, Xinyu Yang, Shivang Gupta, Adetunji Adeniran, Elizabeth Mclaughlin, and Kenneth Koedinger. 2023. When the tutor becomes the student: Design and evaluation of efficient scenario-based lessons for tutors. In LAK23: 13th International Learning Analytics and Knowl...

  27. [35]

    Danielle R Thomas, Jionghao Lin, Shambhavi Bhushan, Ralph Abboud, Erin Gatz, Shivang Gupta, and Kenneth R Koedinger. 2024. Learning and AI Evaluation of Tutors Responding to Students Engaging in Negative Self-Talk. InProceedings of the Eleventh ACM Conference on Learning@ Scal...

  28. [36]

    Meredith Thompson, Kesiena Owho-Ovuakporie, Kevin Robinson, Yoon Jeon Kim, Rachel Slama, and Justin Reich. 2019. Teacher Moments: A digital simulation for preservice teachers to approximate parent–teacher conversations. Journal of Digital Learning in Teacher Education 35, 3 (2...

  29. [37]

    Unicef et al. 2020. How many children and young people have internet access at home?: estimating digital connectivity during the COVID-19 pandemic . Technical Report. Unicef

  30. [38]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35 (2022), 24824–24837

  31. [39]

    Mengxue Zhang, Sami Baral, Neil Heffernan, and Andrew Lan. 2022. Automatic short math answer grading via in-context meta-learning. arXiv preprint arXiv:2205.15219 (2022)

  32. [40]

    Yuan Zhang, Rajat Shah, and Min Chi. 2016. Deep Learning+ Student Modeling+ Clustering: A Recipe for Effective Automatic Short Answer Grading. International Educational Data Mining Society (2016). A Digital Appendix All analysis code, study materials, and log data references c...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.