REVIEW 4 major objections 5 minor 1 cited by
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that on-demand LLM-generated explanatory feedback improves posttest performance in scenario-based tutor lessons, but only among learners who actually request and engage with it—offering the feedback alone produces no…
desk verdict Useful, honest field experiment on on-demand LLM feedback, but the claim that posttest gains reflect learning rather than answer reuse is not yet nailed down. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is an on-demand LLM feedback loop placed inside a predict-observe-explain lesson cycle: after a tutor submits an open response, the model classifies it against a predefined schema and, for incorrect responses, generates a minimally rephrased, research-aligned correction that the learner can request. The causal identification machinery is principal stratification, a method that estimates the treatment effect among learners who would actually use the feedback, combined with an ElasticNet propensity model (a regularized regression with L1 and L2 penalties) that predicts each learner's number of feedback requests from engagement and response features. That predicted propensity is applied to the control group to construct a fairer comparison and to test whether high-propensity learners benefit more.
What would settle it
Re-run the analysis with a logged measure of how long each learner actually spent viewing the LLM feedback: if learners who never open the feedback show the same posttest gain as those who read it, the effect is a selection artifact, not a learning effect. Alternatively, a randomized encouragement design that pushes low-propensity learners to request feedback would settle whether the benefit is caused by the feedback or by the traits that lead learners to seek it.
Extended reading notes
Core claim
Across 2,648 lesson completions by 885 tutor learners, learners who received LLM-generated explanatory feedback on their open responses scored significantly higher at posttest than those who did not receive or use it (0.10 SD, 95% CI [0.01, 0.19], p = .023). Simply being assigned the option to request feedback produced no significant overall gain, which the authors interpret as evidence that actual engagement drives the benefit. After using principal stratification with an ElasticNet propensity model trained on engagement, response, and session features to predict feedback requests, two lessons—Giving Effective Praise and Supporting a Growth Mindset—showed statistically significant propensity-adjusted effects of 0.33 and 0.28 SD, while other lessons showed smaller, non-significant effects. Receiving feedback did not increase lesson completion time (a 9-second average difference), and 94% of learners who rated the feedback called it helpful. The authors conclude that LLM feedback supports learning when learners choose to engage with it, and that its effectiveness is moderated by help-seeking propensity rather than guaranteed by availability.
Load-bearing premise
The load-bearing premise is that the propensity model, trained on engagement and response behaviors, captures every trait that makes learners both more likely to request LLM feedback and more likely to score higher at posttest; if unmeasured traits such as help-seeking skill or motivation remain, the adjusted comparison does not identify a causal effect.
Editorial extensions
If this is right
- Existing online learning systems that already give corrective feedback could add on-demand LLM feedback at near-zero time cost and expect modest posttest gains among learners who request it.
- Offering feedback is not enough; to realize the benefit, systems should encourage help-seeking or otherwise increase the likelihood that learners engage with available support.
- Because lesson-level effects varied from 0.33 SD in Giving Effective Praise to a negative trend in Helping Students Manage Inequity, the benefit of LLM feedback depends on content and task type, not just feedback delivery.
- The lack of completion-time differences implies LLM feedback can be integrated into short scenario-based training without sacrificing efficiency.
Reading between the lines
- If help-seeking propensity is itself a trainable skill, then teaching learners when and how to request feedback could amplify the effect beyond the modest gains reported here.
- A direct robustness check of the copying concern would compare posttest open responses to the LLM feedback text; high similarity would suggest part of the gain is response imitation rather than durable learning.
- The propensity adjustment may under- or over-correct because it uses engagement proxies; adding measured motivational or metacognitive covariates could sharpen the causal estimate and reveal whether high-propensity learners benefit from deeper processing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a randomized-within-subject study of on-demand LLM-generated explanatory feedback in seven scenario-based tutor-training lessons, with 885 learners and 2,648 lesson completions. Learners assigned to the intent-to-treat condition could request GPT-3.5-turbo feedback on an open-response question; all learners received the existing non-LLM corrective feedback. The authors compare posttest performance across learners who received feedback (treatment-on-the-treated), those who declined it, and those without access. They report a TOT association of 0.10 SD (95% CI [0.01, 0.19], p = .023), no significant ITT main effect, a propensity-adjusted analysis showing significant lesson-level benefits for two of seven lessons (Giving Effective Praise, beta = 0.33; Supporting a Growth Mindset, beta = 0.28), no significant completion-time cost, and predominantly positive learner ratings. The paper contributes open datasets, prompts, and rubrics and frames the central conclusion as: LLM feedback helps learning only when learners choose to engage with it.
Significance. If the reported effects are genuine, the paper provides useful evidence about a realistic, low-cost feedback augmentation in existing online tutor training, and its open resources support replication. The within-subject lesson-level assignment, the explicit handling of self-selection via propensity scoring, and the effort to evaluate open responses with rubric-based LLM scoring are notable strengths. However, the central learning claim is only as strong as its weakest load-bearing assumptions: the TOT estimate is observational, the propensity adjustment cannot remove unmeasured confounders, the lesson-level effects are not corrected for multiple comparisons, and the authors concede they cannot rule out posttest answer reuse. The paper therefore advances a plausible but not yet fully established claim, and the currently available evidence does not justify the strength of the concluding statements.
major comments (4)
- [§5 Discussion, posttest contamination] The manuscript's central claim that LLM feedback improves learning rests on posttest open-response scores (Q7/Q9), which are scored by GPT-4o under rubrics rewarding the same key elements that the Q1 LLM feedback reinforces. The paper concedes that learners could copy or adapt the feedback text at posttest and that the only follow-up analysis measured response length, not textual overlap. The statement that such cases were 'rare' is not supported by any reported evidence; low-stakes assessments do not by themselves prevent incidental reuse. Because if reuse inflates the 0.10 SD TOT estimate or the lesson-level effects (0.33 and 0.28), the observed gains would not demonstrate learning, this issue is load-bearing. The authors should report a quantitative textual-overlap analysis between Q1 feedback and Q7/Q9 responses, re-estimate the effects after excluding suspected reuse cases, or otherwise provide direct evidence that the results are not driven by answer reuse.
- [§4.1, propensity adjustment and causal interpretation] The propensity model in §4.1 is trained on engagement, response, and session features from the ITT condition and applied to the control condition to predict the number of LLM feedback requests. The authors acknowledge that unmeasured confounders, such as help-seeking skill or motivation, may remain, and the Discussion correctly hedges that engagement 'may also reflect preexisting learning differences.' However, the abstract and concluding paragraph state as the key empirical finding that 'the effectiveness of such feedback depends on learners' willingness to seek and engage with it.' This causal claim is stronger than the propensity-adjusted analysis supports, since the adjustment cannot fully remove self-selection. The paper should either temper the conclusion to an associational claim or add a sensitivity analysis (for example, reporting how large an unmeasured confounder would need to be to explain the TOT effect, or exploiting the 15.7% API-failure rate as a possible instrument for actual receipt).
- [§4.1, lesson-level comparisons] Seven lesson-level propensity-adjusted treatment effects are reported, with two reaching p < .05. No correction for multiple comparisons is mentioned, and the overall ITT-by-propensity interaction in Table 2 is not significant (beta = 0.04, 95% CI [-0.04, 0.12], p = .307). Given the nonsignificant interaction, the two significant lesson-level effects should be presented as exploratory, and the authors should report adjusted p-values, confidence intervals, or a false-discovery-rate control so readers can assess the strength of the cross-lesson support for the central claim.
- [Abstract and §4.1, interpretation of propensity results] The abstract states that 'Learners with a higher predicted likelihood of engaging with LLM feedback scored significantly higher at posttest than those with lower propensity.' As reported in Table 2, the significant propensity coefficient (beta = 0.06, p = .040) is for the control condition's predicted propensity, not for treated learners. The subsequent sentence about two significant lessons does not mention that the overall ITT-by-propensity interaction was not significant. The abstract should be reworded to match the actual model results, clarifying that the selection effect is observed in the control group and that the treatment-contingent benefit is supported only by the two exploratory lesson-level analyses.
minor comments (5)
- [§3.2, assignment mechanism] The text first says learners were randomly assigned to one of two conditions for each lesson, but later says that for some lessons all learners were assigned to the ITT condition for a period. These statements are in tension; please clarify the actual assignment procedure and whether the analysis accounts for partially non-random assignment periods.
- [§5, response-length finding] The follow-up analysis reporting that feedback recipients produced responses six words longer at Q9 should include an effect size, confidence interval, or p-value, so the reader can judge the strength of the evidence.
- [§3.3, inter-rater reliability citations] The text says IRR was established in prior open-source work for most lessons, but no citation is provided for that prior work; please add the reference or state which lessons came from which source.
- [Global, formatting] Several typographical issues occur, including missing spaces ('tutorsas learners', 'bygpt-3.5-turbo') and a table header that reads 'Feedback Offered: Intent-to-Treat (ITT)'; these should be corrected in the final version.
- [§4.1, model reporting] The lesson-level models are described as 'separate regressions,' but it would be helpful to state whether they include the same random effects as the main model and whether the propensity covariate enters as a continuous variable or a stratifier.
Circularity Check
No circularity: the treatment effects are estimated from independently measured posttest scores, and the propensity model is a nuisance adjustment, not a constructed outcome.
full rationale
This paper is an empirical estimation study rather than a derivation. The posttest outcome is measured independently of any fitted parameter: open responses are scored by GPT-4o under rubrics and multiple-choice items are scored objectively, after the instructional intervention. The propensity model is fit on engagement, timing, and response-length features in the treatment condition and then applied to the control condition to create a covariate; it is not used to define the posttest score or to construct the treatment coefficient. The reported TOT effect (0.10 SD) and the lesson-level propensity-adjusted effects (0.33 and 0.28) come from regressing posttest on condition and propensity, so they are not by construction equal to the fitted propensity values or to any pre-selected parameter. The self-citations in the paper (e.g., refs. 17, 18, 27–29) concern prior lesson design, GPT-based feedback rephrasing, and previously established inter-rater reliability. These are supports for the intervention design and measurement quality, not load-bearing logical premises that force the conclusion. The paper explicitly acknowledges a possible alternative explanation—that learners may have copied or adapted LLM feedback at posttest—and treats it as a limitation rather than dismissing it. That is a validity threat, not a circularity of the kind where Eq. X reduces to Eq. Y or a fitted parameter is renamed as a prediction. No self-definitional, fitted-input-as-prediction, self-citation-load-bearing, uniqueness-imported, ansatz-smuggled, or renamed-known-result pattern is exhibited.
Assumptions & free parameters
free parameters (1)
- ElasticNet propensity model coefficients and hyperparameters (alpha, lambda) =
not reported in paper
assumptions (4)
- domain assumption GPT-4o scoring of posttest open responses is a valid measure of tutor learning
- domain assumption Random assignment balances prior knowledge, so posttest differences can be attributed to learning during the lesson
- domain assumption The propensity model is correctly specified and no unmeasured confounders remain
- domain assumption GPT-3.5-turbo generated feedback is accurate and aligned with the research-based tutoring schema
Cite this review
Pith. "Pith review of LLM-Generated Feedback Supports Learning If Learners Choose to Use It." pith.science (2026). https://pith.science/paper/K4C5HXYF
@misc{pith2026250617006,
author = {Pith},
title = {Pith review of: LLM-Generated Feedback Supports Learning If Learners Choose to Use It},
year = {2026},
howpublished = {\url{https://pith.science/paper/K4C5HXYF}},
note = {Machine review of arXiv:2506.17006}
}
read the original abstract
Large language models (LLMs) are increasingly used to generate feedback, yet their impact on learning remains underexplored, especially compared to existing feedback methods. This study investigates how on-demand LLM-generated explanatory feedback influences learning in seven scenario-based tutor training lessons. Analyzing over 2,600 lesson completions from 885 tutor learners, we compare posttest performance among learners across three groups: learners who received feedback generated by gpt-3.5-turbo, those who declined it, and those without access. All groups received non-LLM corrective feedback. To address potential selection bias-where higher-performing learners may be more inclined to use LLM feedback-we applied propensity scoring. Learners with a higher predicted likelihood of engaging with LLM feedback scored significantly higher at posttest than those with lower propensity. After adjusting for this effect, two out of seven lessons showed statistically significant learning benefits from LLM feedback with standardized effect sizes of 0.28 and 0.33. These moderate effects suggest that the effectiveness of LLM feedback depends on the learners' tendency to seek support. Importantly, LLM feedback did not significantly increase completion time, and learners overwhelmingly rated it as helpful. These findings highlight LLM feedback's potential as a low-cost and scalable way to improve learning on open-ended tasks, particularly in existing systems already providing feedback without LLMs. This work contributes open datasets, LLM prompts, and rubrics to support reproducibility.
Figures
Forward citations
Cited by 1 Pith paper
-
Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
EduAlign trains a three-dimensional reward model (HPC-RM) and uses GRPO to fine-tune Qwen2.5-72B, reporting improved helpfulness, personalization, and creativity on its own and public benchmarks.
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2303.08774 (2023)
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
International Journal of Artificial Intelligence in Education16(2), 101–128 (2006) 14 D
Aleven, V., Mclaren, B., Roll, I., Koedinger, K.: Toward meta-cognitive tutoring: A model of help seeking with a cognitive tutor. International Journal of Artificial Intelligence in Education16(2), 101–128 (2006) 14 D. R. Thomas et al
work page 2006
-
[3]
Computers & Education169, 104194 (2021)
Bardach,L.,Klassen,R.M.,Durksen,T.L.,Rushby,J.V.,Bostwick,K.C.,Sheridan, L.: The power of feedback and reflection: Testing an online scenario-based learning intervention for student teachers. Computers & Education169, 104194 (2021)
work page 2021
-
[4]
Caruccio, L., Cirillo, S., Polese, G., Solimando, G., Sundaramurthy, S., Tortora, G.: Claude 2.0 large language model: Tackling a real-world classification problem with anewiterativepromptengineeringapproach.IntelligentSystemswithApplications 21, 200336 (2024)
work page 2024
-
[5]
Computers and Education: Artificial Intelligence2, 100027 (2021)
Cavalcanti, A.P., Barbosa, A., Carvalho, R., Freitas, F., Tsai, Y.S., Gašević, D., Mello, R.F.: Automatic feedback in online learning environments: A systematic lit- erature review. Computers and Education: Artificial Intelligence2, 100027 (2021)
work page 2021
-
[6]
Chhabra, P., Chine, D., Adeniran, A., Gupta, S., Koedinger, K.: An evaluation of perceptions regarding mentor competencies for technology-based personalized learning.In:SocietyforInformationTechnology&TeacherEducationInternational Conference. pp. 1812–1817. Association for the Advancement of Computing in Education (AACE) (2022)
work page 2022
-
[7]
In: 2023 IEEE International Conference on Advanced Learning Technologies (ICALT)
Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y.S., Gašević, D., Chen, G.: Can large language models provide feedback to students? a case study on chatgpt. In: 2023 IEEE International Conference on Advanced Learning Technologies (ICALT). pp. 323–325. IEEE (2023)
2023
-
[8]
Demszky, D., Liu, J., Hill, H.C., Jurafsky, D., Piech, C.: Can automated feedback improve teachers’ uptake of student ideas? evidence from a randomized controlled trial in a large-scale online course. edworkingpaper no. 21-483. Annenberg Institute for School Reform at Brown University (2021)
work page 2021
Show all 33 references
-
[9]
arXiv preprint arXiv:2306.10509 (2023)
Denny, P., Khosravi, H., Hellas, A., Leinonen, J., Sarsa, S.: Can we trust llm- generated educational content? comparative analysis of human and llm-generated learning resources. arXiv preprint arXiv:2306.10509 (2023)
2023 arXiv
-
[10]
International Journal of Educational Tech- nology in Higher Education20(1), 57 (2023)
Escalante, J., Pack, A., Barrett, A.: Llm-generated feedback on writing: insights into efficacy and enl student preference. International Journal of Educational Tech- nology in Higher Education20(1), 57 (2023)
2023
-
[11]
British Journal of Educational Technology (2024)
Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., Gašević, D.: Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology (2024)
2024
-
[12]
Further Education Unit (1988)
Gibbs, G.: Learning by doing: A guide to teaching and learning methods. Further Education Unit (1988)
1988
-
[13]
Gurung, A., Baral, S., Vanacore, K.P., Mcreynolds, A.A., Kreisberg, H., Botelho, A.F., Shaw, S.T., Heffernan, N.T.: Identification, exploration, and remediation: Can teachers predict common wrong answers? In: LAK23: 13th International Learning Analytics and Knowledge Conferenc...
2023
-
[14]
Review of educational research 77(1), 81–112 (2007)
Hattie, J., Timperley, H.: The power of feedback. Review of educational research 77(1), 81–112 (2007)
2007
-
[15]
Computers and Education: Artificial Intelligence8, 100349 (2025)
Kinder, A., Briese, F.J., Jacobs, M., Dern, N., Glodny, N., Jacobs, S., Leßmann, S.: Effects of adaptive feedback generated by a large language model: A case study in teacher education. Computers and Education: Artificial Intelligence8, 100349 (2025)
2025
-
[16]
Handbook of educational data mining43, 43–56 (2010)
Koedinger,K.R.,Baker,R.S.,Cunningham,K.,Skogsholm,A.,Leber,B.,Stamper, J.: A data repository for the edm community: The pslc datashop. Handbook of educational data mining43, 43–56 (2010)
2010
-
[17]
arXiv preprint arXiv:2405.00291 (2024) LLM-Generated Feedback Supports Learning If Learners Choose to Use It 15
Lin, J., Chen, E., Han, Z., Gurung, A., Thomas, D.R., Tan, W., Nguyen, N.D., Koedinger, K.R.: How can i improve? using gpt to highlight the desired and unde- sired parts of open-ended responses. arXiv preprint arXiv:2405.00291 (2024) LLM-Generated Feedback Supports Learning If...
2024 arXiv
-
[18]
International Journal of Artificial Intelligence in Education pp
Lin, J., Han, Z., Thomas, D.R., Gurung, A., Gupta, S., Aleven, V., Koedinger, K.R.: How can i get it right? using gpt to rephrase incorrect trainee responses. International Journal of Artificial Intelligence in Education pp. 1–27 (2024)
2024
-
[19]
Journal of Educa- tional Psychology114(8), 1743 (2022)
Mertens, U., Finn, B., Lindner, M.A.: Effects of computer-based feedback on lower- and higher-order learning outcomes: A network meta-analysis. Journal of Educa- tional Psychology114(8), 1743 (2022)
2022
-
[20]
Computers and Education: Artificial Intelligence6, 100199 (2024)
Meyer, J., Jansen, T., Schiller, R., Liebenow, L.W., Steinbach, M., Horbach, A., Fleckenstein, J.: Using llms to bring evidence-based feedback into the classroom: Llm-generated feedback increases secondary students’ text revision, motivation, and positive emotions. Computers a...
2024
-
[21]
Plos one19(5), e0304013 (2024)
Pardos,Z.A.,Bhandari,S.:Chatgpt-generatedhelpproduceslearninggainsequiva- lent to human tutor-authored help on mathematics skills. Plos one19(5), e0304013 (2024)
2024
-
[22]
Learn- ing and instruction21(2), 267–280 (2011)
Roll, I., Aleven, V., McLaren, B.M., Koedinger, K.R.: Improving students’ help- seeking skills using metacognitive feedback in an intelligent tutoring system. Learn- ing and instruction21(2), 267–280 (2011)
2011
-
[23]
Sales, A.C., Pane, J.F.: The role of mastery learning in an intelligent tutoring system: Principal stratification on a latent variable (2019)
2019
-
[24]
arXiv preprint arXiv:2212.10406 (2022)
Sales, A.C., Vanacore, K.P., Ottmar, E.R.: Geepers: Principal stratification using principal scores and stacked estimating equations. arXiv preprint arXiv:2212.10406 (2022)
2022 arXiv
-
[25]
In: International Conference on Artificial Intelligence in Education
Stamper, J., Xiao, R., Hou, X.: Enhancing llm-based feedback: Insights from in- telligent tutoring systems and the learning sciences. In: International Conference on Artificial Intelligence in Education. pp. 32–43. Springer (2024)
2024
-
[26]
Computers in human behavior28(5), 1618–1625 (2012)
Sung, E., Mayer, R.E.: When graphics improve liking but not learning from online lessons. Computers in human behavior28(5), 1618–1625 (2012)
2012
-
[27]
In: LAK23: 13th International Learning Analytics and Knowledge Conference
Thomas, D., Yang, X., Gupta, S., Adeniran, A., Mclaughlin, E., Koedinger, K.: When the tutor becomes the student: Design and evaluation of efficient scenario- based lessons for tutors. In: LAK23: 13th International Learning Analytics and Knowledge Conference. pp. 250–261 (2023)
2023
-
[28]
Thomas, D.R., Borchers, C., Kakarla, S., Lin, J., Bhushan, S., Guo, B., Gatz, E., Koedinger, K.R.: Do tutors learn from equity training and can generative ai assess it? arXiv preprint arXiv:2412.11255 (2024)
2024 arXiv
-
[29]
In: Proceedings of the Eleventh ACM Conference on Learning@ Scale
Thomas, D.R., Lin, J., Bhushan, S., Abboud, R., Gatz, E., Gupta, S., Koedinger, K.R.: Learning and ai evaluation of tutors responding to students engaging in negative self-talk. In: Proceedings of the Eleventh ACM Conference on Learning@ Scale. pp. 481–485 (2024)
2024
-
[30]
Journal of Digital Learning in Teacher Education 35(3), 144–164 (2019)
Thompson, M., Owho-Ovuakporie, K., Robinson, K., Kim, Y.J., Slama, R., Reich, J.: Teacher moments: A digital simulation for preservice teachers to approximate parent–teacher conversations. Journal of Digital Learning in Teacher Education 35(3), 144–164 (2019)
2019
-
[31]
Frontiers in psychology 10, 487662 (2020)
Wisniewski, B., Zierer, K., Hattie, J.: The power of feedback revisited: A meta- analysis of educational feedback research. Frontiers in psychology 10, 487662 (2020)
2020
-
[32]
arXiv preprint arXiv:2501.09824 (2025)
Xu, C., Lin, J., Wu, T., Aleven, V., Koedinger, K.R.: Improving automated feed- back systems for tutor training in low-resource scenarios through data augmenta- tion. arXiv preprint arXiv:2501.09824 (2025)
2025
-
[33]
Jour- nal of the Royal Statistical Society Series B: Statistical Methodology67(2), 301– 320 (2005)
Zou, H., Hastie, T.: Regularization and variable selection via the elastic net. Jour- nal of the Royal Statistical Society Series B: Statistical Methodology67(2), 301– 320 (2005)
2005
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.