{"id":"d79ba364-6092-4954-a71b-a80e75ff0825","arxiv_id":"2506.17006","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"On-demand AI feedback produced small posttest gains for tutors who chose to use it, with significant benefits in two of seven lessons and no extra time.","lead":"This study let tutors taking online training lessons choose whether to read AI-generated explanatory feedback on their answers, then compared their final quiz scores with tutors who did not have that option. Tutors who chose to use the AI feedback scored higher on two of seven lessons, with no added time cost.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Posttest contamination from reusing LLM feedback text is not ruled out and is consistent with the authors' own response-length finding, so the central claim that the effect reflects learning is not yet established.","rationale":"The reader's weakest assumption is the propensity model in Section 4.1, which is a real concern. But I see an even more direct threat to the central claim: the outcome itself may be contaminated by the feedback text. The paper openly states in Section 5 that learners who received feedback wrote significantly longer posttest explanations and that copying or adapting the Q1 feedback could have inflated performance, but it does not test for textual overlap. Because the posttest open responses are scored by GPT-4o against rubrics emphasizing the very elements the LLM feedback rephrases, even partial reuse would inflate scores. This would make the reported 0.10 SD TOT effect and the two lesson-level effects (0.33, 0.28) an artifact of answer reuse, not evidence of learning. The proposed overlap test directly settles this. If contamination is ruled out, the propensity concern remains and should also be addressed; if contamination is found, the central claim is not yet supported. The reader's conditional verdict is appropriate, so I leave it unchanged.","tokens_in":10986,"tokens_out":8622,"duration_ms":94249,"concrete_test":"Using the released logs, compute ROUGE-L or 5-gram Jaccard overlap between each learner's displayed LLM feedback text (Q1) and their posttest Q7/Q9 responses. Re-run the main TOT model (Section 4.1) and the two significant lesson-level models with the overlap score as a covariate, and also after excluding completions with overlap above a pre-specified threshold (e.g., 0.5). If the effects persist in the no-overlap subset, the learning claim is supported; if they vanish, contamination is the likely driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLM feedback improves learning—rests on posttest open-response scores (Q7/Q9). The LLM feedback shown on Q1 contains a rephrased, rubric-aligned response. Section 5 acknowledges that learners could copy or adapt that text at posttest and that recipients produced six-word-longer Q9 responses, but the authors only measured length, not textual overlap. Because posttest open responses are scored by GPT-4o under rubrics that reward the same key elements the feedback emphasizes, even partial reuse could inflate scores. The authors' assertion that such cases are 'rare' is not supported by evidence; low-stakes does not prevent incidental reuse. This is load-bearing because if reuse drives the 0.10 SD TOT estimate or the lesson-level effects (0.33, 0.28), the observed gains do not demonstrate learning, only answer reuse. The propensity adjustment in Section 4.1 cannot fix outcome contamination.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a randomized-within-subject study of on-demand LLM-generated explanatory feedback in seven scenario-based tutor-training lessons, with 885 learners and 2,648 lesson completions. Learners assigned to the intent-to-treat condition could request GPT-3.5-turbo feedback on an open-response question; all learners received the existing non-LLM corrective feedback. The authors compare posttest performance across learners who received feedback (treatment-on-the-treated), those who declined it, and those without access. They report a TOT association of 0.10 SD (95% CI [0.01, 0.19], p = .023), no significant ITT main effect, a propensity-adjusted analysis showing significant lesson-level benefits for two of seven lessons (Giving Effective Praise, beta = 0.33; Supporting a Growth Mindset, beta = 0.28), no significant completion-time cost, and predominantly positive learner ratings. The paper contributes open datasets, prompts, and rubrics and frames the central conclusion as: LLM feedback helps learning only when learners choose to engage with it.","tokens_in":11077,"tokens_out":3201,"duration_ms":35832,"significance":"If the reported effects are genuine, the paper provides useful evidence about a realistic, low-cost feedback augmentation in existing online tutor training, and its open resources support replication. The within-subject lesson-level assignment, the explicit handling of self-selection via propensity scoring, and the effort to evaluate open responses with rubric-based LLM scoring are notable strengths. However, the central learning claim is only as strong as its weakest load-bearing assumptions: the TOT estimate is observational, the propensity adjustment cannot remove unmeasured confounders, the lesson-level effects are not corrected for multiple comparisons, and the authors concede they cannot rule out posttest answer reuse. The paper therefore advances a plausible but not yet fully established claim, and the currently available evidence does not justify the strength of the concluding statements.","major_comments":[{"comment":"The manuscript's central claim that LLM feedback improves learning rests on posttest open-response scores (Q7/Q9), which are scored by GPT-4o under rubrics rewarding the same key elements that the Q1 LLM feedback reinforces. The paper concedes that learners could copy or adapt the feedback text at posttest and that the only follow-up analysis measured response length, not textual overlap. The statement that such cases were 'rare' is not supported by any reported evidence; low-stakes assessments do not by themselves prevent incidental reuse. Because if reuse inflates the 0.10 SD TOT estimate or the lesson-level effects (0.33 and 0.28), the observed gains would not demonstrate learning, this issue is load-bearing. The authors should report a quantitative textual-overlap analysis between Q1 feedback and Q7/Q9 responses, re-estimate the effects after excluding suspected reuse cases, or otherwise provide direct evidence that the results are not driven by answer reuse.","section":"§5 Discussion, posttest contamination"},{"comment":"The propensity model in §4.1 is trained on engagement, response, and session features from the ITT condition and applied to the control condition to predict the number of LLM feedback requests. The authors acknowledge that unmeasured confounders, such as help-seeking skill or motivation, may remain, and the Discussion correctly hedges that engagement 'may also reflect preexisting learning differences.' However, the abstract and concluding paragraph state as the key empirical finding that 'the effectiveness of such feedback depends on learners' willingness to seek and engage with it.' This causal claim is stronger than the propensity-adjusted analysis supports, since the adjustment cannot fully remove self-selection. The paper should either temper the conclusion to an associational claim or add a sensitivity analysis (for example, reporting how large an unmeasured confounder would need to be to explain the TOT effect, or exploiting the 15.7% API-failure rate as a possible instrument for actual receipt).","section":"§4.1, propensity adjustment and causal interpretation"},{"comment":"Seven lesson-level propensity-adjusted treatment effects are reported, with two reaching p < .05. No correction for multiple comparisons is mentioned, and the overall ITT-by-propensity interaction in Table 2 is not significant (beta = 0.04, 95% CI [-0.04, 0.12], p = .307). Given the nonsignificant interaction, the two significant lesson-level effects should be presented as exploratory, and the authors should report adjusted p-values, confidence intervals, or a false-discovery-rate control so readers can assess the strength of the cross-lesson support for the central claim.","section":"§4.1, lesson-level comparisons"},{"comment":"The abstract states that 'Learners with a higher predicted likelihood of engaging with LLM feedback scored significantly higher at posttest than those with lower propensity.' As reported in Table 2, the significant propensity coefficient (beta = 0.06, p = .040) is for the control condition's predicted propensity, not for treated learners. The subsequent sentence about two significant lessons does not mention that the overall ITT-by-propensity interaction was not significant. The abstract should be reworded to match the actual model results, clarifying that the selection effect is observed in the control group and that the treatment-contingent benefit is supported only by the two exploratory lesson-level analyses.","section":"Abstract and §4.1, interpretation of propensity results"}],"minor_comments":[{"comment":"The text first says learners were randomly assigned to one of two conditions for each lesson, but later says that for some lessons all learners were assigned to the ITT condition for a period. These statements are in tension; please clarify the actual assignment procedure and whether the analysis accounts for partially non-random assignment periods.","section":"§3.2, assignment mechanism"},{"comment":"The follow-up analysis reporting that feedback recipients produced responses six words longer at Q9 should include an effect size, confidence interval, or p-value, so the reader can judge the strength of the evidence.","section":"§5, response-length finding"},{"comment":"The text says IRR was established in prior open-source work for most lessons, but no citation is provided for that prior work; please add the reference or state which lessons came from which source.","section":"§3.3, inter-rater reliability citations"},{"comment":"Several typographical issues occur, including missing spaces ('tutorsas learners', 'bygpt-3.5-turbo') and a table header that reads 'Feedback Offered: Intent-to-Treat (ITT)'; these should be corrected in the final version.","section":"Global, formatting"},{"comment":"The lesson-level models are described as 'separate regressions,' but it would be helpful to state whether they include the same random effects as the main model and whether the propensity covariate enters as a continuous variable or a stratifier.","section":"§4.1, model reporting"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: this is a careful, honest field experiment, and the most interesting result is conditional—learners who chose to use on-demand LLM feedback did better on the posttest (about 0.10 SD, with significant lesson-level effects in two of seven lessons) and did not spend more time on the lesson. The ITT effect is null, which is important: simply offering the feedback does nothing. That's a useful, practical message for training systems that already provide corrective feedback.\n\nWhat's new: prior LLM-feedback studies covered math and writing; this one applies it to scenario-based tutor training, with on-demand use and a propensity-adjusted TOT analysis. The help-seeking moderator is a real empirical contribution, and the team shares data, prompts, rubrics, and code. That's more than most papers do.\n\nThe main soft spot is exactly the one the authors put in the Discussion: they cannot rule out that learners reused the LLM feedback text when writing their posttest responses. Their own follow-up shows recipients wrote six words longer on Q9, which is consistent with copying or adapting. They say such cases were rare, but they never measured overlap. Since GPT-4o scores the posttest with rubrics that reward the same key elements the feedback emphasizes, even a small number of reused phrases could inflate the TOT estimate. This is load-bearing, because the TOT claim is the paper's centerpiece. Secondary issues: two significant lessons out of seven with no multiple-comparison correction, and the propensity model can't fully remove unmeasured traits like help-seeking skill.\n\nI want to be fair: they flagged the reuse concern themselves, and they're open about the 15.7% API failure rate and the imperfection of propensity scoring. The data availability and transparency make this a reproducible paper even if the causal reading is conditional.\n\nSend it to peer review. A good reviewer will ask for a textual-overlap analysis or a robustness check excluding verbatim matches, and either pre-specify the primary lesson or correct for multiple comparisons. With those, this becomes a solid applied paper. I'd cite it for the design and the null time-cost result.","headline":"Useful, honest field experiment on on-demand LLM feedback, but the claim that posttest gains reflect learning rather than answer reuse is not yet nailed down.","tokens_in":11682,"tokens_out":2264,"would_cite":true,"duration_ms":23322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that on-demand LLM-generated explanatory feedback improves posttest performance in scenario-based tutor lessons, but only among learners who actually request and engage with it—offering the feedback alone produces no…","keywords":["LLM-generated feedback","large language models","scenario-based learning","tutor training","self-selection bias","propensity scoring","intent-to-treat analysis","educational technology"],"falsifier":"Re-run the analysis with a logged measure of how long each learner actually spent viewing the LLM feedback: if learners who never open the feedback show the same posttest gain as those who read it, the effect is a selection artifact, not a learning effect. Alternatively, a randomized encouragement design that pushes low-propensity learners to request feedback would settle whether the benefit is caused by the feedback or by the traits that lead learners to seek it.","tokens_in":10730,"feed_emoji":"🎓","tokens_out":5416,"duration_ms":49365,"temperature":0.7,"pith_summary":"This paper reports a field experiment embedded in seven online scenario-based lessons for college-student tutors. It claims that on-demand explanatory feedback generated by GPT-3.5-turbo improves posttest performance only among learners who actually request and receive it: the intent-to-treat effect of merely offering feedback was not significant, while the treatment-on-the-treated effect was 0.10 SD, and after propensity adjustment two lessons showed significant effects of 0.33 and 0.28 SD. The result matters because LLM feedback added no measurable completion time and was rated helpful by 94% of users, suggesting a low-cost enhancement to existing feedback systems. The paper's central conclusion is conditional: the benefit of LLM feedback depends on learners' willingness to seek support, not just on feedback quality.","feed_headline":"LLM feedback boosts learning only for those who request it","feed_subtitle":"In 2,648 tutor lessons, offering feedback alone did nothing; engaging with it raised posttest scores with no time cost.","key_machinery":"The central mechanism is an on-demand LLM feedback loop placed inside a predict-observe-explain lesson cycle: after a tutor submits an open response, the model classifies it against a predefined schema and, for incorrect responses, generates a minimally rephrased, research-aligned correction that the learner can request. The causal identification machinery is principal stratification, a method that estimates the treatment effect among learners who would actually use the feedback, combined with an ElasticNet propensity model (a regularized regression with L1 and L2 penalties) that predicts each learner's number of feedback requests from engagement and response features. That predicted propensity is applied to the control group to construct a fairer comparison and to test whether high-propensity learners benefit more.","core_discovery":"Across 2,648 lesson completions by 885 tutor learners, learners who received LLM-generated explanatory feedback on their open responses scored significantly higher at posttest than those who did not receive or use it (0.10 SD, 95% CI [0.01, 0.19], p = .023). Simply being assigned the option to request feedback produced no significant overall gain, which the authors interpret as evidence that actual engagement drives the benefit. After using principal stratification with an ElasticNet propensity model trained on engagement, response, and session features to predict feedback requests, two lessons—Giving Effective Praise and Supporting a Growth Mindset—showed statistically significant propensity-adjusted effects of 0.33 and 0.28 SD, while other lessons showed smaller, non-significant effects. Receiving feedback did not increase lesson completion time (a 9-second average difference), and 94% of learners who rated the feedback called it helpful. The authors conclude that LLM feedback supports learning when learners choose to engage with it, and that its effectiveness is moderated by help-seeking propensity rather than guaranteed by availability.","pith_inferences":["If help-seeking propensity is itself a trainable skill, then teaching learners when and how to request feedback could amplify the effect beyond the modest gains reported here.","A direct robustness check of the copying concern would compare posttest open responses to the LLM feedback text; high similarity would suggest part of the gain is response imitation rather than durable learning.","The propensity adjustment may under- or over-correct because it uses engagement proxies; adding measured motivational or metacognitive covariates could sharpen the causal estimate and reveal whether high-propensity learners benefit from deeper processing."],"forward_implications":["Existing online learning systems that already give corrective feedback could add on-demand LLM feedback at near-zero time cost and expect modest posttest gains among learners who request it.","Offering feedback is not enough; to realize the benefit, systems should encourage help-seeking or otherwise increase the likelihood that learners engage with available support.","Because lesson-level effects varied from 0.33 SD in Giving Effective Praise to a negative trend in Helping Students Manage Inequity, the benefit of LLM feedback depends on content and task type, not just feedback delivery.","The lack of completion-time differences implies LLM feedback can be integrated into short scenario-based training without sacrificing efficiency."],"supporting_citations":[{"why":"Supplies the principal stratification method used to estimate the treatment effect among learners who would actually request feedback.","marker":"[23]"},{"why":"Provides the ElasticNet regularized regression used to model each learner's propensity to request LLM feedback.","marker":"[33]"},{"why":"Supplies the cross-validated ROC-AUC procedure used to validate the propensity predictions.","marker":"[24]"},{"why":"Establishes the feedback design principles, including immediacy and correctness, that motivate the study's explanatory feedback approach.","marker":"[14]"},{"why":"Provides evidence that elaborated feedback yields larger learning effects, justifying the choice of explanatory LLM feedback.","marker":"[19]"},{"why":"Shows that ChatGPT-generated help can produce learning gains comparable to human-authored help, providing prior context for LLM feedback effectiveness.","marker":"[21]"},{"why":"Frames help-seeking behavior as a key moderating mechanism, supporting the interpretation that engagement drives the benefit.","marker":"[2]"},{"why":"Raises the concern that copying LLM feedback could inflate posttest performance, which the authors address as an alternative explanation.","marker":"[11]"}],"fun_headline_variants":["LLM feedback boosts learning only when learners ask","AI feedback helps only if students choose to use it","Requesting AI feedback, not just receiving it, lifts scores","On-demand AI feedback supports learning when used"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the propensity model, trained on engagement and response behaviors, captures every trait that makes learners both more likely to request LLM feedback and more likely to score higher at posttest; if unmeasured traits such as help-seeking skill or motivation remain, the adjusted comparison does not identify a causal effect.","fun_headline_variants_meta":{"raw":{"variants":["LLM feedback boosts learning only when learners ask","AI feedback helps only if students choose to use it","Requesting AI feedback, not just receiving it, lifts scores","On-demand AI feedback supports learning when used"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1503,"prompt_tokens":1030,"completion_tokens":473,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":411}},"tokens_in":646,"tokens_out":473,"duration_ms":5298,"temperature":1.0,"reasoning_tokens":411,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:13:38.635326+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis with a logged measure of how long each learner actually spent viewing the LLM feedback: if learners who never open the feedback show the same posttest gain as those who read it, the effect is a selection artifact, not a learning effect. Alternatively, a randomized encouragement design that pushes low-propensity learners to request feedback would settle whether the benefit is caused by the feedback or by the traits that lead learners to seek it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the principal stratification method used to estimate the treatment effect among learners who would actually request feedback."},{"cited_title":"Jour- nal of the Royal Statistical Society Series B: Statistical Methodology67(2), 301– 320 (2005)","cited_arxiv_id":null,"evidence_quote":"Provides the ElasticNet regularized regression used to model each learner's propensity to request LLM feedback."},{"cited_title":"GEEPERs: Principal Stratification using Principal Scores and Stacked Estimating Equations","cited_arxiv_id":"2212.10406","evidence_quote":"Supplies the cross-validated ROC-AUC procedure used to validate the propensity predictions."},{"cited_title":"Journal of Educa- tional Psychology114(8), 1743 (2022)","cited_arxiv_id":null,"evidence_quote":"Provides evidence that elaborated feedback yields larger learning effects, justifying the choice of explanatory LLM feedback."},{"cited_title":"Plos one19(5), e0304013 (2024)","cited_arxiv_id":null,"evidence_quote":"Shows that ChatGPT-generated help can produce learning gains comparable to human-authored help, providing prior context for LLM feedback effectiveness."},{"cited_title":"International Journal of Artificial Intelligence in Education16(2), 101–128 (2006) 14 D","cited_arxiv_id":null,"evidence_quote":"Frames help-seeking behavior as a key moderating mechanism, supporting the interpretation that engagement drives the benefit."},{"cited_title":"British Journal of Educational Technology (2024)","cited_arxiv_id":null,"evidence_quote":"Raises the concern that copying LLM feedback could inflate posttest performance, which the authors address as an alternative explanation."}],"review_version":1}