REVIEW 3 major objections 5 minor 20 references
LLM-supported open-ended self-explanation improves explanation quality on calculus transfer problems that require recognizing missing information, even when students complete far fewer practice problems in the same time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-07-13 15:20 UTC pith:WVQ47OCS
load-bearing objection Fixed-time three-arm calculus study: LLM open-ended self-explanation improves NEI transfer explanation quality despite far fewer problems; package confound and LLM-scored outcome keep the claim narrow. the 3 major comments →
Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a fixed 60-minute practice session, open-ended self-explanation with LLM-generated feedback produced significantly higher-quality written explanations than no self-explanation on “Not Enough Information” (NEI) transfer problems (about 12 percentage points), with only marginal evidence of a broader open-ended transfer explanation advantage and no reliable gains on ordinary post-test accuracy or transfer multiple-choice answers, despite completing substantially fewer practice problems.
What carries the argument
LLM-supported open-ended self-explanation: after solving a practice item correctly, learners write a free-text rationale that an LLM scores against a four-level rubric and returns color-coded iterative feedback (and a reference explanation after two failed attempts).
Load-bearing premise
The gain is assumed to come from open-ended self-explanation as a learning practice, even though the condition also bundled iterative AI scoring and a reference explanation after two failed attempts, which the design cannot separate.
What would settle it
Run the same fixed-time design with an additional arm that gets LLM correctness feedback or the reference explanation without requiring open-ended generation; if the NEI explanation-quality advantage disappears without free-text production, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a between-subjects experiment (N=92) comparing no self-explanation, menu-based self-explanation, and open-ended self-explanation with LLM feedback in a fixed 60-minute calculus practice session on limits and derivatives. All conditions showed positive learning gains with no significant between-condition differences on ordinary post-test problem-solving accuracy. The central positive claim is that the open-ended condition produced higher-quality explanations than control on NEI transfer problems (β=+11.9 percentage points, p=.030), with a marginally significant advantage on combined post-test open-ended transfer explanations (β=+7.3%, p=.057), despite completing far fewer practice problems (16.9 vs 58.9). Transfer MCQ accuracy, including NEI MCQ, did not differ significantly. The authors interpret this as evidence that LLM-supported open-ended self-explanation can improve explanation quality on transfer items that require recognizing insufficient information.
Significance. If the NEI explanation-quality result is robust under independent human scoring and under clearer isolation of generation versus feedback, the paper would be a useful contribution to ITS and self-explanation research. It updates the classic Aleven–Koedinger menu-based and dialogue-based findings with scalable LLM feedback under a fixed-time design that makes the practice-volume tradeoff explicit, and it targets a transfer form (NEI) that is educationally interesting. Strengths include randomization with pre-test covariate ANCOVA, counterbalanced quiz forms, a fixed practice window, step-level logging, and pilot validation of the LLM rubric against human consensus (κ=0.78). The result is not a broad learning-gain claim; it is a narrower claim about explanation quality on a subset of transfer items, which is still of interest if the scoring and attribution concerns are resolved.
major comments (3)
- [Method / Results Table 2] Method (LLM grading validation) and Results (NEI Transfer Open-Ended / Table 2): The headline result is LLM-scored NEI transfer open-ended explanation quality (β=+11.9%, p=.030). Validation used GPT-5.1 on 160 pilot explanations (κ=0.78 vs human consensus), then the same model/rubric scored the experimental post-test. Only the open-ended condition received iterative color-coded LLM feedback on the same 0/0.3–0.7/1.0 scale and a reference explanation after two failures (Fig. 2). Control and menu never saw this feedback style. Without a human re-grade of the experimental NEI OE responses that drive the primary claim, the advantage may partly reflect alignment to the training feedback distribution rather than genuine transfer of explanation quality. A human consensus re-score of the experimental NEI OE items (or a pre-registered subset) is needed before the claim can be treated as establish
- [Discussion / Limitations] Discussion / Limitations: The open-ended condition bundles free-text generation, iterative LLM evaluation, and a reference explanation after two unsuccessful attempts. The Limitations section correctly notes that the design cannot isolate these components, yet the title, abstract, and framing attribute the NEI OE advantage to "LLM-supported open-ended self-explanation." Given that the only significant contrast is this bundled condition versus control, the causal claim should be narrowed to the full package, or the manuscript should add (or clearly plan) a feedback-only / limited-feedback control. As written, the load-bearing attribution to open-ended generation as a learning mechanism is not supported by the design.
- [Results Table 2 / Abstract] Results Table 2 and abstract: Only one primary transfer OE contrast reaches p<.05 (NEI OE); Combined OE is marginal (p=.057), Non-NEI OE is p=.093, and all MCQ transfer contrasts are non-significant (NEI MCQ p=.183). The abstract and discussion already hedge, but the paper still leads with a strong title-level claim. The manuscript should treat NEI OE as a single significant contrast among several related transfer measures, report multiplicity/family-wise considerations or pre-specification more clearly, and avoid overstating breadth of transfer gains.
minor comments (5)
- [Table 1] Table 1 reports mean problems completed but no SDs or distributions; given the large volume difference (58.9 vs 16.9), add dispersion and, if available, time-on-explanation vs time-on-problem breakdowns.
- [Method / Limitations] The menu-based condition uses one correct principle and two misconceptions, which the Limitations note differs from classic Cognitive Tutor principle selection. Make this design difference more prominent earlier so readers do not over-interpret the menu null result as a direct replication failure of Aleven & Koedinger (2002).
- [Results / Analysis] Report exact model degrees of freedom, residual diagnostics, and whether counterbalance interacted with condition for the NEI OE ANCOVA; Adj. R²=.204 is given but full model output would help.
- [Throughout] Minor typography: several compound words are concatenated without spaces in the source text (e.g., "menu-basedself-explanation", "NotEnoughInformation"); clean for camera-ready.
- [Method] Provide the full grading prompt and rubric examples in an appendix or repository, and state whether the experimental post-test OE scores used the identical prompt as the pilot validation.
Circularity Check
No derivation-chain circularity; empirical RCT outcomes are not defined by or fitted to the intervention inputs.
full rationale
This is a between-subjects fixed-time experiment (N=92) reporting ANCOVA contrasts on post-test accuracy and LLM-scored explanation quality. There is no first-principles derivation, no fitted parameter re-labeled as a prediction, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled via self-citation that forces the central claim. The load-bearing result (open-ended vs control NEI OE β=+11.9 pp, p=.030) is an observed group difference under a pre-registered-style design with pre-test covariate; it is not equivalent by construction to any training input. LLM grading was validated against human consensus on a separate pilot sample (κ=0.78) before being applied to experimental responses; shared tooling between practice feedback and post-test scoring is a measurement-validity concern, not a circular reduction of the form Eq. X ≡ Eq. Y. Self-citations to Aleven/Koedinger are ordinary prior-work grounding and do not define the outcome. Per the analyzer rules, score 0 with empty steps is the correct finding for a self-contained empirical comparison.
Axiom & Free-Parameter Ledger
free parameters (3)
- practice_session_duration
- explanation_rubric_score_levels
- max_unsuccessful_explanation_attempts_before_reference
axioms (4)
- domain assumption Self-explanation is a constructive learning process that can promote transfer by abstracting principles (ICAP / Chi / Rittle-Johnson lineage).
- domain assumption LLM scores using the shared rubric are sufficiently valid proxies for human explanation quality on post-test items.
- standard math ANCOVA Outcome ~ Pre-test + Condition + Counterbalance with Control/AB as references is an adequate causal estimator under randomization.
- ad hoc to paper NEI piecewise differentiability items validly measure transfer requiring recognition of insufficient information.
Cite this review
Pith. "Pith review of Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus." pith.science (2026). https://pith.science/paper/WVQ47OCS
@misc{pith2026260400142,
author = {Pith},
title = {Pith review of: Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus},
year = {2026},
howpublished = {\url{https://pith.science/paper/WVQ47OCS}},
note = {Machine review of arXiv:2604.00142}
}
read the original abstract
We conducted a between-subjects experiment (N=92) comparing three conditions in a calculus learning environment: no self-explanation (control), menu-based self-explanation, and open-ended self-explanation with LLM-generated feedback. All conditions showed positive learning gains within a fixed 60-minute practice session, with no significant between-condition differences in post-test performance. On transfer questions, the open-ended condition produced significantly higher-quality explanations than control on "Not Enough Information" (NEI) problems ($\beta$=+11.9 percentage points, $p$=.030), though the corresponding NEI multiple-choice accuracy advantage was not significant ($p$=.183). Moreover, across all post-test open-ended explanations, the open-ended condition showed a marginally significant advantage ($\beta$=+7.3%, $p$=.057). These findings suggest that LLM-supported open-ended self-explanation can improve explanation quality on NEI transfer problems, with weaker evidence across broader transfer explanation measures. Notably, these effects emerged even though learners in the open-ended condition completed substantially fewer practice problems within the same practice time.
Figures
Reference graph
Works this paper leans on
-
[1]
In: Proceedings of the 11th Inter- national Conference on Artificial Intelligence in Education
Aleven, V., Koedinger, K.R., Popescu, O.: A tutorial dialog system to support self-explanation: Evaluation and open questions. In: Proceedings of the 11th Inter- national Conference on Artificial Intelligence in Education. pp. 39–46. IOS Press Amsterdam (2003)
2003
-
[2]
In: Building dialogue systems for tutorial applications, papers of the 2000 AAAI Fall Symposium
Aleven, V., Koedinger, K.R.: The need for tutorial dialog to support self- explanation. In: Building dialogue systems for tutorial applications, papers of the 2000 AAAI Fall Symposium. pp. 65–73 (2000)
2000
-
[3]
In: International Conference on Intelligent Tutoring Systems
Aleven, V., McLaren, B.M., Sewall, J., Koedinger, K.R.: The cognitive tutor au- thoring tools (ctat): Preliminary evaluation of efficiency gains. In: International Conference on Intelligent Tutoring Systems. pp. 61–70. Springer (2006) Practice Less, Explain More: LLM-Supported Self-Explanation 9
2006
-
[4]
In: Pro- ceedings of Artificial Intelligence in Education
Aleven, V., Popescu, O., Koedinger, K.R.: Towards tutorial dialog to support self- explanation: Adding natural language understanding to a cognitive tutor. In: Pro- ceedings of Artificial Intelligence in Education. pp. 246–255 (2001)
2001
-
[5]
Cognitive science 26(2), 147–179 (2002)
Aleven, V.A., Koedinger, K.R.: An effective metacognitive strategy: Learning by doing and explaining with a computer-based cognitive tutor. Cognitive science 26(2), 147–179 (2002)
2002
-
[6]
In: International Conference on Artificial Intelligence in Education
Armfield, D., Chen, E., Omonkulov, A., Tang, X., Lin, J., Thiessen, E., Koedinger, K.: Avalon: a human-in-the-loop llm grading system with instructor calibration and student self-assessment. In: International Conference on Artificial Intelligence in Education. pp. 111–118 (2025)
2025
-
[7]
In: Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024)
Carpenter, D., Min, W., Lee, S., et al.: Assessing student explanations with large language models using fine-tuning and few-shot learning. In: Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024). pp. 403–413. Association for Computational Linguistics (2024)
2024
-
[8]
In: International Conference on Artificial Intelligence in Education
Chen, E., Li, J., Huang, S., Tang, X., Lin, J., Carvalho, P., Koedinger, K.: Iden- tifying effective praise in tutoring: Large language models with transparent expla- nations. In: International Conference on Artificial Intelligence in Education. pp. 157–163 (2025)
2025
-
[9]
arXiv preprint arXiv:2410.11123 (2024)
Chen, E., Wang, D., et al.: A systematic review on prompt engineering in large language models for k-12 stem education. arXiv preprint arXiv:2410.11123 (2024)
Pith/arXiv arXiv 2024
-
[10]
Cognitive Science13(2), 145–182 (1989).https://doi.org/10.1207/s15516709cog1302_1
Chi, M.T.H., Bassok, M., Lewis, M.W., Reimann, P., Glaser, R.: Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science13(2), 145–182 (1989).https://doi.org/10.1207/s15516709cog1302_1
-
[11]
Cognitive science18(3), 439–477 (1994)
Chi, M.T., De Leeuw, N., Chiu, M.H., LaVancher, C.: Eliciting self-explanations improves understanding. Cognitive science18(3), 439–477 (1994)
1994
-
[12]
Educational psychologist49(4), 219–243 (2014)
Chi, M.T., Wylie, R.: The icap framework: Linking cognitive engagement to active learning outcomes. Educational psychologist49(4), 219–243 (2014)
2014
-
[13]
IAFOR Journal of Education7(1), 93–111 (2019)
Hajian, S.: Transfer of learning and teaching: A review of transfer theories and effective instructional practices. IAFOR Journal of Education7(1), 93–111 (2019)
2019
-
[14]
Kumar,H.,Rothschild,D.M.,Goldstein,D.G.,Hofman,J.M.:Matheducationwith large language models: Peril or promise? In: International Conference on Artificial Intelligence in Education. pp. 60–75. Springer (2025)
2025
-
[15]
In: Educational Data Mining Conference 2024 (2024)
Lin, J., Chen, E., Han, Z., Gurung, A., Thomas, D.R., Tan, W., Nguyen, N.D., Koedinger, K.R.: How can i improve? using gpt to highlight the desired and unde- sired parts of open-ended responses. In: Educational Data Mining Conference 2024 (2024)
2024
-
[16]
Journal of experi- mental child psychology104(1), 1–21 (2009)
Matthews, P., Rittle-Johnson, B.: In pursuit of knowledge: Comparing self- explanations, concepts, and procedures as pedagogical tools. Journal of experi- mental child psychology104(1), 1–21 (2009)
2009
-
[17]
British Journal of Educational Psychology83(4), 615–632 (2013)
McEldoon, K.L., et al.: Is self-explanation worth the time? a comparison to addi- tional practice. British Journal of Educational Psychology83(4), 615–632 (2013)
2013
-
[18]
Cognitive science21(1), 1–29 (1997)
Renkl, A.: Learning from worked-out examples: A study on individual differences. Cognitive science21(1), 1–29 (1997)
1997
-
[19]
The Journal of Mathematical Behavior (2024)
Rittle-Johnson, B.: Encouraging students to explain their ideas when learning mathematics: A psychological perspective. The Journal of Mathematical Behavior (2024)
2024
-
[20]
ZDM Mathematics Education49(4),599–611(2017).https://doi.org/10.1007/s11858-017-0834-z
Rittle-Johnson, B., et al.: Promoting self-explanation to improve mathematics learning: A meta-analysis and instructional design principles. ZDM Mathematics Education49(4),599–611(2017).https://doi.org/10.1007/s11858-017-0834-z
This paper was first reviewed by grok-4.5 on July 13, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.