Pith. sign in

REVIEW 3 major objections 5 minor 20 references

LLM-supported open-ended self-explanation improves explanation quality on calculus transfer problems that require recognizing missing information, even when students complete far fewer practice problems in the same time.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-07-13 15:20 UTC pith:WVQ47OCS

load-bearing objection Fixed-time three-arm calculus study: LLM open-ended self-explanation improves NEI transfer explanation quality despite far fewer problems; package confound and LLM-scored outcome keep the claim narrow. the 3 major comments →

arxiv 2604.00142 v2 pith:WVQ47OCS submitted 2026-03-31 cs.HC

Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus

classification cs.HC
keywords self-explanationlarge language modelsintelligent tutoring systemsmathematics educationtransfer of learningcalculusexplanation quality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether asking calculus learners to write free-text explanations and get adaptive LLM feedback helps them transfer better than ordinary practice or menu-based explanations, when total practice time is fixed. In a between-subjects study of 92 adults learning limits and derivatives, all three groups improved on a post-test, with no clear differences in ordinary problem-solving accuracy. The open-ended LLM group, however, wrote higher-quality explanations on transfer items where the right answer was “not enough information,” and showed a weaker advantage across all open-ended transfer explanations. That pattern appeared even though those learners finished only about a quarter as many practice problems as the control group. The claim is that deeper, feedback-supported explanation work can buy better articulated reasoning on hard transfer cases without needing more drills.

Core claim

Under a fixed 60-minute practice session, open-ended self-explanation with LLM-generated feedback produced significantly higher-quality written explanations than no self-explanation on “Not Enough Information” (NEI) transfer problems (about 12 percentage points), with only marginal evidence of a broader open-ended transfer explanation advantage and no reliable gains on ordinary post-test accuracy or transfer multiple-choice answers, despite completing substantially fewer practice problems.

What carries the argument

LLM-supported open-ended self-explanation: after solving a practice item correctly, learners write a free-text rationale that an LLM scores against a four-level rubric and returns color-coded iterative feedback (and a reference explanation after two failed attempts).

Load-bearing premise

The gain is assumed to come from open-ended self-explanation as a learning practice, even though the condition also bundled iterative AI scoring and a reference explanation after two failed attempts, which the design cannot separate.

What would settle it

Run the same fixed-time design with an additional arm that gets LLM correctness feedback or the reference explanation without requiring open-ended generation; if the NEI explanation-quality advantage disappears without free-text production, the central claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a between-subjects experiment (N=92) comparing no self-explanation, menu-based self-explanation, and open-ended self-explanation with LLM feedback in a fixed 60-minute calculus practice session on limits and derivatives. All conditions showed positive learning gains with no significant between-condition differences on ordinary post-test problem-solving accuracy. The central positive claim is that the open-ended condition produced higher-quality explanations than control on NEI transfer problems (β=+11.9 percentage points, p=.030), with a marginally significant advantage on combined post-test open-ended transfer explanations (β=+7.3%, p=.057), despite completing far fewer practice problems (16.9 vs 58.9). Transfer MCQ accuracy, including NEI MCQ, did not differ significantly. The authors interpret this as evidence that LLM-supported open-ended self-explanation can improve explanation quality on transfer items that require recognizing insufficient information.

Significance. If the NEI explanation-quality result is robust under independent human scoring and under clearer isolation of generation versus feedback, the paper would be a useful contribution to ITS and self-explanation research. It updates the classic Aleven–Koedinger menu-based and dialogue-based findings with scalable LLM feedback under a fixed-time design that makes the practice-volume tradeoff explicit, and it targets a transfer form (NEI) that is educationally interesting. Strengths include randomization with pre-test covariate ANCOVA, counterbalanced quiz forms, a fixed practice window, step-level logging, and pilot validation of the LLM rubric against human consensus (κ=0.78). The result is not a broad learning-gain claim; it is a narrower claim about explanation quality on a subset of transfer items, which is still of interest if the scoring and attribution concerns are resolved.

major comments (3)
  1. [Method / Results Table 2] Method (LLM grading validation) and Results (NEI Transfer Open-Ended / Table 2): The headline result is LLM-scored NEI transfer open-ended explanation quality (β=+11.9%, p=.030). Validation used GPT-5.1 on 160 pilot explanations (κ=0.78 vs human consensus), then the same model/rubric scored the experimental post-test. Only the open-ended condition received iterative color-coded LLM feedback on the same 0/0.3–0.7/1.0 scale and a reference explanation after two failures (Fig. 2). Control and menu never saw this feedback style. Without a human re-grade of the experimental NEI OE responses that drive the primary claim, the advantage may partly reflect alignment to the training feedback distribution rather than genuine transfer of explanation quality. A human consensus re-score of the experimental NEI OE items (or a pre-registered subset) is needed before the claim can be treated as establish
  2. [Discussion / Limitations] Discussion / Limitations: The open-ended condition bundles free-text generation, iterative LLM evaluation, and a reference explanation after two unsuccessful attempts. The Limitations section correctly notes that the design cannot isolate these components, yet the title, abstract, and framing attribute the NEI OE advantage to "LLM-supported open-ended self-explanation." Given that the only significant contrast is this bundled condition versus control, the causal claim should be narrowed to the full package, or the manuscript should add (or clearly plan) a feedback-only / limited-feedback control. As written, the load-bearing attribution to open-ended generation as a learning mechanism is not supported by the design.
  3. [Results Table 2 / Abstract] Results Table 2 and abstract: Only one primary transfer OE contrast reaches p<.05 (NEI OE); Combined OE is marginal (p=.057), Non-NEI OE is p=.093, and all MCQ transfer contrasts are non-significant (NEI MCQ p=.183). The abstract and discussion already hedge, but the paper still leads with a strong title-level claim. The manuscript should treat NEI OE as a single significant contrast among several related transfer measures, report multiplicity/family-wise considerations or pre-specification more clearly, and avoid overstating breadth of transfer gains.
minor comments (5)
  1. [Table 1] Table 1 reports mean problems completed but no SDs or distributions; given the large volume difference (58.9 vs 16.9), add dispersion and, if available, time-on-explanation vs time-on-problem breakdowns.
  2. [Method / Limitations] The menu-based condition uses one correct principle and two misconceptions, which the Limitations note differs from classic Cognitive Tutor principle selection. Make this design difference more prominent earlier so readers do not over-interpret the menu null result as a direct replication failure of Aleven & Koedinger (2002).
  3. [Results / Analysis] Report exact model degrees of freedom, residual diagnostics, and whether counterbalance interacted with condition for the NEI OE ANCOVA; Adj. R²=.204 is given but full model output would help.
  4. [Throughout] Minor typography: several compound words are concatenated without spaces in the source text (e.g., "menu-basedself-explanation", "NotEnoughInformation"); clean for camera-ready.
  5. [Method] Provide the full grading prompt and rubric examples in an appendix or repository, and state whether the experimental post-test OE scores used the identical prompt as the pilot validation.

Circularity Check

0 steps flagged

No derivation-chain circularity; empirical RCT outcomes are not defined by or fitted to the intervention inputs.

full rationale

This is a between-subjects fixed-time experiment (N=92) reporting ANCOVA contrasts on post-test accuracy and LLM-scored explanation quality. There is no first-principles derivation, no fitted parameter re-labeled as a prediction, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled via self-citation that forces the central claim. The load-bearing result (open-ended vs control NEI OE β=+11.9 pp, p=.030) is an observed group difference under a pre-registered-style design with pre-test covariate; it is not equivalent by construction to any training input. LLM grading was validated against human consensus on a separate pilot sample (κ=0.78) before being applied to experimental responses; shared tooling between practice feedback and post-test scoring is a measurement-validity concern, not a circular reduction of the form Eq. X ≡ Eq. Y. Self-citations to Aleven/Koedinger are ordinary prior-work grounding and do not define the outcome. Per the analyzer rules, score 0 with empty steps is the correct finding for a self-contained empirical comparison.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

This is an empirical education experiment, not a formal derivation. Load-bearing premises are design and measurement choices: fixed practice time, the three-condition operationalizations, the NEI transfer construct, the four-level explanation rubric, and reliance on LLM scoring validated against pilot human ratings. No new physical entities; free parameters are design constants and scoring thresholds rather than fitted scientific constants.

free parameters (3)
  • practice_session_duration
    Fixed at 60 minutes by design; the practice-volume contrast and all between-condition comparisons depend on this time budget.
  • explanation_rubric_score_levels
    Four-level scores 0, 0.3, 0.7, 1.0 define the continuous explanation-quality outcome; thresholds are design choices that map text to the reported percentage-point effects.
  • max_unsuccessful_explanation_attempts_before_reference
    Set to two attempts before showing a reference explanation in the open-ended condition; this feedback intensity is part of the intervention package.
axioms (4)
  • domain assumption Self-explanation is a constructive learning process that can promote transfer by abstracting principles (ICAP / Chi / Rittle-Johnson lineage).
    Invoked in Introduction to motivate open-ended explanation as the active ingredient for transfer explanation quality.
  • domain assumption LLM scores using the shared rubric are sufficiently valid proxies for human explanation quality on post-test items.
    Supported by pilot human agreement (κ up to 0.78 with consensus) but applied to study outcomes without full human re-grading of the N=92 post-test set.
  • standard math ANCOVA Outcome ~ Pre-test + Condition + Counterbalance with Control/AB as references is an adequate causal estimator under randomization.
    Analysis section; standard linear model assumptions (linearity, residual structure, no unmodeled selection) are not extensively checked in text.
  • ad hoc to paper NEI piecewise differentiability items validly measure transfer requiring recognition of insufficient information.
    Materials define NEI vs EI transfer items; the strongest reported effect is concentrated on this constructed item type.

reviewed 2026-07-13 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus." pith.science (2026). https://pith.science/paper/WVQ47OCS

@misc{pith2026260400142,
  author       = {Pith},
  title        = {Pith review of: Practice Less, Explain More: LLM-Supported Self-Explanation Improves Explanation Quality on Transfer Problems in Calculus},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WVQ47OCS}},
  note         = {Machine review of arXiv:2604.00142}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We conducted a between-subjects experiment (N=92) comparing three conditions in a calculus learning environment: no self-explanation (control), menu-based self-explanation, and open-ended self-explanation with LLM-generated feedback. All conditions showed positive learning gains within a fixed 60-minute practice session, with no significant between-condition differences in post-test performance. On transfer questions, the open-ended condition produced significantly higher-quality explanations than control on "Not Enough Information" (NEI) problems ($\beta$=+11.9 percentage points, $p$=.030), though the corresponding NEI multiple-choice accuracy advantage was not significant ($p$=.183). Moreover, across all post-test open-ended explanations, the open-ended condition showed a marginally significant advantage ($\beta$=+7.3%, $p$=.057). These findings suggest that LLM-supported open-ended self-explanation can improve explanation quality on NEI transfer problems, with weaker evidence across broader transfer explanation measures. Notably, these effects emerged even though learners in the open-ended condition completed substantially fewer practice problems within the same practice time.

Figures

Figures reproduced from arXiv: 2604.00142 by Eason Chen, Elizabeth McLaughlin, Jared Cochrane, Jionghao Lin, Ken Koedinger, Meiyi Chen, Meryam Elmir, Mingyu Yuan, Shyam Agarwal, Tongshuang Wu, Xinyi Tang, Yumo Wang, Yvonne Zhao.

Figure 1
Figure 1. Figure 1: Practice interface (control condition). Top: multiple-choice with correctness feedback and progressive hints. Bottom: short-answer with a piecewise function graph [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Self-explanation conditions. Top: menu-based, selecting from three options (one correct, two misconceptions). Bottom: open-ended with LLM feedback showing itera￾tive revision (red = needs revision, yellow = almost there) and a reference explanation after two unsuccessful attempts. Outcome measures. We distinguish three types of post-test outcomes. First, post-test problem-solving accuracy measured performa… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

20 extracted references · 2 canonical work pages

  1. [1]

    In: Proceedings of the 11th Inter- national Conference on Artificial Intelligence in Education

    Aleven, V., Koedinger, K.R., Popescu, O.: A tutorial dialog system to support self-explanation: Evaluation and open questions. In: Proceedings of the 11th Inter- national Conference on Artificial Intelligence in Education. pp. 39–46. IOS Press Amsterdam (2003)

  2. [2]

    In: Building dialogue systems for tutorial applications, papers of the 2000 AAAI Fall Symposium

    Aleven, V., Koedinger, K.R.: The need for tutorial dialog to support self- explanation. In: Building dialogue systems for tutorial applications, papers of the 2000 AAAI Fall Symposium. pp. 65–73 (2000)

  3. [3]

    In: International Conference on Intelligent Tutoring Systems

    Aleven, V., McLaren, B.M., Sewall, J., Koedinger, K.R.: The cognitive tutor au- thoring tools (ctat): Preliminary evaluation of efficiency gains. In: International Conference on Intelligent Tutoring Systems. pp. 61–70. Springer (2006) Practice Less, Explain More: LLM-Supported Self-Explanation 9

  4. [4]

    In: Pro- ceedings of Artificial Intelligence in Education

    Aleven, V., Popescu, O., Koedinger, K.R.: Towards tutorial dialog to support self- explanation: Adding natural language understanding to a cognitive tutor. In: Pro- ceedings of Artificial Intelligence in Education. pp. 246–255 (2001)

  5. [5]

    Cognitive science 26(2), 147–179 (2002)

    Aleven, V.A., Koedinger, K.R.: An effective metacognitive strategy: Learning by doing and explaining with a computer-based cognitive tutor. Cognitive science 26(2), 147–179 (2002)

  6. [6]

    In: International Conference on Artificial Intelligence in Education

    Armfield, D., Chen, E., Omonkulov, A., Tang, X., Lin, J., Thiessen, E., Koedinger, K.: Avalon: a human-in-the-loop llm grading system with instructor calibration and student self-assessment. In: International Conference on Artificial Intelligence in Education. pp. 111–118 (2025)

  7. [7]

    In: Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024)

    Carpenter, D., Min, W., Lee, S., et al.: Assessing student explanations with large language models using fine-tuning and few-shot learning. In: Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024). pp. 403–413. Association for Computational Linguistics (2024)

  8. [8]

    In: International Conference on Artificial Intelligence in Education

    Chen, E., Li, J., Huang, S., Tang, X., Lin, J., Carvalho, P., Koedinger, K.: Iden- tifying effective praise in tutoring: Large language models with transparent expla- nations. In: International Conference on Artificial Intelligence in Education. pp. 157–163 (2025)

  9. [9]

    arXiv preprint arXiv:2410.11123 (2024)

    Chen, E., Wang, D., et al.: A systematic review on prompt engineering in large language models for k-12 stem education. arXiv preprint arXiv:2410.11123 (2024)

  10. [10]

    Cognitive Science13(2), 145–182 (1989).https://doi.org/10.1207/s15516709cog1302_1

    Chi, M.T.H., Bassok, M., Lewis, M.W., Reimann, P., Glaser, R.: Self-explanations: How students study and use examples in learning to solve problems. Cognitive Science13(2), 145–182 (1989).https://doi.org/10.1207/s15516709cog1302_1

  11. [11]

    Cognitive science18(3), 439–477 (1994)

    Chi, M.T., De Leeuw, N., Chiu, M.H., LaVancher, C.: Eliciting self-explanations improves understanding. Cognitive science18(3), 439–477 (1994)

  12. [12]

    Educational psychologist49(4), 219–243 (2014)

    Chi, M.T., Wylie, R.: The icap framework: Linking cognitive engagement to active learning outcomes. Educational psychologist49(4), 219–243 (2014)

  13. [13]

    IAFOR Journal of Education7(1), 93–111 (2019)

    Hajian, S.: Transfer of learning and teaching: A review of transfer theories and effective instructional practices. IAFOR Journal of Education7(1), 93–111 (2019)

  14. [14]

    Kumar,H.,Rothschild,D.M.,Goldstein,D.G.,Hofman,J.M.:Matheducationwith large language models: Peril or promise? In: International Conference on Artificial Intelligence in Education. pp. 60–75. Springer (2025)

  15. [15]

    In: Educational Data Mining Conference 2024 (2024)

    Lin, J., Chen, E., Han, Z., Gurung, A., Thomas, D.R., Tan, W., Nguyen, N.D., Koedinger, K.R.: How can i improve? using gpt to highlight the desired and unde- sired parts of open-ended responses. In: Educational Data Mining Conference 2024 (2024)

  16. [16]

    Journal of experi- mental child psychology104(1), 1–21 (2009)

    Matthews, P., Rittle-Johnson, B.: In pursuit of knowledge: Comparing self- explanations, concepts, and procedures as pedagogical tools. Journal of experi- mental child psychology104(1), 1–21 (2009)

  17. [17]

    British Journal of Educational Psychology83(4), 615–632 (2013)

    McEldoon, K.L., et al.: Is self-explanation worth the time? a comparison to addi- tional practice. British Journal of Educational Psychology83(4), 615–632 (2013)

  18. [18]

    Cognitive science21(1), 1–29 (1997)

    Renkl, A.: Learning from worked-out examples: A study on individual differences. Cognitive science21(1), 1–29 (1997)

  19. [19]

    The Journal of Mathematical Behavior (2024)

    Rittle-Johnson, B.: Encouraging students to explain their ideas when learning mathematics: A psychological perspective. The Journal of Mathematical Behavior (2024)

  20. [20]

    ZDM Mathematics Education49(4),599–611(2017).https://doi.org/10.1007/s11858-017-0834-z

    Rittle-Johnson, B., et al.: Promoting self-explanation to improve mathematics learning: A meta-analysis and instructional design principles. ZDM Mathematics Education49(4),599–611(2017).https://doi.org/10.1007/s11858-017-0834-z

This paper was first reviewed by grok-4.5 on July 13, 2026.