REVIEW 4 major objections 6 minor 213 references
Off-the-shelf generative AI raises students' unaided test scores by 0.27 SD immediately, and most of that gain still shows up one week later without AI.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 06:16 UTC pith:KZD7FDJ4
load-bearing objection Clean lab RCT: off-the-shelf AI raises unaided test scores ~0.27 SD immediately and one week later, with delayed essay gains that stick for augmentation users; external validity is the real limit, not internal identification. the 4 major comments →
Experimental Evidence on the Learning Impact of Generative AI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Random assignment to off-the-shelf generative AI during a short learning-and-essay phase raises unaided knowledge-test scores by 0.27 SD immediately and by a similar amount about one week later without AI. Higher-order essay quality improves mainly after AI is removed, and those delayed gains are larger for students who use AI to explain concepts than for those who use it to generate text.
What carries the argument
The augmentation-versus-automation classification of ChatGPT conversation logs, which separates students who use AI as a tutor (explain, clarify, feedback) from those who use it to produce draft text. That split organizes both short-run versus retained effects and students' own mental models of how AI affects learning.
Load-bearing premise
That a fixed-time lab session with elite undergraduates on low-prior-knowledge topics, followed by one-week retention, identifies the learning effect students would get when they choose how long to study and often use AI to finish faster.
What would settle it
A field experiment in ordinary coursework where students choose total study time and AI access either raises or lowers total learning once time reallocation is allowed, or where one-week retention gains disappear when topics are already familiar and stakes are real grades.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a proctored, in-person RCT with 211 Middlebury undergraduates who learn an unfamiliar technical topic and write an analytical essay under AI-allowed or AI-forbidden conditions, then take unaided knowledge tests and write unaided essays immediately and about one week later. Random assignment raises ChatGPT use by ~67 pp. The main finding is that AI access raises unaided test scores by 0.27 SD immediately (ITT 6.7 pp on a 56.3% control mean) and by a similar 0.27 SD one week later without AI (~76% of the immediate effect). Session One essays show more AI-detected text and little quality change; Session Two unaided essays improve mainly in writing style/clarity and relevance, with larger delayed quality gains among LLM-classified “augmentation” users than “automation” users. Mechanisms include a shift of time from drafting toward reading/searching and higher reported enjoyment. The paper also reports beliefs and open-ended causal narratives about AI and learning.
Significance. If the estimates hold, this is among the cleanest experimental answers to whether off-the-shelf generative AI builds durable academic human capital rather than only short-run task performance. Strengths include: random assignment with a large first stage; multi-method compliance monitoring; unaided immediate and one-week retention assessments; dual human and AI essay grading plus objective linguistic and AI-detection measures; and a transparent literature meta-comparison. The augmentation-versus-automation heterogeneity and the time-reallocation/enjoyment mechanisms are policy-relevant and map onto students’ own mental models. The design deliberately studies unrestricted chatbots against a no-AI counterfactual, which is more informative for real student use than many scaffolded-tutor designs. External validity to ordinary coursework with endogenous study time remains the main limit, which the authors already flag.
major comments (4)
- Table 7, Panel D and §4.4: AI access raises any integrity violation by 12.6 pp (p=0.005). The back-of-envelope that cheating can explain ~2.2 pp of the 6.7 pp Session One test ITT (~one-third) is load-bearing for interpreting the knowledge gains as learning rather than test-taking contamination. Please report the main test-score ITT/TOT for Sessions One and Two after excluding proctor-flagged and self-reported violators (and a joint “any violation” sample), and clarify whether Session Two retention survives that restriction. If the retention effect is robust, state that prominently; if not, revise the learning interpretation accordingly.
- Table 8 and §5.3: The automation/augmentation split is constructed from LLM labels of treated students’ ChatGPT logs and is endogenous among users (different prompting, time use, and AI-detection rates). The differential fade-out is informative as descriptive heterogeneity, but several passages read as if use mode is a causal treatment. Please reframe Table 8 as non-causal heterogeneity among treated users, report balance of baseline ability/AI experience across use types, and avoid language that implies random assignment of automation vs augmentation.
- Table 6, columns 4–6 and Abstract: Session Two overall quality is +0.31 points (0.20 SD, p=0.143); only writing style and relevance are significant. The abstract’s claim that essay quality “improves in style and relevance” is accurate, but the introduction and §5.2 sometimes elevate this to broader “higher-order skills.” Please align the main text with the dimension-level pattern, report multiple-testing-adjusted inference for the five dimensions (or pre-specify primary essay outcomes), and avoid treating the imprecise overall quality index as established.
- Conclusion and §2: The design holds total learning time roughly fixed (~33 minutes). The paper correctly notes that ordinary AI use often saves time. Because the central policy claim is about learning impact of AI access, please add a short quantitative discussion of how large a reduction in time-on-task would be needed to offset the 0.27 SD retention gain under alternative assumptions, so readers can map the lab ITT to settings with endogenous study time.
minor comments (6)
- Figure 5 / Appendix B.7: State more clearly which of the 22 literature estimates are ITT vs TOT and whether all use unassisted outcomes only, so the grand mean of 0.18 SD is interpretable.
- Table 4: Self-assessed knowledge is essentially flat while objective scores rise. A brief discussion of why subjective knowledge does not track the test gains would help (calibration, ceiling of the 0–10 scale, or different construct).
- Appendix Table A1 and take-up: White students are less likely to use AI among the treated. Given the first-stage is not universal, a short note on whether TOT is driven by particular subgroups would be useful.
- §3.3 / double-lasso: Report the selected controls for the main test-score specifications (or an appendix table) so readers can see what residual imbalance is being adjusted.
- Figure 8 and Appendix C: The narrative coding is interesting but long relative to the experimental contribution; consider moving more of the causal-graph material to the appendix and keeping one summary figure in the main text.
- Typos/clarity: “whereasautomation” missing space in the abstract; ensure consistent Session One/Two capitalization; check that N varies slightly across tables (essay missingness) are explained once in a note.
Circularity Check
No circularity: central claims are experimental ITTs from random assignment, not quantities forced by fitted parameters or self-definition.
full rationale
The paper’s load-bearing results are intent-to-treat (and 2SLS TOT) estimates of random assignment to off-the-shelf AI access on unaided knowledge tests and essays (Eq. 1; Tables 4–6, 8). The 0.27 SD immediate and retention effects are differences in measured outcomes between AI-allowed and AI-forbidden arms; they are not derived from a structural parameter fitted to the same outcomes, nor defined in terms of those outcomes. Augmentation vs. automation is an LLM classification of ChatGPT conversation logs used only for heterogeneity (Appendix B.6; Table 8), not an input that forces the main ITT. Mechanisms (time mix, enjoyment, integrity) and belief/narrative analyses are separately measured descriptive outcomes. Self-citations (Contractor and Reyes 2026 on campus adoption/usage) supply context only and do not underwrite identification. There is no self-definitional loop, fitted-input-as-prediction, uniqueness import, or renaming of a known result as a first-principles derivation. The design is self-contained against its own experimental benchmarks.
Axiom & Free-Parameter Ledger
free parameters (3)
- Double-lasso control selection and strata fixed effects
- LLM conversation classification thresholds/prompts for augmentation vs automation
- Essay quality aggregation (human average + AI grader average)
axioms (5)
- domain assumption Random assignment to AI-allowed vs AI-forbidden identifies the causal effect of AI access (ITT) under SUTVA and no differential attrition.
- domain assumption Unaided multiple-choice tests and analytical essays measure factual/conceptual knowledge and higher-order skills relevant to learning.
- domain assumption One week without resources is a meaningful retention horizon for skill accumulation claims.
- ad hoc to paper LLM labels of ChatGPT logs into augmentation vs automation recover economically meaningful use modes.
- standard math Linear models with heteroskedasticity-robust SEs and double-lasso controls recover average treatment effects of interest.
invented entities (1)
-
Augmentation vs automation user types (student-level)
independent evidence
read the original abstract
We study how generative AI affects student learning in a randomized experiment. In proctored, in-person sessions, undergraduates learn about an unfamiliar topic and write an analytical essay with or without access to off-the-shelf generative AI, then complete unaided assessments immediately and one week later. We measure learning with knowledge tests (factual and conceptual understanding) and open-ended essays (higher-order skills). AI access raises immediate test scores by 0.27 standard deviations. These gains persist one week later. Essay quality, by contrast, changes little while students have AI access but improves in style and relevance one week later, when students write unaided. These delayed gains are larger among augmentation users-who use AI to explain concepts rather than generate text-whereas automation users' short-run quality gains vanish once AI is removed. We find evidence for two mechanisms behind the learning gains: students shift time away from drafting text and toward reading and searching for information, and they report greater learning enjoyment.
Figures
Reference graph
Works this paper leans on
-
[1]
Review of Economic Studies , year =
Andre, Peter and Haaland, Ingar and Roth, Christopher and Wiederholt, Mirko and Wohlfart, Johannes , title =. Review of Economic Studies , year =
-
[2]
Emi, Bradley and Spero, Max , title =. 2024 , institution =. 2402.14873 , archiveprefix =
Pith/arXiv arXiv 2024
-
[3]
2025 , month =
Masrour, Elyas , title =. 2025 , month =
2025
-
[4]
and Spero, Max , title =
Masrour, Elyas and Emi, Bradley N. and Spero, Max , title =. Proceedings of the 1st Workshop on Detecting AI Generated Content (GenAIDetect), COLING , pages =. 2025 , url =
2025
-
[5]
International Conference on Learning Representations (ICLR) , year =
Thai, Katherine and Emi, Bradley and Masrour, Elyas and Iyyer, Mohit , title =. International Conference on Learning Representations (ICLR) , year =
-
[6]
2025 , url =
Jabarian, Brian and Imas, Alex , title =. 2025 , url =
2025
-
[7]
Journal of Economic Perspectives , volume=
Automation and New Tasks: How Technology Displaces and Reinstates Labor , author=. Journal of Economic Perspectives , volume=
-
[8]
Science , volume=
What Can Machine Learning Do? Workforce Implications , author=. Science , volume=
-
[9]
Becker, Joel and Rush, Nate and Barnes, Beth and Rein, David , title =. 2025 , month =. 2507.09089 , archiveprefix =
Pith/arXiv arXiv 2025
-
[10]
2025 , month =
Building an. 2025 , month =
2025
-
[11]
and Hitzig, Zo\"e and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =
Chatterji, Aaron and Cunningham, Thomas and Deming, David J. and Hitzig, Zo\"e and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =. 2025 , type =
2025
-
[12]
The Generative
Str. The Generative. 2026 , url =
2026
-
[13]
VoxDevLit , volume =
Education Technology , author =. VoxDevLit , volume =
-
[14]
2025 , month = aug, howpublished =
Narayanan, Arvind , title =. 2025 , month = aug, howpublished =
2025
-
[15]
2025 , doi =
Kestin, Greg and Miller, Kelly and Klales, Anna and Milbourne, Timothy and Ponti, Gregorio , journal =. 2025 , doi =
2025
-
[16]
The Impact of Generative
Lee, Hao-Ping (Hank) and Sarkar, Advait and Tankelevitch, Lev and Drosos, Ian and Rintel, Sean and Banks, Richard and Wilson, Nicholas , booktitle =. The Impact of Generative. 2025 , publisher =
2025
-
[17]
Reading Between the Lines: Modeling User Behavior and Costs in
Mozannar, Hussein and Bansal, Gagan and Fourney, Adam and Horvitz, Eric , booktitle =. Reading Between the Lines: Modeling User Behavior and Costs in. 2024 , publisher =
2024
-
[18]
2024 , eprint =
Empirical evidence of large language model's influence on human spoken communication , author =. 2024 , eprint =
2024
-
[19]
2025 , eprint =
Ammari, Tawfiq and Chen, Meilun and Zaman, S M Mehedi and Garimella, Kiran , title =. 2025 , eprint =
2025
-
[20]
The New Yorker , year =
Hsu, Hua , title =. The New Yorker , year =
-
[21]
, title =
Walsh, James D. , title =. New York Magazine , year =
-
[22]
2025 , url =
Kunal Handa and Drew Bent and Alex Tamkin and Miles McCain and Esin Durmus and Michael Stern and Mike Schiraldi and Saffron Huang and Stuart Ritchie and Steven Syverud and Kamya Jagadish and Margaret Vo and Matt Bell and Deep Ganguli , title =. 2025 , url =
2025
-
[23]
2026 , url =
Maxim Massenkoff and Eva Lyubich and Peter McCrory and Ruth Appel and Ryan Heller , title =. 2026 , url =
2026
-
[24]
The Quarterly Journal of Economics , volume=
The (perceived) returns to education and the demand for schooling , author=. The Quarterly Journal of Economics , volume=. 2010 , publisher=
2010
-
[25]
The Review of Economic Studies , volume=
Determinants of college major choice: Identification using an information experiment , author=. The Review of Economic Studies , volume=. 2015 , publisher=
2015
-
[26]
Generative
Contractor, Zara and Reyes, Germ. Generative
-
[27]
Handbook of the Economics of Education , editor =
Technology and Education: Computers, Software, and the Internet , author =. Handbook of the Economics of Education , editor =
-
[28]
Journal of Economic Literature , volume=
Upgrading education with technology: Insights from experimental research , author=. Journal of Economic Literature , volume=. 2020 , publisher=
2020
-
[29]
Journal of Memory and Language , volume=
How many words do we read per minute? A review and meta-analysis of reading rate , author=. Journal of Memory and Language , volume=. 2019 , publisher=
2019
-
[30]
Journal of Reading , volume=
Reading rate: Theory, research, and practical implications , author=. Journal of Reading , volume=. 1992 , publisher=
1992
-
[31]
Handbook 1: Cognitive domain , author=
Taxonomy of educational objectives: The classification of educational goals. Handbook 1: Cognitive domain , author=. 1956 , publisher=
1956
-
[32]
2023 , month =
Nam, Jane , title =. 2023 , month =
2023
-
[33]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing Systems , volume =
-
[34]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =
Chiang, Cheng-Han and Lee, Hung-yi , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =. 2023 , publisher =
2023
-
[35]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , publisher =
2023
-
[36]
, title =
Carter, Susan Payne and Greenberg, Kyle and Walker, Michael S. , title =. Economics of Education Review , year =
-
[37]
The Quarterly Journal of Economics , year =
Malamud, Ofer and Pop-Eleches, Cristian , title =. The Quarterly Journal of Economics , year =
-
[38]
Technology and Child Development: Evidence from the
Cristia, Julian and Ibarrar. Technology and Child Development: Evidence from the. American Economic Journal: Applied Economics , year =
-
[39]
and Rush, Mark and Yin, Lu , title =
Figlio, David N. and Rush, Mark and Yin, Lu , title =. Journal of Labor Economics , year =
-
[40]
and Fox, Lindsay and Loeb, Susanna and Taylor, Eric S
Bettinger, Eric P. and Fox, Lindsay and Loeb, Susanna and Taylor, Eric S. , title =. American Economic Review , year =
-
[41]
and Ladd, Helen F
Vigdor, Jacob L. and Ladd, Helen F. and Martinez, Erika , title =. Economic Inquiry , year =
-
[42]
and Goodman, Sarena and Smith, Jonathan , title =
Dettling, Lisa J. and Goodman, Sarena and Smith, Jonathan , title =. The Review of Economics and Statistics , year =
-
[43]
, title =
Caldwell, Jane E. , title =. CBE---Life Sciences Education , year =
-
[44]
Education and Information Technologies , year =
Lewin, Cathy and Somekh, Bridget and Steadman, Stephen , title =. Education and Information Technologies , year =
-
[45]
Journal of the European Economic Association , volume =
Expertise , author =. Journal of the European Economic Association , volume =
-
[46]
Bastani, Hamsa and Bastani, Osbert and Sungu, Alp and Ge, Haosen and Kabakc. Generative. Proceedings of the National Academy of Sciences , volume =. doi:10.1073/pnas.2422633122 , note =
-
[47]
Belloni, Alexandre and Chernozhukov, Victor and Hansen, Christian , year = 2014, month = may, journal =. High-
2014
-
[48]
, year = 2026, journal =
Bick, Alexander and Blandin, Adam and Deming, David J. , year = 2026, journal =. The
2026
-
[49]
Generative
Brynjolfsson, Erik and Li, Danielle and Raymond, Lindsey , year = 2025, month = may, journal =. Generative
2025
-
[50]
Endoscopist
Budzy. Endoscopist. The Lancet Gastroenterology & Hepatology , volume =
-
[51]
Cui, Zheyuan (Kevin) and Demirer, Mert and Jaffe, Sonia and Musolff, Leon and Peng, Sida and Salz, Tobias , year = 2026, journal =. The
2026
-
[52]
Writing Code vs
Demirer, Mert and Musolff, Leon and Yang, Liyuan , institution =. Writing Code vs. Shipping Code: Productivity Effects Across Generations of
-
[53]
Dell'Acqua, Fabrizio and McFowland, Edward and Mollick, Ethan R. and. Navigating the. Organization Science , doi =
-
[54]
and Hauser, Oliver P
Doshi, Anil R. and Hauser, Oliver P. , journal =. Generative
-
[55]
Meincke, Lennart and Nave, Gideon and Terwiesch, Christian , journal =
-
[56]
and Kushlev, Kostadin , journal =
Moon, Kibum and Green, Adam E. and Kushlev, Kostadin , journal =. Homogenizing Effect of Large Language Models (
-
[57]
The Creative Link Between Words and Ideas Is Weakening in the
Moon, Kibum and Kushlev, Kostadin and Bank, Andrew and. The Creative Link Between Words and Ideas Is Weakening in the. doi:10.31234/osf.io/jsz58_v6 , url =
-
[58]
Minnesota Law Review , volume =
Lawyering in the Age of Artificial Intelligence , author =. Minnesota Law Review , volume =
-
[59]
Educational Evaluation and Policy Analysis , volume =
How Big Are Effect Sizes in International Education Studies? , author =. Educational Evaluation and Policy Analysis , volume =
-
[60]
Educational Researcher , volume =
Interpreting Effect Sizes of Education Interventions , author =. Educational Researcher , volume =
-
[61]
and Lavy, Victor , journal =
Angrist, Joshua D. and Lavy, Victor , journal =. Using
-
[62]
Kirabo and Mackevicius, Claire L
Jackson, C. Kirabo and Mackevicius, Claire L. , journal =. What Impacts Can We Expect from School Spending Policy?
-
[63]
The Promise of Tutoring for
Nickow, Andre and Oreopoulos, Philip and Quan, Vincent , journal =. The Promise of Tutoring for
-
[64]
and Rockoff, Jonah E
Chetty, Raj and Friedman, John N. and Rockoff, Jonah E. , journal =. Measuring the Impacts of Teachers
-
[65]
Economics of Education Review , volume =
The Economic Value of Higher Teacher Quality , author =. Economics of Education Review , volume =
-
[66]
Baird, Matthew and Carpanelli, Mar and Xu, Brian and Xu, Kevin , journal =. Firms'
-
[67]
Academy of Management Journal , volume =
When and How Artificial Intelligence Augments Employee Creativity , author =. Academy of Management Journal , volume =
-
[68]
, institution =
Dell'Acqua, Fabrizio and Ayoubi, Charles and Lifshitz, Hila and Sadun, Raffaella and Mollick, Ethan and Mollick, Lilach and Han, Yi and Goldman, Jeff and Nair, Hari and Taub, Stewart and Lakhani, Karim R. , institution =. The Cybernetic Teammate: A Field Experiment on Generative
-
[69]
Does Generative
Cruces, Guillermo and. Does Generative
-
[70]
and Sting, Fabian J
Lehmann, Matthias and Cornelius, Philipp B. and Sting, Fabian J. , year = 2025, month = mar, number =
2025
-
[71]
Experimental
Noy, Shakked and Zhang, Whitney , year = 2023, month = jul, journal =. Experimental
2023
-
[72]
Peng, Sida and Kalliamvakou, Eirini and Cihon, Peter and Demirer, Mert , year = 2023, month = feb, number =. The. 2302.06590 , doi =
Pith/arXiv arXiv 2023
-
[73]
Rav. Higher. 2025 , month = feb, journal =
2025
-
[74]
Perceptions and
St. Perceptions and. Computers and Education: Artificial Intelligence , volume =
-
[75]
Harvard undergraduate survey on generative
Hirabayashi, Shikoh and Jain, Rishab and Jurkovi. Harvard undergraduate survey on generative
-
[76]
, journal =
Goldsmith-Pinkham, Paul and Tan, Chenhao and Zentefis, Alexander K. , journal =. Human-
-
[77]
Kanazawa, Kyogo and Kawaguchi, Daiji and Shigeoka, Hitoshi and Watanabe, Yasutora , journal =
-
[78]
Psychological Methods , volume =
A General Approach to Causal Mediation Analysis , author =. Psychological Methods , volume =
-
[79]
Social Sciences & Humanities Open , volume =
Barcaui, Andr. Social Sciences & Humanities Open , volume =
-
[80]
2412.16429 , archivePrefix =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.