REVIEW 3 major objections 4 minor 12 references
ChatGPT's arrival didn't raise grades in AI-susceptible courses, a 10-year university panel finds.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:08 UTC pith:ASMK44VH
load-bearing objection A transparent, well-measured large-sample null on GenAI grade inflation, but the headline null is not causally identified and the paper internally mislabels its own bounds. the 3 major comments →
Generative AI Availability, Grades, and Student Satisfaction at a Large University
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Using a difference-in-differences design with course, semester, and student fixed effects, and modeling COVID-19 as either a transient window (2020–2022) or a persistent level shift, the paper estimates that fully susceptible courses saw a change of 0.030 to 0.045 grade points (on a 4.0 scale) after ChatGPT, not statistically significant at the 5% level. Effects on withdrawal and failure rates are also null, and self-reported understanding is unchanged; interest shows a modest increase only under the transient-COVID assumption. The authors attribute positive findings in prior studies to two artifacts: anchoring susceptibility to COVID-contaminated years, and measuring susceptibility contempo
What carries the argument
The central object is a course-level 'GenAI susceptibility' measure: the share of a final grade allocated to assessments that can be fully automated by AI—take-home/open exams, homework, papers, and projects—estimated from 36,357 syllabi using a human-validated LLM pipeline (mean absolute error 0.063 on a 0–1 scale). The design anchors this measure to each course's 2019 offering (pre-COVID, pre-ChatGPT) and pairs it with a 'dual-shock' difference-in-differences specification that includes a Susceptibility-by-COVID-Window interaction, so the post-AI coefficient is identified against the pre-COVID baseline rather than the contaminated 2022 semester.
Load-bearing premise
The claim that ChatGPT had no effect rests on the assumption that, in the absence of ChatGPT, high- and low-susceptibility courses would have followed the same outcome trajectories after 2022—an assumption the paper's own event-study test rejects for the grade outcome even after excluding the COVID years.
What would settle it
A replication at another institution where the parallel-trends test for grades passes (COVID-excluded) that finds a significant positive grade effect in susceptible courses would refute the paper's conclusion. Alternatively, within this dataset, demonstrating that the grade pre-trend failure is not attributable to the COVID window (for example, by showing a placebo pre-trend violation in a pre-2016 period with no COVID disruption, or by finding the violation persists when COVID is modeled nonparametrically) would undermine the causal null.
If this is right
- If the null is real, grades in AI-susceptible courses did not lose their signaling value in the first years after ChatGPT, easing immediate concerns about grade inflation.
- Instructors and institutions can use in-class closed-book exams and live demonstrations as relatively AI-proof assessments without expecting large substitution-driven grade shifts.
- The absence of detectable satisfaction declines suggests students did not systematically disengage from susceptible coursework in this setting.
- The findings imply that substitution effects, if any, are context-dependent—varying with student body, institutional AI policies, or how assessments evolve—rather than an immediate universal consequence of AI availability.
- The anchor choice matters: using post-COVID or contemporaneous susceptibility measures can produce spurious positive effects, so future evaluations should fix exposure to a pre-treatment baseline.
Where Pith is reading between the lines
- The paper's own parallel-trends test for the grade outcome fails even after excluding COVID semesters (p = 0.031), so the null grade effect should be read as descriptive rather than strictly causal; the credibility of the null leans on auxiliary assumptions about COVID modeling that the data cannot fully test.
- Because the susceptibility anchor is fixed at 2019, courses that responded to ChatGPT by shifting to in-class exams would be counted as 'susceptible' even if they later reduced AI exposure—potentially masking a protective instructor response.
- If instructors adjusted grading standards to offset AI use, grade-based outcomes could stay flat while actual learning declines; the paper's satisfaction measures are self-reports, not direct learning assessments, so the null does not rule out substitution-induced learning losses.
- A natural extension is to test the same design in settings with cleaner parallel pre-trends (e.g., institutions where 2019 and 2022 assessment structures are demonstrably stable) to see whether the null replicates without relying on the COVID-window assumption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses syllabus, administrative, and course-evaluation data from a large U.S. university (2016–2025) to test whether courses whose grades depend more heavily on GenAI-susceptible assessments—take-home exams, homework, papers, projects—experienced differential changes in grades, withdrawal/failure rates, and student satisfaction after ChatGPT's release. Course susceptibility is measured from 2019 syllabi via an LLM-based annotation pipeline validated against 525 human-consensus labels (MAE = 0.063). The main design is a difference-in-differences estimator with course, semester, and student fixed effects and an explicit COVID-window interaction, reported under two COVID assumptions that are meant to bound the GenAI effect. The authors report null average grade effects, null effects by prior ability tercile, and null or fragile course-evaluation effects, interpreting the results as evidence against the GenAI substitution hypothesis in this setting.
Significance. If the headline result were causally identified, the paper would be an important counterpoint to the two prior catalog-scale studies (Hausman et al. 2025; Chirikov 2026a) and would extend the literature to student satisfaction. The paper has real strengths: a large administrative panel (1.2M student-offering observations), a carefully documented and human-validated LLM measurement pipeline, extensive robustness work across five susceptibility anchors and six ability proxies, a useful replication showing that Hausman et al.'s unanchored treatment produces a spuriously positive effect, and the first university-scale examination of course-evaluation outcomes. The authors also deserve credit for stating their identification limitations candidly. However, the central causal claim on grades is not supported by the paper's own pre-trend tests: the COVID-excluded parallel-trends test fails for the main GPA outcome, and the manuscript explicitly concedes that a strict causal difference-in-differences interpretation is not available. The contribution as it stands is a well-measured descriptive null, not a causally identified one; the framing must be revised accordingly.
major comments (3)
- [§5.1.1, Table 1, §6.1] The primary GPA null is not causally identified. The COVID-excluded parallel pre-trends test rejects (PT p=0.031 in Table 1), and §5.1.1 concedes that this 'precludes a strict causal interpretation of the null result'; §6.1 goes further and says the setting may be 'not suited to a causal difference-in-differences approach.' Visual inspection of Figure 1 and the ability-tercile analysis do not restore identification: Figure 1 is informal, and the middle-ability tercile also fails the COVID-excluded pre-trend test (Table 11, p=0.026). The abstract's 'no significant differential effect of GenAI availability on grades' and the Introduction's causal phrasing therefore overstate what the design can support. The abstract and conclusions should be reframed to present a descriptive null, with the identification failure stated as a headline caveat.
- [§4.3, Table 1, §5.1.1] The labeling of the 'GenAI net COVID' row as the 'conservative' estimate is inconsistent with the paper's own bound rule. §4.3 states that when β1 and β2 have opposite signs, the labels reverse. For GPA, β1=0.030 and β2=−0.017, so the net-COVID linear combination equals 0.045—the upper endpoint of the stated [0.030, 0.045] range, not the lower bound. Under the paper's rule, the conservative estimate for GPA is the post-AI-vs-pre-COVID coefficient 0.030. Table 1 and §5.1.1 instead describe 0.045 as the 'persistent COVID' reading and treat it as conservative. This mislabel propagates to the course-evaluation tables (e.g., Table 12) and should be corrected, with the bound ordering stated consistently for every outcome.
- [§4.1.1, §4.3, Appendix E] The susceptibility anchor and COVID-window choices appear to be the product of a data-dependent specification search. Appendix E shows that COVID-laden anchors produce positive, significant grade estimates (0.115, p<0.05 for 2022; 0.160, p<0.01 for 2020–2022), while clean pre-COVID anchors produce nulls, and the text uses this contrast to justify the 2019 anchor. Likewise, the 2020–2022 COVID window is defended after inspecting event-study dynamics ('the pre-period coefficients in fact settle only once the COVID semesters are set aside'). This does not by itself invalidate the analysis, but it creates a multiple-comparisons risk: the central null may be an artifact of which anchor and window are chosen. The paper should either justify the 2019 anchor on a priori grounds before reporting the alternatives, or present all five anchors with equal emphasis and explicitly discuss the specifica
minor comments (4)
- [Abstract vs. §3.4 vs. §7] Sample sizes are inconsistent: the abstract and conclusion report 138,386 students and 72,730 offerings, while §3.4 reports 122,663 students and 38,754 offerings in the balanced analytic sample. Please reconcile or clarify which sample each number refers to.
- [Throughout] Typos and wording slips: 'satisfcation' in §1; 'offereings' in the Figure 3 caption; 'reoffering' and 'GenAI-susceptible courses--' in the abstract. A careful proofread is needed.
- [Table 1] The 'Susceptibility' main-effects rows report enormous standard errors (e.g., 10,411.457) under fixed effects because these terms are collinear with course fixed effects. Consider suppressing these rows or adding a note that they are not identified in the TWFE columns.
- [§3.3 and Table 12] The text says course evaluations are aggregated at the offering-instructor level, but Table 12 reports N≈938,000, which resembles student-level or enrollment-weighted rows. Clarify whether evaluations are expanded to the student level or enrollment-weighted, and state this in the table notes.
Circularity Check
No significant circularity: the analysis is an observational DiD estimate with transparent reparameterizations; the admitted pre-trend failure is a validity limitation, not a circular step.
full rationale
The paper makes no claim to derive a prediction from first principles; it reports reduced-form difference-in-differences regressions of administrative outcomes on a syllabus-based susceptibility measure. The treatment measure is constructed independently of the outcomes via an LLM pipeline validated against 525 human-consensus labels, so the treatment is not defined in terms of the outcome. The 'GenAI net COVID' row is explicitly a linear combination of the estimated coefficients: the paper states it 'reports the linear combination β̂×Post-AI − β̂×COVIDYear with delta-method standard error' and notes the level-shift specification is 'an exact reparameterization of Equation (1)'. This is transparent algebraic reporting, not a fitted parameter renamed as a prediction. The one self-citation (Gu et al. 2026, which includes a coauthor) is used only as motivational qualitative evidence about instructor perceptions and help-seeking; it is not load-bearing for the grade or satisfaction estimates. The paper openly discloses that the formal parallel-trends test fails for the grade outcome and that a strict causal interpretation is precluded, and Section 6.1 concedes the setting may not be suited to a causal DiD approach. Those are identification limitations, not circularity: they do not make the estimates equal to their inputs by construction. Model-selection choices such as the 2019 susceptibility anchor and the COVID-window specification are data-driven assumptions, but choosing among models after inspecting event studies is not the same as defining the result in terms of itself. No circular step meeting the quoted-evidence standard is present.
Axiom & Free-Parameter Ledger
free parameters (4)
- 2019 susceptibility anchor year =
2019
- COVID disruption window =
2020–2022
- Project-with-presentation split weight =
0.5
- Unknown exam susceptibility weight =
0
axioms (7)
- domain assumption The November 2022 public release of ChatGPT is an exogenous, widespread shock to GenAI availability for students.
- domain assumption High- and low-susceptibility courses would have had parallel outcome trends after 2022 had ChatGPT not been released.
- domain assumption Courses' GenAI susceptibility is stable and the 2019 offering represents pre-AI course design.
- domain assumption COVID-19 disruption effects on the high-low susceptibility gap are either fully transient (2020-2022) or fully persistent, bracketing the true GenAI effect.
- domain assumption Median course-evaluation responses measure student satisfaction (understanding, interest, workload).
- domain assumption Final grades capture assessment performance plus instructor grading adjustments; no instructor policy response fully absorbs GenAI effects.
- domain assumption The LLM-based susceptibility measure is an adequate reconstruction of true assessment weights, with measurement error not systematically correlated with the outcome.
read the original abstract
The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as self-reported understanding and interest in the subject, which prior research links to assessments. We test the substitution hypothesis using syllabus and administrative data from a large U.S. university (2016-2025; 138,386 students; 72,730 course offerings). We measure courses' GenAI susceptibility using a human-validated LLM pipeline to extract assessment types from syllabi, and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient. We find no significant differential effect of GenAI availability on grades overall or among previously lower-performing students. Effects on self-reported understanding are likewise insignificant; effects on interest are significant only assuming transient pandemic effects. Our findings temper concerns that GenAI inflates grades and reduces students' satisfaction.
Figures
Reference graph
Works this paper leans on
-
[1]
Acemoglu, D. (2025). The simple macroeconomics of AI.Economic Policy, 40:13–58. Acemoglu, D. and Restrepo, P. (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor.Journal of Economic Perspectives, 33(2):3–30. Ammari, T., Chen, M., Zaman, S. M. M., and Garimella, K. (2025). How Students (Really) Use ChatGPT: Uncovering Experience...
arXiv 2025
-
[2]
This course advanced my understanding of the subject matter
Durning, S. J., Dong, T., Ratcliffe, T., Schuwirth, L., Artino, A. R., Boulet, J. R., and Eva, K. (2016). Comparing Open-Book and Closed-Book Examinations: A Systematic Review. Academic Medicine, 91(4):583–599. Elliot, A. J. (1999). Approach and avoidance motivation and achievement goals.Educational Psychologist, 34(3):169–189. Ellis, L. (2025). AI Is Tea...
Pith/arXiv arXiv 2016
-
[8]
The courses that do move tend to drift toward higher susceptibility reflecting the 64 Table 7: Within-Course Persistence of Susceptible Assessment Weighting: Year-to-Year Correlations ’15 ’16 ’17 ’18 ’19 ’20 ’21 ’22 ’23 ’24 ’25 ’15 1.00 ’16 0.84 1.00 ’17 0.82 0.85 1.00 ’18 0.78 0.83 0.87 1.00 ’19 0.80 0.79 0.83 0.89 1.00 ’20 0.50 0.55 0.50 0.57 0.61 1.00 ...
2019
-
[9]
The sample is the 1,203 courses offered in 2025 with observed terciles in all three years
Flows trace courses among the Low (below 0.33), Medium (0.33 to 0.67), and High (above 0.67) susceptibility terciles, based on each year’s contemporaneous measure. The sample is the 1,203 courses offered in 2025 with observed terciles in all three years. Node heights are proportional to course counts. E Robustness: Treatment-Measure Comparisons E.1 Altern...
2025
-
[10]
GenAI net COVID
using a different measure of a course’s GenAI suscepti- bility: its 2019 offerings (col. 1, the main-text measure), the pre-COVID period 2015–2019 (col. 2), course-level averages over the full pre-GenAI period 2015–2022 (col. 3), its 2022 offerings (col. 4), and the COVID-affected period 2020–2022 (col. 5). Each column is fit on its own balanced sample (c...
2019
-
[11]
GenAI net COVID
within one tercile ofAbilityRank, the course-residualized within-cohort first-term GPA rank (Sec- tion 4.1.3). The heterogeneity sample is restricted to un- dergraduate, fall-entry students from cohorts entering be- fore the public release of ChatGPT, for whom the proxy is measured pre-treatment. The “GenAI net COVID” row re- ports the conservative, COVID...
2019
-
[12]
GenAI net COVID
Right boxes: averages from COVID-affected terms (shaded band), post-GenAI terms, and the difference between the two with delta-method standard errors. Estimates are the unit difference between fully susceptible and fully non-susceptible courses. Regressions include course and semester fixed effects. Standard errors are clustered on course. Course-evaluati...
2016
-
[40]
take_home
KEY: ‘Performance’ here = demonstrating a physical skill (e.g., CPR, first-aid) in front of instructor. Kinesiology skill demonstrations =performance. This is NOT a test score or course grade. •Presentation 1 + 2 = 50 pts→presentation: 50 •Paper 50 pts→papers: 50 •Final Exam 100 pts (take-home)→open_exam: 100 (take-home→open_exam) •Total: 160 + 40 + 40 + ...
2019
-
[210]
C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K
Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. (2026). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.Organization Science, 37(2):403–423. Demirci, O., Hann...
2026
-
[2019]
The two groups differ substantially at baseline
means for courses in the bottom and top absolute terciles of the 2019-anchored susceptibility measure. The two groups differ substantially at baseline. High susceptibility courses award grades 0.50 points higher than low susceptibility courses, equivalent to half a letter grade of difference. Their withdrawal and failure rates are roughly half as large. E...
2019
-
[2023]
Dashed lines mark the 0.33 and 0.67 cutoffs defining the Low, Medium, and High absolute susceptibility terciles
Each panel plots the contemporaneous susceptibility measure, the share of the final grade allocated to susceptible assessments, across courses in the balanced analytic sample with a syllabus observed in the indicated year (averaged over offereings in that year). Dashed lines mark the 0.33 and 0.67 cutoffs defining the Low, Medium, and High absolute suscep...
2019
-
[2025]
Of these, roughly two-thirds remain in the same tercile at all three points, and nearly three-fourths occupy the same tercile in 2025 as in
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.