Pith. sign in

REVIEW 3 major objections 4 minor 12 references

ChatGPT's arrival didn't raise grades in AI-susceptible courses, a 10-year university panel finds.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:08 UTC pith:ASMK44VH

load-bearing objection A transparent, well-measured large-sample null on GenAI grade inflation, but the headline null is not causally identified and the paper internally mislabels its own bounds. the 3 major comments →

arxiv 2607.21534 v2 pith:ASMK44VH submitted 2026-07-23 cs.CY econ.GNq-fin.EC

Generative AI Availability, Grades, and Student Satisfaction at a Large University

classification cs.CY econ.GNq-fin.EC
keywords generative AIhigher educationgrade inflationdifference-in-differencescourse evaluationsChatGPTCOVID-19assessment design
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tests the 'GenAI substitution hypothesis'—the worry that students offload cognitive effort to AI and earn higher grades without learning. Using administrative and syllabus data from a large U.S. university (138,386 students, 72,730 course offerings, fall 2016–fall 2025), it compares courses that rely heavily on take-home problem sets, essays, and open exams (GenAI-susceptible) with courses that use in-class closed-book exams and live demonstrations. After separating the effects of ChatGPT's November 2022 release from the lingering effects of the COVID-19 pandemic, the paper finds no significant differential change in grades, withdrawal rates, or failure rates in more susceptible courses, and no robust change in students' self-reported understanding or interest. If correct, the findings temper fears that ChatGPT immediately inflated grades or eroded satisfaction in higher education, and suggest that prior positive results may reflect measurement or modeling choices rather than a universal effect.

Core claim

Using a difference-in-differences design with course, semester, and student fixed effects, and modeling COVID-19 as either a transient window (2020–2022) or a persistent level shift, the paper estimates that fully susceptible courses saw a change of 0.030 to 0.045 grade points (on a 4.0 scale) after ChatGPT, not statistically significant at the 5% level. Effects on withdrawal and failure rates are also null, and self-reported understanding is unchanged; interest shows a modest increase only under the transient-COVID assumption. The authors attribute positive findings in prior studies to two artifacts: anchoring susceptibility to COVID-contaminated years, and measuring susceptibility contempo

What carries the argument

The central object is a course-level 'GenAI susceptibility' measure: the share of a final grade allocated to assessments that can be fully automated by AI—take-home/open exams, homework, papers, and projects—estimated from 36,357 syllabi using a human-validated LLM pipeline (mean absolute error 0.063 on a 0–1 scale). The design anchors this measure to each course's 2019 offering (pre-COVID, pre-ChatGPT) and pairs it with a 'dual-shock' difference-in-differences specification that includes a Susceptibility-by-COVID-Window interaction, so the post-AI coefficient is identified against the pre-COVID baseline rather than the contaminated 2022 semester.

Load-bearing premise

The claim that ChatGPT had no effect rests on the assumption that, in the absence of ChatGPT, high- and low-susceptibility courses would have followed the same outcome trajectories after 2022—an assumption the paper's own event-study test rejects for the grade outcome even after excluding the COVID years.

What would settle it

A replication at another institution where the parallel-trends test for grades passes (COVID-excluded) that finds a significant positive grade effect in susceptible courses would refute the paper's conclusion. Alternatively, within this dataset, demonstrating that the grade pre-trend failure is not attributable to the COVID window (for example, by showing a placebo pre-trend violation in a pre-2016 period with no COVID disruption, or by finding the violation persists when COVID is modeled nonparametrically) would undermine the causal null.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the null is real, grades in AI-susceptible courses did not lose their signaling value in the first years after ChatGPT, easing immediate concerns about grade inflation.
  • Instructors and institutions can use in-class closed-book exams and live demonstrations as relatively AI-proof assessments without expecting large substitution-driven grade shifts.
  • The absence of detectable satisfaction declines suggests students did not systematically disengage from susceptible coursework in this setting.
  • The findings imply that substitution effects, if any, are context-dependent—varying with student body, institutional AI policies, or how assessments evolve—rather than an immediate universal consequence of AI availability.
  • The anchor choice matters: using post-COVID or contemporaneous susceptibility measures can produce spurious positive effects, so future evaluations should fix exposure to a pre-treatment baseline.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's own parallel-trends test for the grade outcome fails even after excluding COVID semesters (p = 0.031), so the null grade effect should be read as descriptive rather than strictly causal; the credibility of the null leans on auxiliary assumptions about COVID modeling that the data cannot fully test.
  • Because the susceptibility anchor is fixed at 2019, courses that responded to ChatGPT by shifting to in-class exams would be counted as 'susceptible' even if they later reduced AI exposure—potentially masking a protective instructor response.
  • If instructors adjusted grading standards to offset AI use, grade-based outcomes could stay flat while actual learning declines; the paper's satisfaction measures are self-reports, not direct learning assessments, so the null does not rule out substitution-induced learning losses.
  • A natural extension is to test the same design in settings with cleaner parallel pre-trends (e.g., institutions where 2019 and 2022 assessment structures are demonstrably stable) to see whether the null replicates without relying on the COVID-window assumption.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper uses syllabus, administrative, and course-evaluation data from a large U.S. university (2016–2025) to test whether courses whose grades depend more heavily on GenAI-susceptible assessments—take-home exams, homework, papers, projects—experienced differential changes in grades, withdrawal/failure rates, and student satisfaction after ChatGPT's release. Course susceptibility is measured from 2019 syllabi via an LLM-based annotation pipeline validated against 525 human-consensus labels (MAE = 0.063). The main design is a difference-in-differences estimator with course, semester, and student fixed effects and an explicit COVID-window interaction, reported under two COVID assumptions that are meant to bound the GenAI effect. The authors report null average grade effects, null effects by prior ability tercile, and null or fragile course-evaluation effects, interpreting the results as evidence against the GenAI substitution hypothesis in this setting.

Significance. If the headline result were causally identified, the paper would be an important counterpoint to the two prior catalog-scale studies (Hausman et al. 2025; Chirikov 2026a) and would extend the literature to student satisfaction. The paper has real strengths: a large administrative panel (1.2M student-offering observations), a carefully documented and human-validated LLM measurement pipeline, extensive robustness work across five susceptibility anchors and six ability proxies, a useful replication showing that Hausman et al.'s unanchored treatment produces a spuriously positive effect, and the first university-scale examination of course-evaluation outcomes. The authors also deserve credit for stating their identification limitations candidly. However, the central causal claim on grades is not supported by the paper's own pre-trend tests: the COVID-excluded parallel-trends test fails for the main GPA outcome, and the manuscript explicitly concedes that a strict causal difference-in-differences interpretation is not available. The contribution as it stands is a well-measured descriptive null, not a causally identified one; the framing must be revised accordingly.

major comments (3)
  1. [§5.1.1, Table 1, §6.1] The primary GPA null is not causally identified. The COVID-excluded parallel pre-trends test rejects (PT p=0.031 in Table 1), and §5.1.1 concedes that this 'precludes a strict causal interpretation of the null result'; §6.1 goes further and says the setting may be 'not suited to a causal difference-in-differences approach.' Visual inspection of Figure 1 and the ability-tercile analysis do not restore identification: Figure 1 is informal, and the middle-ability tercile also fails the COVID-excluded pre-trend test (Table 11, p=0.026). The abstract's 'no significant differential effect of GenAI availability on grades' and the Introduction's causal phrasing therefore overstate what the design can support. The abstract and conclusions should be reframed to present a descriptive null, with the identification failure stated as a headline caveat.
  2. [§4.3, Table 1, §5.1.1] The labeling of the 'GenAI net COVID' row as the 'conservative' estimate is inconsistent with the paper's own bound rule. §4.3 states that when β1 and β2 have opposite signs, the labels reverse. For GPA, β1=0.030 and β2=−0.017, so the net-COVID linear combination equals 0.045—the upper endpoint of the stated [0.030, 0.045] range, not the lower bound. Under the paper's rule, the conservative estimate for GPA is the post-AI-vs-pre-COVID coefficient 0.030. Table 1 and §5.1.1 instead describe 0.045 as the 'persistent COVID' reading and treat it as conservative. This mislabel propagates to the course-evaluation tables (e.g., Table 12) and should be corrected, with the bound ordering stated consistently for every outcome.
  3. [§4.1.1, §4.3, Appendix E] The susceptibility anchor and COVID-window choices appear to be the product of a data-dependent specification search. Appendix E shows that COVID-laden anchors produce positive, significant grade estimates (0.115, p<0.05 for 2022; 0.160, p<0.01 for 2020–2022), while clean pre-COVID anchors produce nulls, and the text uses this contrast to justify the 2019 anchor. Likewise, the 2020–2022 COVID window is defended after inspecting event-study dynamics ('the pre-period coefficients in fact settle only once the COVID semesters are set aside'). This does not by itself invalidate the analysis, but it creates a multiple-comparisons risk: the central null may be an artifact of which anchor and window are chosen. The paper should either justify the 2019 anchor on a priori grounds before reporting the alternatives, or present all five anchors with equal emphasis and explicitly discuss the specifica
minor comments (4)
  1. [Abstract vs. §3.4 vs. §7] Sample sizes are inconsistent: the abstract and conclusion report 138,386 students and 72,730 offerings, while §3.4 reports 122,663 students and 38,754 offerings in the balanced analytic sample. Please reconcile or clarify which sample each number refers to.
  2. [Throughout] Typos and wording slips: 'satisfcation' in §1; 'offereings' in the Figure 3 caption; 'reoffering' and 'GenAI-susceptible courses--' in the abstract. A careful proofread is needed.
  3. [Table 1] The 'Susceptibility' main-effects rows report enormous standard errors (e.g., 10,411.457) under fixed effects because these terms are collinear with course fixed effects. Consider suppressing these rows or adding a note that they are not identified in the TWFE columns.
  4. [§3.3 and Table 12] The text says course evaluations are aggregated at the offering-instructor level, but Table 12 reports N≈938,000, which resembles student-level or enrollment-weighted rows. Clarify whether evaluations are expanded to the student level or enrollment-weighted, and state this in the table notes.

Circularity Check

0 steps flagged

No significant circularity: the analysis is an observational DiD estimate with transparent reparameterizations; the admitted pre-trend failure is a validity limitation, not a circular step.

full rationale

The paper makes no claim to derive a prediction from first principles; it reports reduced-form difference-in-differences regressions of administrative outcomes on a syllabus-based susceptibility measure. The treatment measure is constructed independently of the outcomes via an LLM pipeline validated against 525 human-consensus labels, so the treatment is not defined in terms of the outcome. The 'GenAI net COVID' row is explicitly a linear combination of the estimated coefficients: the paper states it 'reports the linear combination β̂×Post-AI − β̂×COVIDYear with delta-method standard error' and notes the level-shift specification is 'an exact reparameterization of Equation (1)'. This is transparent algebraic reporting, not a fitted parameter renamed as a prediction. The one self-citation (Gu et al. 2026, which includes a coauthor) is used only as motivational qualitative evidence about instructor perceptions and help-seeking; it is not load-bearing for the grade or satisfaction estimates. The paper openly discloses that the formal parallel-trends test fails for the grade outcome and that a strict causal interpretation is precluded, and Section 6.1 concedes the setting may not be suited to a causal DiD approach. Those are identification limitations, not circularity: they do not make the estimates equal to their inputs by construction. Model-selection choices such as the 2019 susceptibility anchor and the COVID-window specification are data-driven assumptions, but choosing among models after inspecting event studies is not the same as defining the result in terms of itself. No circular step meeting the quoted-evidence standard is present.

Axiom & Free-Parameter Ledger

4 free parameters · 7 axioms · 0 invented entities

The central null claim rests on seven domain assumptions and several hand-set modeling parameters (2019 anchor, 2020–2022 COVID window, 0.5 project split, unknown exams treated as non-susceptible). No free constants are fitted to the outcome in the sense of a parametric model; the concern is data-dependent specification choice, scored under soundness/red flags.

free parameters (4)
  • 2019 susceptibility anchor year = 2019
    Treatment susceptibility anchored to 2019 offering; chosen because COVID-laden anchors yield positive grade effects (Appendix E). Data-dependent choice affects null.
  • COVID disruption window = 2020–2022
    Dual-shock model assumes COVID differential operates only in 2020–2022 (transient) or persists as level shift; window set by inspection of event-study dynamics (Section 4.3).
  • Project-with-presentation split weight = 0.5
    Projects with a presentation component and unspecified weights split equally between susceptible and non-susceptible categories (Section 4.1.1).
  • Unknown exam susceptibility weight = 0
    Exams whose open/closed status is unknown are treated as non-susceptible, 'the traditional default' (Section 4.1.1); if wrong, susceptibility is mismeasured.
axioms (7)
  • domain assumption The November 2022 public release of ChatGPT is an exogenous, widespread shock to GenAI availability for students.
    Used to define PostAI_t from 2023 onward (Section 4.3); if students/instructors anticipated or adopted other GenAI earlier, treatment timing is mis-specified.
  • domain assumption High- and low-susceptibility courses would have had parallel outcome trends after 2022 had ChatGPT not been released.
    Core DiD identifying assumption; formal test fails for average grade even COVID-excluded (Table 1, p=0.031; Section 5.1.1).
  • domain assumption Courses' GenAI susceptibility is stable and the 2019 offering represents pre-AI course design.
    Susceptibility anchored to 2019 (Section 4.1.1); mean correlation with 2019 is 0.73 but COVID years dip to 0.42-0.68 (Appendix D.1).
  • domain assumption COVID-19 disruption effects on the high-low susceptibility gap are either fully transient (2020-2022) or fully persistent, bracketing the true GenAI effect.
    Dual-shock model in Equation 1 and its level-shift reparameterization; the choice of window is not data-determined (Section 4.3).
  • domain assumption Median course-evaluation responses measure student satisfaction (understanding, interest, workload).
    Outcome proxies defined in Section 4.1.2 and Appendix B; response rates fluctuate and selection into responding may change over time (Section 6.1).
  • domain assumption Final grades capture assessment performance plus instructor grading adjustments; no instructor policy response fully absorbs GenAI effects.
    Acknowledged in Section 6.1: instructors may adjust grading standards.
  • domain assumption The LLM-based susceptibility measure is an adequate reconstruction of true assessment weights, with measurement error not systematically correlated with the outcome.
    Validation MAE=0.063 and category MAEs up to 0.085; the authors note measurement error increases noise and likelihood of nulls (Section 6.1, Appendix C).

pith-pipeline@v1.3.0-alltime-deepseek · 36753 in / 14342 out tokens · 133685 ms · 2026-08-01T07:08:06.227282+00:00 · methodology

0 comments
read the original abstract

The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this "GenAI substitution hypothesis" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as self-reported understanding and interest in the subject, which prior research links to assessments. We test the substitution hypothesis using syllabus and administrative data from a large U.S. university (2016-2025; 138,386 students; 72,730 course offerings). We measure courses' GenAI susceptibility using a human-validated LLM pipeline to extract assessment types from syllabi, and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient. We find no significant differential effect of GenAI availability on grades overall or among previously lower-performing students. Effects on self-reported understanding are likewise insignificant; effects on interest are significant only assuming transient pandemic effects. Our findings temper concerns that GenAI inflates grades and reduces students' satisfaction.

Figures

Figures reproduced from arXiv: 2607.21534 by George Chaney III, Henry Gold, Ivan Bar, James M. Zumel Dumlao, Junyao Hu, Meng Wang, Misha Teplitskiy, Zhonghan Xie.

Figure 1
Figure 1. Figure 1: Event Study: Average Grade. Left box: semester-by-semester interaction coefficients [PITH_FULL_IMAGE:figures/full_fig_p028_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Event Study: Average Grade by AbilityRank Tercile. Left boxes: semester-by [PITH_FULL_IMAGE:figures/full_fig_p031_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of GenAI Susceptibility among courses in 2019, 2021, and 2023. Each [PITH_FULL_IMAGE:figures/full_fig_p064_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Susceptibility Tercile Trajectories, 2019 to 2022 to 2025. Flows trace courses among [PITH_FULL_IMAGE:figures/full_fig_p066_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Event Study: Grade-Distribution Thresholds. Left boxes: semester-by-semester [PITH_FULL_IMAGE:figures/full_fig_p076_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Event Study: Withdrawal and Failure Rates. Left boxes: semester-by-semester [PITH_FULL_IMAGE:figures/full_fig_p077_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Event Study: Course Evaluation Outcomes. Left boxes: semester-by-semester [PITH_FULL_IMAGE:figures/full_fig_p078_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Raw Dynamics: Student Outcomes by Susceptibility Group. Each point is the [PITH_FULL_IMAGE:figures/full_fig_p080_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Raw Dynamics: Average Grade by Susceptibility Group and AbilityRank Tercile. [PITH_FULL_IMAGE:figures/full_fig_p081_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Raw Dynamics: Course Evaluation Outcomes by Susceptibility Group. Each [PITH_FULL_IMAGE:figures/full_fig_p082_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

12 extracted references · 1 linked inside Pith

  1. [1]

    Acemoglu, D. (2025). The simple macroeconomics of AI.Economic Policy, 40:13–58. Acemoglu, D. and Restrepo, P. (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor.Journal of Economic Perspectives, 33(2):3–30. Ammari, T., Chen, M., Zaman, S. M. M., and Garimella, K. (2025). How Students (Really) Use ChatGPT: Uncovering Experience...

  2. [2]

    This course advanced my understanding of the subject matter

    Durning, S. J., Dong, T., Ratcliffe, T., Schuwirth, L., Artino, A. R., Boulet, J. R., and Eva, K. (2016). Comparing Open-Book and Closed-Book Examinations: A Systematic Review. Academic Medicine, 91(4):583–599. Elliot, A. J. (1999). Approach and avoidance motivation and achievement goals.Educational Psychologist, 34(3):169–189. Ellis, L. (2025). AI Is Tea...

  3. [8]

    The courses that do move tend to drift toward higher susceptibility reflecting the 64 Table 7: Within-Course Persistence of Susceptible Assessment Weighting: Year-to-Year Correlations ’15 ’16 ’17 ’18 ’19 ’20 ’21 ’22 ’23 ’24 ’25 ’15 1.00 ’16 0.84 1.00 ’17 0.82 0.85 1.00 ’18 0.78 0.83 0.87 1.00 ’19 0.80 0.79 0.83 0.89 1.00 ’20 0.50 0.55 0.50 0.57 0.61 1.00 ...

  4. [9]

    The sample is the 1,203 courses offered in 2025 with observed terciles in all three years

    Flows trace courses among the Low (below 0.33), Medium (0.33 to 0.67), and High (above 0.67) susceptibility terciles, based on each year’s contemporaneous measure. The sample is the 1,203 courses offered in 2025 with observed terciles in all three years. Node heights are proportional to course counts. E Robustness: Treatment-Measure Comparisons E.1 Altern...

  5. [10]

    GenAI net COVID

    using a different measure of a course’s GenAI suscepti- bility: its 2019 offerings (col. 1, the main-text measure), the pre-COVID period 2015–2019 (col. 2), course-level averages over the full pre-GenAI period 2015–2022 (col. 3), its 2022 offerings (col. 4), and the COVID-affected period 2020–2022 (col. 5). Each column is fit on its own balanced sample (c...

  6. [11]

    GenAI net COVID

    within one tercile ofAbilityRank, the course-residualized within-cohort first-term GPA rank (Sec- tion 4.1.3). The heterogeneity sample is restricted to un- dergraduate, fall-entry students from cohorts entering be- fore the public release of ChatGPT, for whom the proxy is measured pre-treatment. The “GenAI net COVID” row re- ports the conservative, COVID...

  7. [12]

    GenAI net COVID

    Right boxes: averages from COVID-affected terms (shaded band), post-GenAI terms, and the difference between the two with delta-method standard errors. Estimates are the unit difference between fully susceptible and fully non-susceptible courses. Regressions include course and semester fixed effects. Standard errors are clustered on course. Course-evaluati...

  8. [40]

    take_home

    KEY: ‘Performance’ here = demonstrating a physical skill (e.g., CPR, first-aid) in front of instructor. Kinesiology skill demonstrations =performance. This is NOT a test score or course grade. •Presentation 1 + 2 = 50 pts→presentation: 50 •Paper 50 pts→papers: 50 •Final Exam 100 pts (take-home)→open_exam: 100 (take-home→open_exam) •Total: 160 + 40 + 40 + ...

  9. [210]

    C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K

    Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. (2026). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.Organization Science, 37(2):403–423. Demirci, O., Hann...

  10. [2019]

    The two groups differ substantially at baseline

    means for courses in the bottom and top absolute terciles of the 2019-anchored susceptibility measure. The two groups differ substantially at baseline. High susceptibility courses award grades 0.50 points higher than low susceptibility courses, equivalent to half a letter grade of difference. Their withdrawal and failure rates are roughly half as large. E...

  11. [2023]

    Dashed lines mark the 0.33 and 0.67 cutoffs defining the Low, Medium, and High absolute susceptibility terciles

    Each panel plots the contemporaneous susceptibility measure, the share of the final grade allocated to susceptible assessments, across courses in the balanced analytic sample with a syllabus observed in the indicated year (averaged over offereings in that year). Dashed lines mark the 0.33 and 0.67 cutoffs defining the Low, Medium, and High absolute suscep...

  12. [2025]

    Of these, roughly two-thirds remain in the same tercile at all three points, and nearly three-fourths occupy the same tercile in 2025 as in