REVIEW 4 major objections 5 minor 9 references
The paper claims that when students choose how much to learn, AI's sharp performance cutoff produces a discontinuous gap in human ability at the AI frontier.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A theoretical model shows that AI's sharp capability cutoff and hallucination rate create a discontinuous gap in student ability around the AI frontier, and that AI-free assignments can correct students' overestimation of AI accuracy.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Nice idea, but Proposition 1 overreaches as stated; the paper needs boundary conditions and proof fixes before the threshold result can stand. the 4 major comments →
Artificial or Human Intelligence?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Proposition 1 is the paper's central claim: under the model, there exists a unique threshold type T such that every student with type below T chooses ability as(T) < d and uses AI as an independent solver, while every type above T chooses ability ah(T) > d and uses AI only as a helper. At T the student is indifferent, and the induced ability function A(t) jumps at T, skipping the entire interval from as(T) to ah(T), which contains d. The reason is a discontinuity in marginal benefit: a student just below the frontier gains only 1−p from each additional unit of ability, because AI already solves most problems in that range, while a student just above the frontier gains a full 1, since those m
What carries the argument
The machinery is a binary AI frontier (d,p): AI solves problems of difficulty at most d with probability p and none above d, while a human with ability a solves problems up to a with certainty. Students maximize the mass of solvable problems minus a learning cost c(a,t) that has decreasing differences in ability and type. The two key identities are the payoff slopes—marginal benefit 1−p for students below d and 1 for students above d—and the threshold type T defined by indifference between the solver value function and the helper value function. The slope discontinuity is what produces the gap in the induced ability mapping A(t).
Load-bearing premise
The load-bearing assumption is that AI capability is a sharp cliff—it succeeds on every problem up to a fixed difficulty with one probability and on no problem above it—and that students evaluate the AI only by whether it gives the correct final answer; if performance improves gradually or partial credit matters, the marginal benefit of learning is continuous and the predicted ability gap disappears.
What would settle it
Give one group of students unrestricted access to an LLM in a course with a graded exam, then compare the distribution of exam scores around the difficulty level at which that LLM's accuracy drops; if the density of scores is continuous through that level, rather than showing a hollow just below it, the Proposition 1 discontinuity is falsified. A cleaner version: measure students' chosen study time as a function of their distance from the AI's difficulty frontier; the model predicts a strictly positive jump in study time as students cross the frontier.
If this is right
- For an exogenous fixed distribution of abilities, introducing AI helps the least-skilled students most; with endogenous investment, the distribution of ability is no longer continuous and has a gap around the AI frontier.
- Raising AI accuracy p lowers the marginal return to learning for solver-side students and raises the threshold T, so more students use AI as a solver and the ability gap can widen.
- Raising the difficulty threshold d does not change the marginal benefit of learning for anyone but shifts T upward, since the solver option becomes relatively more attractive at the old cutoff.
- If AI reduces learning costs more at higher ability levels, all students choose more ability and T falls; absent such cost benefits, AI advances mainly replace rather than augment human knowledge.
- Putting a share λ = p/p′ of assignments outside AI access aligns misspecified students' incentives with the true technology; the larger the overestimate of AI accuracy, the larger the no-AI weight needed.
Where Pith is reading between the lines
- The discontinuity prediction is a directly testable implication: in courses with universal AI access, ability or exam-score distributions should show a hollow just below the difficulty level where the AI starts failing, and students who self-identify as 'solver' users should cluster on the low side of that hollow.
- Because the comparative statics show the threshold rises with p and d, a calibrated version of the model could predict which course levels—introductory versus advanced—see the largest drops in investment as AI models improve.
- The λ = p/p′ rule suggests a practical calibration procedure: estimate the AI's true accuracy p and students' believed accuracy p′ on a class's problem set, then set the closed-book share accordingly; this follows from Section 6 but is not stated as a measurement protocol in the paper.
- If AI performance turns out to be continuous rather than sharply emergent, the model's main discontinuity would smooth into a conventional continuous trade-off, and the education-policy conclusions would need to be re-derived—the paper itself flags the emergence-mirage evidence as the main challenge to its cutoff assumption.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper builds a model in which students have type-dependent costs of acquiring ability a∈[0,1], and AI can solve all problems of difficulty ≤d with probability p and no problems above d. Students choose ability to maximize the mass of problems they can solve with AI assistance, net of cost. The model separates students into those who use AI as a 'solver' (ability below d) and those who use it as a 'helper' (ability above d). The central claims are: (i) there is a unique threshold type T separating the two groups, with a discontinuous upward jump in ability at T; (ii) advances in p or d shift this threshold and affect ability choices; (iii) if students overestimate AI accuracy, instructors can restore efficient investment by placing weight λ=p/p′ on AI-permitted assignments. The paper motivates these results with a review of hallucination and emergence in LLMs and with recent empirical education studies.
Significance. The topic is timely and the solver/helper distinction is a useful organizing device for thinking about AI in education. The paper is clearly written and connects to a rich empirical literature. Section 6's λ=p/p′ alignment result is elegant and correct. If the stated theorem and comparative statics are repaired, the model would provide a sharp, policy-relevant prediction about a possible knowledge gap around the AI frontier. However, the current version contains a false universal claim in Proposition 1 and a sign error in the main comparative-statics conditions, so the contribution is not yet established as stated.
major comments (4)
- [§4, Proposition 1] As stated, Proposition 1 is false: decreasing differences alone does not imply the existence of an interior threshold T∈(0,1). The proof's inference 'if A(t)=a_s(t) then A(t')=a_s(t') for all t'>t' uses A(t')≥A(t)=a_s(t)≥d, but a_s(t)≤d by definition, so the conclusion does not follow. More importantly, the asserted interior threshold requires boundary conditions that are not stated. Counterexample: take c(a,t)=a^2/[2(t+0.01)], p=0.2, d=0.95, which satisfies all listed assumptions. For every t∈[0,1] the solver branch strictly dominates the helper branch (at t=1, U_s≈0.513 and U_h≈0.505), so no type uses AI as a helper and T∉(0,1). The paper needs explicit conditions ensuring that both branches are chosen—e.g., U_s(0)>U_h(0), U_s(1)<U_h(1), and single crossing—before the headline discontinuity can be claimed.
- [§5.2] The conditions for the threshold to increase in p and d have the inequality reversed. Since Δ(t)=U_s(t)−U_h(t) is decreasing in t at the threshold, T increases iff ∂Δ/∂p>0 or ∂Δ/∂d>0 at T. For p, this gives d−a_s+c_p(a_h)−c_p(a_s)>0, equivalently ∫_{a_s}^{a_h} c_ap da > −(d−a_s), or ∫ −c_ap da < d−a_s. The paper states instead that T grows iff −c_p(a_h)>−c_p(a_s)+d−a_s, which is equivalent to ∫ −c_ap da > d−a_s and implies ∂Δ/∂p<0. The same reversal occurs for d. In the leading case c_p=c_d=0, the paper's condition says T does not grow when p or d increases, even though the solver option directly gains d−a_s (for p) or p (for d); the correct condition says T does grow. The displayed inequalities and the surrounding comparative-static statements need to be reversed.
- [§4, proof of discontinuity] The boundary-optimality inequalities used in the discontinuity proof are stated with the wrong directions. For a_s(T)=d to be optimal on [0,d], the necessary condition is 1−p ≥ c_a(d,T) (the left derivative of the solver objective at d is nonnegative). For a_h(T)=d to be optimal on [d,1], the necessary condition is 1 ≤ c_a(d,T) (the right derivative of the helper objective at d is nonpositive). The contradiction 1−p<1 survives with weak inequalities, but the proof as written—using 1−p>c_a and 1<c_a—is not a valid derivation. Please correct these conditions.
- [§3.2, §4] The discontinuous knowledge gap is a direct consequence of the assumed sharp cutoff at d: the student's payoff has a kink at a=d because the marginal benefit jumps from 1−p to 1. The paper cites Schaeffer et al. (2023) but does not discuss how the result depends on the binary 'can solve/cannot solve' measure. Since the discontinuity is the paper's headline, please state explicitly that the prediction is conditional on the sharp-cutoff assumption, or add a robustness exercise showing how the gap behaves as the AI performance frontier is smoothed.
minor comments (5)
- [§5.3] The equation k(a_h)−k(a_s)=∫_{a_s}^{a_h} k′(a,T)da>0 is accompanied by the assertion 'as a_h(T)<a_s(T)'. At the threshold, however, a_h(T)>d>a_s(T), so the ordering claim is reversed. The conclusion that the threshold decreases is correct once the typo is fixed.
- [§4] In the proof of Proposition 1, 'TThe' should be 'The'.
- [§3.1] Typo: 'hallucianate' should be 'hallucinate'.
- [§2] Typos: 'awy' should be 'away'; 'decision-makes' should be 'decision-makers'; 'to incentive' in §5.1 should be 'to incentivize'.
- [Footnote 1] The footnote refers to 'Section 7' for the discussion of AI-free assignments; the substantive discussion occurs in Section 6. Please update the cross-reference.
Circularity Check
No significant circularity: the model’s derivation chain is self-contained and the main results follow from stated primitives rather than from fitted inputs or self-citations.
full rationale
The paper’s central result, Proposition 1, is derived directly from its explicit primitives: AI has a sharp capability cutoff at difficulty d and hallucination probability p, and students maximize the mass of solvable problems minus a convex cost of ability. The discontinuity in the optimal ability function A(t) follows mathematically from the kink in the payoff at a=d: below d the marginal benefit of ability is 1-p, above d it is 1, and with p<1 no optimum occurs exactly at d. This is an implication of the assumed payoff structure, not an input disguised as an output. No parameters are fitted to data; all comparative statics in Section 5 are computed from first-order conditions and the threshold characterization in Section 5.2. Section 6’s result that setting λ=p/p′ aligns a misspecified student’s objective with the true one is an algebraic derivation, not an assumed equivalence. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation: the sharp-cutoff assumption is explicitly justified by the emergence literature and is an exogenous modeling choice. The paper even acknowledges the alternative continuous-measure view (Schaeffer et al., 2023) and explains why the binary view is appropriate for student behavior. Thus the derivation chain is self-contained and no circular step is present. Any concerns about the existence of T in (0,1) are technical correctness issues, not circularity.
Axiom & Free-Parameter Ledger
axioms (6)
- standard math Problems are uniformly distributed on [0,1] and re-labeling makes this without loss.
- domain assumption AI solves problems of difficulty <= d with probability p and no problems > d.
- domain assumption A student of ability a solves all problems <= a with certainty.
- domain assumption Students maximize the mass of problems they can solve with AI assistance net of learning cost.
- domain assumption Cost function c(a,t|d,p) has all second-order derivatives and decreasing differences in (a,t).
- domain assumption Misspecified students behave as if AI is correct with probability p'>p, and instructors choose weight lambda on AI-permitted assignments.
Cite this review
Pith. "Pith review of Artificial or Human Intelligence?." pith.science (2026). https://pith.science/paper/CD5FFEKM
@misc{pith2026250902879,
author = {Pith},
title = {Pith review of: Artificial or Human Intelligence?},
year = {2026},
howpublished = {\url{https://pith.science/paper/CD5FFEKM}},
note = {Machine review of arXiv:2509.02879}
}
read the original abstract
Artificial intelligence (AI) tools such as large language models (LLMs) are already altering student learning. Unlike previous technologies, LLMs can independently solve problems regardless of student understanding, yet are not always accurate (due to hallucination) and face sharp performance cutoffs (due to emergence). Access to these tools significantly alters a student's incentives to learn, potentially decreasing the sum knowledge of humans and AI. Additionally, the marginal benefit of learning changes depending on which side of the AI frontier a human is on, creating a discontinuous gap between those that know more than or less than AI. This contrasts with downstream models of AI's impact on the labor force which assume continuous ability. Finally, increasing the portion of assignments where AI cannot be used can counteract student mis-specification about AI accuracy, preventing underinvestment. A better understanding of how AI impacts learning and student incentives is crucial for educators to adapt to this new technology.
Figures
Reference graph
Works this paper leans on
-
[1]
Anthropic education report: How educators use Claude
Anthropic Research. 2025a. “Anthropic education report: How educators use Claude.” https://www.anthropic.com/news/ anthropic-education-report-how-educators-use-claude . Anthropic Research. 2025b. “Anthropic Education Report: How Uni- versity Students Use Claude.” https://www.anthropic.com/news/ anthropic-education-report-how-university-students-use-claude...
Pith/arXiv arXiv 2024
-
[4]
The Impact of Artifi- cial Intelligence on Students’ Learning Experience
“The Impact of Artifi- cial Intelligence on Students’ Learning Experience.”Social Science Research Network. 10.2139/ssrn.4716747. Lehmann, Matthias, Philipp B. Cornelius, and Fabian J. Sting.2024. “AI Meets the Classroom: When Do Large Language Models Harm Learning?” arXiv preprint arXiv:2409.09047, https://arxiv.org/abs/2409.09047, Revised version (v2) s...
Pith/arXiv arXiv 2024
-
[5]
Who Benefits from AI? Self-Selection, Skill Gap, and the Hidden Costs of AI Feedback
Noy, Shakked, and Whitney Zhang.2023. “Experimental evidence on the productivity effects of generative artificial intelligence.”Science 381 187–192. 10.1126/science.adh2586. Riedl, Christoph, and Eric Bogert.2024. “Effects of AI Feedback on Learning, the Skill Gap, and Intellectual Diversity.”https://arxiv.org/abs/2409.18660. Schaeffer, Rylan, Brando Mira...
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[6]
The Impact of Artificial Intelli- gence (AI) on Students’ Academic Development
10.48550/arXiv.2304.15004. Vieriu, Aniella Mihaela, and Gabriel Petrea.2025. “The Impact of Artificial Intelli- gence (AI) on Students’ Academic Development.”Education Sciences 15
-
[9]
10.48550/arXiv.2401.11817. Zhai,Chunpeng,SantosoWibowo,andLilyDLi. 2024.“Theeffectsofover-relianceon AI dialogue systems on students’ cognitive abilities: a systematic review.”Smart Learning Environments 11 1–37. 10.1186/s40561-024-00316-7. 20
-
[10]
Hallucination is Inevitable: An Innate Limitation of Large Language Models
10.48550/arXiv.2206.07682. Xu, Ziwei, Sanjay Jain, and Mohan Kankanhalli.2024. “Hallucination is Inevitable: An Innate Limitation of Large Language Models.”
-
[343]
Emergent Abilities of Large Lan- guage Models
10.3390/ educsci15030343. Wei, Jason, Yi Tay, Rishi Bommasani et al.2022. “Emergent Abilities of Large Lan- guage Models.”
work page 2022
-
[2024]
Generative AI Can Harm Learning
“Generative AI Can Harm Learning.”Generative AI Can Harm Learning. 10.2139/ssrn.4895486. Berti, Leonardo, Flavio Giorgi, and Gjergji Kasneci.2025. “Emergent Abilities in Large Language Models: A Survey.”arXiv preprint arXiv:2503.05788, https://arxiv. org/abs/2503.05788, Submitted 28 February 2025; revised 14 March
Pith/arXiv arXiv 2025
-
[2025]
These AI-Skilled 20-Somethings Are Mak- ing Hundreds of Thousands a Year
“These AI-Skilled 20-Somethings Are Mak- ing Hundreds of Thousands a Year.” 08, https://www.wsj.com/tech/ai/ ai-jobs-entry-level-salary-ab2a11c0?gaa_at=eafs&gaa_n=ASWzDAjibQ_ yju9vTitXI-6I46hcqfa7idiXTrgQWu8bE-60506-kqHuLdYHEJKTIjU%3D&gaa_ts= 68b1b9d4&gaa_sig=BG2lKfkjehtZ3_md3Ylo8I7RD2X3H9gHsp-ACZd04-WC5DBtLi_qG9d7_ WzXP8qZ75sRQhnUIWk_5SnC-ECT8Q%3D%3D. Br...
2025
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.