Pith. sign in

REVIEW 4 major objections 5 minor 9 references

The paper claims that when students choose how much to learn, AI's sharp performance cutoff produces a discontinuous gap in human ability at the AI frontier.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A theoretical model shows that AI's sharp capability cutoff and hallucination rate create a discontinuous gap in student ability around the AI frontier, and that AI-free assignments can correct students' overestimation of AI accuracy.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Nice idea, but Proposition 1 overreaches as stated; the paper needs boundary conditions and proof fixes before the threshold result can stand. the 4 major comments →

arxiv 2509.02879 v1 pith:CD5FFEKM submitted 2025-09-02 econ.TH

Artificial or Human Intelligence?

classification econ.TH
keywords AILarge Language ModelsEducationHuman CapitalInvestmentStudent incentivesEmergent abilitiesHallucination
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies how a student's incentive to invest in learning changes when an AI tool can autonomously solve some problems but not others. It models AI as having a sharp difficulty cutoff d and, below that cutoff, a probability p of hallucinating, and lets each student choose how much ability to build given a personal cost of learning. The central result is that students sort into two groups separated by a threshold type: those who use AI as a solver and stop learning below the AI's frontier, and those who use AI as a helper and learn past it. Because the marginal benefit of learning jumps from 1−p to 1 at the frontier, the ability distribution is discontinuous there—a knowledge gap opens at precisely the level AI has mastered. The paper then derives comparative statics and an optimal exam-design rule that can correct students who overestimate AI accuracy.

Core claim

Proposition 1 is the paper's central claim: under the model, there exists a unique threshold type T such that every student with type below T chooses ability as(T) < d and uses AI as an independent solver, while every type above T chooses ability ah(T) > d and uses AI only as a helper. At T the student is indifferent, and the induced ability function A(t) jumps at T, skipping the entire interval from as(T) to ah(T), which contains d. The reason is a discontinuity in marginal benefit: a student just below the frontier gains only 1−p from each additional unit of ability, because AI already solves most problems in that range, while a student just above the frontier gains a full 1, since those m

What carries the argument

The machinery is a binary AI frontier (d,p): AI solves problems of difficulty at most d with probability p and none above d, while a human with ability a solves problems up to a with certainty. Students maximize the mass of solvable problems minus a learning cost c(a,t) that has decreasing differences in ability and type. The two key identities are the payoff slopes—marginal benefit 1−p for students below d and 1 for students above d—and the threshold type T defined by indifference between the solver value function and the helper value function. The slope discontinuity is what produces the gap in the induced ability mapping A(t).

Load-bearing premise

The load-bearing assumption is that AI capability is a sharp cliff—it succeeds on every problem up to a fixed difficulty with one probability and on no problem above it—and that students evaluate the AI only by whether it gives the correct final answer; if performance improves gradually or partial credit matters, the marginal benefit of learning is continuous and the predicted ability gap disappears.

What would settle it

Give one group of students unrestricted access to an LLM in a course with a graded exam, then compare the distribution of exam scores around the difficulty level at which that LLM's accuracy drops; if the density of scores is continuous through that level, rather than showing a hollow just below it, the Proposition 1 discontinuity is falsified. A cleaner version: measure students' chosen study time as a function of their distance from the AI's difficulty frontier; the model predicts a strictly positive jump in study time as students cross the frontier.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • For an exogenous fixed distribution of abilities, introducing AI helps the least-skilled students most; with endogenous investment, the distribution of ability is no longer continuous and has a gap around the AI frontier.
  • Raising AI accuracy p lowers the marginal return to learning for solver-side students and raises the threshold T, so more students use AI as a solver and the ability gap can widen.
  • Raising the difficulty threshold d does not change the marginal benefit of learning for anyone but shifts T upward, since the solver option becomes relatively more attractive at the old cutoff.
  • If AI reduces learning costs more at higher ability levels, all students choose more ability and T falls; absent such cost benefits, AI advances mainly replace rather than augment human knowledge.
  • Putting a share λ = p/p′ of assignments outside AI access aligns misspecified students' incentives with the true technology; the larger the overestimate of AI accuracy, the larger the no-AI weight needed.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The discontinuity prediction is a directly testable implication: in courses with universal AI access, ability or exam-score distributions should show a hollow just below the difficulty level where the AI starts failing, and students who self-identify as 'solver' users should cluster on the low side of that hollow.
  • Because the comparative statics show the threshold rises with p and d, a calibrated version of the model could predict which course levels—introductory versus advanced—see the largest drops in investment as AI models improve.
  • The λ = p/p′ rule suggests a practical calibration procedure: estimate the AI's true accuracy p and students' believed accuracy p′ on a class's problem set, then set the closed-book share accordingly; this follows from Section 6 but is not stated as a measurement protocol in the paper.
  • If AI performance turns out to be continuous rather than sharply emergent, the model's main discontinuity would smooth into a conventional continuous trade-off, and the education-policy conclusions would need to be re-derived—the paper itself flags the emergence-mirage evidence as the main challenge to its cutoff assumption.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper builds a model in which students have type-dependent costs of acquiring ability a∈[0,1], and AI can solve all problems of difficulty ≤d with probability p and no problems above d. Students choose ability to maximize the mass of problems they can solve with AI assistance, net of cost. The model separates students into those who use AI as a 'solver' (ability below d) and those who use it as a 'helper' (ability above d). The central claims are: (i) there is a unique threshold type T separating the two groups, with a discontinuous upward jump in ability at T; (ii) advances in p or d shift this threshold and affect ability choices; (iii) if students overestimate AI accuracy, instructors can restore efficient investment by placing weight λ=p/p′ on AI-permitted assignments. The paper motivates these results with a review of hallucination and emergence in LLMs and with recent empirical education studies.

Significance. The topic is timely and the solver/helper distinction is a useful organizing device for thinking about AI in education. The paper is clearly written and connects to a rich empirical literature. Section 6's λ=p/p′ alignment result is elegant and correct. If the stated theorem and comparative statics are repaired, the model would provide a sharp, policy-relevant prediction about a possible knowledge gap around the AI frontier. However, the current version contains a false universal claim in Proposition 1 and a sign error in the main comparative-statics conditions, so the contribution is not yet established as stated.

major comments (4)
  1. [§4, Proposition 1] As stated, Proposition 1 is false: decreasing differences alone does not imply the existence of an interior threshold T∈(0,1). The proof's inference 'if A(t)=a_s(t) then A(t')=a_s(t') for all t'>t' uses A(t')≥A(t)=a_s(t)≥d, but a_s(t)≤d by definition, so the conclusion does not follow. More importantly, the asserted interior threshold requires boundary conditions that are not stated. Counterexample: take c(a,t)=a^2/[2(t+0.01)], p=0.2, d=0.95, which satisfies all listed assumptions. For every t∈[0,1] the solver branch strictly dominates the helper branch (at t=1, U_s≈0.513 and U_h≈0.505), so no type uses AI as a helper and T∉(0,1). The paper needs explicit conditions ensuring that both branches are chosen—e.g., U_s(0)>U_h(0), U_s(1)<U_h(1), and single crossing—before the headline discontinuity can be claimed.
  2. [§5.2] The conditions for the threshold to increase in p and d have the inequality reversed. Since Δ(t)=U_s(t)−U_h(t) is decreasing in t at the threshold, T increases iff ∂Δ/∂p>0 or ∂Δ/∂d>0 at T. For p, this gives d−a_s+c_p(a_h)−c_p(a_s)>0, equivalently ∫_{a_s}^{a_h} c_ap da > −(d−a_s), or ∫ −c_ap da < d−a_s. The paper states instead that T grows iff −c_p(a_h)>−c_p(a_s)+d−a_s, which is equivalent to ∫ −c_ap da > d−a_s and implies ∂Δ/∂p<0. The same reversal occurs for d. In the leading case c_p=c_d=0, the paper's condition says T does not grow when p or d increases, even though the solver option directly gains d−a_s (for p) or p (for d); the correct condition says T does grow. The displayed inequalities and the surrounding comparative-static statements need to be reversed.
  3. [§4, proof of discontinuity] The boundary-optimality inequalities used in the discontinuity proof are stated with the wrong directions. For a_s(T)=d to be optimal on [0,d], the necessary condition is 1−p ≥ c_a(d,T) (the left derivative of the solver objective at d is nonnegative). For a_h(T)=d to be optimal on [d,1], the necessary condition is 1 ≤ c_a(d,T) (the right derivative of the helper objective at d is nonpositive). The contradiction 1−p<1 survives with weak inequalities, but the proof as written—using 1−p>c_a and 1<c_a—is not a valid derivation. Please correct these conditions.
  4. [§3.2, §4] The discontinuous knowledge gap is a direct consequence of the assumed sharp cutoff at d: the student's payoff has a kink at a=d because the marginal benefit jumps from 1−p to 1. The paper cites Schaeffer et al. (2023) but does not discuss how the result depends on the binary 'can solve/cannot solve' measure. Since the discontinuity is the paper's headline, please state explicitly that the prediction is conditional on the sharp-cutoff assumption, or add a robustness exercise showing how the gap behaves as the AI performance frontier is smoothed.
minor comments (5)
  1. [§5.3] The equation k(a_h)−k(a_s)=∫_{a_s}^{a_h} k′(a,T)da>0 is accompanied by the assertion 'as a_h(T)<a_s(T)'. At the threshold, however, a_h(T)>d>a_s(T), so the ordering claim is reversed. The conclusion that the threshold decreases is correct once the typo is fixed.
  2. [§4] In the proof of Proposition 1, 'TThe' should be 'The'.
  3. [§3.1] Typo: 'hallucianate' should be 'hallucinate'.
  4. [§2] Typos: 'awy' should be 'away'; 'decision-makes' should be 'decision-makers'; 'to incentive' in §5.1 should be 'to incentivize'.
  5. [Footnote 1] The footnote refers to 'Section 7' for the discussion of AI-free assignments; the substantive discussion occurs in Section 6. Please update the cross-reference.

Circularity Check

0 steps flagged

No significant circularity: the model’s derivation chain is self-contained and the main results follow from stated primitives rather than from fitted inputs or self-citations.

full rationale

The paper’s central result, Proposition 1, is derived directly from its explicit primitives: AI has a sharp capability cutoff at difficulty d and hallucination probability p, and students maximize the mass of solvable problems minus a convex cost of ability. The discontinuity in the optimal ability function A(t) follows mathematically from the kink in the payoff at a=d: below d the marginal benefit of ability is 1-p, above d it is 1, and with p<1 no optimum occurs exactly at d. This is an implication of the assumed payoff structure, not an input disguised as an output. No parameters are fitted to data; all comparative statics in Section 5 are computed from first-order conditions and the threshold characterization in Section 5.2. Section 6’s result that setting λ=p/p′ aligns a misspecified student’s objective with the true one is an algebraic derivation, not an assumed equivalence. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation: the sharp-cutoff assumption is explicitly justified by the emergence literature and is an exogenous modeling choice. The paper even acknowledges the alternative continuous-measure view (Schaeffer et al., 2023) and explains why the binary view is appropriate for student behavior. Thus the derivation chain is self-contained and no circular step is present. Any concerns about the existence of T in (0,1) are technical correctness issues, not circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

The model's d and p are exogenous parameters describing AI technology, not fitted to data or chosen ad hoc; the cost function c(a,t|d,p) is a primitive. Core results are qualitative comparative statics, so no fitted values are used. No new entities are introduced; 'AI as a solver' and 'AI as a helper' are descriptive labels for the two regimes.

axioms (6)
  • standard math Problems are uniformly distributed on [0,1] and re-labeling makes this without loss.
    Section 4: 'As problem difficulty is uniformly distributed... Re-labeling problems makes this without loss.'
  • domain assumption AI solves problems of difficulty <= d with probability p and no problems > d.
    Section 4 model primitives; captures hallucination and emergence. This is the load-bearing stylized fact.
  • domain assumption A student of ability a solves all problems <= a with certainty.
    Section 4: 'A human with ability level a can solve problems up to difficulty a with one hundred percent accuracy.'
  • domain assumption Students maximize the mass of problems they can solve with AI assistance net of learning cost.
    Section 4: 'suppose humans simply choose to maximize the mass of problems they can solve with AI assistance.'
  • domain assumption Cost function c(a,t|d,p) has all second-order derivatives and decreasing differences in (a,t).
    Section 4 assumptions 1 and 2; this guarantees monotone comparative statics.
  • domain assumption Misspecified students behave as if AI is correct with probability p'>p, and instructors choose weight lambda on AI-permitted assignments.
    Section 6 model of misspecification.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial or Human Intelligence?." pith.science (2026). https://pith.science/paper/CD5FFEKM

@misc{pith2026250902879,
  author       = {Pith},
  title        = {Pith review of: Artificial or Human Intelligence?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CD5FFEKM}},
  note         = {Machine review of arXiv:2509.02879}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Artificial intelligence (AI) tools such as large language models (LLMs) are already altering student learning. Unlike previous technologies, LLMs can independently solve problems regardless of student understanding, yet are not always accurate (due to hallucination) and face sharp performance cutoffs (due to emergence). Access to these tools significantly alters a student's incentives to learn, potentially decreasing the sum knowledge of humans and AI. Additionally, the marginal benefit of learning changes depending on which side of the AI frontier a human is on, creating a discontinuous gap between those that know more than or less than AI. This contrasts with downstream models of AI's impact on the labor force which assume continuous ability. Finally, increasing the portion of assignments where AI cannot be used can counteract student mis-specification about AI accuracy, preventing underinvestment. A better understanding of how AI impacts learning and student incentives is crucial for educators to adapt to this new technology.

Figures

Figures reproduced from arXiv: 2509.02879 by Eric Gao.

Figure 1
Figure 1. Figure 1: AI as a Solver. In this case, the set of problems AI can solve is not a subset of the problems this human can solve. As such, for some problems (of difficulty 1/2 to 3/4) only the AI can make progress and AI is used as an independent solver. On the other hand, if human ability is greater than d, the AI is never needed to independently solve any problems. This case is plotted in the following figure, where … view at source ↗
Figure 2
Figure 2. Figure 2: AI as a Helper. Next, suppose humans simply choose to maximize the mass of problems they can solve with AI assistance.1 As problem difficulty is uniformly distributed2 the mass of solvable problems is simply the area of the union of the two rectangles (human-solvable and AI￾solvable). As such, humans who use AI as a solver solve: a s (t) ∈ argmax a∈[0,d] {(1 − p)a + dp − c(a, t|d, p)} while humans who use … view at source ↗
Figure 3
Figure 3. Figure 3: Ability Versus Type (Full) Proposition 1. There exists a unique threshold type T ∈ (0, 1) such that all humans of type t < T use AI as a solver while all humans of type t > T use AI as a helper. Then, A(t) = a s (t) for t < T and A(t) = a h (t) for t > T. Furthermore, humans of type T are indifferent between a s (T) and a h (T) and A(t) has a discontinuity at T with lim t→T − A(t) = a s (T) < d < ah (T) = … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

9 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    Anthropic education report: How educators use Claude

    Anthropic Research. 2025a. “Anthropic education report: How educators use Claude.” https://www.anthropic.com/news/ anthropic-education-report-how-educators-use-claude . Anthropic Research. 2025b. “Anthropic Education Report: How Uni- versity Students Use Claude.” https://www.anthropic.com/news/ anthropic-education-report-how-university-students-use-claude...

  2. [4]

    The Impact of Artifi- cial Intelligence on Students’ Learning Experience

    “The Impact of Artifi- cial Intelligence on Students’ Learning Experience.”Social Science Research Network. 10.2139/ssrn.4716747. Lehmann, Matthias, Philipp B. Cornelius, and Fabian J. Sting.2024. “AI Meets the Classroom: When Do Large Language Models Harm Learning?” arXiv preprint arXiv:2409.09047, https://arxiv.org/abs/2409.09047, Revised version (v2) s...

  3. [5]

    Who Benefits from AI? Self-Selection, Skill Gap, and the Hidden Costs of AI Feedback

    Noy, Shakked, and Whitney Zhang.2023. “Experimental evidence on the productivity effects of generative artificial intelligence.”Science 381 187–192. 10.1126/science.adh2586. Riedl, Christoph, and Eric Bogert.2024. “Effects of AI Feedback on Learning, the Skill Gap, and Intellectual Diversity.”https://arxiv.org/abs/2409.18660. Schaeffer, Rylan, Brando Mira...

  4. [6]

    The Impact of Artificial Intelli- gence (AI) on Students’ Academic Development

    10.48550/arXiv.2304.15004. Vieriu, Aniella Mihaela, and Gabriel Petrea.2025. “The Impact of Artificial Intelli- gence (AI) on Students’ Academic Development.”Education Sciences 15

  5. [9]

    Theeffectsofover-relianceon AI dialogue systems on students’ cognitive abilities: a systematic review

    10.48550/arXiv.2401.11817. Zhai,Chunpeng,SantosoWibowo,andLilyDLi. 2024.“Theeffectsofover-relianceon AI dialogue systems on students’ cognitive abilities: a systematic review.”Smart Learning Environments 11 1–37. 10.1186/s40561-024-00316-7. 20

  6. [10]

    Hallucination is Inevitable: An Innate Limitation of Large Language Models

    10.48550/arXiv.2206.07682. Xu, Ziwei, Sanjay Jain, and Mohan Kankanhalli.2024. “Hallucination is Inevitable: An Innate Limitation of Large Language Models.”

  7. [343]

    Emergent Abilities of Large Lan- guage Models

    10.3390/ educsci15030343. Wei, Jason, Yi Tay, Rishi Bommasani et al.2022. “Emergent Abilities of Large Lan- guage Models.”

  8. [2024]

    Generative AI Can Harm Learning

    “Generative AI Can Harm Learning.”Generative AI Can Harm Learning. 10.2139/ssrn.4895486. Berti, Leonardo, Flavio Giorgi, and Gjergji Kasneci.2025. “Emergent Abilities in Large Language Models: A Survey.”arXiv preprint arXiv:2503.05788, https://arxiv. org/abs/2503.05788, Submitted 28 February 2025; revised 14 March

  9. [2025]

    These AI-Skilled 20-Somethings Are Mak- ing Hundreds of Thousands a Year

    “These AI-Skilled 20-Somethings Are Mak- ing Hundreds of Thousands a Year.” 08, https://www.wsj.com/tech/ai/ ai-jobs-entry-level-salary-ab2a11c0?gaa_at=eafs&gaa_n=ASWzDAjibQ_ yju9vTitXI-6I46hcqfa7idiXTrgQWu8bE-60506-kqHuLdYHEJKTIjU%3D&gaa_ts= 68b1b9d4&gaa_sig=BG2lKfkjehtZ3_md3Ylo8I7RD2X3H9gHsp-ACZd04-WC5DBtLi_qG9d7_ WzXP8qZ75sRQhnUIWk_5SnC-ECT8Q%3D%3D. Br...

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.