REVIEW 3 major objections 4 minor 14 references
When workers overestimate AI capability, scaling up the AI can reduce the whole team's realized output—sometimes below human-only levels.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 00:11 UTC pith:NDNNKIE6
load-bearing objection Clean analytical mechanism for how over-perceived AI capability can make joint human-AI performance non-monotone in scale; the central result is real but rests on the unexamined exponential independence assumption, and the policy section is figures only. the 3 major comments →
The Scaling Paradox in Human-AI Collaboration
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is Proposition 2(iii): when a worker over-perceives the AI's capability—the perceived scaling factor exceeds the true one—the expected reward of the human-AI system is not monotone in AI scale over the human-in-the-loop region. Once the ratio of perceived to true scaling factor exceeds a threshold that depends only on the setup time, there is an intermediate scale at which the system earns strictly less than a human-only baseline, because the direct gain from a larger AI is outweighed by the worker's over-confident withdrawal of review effort. The same mechanism, carried to the firm level, makes profit fall even more sharply than reward wheneve
What carries the argument
The load-bearing object is the project success probability p(t; α, s) = 1 − e^{-αs−t}, built by treating AI failure as e^{-αs} (an inverse-power scaling law in physical scale, expressed in log-scale s) and human failure as e^{-t}, with independent failures. The worker maximizes perceived expected reward (1 − e^{-α̂s−t})·C/(t + t0), which makes the optimal per-project effort t the solution of t + t0 + 1 = e^{α̂s + t}, given by the lower Lambert W branch. That equation is the transmission belt: perceived capability determines effort, and actual realized reward is evaluated at that effort with the true scaling factor. Over-perception pushes effort down along this curve, so the comparison betwee
Load-bearing premise
The load-bearing premise is that AI task failure falls exactly as e^{-αs} with no performance floor and that human and AI failures are independent; relax either and the non-monotone scaling paradox may disappear or change sign.
What would settle it
Run a controlled experiment with an AI assistant whose true success probability is 1 − e^{-αs}, while inducing over-perception by telling workers it is more capable than it is; measure realized reward across several scales in the human-in-the-loop region. If reward never decreases with scale and never dips below the human-only baseline, Proposition 2(iii) is false. Alternatively, measure the AI's true failure curve; if it has a nonzero floor or is correlated with human errors, recompute the model and check whether the paradox survives.
If this is right
- When workers overestimate AI, scaling up models can reduce the joint system's expected output over part of the human-in-the-loop range; the realized system can be worse than no AI at that scale.
- Firm profit is even more exposed than system reward: whenever reward declines with scale, profit declines faster because the firm pays per-project AI costs on an increasing number of projects.
- Under-perception does not reverse scaling; it only slows the gains, and a moderate degree of under-perception can improve firm profit by partly offsetting the structural firm-worker misalignment.
- Cost internalization is not one-size-fits-all: shifting AI costs to workers helps when AI is cheap or deployed at small scale, but firms should absorb more cost when deployment is expensive or large to preserve worker adoption.
- Perception alignment reliably helps workers, but only helps firms under over-perception; aligning mildly under-perceiving workers can reduce firm profit.
Where Pith is reading between the lines
- A direct testable extension: in a controlled study where workers are given different capability claims about the same AI, output should trace the predicted non-monotone curve, and measuring effort per task would isolate the effort-withdrawal channel from direct scale effects.
- The model assumes no irreducible AI failure floor and no correlation between human and AI errors. If real deployments have a nonzero floor or correlated mistakes, the paradox may weaken or vanish, so the prediction is sharpest where those assumptions hold.
- Since the paper treats misperception as persistent, an implication of its own logic is that any intervention letting workers learn the true scaling factor—such as transparent error reporting—acts like perception alignment and should restore monotone scaling, which also suggests the paradox is most likely in fast-moving tasks where learning lags.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper builds an analytical model of human-AI collaboration in which a worker reviews AI output with effort t per project, the AI's success probability is 1 - e^{-αs} at scale s, and the joint success probability is p(t; α, s) = 1 - e^{-αs - t} under independence. The worker chooses t to maximize expected throughput, possibly using a misperceived scaling factor α̂. With accurate perception, total reward is non-decreasing in scale. If α̂ > α, the worker under-invests effort; Proposition 2(iii) shows that when α̂/α exceeds a t0-dependent threshold, realized reward is non-monotone in s over the human-in-the-loop interval and eventually falls below the human-only benchmark. If α̂ < α, reward still increases in s but more slowly than under perfect perception. The firm perspective introduces a per-project AI cost cs, generating a firm-worker misalignment that over-perception amplifies and under-perception can partially offset. Section 6 discusses cost internalization and perception alignment, supported by numerical figures. The formal proofs in Appendices A and B are internally consistent, and the central non-monotonicity proof is valid under the stated functional-form assumption.
Significance. If the conclusions hold, the paper makes a valuable contribution to the emerging operations literature on human-AI collaboration: it shows that the empirically observed scaling-law gains of AI need not translate into joint system gains when workers' beliefs about AI capability are biased, and it identifies an asymmetry between over- and under-perception that has direct managerial implications. The formal apparatus is parsimonious, and the paper ships explicit proofs for the main propositions, including the non-monotonicity result, rather than relying on simulation. The model's predictions are falsifiable in principle. However, the central paradox is proved only for a specific failure composition, and the policy section is illustrative rather than formally proven, which limits the generality of the paper's stated contributions.
major comments (3)
- [Section 6.1 and 6.2] The non-monotonicity result is proved only for the exact failure composition p(t;α,s)=1-e^{-αs-t}, which assumes q_A(s)=e^{-αs} with no additive floor and independence of AI and human failure. Section 3.1 itself describes q(N)=N^{-α} as a 'first-order component,' but the proof uses this expression as exact. With a floor q_A(s)=c0+(1-c0)e^{-αs}, the perceived failure has floor c0; if c0 ≥ 1/(1+t0), the worker never fully withdraws effort, and the collapse that drives Proposition 2(iii) at s_b=ln(1+t0)/α̂ does not occur. For intermediate c0, the threshold for the paradox changes and the comparison with the human-only benchmark may fail. The same concern applies to correlated human-AI failures. The paper should either prove the result under a perturbed failure model, give quantitative conditions on the floor/correlation that preserve the paradox, or explicitly restrict the central claim to
- [Section 6.1] The abstract and Section 1.1 claim that firms can 'actively manage' misperception through cost internalization and perception alignment, and the text asserts an 'optimal degree of cost internalization' and direction-dependent effectiveness of alignment. However, Section 6 contains no formal propositions or derivations; the policy conclusions rest entirely on Figs. 1-5 for selected parameter values. In particular, no result characterizes the optimal ρ as a function of (c,r,t0,α,α̂), and the claim that under-perception can make alignment reduce firm profit is only shown numerically. This is a load-bearing gap because the paper's fourth contribution is precisely these policy prescriptions. Either add analytical characterizations (even comparative statics on the profit-maximizing ρ) or soften the claims to 'numerically illustrated' policy insights.
- [Section 6.1] The extended-valued optimizer introduced for cost internalization is not rigorously integrated. When ρcs is large, the objective may have no finite maximizer, and the paper sets t=∞ with anticipated payoff zero. The resulting adoption threshold and payoff functions are discontinuous and non-differentiable, yet the subsequent discussion treats them graphically without specifying the selection rule or proving that the described comparative statics hold. This looseness matters because the discontinuities in Fig. 4 are used to draw the conclusion that over-perceiving workers 'opt out later.' Please formalize the extended-value definition and state the regularity conditions under which the figure-based conclusions are valid.
minor comments (4)
- [Section 3.1] The paper writes q(N)=O(N^{-α}) and then immediately sets q(N)=N^{-α}. Since the proof of Proposition 2 uses exact equality, the big-O notation should be removed or the transition explained as a modeling idealization rather than a first-order approximation.
- [Proposition 4(ii)] The condition c/r (1+αs) ≥ α is introduced as 'the third condition,' but the set I already includes it; the intersection I ∩ S_HIL,firm is therefore redundant. Consider simplifying the statement.
- [Appendix A.3] In the proof of Proposition 2(iii), the sentence 'the reward is initially increasing at s=0' is followed by a limit argument that establishes non-monotonicity. This is correct, but the threshold ¯γ is only shown to exist; it would be helpful to state that ¯γ depends on t0 only, as the proof actually demonstrates.
- [Throughout] Several equations are not numbered (e.g., the definitions of R(t;α,s), R*(α,s), and the cost-internalization payoff). Numbering them would make it easier for readers to follow the derivations in Sections 4-6.
Circularity Check
No significant circularity: the scaling-paradox result is derived from explicit optimizing behavior under a misperceived scaling exponent; the exponential failure form is an assumption, not a fit, and the two self-citations are not load-bearing.
full rationale
The paper's central derivation is self-contained. Section 3.1 posits an AI failure probability q_A(N)=N^{-α}=e^{-αs} and an independent human review failure probability 1-e^{-t}, leading to p(t;α,s)=1-e^{-αs-t}. Section 3.3 then defines the worker's optimization problem, first under true α and then under perceived α̂. Proposition 2(iii) is proved in Appendix A.3 by evaluating the realized reward R̂(α,s) at the perceived threshold s_b=ln(1+t0)/α̂ and comparing with the human-only benchmark; the non-monotonicity follows from continuity and the limit α̂/α→∞. This is a mathematical consequence of the stated optimization model, not an assertion that assumes the conclusion. The functional form p=1-e^{-αs-t} is a modeling assumption, and the result is therefore conditional on it, but an assumption is not circularity: the paper does not define the paradox into existence by fitting a parameter to the outcome. The only self-citations (Guan et al. 2025 and Beer et al. 2026) appear in the literature review and are not used as load-bearing evidence for any proposition; no uniqueness theorem from the authors' prior work is invoked, and no fitted input is relabeled as a prediction. The numerical figures use illustrative parameter values, not calibrated predictions from the model. Accordingly, no step in the derivation chain reduces by construction to its own inputs, and the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (7)
- α (true scaling factor)
- αhat (perceived scaling factor)
- t0 (setup time per project)
- C (worker capacity)
- c (unit AI deployment cost)
- r (firm net reward per success)
- ρ (worker cost share under cost internalization)
axioms (6)
- domain assumption AI and human success probabilities are independent and combine to p = 1 - e^{-αs-t}.
- domain assumption Worker maximizes expected reward using perceived αhat, and this belief is exogenous and persistent.
- domain assumption Worker commits review time t to every project regardless of whether AI output is correct.
- domain assumption Firm pays cs per project and earns r per success; worker pays no AI cost in baseline.
- domain assumption Projects are homogeneous; worker reward per success is normalized to one.
- standard math Lambert W branch properties and implicit differentiation are valid.
Cite this review
Pith. "Pith review of The Scaling Paradox in Human-AI Collaboration." pith.science (2026). https://pith.science/paper/NDNNKIE6
@misc{pith2026260800818,
author = {Pith},
title = {Pith review of: The Scaling Paradox in Human-AI Collaboration},
year = {2026},
howpublished = {\url{https://pith.science/paper/NDNNKIE6}},
note = {Machine review of arXiv:2608.00818}
}
read the original abstract
The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isolation; instead, it often works alongside humans, raising the question of whether these gains persist in human-AI collaboration. In this work, we develop an analytical model to examine when the empirical scaling benefits of AI translate into improved human-AI joint system performance. We demonstrate that the performance of a human-AI system can scale positively as the AI scales up-provided that humans have an accurate perception of the AI's capabilities. Human misperception, however, can fundamentally alter this relationship: i) when humans over-perceive the AI's capabilities, a scaling paradox may arise, in which greater AI scale reduces overall system performance and amplifies firm-level profit losses, and (ii) when humans under-perceive the AI's capabilities, performance still improves with scale but at a substantially slower rate. We further show that firms can actively manage these distortions through operational policies such as cost internalization and perception alignment, whose effectiveness depends on the economics of AI deployment and the direction of human misperception. These findings suggest that organizations may benefit more from managing the human-AI interface than from simply investing in larger, more expensive AI systems. More broadly, our results suggest that AI scaling should be viewed not only as a technological challenge, but also as a behavioral and operational one, and caution against the view that larger AI systems will automatically lead to better operational outcomes. Whether AI scaling creates value ultimately depends on how increased AI capabilities shape human beliefs and collaborative efforts.
Reference graph
Works this paper leans on
-
[1]
Bahri Y, Dyer E, Kaplan J, Lee J, Sharma U (2024) Explaining neural scaling laws.Proc. Natl. Acad. Sci. USA121(27):e2311878121. Bastani H, Cachon GP (2025) The human-AI contracting paradox. SSRN: 5962739. Becker J, Rush N, Barnes B, Rein D (2025) Measuring the impact of early-2025 AI on experienced open- source developer productivity. arXiv: 2507.09089. B...
Pith/arXiv arXiv 2024
-
[5]
Dietvorst BJ, Simmons JP, Massey C (2015) Algorithm aversion: People erroneously avoid algorithms after seeing them err.J. Experiment. Psych. General144(1):114–126. Qi and W ang:The Scaling Paradox in Human-AI Collaboration 37 Doerer K (2025) Klarna changes its AI tune and again recruits humans for customer ser- vice.https://www.customerexperiencedive.com...
work page 2015
-
[6]
Guan X, Qi A, Wang S (2025) Why the best machine may not be the best: Incentivizing human-machine collaboration. SSRN: 5283571. Hestness J, Narang S, Ardalani N, Diamos G, Jun H, Kianinejad H, Patwary MMA, Yang Y, Zhou Y (2017) Deep learning scaling is predictable, empirically. arXiv: 1712.00409. Hoffmann J, Borgeaud S, Mensch A, Buchatskaya E, Cai T, Rut...
Pith/arXiv arXiv 2025
-
[8]
arXiv: 2504.07139. Merali A (2025) Scaling laws for economic productivity: Experimental evidence in LLM-assisted consulting, data analyst, and management tasks. arXiv: 2512.21316. Naddaf M (2025) AI linked to explosion in low-quality biomedical papers.Nature641(5065):1080–1081. National Academies of Sciences, Engineering, and Medicine (2025) How AI is sha...
arXiv 2025
-
[9]
NDTV (2025) Company forces staff to buy AI tools, later refuses reimbursement: ‘isn’t this employee extortion?’.https://www.ndtv.com/offbeat/company-forces-staff-to-buy-ai-tools-later-r efuses-reimbursement-isnt-this-employee-extortion-9598079, accessed: July 20,
work page 2025
-
[10]
Noy S, Zhang W (2023) Experimental evidence on the productivity effects of generative artificial intelligence. Science381(6654):187–192. Parasuraman R, Manzey DH (2010) Complacency and bias in human use of automation: An attentional integration.Human Factors52(3):381–410. Peng S, Kalliamvakou E, Cihon P, Demirer M (2023) The impact of AI on developer prod...
Pith/arXiv arXiv 2023
-
[11]
Reuters (2024b) PwC to become OpenAI’s largest enterprise customer amid genAI boom. URLhttps://www.reuters.com/technology/pwc-become-openai-largest-enterprise-custome r-wsj-reports-2024-05-29/, accessed: July 20,
work page 2024
-
[12]
Sun J, Zhang DJ, Hu H, Van Mieghem JA (2022) Predicting human discretion to adjust algorithmic pre- scription: A large-scale field experiment in warehouse operations.Management Sci.68(2):846–865. The Guardian (2026) Bosses say AI boosts productivity – workers say they’re drowning in ‘workslop’.https://www.theguardian.com/technology/2026/apr/14/ai-producti...
work page 2022
-
[13]
Qi and W ang:The Scaling Paradox in Human-AI Collaboration 39 The New York Times (2026) More! More! More! Tech workers max out their A.I. use. URLhttps://www. nytimes.com/2026/03/20/technology/tokenmaxxing-ai-agents.html, accessed: July 20,
work page 2026
-
[14]
WalkMe (2026) Enterprises lose 51 workdays per employee to technology friction annually despite record AI investment.https://www.walkme.com/news-releases/enterprises-lose-51-workdays-per-emp loyee-to-technology-friction-annually-despite-record-ai-investment-walkme-global-s tudy-of-3750-finds/, State of Digital Adoption Report. Wang L, Ye Z, Zhao J (2025) ...
Pith/arXiv arXiv 2026
-
[35]
Hou T, Li M, Tan YR, Zhao H (2024) Physician adoption of AI assistant.Manufacturing Service Oper. Management26(5):1639–1655. Hu M, Jin Y, Nejati E (2026) AI co-creation. SSRN: 6554898. Huang C, Tang Z, Hu S, Jiang R, Zheng X, Ge D, Wang B, Wang Z (2025) ORLM: A customizable framework in training large models for automated optimization modeling.Oper. Res.7...
Pith/arXiv arXiv 2024
-
[825]
Challapally A, Pease C, Raskar R, Chari P (2025) The GenAI divide: State of AI in business
work page 2025
-
[2025]
Technical report, MIT NANDA. Cui KZ, Demirer M, Jaffe S, Musolff L, Peng S, Salz T (2026) The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers.Management Sci.Forthcoming. De V´ ericourt F, Gurkan H (2026) Is your machine better than you? You may never know.Management Sci. 72(1):558–574. Dell’Acq...
work page 2026
-
[2026]
Caro F, S´ aez de Tejada Cuenca A (2023) Believing in analytics: Managers’ adherence to price recommenda- tions from a DSS.Manufacturing Service Oper. Management25(2):524–542. Castelo N, Bos MW, Lehmann DR (2019) Task-dependent algorithm aversion.J. Marketing Res.56(5):809–
work page 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.