{"id":"fde48052-71be-43ce-a152-0eb86678de60","arxiv_id":"2608.00818","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Over-perception of AI capability can create a scaling paradox in which larger AI reduces human-AI system performance and firm profit.","lead":"An analytical model shows that when workers overestimate an AI tool's abilities, making the AI bigger can reduce total work output and firm profit, even though the AI alone improves. The paper suggests firms should invest in managing human beliefs and incentives, not just in larger AI models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The scaling-paradox result is proved only for the exact failure composition p=1-e^{-αs-t}; without a robustness check for floor/correlated failures, the central claim remains conditional on that functional form.","rationale":"I checked the Lambert-W characterization (Prop. 1), the reward simplification in Corollary 1, and the sign arguments in Prop. 3; the internal mathematics is consistent. My remaining concern is the external robustness of the key non-monotonicity result. The reader identified the same assumption, so I agree on the locus of risk, but I am only partial: my own derivation suggests the paradox may survive many floors/correlations (e.g., an additive floor still leaves a transient effort-withdrawal dip), so the functional form is likely less brittle than the reader's wording suggests. That is precisely why the proposed numerical grid is the decisive check. The paper's policy claims in Section 6 are explicitly numerical illustrations rather than theorems, which supports keeping the CONDITIONAL verdict; my stress test does not change that, hence UNCHANGED.","tokens_in":27609,"tokens_out":19910,"duration_ms":246666,"concrete_test":"Numerically solve the worker's problem under the generalized failure model p=1-[c0+(1-c0)e^{-αs}][η+(1-η)e^{-t}], where the worker maximizes the perceived version with αhat replacing α. Set α=1, r=1, c=0.2, t0∈{1,4,10}; scan c0∈{0,0.1,0.2,0.3,0.5,0.8}, η∈{0,0.2,0.5,0.8}, and αhat/α∈{1.5,2,4,10}. For each cell, compute Rhat(s) on [0,ln(1+t0)/α] and record (i) whether Rhat is non-monotone (has a local max with subsequent decrease), and (ii) whether min_s Rhat(s)<R*(α,0). If both conditions fail in an empirically plausible cell, Proposition 2(iii) is an artifact of the exponential/no-floor/no-correlation form and the central claim must be qualified; if they hold across the grid, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is Proposition 2(iii): over-perception with αhat/α above a threshold makes realized reward non-monotone in scale and eventually below the human-only benchmark. The proof (Appendix A.3) uses two model primitives from Section 3.1: AI failure q_A=e^{-αs} and independent human review failure q_H=e^{-t}, yielding p=1-e^{-αs-t}. Section 3.1 explicitly says q(N)=N^{-α} is a 'first-order component' and 'assumes' independence; no empirical calibration or sensitivity analysis is given for these primitives. The mechanism is that over-perceived αhat drives t*(αhat,s) to 0 at s_b=ln(1+t0)/αhat, while actual AI failure e^{-αs_b} is still near 1, so realized reward collapses. This mechanism is not obviously robust: with a nonzero floor q_A(s)=c0+(1-c0)e^{-αs}, the worker may never fully withdraw effort (if c0>1/(1+t0)), and with uncorrectable human error q_H|A(t)=η+(1-η)e^{-t}, the effort channel weakens. The paper does not show Proposition 2(iii) extends to these perturbed failure models, so the generality of the 'scaling paradox' is exactly as load-bearing as this unverified primitive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper builds an analytical model of human-AI collaboration in which a worker reviews AI output with effort t per project, the AI's success probability is 1 - e^{-αs} at scale s, and the joint success probability is p(t; α, s) = 1 - e^{-αs - t} under independence. The worker chooses t to maximize expected throughput, possibly using a misperceived scaling factor α̂. With accurate perception, total reward is non-decreasing in scale. If α̂ > α, the worker under-invests effort; Proposition 2(iii) shows that when α̂/α exceeds a t0-dependent threshold, realized reward is non-monotone in s over the human-in-the-loop interval and eventually falls below the human-only benchmark. If α̂ < α, reward still increases in s but more slowly than under perfect perception. The firm perspective introduces a per-project AI cost cs, generating a firm-worker misalignment that over-perception amplifies and under-perception can partially offset. Section 6 discusses cost internalization and perception alignment, supported by numerical figures. The formal proofs in Appendices A and B are internally consistent, and the central non-monotonicity proof is valid under the stated functional-form assumption.","tokens_in":27925,"tokens_out":7270,"duration_ms":87041,"significance":"If the conclusions hold, the paper makes a valuable contribution to the emerging operations literature on human-AI collaboration: it shows that the empirically observed scaling-law gains of AI need not translate into joint system gains when workers' beliefs about AI capability are biased, and it identifies an asymmetry between over- and under-perception that has direct managerial implications. The formal apparatus is parsimonious, and the paper ships explicit proofs for the main propositions, including the non-monotonicity result, rather than relying on simulation. The model's predictions are falsifiable in principle. However, the central paradox is proved only for a specific failure composition, and the policy section is illustrative rather than formally proven, which limits the generality of the paper's stated contributions.","major_comments":[{"comment":"The non-monotonicity result is proved only for the exact failure composition p(t;α,s)=1-e^{-αs-t}, which assumes q_A(s)=e^{-αs} with no additive floor and independence of AI and human failure. Section 3.1 itself describes q(N)=N^{-α} as a 'first-order component,' but the proof uses this expression as exact. With a floor q_A(s)=c0+(1-c0)e^{-αs}, the perceived failure has floor c0; if c0 ≥ 1/(1+t0), the worker never fully withdraws effort, and the collapse that drives Proposition 2(iii) at s_b=ln(1+t0)/α̂ does not occur. For intermediate c0, the threshold for the paradox changes and the comparison with the human-only benchmark may fail. The same concern applies to correlated human-AI failures. The paper should either prove the result under a perturbed failure model, give quantitative conditions on the floor/correlation that preserve the paradox, or explicitly restrict the central claim to","section":"Section 6.1 and 6.2"},{"comment":"The abstract and Section 1.1 claim that firms can 'actively manage' misperception through cost internalization and perception alignment, and the text asserts an 'optimal degree of cost internalization' and direction-dependent effectiveness of alignment. However, Section 6 contains no formal propositions or derivations; the policy conclusions rest entirely on Figs. 1-5 for selected parameter values. In particular, no result characterizes the optimal ρ as a function of (c,r,t0,α,α̂), and the claim that under-perception can make alignment reduce firm profit is only shown numerically. This is a load-bearing gap because the paper's fourth contribution is precisely these policy prescriptions. Either add analytical characterizations (even comparative statics on the profit-maximizing ρ) or soften the claims to 'numerically illustrated' policy insights.","section":"Section 6.1"},{"comment":"The extended-valued optimizer introduced for cost internalization is not rigorously integrated. When ρcs is large, the objective may have no finite maximizer, and the paper sets t=∞ with anticipated payoff zero. The resulting adoption threshold and payoff functions are discontinuous and non-differentiable, yet the subsequent discussion treats them graphically without specifying the selection rule or proving that the described comparative statics hold. This looseness matters because the discontinuities in Fig. 4 are used to draw the conclusion that over-perceiving workers 'opt out later.' Please formalize the extended-value definition and state the regularity conditions under which the figure-based conclusions are valid.","section":"Section 6.1"}],"minor_comments":[{"comment":"The paper writes q(N)=O(N^{-α}) and then immediately sets q(N)=N^{-α}. Since the proof of Proposition 2 uses exact equality, the big-O notation should be removed or the transition explained as a modeling idealization rather than a first-order approximation.","section":"Section 3.1"},{"comment":"The condition c/r (1+αs) ≥ α is introduced as 'the third condition,' but the set I already includes it; the intersection I ∩ S_HIL,firm is therefore redundant. Consider simplifying the statement.","section":"Proposition 4(ii)"},{"comment":"In the proof of Proposition 2(iii), the sentence 'the reward is initially increasing at s=0' is followed by a limit argument that establishes non-monotonicity. This is correct, but the threshold ¯γ is only shown to exist; it would be helpful to state that ¯γ depends on t0 only, as the proof actually demonstrates.","section":"Appendix A.3"},{"comment":"Several equations are not numbered (e.g., the definitions of R(t;α,s), R*(α,s), and the cost-internalization payoff). Numbering them would make it easier for readers to follow the derivations in Sections 4-6.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the central analytical machinery is sound under its stated assumptions. The main concern is not internal inconsistency but scope: the flagship Proposition 2(iii) depends sensitively on the exact exponential failure composition, and the policy section is currently only numeric. Both can be addressed in revision: add a robustness analysis (floor/correlation or explicit domain restriction) and either formalize the policy results or clearly label them as simulations. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful way to read this paper is as a clean theoretical mechanism: if workers misperceive the AI scaling factor, their effort allocation can make joint system performance non-monotone in AI scale. The main result, Proposition 2(iii), is a real derivation, not a curve fit, and I checked the Lambert W characterization and the proof chain; it holds. The novelty claim is fair—I don't know of another OM paper that puts the scaling law inside a capacity-allocation model and lets effort respond to perceived capability. The asymmetry between over- and under-perception is a nice touch, and the firm-level analysis in Section 5 is competently done.\n\nThe soft spots are two. First, the paradox depends on the specific failure composition p = 1 - e^{-αs-t}. The authors are explicit that this is a 'first-order component' and that independence is assumed, but they don't test whether the result survives a floor in the AI failure curve or correlated human/AI errors. On a floor alone, with the worker still perceiving no floor, the collapse at the perceived AGI threshold still occurs; but if the worker's model also has a floor, or if human correction is strongly correlated with AI failure, the effort channel weakens and the non-monotonicity can disappear. So the central claim is exactly as general as that primitive, and the paper doesn't say how fragile it is. That should be fixed with a robustness section or an explicit boundary on parameters.\n\nSecond, Section 6's policy claims—cost internalization and perception alignment—are supported only by numerical figures, not propositions. The qualitative takeaways are plausible, but the abstract says 'we show' when the section shows selected parameter figures. This is a gap between contribution claim and evidence. It's a minor overclaim if read charitably, but worth flagging. Also, no code or data are shipped for those figures, so reproducibility is limited to whatever the authors chose to plot.\n\nOne more minor thing: the threshold γ in Prop 2(iii) is existential, not characterized. For practical guidance one would like to know how large over-perception must be to trigger the paradox. Not fatal.\n\nWho is this for? People in OM or behavioral operations working on human-AI systems. The paper gives them a formal mechanism to think with, even if the functional-form assumption keeps the result conditional. It deserves a serious referee; the right verdict after revision would probably be conditional acceptance, with robustness analysis and a tighter Section 6. I'd read it again after those changes.","headline":"Clean analytical mechanism for how over-perceived AI capability can make joint human-AI performance non-monotone in scale; the central result is real but rests on the unexamined exponential independence assumption, and the policy section is figures only.","tokens_in":28438,"tokens_out":5806,"would_cite":true,"duration_ms":61222,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"When workers overestimate AI capability, scaling up the AI can reduce the whole team's realized output—sometimes below human-only levels.","keywords":["human-AI collaboration","scaling law","misperception","over-perception","under-perception","effort allocation","automation bias","algorithm aversion"],"falsifier":"Run a controlled experiment with an AI assistant whose true success probability is 1 − e^{-αs}, while inducing over-perception by telling workers it is more capable than it is; measure realized reward across several scales in the human-in-the-loop region. If reward never decreases with scale and never dips below the human-only baseline, Proposition 2(iii) is false. Alternatively, measure the AI's true failure curve; if it has a nonzero floor or is correlated with human errors, recompute the model and check whether the paradox survives.","tokens_in":27457,"feed_emoji":"📉","tokens_out":7156,"duration_ms":78547,"temperature":0.7,"pith_summary":"The paper sets out to show that the empirical scaling law of AI—better capability with more scale—does not automatically carry over to human-AI teams. In its model, a worker with fixed capacity chooses how much effort to spend reviewing each AI-assisted project, and that choice is driven by the worker's perceived scaling factor rather than the true one. If workers overestimate AI capability, they cut their own effort too aggressively, and past a threshold of over-perception the realized reward becomes non-monotone in AI scale: larger AI can produce less total output, and at some scales less than a system with no AI at all. Underestimation leaves scaling beneficial but weaker. The paper argues this matters because firms can manage the distortion through cost internalization and perception alignment, and often benefit more from managing beliefs than from scaling models further.","feed_headline":"Overestimating AI makes scaling it backfire","feed_subtitle":"Inflated beliefs about AI can push human-AI teams below human-only performance, the model shows.","key_machinery":"The load-bearing object is the project success probability p(t; α, s) = 1 − e^{-αs−t}, built by treating AI failure as e^{-αs} (an inverse-power scaling law in physical scale, expressed in log-scale s) and human failure as e^{-t}, with independent failures. The worker maximizes perceived expected reward (1 − e^{-α̂s−t})·C/(t + t0), which makes the optimal per-project effort t the solution of t + t0 + 1 = e^{α̂s + t}, given by the lower Lambert W branch. That equation is the transmission belt: perceived capability determines effort, and actual realized reward is evaluated at that effort with the true scaling factor. Over-perception pushes effort down along this curve, so the comparison betwee","core_discovery":"On the paper's own terms, the central discovery is Proposition 2(iii): when a worker over-perceives the AI's capability—the perceived scaling factor exceeds the true one—the expected reward of the human-AI system is not monotone in AI scale over the human-in-the-loop region. Once the ratio of perceived to true scaling factor exceeds a threshold that depends only on the setup time, there is an intermediate scale at which the system earns strictly less than a human-only baseline, because the direct gain from a larger AI is outweighed by the worker's over-confident withdrawal of review effort. The same mechanism, carried to the firm level, makes profit fall even more sharply than reward wheneve","pith_inferences":["A direct testable extension: in a controlled study where workers are given different capability claims about the same AI, output should trace the predicted non-monotone curve, and measuring effort per task would isolate the effort-withdrawal channel from direct scale effects.","The model assumes no irreducible AI failure floor and no correlation between human and AI errors. If real deployments have a nonzero floor or correlated mistakes, the paradox may weaken or vanish, so the prediction is sharpest where those assumptions hold.","Since the paper treats misperception as persistent, an implication of its own logic is that any intervention letting workers learn the true scaling factor—such as transparent error reporting—acts like perception alignment and should restore monotone scaling, which also suggests the paradox is most likely in fast-moving tasks where learning lags."],"forward_implications":["When workers overestimate AI, scaling up models can reduce the joint system's expected output over part of the human-in-the-loop range; the realized system can be worse than no AI at that scale.","Firm profit is even more exposed than system reward: whenever reward declines with scale, profit declines faster because the firm pays per-project AI costs on an increasing number of projects.","Under-perception does not reverse scaling; it only slows the gains, and a moderate degree of under-perception can improve firm profit by partly offsetting the structural firm-worker misalignment.","Cost internalization is not one-size-fits-all: shifting AI costs to workers helps when AI is cheap or deployed at small scale, but firms should absorb more cost when deployment is expensive or large to preserve worker adoption.","Perception alignment reliably helps workers, but only helps firms under over-perception; aligning mildly under-perceiving workers can reduce firm profit."],"supporting_citations":[{"why":"Supplies the power-law scaling-law primitive: model error falls as a power of scale, which the model converts into AI failure e^{-αs}.","marker":"Kaplan et al. 2020"},{"why":"Empirical evidence that deep-learning generalization errors follow predictable power-law scaling, grounding the exponential failure form.","marker":"Hestness et al. 2017"},{"why":"Documents automation bias—over-reliance on automated systems—the behavioral basis for over-perception and effort withdrawal.","marker":"Parasuraman and Manzey 2010"},{"why":"Shows algorithm aversion after observed errors, the behavioral basis for under-perception.","marker":"Dietvorst et al. 2015"},{"why":"Shows consumers discount algorithmic advice on subjective tasks, reinforcing the under-perception channel.","marker":"Castelo et al. 2019"},{"why":"Identifies verification bias and selective feedback as reasons workers may never learn true AI capability, supporting the model's persistence assumption.","marker":"De Véricourt and Gurkan 2026"},{"why":"The selective-labels problem explains why outcome data can be missing for non-chosen actions, another reason misperception persists.","marker":"Lakkaraju et al. 2017"},{"why":"Field evidence that experienced developers overestimated AI time savings yet were slower with AI, providing an empirical anchor for the scaling paradox.","marker":"Becker et al. 2025"}],"fun_headline_variants":["Scaling AI can backfire if humans overestimate it","Human overconfidence can turn AI scaling into a net loss","The bigger the AI, the worse the team—if humans overrate it","Overperceiving AI turns scale into a liability"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that AI task failure falls exactly as e^{-αs} with no performance floor and that human and AI failures are independent; relax either and the non-monotone scaling paradox may disappear or change sign.","fun_headline_variants_meta":{"raw":{"variants":["Scaling AI can backfire if humans overestimate it","Human overconfidence can turn AI scaling into a net loss","The bigger the AI, the worse the team—if humans overrate it","Overperceiving AI turns scale into a liability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001266,"raw_usage":{"total_tokens":5058,"prompt_tokens":824,"completion_tokens":4234,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":568,"completion_tokens_details":{"reasoning_tokens":4164}},"tokens_in":568,"tokens_out":4234,"duration_ms":28383,"temperature":1.0,"reasoning_tokens":4164,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T00:11:55.420599+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled experiment with an AI assistant whose true success probability is 1 − e^{-αs}, while inducing over-perception by telling workers it is more capable than it is; measure realized reward across several scales in the human-in-the-loop region. If reward never decreases with scale and never dips below the human-only baseline, Proposition 2(iii) is false. Alternatively, measure the AI's true failure curve; if it has a nonzero floor or is correlated with human errors, recompute the model and check whether the paradox survives.","supporting_citations":[{"cited_title":"Experiment","cited_arxiv_id":null,"evidence_quote":"Shows algorithm aversion after observed errors, the behavioral basis for under-perception."}],"review_version":1}