REVIEW 4 major objections 6 minor 1 cited by
Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks
T0 review · 4 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Across thousands of real labor-market tasks, AI progress looks like a rising tide, not a crashing wave.
desk verdict The flat success–duration slope is probably an artifact of compressing long real-world tasks into short self-contained vignettes, but the underlying dataset and the parallel-shift finding are still worth engaging with seriously. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the estimated logistic curve Pr(success) = Λ(α + β log10 T), where T is human-reported task duration and success means a manager would accept the LLM output with no edits (rating ≥7 on a 1–9 scale). The pooled slope β ≈ −0.31 is the paper's primary evidence for a rising tide: a flat curve means AI performance barely differs between tasks that take minutes and tasks that take days. A micro-foundation casts each task as a chain of N coupled critical steps and model capability as a log-logistic random horizon, making β = −γ/σ, the ratio of how fast serial depth grows with duration to the dispersion of model failure shocks. The release-date term δ enters as a pure inte
What would settle it
Estimate the success–duration slope from the pilot subsample where task time was elicited before the LLM response was shown; a slope substantially below −0.31 (e.g., −0.5 or more negative) would indicate that the rising-tide flatness is an anchoring artifact.
Extended reading notes
Core claim
The central claim is that, across realistic, LLM-addressable labor-market tasks, the success–duration gradient for AI systems is shallow—a tenfold increase in human task time is associated with only about a 0.31 decline in log-odds of 'minimally sufficient without edits', equivalent to a roughly 7.6 percentage-point drop in success at the sample mean. This flat gradient, which holds across model sizes and vintages and across most job families, is the empirical signature of rising-tide automation. Newer model releases shift the curve upward in an approximately parallel fashion, whereas larger models improve short tasks more than long ones. The paper formalizes the slope as reflecting how quic
Load-bearing premise
The entire slope estimate rests on the assumption that evaluator-reported task durations are unbiased—specifically, that asking workers how long a task takes after they have already read the AI's response does not systematically distort their answers; if it does, the flat curve that defines the rising tide is an artifact.
Editorial extensions
If this is right
- If the flat success–duration curve is correct, automation will not arrive as sudden bursts that blindside workers; capability gains will show up on many tasks at once, giving some advance visibility.
- Frontier models already reach 50% success on tasks taking humans about three hours (2024-Q2) and may exceed 70% on most text tasks by 2025-Q3, meaning substantial automation potential is already here.
- At the current pace, the paper projects that most text-based tasks will reach 80–95% success at minimally sufficient quality by 2029–2030, while near-perfect success takes several more years.
- Failure rates for tasks from 5 minutes to 24 hours halve every 2.4–3.2 years, implying annual success-rate gains of 8–11 percentage points.
- Because newer models improve all durations equally but larger models mainly help short tasks, continued model releases will matter more for long-horizon tasks than simply scaling up existing models.
Reading between the lines
- The flat slope may partly reflect the authors' sampling choice: they only kept tasks where an LLM plausibly saves at least 10% of time, so tasks where LLMs are useless are absent; the rising-tide description therefore applies to LLM-addressable text work, not to all labor.
- The paper's forecasts assume the release-date trend in log-odds continues unchanged; any slowdown in compute scaling, algorithmic progress, or hardware gains would push the 2029–2030 projection later, as the authors themselves caution.
- A sharper test of the rising-tide claim would measure the slope on tasks that require external tool use or multi-turn interaction, which the current self-contained prompt design excludes; if that slope is much steeper, the flat curve is partly an artifact of task construction.
- The gap between benchmark-based 'time horizon' estimates (which show steep curves) and this survey's flat curve suggests the two are measuring different task distributions; reconciling them would require evaluating the same models on both deterministic coding tasks and open-ended white-collar tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper uses a large novel survey in which experienced workers evaluate LLM outputs on task instances derived from the O*NET taxonomy. The authors estimate the relationship between task success (manager acceptance without edits) and human task duration, finding a relatively flat success–duration slope (β ≈ −0.31 in logit-log10 duration space) which they interpret as evidence for 'rising tide' rather than 'crashing wave' automation. They further estimate a model-release-date trend that shifts the success–duration curve in an approximately parallel fashion, yielding doubling times of about 3.8 months and projections of 80–95% success on most text-based tasks by 2029. The paper includes multiple robustness checks, quality thresholds, job-family analyses, and a micro-founded interpretation of the logistic slope.
Significance. If the flat success–duration slope and parallel release-date shifts are credible, the paper provides an important correction to benchmark-based evidence (Kwa et al., 2025; METR, 2025) suggesting abrupt capability jumps on narrow task sets. The policy implications are substantial: gradual, broad-based automation would give workers and institutions more adjustment time than a 'crashing wave' scenario. The paper's strengths include a purpose-built evaluation design with domain-expert raters, multiple quality thresholds, and several thoughtful robustness analyses (pre/post duration timing, response-order randomization, complementary log-log specification, model-size/vintage decompositions). However, the central estimate is vulnerable to a construct-validity concern for long-duration tasks, and the manuscript contains unresolved sample-size inconsistencies. These issues are fixable but currently prevent full confidence in the headline forecasts.
major comments (4)
- [§4.1, Table 3, §2.1] The central slope estimate in Eq. (1) requires that the reported task duration T measures the difficulty of the task the LLM actually performs. For long-duration tasks this equivalence is broken: e.g., the '1 week' instance in Table 3 asks the model to 'Describe your program design and evaluation approach' for an Emerging Leaders program, not to perform the multi-step work (interviews, stakeholder meetings, budget reconciliation, iterative review) that would generate a 1-week human completion time. With prompts capped at ~150 words and responses at 700 words, long tasks are compressed into planning/writing tasks, which mechanically inflates success on long tasks and biases β toward zero. Section 3 acknowledges the self-contained constraint but does not quantify its effect on the slope. Please provide a concrete sensitivity analysis: use the Filter 3 'portion_represented' coverage score f
- [Abstract vs. §4.1, Table 2] The sample-size reporting is internally inconsistent. The abstract states 'more than 6,000 text-based, LLM-addressable tasks' and 'over 60,000 evaluations'; the full-text abstract states 'over 3,000 broad-based tasks' and 'more than 17,000 evaluations'; §4.1 reports 11,536 tasks and 69,216 task instances in the survey pool, while Table 2 reports N = 17,205 observations. Since five model responses were evaluated per instance, N = 17,205 implies roughly 3,441 task instances (about 1,720 tasks) in the analyzed sample, not 3,000+ or 11,000+. The discrepancy matters because Figure 7's 2029 forecasts are extrapolations from this subset. Please state precisely which sample underlies each table and figure, and reconcile the abstract numbers.
- [§4.1, §1, §3] The sample is restricted to O*NET tasks for which GPT-4 judged LLM assistance to yield at least 10% time savings. This conditioning may mechanically flatten the success–duration slope by excluding tasks where LLMs are weak, especially long, physical, or heavily interactive tasks. The paper is careful to limit its claims to 'text-based, LLM-addressable' tasks, but the title and introduction frame the conclusion as a general property of AI automation. Please provide sensitivity analyses (e.g., varying the 10% threshold, analyzing the excluded tasks qualitatively, or estimating a selection model) and quantify the share of all labor-market tasks to which the rising-tide conclusion applies.
- [§2.2, Eq. (2), Figure 6, Table A.4] The precision of the release-date trend and doubling-time estimates may be overstated. The regressions cluster standard errors by participant, but the key regressor R_m (release date) varies at the model level, and observations are also nested within task instances and O*NET tasks. Clustering only by participant ignores residual correlation at the task and model levels, which can narrow confidence intervals for the 3.8-month doubling time and the 2029 projections. Please report two-way clustering (participant and task instance) or a mixed model with random intercepts for tasks and models, and confirm that the headline precision survives.
minor comments (6)
- [§4.2] Typo: 'micro-founded rational' should be 'micro-founded rationale'.
- [Table A.5 note] The note says 'the R2 of a logistic regression is not informative as the independent variable varies between 0 and 1'; this likely refers to the dependent variable. Please correct the wording.
- [Appendix F, Table F.1] Several models appear in multiple categories (e.g., GPT-5 in 'Big' and 'Wild Cards — Thinking enabled'; o3 and o4-mini in 'Wild Cards' and in the frontier lists). Clarify whether these are distinct inference modes or accidental duplicates, and ensure the categorization is consistent with the 'large/small' splits in Figure 5.
- [§3, Kwa et al./METR comparison] The discussion of differences with Kwa et al. (2025) and METR (2025) would be strengthened by a small table comparing task selection, evaluation criteria, response formats, and success definitions. Currently the reasons for the contrasting slopes are asserted rather than systematically documented.
- [Figure 7] The many overlapping 'starting probability' curves are difficult to read. Consider labeling only a few representative horizons or presenting small multiples for the 10%, 50%, and 90% starting probabilities.
- [§4.1] Please state the exact data-collection cutoff date used for the analyses and, if possible, the number of unique participants whose evaluations remain after quality filtering. This would help readers map the 'ongoing survey' language to the reported N=17,205.
Circularity Check
No significant circularity: the empirical slope and trend are fitted from survey data, and the 2029 projections are explicit extrapolations rather than quantities built into the model.
full rationale
The paper's central empirical claim is the estimated logistic slope between task success and log human task duration, β ≈ −0.31 in Eq. (1), estimated by maximum likelihood from survey ratings and evaluator-reported durations. The 2029/2030 projections are produced by extrapolating the fitted release-date coefficient δ from Eq. (2) forward; they are not imposed as targets or recovered from the model by construction. The Section 4.2 micro-foundation (Eqs. 3–5) rewrites the same logit specification in terms of an unobserved serial-chain model, but the paper explicitly states this is 'only one possible (plausible) micro-foundation' and that 'Our results do not require this particular interpretation,' so the empirical finding does not reduce to the rationalization. The self-citations (Autor & Thompson 2025; Mertens et al. 2026; Fleming et al. 2024; Svanberg et al. 2024) are used for labor-market interpretation, compute-scaling background, or adoption costs, and none is load-bearing for the fitted slope or the extrapolation. There is no self-cited uniqueness theorem and no ansatz smuggled in via citation: the logistic link follows prior external work (Kwa et al. 2025; Ge et al. 2026). The limitations acknowledged in the paper—self-contained vignettes, post-evaluation duration reports, and the 10% time-savings filter—are genuine measurement/validity concerns, not circular derivations, and Appendix C directly tests the pre/post timing issue. Thus, despite some self-referential framing, the derivation chain is empirically self-contained: the inputs are ratings, durations, and model release dates, and the outputs are fitted parameters and extrapolations.
Assumptions & free parameters
free parameters (8)
- logit intercept alpha =
1.09 (>=7 threshold, pooled)
- logit duration slope beta =
-0.31 (pooled, >=7)
- release-date trend delta =
0.369 (years since Jan 2023, Table A.4)
- Task-inclusion time-savings threshold =
10%
- Quality threshold =
score >=7
- Model size threshold =
100B params
- Frontier-model list per quarter =
varies by quarter
- Micro-foundation parameters N0_d, gamma_d, sigma, eta_m, nu_d =
not separately identified
assumptions (8)
- domain assumption O*NET task statements are a valid representation of labor-market work
- ad hoc to paper GPT-4's >10% time-savings screening correctly identifies LLM-addressable tasks
- domain assumption Evaluator-reported task duration is an unbiased measure of task complexity
- domain assumption A score of >=7 ('minimally sufficient without edits') captures automation potential
- standard math Logistic functional form in Eq. (1) and linear release-date trend in Eq. (2)
- ad hoc to paper Serial-critical-path micro-foundation: success requires N = N0 T^gamma coupled steps and model robustness is log-logistic
- domain assumption Extrapolation assumes the 2024-2025 logit trend continues to 2029-2030
- domain assumption Survey participants with >=6 months experience provide representative expert evaluations
invented entities (1)
-
Serial critical action chain (N_d_js(T))
Cite this review
Pith. "Pith review of Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks." pith.science (2026). https://pith.science/paper/5C4IQ3PW
@misc{pith2026260401363,
author = {Pith},
title = {Pith review of: Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks},
year = {2026},
howpublished = {\url{https://pith.science/paper/5C4IQ3PW}},
note = {Machine review of arXiv:2604.01363}
}
read the original abstract
We characterize AI automation as a continuum between crashing waves, in which capabilities jump abruptly across narrow task sets, and rising tides, in which capabilities improve continuously and broadly. Using evidence from more than 6,000 text-based, LLM-addressable tasks derived from the U.S. Department of Labor's O*NET taxonomy and over 60,000 evaluations by experienced workers, we find little evidence of crashing waves (contrary to existing views). Instead, rising tides are the primary form of AI progress. AI performance is high and improving rapidly across many tasks. In 2024-Q2, models completed text-based tasks that take humans about 1.5 hours to complete with roughly 60% success, rising above 70% by 2025-Q3. If recent trends in AI capability growth persist, frontier LLMs will be able to complete most text-based tasks at minimally sufficient quality with 88%-97% success by 2030.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Affirmative AI-agent insurance with billion-scale limits is achievable by 2030 solely through coordinated industry build-out of an eight-component stack spanning data, CAT models, standards, contracts, underwriting, p...
Reference graph
Works this paper leans on
-
[1]
A detailed reasoning for your assessment
-
[2]
Task Scenario: {{scenario-based question}}
A final exposure label Structure your response as follows: **Basic LLM Exposure (LLME)** Reasoning: [Explain why this task would benefit or not benefit from direct text-only LLM interaction, considering text transformation, code writing, summarization, etc.] Label: [LLME0/LLME1/LLME2/LLME3/LLME4] 42 **LLM+ Tools Exposure (LLMTE)** Reasoning: [Explain how ...
-
[3]
Read the input prompt and identify what is being requested
-
[4]
Analyze if the output explicitly requires action outside of basic text generation (e.g., interacting with external systems, embodied activities, or non-textual media)
-
[5]
Provide detailed, step-by-step reasoning about these considerations
-
[6]
Output your final decision as ‘nontextual_output‘ (‘1‘ for impossible, ‘0‘ for possible), followed by reasoning, then a numeric confidence value as a percentage. 44
-
[7]
nontextual_output
Always deliver JSON fields in this strict order: "nontextual_output", "reasoning", then "confidence ". # Output Format - Output must be a JSON object with fields strictly ordered as: "nontextual_output", "reasoning", " confidence". - "nontextual_output": integer; 1 if impossible for a text LLM, 0 if possible. - "reasoning": string containing detailed, ste...
-
[8]
‘requires_extra_data‘: integer(1 if the example task refers to missing data; 0 if not)
Show all 15 references
-
[9]
‘reasoning‘: string (step-by-step justification)
-
[10]
‘confidence‘: string (likelihood that your evaluation is correct, as a percentage)
-
[11]
occupation
‘error‘: string (only provide an explanation if a required field is missing or ambiguous; otherwise blank) Do not output any commentary, explanation, or formatting outside the JSON object. # Error Handling If any required input is missing or ambiguous, output only the ‘error‘ ...
-
[12]
portion_represented: string (exactly one of the numeric ranges listed above)
-
[13]
reasoning: string (step-by-step justification, starting with your checklist and including explicit comparison of O*NET task requirements to the LLM prompt type and outputs)
-
[14]
confidence: string (how confident you are on a scale of 0-100% in your prediction)
-
[15]
") Only output the JSON object specified. Do not include any text, commentary, or formatting outside the JSON object. # Examples Example 1: Input: {
error: string (only include an error message if a required input is missing or ambiguous; otherwise , set to "") Only output the JSON object specified. Do not include any text, commentary, or formatting outside the JSON object. # Examples Example 1: Input: { "occupation": "Hom...
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.