Pith. sign in

REVIEW 4 major objections 5 minor 90 references

This paper argues that language models' economic value is best assessed task by task, and its simulation-based estimates indicate current chatbots could save substantial time on at least half the tasks in nearly half of U.S. occupations, wi

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 09:46 UTC pith:E6U735JH

load-bearing objection Useful benchmark infrastructure paper whose headline exposure claim rests on an unvalidated LLM self-simulation; the benchmark work deserves attention, the 47% number does not. the 4 major comments →

arxiv 2607.19375 v1 pith:E6U735JH submitted 2026-06-26 cs.CY cs.AI

Economic Evaluations of Language Models

classification cs.CY cs.AI
keywords language modelseconomic evaluationlabor market impactoccupational exposuretask-level benchmarkssynthetic datatime savingsAI adoption
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper replaces hand-waving about AI's labor impact with concrete, task-level measurement. It introduces EconEvals, an open evaluation suite with benchmark queries for every work activity and occupation in the U.S. government's occupational taxonomy, grounded in real user-chatbot conversations where possible and supplemented by synthetic roleplay queries. Alongside the benchmarks, it introduces a 'whitebox' simulation in which a language model roleplays a worker and itemizes, step by step, how much time a chatbot would save on each task. Using that simulation, the paper estimates that current language models could save substantial time on at least half of the tasks in 46.6% of occupations. It also finds that 79.4% of task opportunities judged highly time-saving see little real chatbot use, and that privacy and proprietary systems are the main fixable bottlenecks beyond physical or interactive limits.

Core claim

The paper's central claim is that the connection between language-model capability and labor-market value can be measured directly, task by task, rather than inferred from broad capability lists. It contributes an open evaluation suite with benchmark queries for all 2,087 work activities and all 1,016 occupations in the U.S. occupational taxonomy, built partly from real public user-chatbot conversations and partly from synthetic roleplay queries, at roughly a 500-fold lower query-generation cost than the existing worker-sourced benchmark that covered fewer than 5% of occupations. On the exposure side, it simulates a worker using a chatbot, decomposes each task into timed steps, aggregates sa

What carries the argument

The load-bearing mechanism is the whitebox simulation-based exposure measure. For each occupation and task, a language model roleplays a worker answering a labor economist's questions: first describing the job, then listing the steps of the task with baseline minutes per step, then considering a chatbot response to a synthetic query and marking which steps a chatbot would speed up and by how much. A post-processing filter removes steps that belong to other tasks; three simulation runs are median-averaged; and the aggregate savings are converted into exposure categories. This per-step accounting does two jobs at once: it produces the numeric time-savings estimates behind the 46.6% headline, a

Load-bearing premise

The simulation assumes that a language model roleplaying a worker can accurately estimate how long each task step really takes and how much time a chatbot would save, and no human time-study is used to check those estimates.

What would settle it

Run a controlled time study: have workers in a sample of high-exposure occupations perform their real tasks with and without a current chatbot, with independent time measurement. If the measured minutes saved come out substantially below the simulated step-level savings, or if the steps the simulation marks as automatable prove to require human judgment or integration overhead, then the 46.6% occupation estimate and the bottleneck rankings would be unsupported.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Because the benchmarks cover the full U.S. work taxonomy and cost roughly 500x less to build than worker-sourced ones, model rankings for economically relevant work could be refreshed as new models are released without repeated expensive data collection.
  • The usage-exposure gap implies that capability alone does not determine productivity impact: even if current models could save time on 79.4% of the tasks judged exposed, realized gains depend on complementary integration and adoption.
  • The bottleneck analysis suggests that fixing privacy and proprietary-system constraints, rather than waiting for better models, could unlock a large share of predicted time savings.
  • Synthetic benchmarks that predict worker-grounded scores with an average correlation of 0.67 can serve as a cheap screening tool for deciding which occupations deserve expensive human evaluation.
  • Simulation-based exposure being lower than rubric-based exposure implies that capability-list rubrics systematically overestimate how much current chatbots can help workers.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: If the simulated per-step time accounting is validated, the practical unit of AI planning shifts from whole occupations to individual steps, letting firms and workers target the steps where chatbots help most while leaving interactive or privacy-bound steps to humans.
  • Inference: The bottleneck taxonomy invites a direct experiment: in organizations that add chatbot access to private or proprietary systems (for example, internal data and compliance-approved workflows), the usage gap on currently exposed-but-unused tasks should measurably shrink, which would confirm the paper's causal story rather than just its correlation.
  • Inference: The same roleplay-and-itemize simulation could be run on other countries' job taxonomies or on an individual firm's task lists, producing localized exposure maps at low cost; the method is not tied to the U.S. taxonomy.
  • Inference: Because the simulation is produced by the very kind of model being evaluated, it may be systematically optimistic about step-level savings; comparing its estimates to randomized field measurements is the natural next test.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces EconEvals, an open-source evaluation infrastructure for measuring language-model capabilities on U.S. labor-economy tasks. It builds DWA-level benchmarks from real user–chatbot conversations (143 DWAs) and augments these with synthetically generated queries for a broader task space, reporting 226 benchmark evaluations in total, including 43 occupation-level benchmarks aligned with OpenAI's GDPval. The paper also proposes a 'whitebox' simulation-based exposure measure in which an LM roleplays a worker, decomposes each O*NET task into steps with baseline times, and estimates per-step time savings from current chatbot capabilities. The headline empirical claims are that current models could save substantial time on at least half of tasks in 46.6% of U.S. occupations, that 79.4% of predicted-high-exposure tasks show little current Claude usage, and that the synthetic occupation-level benchmarks predict GDPval scores with mean Spearman correlation 0.67.

Significance. If the claims hold, the paper would be a substantial contribution: it provides the first open benchmark suite attempting to cover the full O*NET taxonomy, a transparent per-task exposure accounting with reasoning traces, and a much cheaper synthetic-data alternative to human-sourced GDPval-style benchmarks. The paper's strengths include its detailed pipeline documentation, explicit cost reporting, a labeled precision estimate (0.91 on n=50) for the real-query mapping, correlation checks against GDPval, and the release of per-task exposure justifications. However, two load-bearing claims are currently unsupported: the assertion that benchmarks exist for all 2,087 DWAs / 1,016 occupations, and the validity of the simulation-based exposure headline. These need substantial revision or re-scoping before the central conclusions can be accepted.

major comments (4)
  1. [§2.2, §3, and Introduction] The Introduction states that the pipeline 'result[s] in benchmarks for all 2,087 DWAs, spanning all 1,016 occupations,' and the abstract implies full-coverage evaluation. But §3 reports results for only 226 benchmarks: 143 real DWA-level, 40 synthetic DWA-level, and 43 occupation-level. The other ~2,000 DWAs have synthetically generated queries, but no model scores are reported. The paper should clearly distinguish 'query generation coverage' from 'evaluated benchmark coverage' and revise the claims that ECONEVALS 'provides benchmarks for all 2,087 DWAs.' This is load-bearing because full coverage is a central advertised advantage over GDPval.
  2. [§4 and Appendix D.1] The exposure headline (46.6% of occupations with at least half of tasks substantially exposed) rests entirely on an unvalidated self-simulation. In Appendix D.1, the LM generates the worker background, the step decomposition, the baseline time per step, and the time saved per step; there is no human time-use study, no direct observation, and no calibration against measured time savings. Unlike the query-mapping pipeline, no precision or recall is reported for these time estimates. Figure 6 also excludes 159 of 1,016 occupations due to 'data generation errors,' so the 46.6% figure is computed on 857 occupations, not the full set. The paper needs either external validation (e.g., human time studies, comparison to actual task-completion times) or a clear reframing of these numbers as 'simulation-based potential' with the denominator limitation stated in the abstract and main text.
  3. [Appendix B and Appendix D.1] The synthetic data generation and the exposure simulation share the same conditioning on a time-savings ladder: the worker roleplay is asked to produce a prompt that saves a target percentage of time, and a verifier checks only realism and answerability, not the truth of the time-savings claim. The exposure simulation then uses that same synthetic prompt and a model response to estimate time savings. This creates a potential circularity: the LM is asked to estimate savings for a prompt it generated to justify a target savings level. The paper should provide a sensitivity analysis with independently constructed prompts, or demonstrate that estimated time savings are not inflated by the generation procedure.
  4. [§3 and Appendix C.1] The abstract's '500× lower cost' claim is based on a rough, lower-bound estimate. Appendix C.1 uses the GDPval mean review time (109 minutes) as a proxy for per-reviewer task-creation time, explicitly excludes drafting and editing time, and rates the estimate as low-to-medium confidence. The comparison also mixes LM-generated synthetic queries against human worker queries, which the paper itself acknowledges are higher quality along several dimensions. The cost factor should be qualified as a very rough estimate, not stated as a precise 500× advantage in the abstract and introduction.
minor comments (5)
  1. [Abstract / §4] The abstract reports '47% of occupations' while §4 reports 46.6%; the rounding is acceptable, but the paper should be consistent and should state the effective denominator (857 occupations) near the headline.
  2. [§2.1] The sentence 'We prioritize recall to improve task coverage' appears to contradict the immediately preceding reported recall of 0.18. This is likely a typo (perhaps 'coverage' was meant), but as written it is confusing.
  3. [Table 3] The table caption refers to 'lines in grey' that are filtered out, but in the provided text all rows appear identical. Please ensure the grey shading is visible in the final version.
  4. [§3] The text says 'We present results for 226 benchmarks' and then lists 143 + 40 + 43 = 226, but the synthetic DWA evaluation is only for 40 DWAs. The paper should explicitly note that the remaining DWAs have generated queries but no reported benchmark scores, to avoid the impression of full evaluation.
  5. [General] The paper promises open-source release but no repository or data URL is provided in the text. Please include a link or state the planned release venue.

Circularity Check

1 steps flagged

Exposure headline is anchored by construction: synthetic prompts are generated conditional on target time-savings, then the same LM estimates savings from those prompts.

specific steps
  1. fitted input called prediction [Appendix B (Synthetic User Queries) + §D.1 (Predicting exposure through simulation); feeds §4 headline (46.6%)]
    "Crucially, we condition in this prompt the amount of time-savings that prompt provides for this task. This lets us generate prompts with varying degrees of 'comprehensiveness': we can generate a complex prompt that performs the entire task if we ask the 'worker' to send a prompt with 90-100% time-savings; we can also generate a simpler prompt with 1% time-savings. ... Given a synthetically generated prompt for the work task (see Appendix B) and an LM response to the prompt, walks through the previously generated step-by-step breakdown and notes instances when a chatbot with the demonstrated ca"

    The exposure estimate is not an independent measurement of LM time-savings. The synthetic prompt given to the worker simulation is produced by Appendix B's loop that conditions on a specific time-savings target, starting at 90–100% and lowering the target only until a prompt passes the paper's own LM verifiers. The same simulation pipeline then asks an LM to estimate how much time that prompt would save and aggregates the result into the 46.6% occupation-level headline. Thus the predicted quantity is baked into the construction of the input prompt; the only checks are LM self-consistency (realism, answerability), not human time-use data. This is a fitted input called a prediction, not an out-of-sample estimate.

full rationale

The paper has two largely independent parts. The benchmark part is not circular: real-query benchmarks are precision-audited (0.91), synthetic occupation benchmarks are checked against GDPval (Spearman 0.67), and the cost/coverage claims are cost-accounting comparisons. The exposure part, however, contains a concrete construction loop. Appendix B generates the synthetic user prompt by conditioning on the amount of time-savings the prompt is supposed to provide, starting from 90–100% and backing off only as needed. §D.1 then feeds that same synthetic prompt into the worker-roleplay simulation and asks the LM to estimate the time-savings it would enable. The 46.6% headline thus summarizes simulated LMs rating prompts that were themselves generated to justify high savings; there is no human time-study or independent time-use data in the loop. This is a partial circularity: the quantity being predicted is inserted into the input-generation process. I do not count the Park et al. self-citation as circular, since it is an independent published result on LM simulation. The final score is 6 rather than higher because the exposure measure does produce lower estimates than the rubric baseline and is compared to external Claude usage, so the loop is not the only evidence in the paper.

Axiom & Free-Parameter Ledger

7 free parameters · 9 axioms · 0 invented entities

The core measurement depends on O*NET being an accurate map of US work, on public chatbot data being representative of work usage, on LLM classifiers/judges/verifiers being accurate, and on LLM self-simulation of worker time being valid. The exposure claim adds no external time-use data; thresholds (25%, 50%, 0.0025%, 50 instances) are hand-chosen and drive headline numbers.

free parameters (7)
  • Substantial time-savings threshold = 25% time saved
    A task is 'exposed' if predicted time savings is at least 25%; the headline '46.6% of occupations' depends on this hand-chosen cutoff.
  • High-exposure threshold = 50% time saved
    Used to define 'high exposure' in Figure 6; changing this threshold changes the occupation-level exposure counts.
  • Significant usage threshold = 0.0025% of Anthropic work-related Claude.ai traffic
    Tasks below this share are labeled 'minimal to no Claude usage'; this drives the 79.4% usage-lag claim.
  • DWA coverage threshold = 50 instances per DWA
    A DWA is 'covered' only if at least 50 queries are retrieved; this determines the 6.9% and 38.6% coverage numbers.
  • Retrieval design choices = top-200 DWAs; top-300 conversations per task; shallow retrieval k=10
    These choices set the precision/recall tradeoff (precision 0.91, recall 0.18) and the 300x cost savings; different choices would change benchmark coverage.
  • Simulation repetition count = 3 runs, median
    Exposure estimates are the median of three LLM simulations; no uncertainty intervals are reported.
  • Synthetic generation time-savings ladder = 90-100% down to 1%, backing off until verifier passes
    Controls the difficulty distribution of synthetic benchmark queries and indirectly affects all synthetic benchmark scores.
axioms (9)
  • domain assumption O*NET taxonomy is an accurate and complete codification of economically valuable US work.
    The entire benchmark and exposure framework is anchored to O*NET's 1,016 occupations, 2,087 DWAs, and 18,796 tasks (Section 1).
  • domain assumption Public chatbot conversations (WildChat, LMSYS, Chatbot Arena) are representative of real-world work-related LM usage.
    Used as the real-data grounding for benchmarks (Section 2.1); these are convenience samples from free chat services, not a random sample of US work.
  • domain assumption The LM-based retrieval and classification pipeline maps conversations to DWAs with sufficient accuracy.
    Precision is estimated on only 50 labeled pairs (0.91) and recall is extrapolated (0.18); mapping errors propagate into all real-data benchmarks (Section 2.1, Appendix A.2).
  • domain assumption A language-model judge can reliably rank model responses on economic work tasks.
    The paper uses LLM binary preference judgments with a reference model but provides no human validation of judge accuracy on these tasks (Section 3).
  • domain assumption Synthetic prompts generated by GPT-5-mini roleplay and verified by GPT-5.2 are realistic, answerable work queries.
    Initial simple prompts were rejected as low quality, so roleplay and verifier prompts were used; no human evaluation of synthetic prompt realism is reported (Appendix B).
  • domain assumption An LLM roleplaying a worker can produce accurate baseline task step times and time-savings, so the aggregate exposure reflects actual human time savings.
    This is the load-bearing premise for the headline exposure estimates; no human time-motion study or ground-truth comparison is provided (Appendix D.1).
  • domain assumption Anthropic's Claude Economic Index usage data is representative of actual economy-wide LM usage.
    Used to compare predicted exposure with observed usage; the data is proprietary and thresholds require at least 15 conversations or 5 unique accounts (Section 2.1, Section 4).
  • domain assumption GDPval scores are a valid external benchmark for economically valuable task performance.
    Used to validate synthetic occupation-level benchmarks; GDPval covers 44 worker-selected occupations and the paper removes one occupation from the comparison (Section 3).
  • standard math Standard statistical tools (Spearman correlation, median) are appropriate for the comparisons made.
    Correlation analyses between benchmark scores, GDPval, and capability benchmarks are standard rank-based comparisons (Section 3, Appendix C).

pith-pipeline@v1.3.0-alltime-deepseek · 26477 in / 17446 out tokens · 155944 ms · 2026-08-02T09:46:16.146802+00:00 · methodology

0 comments
read the original abstract

Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.

Figures

Figures reproduced from arXiv: 2607.19375 by Alexander Wan, Percy Liang, Rishi Bommasani, Stephane Hatgis-Kessell, Tom\'as Aguirre.

Figure 1
Figure 1. Figure 1: O*NET work taxonomy and whitebox simulation-based task-level exposure estimates. O*NET links occupations to DWAs and tasks as depicted for Software Developers (15-1252.00). Our method estimates task-level exposure across tasks using simulations and provides justifications. [2026] quantifies this distribution shift: existing benchmarks are concentrated on topics like math and coding (39.7% of benchmarking e… view at source ↗
Figure 2
Figure 2. Figure 2: Comparing economic benchmarks to raw public usage benchmarks. The left subfigure depicts model performance on 500 randomly sampled public user queries whereas the right depicts model performance across all our 143 50-instance DWA-level benchmarks. 2.2 Synthetic User Queries Public usage data has limited work coverage and even proprietary usage is fundamentally constrained to how AI is currently used. Howev… view at source ↗
Figure 3
Figure 3. Figure 3: Average proportion of moderate or highly exposed tasks versus average proportion of tasks with significant observed usage. We average the proportion of tasks that are exposed/have usage for each occupation, then average the occupation-level values for each major group. Across all major groups, both exposure predictions exceed actual Claude usage, indicating that AI could save time on far more tasks than it… view at source ↗
Figure 4
Figure 4. Figure 4: Prior blackbox rubric-based vs. our whitebox simulation-based exposure measures. The left side shows an excerpt from the blackbox rubric-based measure from prior work [Hos￾seini Maasoum and Lichtinger, 2026]. The right side shows our decomposition of the task into steps with step-wise accounting to determine the overall time-savings based on our worker simulation. recommendations, they would also need to h… view at source ↗
Figure 5
Figure 5. Figure 5: Simulation-based exposure estimates by occupational group and Claude usage. We plot the exposure estimate for each SOC major group aggregated across all tasks within the group. Results are stratified based on whether the underlying tasks have significant reported Claude usage (dark; at least 0.0025% of Anthropic’s work-related Claude.ai traffic) or not (light). 0.0 0.2 0.4 0.6 0.8 1.0 Share of tasks expose… view at source ↗
Figure 6
Figure 6. Figure 6: Simulation-based vs. rubric-based task-level exposure estimates. We plot the occupation-level exposure for our method and prior work as a function of the minimum proportion of tasks exposed. At both the high and moderate-to-high threshold, the simulation-based method predicts lower exposure as it surfaces bottlenecks to adoption that only arise from a more detailed accounting of AI-use. platform tools like… view at source ↗
Figure 7
Figure 7. Figure 7: Explaining time savings via reasoning trace analysis. The left side shows bottlenecks predicted by our simulation-based exposure predictions that limit time savings below 25%. The right side shows the predicted uses that enable time savings above 25%. for a substantial share of tasks that are labeled as not exposed by our simulation-based exposure measure. Rubric-based measures, which assume a fixed list o… view at source ↗
Figure 8
Figure 8. Figure 8: Results from analyzing the coverage of the retrieval pipeline. 21 [PITH_FULL_IMAGE:figures/full_fig_p021_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: DWA-level synthetic benchmark results. The top-half correspond to DWAs also found in retrieved data. The bottom-half correspond to DWAs not found in retrieved data. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Spearman correlation between leaderboards where the DWA is found in real usage versus not found in real usage. 23 [PITH_FULL_IMAGE:figures/full_fig_p023_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Variance in DWA-level synthetic benchmark results. We rank DWAs by the variance they induce in model performance for the 40 DWA-level synthetic benchmarks. 24 [PITH_FULL_IMAGE:figures/full_fig_p024_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: DWA-level synthetic benchmark results compared to real benchmark results. 25 [PITH_FULL_IMAGE:figures/full_fig_p025_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Occupation-level synthetic benchmark results compared to GDPval benchmark results. 26 [PITH_FULL_IMAGE:figures/full_fig_p026_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Correlation between synthetic benchmark results and GDPval benchmark results versus generic capabilities benchmarks. 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Proportion Task has high stakes Task requires real-time monitoring Task requires proprietary systems Chatbot consultation has overhead Task requires synchronous interaction Task requires visual judgment Chatbot outputs need verification Task involves live … view at source ↗
Figure 15
Figure 15. Figure 15: The factors that result in the simulation-based exposure measure to predict substantial/non-substantial time-savings when Claude usage data suggests the opposite. 1-2 3 4 5 O*NET Job Zone 0.0 0.2 0.4 0.6 0.8 1.0 Mean fraction of tasks exposed to AI over occupations in zone Simulation-Based-Exposure (moderate + high exposure tasks) Simulation-Based-Exposure (high exposure tasks only) Rubric-Based-Exposure … view at source ↗
Figure 16
Figure 16. Figure 16: Predicted exposure by O*NET Job Zone, under simulation-based and rubric-based exposure measures. Job Zones group occupations by the education, experience, and training required for entry, ranging from Zone 1–2 (under one year of preparation) to Zone 5 (4+ years). Both measures exhibit the same pattern: exposure rises from Job Zone 1 through Job Zone 4 and then plateaus or declines at Job Zone 5. 27 [PITH… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

90 extracted references · 1 canonical work pages

  1. [1]

    Generative

    Gmyrek, Pawe. Generative. 2025 , publisher=

  2. [2]

    Science , volume =

    Tyna Eloundou and Sam Manning and Pamela Mishkin and Daniel Rock , title =. Science , volume =. 2024 , doi =

  3. [3]

    Artificial Intelligence and the Great Divergence , year =

  4. [4]

    and Karger, Ezra , title =

    Murphy, Connacher and Rosenberg, Josh and Canedy, Jordan and Jacobs, Zach and Flechner, Nadja and Britt, Rhiannon and Pan, Alexa and Rogers-Smith, Charlie and Mayland, Dan and Buffington, Cathy and Kučinskas, Simas and Coston, Amanda and Kerner, Hannah and Pierson, Emma and Rabbany, Reihaneh and Salganik, Matthew and Seamans, Robert and Su, Yu and Tramèr,...

  5. [5]

    Tetlock , title =

    Ezra Karger and Otto Kuusela and Jason Abaluck and Kevin Bryan and Basil Halperin and Todd Jones and Connacher Murphy and Phil Trammell and Matt Reynolds and Dan Mayland and Ria Viswanathan and Ananaya Mittal and Rebecca Ceppas de Castro and Josh Rosenberg and Philip E. Tetlock , title =. 2026 , month = mar, url =

  6. [6]

    2026 , eprint=

    How Well Does Agent Development Reflect Real-World Work? , author=. 2026 , eprint=

  7. [7]

    Task-Completion Time Horizons of Frontier

  8. [8]

    Ziegler and Elizabeth Barnes and Lawrence Chan , booktitle =

    Thomas Kwa and Ben West and Joel Becker and Amy Deng and Katharyn Garcia and Max Hasin and Sami Jawhar and Megan Kinniment and Nate Rush and Sydney Von Arx and Ryan Bloom and Thomas Broadley and Haoxing Du and Brian Goodrich and Nikola Jurkovic and Luke Harold Miles and Seraphina Nix and Tao Lin and Neev Parikh and David Rein and Lucas Jun Koba Sato and H...

  9. [12]

    2025 , month = apr, day =

    Arvind Narayanan and Sayash Kapoor , institution =. 2025 , month = apr, day =

  10. [13]

    Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation , author =

  11. [14]

    Towards a Science of

    Stephan Rabanser and Sayash Kapoor and Peter Kirgis and Kangheng Liu and Saiteja Utpala and Arvind Narayanan , journal =. Towards a Science of

  12. [15]

    2024 , month = may, doi =

    Daron Acemoglu , title =. 2024 , month = may, doi =

  13. [16]

    2026 , month = jan, url =

    Dario Amodei , title =. 2026 , month = jan, url =

  14. [17]

    2025 , month = oct, day =

    Bharat Chandar , title =. 2025 , month = oct, day =

  15. [18]

    2026 , month = jan, day =

    Alex Imas , title =. 2026 , month = jan, day =

  16. [19]

    Ipsos Predictions Survey 2025: Positivity about how this year has gone highest since before the pandemic , year =

  17. [20]

    2025 , url =

    Human Development Report 2025: A Matter of Choice: People and Possibilities in the Age of AI , institution =. 2025 , url =

  18. [21]

    2025 , url =

    AI Index Report 2025: Public Opinion , institution =. 2025 , url =

  19. [22]

    and Raj, Manav and Seamans, Robert , title =

    Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. AEA Papers and Proceedings , volume =. 2018 , doi =

  20. [23]

    and Raj, Manav and Seamans, Robert , title =

    Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. Strategic Management Journal , volume =. 2021 , doi =

  21. [24]

    2020 , doi =

    Webb, Michael , title =. 2020 , doi =

  22. [25]

    Measuring the Occupational Impact of

    Tolan, Song. Measuring the Occupational Impact of. Journal of Artificial Intelligence Research , volume =. 2021 , doi =

  23. [26]

    and Raj, Manav and Seamans, Robert , title =

    Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. SSRN Electronic Journal , year =

  24. [27]

    and Cazzaniga, Mauro and Li, Longji , title =

    Pizzinelli, Carlo and Panton, Augustus and Tavares, Marina M. and Cazzaniga, Mauro and Li, Longji , title =

  25. [28]

    Journal of Economic Perspectives , volume =

    Acemoglu, Daron and Restrepo, Pascual , title =. Journal of Economic Perspectives , volume =. 2019 , doi =

  26. [29]

    Generative

    Gmyrek, Pawel and Berg, Janine and Kami. Generative. 2025 , doi =

  27. [30]

    Mirror, Mirror on the Wall: Which Jobs Will AI Replace After All? A New Index of Occupational Exposure , institution =

    Ben. Mirror, Mirror on the Wall: Which Jobs Will AI Replace After All? A New Index of Occupational Exposure , institution =

  28. [31]

    2026 , url =

    Massenkoff, Maxim and McCrory, Peter , title =. 2026 , url =

  29. [32]

    2026 , month = apr, url =

    Sha Sajadieh and Loredana Fattorini and Raymond Perrault and Yolanda Gil and Vanessa Parli and Lapo Santarlasci and Juan Pava and Nestor Maslej and Russ Altman and Erik Brynjolfsson and Carla Brodley and Jack Clark and Virginia Dignum and Vipin Kumar and James Landay and Terah Lyons and James Manyika and Juan Carlos Niebles and Yoav Shoham and Elham Tabas...

  30. [34]

    Yoshua Bengio and Stephen Clare and Carina Prunkl and Maksym Andriushchenko and Ben Bucknall and Malcolm Murray and Rishi Bommasani and Stephen Casper and Tom Davidson and Raymond Douglas and David Duvenaud and Philip Fox and Usman Gohar and Rose Hadshar and Anson Ho and Tiancheng Hu and Cameron Jones and Sayash Kapoor and Atoosa Kasirzadeh and Sam Mannin...

  31. [36]

    2026 , month = jan, howpublished =

    Thomas Kwa , title =. 2026 , month = jan, howpublished =

  32. [37]

    Generative

    Hosseini Maasoum, Seyed Mahdi and Lichtinger, Guy , year =. Generative

  33. [38]

    and Levy, Frank and Murnane, Richard J

    Autor, David H. and Levy, Frank and Murnane, Richard J. , title =. The Quarterly Journal of Economics , volume =. 2003 , month = nov, doi =

  34. [39]

    Handbook of Labor Economics , editor =

    Acemoglu, Daron and Autor, David , title =. Handbook of Labor Economics , editor =. 2011 , doi =

  35. [40]

    and Raj, Manav and Seamans, Robert , title =

    Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. arXiv preprint arXiv:2303.01157 , year =. 2303.01157 , archivePrefix =

  36. [41]

    AEA Papers and Proceedings , volume =

    Brynjolfsson, Erik and Mitchell, Tom and Rock, Daniel , title =. AEA Papers and Proceedings , volume =. 2018 , doi =

  37. [43]

    arXiv preprint arXiv:2507.07935 , year =

    Tomlinson, Kiran and Jaffe, Sonia and Wang, Will and Counts, Scott and Suri, Siddharth , title =. arXiv preprint arXiv:2507.07935 , year =. 2507.07935 , archivePrefix =

  38. [44]

    and Kamphorst, Jonne and Egan, Niles and Shaw, Aaron and Hill, Benjamin Mako and Cai, Carrie and Morris, Meredith Ringel and Liang, Percy and Willer, Robb and Bernstein, Michael S

    Park, Joon Sung and Zou, Carolyn Q. and Kamphorst, Jonne and Egan, Niles and Shaw, Aaron and Hill, Benjamin Mako and Cai, Carrie and Morris, Meredith Ringel and Liang, Percy and Willer, Robb and Bernstein, Michael S. , title =. arXiv preprint arXiv:2411.10109 , year =. 2411.10109 , archivePrefix =

  39. [45]

    and Cai, Carrie J

    Park, Joon Sung and O'Brien, Joseph C. and Cai, Carrie J. and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual. 2023 , doi =

  40. [46]

    and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =

    Chatterji, Aaron and Cunningham, Thomas and Deming, David J. and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =. 2025 , month = sep, doi =

  41. [47]

    International Conference on Learning Representations (

    Zhao, Wenting and Ren, Xiang and Hessel, Jack and Cardie, Claire and Choi, Yejin and Deng, Yuntian , title =. International Conference on Learning Representations (. 2024 , eprint =

  42. [49]

    and Stoica, Ion , title =

    Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos, Anastasios Nikolas and Li, Tianle and Li, Dacheng and Zhang, Hao and Zhu, Banghua and Jordan, Michael and Gonzalez, Joseph E. and Stoica, Ion , title =. 2024 , eprint =

  43. [51]

    How Adaptable Are American Workers to

    Sam Manning and Tom\'. How Adaptable Are American Workers to. 2026 , month = jan, url =

  44. [52]

    2025 , month = jun, url =

    David Autor and Neil Thompson , title =. 2025 , month = jun, url =

  45. [56]

    Ruth Appel and Maxim Massenkoff and Peter McCrory and Miles McCain and Ryan Heller and Tyler Neylon and Alex Tamkin , title =

  46. [57]

    2026 , eprint=

    Agents' Last Exam , author=. 2026 , eprint=

  47. [58]

    Ipsos predictions survey 2025: Positivity about how this year has gone highest since before the pandemic, December 2024

    Ipsos . Ipsos predictions survey 2025: Positivity about how this year has gone highest since before the pandemic, December 2024. URL https://www.ipsos.com/en-us/ipsos-predictions-2025. Reports that globally 65\

  48. [59]

    Human development report 2025: A matter of choice: People and possibilities in the age of ai

    United Nations . Human development report 2025: A matter of choice: People and possibilities in the age of ai. Technical report, United Nations Development Programme, 2025. URL https://hdr.undp.org/system/files/documents/global-report-document/hdr2025reporten.pdf. Uses the 2025 Global Survey on AI and Human Development; reports that 57\

  49. [60]

    Ai index report 2025: Public opinion

    Stanford HAI . Ai index report 2025: Public opinion. Technical report, Stanford HAI, 2025. URL https://hai.stanford.edu/ai-index/2025-ai-index-report/public-opinion. Reports that globally 60\

  50. [61]

    Artificial intelligence and the great divergence, January 2026

    The White House . Artificial intelligence and the great divergence, January 2026. URL https://www.whitehouse.gov/research/2026/01/artificial-intelligence-and-the-great-divergence/. Accessed: 2026-04-09

  51. [62]

    Tetlock, and Ezra Karger

    Connacher Murphy, Josh Rosenberg, Jordan Canedy, Zach Jacobs, Nadja Flechner, Rhiannon Britt, Alexa Pan, Charlie Rogers-Smith, Dan Mayland, Cathy Buffington, Simas Kučinskas, Amanda Coston, Hannah Kerner, Emma Pierson, Reihaneh Rabbany, Matthew Salganik, Robert Seamans, Yu Su, Florian Tramèr, Tatsunori Hashimoto, Arvind Narayanan, Philip E. Tetlock, and E...

  52. [63]

    Ezra Karger, Otto Kuusela, Jason Abaluck, Kevin Bryan, Basil Halperin, Todd Jones, Connacher Murphy, Phil Trammell, Matt Reynolds, Dan Mayland, Ria Viswanathan, Ananaya Mittal, Rebecca Ceppas de Castro, Josh Rosenberg, and Philip E. Tetlock. Forecasting the economic effects of AI , March 2026. URL https://static1.squarespace.com/static/635693acf15a3e2a14a...

  53. [64]

    The simple macroeconomics of AI

    Daron Acemoglu. The simple macroeconomics of AI . Working Paper 32487, National Bureau of Economic Research, May 2024. URL https://www.nber.org/papers/w32487

  54. [65]

    The adolescence of technology, January 2026

    Dario Amodei. The adolescence of technology, January 2026. URL https://www.darioamodei.com/essay/the-adolescence-of-technology. Accessed: 2026-04-09

  55. [66]

    AI and labor markets: What we know and don't know, October 2025

    Bharat Chandar. AI and labor markets: What we know and don't know, October 2025. URL https://digitaleconomy.stanford.edu/news/ai-and-labor-markets-what-we-know-and-dont-know/. Accessed: 2026-04-09

  56. [67]

    What is the impact of AI on productivity? Ghosts of Electricity (Substack), January 2026

    Alex Imas. What is the impact of AI on productivity? Ghosts of Electricity (Substack), January 2026. URL https://aleximas.substack.com/p/what-is-the-impact-of-ai-on-productivity. Accessed: 2026-04-09

  57. [68]

    Maria del Rio-Chanona, Ekkehard Ernst, Rossana Merola, Daniel Samaan, and Ole Teutloff

    R. Maria del Rio-Chanona, Ekkehard Ernst, Rossana Merola, Daniel Samaan, and Ole Teutloff. AI and jobs. A review of theory, estimates, and evidence, 2025. URL https://arxiv.org/abs/2509.15265

  58. [69]

    Yoshua Bengio, Stephen Clare, Carina Prunkl, Maksym Andriushchenko, Ben Bucknall, Malcolm Murray, Rishi Bommasani, Stephen Casper, Tom Davidson, Raymond Douglas, David Duvenaud, Philip Fox, Usman Gohar, Rose Hadshar, Anson Ho, Tiancheng Hu, Cameron Jones, Sayash Kapoor, Atoosa Kasirzadeh, Sam Manning, Nestor Maslej, Vasilios Mavroudis, Conor McGlynn, Rich...

  59. [70]

    Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo...

  60. [71]

    Artificial intelligence index report 2026

    Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, and Dan Weld. Artificial in...

  61. [72]

    Ziegler, Elizabeth Barnes, and Lawrence Chan

    Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, and Lawrence ...

  62. [73]

    Task-completion time horizons of frontier AI models

    METR . Task-completion time horizons of frontier AI models. https://metr.org/time-horizons/, March 2026

  63. [74]

    Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek

    Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Sim \'o n Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek. GDPval : Evaluating AI model performance on real...

  64. [75]

    Remote labor index: Measuring AI automation of remote work

    Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, Vinaya Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevi...

  65. [76]

    Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski

    Bertie Vidgen, Abby Fennelly, Evan Pinnix, Julien Benchek, Daniyal Khan, Zach Richards, Austin Bridges, Calix Huang, Kanishka Sahu, Abhishek Kottamasu, Bo Ma, Ben Hunsberger, Isaac Robinson, Akul Datta, Chirag Mahapatra, Dominic Barton, Cass R. Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski. The AI P roductivity I ndex ( APEX ), 2025. URL https://...

  66. [77]

    Agents' last exam, 2026

    Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, et al. Agents' last exam, 2026. URL https://arxiv.org/abs/2606.05405

  67. [78]

    How well does agent development reflect real-world work?, 2026

    Zora Zhiruo Wang, Sanidhya Vijayvargiya, Aspen Chen, Hanmo Zhang, Venu Arvind Arangarajan, Jett Chen, Valerie Chen, Diyi Yang, Daniel Fried, and Graham Neubig. How well does agent development reflect real-world work?, 2026. URL https://arxiv.org/abs/2603.01203

  68. [79]

    AI as normal technology: An alternative to the vision of AI as a potential superintelligence

    Arvind Narayanan and Sayash Kapoor. AI as normal technology: An alternative to the vision of AI as a potential superintelligence. Technical report, Knight First Amendment Institute, April 2025. URL https://knightcolumbia.org/content/ai-as-normal-technology

  69. [80]

    Shorrer, and Yannai A

    Sara Fish, Julia Shephard, Minkai Li, Ran I. Shorrer, and Yannai A. Gonczarowski. EconEvals : Benchmarks and litmus tests for economic decision-making by llm agents, 2026. URL https://arxiv.org/abs/2503.18825

  70. [81]

    WildChat : 1 M ChatGPT interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat : 1 M ChatGPT interaction logs in the wild. In International Conference on Learning Representations ( ICLR ) , 2024. URL https://arxiv.org/abs/2405.01470

  71. [82]

    Xing, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. LMSYS-Chat-1M : A large-scale real-world LLM conversation dataset. arXiv preprint arXiv:2309.11998, 2023. URL https://arxiv.org/abs/2309.11998

  72. [83]

    Anthropic economic index report: economic primitives, 2026

    Ruth Appel, Maxim Massenkoff, Peter McCrory, Miles McCain, Ryan Heller, Tyler Neylon, and Alex Tamkin. Anthropic economic index report: economic primitives, 2026

  73. [84]

    Autor, Frank Levy, and Richard J

    David H. Autor, Frank Levy, and Richard J. Murnane. The skill content of recent technological change: An empirical exploration. The Quarterly Journal of Economics, 118 0 (4): 0 1279--1333, November 2003. doi:10.1162/003355303322552801. URL https://academic.oup.com/qje/article-abstract/118/4/1279/1925105

  74. [85]

    GPTs are GPTs : Labor market impact potential of LLMs

    Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. GPTs are GPTs : Labor market impact potential of LLMs . Science, 384 0 (6702): 0 1306--1308, 2024. doi:10.1126/science.adj0998. URL https://www.science.org/doi/abs/10.1126/science.adj0998

  75. [86]

    O'Brien, Carrie J

    Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology ( UIST '23) , 2023. doi:10.1145/3586183.3606763. URL https://dl.acm.org/doi/10.1145/3586183.3606763

  76. [87]

    Gonzalez, and Ion Stoica

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference, 2024. URL https://arxiv.org/abs/2403.04132

  77. [88]

    The Faiss library

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazar \'e , Maria Lomeli, Lucas Hosseini, and Herv \'e J \'e gou. The Faiss library. arXiv preprint arXiv:2401.08281, 2024. URL https://arxiv.org/abs/2401.08281

  78. [89]

    Gonzalez, and Ion Stoica

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: A rena- H ard and B ench B uilder pipeline, 2024. URL https://arxiv.org/abs/2406.11939

  79. [90]

    A rosetta stone for AI benchmarks, 2025

    Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. A rosetta stone for AI benchmarks, 2025. URL https://arxiv.org/abs/2512.00193

  80. [91]

    Skills, tasks and technologies: Implications for employment and earnings

    Daron Acemoglu and David Autor. Skills, tasks and technologies: Implications for employment and earnings. In Orley Ashenfelter and David Card, editors, Handbook of Labor Economics, volume 4, chapter 12, pages 1043--1171. Elsevier, Amsterdam, 2011. doi:10.1016/S0169-7218(11)02410-5. Part B

Showing first 80 references.