REVIEW 4 major objections 5 minor 90 references
This paper argues that language models' economic value is best assessed task by task, and its simulation-based estimates indicate current chatbots could save substantial time on at least half the tasks in nearly half of U.S. occupations, wi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 09:46 UTC pith:E6U735JH
load-bearing objection Useful benchmark infrastructure paper whose headline exposure claim rests on an unvalidated LLM self-simulation; the benchmark work deserves attention, the 47% number does not. the 4 major comments →
Economic Evaluations of Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the connection between language-model capability and labor-market value can be measured directly, task by task, rather than inferred from broad capability lists. It contributes an open evaluation suite with benchmark queries for all 2,087 work activities and all 1,016 occupations in the U.S. occupational taxonomy, built partly from real public user-chatbot conversations and partly from synthetic roleplay queries, at roughly a 500-fold lower query-generation cost than the existing worker-sourced benchmark that covered fewer than 5% of occupations. On the exposure side, it simulates a worker using a chatbot, decomposes each task into timed steps, aggregates sa
What carries the argument
The load-bearing mechanism is the whitebox simulation-based exposure measure. For each occupation and task, a language model roleplays a worker answering a labor economist's questions: first describing the job, then listing the steps of the task with baseline minutes per step, then considering a chatbot response to a synthetic query and marking which steps a chatbot would speed up and by how much. A post-processing filter removes steps that belong to other tasks; three simulation runs are median-averaged; and the aggregate savings are converted into exposure categories. This per-step accounting does two jobs at once: it produces the numeric time-savings estimates behind the 46.6% headline, a
Load-bearing premise
The simulation assumes that a language model roleplaying a worker can accurately estimate how long each task step really takes and how much time a chatbot would save, and no human time-study is used to check those estimates.
What would settle it
Run a controlled time study: have workers in a sample of high-exposure occupations perform their real tasks with and without a current chatbot, with independent time measurement. If the measured minutes saved come out substantially below the simulated step-level savings, or if the steps the simulation marks as automatable prove to require human judgment or integration overhead, then the 46.6% occupation estimate and the bottleneck rankings would be unsupported.
If this is right
- Because the benchmarks cover the full U.S. work taxonomy and cost roughly 500x less to build than worker-sourced ones, model rankings for economically relevant work could be refreshed as new models are released without repeated expensive data collection.
- The usage-exposure gap implies that capability alone does not determine productivity impact: even if current models could save time on 79.4% of the tasks judged exposed, realized gains depend on complementary integration and adoption.
- The bottleneck analysis suggests that fixing privacy and proprietary-system constraints, rather than waiting for better models, could unlock a large share of predicted time savings.
- Synthetic benchmarks that predict worker-grounded scores with an average correlation of 0.67 can serve as a cheap screening tool for deciding which occupations deserve expensive human evaluation.
- Simulation-based exposure being lower than rubric-based exposure implies that capability-list rubrics systematically overestimate how much current chatbots can help workers.
Where Pith is reading between the lines
- Inference: If the simulated per-step time accounting is validated, the practical unit of AI planning shifts from whole occupations to individual steps, letting firms and workers target the steps where chatbots help most while leaving interactive or privacy-bound steps to humans.
- Inference: The bottleneck taxonomy invites a direct experiment: in organizations that add chatbot access to private or proprietary systems (for example, internal data and compliance-approved workflows), the usage gap on currently exposed-but-unused tasks should measurably shrink, which would confirm the paper's causal story rather than just its correlation.
- Inference: The same roleplay-and-itemize simulation could be run on other countries' job taxonomies or on an individual firm's task lists, producing localized exposure maps at low cost; the method is not tied to the U.S. taxonomy.
- Inference: Because the simulation is produced by the very kind of model being evaluated, it may be systematically optimistic about step-level savings; comparing its estimates to randomized field measurements is the natural next test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EconEvals, an open-source evaluation infrastructure for measuring language-model capabilities on U.S. labor-economy tasks. It builds DWA-level benchmarks from real user–chatbot conversations (143 DWAs) and augments these with synthetically generated queries for a broader task space, reporting 226 benchmark evaluations in total, including 43 occupation-level benchmarks aligned with OpenAI's GDPval. The paper also proposes a 'whitebox' simulation-based exposure measure in which an LM roleplays a worker, decomposes each O*NET task into steps with baseline times, and estimates per-step time savings from current chatbot capabilities. The headline empirical claims are that current models could save substantial time on at least half of tasks in 46.6% of U.S. occupations, that 79.4% of predicted-high-exposure tasks show little current Claude usage, and that the synthetic occupation-level benchmarks predict GDPval scores with mean Spearman correlation 0.67.
Significance. If the claims hold, the paper would be a substantial contribution: it provides the first open benchmark suite attempting to cover the full O*NET taxonomy, a transparent per-task exposure accounting with reasoning traces, and a much cheaper synthetic-data alternative to human-sourced GDPval-style benchmarks. The paper's strengths include its detailed pipeline documentation, explicit cost reporting, a labeled precision estimate (0.91 on n=50) for the real-query mapping, correlation checks against GDPval, and the release of per-task exposure justifications. However, two load-bearing claims are currently unsupported: the assertion that benchmarks exist for all 2,087 DWAs / 1,016 occupations, and the validity of the simulation-based exposure headline. These need substantial revision or re-scoping before the central conclusions can be accepted.
major comments (4)
- [§2.2, §3, and Introduction] The Introduction states that the pipeline 'result[s] in benchmarks for all 2,087 DWAs, spanning all 1,016 occupations,' and the abstract implies full-coverage evaluation. But §3 reports results for only 226 benchmarks: 143 real DWA-level, 40 synthetic DWA-level, and 43 occupation-level. The other ~2,000 DWAs have synthetically generated queries, but no model scores are reported. The paper should clearly distinguish 'query generation coverage' from 'evaluated benchmark coverage' and revise the claims that ECONEVALS 'provides benchmarks for all 2,087 DWAs.' This is load-bearing because full coverage is a central advertised advantage over GDPval.
- [§4 and Appendix D.1] The exposure headline (46.6% of occupations with at least half of tasks substantially exposed) rests entirely on an unvalidated self-simulation. In Appendix D.1, the LM generates the worker background, the step decomposition, the baseline time per step, and the time saved per step; there is no human time-use study, no direct observation, and no calibration against measured time savings. Unlike the query-mapping pipeline, no precision or recall is reported for these time estimates. Figure 6 also excludes 159 of 1,016 occupations due to 'data generation errors,' so the 46.6% figure is computed on 857 occupations, not the full set. The paper needs either external validation (e.g., human time studies, comparison to actual task-completion times) or a clear reframing of these numbers as 'simulation-based potential' with the denominator limitation stated in the abstract and main text.
- [Appendix B and Appendix D.1] The synthetic data generation and the exposure simulation share the same conditioning on a time-savings ladder: the worker roleplay is asked to produce a prompt that saves a target percentage of time, and a verifier checks only realism and answerability, not the truth of the time-savings claim. The exposure simulation then uses that same synthetic prompt and a model response to estimate time savings. This creates a potential circularity: the LM is asked to estimate savings for a prompt it generated to justify a target savings level. The paper should provide a sensitivity analysis with independently constructed prompts, or demonstrate that estimated time savings are not inflated by the generation procedure.
- [§3 and Appendix C.1] The abstract's '500× lower cost' claim is based on a rough, lower-bound estimate. Appendix C.1 uses the GDPval mean review time (109 minutes) as a proxy for per-reviewer task-creation time, explicitly excludes drafting and editing time, and rates the estimate as low-to-medium confidence. The comparison also mixes LM-generated synthetic queries against human worker queries, which the paper itself acknowledges are higher quality along several dimensions. The cost factor should be qualified as a very rough estimate, not stated as a precise 500× advantage in the abstract and introduction.
minor comments (5)
- [Abstract / §4] The abstract reports '47% of occupations' while §4 reports 46.6%; the rounding is acceptable, but the paper should be consistent and should state the effective denominator (857 occupations) near the headline.
- [§2.1] The sentence 'We prioritize recall to improve task coverage' appears to contradict the immediately preceding reported recall of 0.18. This is likely a typo (perhaps 'coverage' was meant), but as written it is confusing.
- [Table 3] The table caption refers to 'lines in grey' that are filtered out, but in the provided text all rows appear identical. Please ensure the grey shading is visible in the final version.
- [§3] The text says 'We present results for 226 benchmarks' and then lists 143 + 40 + 43 = 226, but the synthetic DWA evaluation is only for 40 DWAs. The paper should explicitly note that the remaining DWAs have generated queries but no reported benchmark scores, to avoid the impression of full evaluation.
- [General] The paper promises open-source release but no repository or data URL is provided in the text. Please include a link or state the planned release venue.
Circularity Check
Exposure headline is anchored by construction: synthetic prompts are generated conditional on target time-savings, then the same LM estimates savings from those prompts.
specific steps
-
fitted input called prediction
[Appendix B (Synthetic User Queries) + §D.1 (Predicting exposure through simulation); feeds §4 headline (46.6%)]
"Crucially, we condition in this prompt the amount of time-savings that prompt provides for this task. This lets us generate prompts with varying degrees of 'comprehensiveness': we can generate a complex prompt that performs the entire task if we ask the 'worker' to send a prompt with 90-100% time-savings; we can also generate a simpler prompt with 1% time-savings. ... Given a synthetically generated prompt for the work task (see Appendix B) and an LM response to the prompt, walks through the previously generated step-by-step breakdown and notes instances when a chatbot with the demonstrated ca"
The exposure estimate is not an independent measurement of LM time-savings. The synthetic prompt given to the worker simulation is produced by Appendix B's loop that conditions on a specific time-savings target, starting at 90–100% and lowering the target only until a prompt passes the paper's own LM verifiers. The same simulation pipeline then asks an LM to estimate how much time that prompt would save and aggregates the result into the 46.6% occupation-level headline. Thus the predicted quantity is baked into the construction of the input prompt; the only checks are LM self-consistency (realism, answerability), not human time-use data. This is a fitted input called a prediction, not an out-of-sample estimate.
full rationale
The paper has two largely independent parts. The benchmark part is not circular: real-query benchmarks are precision-audited (0.91), synthetic occupation benchmarks are checked against GDPval (Spearman 0.67), and the cost/coverage claims are cost-accounting comparisons. The exposure part, however, contains a concrete construction loop. Appendix B generates the synthetic user prompt by conditioning on the amount of time-savings the prompt is supposed to provide, starting from 90–100% and backing off only as needed. §D.1 then feeds that same synthetic prompt into the worker-roleplay simulation and asks the LM to estimate the time-savings it would enable. The 46.6% headline thus summarizes simulated LMs rating prompts that were themselves generated to justify high savings; there is no human time-study or independent time-use data in the loop. This is a partial circularity: the quantity being predicted is inserted into the input-generation process. I do not count the Park et al. self-citation as circular, since it is an independent published result on LM simulation. The final score is 6 rather than higher because the exposure measure does produce lower estimates than the rubric baseline and is compared to external Claude usage, so the loop is not the only evidence in the paper.
Axiom & Free-Parameter Ledger
free parameters (7)
- Substantial time-savings threshold =
25% time saved
- High-exposure threshold =
50% time saved
- Significant usage threshold =
0.0025% of Anthropic work-related Claude.ai traffic
- DWA coverage threshold =
50 instances per DWA
- Retrieval design choices =
top-200 DWAs; top-300 conversations per task; shallow retrieval k=10
- Simulation repetition count =
3 runs, median
- Synthetic generation time-savings ladder =
90-100% down to 1%, backing off until verifier passes
axioms (9)
- domain assumption O*NET taxonomy is an accurate and complete codification of economically valuable US work.
- domain assumption Public chatbot conversations (WildChat, LMSYS, Chatbot Arena) are representative of real-world work-related LM usage.
- domain assumption The LM-based retrieval and classification pipeline maps conversations to DWAs with sufficient accuracy.
- domain assumption A language-model judge can reliably rank model responses on economic work tasks.
- domain assumption Synthetic prompts generated by GPT-5-mini roleplay and verified by GPT-5.2 are realistic, answerable work queries.
- domain assumption An LLM roleplaying a worker can produce accurate baseline task step times and time-savings, so the aggregate exposure reflects actual human time savings.
- domain assumption Anthropic's Claude Economic Index usage data is representative of actual economy-wide LM usage.
- domain assumption GDPval scores are a valid external benchmark for economically valuable task performance.
- standard math Standard statistical tools (Spearman correlation, median) are appropriate for the comparisons made.
read the original abstract
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an open-source evaluation suite to measure capabilities relevant to tasks, work activities, and occupations in the US labor economy. We ground the evaluation suite in real user queries to language models where possible, and supplement these with synthetic data. Our evaluations improve coverage over OpenAI's GDPval benchmark, which is the existing state-of-the-art that covers 5% of US occupations, at 500x lower cost. Alongside benchmarks, we also introduce a simulation-based exposure measure to estimate how much time current language model capabilities could save across all tasks belonging to all US occupations, with detailed accounting for each estimate. Our estimates indicate that current models could save workers substantial time on at least half of their tasks in 47% of occupations. However, for 79% of tasks where we predict substantial time savings, observed Claude usage is low, suggesting that existing usage lags potential. Beyond inherent constraints of language model chatbots, our data identifies privacy and proprietary systems as the principal bottlenecks limiting further time savings from AI. Overall, we introduce adaptable infrastructure that grounds inferences about language models' labor-market impact in their current capabilities, which can be continually updated as capabilities improve.
Figures
Reference graph
Works this paper leans on
-
[1]
Generative
Gmyrek, Pawe. Generative. 2025 , publisher=
2025
-
[2]
Science , volume =
Tyna Eloundou and Sam Manning and Pamela Mishkin and Daniel Rock , title =. Science , volume =. 2024 , doi =
2024
-
[3]
Artificial Intelligence and the Great Divergence , year =
-
[4]
and Karger, Ezra , title =
Murphy, Connacher and Rosenberg, Josh and Canedy, Jordan and Jacobs, Zach and Flechner, Nadja and Britt, Rhiannon and Pan, Alexa and Rogers-Smith, Charlie and Mayland, Dan and Buffington, Cathy and Kučinskas, Simas and Coston, Amanda and Kerner, Hannah and Pierson, Emma and Rabbany, Reihaneh and Salganik, Matthew and Seamans, Robert and Su, Yu and Tramèr,...
-
[5]
Tetlock , title =
Ezra Karger and Otto Kuusela and Jason Abaluck and Kevin Bryan and Basil Halperin and Todd Jones and Connacher Murphy and Phil Trammell and Matt Reynolds and Dan Mayland and Ria Viswanathan and Ananaya Mittal and Rebecca Ceppas de Castro and Josh Rosenberg and Philip E. Tetlock , title =. 2026 , month = mar, url =
2026
-
[6]
2026 , eprint=
How Well Does Agent Development Reflect Real-World Work? , author=. 2026 , eprint=
2026
-
[7]
Task-Completion Time Horizons of Frontier
-
[8]
Ziegler and Elizabeth Barnes and Lawrence Chan , booktitle =
Thomas Kwa and Ben West and Joel Becker and Amy Deng and Katharyn Garcia and Max Hasin and Sami Jawhar and Megan Kinniment and Nate Rush and Sydney Von Arx and Ryan Bloom and Thomas Broadley and Haoxing Du and Brian Goodrich and Nikola Jurkovic and Luke Harold Miles and Seraphina Nix and Tao Lin and Neev Parikh and David Rein and Lucas Jun Koba Sato and H...
2025
-
[12]
2025 , month = apr, day =
Arvind Narayanan and Sayash Kapoor , institution =. 2025 , month = apr, day =
2025
-
[13]
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation , author =
-
[14]
Towards a Science of
Stephan Rabanser and Sayash Kapoor and Peter Kirgis and Kangheng Liu and Saiteja Utpala and Arvind Narayanan , journal =. Towards a Science of
-
[15]
2024 , month = may, doi =
Daron Acemoglu , title =. 2024 , month = may, doi =
2024
-
[16]
2026 , month = jan, url =
Dario Amodei , title =. 2026 , month = jan, url =
2026
-
[17]
2025 , month = oct, day =
Bharat Chandar , title =. 2025 , month = oct, day =
2025
-
[18]
2026 , month = jan, day =
Alex Imas , title =. 2026 , month = jan, day =
2026
-
[19]
Ipsos Predictions Survey 2025: Positivity about how this year has gone highest since before the pandemic , year =
2025
-
[20]
2025 , url =
Human Development Report 2025: A Matter of Choice: People and Possibilities in the Age of AI , institution =. 2025 , url =
2025
-
[21]
2025 , url =
AI Index Report 2025: Public Opinion , institution =. 2025 , url =
2025
-
[22]
and Raj, Manav and Seamans, Robert , title =
Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. AEA Papers and Proceedings , volume =. 2018 , doi =
2018
-
[23]
and Raj, Manav and Seamans, Robert , title =
Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. Strategic Management Journal , volume =. 2021 , doi =
2021
-
[24]
2020 , doi =
Webb, Michael , title =. 2020 , doi =
2020
-
[25]
Measuring the Occupational Impact of
Tolan, Song. Measuring the Occupational Impact of. Journal of Artificial Intelligence Research , volume =. 2021 , doi =
2021
-
[26]
and Raj, Manav and Seamans, Robert , title =
Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. SSRN Electronic Journal , year =
-
[27]
and Cazzaniga, Mauro and Li, Longji , title =
Pizzinelli, Carlo and Panton, Augustus and Tavares, Marina M. and Cazzaniga, Mauro and Li, Longji , title =
-
[28]
Journal of Economic Perspectives , volume =
Acemoglu, Daron and Restrepo, Pascual , title =. Journal of Economic Perspectives , volume =. 2019 , doi =
2019
-
[29]
Generative
Gmyrek, Pawel and Berg, Janine and Kami. Generative. 2025 , doi =
2025
-
[30]
Mirror, Mirror on the Wall: Which Jobs Will AI Replace After All? A New Index of Occupational Exposure , institution =
Ben. Mirror, Mirror on the Wall: Which Jobs Will AI Replace After All? A New Index of Occupational Exposure , institution =
-
[31]
2026 , url =
Massenkoff, Maxim and McCrory, Peter , title =. 2026 , url =
2026
-
[32]
2026 , month = apr, url =
Sha Sajadieh and Loredana Fattorini and Raymond Perrault and Yolanda Gil and Vanessa Parli and Lapo Santarlasci and Juan Pava and Nestor Maslej and Russ Altman and Erik Brynjolfsson and Carla Brodley and Jack Clark and Virginia Dignum and Vipin Kumar and James Landay and Terah Lyons and James Manyika and Juan Carlos Niebles and Yoav Shoham and Elham Tabas...
2026
-
[34]
Yoshua Bengio and Stephen Clare and Carina Prunkl and Maksym Andriushchenko and Ben Bucknall and Malcolm Murray and Rishi Bommasani and Stephen Casper and Tom Davidson and Raymond Douglas and David Duvenaud and Philip Fox and Usman Gohar and Rose Hadshar and Anson Ho and Tiancheng Hu and Cameron Jones and Sayash Kapoor and Atoosa Kasirzadeh and Sam Mannin...
2026
-
[36]
2026 , month = jan, howpublished =
Thomas Kwa , title =. 2026 , month = jan, howpublished =
2026
-
[37]
Generative
Hosseini Maasoum, Seyed Mahdi and Lichtinger, Guy , year =. Generative
-
[38]
and Levy, Frank and Murnane, Richard J
Autor, David H. and Levy, Frank and Murnane, Richard J. , title =. The Quarterly Journal of Economics , volume =. 2003 , month = nov, doi =
2003
-
[39]
Handbook of Labor Economics , editor =
Acemoglu, Daron and Autor, David , title =. Handbook of Labor Economics , editor =. 2011 , doi =
2011
-
[40]
and Raj, Manav and Seamans, Robert , title =
Felten, Edward W. and Raj, Manav and Seamans, Robert , title =. arXiv preprint arXiv:2303.01157 , year =. 2303.01157 , archivePrefix =
-
[41]
AEA Papers and Proceedings , volume =
Brynjolfsson, Erik and Mitchell, Tom and Rock, Daniel , title =. AEA Papers and Proceedings , volume =. 2018 , doi =
2018
-
[43]
arXiv preprint arXiv:2507.07935 , year =
Tomlinson, Kiran and Jaffe, Sonia and Wang, Will and Counts, Scott and Suri, Siddharth , title =. arXiv preprint arXiv:2507.07935 , year =. 2507.07935 , archivePrefix =
-
[44]
Park, Joon Sung and Zou, Carolyn Q. and Kamphorst, Jonne and Egan, Niles and Shaw, Aaron and Hill, Benjamin Mako and Cai, Carrie and Morris, Meredith Ringel and Liang, Percy and Willer, Robb and Bernstein, Michael S. , title =. arXiv preprint arXiv:2411.10109 , year =. 2411.10109 , archivePrefix =
-
[45]
and Cai, Carrie J
Park, Joon Sung and O'Brien, Joseph C. and Cai, Carrie J. and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual. 2023 , doi =
2023
-
[46]
and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =
Chatterji, Aaron and Cunningham, Thomas and Deming, David J. and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =. 2025 , month = sep, doi =
2025
-
[47]
International Conference on Learning Representations (
Zhao, Wenting and Ren, Xiang and Hessel, Jack and Cardie, Claire and Choi, Yejin and Deng, Yuntian , title =. International Conference on Learning Representations (. 2024 , eprint =
2024
-
[49]
and Stoica, Ion , title =
Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos, Anastasios Nikolas and Li, Tianle and Li, Dacheng and Zhang, Hao and Zhu, Banghua and Jordan, Michael and Gonzalez, Joseph E. and Stoica, Ion , title =. 2024 , eprint =
2024
-
[51]
How Adaptable Are American Workers to
Sam Manning and Tom\'. How Adaptable Are American Workers to. 2026 , month = jan, url =
2026
-
[52]
2025 , month = jun, url =
David Autor and Neil Thompson , title =. 2025 , month = jun, url =
2025
-
[56]
Ruth Appel and Maxim Massenkoff and Peter McCrory and Miles McCain and Ryan Heller and Tyler Neylon and Alex Tamkin , title =
-
[57]
2026 , eprint=
Agents' Last Exam , author=. 2026 , eprint=
2026
-
[58]
Ipsos predictions survey 2025: Positivity about how this year has gone highest since before the pandemic, December 2024
Ipsos . Ipsos predictions survey 2025: Positivity about how this year has gone highest since before the pandemic, December 2024. URL https://www.ipsos.com/en-us/ipsos-predictions-2025. Reports that globally 65\
2025
-
[59]
Human development report 2025: A matter of choice: People and possibilities in the age of ai
United Nations . Human development report 2025: A matter of choice: People and possibilities in the age of ai. Technical report, United Nations Development Programme, 2025. URL https://hdr.undp.org/system/files/documents/global-report-document/hdr2025reporten.pdf. Uses the 2025 Global Survey on AI and Human Development; reports that 57\
2025
-
[60]
Ai index report 2025: Public opinion
Stanford HAI . Ai index report 2025: Public opinion. Technical report, Stanford HAI, 2025. URL https://hai.stanford.edu/ai-index/2025-ai-index-report/public-opinion. Reports that globally 60\
2025
-
[61]
Artificial intelligence and the great divergence, January 2026
The White House . Artificial intelligence and the great divergence, January 2026. URL https://www.whitehouse.gov/research/2026/01/artificial-intelligence-and-the-great-divergence/. Accessed: 2026-04-09
2026
-
[62]
Tetlock, and Ezra Karger
Connacher Murphy, Josh Rosenberg, Jordan Canedy, Zach Jacobs, Nadja Flechner, Rhiannon Britt, Alexa Pan, Charlie Rogers-Smith, Dan Mayland, Cathy Buffington, Simas Kučinskas, Amanda Coston, Hannah Kerner, Emma Pierson, Reihaneh Rabbany, Matthew Salganik, Robert Seamans, Yu Su, Florian Tramèr, Tatsunori Hashimoto, Arvind Narayanan, Philip E. Tetlock, and E...
2025
-
[63]
Ezra Karger, Otto Kuusela, Jason Abaluck, Kevin Bryan, Basil Halperin, Todd Jones, Connacher Murphy, Phil Trammell, Matt Reynolds, Dan Mayland, Ria Viswanathan, Ananaya Mittal, Rebecca Ceppas de Castro, Josh Rosenberg, and Philip E. Tetlock. Forecasting the economic effects of AI , March 2026. URL https://static1.squarespace.com/static/635693acf15a3e2a14a...
arXiv 2026
-
[64]
The simple macroeconomics of AI
Daron Acemoglu. The simple macroeconomics of AI . Working Paper 32487, National Bureau of Economic Research, May 2024. URL https://www.nber.org/papers/w32487
2024
-
[65]
The adolescence of technology, January 2026
Dario Amodei. The adolescence of technology, January 2026. URL https://www.darioamodei.com/essay/the-adolescence-of-technology. Accessed: 2026-04-09
2026
-
[66]
AI and labor markets: What we know and don't know, October 2025
Bharat Chandar. AI and labor markets: What we know and don't know, October 2025. URL https://digitaleconomy.stanford.edu/news/ai-and-labor-markets-what-we-know-and-dont-know/. Accessed: 2026-04-09
2025
-
[67]
What is the impact of AI on productivity? Ghosts of Electricity (Substack), January 2026
Alex Imas. What is the impact of AI on productivity? Ghosts of Electricity (Substack), January 2026. URL https://aleximas.substack.com/p/what-is-the-impact-of-ai-on-productivity. Accessed: 2026-04-09
2026
-
[68]
Maria del Rio-Chanona, Ekkehard Ernst, Rossana Merola, Daniel Samaan, and Ole Teutloff
R. Maria del Rio-Chanona, Ekkehard Ernst, Rossana Merola, Daniel Samaan, and Ole Teutloff. AI and jobs. A review of theory, estimates, and evidence, 2025. URL https://arxiv.org/abs/2509.15265
arXiv 2025
-
[69]
Yoshua Bengio, Stephen Clare, Carina Prunkl, Maksym Andriushchenko, Ben Bucknall, Malcolm Murray, Rishi Bommasani, Stephen Casper, Tom Davidson, Raymond Douglas, David Duvenaud, Philip Fox, Usman Gohar, Rose Hadshar, Anson Ho, Tiancheng Hu, Cameron Jones, Sayash Kapoor, Atoosa Kasirzadeh, Sam Manning, Nestor Maslej, Vasilios Mavroudis, Conor McGlynn, Rich...
2026
-
[70]
Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo...
Pith/arXiv arXiv 2026
-
[71]
Artificial intelligence index report 2026
Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, and Dan Weld. Artificial in...
2026
-
[72]
Ziegler, Elizabeth Barnes, and Lawrence Chan
Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, and Lawrence ...
Pith/arXiv arXiv 2025
-
[73]
Task-completion time horizons of frontier AI models
METR . Task-completion time horizons of frontier AI models. https://metr.org/time-horizons/, March 2026
2026
-
[74]
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Sim \'o n Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek. GDPval : Evaluating AI model performance on real...
Pith/arXiv arXiv 2025
-
[75]
Remote labor index: Measuring AI automation of remote work
Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, Vinaya Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevi...
arXiv 2025
-
[76]
Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski
Bertie Vidgen, Abby Fennelly, Evan Pinnix, Julien Benchek, Daniyal Khan, Zach Richards, Austin Bridges, Calix Huang, Kanishka Sahu, Abhishek Kottamasu, Bo Ma, Ben Hunsberger, Isaac Robinson, Akul Datta, Chirag Mahapatra, Dominic Barton, Cass R. Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski. The AI P roductivity I ndex ( APEX ), 2025. URL https://...
arXiv 2025
-
[77]
Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, et al. Agents' last exam, 2026. URL https://arxiv.org/abs/2606.05405
arXiv 2026
-
[78]
How well does agent development reflect real-world work?, 2026
Zora Zhiruo Wang, Sanidhya Vijayvargiya, Aspen Chen, Hanmo Zhang, Venu Arvind Arangarajan, Jett Chen, Valerie Chen, Diyi Yang, Daniel Fried, and Graham Neubig. How well does agent development reflect real-world work?, 2026. URL https://arxiv.org/abs/2603.01203
arXiv 2026
-
[79]
AI as normal technology: An alternative to the vision of AI as a potential superintelligence
Arvind Narayanan and Sayash Kapoor. AI as normal technology: An alternative to the vision of AI as a potential superintelligence. Technical report, Knight First Amendment Institute, April 2025. URL https://knightcolumbia.org/content/ai-as-normal-technology
2025
-
[80]
Sara Fish, Julia Shephard, Minkai Li, Ran I. Shorrer, and Yannai A. Gonczarowski. EconEvals : Benchmarks and litmus tests for economic decision-making by llm agents, 2026. URL https://arxiv.org/abs/2503.18825
arXiv 2026
-
[81]
WildChat : 1 M ChatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. WildChat : 1 M ChatGPT interaction logs in the wild. In International Conference on Learning Representations ( ICLR ) , 2024. URL https://arxiv.org/abs/2405.01470
Pith/arXiv arXiv 2024
-
[82]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. LMSYS-Chat-1M : A large-scale real-world LLM conversation dataset. arXiv preprint arXiv:2309.11998, 2023. URL https://arxiv.org/abs/2309.11998
Pith/arXiv arXiv 2023
-
[83]
Anthropic economic index report: economic primitives, 2026
Ruth Appel, Maxim Massenkoff, Peter McCrory, Miles McCain, Ryan Heller, Tyler Neylon, and Alex Tamkin. Anthropic economic index report: economic primitives, 2026
2026
-
[84]
Autor, Frank Levy, and Richard J
David H. Autor, Frank Levy, and Richard J. Murnane. The skill content of recent technological change: An empirical exploration. The Quarterly Journal of Economics, 118 0 (4): 0 1279--1333, November 2003. doi:10.1162/003355303322552801. URL https://academic.oup.com/qje/article-abstract/118/4/1279/1925105
-
[85]
GPTs are GPTs : Labor market impact potential of LLMs
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. GPTs are GPTs : Labor market impact potential of LLMs . Science, 384 0 (6702): 0 1306--1308, 2024. doi:10.1126/science.adj0998. URL https://www.science.org/doi/abs/10.1126/science.adj0998
-
[86]
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology ( UIST '23) , 2023. doi:10.1145/3586183.3606763. URL https://dl.acm.org/doi/10.1145/3586183.3606763
arXiv 2023
-
[87]
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. Chatbot arena: An open platform for evaluating LLMs by human preference, 2024. URL https://arxiv.org/abs/2403.04132
Pith/arXiv arXiv 2024
-
[88]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazar \'e , Maria Lomeli, Lucas Hosseini, and Herv \'e J \'e gou. The Faiss library. arXiv preprint arXiv:2401.08281, 2024. URL https://arxiv.org/abs/2401.08281
Pith/arXiv arXiv 2024
-
[89]
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: A rena- H ard and B ench B uilder pipeline, 2024. URL https://arxiv.org/abs/2406.11939
Pith/arXiv arXiv 2024
-
[90]
A rosetta stone for AI benchmarks, 2025
Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. A rosetta stone for AI benchmarks, 2025. URL https://arxiv.org/abs/2512.00193
arXiv 2025
-
[91]
Skills, tasks and technologies: Implications for employment and earnings
Daron Acemoglu and David Autor. Skills, tasks and technologies: Implications for employment and earnings. In Orley Ashenfelter and David Card, editors, Handbook of Labor Economics, volume 4, chapter 12, pages 1043--1171. Elsevier, Amsterdam, 2011. doi:10.1016/S0169-7218(11)02410-5. Part B
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.