{"id":"3cc03aae-98b4-4115-b48e-85e5b8f15e67","arxiv_id":"2507.08244","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Occupations with larger LLM-assessed increases in AI exposure saw larger employment declines, unemployment rises, and work-hour reductions in the US between late 2022 and early 2025.","lead":"This economics paper builds a new measure of how exposed each US occupation is to AI, based on what ChatGPT-4o and Claude 3.5 say they can do on O*NET task lists, and links it to monthly U.S. Current Population Survey data. It finds that occupations with larger increases in AI exposure between 2022 and 2025 experienced larger employment declines, higher unemployment, and shorter work hours.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated LLM self-reports are the sole source of the exposure regressor; cross-model agreement is inter-rater reliability, not accuracy.","rationale":"I reviewed the construction of OAIES in Section 2.1, the first-differenced specification in Section 4.1, the main results in Table 5, the robustness checks (OCC2010, WFH, sample thresholds, placebos), and the Limitations section. The most load-bearing assumption is indeed that the LLM self-reported task performability scores measure AI capability accurately enough to rank occupations by exposure change. The paper's own limitation statement concedes this is unvalidated. The cross-model correlations are strong but only establish reliability, not validity, since both models share training data and prompt structure. The placebo tests address pre-existing trends, not classical or systematic measurement error in the regressor. I also note the baseline specification shows no association until demographic and task controls are added, which increases the stakes for the control specification but is not itself a disqualifying flaw. Given the authors' transparency and the many robustness checks, the conditional verdict is appropriate: the paper provides an early signal, but magnitudes should not be treated as established until the exposure measure is externally validated. Therefore I recommend no change to the reader's verdict.","tokens_in":43545,"tokens_out":4381,"duration_ms":55932,"concrete_test":"Select a stratified random sample of 100 O*NET tasks. For each stage, run actual historical or hypothetical model checkpoints under stage-appropriate constraints (e.g., GPT-3 for Stage 2, GPT-4o for Stage 3, o1 for Stage 4) and have independent human raters judge whether the model completes each task. Compare measured success rates to the model self-reported percentages used in OAIES. If self-reports deviate systematically by task type or stage, recalibrate the exposure scores and re-estimate Equation (1). If the Table 5 Panel C coefficient on log employment changes by more than its standard error, the central claim is not robust to measurement error in the key regressor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 builds OAIES by asking ChatGPT-4o and Claude 3.5 Sonnet to estimate the percentage of each O*NET task they could perform at five hypothetical capability stages, including retroactive Stage 1 and forward-looking Stage 5. Equation (1) and Table 5 Panel C then use the difference S3-S1 as the key regressor. The central claim in the reader's strongest_claim inherits every bias in that self-report. The paper concedes in the Limitations (Section 5) that the scores 'do not validate the accuracy of the scores themselves.' Cross-model agreement (Appendix Table 12: Pearson 0.89-0.96) is evidence of inter-rater reliability only; both models share training distributions and were given the same prompting frame, so agreement cannot establish ground truth. If LLM self-assessments are systematically over- or under-confident for occupations that also have differential employment trends (e.g., cognitive vs manual jobs), the estimated coefficients are biased. The placebo tests in Figure 8 check pre-trends, not measurement error in the regressor. A related fragility: the raw association in Table 5 Panel A is near zero (-0.07), and the headline coefficient appears only after demographic and task controls; this makes the result depend on the very controls and on the unvalidated exposure measure. Thus the employment effect is conditionally plausible, but the load-bearing premise is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper constructs a dynamic Occupational AI Exposure Score (OAIES) by asking two LLMs (ChatGPT-4o and Claude 3.5 Sonnet) to self-assess, for each O*NET task, the percentage the model could perform at each of five hypothetical AI capability stages. The scores are aggregated to 513 Census-SOC occupations and linked to occupation-level outcomes from the CPS. Using first-differenced regressions between October 2022–March 2023 and October 2024–March 2025, the paper reports that a 10-point increase in the S3–S1 exposure change is associated with a 5.6–8.5 percentage point decline in scaled log employment, a 0.64–0.68 percentage point increase in the unemployment rate, and reductions in main-job hours, with heterogeneous effects by age, gender, education, and task content. The authors explicitly interpret the results as associations, not causal effects.","tokens_in":43776,"tokens_out":5273,"duration_ms":64376,"significance":"If the OAIES is a valid measure of AI capability exposure, this is a valuable and timely contribution: it is among the first near-real-time occupation-level analyses of AI and labor outcomes using the CPS, and it extends the outcome set beyond employment to hours, full-time status, and secondary jobs. The empirical work is careful in several respects: the analysis uses first differences, demographic and task controls, placebo periods before ChatGPT, an alternative harmonized occupational classification (OCC2010), a work-from-home robustness check, and varying occupation sample-size thresholds. These design choices are appropriate for a descriptive association study. However, the central claim rests entirely on an unvalidated LLM self-assessment measure, and the paper's own Limitations section concedes that the scores 'do not validate the accuracy of the scores themselves.' Cross-model agreement is evidence of inter-rater reliability, not validity. The significance of the headline findings is therefore conditional on an assumption that the paper does not yet establish.","major_comments":[{"comment":"The key regressor, ΔExp(S3–S1), is constructed from LLM self-assessments of task performance with no external benchmark. The paper's own Limitations (Section 5) state that the scores 'do not validate the accuracy of the scores themselves.' The high cross-model correlations in Table 12 (Pearson 0.89–0.96 for S3–S1) establish inter-rater reliability, not validity: both models share training distributions and were given the same prompting frame. If LLM self-assessments are systematically over- or under-confident for occupations that also have differential employment trends (e.g., cognitive versus manual jobs), the coefficients in Table 5 Panel C inherit that bias. The placebo tests in Figure 8 check pre-trends in outcomes, not measurement error in the regressor. I would like to see validation against an external benchmark—for example, human expert task ratings, the Eloundou et al. (2024) task-level ratings, or Webb's (2019) patent-based exposure measure—and re-estimation of the main specification with that alternative.","section":"Section 2.1, Eq. (1), Table 12"},{"comment":"The raw association between ΔExp(S3–S1) and changes in log employment is essentially zero in Panel A (β = -0.07, s.e. 0.13 for ChatGPT; β = -0.07, s.e. 0.14 for Claude). The headline coefficients (-0.558 and -0.846) emerge only after demographic controls and task indices are added. This makes the central result heavily dependent on the control specification. Because the controls are measured in Period 2 and may themselves be affected by early AI adoption or by occupational compositional changes correlated with exposure, the estimates could reflect selection on controls rather than a robust exposure effect. The authors should present a structured sensitivity analysis—for example, adding controls one at a time or reporting coefficient stability measures such as Oster's (2019) delta—and justify why the P2 demographic shares are appropriate controls in a first-differenced design.","section":"Table 5, Panels A–C"},{"comment":"The exposure regressor ΔExp(S3–S1) is a capability-stage contrast: Stage 3 begins in October 2023, while the outcome window is the change from Period 2 (October 2022–March 2023) to Period 4 (October 2024–March 2025), which includes Stage 4. The paper's own Figure 9 shows that the standardized coefficient varies substantially across stage differences, with the strongest associations for S2–S1 and weaker or different patterns for S4–S1. Using S3–S1 as 'the' exposure change is therefore not neutral. The authors should either align the exposure change with the outcome window (e.g., S4–S1 or a time-varying exposure measure) or provide a substantive justification for why the S3–S1 contrast is the relevant one for the P4–P2 outcome change.","section":"Section 4.1–4.2, Eq. (1), Figure 9"}],"minor_comments":[{"comment":"The text states that 'The full prompt text is included in the Appendix,' but the appendix as presented contains only figures and tables, not the prompt. Please include the complete prompt so that the exposure construction is reproducible.","section":"Section 2.1"},{"comment":"The sentence 'These patterns align with expectations that generative AI primarily affects cognitive information-processing tasks, s to earlier technologies like robotics' contains a typo; it should read 'compared to earlier technologies like robotics.'","section":"Section 3.1"},{"comment":"The demographic heterogeneity model is introduced as 'equation (3)' but the displayed equation is labeled '(2)'; the equation numbering should be corrected.","section":"Section 4.6"},{"comment":"The paper refers to 'Dingel and Nieman' in the text; the correct spelling is 'Neiman' (Dingel and Neiman, 2020).","section":"Section 4.2 and References"},{"comment":"Section 2.2 says the analysis uses robust standard errors that do not account for the CPS complex survey design, but Section 5's Limitations state 'we include clustered standard errors.' Please clarify which standard errors are actually reported and, if clustering is used, state the clustering unit.","section":"Sections 2.2 and 5"},{"comment":"The note says Capped Earnings are 'normalized to constant 2010 dollars' but the cap of $2,884.61 appears to be a nominal cap; please clarify whether the cap is applied before or after the CPI adjustment so the wage results are interpretable.","section":"Table 19 note"}],"recommendation":"major_revision","confidential_remarks":"The paper is timely and the descriptive empirical work is careful, but the core exposure measure is the central risk. I am not asking for causal identification, but the authors must either externally validate the OAIES against an independent benchmark or substantially reframe the headline claims as conditional on a measure whose validity is unknown. The paper's own Limitations section acknowledges this gap, which is to the authors' credit, but the acknowledgment does not resolve the issue. The dependence of the main result on the control specification and the misalignment between the S3–S1 exposure contrast and the P4–P2 outcome window are additional reasons why the current version is not yet ready for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious, careful empirical paper with a genuinely new measurement approach, but the key regressor is an unvalidated LLM self-report, and the headline associations only appear after controls. Treat it as suggestive early evidence, not as a measurement-based fact.\n\nWhat's new: the five-stage OAIES that evolves with model capabilities and its link to monthly CPS data is a real step beyond static exposure measures. The first-differenced design with demographic and task controls, placebo periods, sample-size thresholds, alternative occupational classifications, and the WFH control is more thorough than most applied papers in this space. The authors are honest about non-causality and about standard-error limitations.\n\nWhere it's soft: the exposure score is built entirely from ChatGPT-4o and Claude 3.5 rating what percentage of each O*NET task they could perform at five capability stages. No external benchmark checks whether those self-assessments are accurate. Cross-model agreement (Pearson 0.89-0.96) shows inter-rater reliability, not ground truth. Both models share training distributions and the same prompting frame, so they can be wrong in the same way. The paper's own Limitations section concedes the scores 'do not validate the accuracy of the scores themselves.' That's the load-bearing premise, and it is unverified. Also, the raw association in Panel A of Table 5 is near zero; the employment coefficient appears only after demographic and task controls. That makes the result depend on the controls and on the unvalidated measure. The placebo tests address pre-trends, not measurement error.\n\nProportionately: these are real weaknesses, but they are acknowledged, and the empirical work is transparent. The paper does not overclaim causality. It is exactly the kind of conditional, early-evidence paper that merits peer review, with referees pushing for external validation or at least a clear statement that the measure is a belief-elicitation instrument.\n\nFor whom: labor economists and AI-and-work researchers will want to engage. A good referee could strengthen it substantially. I'd send it to review.","headline":"Dynamic AI exposure scores tied to CPS data are genuinely new, but the unvalidated LLM self-reports make the headline estimates conditional on measurement assumptions the paper acknowledges but does not resolve.","tokens_in":44320,"tokens_out":1743,"would_cite":false,"duration_ms":20560,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that U.S. occupations with larger increases in AI task-level exposure between late 2022 and early 2025 experienced measurable employment declines, higher unemployment, and shorter work hours.","keywords":["AI capabilities","AI exposure","occupations","employment","work hours","Current Population Survey","large language models","task-level measurement"],"falsifier":"Collect ground-truth performance data: have independent human experts or standardized benchmark suites rate the same O*NET tasks at each of the five stages, then compare those ratings with the LLM self-assessments. If the two diverge systematically—for example, if models overstate their ability on tasks that require physical presence, tacit knowledge, or accountability—the exposure regressor is mismeasured and the estimated employment and unemployment associations are not trustworthy. A simpler version: rerun the analysis with exposure scores built from expert ratings instead of LLM self-assessments and check whether the signs and magnitudes survive.","tokens_in":43283,"feed_emoji":"📉","tokens_out":5191,"duration_ms":52444,"temperature":0.7,"pith_summary":"This paper tries to establish that recent advances in generative AI capabilities are already visible in U.S. labor market data. It builds a dynamic Occupational AI Exposure Score by asking ChatGPT-4o and Claude 3.5 Sonnet to rate, task by task, how much of each occupation's work AI could perform at five stages of capability, then links those scores to monthly Current Population Survey records. Comparing October 2022–March 2023 with October 2024–March 2025, the paper finds that occupations with larger exposure increases saw bigger employment declines, higher unemployment, and fewer hours at the main job. If the relationship holds, 2025 may mark the start of measurable AI-driven labor disruption rather than just anecdotal reports.","feed_headline":"AI-exposed occupations lost jobs and hours since ChatGPT","feed_subtitle":"A five-stage task-level exposure score links the 2022–2025 CPS to employment, unemployment, and work-hour declines.","key_machinery":"The central object is the five-stage Occupational AI Exposure Score (OAIES), a 0–100 measure of the share of an occupation's O*NET tasks that a frontier LLM reports being able to perform. Stage 1 is pre-LLM machine learning; Stage 2 early LLMs; Stage 3 multimodal models; Stage 4 reasoning models; Stage 5 agentic AI. The scoring works by prompting ChatGPT-4o and Claude 3.5 Sonnet to estimate the percentage of each task performable at each stage, weighting task estimates by O*NET relevance, and aggregating to Census-SOC occupations. The empirical engine is an occupation-level first-differenced regression of labor outcome changes (employment, unemployment rate, hours, part-time, second jobs) between Periods 2 and 4 on the exposure change from Stage 1 to Stage 3, with demographic composition and task indices as controls. The five-stage design is what makes the exposure measure dynamic rather than a static snapshot.","core_discovery":"The paper's central claim is that rising occupational exposure to AI capability is associated with deterioration in labor outcomes in the United States between the late-2022 launch of ChatGPT and early 2025. In the fully controlled specification, a 10-point increase in exposure between Stage 1 (pre-LLM machine learning) and Stage 3 (multimodal LLMs) is associated with a 5.6 percentage point decline in occupational employment and a 0.64 percentage point rise in the unemployment rate using ChatGPT-generated scores, and an 8.5 percentage point employment decline and 0.68 percentage point unemployment rise using Claude-generated scores. The same exposure change is tied to shorter main-job hours and, among several subgroups, more secondary job holding and less full-time work. The authors interpret these as associations, not causal effects, and read them as evidence that AI-driven labor shifts are appearing on both the extensive margin (fewer jobs) and the intensive margin (fewer hours).","pith_inferences":["Editorial inference: the exposure measure treats task overlap as potential substitution, but a task that AI can perform may instead be a complement that raises demand for the worker; the paper's design cannot distinguish substitution from complementarity, so the negative associations may understate or overstate true displacement.","Editorial inference: because the two LLMs produce correlated but different magnitudes (Claude's employment coefficient is about 50% larger), the quantitative size of the effect is model-dependent; a validation exercise against observed firm-level adoption or layoff announcements would identify which score is closer to reality.","Editorial inference: the finding that women had larger exposure increases yet men saw larger adverse outcomes suggests exposure level alone is not the driver; occupation-level task context and labor-market power likely mediate the effect, a mechanism the paper notes but does not test."],"forward_implications":["If the central association is correct, occupations with high exposure change should continue to show employment and hours losses in later CPS releases as Stage 4 and Stage 5 capabilities diffuse.","Policymakers monitoring unemployment by occupation can treat the exposure score as an early-warning signal of future employment decline.","The paper's period decomposition implies the first visible sign of AI disruption should be rising unemployment and secondary-job activity, followed a year or two later by larger employment losses.","College-educated workers' smaller employment losses but larger shifts in hours and full-time status imply workforce adjustment will show up first as job restructuring rather than layoffs in highly educated occupations.","Manual and routine-manual occupations are predicted to be relatively insulated in the near term, with possible employment gains."],"supporting_citations":[{"why":"Pioneered the prompt-engineering method of asking LLMs to rate task-level AI capability that the paper's OAIES construction directly builds on.","marker":"Eloundou et al. (2024)"},{"why":"Supplies the task indices used as controls and heterogeneity dimensions, including the non-routine cognitive analytical and routine manual measures.","marker":"Acemoglu and Autor (2011)"},{"why":"Provides the work-from-home indicator used in robustness checks to rule out post-COVID in-person rebound as the driver.","marker":"Dingel and Neiman (2020)"},{"why":"Documents that CPS public-use data without design variables understate standard errors, which the paper cites when cautioning about inference.","marker":"Davern et al. (2006, 2007)"},{"why":"Earlier expert-based automation probability measure that this paper contrasts with its dynamic, task-level exposure scores.","marker":"Frey and Osborne (2017)"},{"why":"Patent-based occupational exposure measure that the paper contrasts as static and innovation-oriented rather than capability-dynamic.","marker":"Webb (2019)"},{"why":"Provides field-experimental evidence of generative AI productivity gains in customer service, motivating the paper's focus on intensive-margin adjustment.","marker":"Brynjolfsson et al. (2025)"},{"why":"Uses millions of real Claude conversations to show which economic tasks are performed with AI, an updated, usage-based counterpart to the paper's capability-exposure approach.","marker":"Handa et al. (2025)"}],"fun_headline_variants":["AI exposure tied to job losses and shorter hours","Higher AI exposure linked to unemployment and fewer hours","AI capability advances tied to labor market declines","Jobs and hours fall for AI-exposed occupations","Study: AI exposure linked to labor force losses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an LLM's self-reported percentage of each occupation's tasks it can perform is a valid measure of real AI capability exposure; the paper itself concedes the scores do not validate their own accuracy.","fun_headline_variants_meta":{"raw":{"variants":["AI exposure tied to job losses and shorter hours","Higher AI exposure linked to unemployment and fewer hours","AI capability advances tied to labor market declines","Jobs and hours fall for AI-exposed occupations","Study: AI exposure linked to labor force losses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000695,"raw_usage":{"total_tokens":3193,"prompt_tokens":1044,"completion_tokens":2149,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":2079}},"tokens_in":660,"tokens_out":2149,"duration_ms":17552,"temperature":1.0,"reasoning_tokens":2079,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:22:31.039408+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect ground-truth performance data: have independent human experts or standardized benchmark suites rate the same O*NET tasks at each of the five stages, then compare those ratings with the LLM self-assessments. If the two diverge systematically—for example, if models overstate their ability on tasks that require physical presence, tacit knowledge, or accountability—the exposure regressor is mismeasured and the estimated employment and unemployment associations are not trustworthy. A simpler version: rerun the analysis with exposure scores built from expert ratings instead of LLM self-assessments and check whether the signs and magnitudes survive.","supporting_citations":[],"review_version":1}