{"id":"79b78d9f-0522-4644-a390-cbc97595cab5","arxiv_id":"2607.08920","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Deep AI adoption in S&P 500 firms hit 11% (score 5) and 21% (scores 4–5) by 2025 after quadrupling from 2022, with a profitability J-curve, tech dominance, and no robust capex or revenue-per-employee gains.","lead":"S&P 500 firms' deep AI integration reached 11% in 2025 (21% at production or deeper), more than quadrupling since 2022, led by tech. A new 10-K-based rubric shows a profitability J-curve but no clear capex or productivity links.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 10-K LLM scores may systematically mis-rank deep integration versus selective disclosure, so the 11%/21% rates and J-curve rest on an under-validated classifier.","rationale":"The reader correctly isolates the measurement step as the weakest assumption. The paper is careful on causality and reports useful descriptive patterns, but every quantitative claim in the abstract and strongest_claim is a direct function of the LLM scores. Industry-level correlations with BTOS/Ramp are encouraging yet do not validate firm-year ordinal rankings or the critical 3/4/5 distinctions that drive the J-curve and the 11%/21% figures. A modest human re-coding exercise is feasible, falsifiable, and would either shore up or materially qualify the contribution. No stronger internal inconsistency appears; residual endogeneity is already disclaimed. Therefore the CONDITIONAL verdict stands, with the same high-confidence measurement caveat the reader identified.","tokens_in":31812,"tokens_out":638,"duration_ms":19253,"concrete_test":"Draw a stratified random sample of ~150 firm-years (balanced across scores 1–5 and tech/non-tech). Have two independent human coders, blinded to the LLM score, apply the exact Table-2 rubric to the same extracted paragraphs and assign 1–5. Report Cohen’s kappa and the confusion matrix, especially 3-vs-4 and 4-vs-5. If agreement with GPT-5-mini is <0.6 or >25% of LLM 4/5 labels are human-coded ≤3, recompute the 2025 11%/21% rates and the Table-5 J-curve coefficients on the human-validated subset; material shifts would weaken the headline claims.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central descriptive claims (11% score-5 and 21% score-4/5 in 2025; quadrupling since 2022; tech two-thirds of deep adoption; profitability J-curve) all rest on GPT-5-mini ordinal scores of keyword-filtered 10-K paragraphs against the authors' 1–5 rubric (Section 2.1, Table 2). The legal claim that 10-Ks cannot contain material falsehoods does not prevent selective emphasis, aspirational language, or boilerplate that the model can still map to high scores. Validation is only industry-level correlations with BTOS (0.76) and Ramp (0.72–0.87) (Section 2.1.1, Fig. 1); there is no firm-level gold-standard sample, inter-rater reliability, or confusion matrix between scores 3/4/5. If the classifier systematically over-assigns 4–5 to firms that merely discuss AI strategy or under-assigns true production users who disclose little, both the adoption rates and the outcome regressions (Tables 5–9) become unreliable. The authors note manual spot-checks and Magnificent-Seven consistency, but that is insufficient for the load-bearing measurement.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper constructs a firm-year AI adoption measure for S&P 500 firms (2016–2025) by keyword-filtering 10-K paragraphs and scoring them 1–5 with GPT-5-mini against a rubric that distinguishes no mention, exploration, pilots, production use, and deep process integration. It reports that by 2025, 11% of firms score 5 and 21% score 4 or 5 (up from ~5% in 2022), with technology firms accounting for about two-thirds of deep integration. Fixed-effects regressions of Compustat outcomes on these scores recover a J-curve in net profit margin (especially for non-tech), little association with capex-to-revenue or revenue-per-employee, and positive associations of adoption with headcount and Tobin’s q mainly among technology firms. The authors emphasize descriptive correlations and flag endogeneity.","tokens_in":32158,"tokens_out":1347,"duration_ms":13701,"significance":"If the 10-K scores are reliable, the paper supplies a timely, transparent, enterprise-level panel of deep AI adoption for the largest U.S. public firms, with external industry-level validation against BTOS (corr 0.76) and Ramp (0.72–0.87) and clear sector and time patterns that are hard to obtain from surveys alone. The measurement contribution and the descriptive facts on post-ChatGPT acceleration, tech concentration, and the profitability J-curve would be useful inputs for productivity, labor, and finance work on AI. Strengths include the a priori rubric and keyword list, public-filing basis, and explicit non-causal framing of the outcome regressions.","major_comments":[{"comment":"Section 2.1 and 2.1.1: The central claims (11%/21% rates, quadrupling, tech share of deep adoption, and all Tables 5–9 results) rest on GPT-5-mini ordinal scores of keyword-filtered 10-K text. Validation is only industry-level correlations with BTOS and Ramp; there is no firm-level gold-standard sample, inter-rater reliability, or confusion matrix for scores 3 vs 4 vs 5. Manual spot-checks and Magnificent-Seven consistency are insufficient for a load-bearing classifier. A modest human-coded subsample (or multi-model agreement) with reported precision/recall by score is needed before the rates and J-curve can be treated as reliable.","section":"Section 2.1, 2.1.1"},{"comment":"Table 5 (preferred cols 4 and 6) and Fig. 3: The non-tech J-curve for net profit margin is driven by a small number of score-5 non-tech observations (text notes ~19 firms and wide CIs). With firm FE, the score-5 coefficient is large and significant, but the cell is thin and sensitive to outliers (as the authors themselves show for Tobin’s q with ALGN/PAYC in Table 7). Report cell counts by sector×score, re-estimate with winsorization or leave-one-out, and qualify the non-tech deep-integration claim accordingly.","section":"Table 5, Fig. 3"},{"comment":"Eqs. (1)–(3) and Section 3: The authors correctly state that estimates are correlations, not causal. The abstract and conclusion still lead with the J-curve and “no differences in capex or productivity” as if they are robust adoption effects. Tighten abstract/conclusion language so that measurement facts are primary and outcome associations are clearly descriptive; the lagged-AI and sector-year FE checks help but do not resolve reverse causality or time-varying confounders.","section":"Abstract; Eqs. (1)–(3); Section 5"}],"minor_comments":[{"comment":"Manual reclassification of Amazon and Tesla (and other GICS adjustments) into technology is consequential for the “tech two-thirds” claim; document the full list of reclassifications and show robustness without them.","section":"Section 2 (sector definition)"},{"comment":"Table 1 and Appendix C: report N by AI score×sector×year so readers can see how thin the high-score cells are, especially non-tech score 5.","section":"Table 1; Appendix C"},{"comment":"Fig. 2 / Table 12: clarify whether bars are shares of the contemporaneous S&P 500 panel or a fixed firm set; index composition changes over 2016–2025.","section":"Fig. 2; Table 12"},{"comment":"Prompt and model: state temperature/seed (or determinism settings) and whether scores were re-run for stability; Appendix B keyword list is useful but coverage of generative-AI terms post-2022 could be noted.","section":"Section 2.1; Appendix B"},{"comment":"Minor typos and consistency: e.g., “non-techology,” “regress-ing,” and mixed use of “genAI” vs “AI”; ensure 2025 10-K coverage is complete or flag partial-year filings.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The measurement idea is timely and the descriptive adoption facts would be valuable if the classifier is better validated. Without firm-level validation of the 3/4/5 cut, I would not treat the 11%/21% rates or the non-tech J-curve as ready for a top field journal. Scope is a good fit for applied micro / industrial organization / productivity outlets if the validation bar is met; otherwise the paper is closer to a data note."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new thing here is a firm-year 1–5 deep-integration score for the S&P 500 2016–2025, built from keyword-filtered 10-K paragraphs and a production-oriented rubric, not chatbot use. That panel is what people will actually use. In 2025 they get 11% score 5 and 21% score 4/5, more than a quadrupling since 2022, with tech carrying two-thirds of deep adoption. Those headline rates, the sector tiers in Table 4, and the Mag-7 consistency checks are the paper’s real product.\n\nWhat they do well: the rubric is explicit (Table 2), the legal-text motivation is sensible, and they validate at industry level against BTOS (0.76) and Ramp (0.72–0.87). They correctly flag endogeneity and reverse causality, run firm and sector-year fixed effects, and report nulls on capex and revenue-per-employee that match the “most firms buy AI as opex” story. The J-curve in net profit margin for non-tech (early dip, deep-integration lift) and the size/q gradient only in tech are cleanly reported and useful as descriptive facts. Citations sit in the right places (Brynjolfsson J-curve, Bresnahan org complements, Acemoglu risk/adoption, Census BTOS).\n\nSoft spot, in proportion: the load-bearing step is GPT-5-mini scoring of filtered paragraphs. There is no firm-level gold sample, no confusion matrix for 3 vs 4 vs 5, and no inter-rater reliability. Industry correlations and spot-checks (including Mag-7) are better than nothing, but selective disclosure and aspirational language can still inflate high scores. If that measurement error is systematic, the 11%/21% rates and the outcome tables move together. That is the main reason this is conditional rather than unconditional. Code and data are not released, which is a practical friction for a measurement paper.\n\nThis is for people who need a credible large-firm AI diffusion series for productivity accounting, labor forecasts, or valuation work—not for anyone looking for causal identification or a new growth model. The descriptive core is solid enough that a serious editor should send it to referees; the measurement validation is the natural revision ask. I would engage with the work and cite the adoption rates with the usual caveat on LLM scoring.","headline":"Usable S&P 500 deep-AI panel from 10-Ks; rates and J-curve are real contributions, but the LLM classifier is only lightly validated.","tokens_in":32773,"tokens_out":606,"would_cite":true,"duration_ms":8177,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"By 2025 only 21% of S&P 500 firms had AI in production or deep process integration, up fourfold since 2022, with tech firms driving most of the deep cases and profitability tracing a J-curve.","keywords":["AI adoption","S&P 500","10-K filings","enterprise AI","J-curve","productivity","Tobin's q","technology sector"],"falsifier":"Independent audits or structured surveys of the same S&P 500 firms that show systematically different rank orderings of deep integration versus the 10-K rubric scores, especially for non-technology firms claiming score 4 or 5.","tokens_in":32700,"feed_emoji":"📈","tokens_out":1080,"duration_ms":15590,"temperature":0.7,"pith_summary":"This paper builds a firm-year measure of deep enterprise AI adoption from SEC 10-K filings for S&P 500 companies from 2016 to 2025. It separates real integration of AI into business processes from mere exploration or hype, scoring firms from no mention through pilot use to production deployment and full strategic embedding. The central finding is that deep adoption remains limited: in 2025 only 11% of firms scored at full integration and another 10% used AI in production of goods or services, yet the combined share more than quadrupled from 2022 levels. Technology firms account for roughly two-thirds of the deepest adoption and show aggressive uptake, while non-technology firms move more slowly. Profitability follows a J-curve—early stages associate with lower margins, deep integration with higher ones—while capital expenditure intensity and revenue-per-employee productivity show no clear link. The work matters because large public firms are bellwethers for broader economy-wide productivity and labor-market effects of AI.","feed_headline":"Only 21% of S&P 500 firms run AI in production by 2025","feed_subtitle":"Deep integration quadrupled since 2022; tech leads, profits follow a J-curve, capex and productivity do not","key_machinery":"A two-step 10-K text measure: keyword filtering of AI-related paragraphs followed by GPT classification against an ordinal rubric (1 = no mention, 2 = exploration, 3 = pilot, 4 = used in production with financial expectations, 5 = deep strategic embedding). Legal constraints on 10-K accuracy are used to distinguish integration from hype.","core_discovery":"Using a novel 1–5 rubric applied to AI-related paragraphs extracted from SEC 10-K filings, the authors show that in 2025 eleven percent of S&P 500 enterprises had AI deeply integrated into business processes and a further ten percent used AI in production of goods and delivery of services. Combined advanced adoption more than quadrupled from five percent in 2022. Technology-sector firms account for two-thirds of deep integration; non-technology firms adopt more slowly. Across firms, net profit margins display a J-curve from no adoption to deep adoption, while capital-expenditure intensity and revenue-per-employee productivity show no significant differences. Among technology firms only, deep","pith_inferences":["If the J-curve is real, non-technology firms that stay at score 3 for several years may show persistent margin compression until process redesign catches up.","The sharp post-2022 jump implies that foundation-model APIs lowered the fixed-cost barrier for production use, so further model cost declines should accelerate non-tech score-4 transitions.","Because deep adoption is still rare outside tech, cross-industry production-network risk from AI failures remains limited for now but will rise as score-4/5 shares grow.","Revenue-per-employee’s lack of association suggests future work should track task-level automation and skill mix rather than headcount alone."],"forward_implications":["Aggregate productivity gains from AI will remain modest until non-technology sectors move beyond pilots into production use.","Observed profitability J-curves imply early adopters may report weaker margins before later gains appear, so short-run financial screens can mis-rank AI progress.","Capex intensity is not a reliable proxy for most firms’ AI adoption because most buy models and cloud services as operating expenses rather than capital assets.","Labor-market effects are likely to concentrate first among large technology employers where deep adoption and headcount already co-vary.","Market valuations (Tobin’s q) currently price AI more as a technology-sector phenomenon than as firm-specific deep integration outside tech."],"fun_headline_variants":["21% of S&P 500 firms run AI in production or deeper by 2025","Deep AI use in S&P 500 quadruples to 21% advanced adoption since 2022","Tech sector claims two-thirds of deep AI integration among S&P 500","S&P 500 AI adoption shows profit J-curve, no capex or productivity shift","11% of S&P 500 deeply integrate AI; another 10% use it in production"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The automated scores of filtered 10-K paragraphs against the five-level rubric recover true enterprise-level deep AI integration rather than selective disclosure, marketing language, or residual hype.","fun_headline_variants_meta":{"raw":{"variants":["21% of S&P 500 firms run AI in production or deeper by 2025","Deep AI use in S&P 500 quadruples to 21% advanced adoption since 2022","Tech sector claims two-thirds of deep AI integration among S&P 500","S&P 500 AI adoption shows profit J-curve, no capex or productivity shift","11% of S&P 500 deeply integrate AI; another 10% use it in production"]},"model":"grok-4.5","effort":"low","cost_usd":0.005642,"raw_usage":{"total_tokens":1589,"prompt_tokens":877,"num_sources_used":0,"completion_tokens":123,"cost_in_usd_ticks":56420000,"prompt_tokens_details":{"text_tokens":877,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":589,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":877,"tokens_out":123,"duration_ms":6208,"temperature":1.0,"reasoning_tokens":589,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T05:47:30.445672+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Independent audits or structured surveys of the same S&P 500 firms that show systematically different rank orderings of deep integration versus the 10-K rubric scores, especially for non-technology firms claiming score 4 or 5.","supporting_citations":[],"review_version":1}