Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

AI Adoption in S&P 500 Firms

T0 review · 3 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read By 2025 only 21% of S&P 500 firms had AI in production or deep process integration, up fourfold since 2022, with tech firms driving most of the deep cases and profitability tracing a J-curve.

desk verdict Usable S&P 500 deep-AI panel from 10-Ks; rates and J-curve are real contributions, but the LLM classifier is only lightly validated. read the letter →

arxiv 2607.08920 v1 pith:AWGBRWYL submitted 2026-07-09 econ.GN q-fin.EC

classification econ.GNq-fin.EC
keywords AIadoptionS&P50010-KfilingsenterpriseJ-curveproductivityTobin'sqtechnologysector
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds a firm-year measure of deep enterprise AI adoption from SEC 10-K filings for S&P 500 companies from 2016 to 2025. It separates real integration of AI into business processes from mere exploration or hype, scoring firms from no mention through pilot use to production deployment and full strategic embedding. The central finding is that deep adoption remains limited: in 2025 only 11% of firms scored at full integration and another 10% used AI in production of goods or services, yet the combined share more than quadrupled from 2022 levels. Technology firms account for roughly two-thirds of the deepest adoption and show aggressive uptake, while non-technology firms move more slowly. Profitability follows a J-curve—early stages associate with lower margins, deep integration with higher ones—while capital expenditure intensity and revenue-per-employee productivity show no clear link. The work matters because large public firms are bellwethers for broader economy-wide productivity and labor-market effects of AI.

What carries the argument

A two-step 10-K text measure: keyword filtering of AI-related paragraphs followed by GPT classification against an ordinal rubric (1 = no mention, 2 = exploration, 3 = pilot, 4 = used in production with financial expectations, 5 = deep strategic embedding). Legal constraints on 10-K accuracy are used to distinguish integration from hype.

What would settle it

Independent audits or structured surveys of the same S&P 500 firms that show systematically different rank orderings of deep integration versus the 10-K rubric scores, especially for non-technology firms claiming score 4 or 5.

Watch

Extended reading notes

Core claim

Using a novel 1–5 rubric applied to AI-related paragraphs extracted from SEC 10-K filings, the authors show that in 2025 eleven percent of S&P 500 enterprises had AI deeply integrated into business processes and a further ten percent used AI in production of goods and delivery of services. Combined advanced adoption more than quadrupled from five percent in 2022. Technology-sector firms account for two-thirds of deep integration; non-technology firms adopt more slowly. Across firms, net profit margins display a J-curve from no adoption to deep adoption, while capital-expenditure intensity and revenue-per-employee productivity show no significant differences. Among technology firms only, deep

Load-bearing premise

The automated scores of filtered 10-K paragraphs against the five-level rubric recover true enterprise-level deep AI integration rather than selective disclosure, marketing language, or residual hype.

Editorial extensions

If this is right

  • Aggregate productivity gains from AI will remain modest until non-technology sectors move beyond pilots into production use.
  • Observed profitability J-curves imply early adopters may report weaker margins before later gains appear, so short-run financial screens can mis-rank AI progress.
  • Capex intensity is not a reliable proxy for most firms’ AI adoption because most buy models and cloud services as operating expenses rather than capital assets.
  • Labor-market effects are likely to concentrate first among large technology employers where deep adoption and headcount already co-vary.
  • Market valuations (Tobin’s q) currently price AI more as a technology-sector phenomenon than as firm-specific deep integration outside tech.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the J-curve is real, non-technology firms that stay at score 3 for several years may show persistent margin compression until process redesign catches up.
  • The sharp post-2022 jump implies that foundation-model APIs lowered the fixed-cost barrier for production use, so further model cost declines should accelerate non-tech score-4 transitions.
  • Because deep adoption is still rare outside tech, cross-industry production-network risk from AI failures remains limited for now but will rise as score-4/5 shares grow.
  • Revenue-per-employee’s lack of association suggests future work should track task-level automation and skill mix rather than headcount alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper constructs a firm-year AI adoption measure for S&P 500 firms (2016–2025) by keyword-filtering 10-K paragraphs and scoring them 1–5 with GPT-5-mini against a rubric that distinguishes no mention, exploration, pilots, production use, and deep process integration. It reports that by 2025, 11% of firms score 5 and 21% score 4 or 5 (up from ~5% in 2022), with technology firms accounting for about two-thirds of deep integration. Fixed-effects regressions of Compustat outcomes on these scores recover a J-curve in net profit margin (especially for non-tech), little association with capex-to-revenue or revenue-per-employee, and positive associations of adoption with headcount and Tobin’s q mainly among technology firms. The authors emphasize descriptive correlations and flag endogeneity.

Significance. If the 10-K scores are reliable, the paper supplies a timely, transparent, enterprise-level panel of deep AI adoption for the largest U.S. public firms, with external industry-level validation against BTOS (corr 0.76) and Ramp (0.72–0.87) and clear sector and time patterns that are hard to obtain from surveys alone. The measurement contribution and the descriptive facts on post-ChatGPT acceleration, tech concentration, and the profitability J-curve would be useful inputs for productivity, labor, and finance work on AI. Strengths include the a priori rubric and keyword list, public-filing basis, and explicit non-causal framing of the outcome regressions.

major comments (3)
  1. [Section 2.1, 2.1.1] Section 2.1 and 2.1.1: The central claims (11%/21% rates, quadrupling, tech share of deep adoption, and all Tables 5–9 results) rest on GPT-5-mini ordinal scores of keyword-filtered 10-K text. Validation is only industry-level correlations with BTOS and Ramp; there is no firm-level gold-standard sample, inter-rater reliability, or confusion matrix for scores 3 vs 4 vs 5. Manual spot-checks and Magnificent-Seven consistency are insufficient for a load-bearing classifier. A modest human-coded subsample (or multi-model agreement) with reported precision/recall by score is needed before the rates and J-curve can be treated as reliable.
  2. [Table 5, Fig. 3] Table 5 (preferred cols 4 and 6) and Fig. 3: The non-tech J-curve for net profit margin is driven by a small number of score-5 non-tech observations (text notes ~19 firms and wide CIs). With firm FE, the score-5 coefficient is large and significant, but the cell is thin and sensitive to outliers (as the authors themselves show for Tobin’s q with ALGN/PAYC in Table 7). Report cell counts by sector×score, re-estimate with winsorization or leave-one-out, and qualify the non-tech deep-integration claim accordingly.
  3. [Abstract; Eqs. (1)–(3); Section 5] Eqs. (1)–(3) and Section 3: The authors correctly state that estimates are correlations, not causal. The abstract and conclusion still lead with the J-curve and “no differences in capex or productivity” as if they are robust adoption effects. Tighten abstract/conclusion language so that measurement facts are primary and outcome associations are clearly descriptive; the lagged-AI and sector-year FE checks help but do not resolve reverse causality or time-varying confounders.
minor comments (5)
  1. [Section 2 (sector definition)] Manual reclassification of Amazon and Tesla (and other GICS adjustments) into technology is consequential for the “tech two-thirds” claim; document the full list of reclassifications and show robustness without them.
  2. [Table 1; Appendix C] Table 1 and Appendix C: report N by AI score×sector×year so readers can see how thin the high-score cells are, especially non-tech score 5.
  3. [Fig. 2; Table 12] Fig. 2 / Table 12: clarify whether bars are shares of the contemporaneous S&P 500 panel or a fixed firm set; index composition changes over 2016–2025.
  4. [Section 2.1; Appendix B] Prompt and model: state temperature/seed (or determinism settings) and whether scores were re-run for stability; Appendix B keyword list is useful but coverage of generative-AI terms post-2022 could be noted.
  5. [Throughout] Minor typos and consistency: e.g., “non-techology,” “regress-ing,” and mixed use of “genAI” vs “AI”; ensure 2025 10-K coverage is complete or flag partial-year filings.

Circularity Check

0 steps flagged · score 1.0 of 10

Measurement-plus-correlation paper with no derivation that folds its own inputs back into the claim; only minor self-citation of prior AI-risk work that is not load-bearing for the adoption rates or J-curve.

full rationale

The paper constructs an a-priori 1–5 rubric (Table 2) and keyword list (Appendix B), extracts AI-related paragraphs from 10-Ks, scores them with GPT-5-mini, and then reports descriptive adoption rates and fixed-effects correlations with Compustat outcomes (Eqs. 1–3, Tables 5–9). The scores are not fitted to the outcome variables, nor are any parameters estimated from the same data later re-labeled as predictions. Validation is external (industry-level BTOS correlation 0.76 and Ramp 0.72–0.87). Self-citations (Acemoglu et al. 2022 for keywords; authors’ own AI-risk and computer-vision papers for motivation) supply background only; none is used as a uniqueness theorem or as the sole justification for the central descriptive claims. The authors repeatedly state that results are correlational, not causal. Consequently there is no circular reduction of the form ‘X derives Y but X is defined from Y’ or ‘fitted input called prediction.’ Score 1 reflects only the ordinary presence of non-load-bearing self-citation.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

Empirical measurement paper. Load-bearing elements are the constructed ordinal score, the keyword filter, the LLM classifier, the legal-text reliability premise, and standard panel FE assumptions. No free physical constants or new particles; free parameters are design choices of the measurement system.

free parameters (3)
  • AI adoption rubric cut-points (scores 1–5)
    Ordinal thresholds that map textual evidence into discrete levels; chosen by authors and refined manually; directly determine all reported adoption rates and regression bins.
  • AI keyword list for paragraph extraction
    Pre-filter that decides which text reaches the classifier; drawn from Acemoglu et al. (2022) plus NeurIPS/Gao-Wang terms and manual additions; incomplete coverage would bias scores downward.
  • GPT-5-mini classification temperature/seed and prompt wording
    Stochastic model choices that affect score assignment; not fully specified for exact replication.
assumptions (4)
  • domain assumption SEC 10-K language is a reliable, low-hype signal of material AI adoption because regulations prohibit materially false or misleading statements.
    Stated in abstract and introduction; underpins the claim that the measure distinguishes deep integration from hype.
  • domain assumption Firm and year (or sector-year) fixed effects absorb time-invariant firm traits and common shocks sufficiently for the reported associations to be informative.
    Standard panel assumption invoked in equations (1)–(3); authors note residual endogeneity remains.
  • domain assumption Revenue per employee is a usable proxy for firm-level productivity in this setting.
    Section 2.2.6; used for the null productivity result.
  • ad hoc to paper Manual reclassification of Amazon and Tesla into the technology sector (and other GICS adjustments) correctly reflects economic substance.
    Section 2; affects the tech vs non-tech split that drives many heterogeneous results.
invented entities (1)
  • 1–5 deep AI adoption score (rubric-based) independent evidence
    purpose: Ordinal measure of enterprise AI integration from no mention to core strategic embedding, constructed to separate production use from pilots and hype.
    Central invented object; all headline rates and regressions are defined on it. Independent evidence is partial via BTOS/Ramp correlations, but the exact ordinal mapping is paper-specific.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Adoption in S&P 500 Firms." pith.science (2026). https://pith.science/paper/AWGBRWYL

@misc{pith2026260708920,
  author       = {Pith},
  title        = {Pith review of: AI Adoption in S&P 500 Firms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AWGBRWYL}},
  note         = {Machine review of arXiv:2607.08920}
}
read the original abstract

The adoption of artificial intelligence (AI) by large enterprises is an important potential source of aggregate productivity improvement and labor market impact. We study AI adoption of S&P 500 firms over the period 2016 to 2025, estimating adoption at the enterprise level. While generative AI tools are useful for personal and professional applications, our focus is on the deep integration of AI in the business processes of large enterprises which are bellwethers for firm adoption more broadly. We develop a novel measure to assess deep AI adoption (and distinguish it from AI hype) that is based on SEC 10-K filings, where laws and regulations ``prohibit companies from making materially false or misleading statements." In 2025, 11% of S&P 500 enterprises had AI deeply integrated into their business processes, and a further 10% were using AI in the production of goods and delivery of services. AI adoption has more than quadrupled from 5% in 2022 with slowly accelerating adoption among non-technology firms but very aggressive adoption in the technology sector which accounts for two-thirds of deeply integrated enterprise adoption. Firm profitability shows a "J-curve" as firms move from no adoption to deep adoption, but we observe no differences in capex or productivity. Among technology firms, but not others, AI adoption is higher for firms with more employees and higher values of Tobin's q.

Figures

Figures reproduced from arXiv: 2607.08920 by the authors.

Figure 1
Figure 1. Industry-Level AI Adoption Validations (a) 10-K vs Business Trends and Outlook Survey (b) 10-K vs Ramp An alternative data source for validation is Ramp, a US-based fintech company that provides corporate card and pay bill platforms to tens of thousands of businesses in the 10 [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 3
Figure 3. reports the mean net profit margin of firms at each level of AI adoption from 2016 to 2025. In the technology sector, firms at AI score 1 (no or very limited adoption), report an average net profit margin of 14.4%. Among those firms with an AI score of 3, the product￾level deployment phase, the average profit margin is 12.4%. However, among those firms with deep adoption and score a 5 the profit margin is 17.4%. On … view at source ↗
Figure 4
Figure 4. AI Adoption and Capex-to-Revenue Note: This figure reports the mean capex-to-revenue ratio of firms at each level of AI adoption from 2016 to 2025. Error bars represent one standard error of the mean estimates. The error bar for AI score 5 in the non-technology sector is notably wide due to the small number of observations in that category. Since AI-related capital expenditures are likely concentrated among a small … view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Capex-to-Revenue Ratio of Magnificent Seven Note: This figure reports the mean capex-to-revenue ratio of firms inside and outside the Magnificent Seven from 2016 to 2025, the error bars represent one standard error. 11https://finance.yahoo.com/news/heres-why-metas-135-…
Figure 6
Figure 6. Figure 6: reports the mean Tobin’s Q of firms at each level of AI adoption from 2016 to 2025. We observe two salient patterns. First, technology firms on average have a higher Tobin’s Q than non-technology firms across all AI adoption levels. Second, Tobin’s Q tends to be higher…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Organizations Use AI: Evidence from ChatGPT

    econ.GN 2026-08 conditional novelty 6.0 of 10

    Using proprietary telemetry from ChatGPT Enterprise linked to Compustat, the authors document that enterprise AI adoption favors large intangible-intensive firms, that use spreads across job functions, and that early-...

Reference graph

Works this paper leans on

29 extracted references · 1 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Management Science , volume=

    A text-based analysis of corporate innovation , author=. Management Science , volume=. 2021 , publisher=

  2. [2]

    2005 , publisher =

    Managing Innovation: Integrating Technological, Market and Organizational Change , author =. 2005 , publisher =

  3. [3]

    Journal of Labor Economics , volume=

    Artificial intelligence and jobs: Evidence from online vacancies , author=. Journal of Labor Economics , volume=. 2022 , publisher=

  4. [4]

    Nature human behaviour , volume=

    Quantifying the use and potential benefits of artificial intelligence in scientific research , author=. Nature human behaviour , volume=. 2024 , publisher=

  5. [5]

    2023 , doi =

    Regulating Transformative Technologies , author =. 2023 , doi =

  6. [6]

    The Quarterly Journal of Economics , volume =

    Competition and Innovation: An Inverted-U Relationship , author =. The Quarterly Journal of Economics , volume =

  7. [7]

    The Review of Economic Studies , volume =

    Competition, Imitation and Growth with Step-by-Step Innovation , author =. The Review of Economic Studies , volume =

  8. [8]

    Econometrica , volume =

    A Model of Growth through Creative Destruction , author =. Econometrica , volume =

Show all 29 references
  1. [9]

    Journal of Economic Perspectives , volume =

    Why Do Management Practices Differ across Firms and Countries? , author =. Journal of Economic Perspectives , volume =

  2. [10]

    The Quarterly Journal of Economics , volume =

    Information Technology, Workplace Organization, and the Demand for Skilled Labor: Firm-Level Evidence , author =. The Quarterly Journal of Economics , volume =

  3. [11]

    Journal of Economic Perspectives , volume =

    From Micro to Macro via Production Networks , author =. Journal of Economic Perspectives , volume =

  4. [12]

    Science , volume =

    GPTs Are GPTs: Labor Market Impact Potential of LLMs , author =. Science , volume =

  5. [13]

    2024 , month = aug, day =

    The Last Mile Problem in AI: Why Job Automation Will Be Slower than Technological Progress Suggests , author =. 2024 , month = aug, day =

  6. [14]

    2025 , month = jun, day =

    Generative AI Part XI: Agentic AI Expands the APP Software TAM , author =. 2025 , month = jun, day =

  7. [15]

    Productivity: Information Technology and the American Growth Resurgence , author =

  8. [16]

    2025 , month = dec, day =

    Organizational Transmission of AI: Managerial Influence on Generative AI Adoption , author =. 2025 , month = dec, day =

  9. [17]

    2019 , howpublished =

    A Constructive Prediction of the Generalization Error across Scales , author =. 2019 , howpublished =

  10. [18]

    Workforce , author =

    Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce , author =. 2025 , howpublished =

  11. [19]

    2024 , howpublished =

    Beyond AI Exposure: Which Tasks Are Cost-Effective to Automate with Computer Vision? , author =. 2024 , howpublished =

  12. [20]

    Journal of Economic Literature , volume=

    The Economic Impacts of Artificial Intelligence: A Multidisciplinary, Multi-book Review , author=. Journal of Economic Literature , volume=. 2026 , publisher=

  13. [21]

    American Economic Journal: Macroeconomics , volume=

    The productivity J-curve: How intangibles complement general purpose technologies , author=. American Economic Journal: Macroeconomics , volume=. 2021 , publisher=

  14. [22]

    2025 , publisher=

    The Rise of Industrial AI in America: Microfoundations of the Productivity J-curve (s) , author=. 2025 , publisher=

  15. [23]

    Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks , author=

    Crashing Waves vs. Rising Tides: Preliminary Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks , author=. arXiv preprint arXiv:2604.01363 , year=

  16. [24]

    Journal of Economic literature , volume=

    What determines productivity? , author=. Journal of Economic literature , volume=. 2011 , publisher=

  17. [25]

    American Economic Review: Insights , volume=

    Regulating transformative technologies , author=. American Economic Review: Insights , volume=. 2024 , publisher=

  18. [26]

    arXiv preprint arXiv:2408.12622 , year =

    The AI Risk Repository: A Comprehensive Meta-Review, Database, and Taxonomy of Risks From Artificial Intelligence , author =. arXiv preprint arXiv:2408.12622 , year =

  19. [27]

    2026 , url =

    Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts , author =. 2026 , url =

  20. [28]

    Princeton University , year=

    Robot adoption and labor market dynamics , author=. Princeton University , year=

  21. [29]

    The Aggregation Paradox of AI: Why do micro-economic productivity gains from AI disappear at scale , author=

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.