{"id":"880efae7-9de1-46b6-b4a8-bc8b7f1a4eb5","arxiv_id":"2608.12236","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using proprietary telemetry from ChatGPT Enterprise linked to Compustat, the authors document that enterprise AI adoption favors large intangible-intensive firms, that use spreads across job functions, and that early-career workers are the most intensive users.","lead":"This paper links internal ChatGPT Enterprise records to company financial data, job titles, and message-level task classifications to document how over 1,500 organizations use generative AI at work. It finds that adoption concentrates in larger, R&D-intensive firms, while within firms early-career workers and a broad mix of knowledge tasks account for the heaviest use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The firm-level adoption facts rest on an unvalidated LLM account-to-ticker crosswalk; if its recall varies with firm size, the size and intangibles gradients in Tables 1, 3, and 4 are not identified.","rationale":"Good faith reading: the paper is a careful descriptive study with clear limitations, and the four facts are plausible. But the most consequential evidence is the public-company adoption gradient. All of it flows through the unvalidated ticker bridge. The central question is not whether the bridge is perfect (no bridge is), but whether its error is classical. Classical false matches attenuate; non-classical, size-correlated recall biases. The paper's own footnote assumes the former without evidence. This is a demand for a specific validation, not a claim of fraud or sloppiness. The worker-level sample representativeness and task classifier accuracy are real but secondary; the paper flags those limitations clearly and the facts are phrased conditionally. The crosswalk issue is less flagged. If validation shows flat recall, no change. If it shows differential recall, Tables 1/3/4 need qualification. Since the reader already made acceptance conditional on such validation, I recommend no change to the verdict.","tokens_in":26562,"tokens_out":7501,"duration_ms":73415,"concrete_test":"Internally, draw a stratified random sample of 300 accounts that the bridge maps to tickers and 300 public firms with no match, stratified by revenue decile. Have annotators blind to the bridge verify each account's public-company identity from official domains, SEC filings, and public registries. Compute bridge precision and recall by firm-size tercile. Then re-estimate Tables 1, 3, and 4 under two scenarios: (i) using only verified matches, and (ii) applying non-adopter-to-adopter misclassification rates that vary by size, e.g., inverse-probability weighting or Lee bounds. If the positive size, R&D, and SG&A coefficients persist under both, the concern is resolved; if they shrink or flip, Fact 2 must be restated as conditional on the bridge's differential recall. Report the recall-by-size table in the paper.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central firm-level claim (Fact 2, and the adoption regressions in Tables 1, 3, and 4) depends on the Section 3.3 account-to-ticker crosswalk and on defining non-adopters as public firms with no bridge match. The paper reports no precision or recall for this LLM-assisted bridge. The probability that a true adopter is observed as an adopter is the bridge's recall, and recall is likely correlated with firm observables: large firms have more public coverage, more subsidiaries, and cleaner name-to-ticker mappings. If recall rises with revenue, the matched adopter sample overrepresents large, high-intangible firms even when adoption is unrelated to those characteristics. Footnote 10 only discusses attenuation from non-adopters using other AI products; it does not bound differential measurement error in the bridge. Since 410 adopters are compared with 11,784 unmatched firms, even a small false-negative rate concentrated among smaller firms can generate the reported gradients. The worker-level and task facts are secondary because they are explicitly framed as conditional on the covered sample; the public-company facts are presented as unconditional facts about adoption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses administrative telemetry from ChatGPT Enterprise, linked to employee job titles, a message-level task classifier, and Compustat financials, to document four descriptive facts: (i) enterprise usage grew rapidly between June 2025 and March 2026, with substantial growth within fixed adoption cohorts; (ii) U.S. public-company adopters are larger, higher-revenue-per-employee, and more R&D- and SG&A-intensive than non-adopters; (iii) among firms with usable title data, active use spans job functions and seniority levels, with early-career workers showing the highest messages-per-active-user; and (iv) task composition is broad, with writing, communication, and information tasks most common and with industry and role variation. The authors are careful to frame the regressions as conditional associations and to state several data limitations.","tokens_in":26735,"tokens_out":7112,"duration_ms":64087,"significance":"If the measurement concerns can be addressed, this is an unusually rich descriptive contribution: it moves beyond surveys to observe actual enterprise usage at scale, with 1,764 organizations and 17.4 million messages in the worker sample, and it connects usage to firm financials, roles, and tasks. The paper is also commendably explicit about what it does not measure: no workforce denominators, no downstream outcomes, single-vendor scope, and incomplete title coverage. The four facts, especially the within-firm worker and task heterogeneity, would be useful calibration for theories of GPT diffusion and for future work on complements. The main risk is that Fact 2, and to a lesser extent Facts 3 and 4, could be partly generated by measurement error in the account-to-ticker bridge and in title coverage rather than by true adoption patterns.","major_comments":[{"comment":"The central firm-level results rest on the LLM-assisted account-to-ticker crosswalk, but the paper reports no precision or recall for this bridge. Because non-adopters are defined as public firms with no bridge match, and because only 410 matched tickers have positive usage versus 11,784 unmatched tickers, even a small false-negative rate that is correlated with size, revenue, or intangible intensity can generate the documented gradients in Tables 1, 3, and 4. Footnote 10 discusses attenuation from non-adopters using other AI products, which is a different source of error and does not bound differential measurement error in the bridge itself. Please report validation results for the crosswalk, for example a hand-audited sample with precision and recall by firm size and industry, and provide a sensitivity analysis that bounds the adoption probability under alternative assumptions about false-negative rates.","section":"Section 3.3, Tables 1, 3, and 4"},{"comment":"The worker-level facts in Section 4.3 are stated as general facts about enterprise AI use, but the underlying sample requires high-quality job title information and an active organization-week at week 26. The paper notes in Section 3.2 that title coverage is incomplete, yet it never reports the coverage rate, how it varies by organization size or industry, or whether the 1,764-organization sample is representative of the full Enterprise population. If title coverage is higher among firms with more IT administration, or among particular roles, the composition shares in Figure 5 and the seniority gradient in Panel B of Figure 6 could reflect selection. Please report the fraction of active users with usable titles, decompose missingness by role and firm size, and re-estimate the main figures on a high-coverage subsample.","section":"Section 3.2, Figures 5 and 6"},{"comment":"The task-classification analysis in Section 4.4 relies on a classifier available only from October 30, 2025, and measures use at week 26 after adoption. This combination implies that the task sample of 973 organizations can only include firms whose week-26 horizon falls after that date, effectively restricting the analysis to organizations that adopted in a particular window (roughly May through October 2025). The paper does not report the adoption-date distribution of the task sample or compare it with the worker sample or the full Enterprise sample. Because Section 4.1 documents a simultaneous acceleration across cohorts in early 2026, the task composition measured in this window may not generalize to earlier or later adopters. Please document the cohort composition of the task subsample and, if possible, show task results for multiple adoption horizons.","section":"Section 3.2 and Section 4.4"}],"minor_comments":[{"comment":"The random subsample of matched accounts is not described in terms of the sampling rate or whether sampling weights are used; please clarify whether all regression analyses use the unweighted random sample and whether any disclosure-related subsampling affects precision.","section":"Section 3.3"},{"comment":"The SG&A and R&D stocks rely on assumed depreciation rates of 20 percent and 15 percent and a zero-growth steady-state opening stock; please report sensitivity to alternative depreciation rates and opening-stock assumptions.","section":"Section 4.2.4 and Table 4"},{"comment":"The category 'Other / unknown' pools missing, malformed, and genuinely unclassifiable job titles; because the reported shares for substantive categories are sensitive to the size of this residual, please report the unclassified share explicitly or show results excluding unclassified users.","section":"Figures 5 and 6"},{"comment":"The sevenfold and fourfold growth figures are stated in the text but the underlying indexed series are not tabulated; please include a small table of index values in an appendix.","section":"Section 4.1 and Figure 3"},{"comment":"The figure notes describe firm-bootstrap confidence intervals in Figure 5 and firm-clustered standard errors in Figure 6; please specify the bootstrap procedure, including the number of replications and whether resampling is clustered by firm.","section":"Section 4.3, figure notes"}],"recommendation":"major_revision","confidential_remarks":"The authors are employees or contractors of OpenAI and the data are proprietary, and the paper cites several related vendor-telemetry papers. This is not by itself a flaw, since disclosure of vendor data is inherently limited, but it increases the importance of validating the crosswalk and providing enough detail for external auditing. I would not reject on this basis, but I would ask the editor to ensure the revision includes the validation and sensitivity material described in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first large-sample administrative look at how organizations actually use ChatGPT Enterprise, and the strongest part of the paper is the within-firm stuff—who uses it, how intensively, and for what tasks. The early-career intensity gradient and the breadth of task use across job functions will be the facts people remember. The second thing: the firm-level adoption results in Tables 1, 3, and 4 rest on an account-to-ticker crosswalk whose precision and recall are not reported, and that is a real problem, not a nitpick.\n\nWhat the paper does well: the data are genuinely new. 1,500+ organizations, 17 million messages, linked to administrative job titles and task classifications. The four facts are clearly written and mostly honestly caveated. The authors repeatedly note that the worker-level results are conditional on covered samples and lack workforce denominators, and they present the regressions as conditional associations rather than causal claims. That is the right posture.\n\nWhere the soft spots are: the Compustat-company analysis is the load-bearing part for Fact 2. The paper defines non-adopters as public firms with no match in the bridge, but never tells us how good the bridge is. If recall is lower for small firms—which is likely, since they have less public coverage and messier name-to-ticker mappings—then the matched adopter sample overrepresents large firms and the size and intangibles gradients are inflated. Footnote 10 discusses attenuation from non-adopters who use other AI products, but it does not address differential measurement error in the bridge itself. With 410 adopters against 11,784 unmatched firms, even a small false-negative rate concentrated among small firms could move the coefficients noticeably. That does not mean the qualitative pattern is false, but the paper currently cannot rule out that the magnitude is driven by the bridge.\n\nTwo slightly smaller points: the task classifier is described as validated against an internal benchmark, but no numbers are given, and the job-title validation is only top-five lists per category. These are minor relative to the bridge, but they are easy fixes.\n\nBottom line: this is a paper for anyone studying AI adoption or the future of work, and the descriptive facts will be cited. It deserves a serious referee. The referee should return with a request for validation of the crosswalk—either a hand-labeled precision/recall sample or bounds under differential recall. I would want that before treating the firm-size facts as established.\n\nSerious thinker: yes. The authors are honest about the scope of the data and the limits of what they can see. The missing validation is a gap in the paper, not a sign of sloppiness.","headline":"A serious descriptive paper whose worker-level and task facts are likely to stick, but whose firm-size gradient depends on an unvalidated account-to-ticker crosswalk that the authors need to address.","tokens_in":27314,"tokens_out":3805,"would_cite":true,"duration_ms":33239,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Enterprise AI use grows fast, skews to large R&D-intensive firms, and spans every job level.","keywords":["enterprise AI adoption","ChatGPT Enterprise","generative AI","usage telemetry","job title classification","task classification","knowledge work","intangible capital"],"falsifier":"Audit the crosswalk: take a random sample of matched and unmatched U.S. public firms, verify against public procurement records or a separate enterprise-customer list whether each truly runs a ChatGPT Enterprise workspace, and compare the resulting error rates; if unmatched firms are often actual adopters, the adopter/non-adopter regressions in Tables 1, 3, and 4 would be biased.","tokens_in":26335,"feed_emoji":"📈","tokens_out":8616,"duration_ms":69586,"temperature":0.7,"pith_summary":"This paper uses de-identified ChatGPT Enterprise telemetry—more than 1,500 organizations and 17 million messages at the six-month adoption horizon—to establish four descriptive facts about how firms actually use frontier generative AI. Enterprise usage grew roughly sevenfold between June 2025 and March 2026, through both new adoption and deepening use inside existing customers. U.S. public-company adopters are larger, more valuable, and more R&D- and SG&A-intensive than non-adopters, and within adopting firms active use spans job functions and seniority levels, with early-career workers the most intensive users. Usage also covers a broad range of knowledge-work tasks rather than a single workflow. The point of these facts is to show that adoption is only the beginning of deployment: firms are still learning where AI belongs in their workflows.","feed_headline":"Enterprise ChatGPT use grew sevenfold in nine months","feed_subtitle":"New telemetry shows adoption spans every job level and task, with early-career workers using it most.","key_machinery":"The central machinery is the linked telemetry panel: an organization-week record of ChatGPT Enterprise adoption and usage joined to worker-level job-title metadata, a message-level task classifier, and, for public firms, annual financial data through an LLM-assisted account-to-ticker crosswalk. The job-title classifier maps raw administrative titles to departments, seniority levels, and manager status; the task classifier assigns each message to one of 60 work-task categories. These links allow the paper to measure the extensive margin (which firms and workers adopt), the intensive margin (how much they use), and task breadth within a single privacy-preserving dataset, which is what carries all four facts.","core_discovery":"The paper's central claim is that four stylized facts about enterprise AI adoption and use hold in the ChatGPT Enterprise workspaces it observes. First, output tokens generated by enterprise customers grew roughly sevenfold between June 2025 and March 2026, and about fourfold within a fixed cohort of firms that had already adopted by June 2025, so roughly half of aggregate growth came from deepening use rather than new customers. Second, among U.S. public companies, adopters are larger, more valuable, and more intensive in R&D and SG&A spending per employee, with the largest firms in each industry significantly more likely to adopt. Third, six months after adoption, active users are spread across job functions and seniority levels, but intensity is uneven: early-career workers and trainees send roughly eight to nine more weekly messages than the average active user in the same firm, while executives send fewer. Fourth, message-level classification shows a long tail of knowledge-work tasks—writing, technical work, communication, and synthesis—rather than a single dominant use, with task mix varying across industries and roles. The authors emphasize that these are conditional associations, not causal effects.","pith_inferences":["A natural next step the authors leave open is following the same worker cohorts beyond six months to see whether the early-career intensity gap is a learning phase that fades or a stable feature of how junior staff use AI.","A testable prediction from the complements result: among adopters, firms with larger SG&A stocks should show steeper growth in per-employee message volume after adoption, because the same capabilities that predict entry should also predict deployment success.","The paper's task taxonomy could be applied to usage from other enterprise AI products to check whether the observed breadth is specific to ChatGPT Enterprise or common to workplace LLM tools generally.","If the negative seniority gradient in message volume is causal, then productivity studies that average over all workers will understate early-career gains; measuring effects separately by seniority would be a sharper test."],"forward_implications":["Firm-level 'adoption' measured by a signed contract understates deployment: a large share of enterprise growth comes from existing customers using the product more, so studies that count adopters miss the intensive margin.","Early AI adoption is tied to pre-existing intangible capital, so if the pattern holds, generative AI diffusion may widen productivity and value gaps between large, capability-rich firms and the rest.","Because early-career workers are the heaviest users, the productivity and employment effects of generative AI are most likely to show up first among junior employees.","Task breadth across industries and roles supports treating generative AI as a general purpose technology for knowledge work, implying that complements—training, workflow redesign, and organizational change—will determine realized value.","Comparisons of AI use across firms should separate the task-prevalence margin (how many workers try a task) from the message-share margin (where the volume sits), since a task can be widespread but low-volume or narrow but high-volume."],"supporting_citations":[{"why":"Closest prior work on firm AI adoption versus deployment; the paper positions its worker- and task-level facts as an advance over this distinction.","marker":"Bonney et al. (2026)"},{"why":"Supplies the task-based exposure framework and the general-purpose-technology framing used to interpret the breadth of enterprise usage.","marker":"Eloundou et al. 2024"},{"why":"Baseline evidence on how individuals use ChatGPT; the enterprise results extend this to organizational settings.","marker":"Chatterji et al. 2025"},{"why":"Comparable workplace telemetry from another enterprise AI product; provides the benchmark for aggregate workplace use and task composition.","marker":"Counts et al. (2026)"},{"why":"The job-title classifier and seniority analysis are shared with this paper, making it the methodological source for the worker-level results.","marker":"Johnston et al. (2026)"},{"why":"Survey evidence on firm AI adoption that motivates the need for telemetry-based measurement and provides a comparison point.","marker":"McElheran et al. 2024"},{"why":"Provides the industry-year relative scale approach used in Table 3 to show adoption is concentrated among the largest firms even within industries.","marker":"Autor et al. 2020"},{"why":"Frames intangibles as complements to general purpose technologies, motivating the SG&A and R&D stock regressions in Table 4.","marker":"Brynjolfsson et al. 2021"}],"fun_headline_variants":["Enterprise ChatGPT use grew sevenfold in nine months","Large firms lead ChatGPT adoption, early-career workers use it most","ChatGPT at work: sevenfold growth, broad task coverage","Enterprise AI adoption: four facts from ChatGPT telemetry","ChatGPT enterprise use: bigger firms, more messages, wider tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the account-to-ticker crosswalk correctly identifies which public firms have ChatGPT Enterprise, so that labeling a firm a non-adopter because it has no match does not systematically misclassify enterprise users.","fun_headline_variants_meta":{"raw":{"variants":["Enterprise ChatGPT use grew sevenfold in nine months","Large firms lead ChatGPT adoption, early-career workers use it most","ChatGPT at work: sevenfold growth, broad task coverage","Enterprise AI adoption: four facts from ChatGPT telemetry","ChatGPT enterprise use: bigger firms, more messages, wider tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":1969,"prompt_tokens":975,"completion_tokens":994,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":911}},"tokens_in":591,"tokens_out":994,"duration_ms":9513,"temperature":1.0,"reasoning_tokens":911,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:11:10.481351+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit the crosswalk: take a random sample of matched and unmatched U.S. public firms, verify against public procurement records or a separate enterprise-customer list whether each truly runs a ChatGPT Enterprise workspace, and compare the resulting error rates; if unmatched firms are often actual adopters, the adopter/non-adopter regressions in Tables 1, 3, and 4 would be biased.","supporting_citations":[{"cited_title":"and Dinlersoz, Emin and Foster, Lucia S","cited_arxiv_id":null,"evidence_quote":"Closest prior work on firm AI adoption versus deployment; the paper positions its worker- and task-level facts as an advance over this distinction."},{"cited_title":"and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =","cited_arxiv_id":null,"evidence_quote":"Baseline evidence on how individuals use ChatGPT; the enterprise results extend this to organizational settings."}],"review_version":1}