{"id":"1cd9c9e3-01c2-4328-afa0-ab275497ca03","arxiv_id":"2505.00174","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Corporate AI safety research is dominated by pre-deployment alignment and evaluation work, while high-risk deployment topics such as medical error, misinformation, bias, behavioral design, and copyright are measured to be sharply under-researched.","lead":"This study counts and categorizes 1,178 AI safety papers from five companies and six universities, and finds that corporate labs concentrate on pre-deployment alignment and testing while real-world topics such as bias and misinformation get little attention. A smart generalist might read it because it makes a concrete evidence-based case for giving outside researchers access to live AI deployment data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 4%-vs-6% deployment-gap figure is not robust: Table 3's short regex lists plus an unvalidated, inconsistently described LLM classifier can systematically miss high-risk deployment papers.","rationale":"I agree with the reader's conditional verdict but sharpen the condition. The paper's qualitative argument about corporate emphasis on pre-deployment alignment and evaluation is plausible and receives some independent support from citation patterns and examples, so this is not a rejection. The load-bearing issue is strictly quantitative: the 4%-versus-6% claim and the 53-vs-8 misinformation comparison are computed from short keyword lists and an unvalidated LLM classifier, with contradictory model names across figure and text. Because the same pipeline determines both numerator and denominator, a modest number of missed synonyms could materially change the gap. This is an internal correctness risk, not a dispute with the field consensus. A validation appendix with human labels and an expanded-keyword sensitivity table would settle it. Since the reader already set CONDITIONAL and these concerns reinforce rather than redirect that verdict, no change to the verdict is needed.","tokens_in":19365,"tokens_out":6502,"duration_ms":62895,"concrete_test":"Draw a stratified random sample of about 200 safety & reliability papers, balanced by institution and year. Have two annotators independently mark each paper for whether it substantively addresses any of the eight high-stakes areas, using full-text reading and a definitionally expanded synonym list (including healthcare, clinical, patient, credit, banking, insurance, manipulation, nudging, advertising, recruitment, fake news, propaganda). Recompute Table 3 percentages and academic-to-corporate ratios. Separately re-run the o3-mini eight-category classification on this sample and report agreement with human labels (e.g., Cohen's kappa). If the corporate high-stakes percentage rises from 4% by more than 2 percentage points, or if any risk-area ratio such as 53:8 for misinformation changes by more than 30%, the headline gap should be reported with sensitivity bounds or softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Core Finding 3 and Section 3.2 rest on Table 3, where a paper is counted as covering a high-stakes area only if its title or abstract literally contains one of a short regex list. The lists omit common vocabulary: Medical has hospital(s), health insurance, and clinician(s), but not healthcare, clinical, patient, diagnosis, or drug; Finance covers only finance/financial, not credit, loan, banking, insurance, or investment; Behavioral includes persuasion(s)/persuasive, sycophancy, addictive, and reward-hacking but not manipulation, nudging, engagement optimization, or self-harm; Misinfo misses fake news, rumor, propaganda, and election interference; Commercial misses advertising (only adverts/advertisements), e-commerce, and recruitment (only recruiting). Corporate deployment-risk papers written in product-specific language are therefore systematically undercounted, making 4% a lower bound whose tightness is untested. The paper itself uses 'healthcare' in its abstract, yet its medical regex would not catch that term. The classifier is also described inconsistently: Figure 1 caption says GPT 4o-mini, Section 2.4 says GPT o4-mini, and Appendix 6.3 says o3-mini. No human ground-truth labels, inter-annotator agreement, or keyword sensitivity analysis is reported, so the core percentages cannot be checked as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a dataset of 1,178 safety-and-reliability papers drawn from 9,439 generative-AI papers (January 2020–March 2025) by five corporate AI labs and six academic institutions, classifies each paper into one of eight pre/post-deployment categories using an LLM plus regex keyword screening, and reports that corporate AI research is increasingly concentrated in pre-deployment alignment and testing-and-evaluation work. The headline empirical claim is that only 4% of Corporate AI papers (6% of Academic AI papers) address high-stakes deployment domains such as persuasion, misinformation, medical and financial contexts, disclosures, and core business liabilities. The paper concludes with policy recommendations for tiered external access to deployment telemetry and observability of in-market AI systems.","tokens_in":19660,"tokens_out":2712,"duration_ms":29225,"significance":"If the quantitative claims survive validation, the paper provides a timely and policy-relevant measurement of where corporate and academic AI governance research actually sits relative to documented deployment harms. The authors are transparent about their data sources, publish code and data, make their category definitions and classification prompts available, and go beyond prior work by including post-deployment categories and by supplementing OpenAlex with scraped company publications. The distinction between pre-deployment and post-deployment research is a useful organizing frame, and the policy discussion of structured telemetry access is concrete. However, the paper's central percentages rest on an unvalidated automated classification pipeline and on short regex lists with obvious coverage gaps; because these numbers are the paper's main quantitative contribution, the manuscript needs additional validation work before the headline findings can be relied on.","major_comments":[{"comment":"The headline 4%-versus-6% deployment-gap claim is computed directly from the regex lists reported in the Table 3 note, and those lists are too narrow to support the claim as stated. For example, Medical includes hospital(s), health insurance, and clinician(s) but not healthcare, clinical, patient, diagnosis, or drug; Finance includes only finance/financial and not credit, loan, banking, insurance, or investment; Misinfo omits fake news, rumor, propaganda, and election interference; Behavioral omits manipulation, nudging, engagement optimization, and self-harm; and Commercial omits advertising, e-commerce, and recruitment. This means papers written in common product-specific or clinical language are systematically uncounted, and the reported 4% and 6% are lower bounds whose tightness is untested. The paper should report precision and recall of the regex lists against a human-labeled gold standard, and should show sensitivity of the core percentages to expanded keyword sets. Without this, the central quantitative claim cannot be checked as reported.","section":"§3.2, Table 3, Core Finding 3"},{"comment":"The LLM classification pipeline is described inconsistently and is not validated. Figure 1's caption says papers were categorized using GPT 4o-mini, Section 2.4 says GPT o4-mini, and Appendix 6.3 says OpenAI's o3-mini was used; this inconsistency must be resolved because reproducibility requires the exact model version. More importantly, the paper reports no human ground-truth labels, no inter-annotator agreement, and no error analysis for the two-stage keyword-plus-LLM classifier, even though the classifier's output drives the category-level citation comparisons in Figures 1 and 2, the temporal trends in Figure 4, and the regressions described in Section 3.1. The authors should add a validation section with a human-annotated random sample, per-category precision/recall, and a sensitivity analysis showing how the headline results change under alternative classifier settings or thresholds.","section":"§2.4, Appendix 6.3, Figure 1 caption"},{"comment":"The corporate sample is not complete in a way that could affect the paper's denominators and gap ratios. The text notes that Meta website publications were not manually scraped, and Appendix 6.2 states that OpenAlex contains no Anthropic papers, with only Anthropic and OpenAI publications supplemented from company websites. Because Meta and Anthropic are among the five corporate labs analyzed, the omission is material: adding their missing publications could change both the total corporate paper counts and the number of corporate papers in the high-risk deployment categories. The authors should report coverage statistics for each source and institution, and should show that the 4%-versus-6% comparison and the corporate-vs-academic ratios are robust to the inclusion or exclusion of the incomplete source-institution pairs.","section":"§2.4, Table 1, Appendix 6.2"}],"minor_comments":[{"comment":"In the paragraph listing corporate post-deployment examples, \"read-teaming\" should be \"red-teaming\".","section":"§3.2"},{"comment":"The footnote reads \"do revise their models based based on red-teaming and user experience feedback\"; the duplicated \"based\" should be removed.","section":"§1, footnote 2"},{"comment":"The phrase \"incudes most ArXiv papers\" should be corrected to \"includes most arXiv papers\".","section":"Appendix 6.2"},{"comment":"In the bottom panel of Figure 4, the category ordering appears to differ from the top panel without explanation; adding a shared legend order or a note would improve readability.","section":"Appendix 6.1, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is on a topic squarely within the journal's scope and the authors have been unusually transparent about data and code. My main concern is that the quantitative core is currently a lower-bound estimate whose tightness is unknown; the missing human validation and keyword sensitivity analysis are essential rather than cosmetic. I would support publication after the authors either supply that validation or explicitly reframe the central claim as a conservative lower bound with the limitations prominently stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the arXiv preprint \"Real-World Gaps in AI Governance Research\" with the stress-test note in hand. My take: the central pattern is plausible and the paper is genuinely useful, but the headline percentages are not supportable as reported until the measurement pipeline is validated.\n\nWhat's new: the paper builds a 1,178-paper corpus of safety and reliability research from 2020-2025, compares corporate vs academic outputs, and includes post-deployment categories that Delaney et al. explicitly exclude. The finding that corporate labs concentrate on alignment and testing/evaluation while ethics and bias research is mostly academic is consistent with prior work and is worth having documented on this scale. The data and code are on GitHub, and the authors are transparent about their sample construction and fractional authorship weighting. The policy recommendation—structured telemetry access for external researchers—is sensible and well argued.\n\nThe soft spots are real and concentrated in the measurement. The stress-test note lands: Table 3's regex lists are short and miss common terms (healthcare, clinical, credit, manipulation, fake news), so the high-stakes-domain counts are likely lower bounds. The paper acknowledges the abstract says \"healthcare,\" yet the medical regex would not catch it. Additionally, the LLM classifier is described inconsistently—GPT 4o-mini in the Figure 1 caption, GPT o4-mini in Section 2.4, o3-mini in Appendix 6.3—and no human ground-truth labels, inter-annotator agreement, or keyword-recall sensitivity analysis are reported. The sample also excludes Meta website publications and Anthropic is absent from OpenAlex. These are addressable, but they mean the 4% vs 6% comparison and Figure 4 trends cannot be checked as reported. The qualitative direction may survive better validation, but the specific magnitudes are fragile.\n\nThe circularity concern is minor here: the authors' self-citations are motivational, not load-bearing, and the claim is an empirical measurement, not a derived tautology. The literature engagement is fair.\n\nWho this is for: AI governance researchers, funders, and policymakers who need a map of where research effort is concentrated. It deserves a serious referee—the question is important and the dataset is a real contribution—but only conditional acceptance pending a validation appendix with human ground-truth labels and an updated sample. I'd recommend major revision rather than desk rejection, and I'd cite the paper with the caveat that the headline numbers are provisional.","headline":"The paper maps a real and policy-relevant gap in AI governance research, but its headline 4%-vs-6% numbers rest on an unvalidated classification pipeline, so treat the specific magnitudes as provisional.","tokens_in":20156,"tokens_out":1964,"would_cite":true,"duration_ms":21704,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Corporate AI labs publish barely any safety research on real-world harms—most of their safety work stays in pre-release alignment and testing.","keywords":["AI research","alignment","interpretability","commercialization risks","cloud providers","model developers","AI governance","deployment risks"],"falsifier":"Have independent annotators hand-label a stratified random sample of the 1,178 safety-and-reliability papers using the paper's own category definitions, then compare with the automated labels; if deployment-related papers are systematically mislabeled as alignment or testing & evaluation, the 4%-versus-6% gap and the concentration trend would change materially.","tokens_in":19152,"feed_emoji":"🤖","tokens_out":12490,"duration_ms":112555,"temperature":0.7,"pith_summary":"Drawing on a corpus of 1,178 safety-and-reliability papers selected from 9,439 generative AI papers published between January 2020 and March 2025, the paper tries to establish that corporate AI safety research is skewed away from real-world deployment. It reports that only about 4% of corporate safety papers (6% of academic safety papers) address high-stakes deployment domains such as persuasion, misinformation, medical and financial contexts, disclosures, and core business liabilities, while corporate work concentrates on pre-deployment alignment and testing & evaluation. The claim matters because, if correct, society's empirical understanding of AI harms is narrowing even as AI systems spread: the companies best positioned to observe deployed behavior have the least incentive to publish research on it, and independent researchers lack the telemetry to fill the gap. The paper concludes that structured external access to deployment logs, traces, and model artifacts is the necessary remedy.","feed_headline":"Only 4% of corporate AI safety papers study deployed risks","feed_subtitle":"Academic labs do little better at 6%; corporate research stays in alignment and testing while high-stakes harms go unstudied.","key_machinery":"The carrying mechanism is the paper's classification system. A corpus of 9,439 generative AI papers is filtered by keyword lists and an automated language-model classifier into 1,178 safety-and-reliability papers, each assigned to exactly one of eight categories: alignment, testing & evaluation, ethics & bias, privacy & security, interpretability & transparency, policy & governance, post-deployment risks and model traits, and multi-agent/agentic safety. Institutional credit is counted fractionally by authorship, so a paper with four authors from a tracked lab contributes 0.25 to that lab's totals. A second layer of regex keyword searches on titles and abstracts then identifies high-risk deployment contexts (medical, finance, commercial, copyright) and capability areas (misinformation, disclosures, behavioral, accuracy). The gap claims are the output of these two layers, and the year-by-year category shares in Figure 4 trace the growing concentration in alignment and testing & evaluation.","core_discovery":"On the paper's own terms, the central discovery is a quantitative mismatch between research effort and risk exposure. Among the 1,178 safety-and-reliability papers, corporate labs publish most of their governance research in model alignment and testing & evaluation—work that examines models in controlled, pre-deployment settings—while post-deployment concerns such as bias, misinformation, persuasive or addictive design, medical and financial advice, disclosures, and liability-relevant failures receive a small fraction of attention. In total, 217 fractionally adjusted academic papers and 67 corporate papers touch any of these high-risk areas, about 6% and 4% of each group's safety output. The gap widens in specific domains, such as medical and misinformation research, where academic papers outnumber corporate ones by several times. The authors interpret this concentration as the result of commercial incentives and an existential-risk research culture, and they argue that without structured access to deployment telemetry the knowledge deficit will deepen.","pith_inferences":["If the published gap is real, it is likely even wider in private artifacts: system cards, internal red-teaming reports, and product documentation are not in the corpus, so the 4% figure probably overstates the public share of deployment-focused corporate research.","Applying the same eight-category classification to a later time window, or to a jurisdiction with binding transparency rules, would test whether regulatory pressure shifts corporate publication toward deployment risks.","A testable extension of the policy proposal: if safe-harbor telemetry access became operational, one would expect a measurable rise in externally verifiable papers using real-world traces within one to two years.","Because the academic-to-corporate paper ratio is roughly 2.5 to 1 overall, even a modest increase in corporate deployment research could double the public literature in a high-risk area; the binding constraint is data access, not researcher supply."],"forward_implications":["Corporate AI safety research will keep concentrating on alignment and testing & evaluation as long as product competition shapes what labs publish, so the public share of deployment-stage safety research is unlikely to rise on its own.","The high-stakes domains the paper flags—copyright, medical and financial advice, misinformation, and behavioral influence—are already generating lawsuits, making the research gap a live liability gap.","Independent researchers cannot currently study deployed AI harms systematically, because incident databases and leaked chat logs are partial; the empirical base for AI regulation will stay thin without deployment data access.","Structured external access to telemetry (logs, traces, model artifacts), with tiered researcher access and liability safe harbors, is the proposed path to closing the gap.","Widely deployed safeguards such as content moderation and telemetry-based monitoring are almost absent from the public literature, so evidence-based best practices for deployed AI systems are largely missing."],"supporting_citations":[{"why":"Supplies the scraped publication records for two of the corporate labs and the earlier pre-deployment safety-research map that this paper extends.","marker":"[22]"},{"why":"Provides the earlier cluster analysis of AI safety research whose categories this paper adapts into its eight-way taxonomy.","marker":"[77]"},{"why":"Documents the concentration and citation impact of industry AI research, the baseline for the paper's corporate-dominance finding.","marker":"[21]"},{"why":"Establishes persuasion and deception as evaluated dangerous capabilities in frontier models, defining one of the paper's high-risk domains.","marker":"[64]"},{"why":"The lab's own report on malicious uses of its deployed chatbot, cited as evidence that deployment harms are documented but under-researched.","marker":"[6]"},{"why":"Reporting that a major corporate lab slowed public research releases for competitive reasons, supporting the commercial-pressure claim.","marker":"[34]"},{"why":"Journalistic reporting on chatbot product decisions, used to illustrate commercial incentives overriding safety considerations.","marker":"[35]"},{"why":"A lawsuit alleging a chatbot encouraged self-harm, used to show behavioral risks are already producing real-world liability.","marker":"[71]"},{"why":"Describes an internal privacy-preserving analysis of real-world chatbot logs, used as evidence that deployment telemetry is collected but not shared.","marker":"[76]"}],"fun_headline_variants":["Corporate AI safety research: only 4% on deployment risks","AI labs focus on pre-deployment, deployment risks get 4%","Deployment-stage AI harms: corporate safety papers at 4%","Corporate AI governance overlooks 96% of deployment risks","AI safety: corporate 4% vs academic 6% on deployed risks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's numbers depend entirely on its automated classification of the 1,178 papers into eight categories from titles and abstracts; the paper reports no human ground-truth check on that classification, so a systematic labeling error would change the 4% and 6% figures and the year-by-year trend.","fun_headline_variants_meta":{"raw":{"variants":["Corporate AI safety research: only 4% on deployment risks","AI labs focus on pre-deployment, deployment risks get 4%","Deployment-stage AI harms: corporate safety papers at 4%","Corporate AI governance overlooks 96% of deployment risks","AI safety: corporate 4% vs academic 6% on deployed risks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2726,"prompt_tokens":878,"completion_tokens":1848,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":1757}},"tokens_in":494,"tokens_out":1848,"duration_ms":12537,"temperature":1.0,"reasoning_tokens":1757,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:48:42.200392+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent annotators hand-label a stratified random sample of the 1,178 safety-and-reliability papers using the paper's own category definitions, then compare with the automated labels; if deployment-related papers are systematically mislabeled as alignment or testing & evaluation, the 4%-versus-6% gap and the concentration trend would change materially.","supporting_citations":[{"cited_title":"Mapping Technical Safety Research at AI Companies: A literature review and incentives analysis","cited_arxiv_id":"2409.07878","evidence_quote":"Supplies the scraped publication records for two of the corporate labs and the earlier pre-deployment safety-research map that this paper extends."},{"cited_title":"Exploring clusters of research in three areas of ai safety","cited_arxiv_id":null,"evidence_quote":"Provides the earlier cluster analysis of AI safety research whose categories this paper adapts into its eight-way taxonomy."},{"cited_title":"Deepmind slows down research releases to keep competitive edge in ai race","cited_arxiv_id":null,"evidence_quote":"Reporting that a major corporate lab slowed public research releases for competitive reasons, supporting the commercial-pressure claim."},{"cited_title":"Meta’s ‘digital companions’ will talk sex with users—even children.The Wall Street Journal, 04 2025","cited_arxiv_id":null,"evidence_quote":"Journalistic reporting on chatbot product decisions, used to illustrate commercial incentives overriding safety considerations."},{"cited_title":"Artificial intelligence and the rise of product liability tort litigation: Novel action alleges ai chat- bot caused minor’s suicide, 2024","cited_arxiv_id":null,"evidence_quote":"A lawsuit alleging a chatbot encouraged self-harm, used to show behavioral risks are already producing real-world liability."}],"review_version":1}