Pith. sign in

REVIEW 3 major objections 4 minor 96 references

Real-World Gaps in AI Governance Research

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Corporate AI labs publish barely any safety research on real-world harms—most of their safety work stays in pre-release alignment and testing.

desk verdict The paper maps a real and policy-relevant gap in AI governance research, but its headline 4%-vs-6% numbers rest on an unvalidated classification pipeline, so treat the specific magnitudes as provisional. read the letter →

arxiv 2505.00174 v2 pith:MSUKCMIP submitted 2025-04-30 cs.AI

classification cs.AI
keywords AIresearchalignmentinterpretabilitycommercializationriskscloudprovidersmodeldevelopersgovernancedeployment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Drawing on a corpus of 1,178 safety-and-reliability papers selected from 9,439 generative AI papers published between January 2020 and March 2025, the paper tries to establish that corporate AI safety research is skewed away from real-world deployment. It reports that only about 4% of corporate safety papers (6% of academic safety papers) address high-stakes deployment domains such as persuasion, misinformation, medical and financial contexts, disclosures, and core business liabilities, while corporate work concentrates on pre-deployment alignment and testing & evaluation. The claim matters because, if correct, society's empirical understanding of AI harms is narrowing even as AI systems spread: the companies best positioned to observe deployed behavior have the least incentive to publish research on it, and independent researchers lack the telemetry to fill the gap. The paper concludes that structured external access to deployment logs, traces, and model artifacts is the necessary remedy.

What carries the argument

The carrying mechanism is the paper's classification system. A corpus of 9,439 generative AI papers is filtered by keyword lists and an automated language-model classifier into 1,178 safety-and-reliability papers, each assigned to exactly one of eight categories: alignment, testing & evaluation, ethics & bias, privacy & security, interpretability & transparency, policy & governance, post-deployment risks and model traits, and multi-agent/agentic safety. Institutional credit is counted fractionally by authorship, so a paper with four authors from a tracked lab contributes 0.25 to that lab's totals. A second layer of regex keyword searches on titles and abstracts then identifies high-risk deployment contexts (medical, finance, commercial, copyright) and capability areas (misinformation, disclosures, behavioral, accuracy). The gap claims are the output of these two layers, and the year-by-year category shares in Figure 4 trace the growing concentration in alignment and testing & evaluation.

What would settle it

Have independent annotators hand-label a stratified random sample of the 1,178 safety-and-reliability papers using the paper's own category definitions, then compare with the automated labels; if deployment-related papers are systematically mislabeled as alignment or testing & evaluation, the 4%-versus-6% gap and the concentration trend would change materially.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a quantitative mismatch between research effort and risk exposure. Among the 1,178 safety-and-reliability papers, corporate labs publish most of their governance research in model alignment and testing & evaluation—work that examines models in controlled, pre-deployment settings—while post-deployment concerns such as bias, misinformation, persuasive or addictive design, medical and financial advice, disclosures, and liability-relevant failures receive a small fraction of attention. In total, 217 fractionally adjusted academic papers and 67 corporate papers touch any of these high-risk areas, about 6% and 4% of each group's safety output. The gap widens in specific domains, such as medical and misinformation research, where academic papers outnumber corporate ones by several times. The authors interpret this concentration as the result of commercial incentives and an existential-risk research culture, and they argue that without structured access to deployment telemetry the knowledge deficit will deepen.

Load-bearing premise

The paper's numbers depend entirely on its automated classification of the 1,178 papers into eight categories from titles and abstracts; the paper reports no human ground-truth check on that classification, so a systematic labeling error would change the 4% and 6% figures and the year-by-year trend.

Editorial extensions

If this is right

  • Corporate AI safety research will keep concentrating on alignment and testing & evaluation as long as product competition shapes what labs publish, so the public share of deployment-stage safety research is unlikely to rise on its own.
  • The high-stakes domains the paper flags—copyright, medical and financial advice, misinformation, and behavioral influence—are already generating lawsuits, making the research gap a live liability gap.
  • Independent researchers cannot currently study deployed AI harms systematically, because incident databases and leaked chat logs are partial; the empirical base for AI regulation will stay thin without deployment data access.
  • Structured external access to telemetry (logs, traces, model artifacts), with tiered researcher access and liability safe harbors, is the proposed path to closing the gap.
  • Widely deployed safeguards such as content moderation and telemetry-based monitoring are almost absent from the public literature, so evidence-based best practices for deployed AI systems are largely missing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the published gap is real, it is likely even wider in private artifacts: system cards, internal red-teaming reports, and product documentation are not in the corpus, so the 4% figure probably overstates the public share of deployment-focused corporate research.
  • Applying the same eight-category classification to a later time window, or to a jurisdiction with binding transparency rules, would test whether regulatory pressure shifts corporate publication toward deployment risks.
  • A testable extension of the policy proposal: if safe-harbor telemetry access became operational, one would expect a measurable rise in externally verifiable papers using real-world traces within one to two years.
  • Because the academic-to-corporate paper ratio is roughly 2.5 to 1 overall, even a modest increase in corporate deployment research could double the public literature in a high-risk area; the binding constraint is data access, not researcher supply.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper constructs a dataset of 1,178 safety-and-reliability papers drawn from 9,439 generative-AI papers (January 2020–March 2025) by five corporate AI labs and six academic institutions, classifies each paper into one of eight pre/post-deployment categories using an LLM plus regex keyword screening, and reports that corporate AI research is increasingly concentrated in pre-deployment alignment and testing-and-evaluation work. The headline empirical claim is that only 4% of Corporate AI papers (6% of Academic AI papers) address high-stakes deployment domains such as persuasion, misinformation, medical and financial contexts, disclosures, and core business liabilities. The paper concludes with policy recommendations for tiered external access to deployment telemetry and observability of in-market AI systems.

Significance. If the quantitative claims survive validation, the paper provides a timely and policy-relevant measurement of where corporate and academic AI governance research actually sits relative to documented deployment harms. The authors are transparent about their data sources, publish code and data, make their category definitions and classification prompts available, and go beyond prior work by including post-deployment categories and by supplementing OpenAlex with scraped company publications. The distinction between pre-deployment and post-deployment research is a useful organizing frame, and the policy discussion of structured telemetry access is concrete. However, the paper's central percentages rest on an unvalidated automated classification pipeline and on short regex lists with obvious coverage gaps; because these numbers are the paper's main quantitative contribution, the manuscript needs additional validation work before the headline findings can be relied on.

major comments (3)
  1. [§3.2, Table 3, Core Finding 3] The headline 4%-versus-6% deployment-gap claim is computed directly from the regex lists reported in the Table 3 note, and those lists are too narrow to support the claim as stated. For example, Medical includes hospital(s), health insurance, and clinician(s) but not healthcare, clinical, patient, diagnosis, or drug; Finance includes only finance/financial and not credit, loan, banking, insurance, or investment; Misinfo omits fake news, rumor, propaganda, and election interference; Behavioral omits manipulation, nudging, engagement optimization, and self-harm; and Commercial omits advertising, e-commerce, and recruitment. This means papers written in common product-specific or clinical language are systematically uncounted, and the reported 4% and 6% are lower bounds whose tightness is untested. The paper should report precision and recall of the regex lists against a human-labeled gold standard, and should show sensitivity of the core percentages to expanded keyword sets. Without this, the central quantitative claim cannot be checked as reported.
  2. [§2.4, Appendix 6.3, Figure 1 caption] The LLM classification pipeline is described inconsistently and is not validated. Figure 1's caption says papers were categorized using GPT 4o-mini, Section 2.4 says GPT o4-mini, and Appendix 6.3 says OpenAI's o3-mini was used; this inconsistency must be resolved because reproducibility requires the exact model version. More importantly, the paper reports no human ground-truth labels, no inter-annotator agreement, and no error analysis for the two-stage keyword-plus-LLM classifier, even though the classifier's output drives the category-level citation comparisons in Figures 1 and 2, the temporal trends in Figure 4, and the regressions described in Section 3.1. The authors should add a validation section with a human-annotated random sample, per-category precision/recall, and a sensitivity analysis showing how the headline results change under alternative classifier settings or thresholds.
  3. [§2.4, Table 1, Appendix 6.2] The corporate sample is not complete in a way that could affect the paper's denominators and gap ratios. The text notes that Meta website publications were not manually scraped, and Appendix 6.2 states that OpenAlex contains no Anthropic papers, with only Anthropic and OpenAI publications supplemented from company websites. Because Meta and Anthropic are among the five corporate labs analyzed, the omission is material: adding their missing publications could change both the total corporate paper counts and the number of corporate papers in the high-risk deployment categories. The authors should report coverage statistics for each source and institution, and should show that the 4%-versus-6% comparison and the corporate-vs-academic ratios are robust to the inclusion or exclusion of the incomplete source-institution pairs.
minor comments (4)
  1. [§3.2] In the paragraph listing corporate post-deployment examples, "read-teaming" should be "red-teaming".
  2. [§1, footnote 2] The footnote reads "do revise their models based based on red-teaming and user experience feedback"; the duplicated "based" should be removed.
  3. [Appendix 6.2] The phrase "incudes most ArXiv papers" should be corrected to "includes most arXiv papers".
  4. [Appendix 6.1, Figure 4] In the bottom panel of Figure 4, the category ordering appears to differ from the top panel without explanation; adding a shared legend order or a note would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central 4%-vs-6% gap claim is an empirical corpus measurement, not a derivation from fitted inputs or self-citations.

full rationale

The paper's central claim—that only 4% of Corporate AI papers (6% of Academic AI papers) address high-stakes deployment areas—is an empirical measurement computed from an explicitly described corpus and classification procedure, not a derivation that reduces to its own inputs. Table 3's percentages are generated by a stated regex keyword protocol over titles and abstracts, and the eight safety & reliability categories are defined independently in Appendix 6.3; nothing in the keyword lists is fitted to produce the 4%/6% totals, and no model parameter is renamed as a prediction. The main validity concerns—possible undercounts from short regex lists, the inconsistent classifier label (GPT 4o-mini vs. o4-mini vs. o3-mini), and the absence of human ground-truth labels—are measurement reliability and reproducibility threats, not logical circularity: an inaccurate count is still not a count made true by definition. The self-citations by the authors (e.g., refs. 61, 73, 74) are used to motivate the policy framing and the deployment-risk taxonomy, not to establish the quantitative findings, and the cited prior classification work by Delaney et al. is external to this paper's authors. Because the load-bearing numerical claims are data-derived rather than assumption-derived, no circular step can be exhibited, and the honest finding is no significant circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The measurement rests on no mathematical derivation, so the ledger records the data and classification choices that act as assumptions: sample coverage, LLM labeling, regex keyword lists, citation retrieval, and fractional authorship. Each is a choice the reader must buy without independent verification.

free parameters (3)
  • Keyword regex lists for high-risk domains
    The lists used in Table 3 (medical, finance, misinformation, behavioral, copyright, disclosure) were hand-authored and determine which papers count; different lists change the headline percentages, and no sensitivity analysis is reported.
  • LLM category definitions (eight categories)
    The eight mutually exclusive categories and their descriptive prompts were authored by the researchers and applied by a proprietary model; the central comparisons in Figures 1, 2, and 4 depend on these definitions.
  • Fractional authorship weighting rule
    Each author contributes equally to an institution's paper count, with no author-order weighting, and authors with multiple affiliations are assigned to one target institution; alternative weighting changes institutional counts and citation shares.
assumptions (5)
  • domain assumption OpenAlex and scraped company websites together provide a representative sample of generative AI research from the selected institutions.
    The sample is built from these sources; OpenAlex contains no Anthropic papers, and Meta papers were not manually scraped, so coverage is incomplete by the authors' own account.
  • ad hoc to paper The o3-mini LLM's assignment of papers to safety categories is accurate enough for the aggregate claims.
    Section 2.4 and Appendix 6.3 describe classification by OpenAI's o3-mini with author-written category descriptions; no human ground-truth set or agreement metric is provided, yet all counts in Figures 1, 2, and 4 rely on these labels.
  • ad hoc to paper Regex matching of titles and abstracts reliably identifies papers addressing high-risk deployment domains.
    Table 3 counts are based on keyword lists for medical, finance, misinformation, behavioral, and copyright; papers that study these topics without using the listed terms are missed, and no recall check is reported.
  • domain assumption Google Scholar citation counts retrieved via SerpApi approximate research impact.
    Citation-based claims in Section 3 and Figures 1 and 2 rest on these counts; 43 citation counts are missing and the retrieval method is heuristic, as described in Appendix 6.2.
  • domain assumption Fractional authorship provides a valid measure of institutional research output.
    The paper assigns equal credit to every author and one institution per author; this choice affects all totals in Tables 2 and 4 and is not validated against other weighting schemes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Real-World Gaps in AI Governance Research." pith.science (2026). https://pith.science/paper/MSUKCMIP

@misc{pith2026250500174,
  author       = {Pith},
  title        = {Pith review of: Real-World Gaps in AI Governance Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MSUKCMIP}},
  note         = {Machine review of arXiv:2505.00174}
}
read the original abstract

Drawing on 1,178 safety and reliability papers from 9,439 generative AI papers (January 2020 - March 2025), we compare research outputs of leading AI companies (Anthropic, Google DeepMind, Meta, Microsoft, and OpenAI) and AI universities (CMU, MIT, NYU, Stanford, UC Berkeley, and University of Washington). We find that corporate AI research increasingly concentrates on pre-deployment areas -- model alignment and testing & evaluation -- while attention to deployment-stage issues such as model bias has waned. Significant research gaps exist in high-risk deployment domains, including healthcare, finance, misinformation, persuasive and addictive features, hallucinations, and copyright. Without improved observability into deployed AI, growing corporate concentration could deepen knowledge deficits. We recommend expanding external researcher access to deployment data and systematic observability of in-market AI behaviors.

Figures

Figures reproduced from arXiv: 2505.00174 by the authors.

Figure 1
Figure 1. Total Citations for Safety & Reliability Research 1557 543 274 4390 1760 222 409 3128 3687 379 288 569 420 389 695 898 539 490 1458 1522 347 271 512 996 193 874 404 325 280 444 194 318 275 456 Corporate AI Academic AI Anthropic OpenAI Google DeepMind Microsoft Meta U. of Washington CMU Stanford MIT UC Berkeley NYU 0 2000 4000 6000 Total Citations Alignment Ethics & Bias Interpretability & Transparency Multi−Agent & … view at source ↗
Figure 2
Figure 2. Number of AI Safety & Reliability Papers 52 27 13 21 12 13 12 22 11 12 36 13 13 30 38 16 12 22 21 13 21 29 12 19 10 11 13 Corporate AI Academic AI Google DeepMind Anthropic Microsoft OpenAI Meta Stanford CMU U. of Washington NYU MIT UC Berkeley 0 50 100 Number of Papers Alignment Ethics & Bias Interpretability & Transparency Multi−Agent & Agentic Policy & Governance Post−Deployment & Model Traits Privacy & Security … view at source ↗
Figure 3
Figure 3. All Generative AI Publications by Institution (2020-2024) Stanford UC Berkeley University of Washington Microsoft MIT New York University OpenAI Anthropic CMU Google DeepMind Meta 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023 2024 2021 2022 2023 2024 2020 2021 2022 2023 2024 2020 2021 2022 2023… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: AI Governance Areas by Total Paper Numbers (by Year) - Top Graph; and by Total Citations (Fractionally Adjusted) - Bottom Graph. Multi−Agent & Agentic Policy & Governance Post−Deployment & Model Traits Ethics & Bias Interpretability & Transparency Privacy & Security Al…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 49 canonical work pages

  1. [1]

    Securing large language models: Threats, vulnerabilities and responsible practices

    Sara Abdali, Richard Anarfi, CJ Barberan, and Jia He. Securing large language models: Threats, vulnerabilities and responsible practices. arXiv preprint arXiv:2403.12503, 2024

  2. [2]

    The urgency of interpretability, 04 2025

    Dario Amodei. The urgency of interpretability, 04 2025. URL https://www. darioamodei.com/post/the-urgency-of-interpretability. Accessed: 2025-04-28

  3. [3]

    Faulty reward functions in the wild, 12 2016

    Dario Amodei and Jack Clark. Faulty reward functions in the wild, 12 2016. URL https://openai.com/blog/faulty-reward-functions/. OpenAI Blog

  4. [4]

    Frontier ai regulation: Managing emerging risks to public safety.arXiv preprint arXiv:2307.03718, 2023

    Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, et al. Frontier ai regulation: Managing emerging risks to public safety.arXiv preprint arXiv:2307.03718, 2023

  5. [5]

    Anthropic economic index: Insights from claude 3.7 sonnet, 03 2025

    Anthropic. Anthropic economic index: Insights from claude 3.7 sonnet, 03 2025. URL https://www.anthropic.com/news/ anthropic-economic-index-insights-from-claude-sonnet-3-7 . Accessed: 2025-04-28

  6. [6]

    Detecting and countering malicious uses of claude: March

    Anthropic. Detecting and countering malicious uses of claude: March

  7. [7]

    Exploring model welfare, April 2025

    Anthropic. Exploring model welfare, April 2025. URLhttps://www.anthropic.com/ research/exploring-model-welfare. Accessed: 2025-04-25

  8. [8]

    Ai liability along the value chain, 2025

    Beatriz Botero Arcila. Ai liability along the value chain, 2025. URLhttps://blog. mozilla.org/netpolicy/files/2025/03/AI-Liability-Along-the-Value-Chain_ Beatriz-Arcila.pdf. Mozilla

Show all 96 references
  1. [9]

    What is LLMOps?: large language models in production

    Abi Aryan. What is LLMOps?: large language models in production. O’Reilly Media, Inc., 2024

  2. [10]

    Constitutional ai: Harmlessness from ai feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022. 19

  3. [11]

    Factbench: A dy- namic benchmark for in-the-wild language model factuality evaluation.arXiv preprint arXiv:2410.22257, 2025

    Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, and Lu Wang. Factbench: A dy- namic benchmark for in-the-wild language model factuality evaluation.arXiv preprint arXiv:2410.22257, 2025

  4. [12]

    Superintelligence: Paths, dangers, strategies, 2014

    Nick Bostrom. Superintelligence: Paths, dangers, strategies, 2014

  5. [13]

    Lessons learned on lan- guage model safety and misuse, 3 2022

    Miles Brundage, Katie Mayer, Tyna Eloundou, Sandhini Agarwal, Steven Adler, Gretchen Krueger, Jan Leike, and Pamela Mishkin. Lessons learned on lan- guage model safety and misuse, 3 2022. URL https://openai.com/index/ language-model-safety-and-misuse/. Accessed: 2025-01-23

  6. [14]

    Lessons from red teaming 100 generative ai products.arXiv preprint arXiv:2501.07238, 2025

    Blake Bullwinkel, Amanda Minnich, Shiven Chawla, Gary Lopez, Martin Pouliot, Whit- ney Maxwell, Joris de Gruyter, Katherine Pratt, Saphir Qi, Nina Chikanov, et al. Lessons from red teaming 100 generative ai products.arXiv preprint arXiv:2501.07238, 2025

  7. [15]

    Cfr part 43; rin 3038-ad08: Real-time public reporting of swap transaction data

    US CFTC. Cfr part 43; rin 3038-ad08: Real-time public reporting of swap transaction data. Federal Register, 77(5):1182–266, 2012

  8. [16]

    Draft report of the joint california policy working group on ai frontier models

    Jennifer Tour Chayes, Mariano-Florentino Cuèllar, and Fei-Fei Li. Draft report of the joint california policy working group on ai frontier models. Technical re- port, Joint California Policy Working Group on AI Frontier Models, 3 2025. URL https://www.cafrontieraigov.org/wp-co...

  9. [17]

    Humans or llms as the judge? a study on judgement biases.arXiv preprint arXiv:2402.10669, 2024

    GuimingHardyChen, ShunianChen, ZicheLiu, FengJiang, andBenyouWang. Humans or llms as the judge? a study on judgement biases.arXiv preprint arXiv:2402.10669, 2024

  10. [18]

    Realm dataset dashboard, 03 2025

    Jingwen Cheng, Kshitish Ghate, Wenyue Hua, William Yang Wang, Hong Shen, and Fei Fang. Realm dataset dashboard, 03 2025. URLhttps://realm-e7682.web.app/. Accessed: 2025-04-28

  11. [19]

    Deep reinforcement learning from human preferences

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017

  12. [20]

    Regulatory markets for ai safety.arXiv preprint arXiv:2001.00078, 2019

    Jack Clark and Gillian K Hadfield. Regulatory markets for ai safety.arXiv preprint arXiv:2001.00078, 2019. 20

  13. [21]

    Who is leading in ai? an analysis of industry ai research.arXiv preprint arXiv:2312.00043, 2023

    Ben Cottier, Tamay Besiroglu, and David Owen. Who is leading in ai? an analysis of industry ai research.arXiv preprint arXiv:2312.00043, 2023

  14. [22]

    Mapping technical safety re- search at ai companies: A literature review and incentives analysis

    Oscar Delaney, Oliver Guest, and Zoe Williams. Mapping technical safety re- search at ai companies: A literature review and incentives analysis. arXiv preprint arXiv:2409.07878, 2024

  15. [23]

    Assessing bias in metric models for llm open-ended generation bias benchmarks

    Nathaniel Demchak, Xin Guan, Zekun Wu, Ziyi Xu, Adriano Koshiyama, and Emre Kazim. Assessing bias in metric models for llm open-ended generation bias benchmarks. arXiv preprint arXiv:2410.11059, 2024

  16. [24]

    Sycophancy to subterfuge: Investigating reward-tampering in large language models

    Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, et al. Sycophancy to subterfuge: Investigating reward-tampering in large language models. arXiv preprint arXiv:2406.10162, 2024

  17. [25]

    How ai can help learn lessons from incident reporting systems

    Robin Dillon, Peter Madsen, Brian Holland, and Danniel Cao. How ai can help learn lessons from incident reporting systems. In 2024 IEEE Aerospace Conference, pages 1–15. IEEE, 2024

  18. [26]

    Openai’s new ai image generator is potent and bound to pro- voke

    Benj Edwards. Openai’s new ai image generator is potent and bound to pro- voke. Ars Technica, 03 2025. URL https://arstechnica.com/ai/2025/03/ openais-new-ai-image-generator-is-potent-and-bound-to-provoke/

  19. [27]

    Analyzingtheimpactofcompaniesonairesearch based on publications

    MichaelFarberandLazarosTampakis. Analyzingtheimpactofcompaniesonairesearch based on publications. arXiv preprint, 10 2023. URLhttps://arxiv.org/pdf/2310. 20444. Accessed: 2025-01-10

  20. [28]

    Collaborative gatekeepers.Wash

    Stavros Gadinis and Colby Mangels. Collaborative gatekeepers.Wash. & Lee L. Rev., 73:797, 2016

  21. [29]

    Chal- lenges in evaluating ai systems, 2023

    Deep Ganguli, Nicholas Schiefer, Favarom Marina, and Jack Clark. Chal- lenges in evaluating ai systems, 2023. URL https://www.anthropic.com/news/ evaluating-ai-systems. Accessed: 2025-01-23

  22. [30]

    Technological triggers to tort revolutions: steam locomotives, au- tonomous vehicles, and accident compensation.Journal of tort law, 11(1):71–143, 2018

    Donald G Gifford. Technological triggers to tort revolutions: steam locomotives, au- tonomous vehicles, and accident compensation.Journal of tort law, 11(1):71–143, 2018

  23. [31]

    Im- 21 proving alignment of dialogue agents via targeted human feedback

    Andreas Glaese, Natasha McAleese, Julian Aslanides, Andy Huang, Laura Rimell, Jonathan Uesato, Jack Rae, Long Ouyang, Joe Mellor, Isaac Caswell, et al. Im- 21 proving alignment of dialogue agents via targeted human feedback. arXiv preprint arXiv:2209.14375, 2022. URL https://a...

  24. [32]

    Deliberative alignment: Reasoning enables safer language models.arXiv preprint arXiv:2412.16339, 2024

    Melody Y Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Heylar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, et al. Deliberative alignment: Reasoning enables safer language models.arXiv preprint arXiv:2412.16339, 2024

  25. [33]

    Regulatory markets: The future of ai governance

    Gillian K Hadfield and Jack Clark. Regulatory markets: The future of ai governance. arXiv preprint arXiv:2304.04914, 2023

  26. [34]

    Deepmind slows down research releases to keep competitive edge in ai race

    Melissa Heikkilä and Stephen Morris. Deepmind slows down research releases to keep competitive edge in ai race. Financial Times, 04 2025. URL https://www.ft.com/ content/2ee1ffde-008e-4ea4-861b-24f15b25cf54. Accessed: 2025-04-10

  27. [35]

    Meta’s ‘digital companions’ will talk sex with users—even children.The Wall Street Journal, 04 2025

    Jeff Horwitz and Georgia Wells. Meta’s ‘digital companions’ will talk sex with users—even children.The Wall Street Journal, 04 2025. URLhttps://www.wsj.com/ tech/ai/meta-ai-chatbots-sex-a25311bf

  28. [36]

    nipitinthebud

    JaneHsieh, JoselynKim, LauraDabbish, andHaiyiZhu. "nipitinthebud": Moderation strategies in open source software projects and the role of bots.Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2):1–29, 2023

  29. [37]

    Values in the wild: Discovering and analyzing values in real-world language model interactions

    Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. Values in the wild: Discovering and analyzing values in real-world language model interactions. arXiv preprint arXiv:2504.15236, 2025

  30. [38]

    Can we govern ai without breaking it?, 02 2025

    Chris Hughes. Can we govern ai without breaking it?, 02 2025. URL https: //chrishughes.substack.com/p/can-we-govern-ai-without-breaking. Accessed: 2025-04-28

  31. [39]

    Beyond static ai evaluations: advancing human interaction evaluations for llm harms and risks.arXiv preprint arXiv:2405.10632, 2024

    Lujain Ibrahim, Saffron Huang, Lama Ahmad, and Markus Anderljung. Beyond static ai evaluations: advancing human interaction evaluations for llm harms and risks.arXiv preprint arXiv:2405.10632, 2024

  32. [41]

    Openai o1 system card

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  33. [42]

    Randomness, not repre- sentation: The unreliability of evaluating cultural alignment in llms

    Ariba Khan, Stephen Casper, and Dylan Hadfield-Menell. Randomness, not repre- sentation: The unreliability of evaluating cultural alignment in llms. arXiv preprint arXiv:2503.08688, 2025

  34. [43]

    Expanding academia’s role in public sector ai

    Kevin Klyman, Caroline Meinhardt, Daniel Zhang, Elena Cryst, Russell Wald, and Aaron Bao. Expanding academia’s role in public sector ai. Issue brief, Stanford Institute for Human-Centered Artificial Intelligence, Stanford Univer- sity, Stanford, CA, December 2024. URL https://...

  35. [44]

    Every ai copyright lawsuit in the us, visualized.WIRED, 03 2025

    Kate Knibbs. Every ai copyright lawsuit in the us, visualized.WIRED, 03 2025. URL https://www.wired.com/story/ai-copyright-case-tracker/

  36. [45]

    Lessons from the fda for ai, 08 2024

    Anna Lenhart and Sarah Myers West. Lessons from the fda for ai, 08 2024. URLhttps: //ainowinstitute.org/publications/research/lessons-from-the-fda-for-ai

  37. [46]

    A safe harbor for ai evaluation and red teaming.arXiv preprint arXiv:2403.04893, 2024

    Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bom- masani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, et al. A safe harbor for ai evaluation and red teaming.arXiv preprint arXiv:2403.04893, 2024

  38. [47]

    Zero-shot image moderation in google ads with llm-assisted textual descriptions and cross-modal co-embeddings

    Enming Luo, Wei Qiao, Katie Warren, Jingxiang Li, Eric Xiao, Krishna Viswanathan, Yuan Wang, Yintao Liu, Jimin Li, and Ariel Fuxman. Zero-shot image moderation in google ads with llm-assisted textual descriptions and cross-modal co-embeddings. In Proceedings of the Eighteenth ...

  39. [48]

    f**k it, we’re doing imperson- ation now

    Alexios Mantzarlis. Openai says "f**k it, we’re doing imperson- ation now". Faked Up , 04 2025. URL https://fakedup.org/ openai-says-fk-it-were-doing-impersonation-now/

  40. [49]

    Generative ai misuse: A taxonomy of tactics and insights from real- world data.arXiv preprint arXiv:2406.13843, 2024

    Nahema Marchal, Rachel Xu, Rasmi Elasmar, Iason Gabriel, Beth Goldberg, and William Isaac. Generative ai misuse: A taxonomy of tactics and insights from real- world data.arXiv preprint arXiv:2406.13843, 2024. 23

  41. [50]

    Consolidated audit trail: Strategic planning and best practices.Journal of Securities Operations & Custody, 10(1):77–83, 2018

    Michael Martinen, George Black, Ripple Bhullar, and Victor Marranca. Consolidated audit trail: Strategic planning and best practices.Journal of Securities Operations & Custody, 10(1):77–83, 2018

  42. [51]

    Ai ’hallucinations’ in court papers spell trouble for lawyers

    Sara Merken. Ai ’hallucinations’ in court papers spell trouble for lawyers. Reuters, 02 2025. URL https://www.reuters.com/technology/artificial-intelligence/ ai-hallucinations-court-papers-spell-trouble-lawyers-2025-02-18/ . Ac- cessed: 2025-04-28

  43. [52]

    Microsoft digital defense report 2024

    Microsoft Corporation. Microsoft digital defense report 2024. Tech- nical report, Microsoft Corporation, 10 2024. URL https://www. microsoft.com/en-us/security/security-insider/intelligence-reports/ microsoft-digital-defense-report-2024. Accessed: 2025-04-28

  44. [53]

    Mit ai incident tracker, 2024

    Simon Mylius. Mit ai incident tracker, 2024. URL https://airisk.mit.edu/ ai-incident-tracker. Accessed: February 6, 2025

  45. [54]

    Scalable ai incident classification, 2024

    Simon Mylius and Jamie Bernadi. Scalable ai incident classification, 2024. URLhttps: //simonmylius.com/blog/incident-classification. Blog post

  46. [55]

    The alignment problem from a deep learning perspective.arXiv preprint arXiv:2209.00626, 2022

    Richard Ngo, Lawrence Chan, and Sören Mindermann. The alignment problem from a deep learning perspective.arXiv preprint arXiv:2209.00626, 2022

  47. [56]

    Supremacy: AI, ChatGPT, and the Race that Will Change the World

    Parmy Olson. Supremacy: AI, ChatGPT, and the Race that Will Change the World. St. Martin’s Press, 2024

  48. [57]

    How to implement llm guardrails, 2023

    OpenAI. How to implement llm guardrails, 2023. URL https://cookbook.openai. com/examples/how_to_use_guardrails. Accessed: 2025-04-29

  49. [58]

    Influence and cyber operations: An up- date, 10 2024

    OpenAI. Influence and cyber operations: An up- date, 10 2024. URL https://openai.com/index/ disrupting-deceptive-uses-of-AI-by-covert-influence-operations

  50. [59]

    Safety best practices

    OpenAI. Safety best practices. https://platform.openai.com/docs/guides/ safety-best-practices, 2024. Accessed: 2025-04-29

  51. [60]

    Disrupting malicious uses of our models: Febru- ary 2025 update

    OpenAI. Disrupting malicious uses of our models: Febru- ary 2025 update. Technical report, OpenAI, 02 2025. URL https://cdn.openai.com/threat-intelligence-reports/ disrupting-malicious-uses-of-our-models-february-2025-update.pdf . 24

  52. [61]

    What auto safety teaches us about ai safety

    Tim O’Reilly. What auto safety teaches us about ai safety. Substack post, 11 2024. URL https://asimovaddendum.substack.com/p/what-auto-safety-teaches-us-about

  53. [62]

    Train- ing language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Train- ing language models to follow instructions with human feedback.Advances in neural information processing systems, 35:277...

  54. [63]

    Discovering language model behaviors with model-written evaluations.arXiv preprint arXiv:2212.09251, 2022

    Ethan Perez, Sam Ringer, Kamil˙ e Lukoši¯ ut˙ e, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al. Discovering language model behaviors with model-written evaluations.arXiv preprint arXiv:2212.09251, 2022

  55. [64]

    Evaluating frontier models for dangerous capabilities.arXiv preprint arXiv:2403.13793, 2024

    Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Vic- toria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, et al. Evaluating frontier models for dangerous capabilities.arXiv preprint arXiv:2403.13793, 2024

  56. [65]

    Scaling up llm re- views for google ads content moderation

    Wei Qiao, Tushar Dogra, Otilia Stretcu, Yu-Han Lyu, Tiantian Fang, Dongjin Kwon, Chun-Ta Lu, Enming Luo, Yuan Wang, Chih-Chun Chia, et al. Scaling up llm re- views for google ads content moderation. InProceedings of the 17th ACM International Conference on Web Search and Data ...

  57. [66]

    Bet- ter language models and their implications, 2019

    AlecRadford, JeffreyWu, DarioAmodei, DaniellaAmodei, JackClark, MilesBrundage, Ilya Sutskever, Amanda Askell, David Lansky, Danny Hernandez, and David Luan. Bet- ter language models and their implications, 2019. URLhttps://openai.com/index/ better-language-models/. Accessed: 2...

  58. [67]

    Kevin Roose. If a.i. systems become conscious, should they have rights? The New York Times. URL https://www.nytimes.com/2025/04/24/technology/ ai-welfare-anthropic-claude.html. Accessed: 2025-04-28

  59. [68]

    Sharegpt vicuna unfiltered

    ShareGPT. Sharegpt vicuna unfiltered. https://huggingface.co/datasets/ anon8231489123/ShareGPT_Vicuna_unfiltered, 2023. Apache 2.0 License

  60. [69]

    Towards understanding sycophancy in language models

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R John- ston, et al. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548, 2023. 25

  61. [70]

    Release strategies and the social impacts of language models

    Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, et al. Release strategies and the social impacts of language models. arXiv preprint, arXiv:1908.09203, 2019. URL ht...

  62. [71]

    Artificial intelligence and the rise of product liability tort litigation: Novel action alleges ai chat- bot caused minor’s suicide, 2024

    Katy Spicer, Julia Jacbson, Daniel Stephen, Naija Perry, and Aden Hochrun. Artificial intelligence and the rise of product liability tort litigation: Novel action alleges ai chat- bot caused minor’s suicide, 2024. URL https://www.privacyworld.blog/2024/11/ artificial-intellige...

  63. [72]

    Learning to summarize with human feedback

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in neural information processing systems, 33:3008–3021, 2020

  64. [73]

    Ai is entirely new, ai is exactly the same: Thoughts on the new white house ai memorandum, 10 2024

    Ilan Strauss and Tim O’Reilly. Ai is entirely new, ai is exactly the same: Thoughts on the new white house ai memorandum, 10 2024. Asimov’s Addendum

  65. [74]

    Risk without uncertainty? openai would like us to think so...Asimov’s Addendum, 2024

    Ilan Strauss and Tim O’Reilly. Risk without uncertainty? openai would like us to think so...Asimov’s Addendum, 2024. URLhttps://asimovaddendum.substack.com/ p/can-we-have-ai-model-risk-evaluation . AI model evaluations, such as those conducted by OpenAI in its GPT system cards...

  66. [75]

    Faiz Surani and Daniel E. Ho. Ai on trial: Legal models hal- lucinate in 1 out of 6 (or more) benchmarking queries. Stan- ford HAI , 05 2024. URL https://hai.stanford.edu/news/ ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries . Accessed: 2025-04-28

  67. [76]

    Clio: Privacy-preserving insights into real-world ai use.arXiv preprint arXiv:2412.13678, 2024

    Alex Tamkin and et al. Clio: Privacy-preserving insights into real-world ai use.arXiv preprint arXiv:2412.13678, 2024

  68. [77]

    Exploring clusters of research in three areas of ai safety

    Helen Toner and Ashwin Acharya. Exploring clusters of research in three areas of ai safety. Center for Security and Emerging Technol- ogy, 2022. URL https://cset.georgetown.edu/wp-content/uploads/ Exploring-Clusters-of-Research-in-Three-Areas-of-AI-Safety.pdf . 26

  69. [78]

    Who do we become when we talk to machines?, 2024

    Sherry Turkle. Who do we become when we talk to machines?, 2024. URLhttps: //www.youtube.com/watch?v=yYlfGc0YR3Y

  70. [79]

    Sociotechnical safety evaluation of generative ai systems.arXiv preprint arXiv:2310.11986, 2023

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, et al. Sociotechnical safety evaluation of generative ai systems.arXiv preprint arXiv:2310.11986, 2023

  71. [80]

    Holistic safety and responsibility evaluations of advanced ai models

    Laura Weidinger, Joslyn Barnhart, Jenny Brennan, Christina Butterfield, Susie Young, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chang, Mikel Ro- driguez, et al. Holistic safety and responsibility evaluations of advanced ai models. arXiv preprint arXiv:2404.14068, 2024

  72. [81]

    Toward an evaluation science for generative ai systems

    Laura Weidinger, Deb Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Sayash Kapoor, Deep Ganguli, Sanmi Koyejo, et al. Toward an evaluation science for generative ai systems. arXiv preprint arXiv:2503.05336, 2025

  73. [82]

    Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359, 2021

    Laura Weidinger et al. Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359, 2021

  74. [83]

    Responsible ai: A guide to guardrails and scorers.https://wandb

    Weights & Biases. Responsible ai: A guide to guardrails and scorers.https://wandb. ai/site/articles/ai-guardrails/, 2025. Accessed: 2025-04-29

  75. [84]

    Meta’s ’digital companions’ will talk sex with users—even children

    Georgia Wells, Jeff Horwitz, and Deepa Seetharaman. Meta’s ’digital companions’ will talk sex with users—even children. The Wall Street Journal, 04 2025. URL https: //www.wsj.com/tech/ai/meta-ai-chatbots-sex-a25311bf

  76. [85]

    Google folds more ai teams into deepmind to ‘accelerate the research- to-developer pipeline’, 01 2025

    Kyle Wiggers. Google folds more ai teams into deepmind to ‘accelerate the research- to-developer pipeline’, 01 2025. URL https://techcrunch.com/2025/01/09/ google-folds-more-ai-teams-into-deepmind-to-accelerate-the-research-to-developer-pipeline/ . Accessed: 2025-01-16

  77. [86]

    OWASP Top Ten.https://owasp.org/www-project-top-ten/, 2024

    Steve Willison. OWASP Top Ten.https://owasp.org/www-project-top-ten/, 2024. Accessed: 2025-04-21

  78. [87]

    The Developer’s Playbook for Large Language Model Security

    Steve Wilson. The Developer’s Playbook for Large Language Model Security. O’Reilly Media, Incorporated, 2024

  79. [88]

    Measuring models’ special interests

    Zack Witten. Measuring models’ special interests. https://zswitten.github.io/ 2025/04/14/model-special-interests.html, 2025. Accessed: 2025-04-18. 27

  80. [89]

    Google’s ai unit reorganizes product work, an- nounces changes to gemini app team

    Erin Woo. Google’s ai unit reorganizes product work, an- nounces changes to gemini app team. The Information , 03 2025. URL https://www.theinformation.com/briefings/ googles-ai-unit-reorganizes-product-work-announces-changes-to-gemini-app-team? rc=7em78a. Accessed: 2025-04-18

  81. [90]

    The AI-box experiment, 2002

    Eliezer Yudkowsky. The AI-box experiment, 2002. URL http://yudkowsky.net/ singularity/aibox. Accessed: 2025-02-03

  82. [91]

    The sequences (lesswrong)

    Eliezer Yudkowsky. The sequences (lesswrong). https://www.lesswrong.com/tag/ sequences, 2020. Accessed: 2025-02-03

  83. [92]

    Openai’s new reasoning ai models hallucinate more, 04 2025

    Maxwell Zeff. Openai’s new reasoning ai models hallucinate more, 04 2025. URL https://techcrunch.com/2025/04/18/ openais-new-reasoning-ai-models-hallucinate-more/ . TechCrunch, accessed April 23, 2025

  84. [93]

    thinking slow

    Yiming Zhang, Sravani Nanduri, Liwei Jiang, Tongshuang Wu, and Maarten Sap. Bi- asx:" thinking slow" in toxic content moderation with explanations of implied social biases. arXiv preprint arXiv:2305.13589, 2023

  85. [94]

    Wildchat: 1m chatGPT interaction logs in the wild

    Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wildchat: 1m chatGPT interaction logs in the wild. InThe Twelfth International Con- ference on Learning Representations, 2024. URLhttps://openreview.net/forum?id= Bl8u7ZRlbM

  86. [95]

    Halueval-wild: Evaluating hallucinations of language models in the wild.arXiv preprint arXiv:2403.04307, 2024

    Zhiying Zhu, Yiming Yang, and Zhiqing Sun. Halueval-wild: Evaluating hallucinations of language models in the wild.arXiv preprint arXiv:2403.04307, 2024

  87. [96]

    language model*

    Georg Zoeller. Comment on ethan mollick’s post about model preferences and claude’s behavior. https://www.linkedin.com, 04 2025. LinkedIn post, April 15, 2025. Ac- cessed via Ethan Mollick’s public post. 28 6 Appendix 6.1 Additional Analysis T able 4. Dataset Adjusted for Auth...

  88. [2025]

    URL https://www.anthropic.com/news/ detecting-and-countering-malicious-uses-of-claude-march-2025

    Anthropic News, 04 2025. URL https://www.anthropic.com/news/ detecting-and-countering-malicious-uses-of-claude-march-2025 . Accessed: 2025-04-28

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.