Pith. sign in

REVIEW 3 major objections 4 minor 28 cited by

International AI Safety Report

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read An international scientific synthesis concludes that harms from general-purpose AI are already well established and that risk-management techniques remain nascent.

desk verdict A genuinely useful, policy-grade synthesis that should be read for its evidence map and honest caveats, though its 'well-established harms' headline slightly overstates what the body actually supports. read the letter →

arxiv 2501.17805 v1 pith:VVBXBU7O submitted 2025-01-29 cs.CY cs.AIcs.LG

Yoshua Bengio , Sören Mindermann , Daniel Privitera , Tamay Besiroglu , Rishi Bommasani , Stephen Casper , Yejin Choi , Philip Fox
show 88 more authors
This is my paper · ORCID
classification cs.CYcs.AIcs.LG
keywords general-purposeAIsafetyfrontierriskassessmentevidencesynthesisdeepfakespolicyopen-weightmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is the first internationally mandated scientific synthesis of what is known about the safety of general-purpose AI—AI that can perform a wide variety of tasks. Its central finding is that several harms from such systems are already well established, including scams built on fake content, non-consensual intimate imagery, child sexual abuse material, biased outputs, reliability failures, and privacy violations. It also reports that as capabilities advance, evidence is gradually emerging for larger risks such as AI-enabled cyber attacks, biological attacks, labour-market disruption, and loss of control, though experts disagree about how soon these will materialize. On the management side, the paper finds that risk identification and mitigation techniques are real but nascent: current evaluations are spot checks that can miss hazards, and no combination of methods fully resolves even established harms. The report's purpose is to give policymakers a shared evidence base for decisions, while explicitly declining to recommend specific policies.

What carries the argument

The central object is general-purpose AI, defined as an AI model or system that can perform, or be adapted to perform, a wide variety of tasks. The argument is carried by a three-part organizing structure: a capabilities assessment, a risk taxonomy that separates malicious use, malfunctions, and systemic risks, and a review of the AI development lifecycle (data collection, pre-training, fine-tuning, system integration, deployment, and monitoring). This scaffolding lets the report connect each risk to a stage where interventions could act, and it motivates the 'evidence dilemma': capability advances can be rapid and hard to predict, while reliable evidence about harms lags behind, so risk-management decisions must be made under uncertainty.

What would settle it

A single decisive check would be to pre-register a replication of the report's load-bearing capability and harm measurements on held-out data: re-run the o1-era GPQA, SWE-bench, and vulnerability-discovery evaluations on problems published after the report's cutoff, and audit the deepfake-exposure survey for non-response bias. If these replications show much smaller capability gains or much lower harm prevalence than the cited figures, the report's central risk assessment would need major revision.

Watch

Extended reading notes

Core claim

On the report's own terms, the central discovery is that the evidence base on advanced AI has matured enough to support a two-part claim: (1) general-purpose AI already produces measurable, well-documented harms to individuals and society, and (2) the same capability trends that drive benefits are generating credible, though contested, evidence for future risks on a larger scale. The report assembles this evidence across three risk categories—malicious use, malfunctions, and systemic risks—and evaluates the technical toolbox for managing those risks, concluding that methods such as stress-testing, watermarking, bias mitigation, interpretability, and monitoring are available but severely limited. It frames the policy problem as an 'evidence dilemma': risks can emerge in leaps, so waiting for conclusive evidence may leave society unprepared, while acting early on limited evidence may prove unnecessary. The report does not recommend policies; it aims to provide a scientific foundation for choices that will determine whether the technology's wide range of possible futures ends up positive or negative.

Load-bearing premise

The synthesis assumes that the evidence base assembled by the international expert panel—including non-peer-reviewed studies and expert judgement—is representative and unbiased enough to support the report's risk conclusions; if a different selection of evidence would shift the balance, the central claims weaken.

Editorial extensions

If this is right

  • Governments and companies should treat AI safety as a current problem, not a hypothetical one: several harm categories already have documented incidents, and no existing mitigation fully removes them.
  • Because current evaluations are spot checks that can miss hazards, model-release decisions based on such tests carry a risk of false assurance; the report's implied standard is to combine multiple evaluation approaches and monitor post-deployment.
  • Open-weight model releases should be assessed by marginal risk—whether release increases or decreases a given risk relative to existing alternatives—rather than treated as categorically safe or dangerous.
  • If capabilities continue scaling at recent rates, decision-makers should expect the evidence dilemma to intensify, making pre-commitment to trigger-based mitigations more valuable.
  • The wide range of possible futures means the trajectory is not predetermined; the report's corollary is that investment in risk research and international coordination can shift the balance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The report's own framing implies that a policy of waiting for conclusive evidence is not neutral: for fast-moving risks it systematically sacrifices preparedness, and trigger-based frameworks that bind mitigations to observed capability thresholds should outperform purely reactive approaches in simulations of capability jumps.
  • Because the report's evidence selection includes non-peer-reviewed sources and expert judgment, its risk balance inherits the blind spots of the literature it synthesizes; a differently constituted expert panel could plausibly shift the assessment, and auditing the source selection against a pre-registered search protocol would test this.
  • The report's finding that current evaluations rarely replicate implies a concrete research agenda: build dynamic, held-out benchmark suites that are refreshed over time, so capability claims can be verified rather than taken from developer reports.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The International AI Safety Report is a government-mandated synthesis, led by Professor Yoshua Bengio with input from 96 experts and an Expert Advisory Panel nominated by 30 countries, the UN, the EU, and the OECD. It reviews evidence on the capabilities of general-purpose AI, associated risks (malicious use, malfunctions, and systemic risks), and technical approaches to risk management. Its central claims are that capabilities have advanced rapidly; that several harms from general-purpose AI are already 'well established' (scams, non-consensual intimate imagery, CSAM, biased outputs, reliability failures, and privacy violations); that evidence of additional risks (labour market, cyber, biological, loss of control) is gradually emerging; and that risk-management techniques are nascent and limited. The report is deliberately non-prescriptive and repeatedly emphasizes expert disagreement and evidence gaps.

Significance. If its central claims hold, the report is a significant international policy reference: it assembles a broad expert consensus across governments and disciplines, transparently flags areas of disagreement and missing evidence, and includes a post-writing Chair's note updating the capability picture in light of o3 and DeepSeek R1. Its strengths are the breadth of international expert participation, the consistent hedging around future trajectories, and the explicit identification of evidence gaps. However, the report's evidence base is assembled through qualitative expert selection rather than a documented systematic review, many cited sources are non-peer-reviewed, and the headline 'well-established harms' claim is not calibrated to the body's own repeated caveats about lacking prevalence statistics. These issues are load-bearing because the report's policy relevance depends on its credibility as a scientific synthesis.

major comments (3)
  1. [Key findings (p.13) and Executive Summary (p.17)] The headline claim that 'several harms from general-purpose AI are already well established' is not calibrated to the body's own evidence-strength statements. Section 2.3.5 states that 'researchers have not found evidence of widespread privacy violations associated with general-purpose AI'; Section 2.1.1 states that 'reliable statistics on the frequency and impact of these incidents are lacking'; and Section 2.1.2 reports that 'evidence on how prevalent and how effective such efforts are remains limited.' The report never defines what 'well established' means operationally. If it means 'documented cases exist,' it is compatible with the body but likely to be misread by policymakers; if it means 'robust representative evidence,' it is internally inconsistent. Please either add an operational definition and calibrate the Key Findings to the body's confidence levels, or rephrase the bullet to say 'harms for which documented cases exist, with unknown prevalence.'
  2. [Introduction (pp.26-28)] The report's evidence-selection method is not auditable. The Introduction lists qualitative quality criteria (original contribution, comprehensive engagement with the literature, good-faith discussion of objections, described methods, stated limitations, and influence in the scientific community) but does not document a search protocol, inclusion/exclusion rules, or inter-rater reliability for screening sources. Many cited sources are non-peer-reviewed industry reports, model cards, and preprints. Because the 'well-established harms' claim depends on the representativeness of this corpus, a different source selection or panel composition could shift the balance of risk conclusions. The report should add a methods appendix describing how sources were identified, screened, and adjudicated, or explicitly scope the claim as 'based on the sources available to and selected by the expert panel.'
  3. [Section 2.1.1, CSAM paragraph (p.64)] The claim that AI-generated CSAM is a 'well-established' harm rests on thin evidence as cited: a 2019 study of deepfake videos (ref. 282), an academic investigation of one training dataset (ref. 285), and a UK survey in which 17% of adults exposed to sexual deepfakes believed some depicted minors (ref. 286). These sources establish that AI-generated CSAM exists, not that it is a well-established harm of general-purpose AI in the same evidentiary sense as, for example, biased outputs with multiple replicated studies. The report should either strengthen the citation base for this item or qualify the Key Findings bullet to match the evidence level presented in the body.
minor comments (4)
  1. [Page 11, heading] The running header on page 11 contains a duplicated and malformed line: 'Update on latest AI h Update on latest AI advances after the writing of this report: Chair's note.' Please correct the heading.
  2. [Introduction, independence wording (p.10)] The report says the experts 'collectively had full discretion over its content' while the Expert Advisory Panel members were nominated by governments; consider adding one sentence clarifying how panel nomination relates to independence, so readers do not infer governmental control of content.
  3. [Figure 1.5 note (p.50)] The note says o1-mini writes chains of thought 'that users cannot access before producing a final answer'; please clarify whether this refers to hidden reasoning tokens visible only to the API provider or to a lack of user-facing transparency, because the phrase is ambiguous.
  4. [General presentation] Several 'Key Definitions' boxes repeat nearly identical definitions across sections (e.g., 'AI agent' appears in sections 1.2, 1.3, and 2.1.2). Consolidating them would reduce redundancy and improve readability, although this is not a substantive issue.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the report is an expert synthesis, not a derivation from fitted parameters or a self-citation chain.

full rationale

The report does not contain a mathematical derivation chain, fitted parameters, or predictions that reduce by construction to their inputs. Its central claims, such as 'Several harms from general-purpose AI are already well established' and 'As general-purpose AI becomes more capable, evidence of additional risks is gradually emerging', are qualitative judgments supported by cited empirical studies and structured expert deliberation, not consequences of the report's own definitions or of self-citation. The only self-referential elements are procedural and presentational: the Expert Advisory Panel is nominated by governments, some figures are credited to 'International AI Safety Report', and some cited studies are authored by contributors. None of these are load-bearing in the sense of making a prediction equal to its input. The report repeatedly acknowledges evidence limitations, for example that 'reliable statistics on the frequency and impact of these incidents are lacking' (§2.1.1) and that 'researchers have not found evidence of widespread privacy violations' (§2.3.5). Those caveats may weaken the strength of the headline claims, and concerns about source selection and calibration are legitimate, but they are not circularity. The capability projections in §1.3 are explicitly framed as trend extrapolations, not as deductions from the phenomenon they purport to predict. Overall, the report is self-contained as a synthesis and does not exhibit circular reasoning.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The report introduces no free parameters or invented entities. Its conclusions rest on definitional and methodological assumptions about what counts as general-purpose AI, what counts as high-quality evidence, and whether the invited expert panel represents the broader scientific community.

assumptions (3)
  • domain assumption General-purpose AI is a meaningful category worth analyzing separately from narrow AI and AGI.
    The report's scope, risk classification, and policy relevance depend on this definitional choice, set out in the Introduction and 'About this report'.
  • domain assumption Expert-selected, partly non-peer-reviewed sources can serve as reliable scientific evidence for risk assessment.
    The Introduction lists qualitative quality criteria but no systematic review protocol, so the synthesis rests on the judgment of the expert panel.
  • domain assumption The views of the 96 invited experts adequately represent the range of scientific opinion on AI safety.
    The report treats expert disagreement and consensus as evidence, and this assumes the panel is representative, as stated in the Foreword and 'About this report'.

how reviews work

0 comments
Cite this review

Pith. "Pith review of International AI Safety Report." pith.science (2026). https://pith.science/paper/VVBXBU7O

@misc{pith2026250117805,
  author       = {Pith},
  title        = {Pith review of: International AI Safety Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VVBXBU7O}},
  note         = {Machine review of arXiv:2501.17805}
}
read the original abstract

The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by the nations attending the AI Safety Summit in Bletchley, UK. Thirty nations, the UN, the OECD, and the EU each nominated a representative to the report's Expert Advisory Panel. A total of 100 AI experts contributed, representing diverse perspectives and disciplines. Led by the report's Chair, these independent experts collectively had full discretion over the report's content.

Figures

Figures reproduced from arXiv: 2501.17805 by the authors.

Figure 0.1
Figure 0.1. Scores of notable general-purpose AI models on key benchmarks from June 2023 to December 2024. o3 showed significantly improved performance compared to the previous state of the art (shaded region). These benchmarks are some of the field’s most challenging tests of programming, abstract reasoning, and scientific reasoning. For the unreleased o3, the announcement date is shown; for the other models, the release date … view at source ↗
Figure 1.1
Figure 1.1. Today's general-purpose AI models are neural networks, which are inspired by the animal brain. These networks are composed of connected nodes, where the strength of connections between nodes are called 'weights'. Weights are updated through iterative training with large quantities of data. Source: International AI Safety Report. There are many different types of general-purpose AI, but they are developed using commo… view at source ↗
Figure 1.2
Figure 1.2. The process of developing and deploying general-purpose AI follows a series of distinct stages, from data collection and pre-processing to post-deployment monitoring. Source: International AI Safety Report. Before training a general-purpose AI model, developers collect and prepare suitable data, which is a large-scale operation. Creating high-quality training datasets involves complex pipelines of data collection, c… view at source ↗
Figures from the paper (24 more)
Figure 1.3
Figure 1.3. Figure 1.3: Since the publication of the Interim Report (May 2024), general-purpose AI models have seen rapid performance increases in answering PhD-level science questions. Researchers have been testing models on GPQA Diamond, a collection of challenging multiple-choice questio…
Figure 1
Figure 1. Figure 1: ). For example, consider the MATH benchmark [PITH_FULL_IMAGE:figures/full_fig_p048_1.png]
Figure 1.4
Figure 1.4. Figure 1.4: Performance of AI models on various benchmarks has advanced rapidly between 1998 to 2024. Note that some earlier results used machine learning AI models that are not general-purpose models. On some recent benchmarks, models progressed within a short period of time fr…
Figure 1.5
Figure 1.5. Figure 1.5: This graph shows how general-purpose language models have become markedly more cost efficient to use, measured by the number of words generated per dollar while maintaining a given performance level on the MMLU benchmark. The version of GPT-3 175B released after Sept…
Figure 1.6
Figure 1.6. Figure 1.6: Performance (as measured by ‘training loss’) improves predictably as AI developers use more compute for training (lower ‘training loss’ means better performance) (157*). In this experiment, additional compute was allocated to training larger language models (more par…
Figure 1.7
Figure 1.7. Figure 1.7: AI developers have consistently used more compute to train notable machine learning models over time, at an increasing pace since 2010 (26, 197). Computation is measured in total FLOP (floating point operations) estimated from AI literature — this refers to the numbe…
Figure 1.8
Figure 1.8. Figure 1.8: Four physical constraints to training general-purpose AI models using more compute by 2030. There are many orders of magnitude of uncertainty in the overall estimates, but training runs using 10,000x more computation than GPT-4 (released in 2023), which is in line wi…
Figure 1.9
Figure 1.9. Figure 1.9: In a set of experiments, LLM-based AI agents released in 2024 performed better than expert human engineers at open-ended AI research engineering tasks when both were given two hours or less to complete the work. Conversely, human experts performed better when given e…
Figure 2.1
Figure 2.1. Figure 2.1: Multiple stages lie between an initial intent to manipulate public opinion and a potential impact on society. While there is strong evidence for technical capability to create AI-generated content, evidence becomes sparse at later stages, reflecting research gaps rat…
Figure 2.2
Figure 2.2. Figure 2.2: Recent advances in AI models' ability to find and exploit cybersecurity vulnerabilities autonomously has grown across multiple benchmarks. On DARPA and ARPA-H's AI Cyber Challenge (353, 359), OpenAI's new o1 model (Sept 2024) substantially outperformed GPT-4o (May 20…
Figure 2
Figure 2. Figure 2: illustrates the [PITH_FULL_IMAGE:figures/full_fig_p074_2.png]
Figure 2.3
Figure 2.3. Figure 2.3: Dual-use capabilities in biology have been increasing over time for LLMs (2*), biological general-purpose AI such as AlphaFold3 (23) and specialised (not general-purpose) models relevant to pathogens (390). This chart shows performance scores, calculated as percentag…
Figure 2.4
Figure 2.4. Figure 2.4: Overview of a typical chemical and biological product development pipeline, which parallels the process used for creating chemical and biological weapons. LLMs can aid in the planning and acquisition stages, advise on performing laboratory work to build and test a de…
Figure 2.5
Figure 2.5. Figure 2.5: There are multiple kinds of ‘loss of control’ scenarios, depending on whether or not AI systems actively undermine human control and, if they do, whether or not they have been actively designed or instructed to do so. So far, ‘active’ and unintentional loss of contro…
Figure 2.6
Figure 2.6. Figure 2.6: Large Language Models (LLMs) have an unequal economic impact on different parts of the income distribution. Exposure is highest for worker tasks at the upper end of annual wages, peaking at approximately $90,000/year in the US, while low and middle incomes are signif…
Figure 2.7
Figure 2.7. Figure 2.7: So far, generative AI appears to have been adopted at a faster pace than PCs or the internet in the US. Faster adoption compared with PCs is driven by much greater use outside of work, probably due to differences in portability and cost. Source: Bick et al., 2024 (66…
Figure 2.8
Figure 2.8. Figure 2.8: Estimated training costs of AI models have sharply increased over the past few years. Only a few companies can afford to train models at such high cost, further increasing market concentration. Source: Maslej et al., 2024a (730). Access to massive datasets is crucial…
Figure 2.9
Figure 2.9. Figure 2.9: Amazon (AWS), Microsoft (Azure) and Google together control over ⅔ of global cloud computing services, concentrating power over essential AI training and deployment infrastructure in just three companies. Source: Richter, 2024 (756) [PITH_FULL_IMAGE:figures/full_fig…
Figure 2.10
Figure 2.10. Figure 2.10: US data centre energy use is projected to grow rapidly, reaching between 270–930 TWh annually by 2030. This wide range in projections (varying by over 600 TWh, equivalent to >10% of total US 2022 energy use) stems from rapidly evolving technology and limited histori…
Figure 2.11
Figure 2.11. Figure 2.11: Risks to privacy from general-purpose AI fall into three risk groups: 1. Training risks: risks associated with training on sensitive data, 2. Use risks: risks related to handling sensitive information during the use of general-purpose AI, and 3. Intentional harm ris…
Figure 2.12
Figure 2.12. Figure 2.12: The benefits of using large quantities of training data can have cascading consequences for data transparency, web crawling, and the norms of sharing information on the web. Source: International AI Safety Report [PITH_FULL_IMAGE:figures/full_fig_p145_2_12.png]
Figure 3.1
Figure 3.1. Figure 3.1: Technical challenges for managing general-purpose AI risks can be divided into two types: challenges with training and evaluating systems, and challenges with deploying them. This section discusses six broad challenges that apply to many risks. Source: International …
Figure 3.2
Figure 3.2. Figure 3.2: Monitoring and intervention techniques are system-level safeguards that can be applied to general-purpose AI system inputs, outputs, and models themselves in order to help researchers and developers monitor AI behaviour and, if necessary, intervene. Source: Internati…
Figure 3.3
Figure 3.3. Figure 3.3: Actionable methods exist for mitigating privacy harms from general-purpose AI systems, including removal of PII from training data, using on-device models, and strengthening cybersecurity. The methods are ranked based on their relative feasibility within each risk gr…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

    cs.CL 2026-07 accept novelty 7.0 of 10

    Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.

  2. Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment

    cs.LG 2026-01 conditional novelty 7.0 of 10

    Static black-box alignment evaluation cannot certify post-update safety, because overparameterized models can conceal misbehavior that a single benign gradient step activates.

  3. SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

    cs.AI 2025-05 conditional novelty 7.0 of 10

    Frontier LLMs pass fewer than 58% of systematically varied safety-fact scenarios, revealing weak generalization of critical safety knowledge to naive user queries.

  4. Hardware Mechanisms to Dynamically Throttle AI Performance

    cs.AR 2026-07 conditional novelty 6.0 of 10

    Dynamic microarchitecture throttling of GPU memory resources can cut LLM inference performance by up to 80% with low hardware overhead, giving architects a continuous, hardware-enforced AI capability control.

  5. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

  6. Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A game-theoretic model shows media can act as a soft regulator of AI safety, but only when media signals are reliable and costs are low; otherwise defection can persist.

  7. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  8. The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

    cs.CY 2025-07 conditional novelty 6.0 of 10

    On Llama-3 and Qwen-2.5, removing safety guardrails sharply raises compliance with dangerous bio, chem, and cyber requests, and the resulting safety gap grows with model scale.

  9. Technical Options for Flexible Hardware-Enabled Guarantees

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A hardware 'interlock' placed on AI accelerator network paths could provide privacy-preserving, verifiable guarantees about AI compute usage, according to a design analysis that sketches FLOP-counting and update protocols.

  10. Reconsidering LLM Uncertainty Estimation Methods in the Wild

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.

  11. The AI Agent Index

    cs.SE 2025-02 accept novelty 6.0 of 10

    The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.

  12. How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

    cs.CR 2026-07 conditional novelty 5.0 of 10

    Attacks that break LLMs best are not the ones that improve safety most; a Shapley- and greedy-based framework that selects attack subsets by downstream defender utility outperforms attacker-centric and attribution-onl...

  13. On the Principles of Deep Feedforward ReLU Networks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.

  14. The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

    cs.CY 2026-06 conditional novelty 5.0 of 10

    AI research agents remove the human accountability backstop that prior science verification relied on, so the paper proposes observable-by-default workflows, tiered verification, and AI attribution standards to preser...

  15. Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?

    cs.CL 2025-09 reject novelty 5.0 of 10

    Across 23 models from 135M to 32B parameters, internal factual knowledge scales about twice as fast with model size as linguistic competence, supporting modular small-model-plus-retrieval systems.

  16. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  17. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

  18. A Taxonomy of Omnicidal Futures Involving Artificial Intelligence

    cs.AI 2025-07 accept novelty 5.0 of 10

    A conceptual taxonomy dividing AI-driven omnicide into unintentional, intentional-by-state, intentional-by-institution, intentional-by-individual, and intentional-by-AI scenarios.

  19. LLM Agents Should Employ Security Principles

    cs.CR 2025-05 conditional novelty 5.0 of 10

    A position paper proposing AgentSandbox, a framework that applies Saltzer-Schroeder security principles to LLM agents and reports large attack-success-rate reductions on AgentDojo.

  20. Mitigating Deceptive Alignment via Self-Monitoring

    cs.AI 2025-05 conditional novelty 5.0 of 10

    CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.

  21. Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper introduces an Explanatory Virtues Framework and argues, via a qualitative rubric, that Compact Proofs are the most promising method for mechanistic interpretability.

  22. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0 of 10

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

  23. Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 4.0 of 10

    A survey mapping how neurosymbolic AI could address safety, regulatory, and operational challenges in advanced air mobility.

  24. Persuasion and Safety in the Era of Generative AI

    cs.CY 2025-05 unverdicted novelty 4.0 of 10

    This is a student dissertation proposal outlining future work to build a taxonomy, dataset, and LLM benchmark for distinguishing rational persuasion from manipulation; no experiments or results are reported.

  25. AI Awareness

    cs.AI 2025-04 accept novelty 4.0 of 10

    A review arguing that AI awareness is a measurable, four-dimensional functional capacity (metacognition, self, social, situational) that current LLMs partially exhibit and that both improves AI and creates safety risks.

  26. EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models

    cs.CL 2025-04 conditional novelty 4.0 of 10

    EasyEdit2 packages test-time steering methods into one configurable framework, adds merging of steering vectors for multi-objective control, and reports safety and sentiment results on Gemma-2-9B and Qwen2.5-7B.

  27. Aligning Generalisation Between Humans and Machines

    cs.AI 2024-11 unverdicted novelty 4.0 of 10

    A perspective paper maps how humans and machines generalize differently and argues that aligning these generalization behaviors is essential for human-AI teaming.

  28. From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs

    cs.CR 2025-06 conditional novelty 3.0 of 10

    LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.

Reference graph

Works this paper leans on

1 extracted references · 1 canonical work pages · cited by 28 Pith papers

  1. [1]

    298 Any enquiries regarding this publication should be sent to: secretariat.aistateofscience@dsit.gov.uk Research series number: DSIT 2024/000 Published in January 2025 by the UK Government @Crown copyright 2025

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.