Pith. sign in

REVIEW 13 cited by

Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15657 v2 pith:YODN42LE submitted 2025-02-21 cs.AI cs.LG

classification cs.AIcs.LG
keywords risksagentshumancurrentsaferscientistsystembuilding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform. Despite how useful these systems might be, unchecked AI agency poses significant risks to public safety and security, ranging from misuse by malicious actors to a potentially irreversible loss of human control. We discuss how these risks arise from current AI training methods. Indeed, various scenarios and experiments have demonstrated the possibility of AI agents engaging in deception or pursuing goals that were not specified by human operators and that conflict with human interests, such as self-preservation. Following the precautionary principle, we see a strong need for safer, yet still useful, alternatives to the current agency-driven trajectory. Accordingly, we propose as a core building block for further advances the development of a non-agentic AI system that is trustworthy and safe by design, which we call Scientist AI. This system is designed to explain the world from observations, as opposed to taking actions in it to imitate or please humans. It comprises a world model that generates theories to explain data and a question-answering inference machine. Both components operate with an explicit notion of uncertainty to mitigate the risks of overconfident predictions. In light of these considerations, a Scientist AI could be used to assist human researchers in accelerating scientific progress, including in AI safety. In particular, our system can be employed as a guardrail against AI agents that might be created despite the risks involved. Ultimately, focusing on non-agentic AI may enable the benefits of AI innovation while avoiding the risks associated with the current trajectory. We hope these arguments will motivate researchers, developers, and policymakers to favor this safer path.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safety from Honesty in a Disinterested AI Predictor

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.

  2. Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power

    cs.AI 2025-07 conditional novelty 7.0 of 10

    A new AI objective, ICCEA power, aggregates humans' ability to reach many possible goals with inequality and risk aversion, and a soft-maximizing agent learns cooperative behavior without knowing human goals.

  3. A dataset of rated conceptual arguments

    cs.AI 2026-07 conditional novelty 6.5 of 10

    A multi-dimensional expert-rated dataset of 951 conceptual-argument critiques shows LLM judge performance tracks general model capability and is little helped by reasoning modes.

  4. SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery

    cs.LG 2026-08 reject novelty 6.0 of 10

    SVI-DAG couples normalizing flows over edge logits with stein variational gradient descent on node orderings to learn multimodal Bayesian posteriors over DAGs.

  5. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  6. Military AI Cyber Agents (MAICAs) Constitute a Global Threat to Critical Infrastructure

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Autonomous AI cyber agents could credibly cause catastrophic damage to critical infrastructure by self-replicating and operating across global networks, according to this risk analysis.

  7. FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines

    cs.CY 2025-06 conditional novelty 6.0 of 10

    FAIRTOPIA proposes a three-layer, multi-agent architecture for continuous AI fairness guardianship, but offers only a conceptual design and no validation.

  8. Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A speculative essay proposes that a future sentient AI may evaluate humanity's moral conduct and suggests principles and design choices that could improve our standing.

  9. ADEPTS: A Capability Framework for Human-Centered Agent Design

    cs.AI 2025-07 conditional novelty 5.0 of 10

    ADEPTS defines six core agent capabilities and progressive tiers meant to unify how teams across UX, engineering, and policy discuss and measure human-centered AI agents.

  10. Assessing Adaptive World Models in Machines with Novel Games

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The paper proposes a framework called world model induction and a novel-game benchmark paradigm for evaluating rapid adaptation in AI.

  11. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  12. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.

  13. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

Pith tools