Pith. sign in

REVIEW 15 cited by

Safety cases for frontier AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21572 v1 pith:AWEDBIHF submitted 2024-10-28 cs.CY

classification cs.CY
keywords safetycasesfrontierexplainsafesystemsalreadyargument
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured argument, supported by evidence, that a system is safe enough in a given operational context. Safety cases are already common in other safety-critical industries such as aviation and nuclear power. In this paper, we explain why they may also be a useful tool in frontier AI governance, both in industry self-regulation and government regulation. We then discuss the practicalities of safety cases, outlining how to produce a frontier AI safety case and discussing what still needs to happen before safety cases can substantially inform decisions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements

    cs.CY 2026-06 conditional novelty 6.0 of 10

    Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.

  2. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  3. Levels of Autonomy for AI Agents

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.

  4. An Example Safety Case for Safeguards Against Misuse

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A proposed framework, built around an 'uplift model' that translates red-team safeguard-evasion data into estimated risk, for justifying that AI misuse safeguards keep large-scale harm risk below a threshold.

  5. Evaluating Frontier Models for Stealth and Situational Awareness

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Five frontier AI models fail most of a new suite of stealth and situational awareness tests, which the authors use to argue that current models likely cannot cause severe harm through scheming.

  6. Assessing confidence in frontier AI safety cases

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Applying Assurance 2.0 to a cyber-misuse safety case, the authors show that high top-level confidence requires extremely high confidence in every component, and propose an LLM-based Delphi for eliciting those componen...

  7. Towards Frontier Safety Policies Plus

    cs.CY 2025-01 conditional novelty 6.0 of 10

    Frontier safety policies should be rebuilt around a standardized taxonomy of precursory capabilities and a mutual feedback mechanism with AI safety cases.

  8. Safety case template for frontier AI: A cyber inability argument

    cs.CY 2024-11 accept novelty 6.0 of 10

    A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.

  9. Systematic Hazard Analysis for Frontier AI using STPA

    cs.CY 2025-06 conditional novelty 5.0 of 10

    Applying STPA to the AI Control scenario produces structured unsafe control actions and loss scenarios, supporting an argument that systematic hazard analysis can improve frontier AI safety assurance.

  10. Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

    cs.AI 2025-05 reject novelty 5.0 of 10

    SafetyNet is an ensemble of standard outlier detectors for LLM backdoor monitoring, but its key mechanistic claim and headline numbers are contradicted by inconsistent tables and a mismatched abstract.

  11. Justified Evidence Collection for Argument-based AI Fairness Assurance

    cs.HC 2025-05 conditional novelty 5.0 of 10

    The paper introduces a dynamic argument-based assurance framework that gathers evidence from model, data, and use case transparency artefacts to support fairness claims, illustrated on a finance sentiment analysis use case.

  12. Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development

    cs.CY 2025-01 conditional novelty 5.0 of 10

    Gradual AI progress could permanently remove human influence from society's core systems, an underappreciated existential risk path.

  13. Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.

  14. Accountability Asymmetry and Structural Trust in Autonomous AI Systems

    cs.CY 2026-08 accept novelty 4.0 of 10

    Accountability asymmetry means autonomous AI should be governed like infrastructure, with independent review and audit, not treated as moral actors.

  15. Catastrophic Liability: Managing Systemic Risks in Frontier AI Development

    cs.CY 2025-05 conditional novelty 4.0 of 10

    Voluntary AI safety standards may already bind frontier labs through US tort law, making thorough safety documentation a liability shield.

Pith tools