REVIEW 15 cited by
Safety cases for frontier AI
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured argument, supported by evidence, that a system is safe enough in a given operational context. Safety cases are already common in other safety-critical industries such as aviation and nuclear power. In this paper, we explain why they may also be a useful tool in frontier AI governance, both in industry self-regulation and government regulation. We then discuss the practicalities of safety cases, outlining how to produce a frontier AI safety case and discussing what still needs to happen before safety cases can substantially inform decisions.
Forward citations
Cited by 15 Pith papers
-
How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.
-
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.
-
Levels of Autonomy for AI Agents
A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.
-
An Example Safety Case for Safeguards Against Misuse
A proposed framework, built around an 'uplift model' that translates red-team safeguard-evasion data into estimated risk, for justifying that AI misuse safeguards keep large-scale harm risk below a threshold.
-
Evaluating Frontier Models for Stealth and Situational Awareness
Five frontier AI models fail most of a new suite of stealth and situational awareness tests, which the authors use to argue that current models likely cannot cause severe harm through scheming.
-
Assessing confidence in frontier AI safety cases
Applying Assurance 2.0 to a cyber-misuse safety case, the authors show that high top-level confidence requires extremely high confidence in every component, and propose an LLM-based Delphi for eliciting those componen...
-
Towards Frontier Safety Policies Plus
Frontier safety policies should be rebuilt around a standardized taxonomy of precursory capabilities and a mutual feedback mechanism with AI safety cases.
-
Safety case template for frontier AI: A cyber inability argument
A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.
-
Systematic Hazard Analysis for Frontier AI using STPA
Applying STPA to the AI Control scenario produces structured unsafe control actions and loss scenarios, supporting an argument that systematic hazard analysis can improve frontier AI safety assurance.
-
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
SafetyNet is an ensemble of standard outlier detectors for LLM backdoor monitoring, but its key mechanistic claim and headline numbers are contradicted by inconsistent tables and a mismatched abstract.
-
Justified Evidence Collection for Argument-based AI Fairness Assurance
The paper introduces a dynamic argument-based assurance framework that gathers evidence from model, data, and use case transparency artefacts to support fairness claims, illustrated on a finance sentiment analysis use case.
-
Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development
Gradual AI progress could permanently remove human influence from society's core systems, an underappreciated existential risk path.
-
Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations
A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.
-
Accountability Asymmetry and Structural Trust in Autonomous AI Systems
Accountability asymmetry means autonomous AI should be governed like infrastructure, with independent review and audit, not treated as moral actors.
-
Catastrophic Liability: Managing Systemic Risks in Frontier AI Development
Voluntary AI safety standards may already bind frontier labs through US tort law, making thorough safety documentation a liability shield.
Discussion (0). Continue with ORCID to comment.