Pith. sign in

REVIEW 3 cited by

Safety case template for frontier AI: A cyber inability argument

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.08088 v1 pith:5GPJQJ7R submitted 2024-11-12 cs.CY cs.CR

classification cs.CYcs.CR
keywords safetytemplatecaseriskargumentscyberfrontiermodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offering such assurances is through a safety case: a structured, evidence-based argument aimed at demonstrating why the risk associated with a safety-critical system is acceptable. In this article, we propose a safety case template for offensive cyber capabilities. We illustrate how developers could argue that a model does not have capabilities posing unacceptable cyber risks by breaking down the main claim into progressively specific sub-claims, each supported by evidence. In our template, we identify a number of risk models, derive proxy tasks from the risk models, define evaluation settings for the proxy tasks, and connect those with evaluation results. Elements of current frontier safety techniques - such as risk models, proxy tasks, and capability evaluations - use implicit arguments for overall system safety. This safety case template integrates these elements using the Claims Arguments Evidence (CAE) framework in order to make safety arguments coherent and explicit. While uncertainties around the specifics remain, this template serves as a proof of concept, aiming to foster discussion on AI safety cases and advance AI assurance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  2. Levels of Autonomy for AI Agents

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.

  3. Systematic Hazard Analysis for Frontier AI using STPA

    cs.CY 2025-06 conditional novelty 5.0 of 10

    Applying STPA to the AI Control scenario produces structured unsafe control actions and loss scenarios, supporting an argument that systematic hazard analysis can improve frontier AI safety assurance.

Pith tools