Pith. sign in

REVIEW 4 cited by

A sketch of an AI control safety case

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17315 v1 pith:FYU3FTRQ submitted 2025-01-28 cs.AI cs.CRcs.SE

classification cs.AIcs.CRcs.SE
keywords controlcasesketchsafetydatadeploymentdevelopersexfiltrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a structured argument that models are incapable of subverting control measures in order to cause unacceptable outcomes. As a case study, we sketch an argument that a hypothetical LLM agent deployed internally at an AI company won't exfiltrate sensitive information. The sketch relies on evidence from a "control evaluation,"' where a red team deliberately designs models to exfiltrate data in a proxy for the deployment environment. The safety case then hinges on several claims: (1) the red team adequately elicits model capabilities to exfiltrate data, (2) control measures remain at least as effective in deployment, and (3) developers conservatively extrapolate model performance to predict the probability of data exfiltration in deployment. This safety case sketch is a step toward more concrete arguments that can be used to show that a dangerously capable LLM agent is safe to deploy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  2. Levels of Autonomy for AI Agents

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.

  3. Towards Measurement Theory for Artificial Intelligence

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A formal measurement theory for AI, built from representational measurement theory, measure theory, metrology, and psychometrics, would make evaluations of AI systems commensurable and scientifically grounded.

  4. Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper proposing that AI alignment adopt formal optimal control and a ten-layer Alignment Control Stack for organizing and interoperating control interventions.

Pith tools