Pith. sign in

REVIEW 12 cited by

A sketch of an AI control safety case

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17315 v1 pith:FYU3FTRQ submitted 2025-01-28 cs.AI cs.CRcs.SE

classification cs.AIcs.CRcs.SE
keywords controlcasesketchsafetydatadeploymentdevelopersexfiltrate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a structured argument that models are incapable of subverting control measures in order to cause unacceptable outcomes. As a case study, we sketch an argument that a hypothetical LLM agent deployed internally at an AI company won't exfiltrate sensitive information. The sketch relies on evidence from a "control evaluation,"' where a red team deliberately designs models to exfiltrate data in a proxy for the deployment environment. The safety case then hinges on several claims: (1) the red team adequately elicits model capabilities to exfiltrate data, (2) control measures remain at least as effective in deployment, and (3) developers conservatively extrapolate model performance to predict the probability of data exfiltration in deployment. This safety case sketch is a step toward more concrete arguments that can be used to show that a dangerously capable LLM agent is safe to deploy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A framework paper that adapts AI safety case methodology to the specific threat of manipulation attacks by internally deployed misaligned AI.

  2. Levels of Autonomy for AI Agents

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A user-role-based five-level framework for designing, certifying, and evaluating AI agent autonomy as a choice independent of agent capability.

  3. An Example Safety Case for Safeguards Against Misuse

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A proposed framework, built around an 'uplift model' that translates red-team safeguard-evasion data into estimated risk, for justifying that AI misuse safeguards keep large-scale harm risk below a threshold.

  4. AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A MIRI governance agenda argues for an internationally coordinated halt to dangerous AI development and catalogs around 400 research questions across four strategic scenarios.

  5. An alignment safety case sketch based on debate

    cs.AI 2025-05 unverdicted novelty 6.0 of 10

    The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices fo...

  6. Evaluating Frontier Models for Stealth and Situational Awareness

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Five frontier AI models fail most of a new suite of stealth and situational awareness tests, which the authors use to argue that current models likely cannot cause severe harm through scheming.

  7. AI Behind Closed Doors: a Primer on The Governance of Internal Deployment

    cs.CY 2025-04 conditional novelty 6.0 of 10

    Internal deployment of frontier AI systems is an under-governed risk area; the paper provides a conceptual map, a legal review, lessons from safety-critical industries, and a defense-in-depth governance blueprint.

  8. Systematic Hazard Analysis for Frontier AI using STPA

    cs.CY 2025-06 conditional novelty 5.0 of 10

    Applying STPA to the AI Control scenario produces structured unsafe control actions and loss scenarios, supporting an argument that systematic hazard analysis can improve frontier AI safety assurance.

  9. Towards Measurement Theory for Artificial Intelligence

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A formal measurement theory for AI, built from representational measurement theory, measure theory, metrology, and psychometrics, would make evaluations of AI systems commensurable and scientifically grounded.

  10. Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A position paper proposing that AI alignment adopt formal optimal control and a ten-layer Alignment Control Stack for organizing and interoperating control interventions.

  11. A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A synthesis of established risk management practices into a structured framework for frontier AI developers, centered on explicit risk tolerance, KRI/KCI thresholds, and governance.

  12. You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

    cs.LG 2025-02 conditional novelty 3.0 of 10

    To align powerful AI, researchers must understand how statistical patterns in training data shape the internal structure of models, because that structure, not eval scores, determines generalization.

Pith tools