Pith. sign in

REVIEW 6 cited by

Deployment Corrections: An incident response framework for frontier AI models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.00328 v1 pith:N4CKL34A submitted 2023-09-30 cs.CY

classification cs.CY
keywords modelsdeploymentcorrectionsdevelopersfrontierincidentresponsedangerous
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A comprehensive approach to addressing catastrophic risks from AI models should cover the full model lifecycle. This paper explores contingency plans for cases where pre-deployment risk management falls short: where either very dangerous models are deployed, or deployed models become very dangerous. Informed by incident response practices from industries including cybersecurity, we describe a toolkit of deployment corrections that AI developers can use to respond to dangerous capabilities, behaviors, or use cases of AI models that develop or are detected after deployment. We also provide a framework for AI developers to prepare and implement this toolkit. We conclude by recommending that frontier AI developers should (1) maintain control over model access, (2) establish or grow dedicated teams to design and maintain processes for deployment corrections, including incident response plans, and (3) establish these deployment corrections as allowable actions with downstream users. We also recommend frontier AI developers, standard-setting organizations, and regulators should collaborate to define a standardized industry-wide approach to the use of deployment corrections in incident response. Caveat: This work applies to frontier AI models that are made available through interfaces (e.g., API) that provide the AI developer or another upstream party means of maintaining control over access (e.g., GPT-4 or Claude). It does not apply to management of catastrophic risk from open-source models (e.g., BLOOM or Llama-2), for which the restrictions we discuss are largely unenforceable.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Third-party compliance reviews for frontier AI safety frameworks

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Independent third-party reviews can verify frontier AI companies' adherence to their safety frameworks, with practical design options for reviewer type, information access, assessment, disclosure, enforcement, and timing.

  2. Adapting Probabilistic Risk Assessment for AI

    cs.AI 2025-04 conditional novelty 5.0 of 10

    The paper introduces PRA for AI, a hazard taxonomy-driven framework and workbook tool that produces banded likelihood and severity risk estimates for AI systems.

  3. Bare Minimum Mitigations for Autonomous AI Development

    cs.CY 2025-04 conditional novelty 5.0 of 10

    A position paper proposing two thresholds and four minimum safeguards for frontier AI labs before AI agents automate AI R&D.

  4. Authenticated Delegation and Authorized AI Agents

    cs.CY 2025-01 conditional novelty 5.0 of 10

    A framework extending OAuth 2.0 and OpenID Connect with agent-ID and delegation tokens so AI agents can act on behalf of verified humans with auditable, limited permissions.

  5. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  6. Generative AI in Financial Institution: A Global Survey of Opportunities, Threats, and Regulation

    cs.CR 2025-04 unverdicted

    A survey of generative AI applications, cyber threats, and regulatory approaches in global finance, with practical recommendations but no new empirical or theoretical contribution.

Pith tools