REVIEW 6 cited by
Deployment Corrections: An incident response framework for frontier AI models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
A comprehensive approach to addressing catastrophic risks from AI models should cover the full model lifecycle. This paper explores contingency plans for cases where pre-deployment risk management falls short: where either very dangerous models are deployed, or deployed models become very dangerous. Informed by incident response practices from industries including cybersecurity, we describe a toolkit of deployment corrections that AI developers can use to respond to dangerous capabilities, behaviors, or use cases of AI models that develop or are detected after deployment. We also provide a framework for AI developers to prepare and implement this toolkit. We conclude by recommending that frontier AI developers should (1) maintain control over model access, (2) establish or grow dedicated teams to design and maintain processes for deployment corrections, including incident response plans, and (3) establish these deployment corrections as allowable actions with downstream users. We also recommend frontier AI developers, standard-setting organizations, and regulators should collaborate to define a standardized industry-wide approach to the use of deployment corrections in incident response. Caveat: This work applies to frontier AI models that are made available through interfaces (e.g., API) that provide the AI developer or another upstream party means of maintaining control over access (e.g., GPT-4 or Claude). It does not apply to management of catastrophic risk from open-source models (e.g., BLOOM or Llama-2), for which the restrictions we discuss are largely unenforceable.
Forward citations
Cited by 6 Pith papers
-
Third-party compliance reviews for frontier AI safety frameworks
Independent third-party reviews can verify frontier AI companies' adherence to their safety frameworks, with practical design options for reviewer type, information access, assessment, disclosure, enforcement, and timing.
-
Adapting Probabilistic Risk Assessment for AI
The paper introduces PRA for AI, a hazard taxonomy-driven framework and workbook tool that produces banded likelihood and severity risk estimates for AI systems.
-
Bare Minimum Mitigations for Autonomous AI Development
A position paper proposing two thresholds and four minimum safeguards for frontier AI labs before AI agents automate AI R&D.
-
Authenticated Delegation and Authorized AI Agents
A framework extending OAuth 2.0 and OpenID Connect with agent-ID and delegation tokens so AI agents can act on behalf of verified humans with auditable, limited permissions.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
Generative AI in Financial Institution: A Global Survey of Opportunities, Threats, and Regulation
A survey of generative AI applications, cyber threats, and regulatory approaches in global finance, with practical recommendations but no new empirical or theoretical contribution.
Discussion (0). Continue with ORCID to comment.