Chain-of-thought monitoring detects bad reasoning when the task is hard enough that the model must think aloud, and current models can only evade it with significant external help.
Safety Cases: A Scalable Approach to Frontier AI Safety
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g. boards, customers, third parties) that a system is safe. In this paper, we cover how and why frontier AI developers might also want to use safety cases. We then argue that writing and reviewing safety cases would substantially assist in the fulfilment of many of the Frontier AI Safety Commitments. Finally, we outline open research questions on the methodology, implementation, and technical details of safety cases.
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
Chain-of-thought monitoring detects bad reasoning when the task is hard enough that the model must think aloud, and current models can only evade it with significant external help.