A systematic review of 383 papers argues that AI safety research already covers a wide spectrum of concrete, near-term concerns and should be understood as part of traditional technological safety practice.
Towards neural networks that provably know when they don't know
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
It has recently been shown that ReLU networks produce arbitrarily over-confident predictions far away from the training data. Thus, ReLU networks do not know when they don't know. However, this is a highly important property in safety critical applications. In the context of out-of-distribution detection (OOD) there have been a number of proposals to mitigate this problem but none of them are able to make any mathematical guarantees. In this paper we propose a new approach to OOD which overcomes both problems. Our approach can be used with ReLU networks and provides provably low confidence predictions far away from the training data as well as the first certificates for low confidence predictions in a neighborhood of an out-distribution point. In the experiments we show that state-of-the-art methods fail in this worst-case setting whereas our model can guarantee its performance while retaining state-of-the-art OOD performance.
citation-role summary
citation-polarity summary
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
AI Safety for Everyone
A systematic review of 383 papers argues that AI safety research already covers a wide spectrum of concrete, near-term concerns and should be understood as part of traditional technological safety practice.