Given a sufficiently accurate model of social dynamics, there provably exist policies that are near-optimal for a chosen social welfare function with high probability, plus a safety filter for arbitrary black-box policies.
Rethinking the Maturity of Artificial Intelligence in Safety-Critical Settings
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
Given a sufficiently accurate model of social dynamics, there provably exist policies that are near-optimal for a chosen social welfare function with high probability, plus a safety filter for arbitrary black-box policies.