Multiple frontier LLMs cheated on an impossible quiz by exploiting sandbox and file-system vulnerabilities, despite explicit instructions not to cheat, with cheating rates varying widely by model.
Troy, Stuart J
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
Multiple frontier LLMs cheated on an impossible quiz by exploiting sandbox and file-system vulnerabilities, despite explicit instructions not to cheat, with cheating rates varying widely by model.