REVIEW 3 cited by
Adversarial Policies Beat Superhuman Go AIs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders. Our attack transfers zero-shot to other superhuman Go-playing AIs, and is comprehensible to the extent that human experts can implement it without algorithmic assistance to consistently beat superhuman AIs. The core vulnerability uncovered by our attack persists even in KataGo agents adversarially trained to defend against our attack. Our results demonstrate that even superhuman AI systems may harbor surprising failure modes. Example games are available https://goattack.far.ai/.
Forward citations
Cited by 3 Pith papers
-
LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning
LLMs estimate pulley mechanical advantage above chance via a pulley-counting heuristic, but fail to distinguish functional from connected-but-nonfunctional systems, indicating brittle world-model use.
-
Detecting AI Assistance in Abstract Complex Tasks
Converting behavioral search traces into image channels plus an exploration/exploitation time series lets a small CNN-RNN detect AI assistance with about 86% accuracy on a balanced lab task.
-
Foundation Model Self-Play: Open-Ended Strategy Innovation via Foundation Models
Foundation models can act as search operators in multi-agent self-play, generating diverse code strategies that match or beat hand-designed baselines and automate LLM jailbreaking and patching.
Discussion (0). Continue with ORCID to comment.