Pith. sign in

REVIEW 3 cited by

Adversarial Policies Beat Superhuman Go AIs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.00241 v4 pith:EEODD2YC submitted 2022-11-01 cs.LG cs.AIcs.CRstat.ML

classification cs.LGcs.AIcs.CRstat.ML
keywords superhumanattackkatagoadversarialbeatevengo-playingpolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders. Our attack transfers zero-shot to other superhuman Go-playing AIs, and is comprehensible to the extent that human experts can implement it without algorithmic assistance to consistently beat superhuman AIs. The core vulnerability uncovered by our attack persists even in KataGo agents adversarially trained to defend against our attack. Our results demonstrate that even superhuman AI systems may harbor surprising failure modes. Example games are available https://goattack.far.ai/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    LLMs estimate pulley mechanical advantage above chance via a pulley-counting heuristic, but fail to distinguish functional from connected-but-nonfunctional systems, indicating brittle world-model use.

  2. Detecting AI Assistance in Abstract Complex Tasks

    cs.AI 2025-07 reject novelty 6.0 of 10

    Converting behavioral search traces into image channels plus an exploration/exploitation time series lets a small CNN-RNN detect AI assistance with about 86% accuracy on a balanced lab task.

  3. Foundation Model Self-Play: Open-Ended Strategy Innovation via Foundation Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Foundation models can act as search operators in multi-agent self-play, generating diverse code strategies that match or beat hand-designed baselines and automate LLM jailbreaking and patching.

Pith tools