Pith. sign in

REVIEW

Understanding Game-Playing Agents with Natural Language Annotations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07531 v1 pith:NTSRQT75 submitted 2022-04-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords annotationslanguagenaturalagentsdomain-specificgame-playinglearningmentions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a new dataset containing 10K human-annotated games of Go and show how these natural language annotations can be used as a tool for model interpretability. Given a board state and its associated comment, our approach uses linear probing to predict mentions of domain-specific terms (e.g., ko, atari) from the intermediate state representations of game-playing agents like AlphaGo Zero. We find these game concepts are nontrivially encoded in two distinct policy networks, one trained via imitation learning and another trained via reinforcement learning. Furthermore, mentions of domain-specific terms are most easily predicted from the later layers of both models, suggesting that these policy networks encode high-level abstractions similar to those used in the natural language annotations.

Discussion (0). Sign in to comment.

Pith tools