Pith. sign in

REVIEW 1 cited by

Information based explanation methods for deep learning agents -- with applications on large open-source chess models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09702 v1 pith:I7C5ZFG5 submitted 2023-09-18 cs.LG cs.AI

classification cs.LGcs.AI
keywords chessmodelsinformationlargemethodopen-sourcealphazeromodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With large chess-playing neural network models like AlphaZero contesting the state of the art within the world of computerised chess, two challenges present themselves: The question of how to explain the domain knowledge internalised by such models, and the problem that such models are not made openly available. This work presents the re-implementation of the concept detection methodology applied to AlphaZero in McGrath et al. (2022), by using large, open-source chess models with comparable performance. We obtain results similar to those achieved on AlphaZero, while relying solely on open-source resources. We also present a novel explainable AI (XAI) method, which is guaranteed to highlight exhaustively and exclusively the information used by the explained model. This method generates visual explanations tailored to domains characterised by discrete input spaces, as is the case for chess. Our presented method has the desirable property of controlling the information flow between any input vector and the given model, which in turn provides strict guarantees regarding what information is used by the trained model during inference. We demonstrate the viability of our method by applying it to standard 8x8 chess, using large open-source chess models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Gradient Steering: Interventions from Behavioral Objectives

    cs.LG 2026-07 conditional novelty 6.0 of 10

    PGS builds a removable activation offset from return-weighted action-score gradients and steers frozen policies across gridworld, chess, and football.

Pith tools