Pith. sign in

REVIEW 2 cited by

Explainable Reinforcement Learning Through a Causal Lens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.10958 v2 pith:NQULUVLC submitted 2019-05-27 cs.LG cs.AIcs.HCstat.ML

classification cs.LGcs.AIcs.HCstat.ML
keywords causalexplanationsmodelagentsbehaviourlearningmodelsreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prevalent theories in cognitive science propose that humans understand and represent the knowledge of the world through causal relationships. In making sense of the world, we build causal models in our mind to encode cause-effect relations of events and use these to explain why new events happen. In this paper, we use causal models to derive causal explanations of behaviour of reinforcement learning agents. We present an approach that learns a structural causal model during reinforcement learning and encodes causal relationships between variables of interest. This model is then used to generate explanations of behaviour based on counterfactual analysis of the causal model. We report on a study with 120 participants who observe agents playing a real-time strategy game (Starcraft II) and then receive explanations of the agents' behaviour. We investigated: 1) participants' understanding gained by explanations through task prediction; 2) explanation satisfaction and 3) trust. Our results show that causal model explanations perform better on these measures compared to two other baseline explanation models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  2. Abstract Counterfactuals for Language Model Agents

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Counterfactuals for LM agents computed over a high-level abstraction of the action, instead of its tokens, preserve the observed action's meaning across counterfactual contexts far more often than token-level counterfactuals.

Pith tools