Pith. sign in

REVIEW 1 cited by

Explaining Reinforcement Learning Agents Through Counterfactual Action Outcomes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11118 v1 pith:XQ7D7FI3 submitted 2023-12-18 cs.AI cs.LG

classification cs.AIcs.LG
keywords agentlocalactionexplanationsmethodoutcomesagentscounterfactual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Explainable reinforcement learning (XRL) methods aim to help elucidate agent policies and decision-making processes. The majority of XRL approaches focus on local explanations, seeking to shed light on the reasons an agent acts the way it does at a specific world state. While such explanations are both useful and necessary, they typically do not portray the outcomes of the agent's selected choice of action. In this work, we propose ``COViz'', a new local explanation method that visually compares the outcome of an agent's chosen action to a counterfactual one. In contrast to most local explanations that provide state-limited observations of the agent's motivation, our method depicts alternative trajectories the agent could have taken from the given state and their outcomes. We evaluated the usefulness of COViz in supporting people's understanding of agents' preferences and compare it with reward decomposition, a local explanation method that describes an agent's expected utility for different actions by decomposing it into meaningful reward types. Furthermore, we examine the complementary benefits of integrating both methods. Our results show that such integration significantly improved participants' performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning

    cs.AI 2025-06 reject novelty 6.0 of 10

    A proposed AR framework, Arvolution, visualizes past failed RL policies as ghosts to support failure analysis and a dual human-agent learning loop.

Pith tools