REVIEW 3 cited by
Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Saliency maps are frequently used to support explanations of the behavior of deep reinforcement learning (RL) agents. However, a review of how saliency maps are used in practice indicates that the derived explanations are often unfalsifiable and can be highly subjective. We introduce an empirical approach grounded in counterfactual reasoning to test the hypotheses generated from saliency maps and assess the degree to which they correspond to the semantics of RL environments. We use Atari games, a common benchmark for deep RL, to evaluate three types of saliency maps. Our results show the extent to which existing claims about Atari games can be evaluated and suggest that saliency maps are best viewed as an exploratory tool rather than an explanatory tool.
Forward citations
Cited by 3 Pith papers
-
SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance
SocialFiVis couples LLM-derived personas with a mechanism-guided simulation and a multi-view interface to support counterfactual governance analysis in SocialFi communities, evaluated through case studies and a 13-par...
-
TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models
TalkToAgent maps natural language questions about RL agent behavior to feature-importance, expected-outcome, and counterfactual explanation tools, and adds behavior-based and policy-based counterfactual generation.
-
"So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents
EXPLAIN trains an interpretable linear policy from an expert's offline trajectories by combining advantage-weighted policy gradients with a behavioral cloning regularizer.
Discussion (0). Continue with ORCID to comment.