Pith. sign in

REVIEW 2 cited by

Counterfactual Off-Policy Evaluation with Gumbel-Max Structural Causal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.05824 v3 pith:V2TRBCZY submitted 2019-05-14 cs.LG stat.ML

classification cs.LGstat.ML
keywords episodescounterfactualoff-policypolicyprocedurecausaldifferenceevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce an off-policy evaluation procedure for highlighting episodes where applying a reinforcement learned (RL) policy is likely to have produced a substantially different outcome than the observed policy. In particular, we introduce a class of structural causal models (SCMs) for generating counterfactual trajectories in finite partially observable Markov Decision Processes (POMDPs). We see this as a useful procedure for off-policy "debugging" in high-risk settings (e.g., healthcare); by decomposing the expected difference in reward between the RL and observed policy into specific episodes, we can identify episodes where the counterfactual difference in reward is most dramatic. This in turn can be used to facilitate review of specific episodes by domain experts. We demonstrate the utility of this procedure with a synthetic environment of sepsis management.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. To Measure or Not: A Cost-Sensitive, Selective Measuring Environment for Agricultural Management Decisions with Reinforcement Learning

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A cost-sensitive reinforcement learning environment shows an agent can learn when to pay for crop measurements to guide nitrogen fertilization in winter wheat.

  2. CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Using imperfect counterfactual annotations only in the reward model part of a doubly robust estimator is the theoretically and empirically safest way to incorporate them.

Pith tools