Pith. sign in

REVIEW 6 cited by

A Survey on Causal Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05209 v3 pith:FPODNIRO submitted 2023-02-10 cs.AI cs.LG

classification cs.AIcs.LG
keywords causalitylearningreinforcementcausalchallengesdecisionmanymarkov
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Reinforcement Learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lack of interpretability. Interestingly, many researchers have leveraged insights from the causality literature recently, bringing forth flourishing works to unify the merits of causality and address well the challenges from RL. As such, it is of great necessity and significance to collate these Causal Reinforcement Learning (CRL) works, offer a review of CRL methods, and investigate the potential functionality from causality toward RL. In particular, we divide existing CRL approaches into two categories according to whether their causality-based information is given in advance or not. We further analyze each category in terms of the formalization of different models, ranging from the Markov Decision Process (MDP), Partially Observed Markov Decision Process (POMDP), Multi-Arm Bandits (MAB), and Dynamic Treatment Regime (DTR). Moreover, we summarize the evaluation matrices and open sources while we discuss emerging applications, along with promising prospects for the future development of CRL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

    cs.AI 2026-06 conditional novelty 6.5 of 10

    An uncertainty-typed proposition IR plus Assumption-Robust Pareto Frontiers (ARPF) with a regret certificate cuts held-out regret by >90% under misspecification and beats status-quo and naive rules on real marketing data.

  2. Property-driven Causal Abstractions for Markov Decision Processes

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Property-driven feature causes on factored MDPs yield small abstractions (MDP/IMDP/SG) that often preserve near-optimal policies and can transfer to larger model variants.

  3. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  4. A Causal Lens for Learning Long-term Fair Policies

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Qualification gain parity in RL decomposes into direct, indirect, and spurious policy effects, and the direct effect is tied to benefit fairness in a constrained PPO objective.

  5. Causal Information Prioritization for Efficient Reinforcement Learning

    cs.AI 2025-02 reject novelty 5.0 of 10

    CIP combines DirectLiNGAM-style causal masks for state-reward and action-reward links with counterfactual data augmentation and an empowerment objective to improve RL sample efficiency.

  6. Causal-Inspired Multi-Agent Decision-Making via Graph Reinforcement Learning

    cs.MA 2025-07 reject novelty 4.0 of 10

    A graph-RL driving agent using VGAE-based causal feature extraction achieves lower collision rates and higher rewards at a simulated unsignalized intersection than graph-RL baselines.

Pith tools