REVIEW 6 cited by
A Survey on Causal Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Reinforcement Learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lack of interpretability. Interestingly, many researchers have leveraged insights from the causality literature recently, bringing forth flourishing works to unify the merits of causality and address well the challenges from RL. As such, it is of great necessity and significance to collate these Causal Reinforcement Learning (CRL) works, offer a review of CRL methods, and investigate the potential functionality from causality toward RL. In particular, we divide existing CRL approaches into two categories according to whether their causality-based information is given in advance or not. We further analyze each category in terms of the formalization of different models, ranging from the Markov Decision Process (MDP), Partially Observed Markov Decision Process (POMDP), Multi-Arm Bandits (MAB), and Dynamic Treatment Regime (DTR). Moreover, we summarize the evaluation matrices and open sources while we discuss emerging applications, along with promising prospects for the future development of CRL.
Forward citations
Cited by 6 Pith papers
-
YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate
An uncertainty-typed proposition IR plus Assumption-Robust Pareto Frontiers (ARPF) with a regret certificate cuts held-out regret by >90% under misspecification and beats status-quo and naive rules on real marketing data.
-
Property-driven Causal Abstractions for Markov Decision Processes
Property-driven feature causes on factored MDPs yield small abstractions (MDP/IMDP/SG) that often preserve near-optimal policies and can transfer to larger model variants.
-
Training Large Language Models for Self-Explanation Faithfulness
RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.
-
A Causal Lens for Learning Long-term Fair Policies
Qualification gain parity in RL decomposes into direct, indirect, and spurious policy effects, and the direct effect is tied to benefit fairness in a constrained PPO objective.
-
Causal Information Prioritization for Efficient Reinforcement Learning
CIP combines DirectLiNGAM-style causal masks for state-reward and action-reward links with counterfactual data augmentation and an empowerment objective to improve RL sample efficiency.
-
Causal-Inspired Multi-Agent Decision-Making via Graph Reinforcement Learning
A graph-RL driving agent using VGAE-based causal feature extraction achieves lower collision rates and higher rewards at a simulated unsignalized intersection than graph-RL baselines.
Discussion (0). Sign in to comment.