Pith. sign in

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learning and develop a two-way deconfounder algorithm that devises a neural tensor network to simultaneously learn both the unmeasured confounders and the system dynamics, based on which a model-based estimator can be constructed for consistent policy value estimation. We illustrate the effectiveness of the proposed estimator through theoretical results and numerical experiments.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Training Large Language Models for Self-Explanation Faithfulness

cs.LG · 2026-07-23 · conditional · novelty 6.0

RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Training Large Language Models for Self-Explanation Faithfulness cs.LG · 2026-07-23 · conditional · none · ref 104 · internal anchor

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.