CSI converts offline contextual bandit learning into a binary classification problem by comparing the logged action against a counterfactual action sampled from the logging policy, and the argmax of the resulting classifier provably matches the argmax of expected reward.
Regularization and confounding in linear regression for treatment effect estimation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper investigates the use of regularization priors in the context of treatment effect estimation using observational data where the number of control variables is large relative to the number of observations. First, the phenomenon of regularization-induced confounding is introduced, which refers to the tendency of regularization priors to adversely bias treatment effect estimates by over-shrinking control variable regression coefficients. Then, a simultaneous regression model is presented which permits regularization priors to be specified in a way that avoids this unintentional re-confounding. The new model is illustrated on synthetic and empirical data.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Offline Contextual Bandit with Counterfactual Sample Identification
CSI converts offline contextual bandit learning into a binary classification problem by comparing the logged action against a counterfactual action sampled from the logging policy, and the argmax of the resulting classifier provably matches the argmax of expected reward.