Pith. sign in

Personalized Pricing with Invalid Instrumental Variables: Identification, Estimation, and Policy Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Pricing based on individual customer characteristics is widely used to maximize sellers' revenues. This work studies offline personalized pricing under endogeneity using an instrumental variable approach. Standard instrumental variable methods in causal inference/econometrics either focus on a discrete treatment space or require the exclusion restriction of instruments from having a direct effect on the outcome, which limits their applicability in personalized pricing. In this paper, we propose a new policy learning method for Personalized pRicing using Invalid iNsTrumental variables (PRINT) for continuous treatment that allow direct effects on the outcome. Specifically, relying on the structural models of revenue and price, we establish the identifiability condition of an optimal pricing strategy under endogeneity with the help of invalid instrumental variables. Based on this new identification, which leads to solving conditional moment restrictions with generalized residual functions, we construct an adversarial min-max estimator and learn an optimal pricing strategy. Furthermore, we establish an asymptotic regret bound to find an optimal pricing strategy. Finally, we demonstrate the effectiveness of the proposed method via extensive simulation studies as well as a real data application from an US online auto loan company.

fields

stat.ML 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Quantile-Optimal Policy Learning under Unmeasured Confounding

stat.ML · 2025-06-08 · conditional · novelty 7.0

Under instrumental-variable or negative-control assumptions, the authors prove a pessimism-based policy learning method achieves about 1/sqrt(n)-type regret for quantile reward objectives with unmeasured confounders.

citing papers explorer

Showing 1 of 1 citing paper.

  • Quantile-Optimal Policy Learning under Unmeasured Confounding stat.ML · 2025-06-08 · conditional · none · ref 45 · internal anchor

    Under instrumental-variable or negative-control assumptions, the authors prove a pessimism-based policy learning method achieves about 1/sqrt(n)-type regret for quantile reward objectives with unmeasured confounders.