REVIEW 4 cited by
POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -- most of which rely crucially on reward-regression models or importance-weighted policy gradients -- fail due to excessive bias or variance. To overcome these issues in OPL, we propose a novel two-stage algorithm, called Policy Optimization via Two-Stage Policy Decomposition (POTEC). It leverages clustering in the action space and learns two different policies via policy- and regression-based approaches, respectively. In particular, we derive a novel low-variance gradient estimator that enables to learn a first-stage policy for cluster selection efficiently via a policy-based approach. To select a specific action within the cluster sampled by the first-stage policy, POTEC uses a second-stage policy derived from a regression-based approach within each cluster. We show that a local correctness condition, which only requires that the regression model preserves the relative expected reward differences of the actions within each cluster, ensures that our policy-gradient estimator is unbiased and the second-stage policy is optimal. We also show that POTEC provides a strict generalization of policy- and regression-based approaches and their associated assumptions. Comprehensive experiments demonstrate that POTEC provides substantial improvements in OPL effectiveness particularly in large and structured action spaces.
Forward citations
Cited by 4 Pith papers
-
Hadronic screening masses in thermal QCD up to the electroweak scale
Lattice results for hadronic screening masses (including baryonic and preliminary non-static mesonic modes) up to the electroweak scale show persistent higher-order and non-perturbative deviations from 3D effective-th...
-
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
A new importance-weighted estimator, OPFV, estimates and optimizes future policy value in non-stationary bandit environments by leveraging recurring time features in historical logs.
-
A General Framework for Off-Policy Learning with Partially-Observed Reward
HyPeR is a doubly robust policy-gradient estimator that uses secondary rewards to reduce variance when target rewards are only partially observed, with data-driven tuning of the mixing weight.
-
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
The paper claims a new family of embedding-based off-policy estimators for ranking policies, but its central unbiasedness theorem is false under the stated assumptions.
Discussion (0). Continue with ORCID to comment.