Pith. sign in

REVIEW 4 cited by

POTEC: Off-Policy Learning for Large Action Spaces via Two-Stage Policy Decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06151 v1 pith:KSNFZQFT submitted 2024-02-09 stat.ML cs.LG

classification stat.MLcs.LG
keywords policyactionpotecclusterlargeregression-basedspacestwo-stage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study off-policy learning (OPL) of contextual bandit policies in large discrete action spaces where existing methods -- most of which rely crucially on reward-regression models or importance-weighted policy gradients -- fail due to excessive bias or variance. To overcome these issues in OPL, we propose a novel two-stage algorithm, called Policy Optimization via Two-Stage Policy Decomposition (POTEC). It leverages clustering in the action space and learns two different policies via policy- and regression-based approaches, respectively. In particular, we derive a novel low-variance gradient estimator that enables to learn a first-stage policy for cluster selection efficiently via a policy-based approach. To select a specific action within the cluster sampled by the first-stage policy, POTEC uses a second-stage policy derived from a regression-based approach within each cluster. We show that a local correctness condition, which only requires that the regression model preserves the relative expected reward differences of the actions within each cluster, ensures that our policy-gradient estimator is unbiased and the second-stage policy is optimal. We also show that POTEC provides a strict generalization of policy- and regression-based approaches and their associated assumptions. Comprehensive experiments demonstrate that POTEC provides substantial improvements in OPL effectiveness particularly in large and structured action spaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hadronic screening masses in thermal QCD up to the electroweak scale

    hep-lat 2026-03 unverdicted novelty 6.0 of 10

    Lattice results for hadronic screening masses (including baryonic and preliminary non-static mesonic modes) up to the electroweak scale show persistent higher-order and non-perturbative deviations from 3D effective-th...

  2. Off-Policy Evaluation and Learning for the Future under Non-Stationarity

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A new importance-weighted estimator, OPFV, estimates and optimizes future policy value in non-stationary bandit environments by leveraging recurring time features in historical logs.

  3. A General Framework for Off-Policy Learning with Partially-Observed Reward

    cs.LG 2025-06 conditional novelty 6.0 of 10

    HyPeR is a doubly robust policy-gradient estimator that uses secondary rewards to reduce variance when target rewards are only partially observed, with data-driven tuning of the mixing weight.

  4. Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling

    stat.ML 2025-05 reject novelty 5.0 of 10

    The paper claims a new family of embedding-based off-policy estimators for ranking policies, but its central unbiasedness theorem is false under the stated assumptions.

Pith tools