Pith. sign in

REVIEW 1 cited by

Non-Stationary Off-Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.08236 v3 pith:AUCDJWPL submitted 2020-06-15 cs.LG cs.AIstat.ML

Non-Stationary Off-Policy Optimization

classification cs.LG cs.AIstat.ML
keywords off-policyapproachoptimizationdatadeploymentlearnedlearningnon-stationary
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Off-policy learning is a framework for evaluating and optimizing policies without deploying them, from data collected by another policy. Real-world environments are typically non-stationary and the offline learned policies should adapt to these changes. To address this challenge, we study the novel problem of off-policy optimization in piecewise-stationary contextual bandits. Our proposed solution has two phases. In the offline learning phase, we partition logged data into categorical latent states and learn a near-optimal sub-policy for each state. In the online deployment phase, we adaptively switch between the learned sub-policies based on their performance. This approach is practical and analyzable, and we provide guarantees on both the quality of off-policy optimization and the regret during online deployment. To show the effectiveness of our approach, we compare it to state-of-the-art baselines on both synthetic and real-world datasets. Our approach outperforms methods that act only on observed context.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Metriplectic Conditional Flow Matching for Dissipative Dynamics

    cs.LG 2025-09 conditional novelty 6.0

    Metriplectic conditional flow matching combines a conservative-dissipative vector field decomposition with flow-matching training and a Strang-prox sampler to keep damped-pendulum rollouts nearly energy-monotone.