Pith. sign in

REVIEW 1 cited by

Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.11531 v2 pith:36NM57SX submitted 2023-10-17 cs.LG cs.AIcs.SYeess.SYstat.ML

classification cs.LGcs.AIcs.SYeess.SYstat.ML
keywords learningalgorithmhorizoninfiniteofflineonlineregretbayesian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, we study the problem of efficient online reinforcement learning in the infinite horizon setting when there is an offline dataset to start with. We assume that the offline dataset is generated by an expert but with unknown level of competence, i.e., it is not perfect and not necessarily using the optimal policy. We show that if the learning agent models the behavioral policy (parameterized by a competence parameter) used by the expert, it can do substantially better in terms of minimizing cumulative regret, than if it doesn't do that. We establish an upper bound on regret of the exact informed PSRL algorithm that scales as $\tilde{O}(\sqrt{T})$. This requires a novel prior-dependent regret analysis of Bayesian online learning algorithms for the infinite horizon setting. We then propose the Informed RLSVI algorithm to efficiently approximate the iPSRL algorithm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deconfounded Warm-Start Thompson Sampling with Applications to Precision Medicine

    stat.ML 2025-05 conditional novelty 4.0 of 10

    DWTS debiases and selects features from observational data, then warm-starts Thompson sampling with those estimates, achieving lower cumulative regret than LinTS in simulations.

Pith tools