Pith. sign in

REVIEW 1 cited by

RL for Latent MDPs: Regret Guarantees and a Lower Bound

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.04939 v1 pith:FDLQ2U6P submitted 2021-02-09 cs.LG

classification cs.LG
keywords regretassumptionsconsiderepisodesgivengoodguaranteeinitialization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this work, we consider the regret minimization problem for reinforcement learning in latent Markov Decision Processes (LMDP). In an LMDP, an MDP is randomly drawn from a set of $M$ possible MDPs at the beginning of the interaction, but the identity of the chosen MDP is not revealed to the agent. We first show that a general instance of LMDPs requires at least $\Omega((SA)^M)$ episodes to even approximate the optimal policy. Then, we consider sufficient assumptions under which learning good policies requires polynomial number of episodes. We show that the key link is a notion of separation between the MDP system dynamics. With sufficient separation, we provide an efficient algorithm with local guarantee, {\it i.e.,} providing a sublinear regret guarantee when we are given a good initialization. Finally, if we are given standard statistical sufficiency assumptions common in the Predictive State Representation (PSR) literature (e.g., Boots et al.) and a reachability assumption, we show that the need for initialization can be removed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow-based Domain Randomization for Learning and Sequencing Robotic Skills

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A normalizing-flow sampling distribution, trained with entropy-regularized reward maximization, improves domain coverage and sim-to-real transfer over Gaussian, beta, and interval-based learned domain randomization.

Pith tools