Pith. sign in

REVIEW 2 cited by

The Role of Environment Access in Agnostic Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.05405 v1 pith:DIOR6XV4 submitted 2025-04-07 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords policylearningaccessagnosticfunctionresetstateapproximation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study Reinforcement Learning (RL) in environments with large state spaces, where function approximation is required for sample-efficient learning. Departing from a long history of prior work, we consider the weakest possible form of function approximation, called agnostic policy learning, where the learner seeks to find the best policy in a given class $\Pi$, with no guarantee that $\Pi$ contains an optimal policy for the underlying task. Although it is known that sample-efficient agnostic policy learning is not possible in the standard online RL setting without further assumptions, we investigate the extent to which this can be overcome with stronger forms of access to the environment. Specifically, we show that: 1. Agnostic policy learning remains statistically intractable when given access to a local simulator, from which one can reset to any previously seen state. This result holds even when the policy class is realizable, and stands in contrast to a positive result of [MFR24] showing that value-based learning under realizability is tractable with local simulator access. 2. Agnostic policy learning remains statistically intractable when given online access to a reset distribution with good coverage properties over the state space (the so-called $\mu$-reset setting). We also study stronger forms of function approximation for policy learning, showing that PSDP [BKSN03] and CPI [KL02] provably fail in the absence of policy completeness. 3. On a positive note, agnostic policy learning is statistically tractable for Block MDPs with access to both of the above reset models. We establish this via a new algorithm that carefully constructs a policy emulator: a tabular MDP with a small state space that approximates the value functions of all policies $\pi \in \Pi$. These values are approximated without any explicit value function class.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Sample Complexity of Policy Learning with Mu-Resets

    cs.LG 2026-08 conditional novelty 8.0 of 10

    Under mu-resets and policy realizability, sample complexity is exp(Theta(H)) under all-policy concentrability and exp(Theta(sqrt H)) under pushforward concentrability.

  2. Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning

    cs.LG 2025-07 reject novelty 6.0 of 10

    Under variational gradient dominance, the paper derives state-space-independent sample complexity bounds for SDPO, CPI, DA-CPI, and PMD in agnostic policy learning.

Pith tools