Pith. sign in

REVIEW 2 cited by

Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.03016 v4 pith:VRWPRI2Y submitted 2019-10-07 cs.LG cs.AImath.OCstat.ML

classification cs.LGcs.AImath.OCstat.ML
keywords learningreinforcementrepresentationefficientgoodsamplevalue-basedapproximation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern deep learning methods provide effective means to learn good representations. However, is a good representation itself sufficient for sample efficient reinforcement learning? This question has largely been studied only with respect to (worst-case) approximation error, in the more classical approximate dynamic programming literature. With regards to the statistical viewpoint, this question is largely unexplored, and the extant body of literature mainly focuses on conditions which permit sample efficient reinforcement learning with little understanding of what are necessary conditions for efficient reinforcement learning. This work shows that, from the statistical viewpoint, the situation is far subtler than suggested by the more traditional approximation viewpoint, where the requirements on the representation that suffice for sample efficient RL are even more stringent. Our main results provide sharp thresholds for reinforcement learning methods, showing that there are hard limitations on what constitutes good function approximation (in terms of the dimensionality of the representation), where we focus on natural representational conditions relevant to value-based, model-based, and policy-based learning. These lower bounds highlight that having a good (value-based, model-based, or policy-based) representation in and of itself is insufficient for efficient reinforcement learning, unless the quality of this approximation passes certain hard thresholds. Furthermore, our lower bounds also imply exponential separations on the sample complexity between 1) value-based learning with perfect representation and value-based learning with a good-but-not-perfect representation, 2) value-based learning and policy-based learning, 3) policy-based learning and supervised learning and 4) reinforcement learning and imitation learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A tool constructs ε-optimal decision-tree controllers for MDPs, yielding trees orders of magnitude smaller than existing tools.

  2. Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Orthogonal bottlenecks constrain RL encoder features to low-dimensional subspaces while preserving expressivity and gradient dynamics under linear realizability when dimension exceeds the value function's intrinsic rank.

Pith tools