Pith. sign in

REVIEW 10 cited by

Foundations of Reinforcement Learning and Interactive Decision Making

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.16730 v1 pith:5END7BQI submitted 2023-12-27 cs.LG math.OCmath.STstat.MLstat.TH

Foundations of Reinforcement Learning and Interactive Decision Making

classification cs.LG math.OCmath.STstat.MLstat.TH
keywords learningdecisionmakingreinforcementbanditsfoundationsinteractiveaddressing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

These lecture notes give a statistical perspective on the foundations of reinforcement learning and interactive decision making. We present a unifying framework for addressing the exploration-exploitation dilemma using frequentist and Bayesian approaches, with connections and parallels between supervised learning/estimation and decision making as an overarching theme. Special attention is paid to function approximation and flexible model classes such as neural networks. Topics covered include multi-armed and contextual bandits, structured bandits, and reinforcement learning with high-dimensional feedback.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles

    cs.LG 2026-07 conditional novelty 7.0

    An OCO algorithm with only O(√T) static regret, pluggable as a preconditioner selector, recovers the classical O(1/√T) stationarity rate on smooth stochastic nonconvex problems and the O(T^{-2/7}) rate on nonsmooth ones.

  2. When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

    cs.LG 2026-06 unverdicted novelty 7.0

    Online IL overcomes an information-theoretic bottleneck that offline IL faces in non-realizable settings even at horizon 1, under a new structural characterization of reward-relative misspecification.

  3. Efficient Exploration for Iterative Nash Preference Optimization

    cs.LG 2026-05 unverdicted novelty 7.0

    An explicitly exploratory iterative NLHF method achieves O(sqrt(T)) regret for Nash equilibria under general preference models, removing the exponential KL dependence that plagues standard iterative approaches.

  4. A Geometric Approach to Constrained Online Learning

    cs.LG 2026-05 conditional novelty 7.0

    A nested-projection gradient algorithm attains O(log T) regret with O(log T) cumulative constraint violation for strongly convex losses, and O(√T) for both with convex losses; the body's proof is coherent, though the ...

  5. A Geometric Approach to Constrained Online Learning

    cs.LG 2026-05 unverdicted novelty 7.0

    NP-OGD attains O(log T) regret and O(log T) CCV for strongly convex losses, and O(√T) for both under convex losses, with complementary geometric lower bounds.

  6. Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability

    cs.LG 2026-05 unverdicted novelty 7.0

    The paper establishes the first tilde O(epsilon^{-1}) upper bounds and matching lower bounds for forward-KL-regularized offline contextual bandits under single-policy concentrability in both tabular and general functi...

  7. Constrained Contextual Bandits with Adversarial Contexts

    cs.LG 2026-05 unverdicted novelty 7.0

    A modular reduction from budget-constrained contextual bandits with adversarial contexts to unconstrained bandits via surrogate rewards, yielding improved guarantees and an efficient algorithm based on SquareCB.

  8. A Jointly Efficient and Optimal Algorithm for Heteroskedastic Generalized Linear Bandits with Adversarial Corruptions

    cs.LG 2026-02 conditional novelty 7.0

    A per-round O(1) algorithm for generalized linear bandits achieves near-optimal regret with time-varying dispersion and adversarial corruptions, up to a κ factor.

  9. When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

    stat.ML 2026-06 unverdicted novelty 6.0

    Proposes OPAC for trajectory-level offline RL achieving 𝓣O(H^{2}√(C_sa(π*)/n)) bounds with matching lower bound, plus conditions for tractability in generalized nonlinear outcome settings.

  10. A Geometric Approach to Constrained Online Learning

    cs.LG 2026-05 unverdicted novelty 6.0

    A projection-based algorithm for COCO achieves O(log T) regret and O(log T) CCV for strongly convex losses and O(sqrt(T)) for convex losses by leveraging self-contracted curves.