Pith. sign in

REVIEW 7 cited by

RvS: What is Essential for Offline RL via Supervised Learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.10751 v2 pith:7QRFGGDJ submitted 2021-12-20 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningsupervisedofflinechoosingessentialmethodsalgorithmicalone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that supervised learning alone, without temporal difference (TD) learning, can be remarkably effective for offline RL. When does this hold true, and which algorithmic components are necessary? Through extensive experiments, we boil supervised learning for offline RL down to its essential elements. In every environment suite we consider, simply maximizing likelihood with a two-layer feedforward MLP is competitive with state-of-the-art results of substantially more complex methods based on TD learning or sequence modeling with Transformers. Carefully choosing model capacity (e.g., via regularization or architecture) and choosing which information to condition on (e.g., goals or rewards) are critical for performance. These insights serve as a field guide for practitioners doing Reinforcement Learning via Supervised Learning (which we coin "RvS learning"). They also probe the limits of existing RvS methods, which are comparatively weak on random data, and suggest a number of open problems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nonreciprocal current induced by dissipation in time-reversal symmetric systems

    cond-mat.mes-hall 2026-04 unverdicted novelty 7.0 of 10

    Dissipation induces nonreciprocal current in time-reversal symmetric noncentrosymmetric systems via interband processes, inversely proportional to lifetime and linked to the shift vector.

  2. Freeform Preference Learning for Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.5 of 10

    Language-conditioned multi-axis human preferences yield denser rewards and steerable robot policies that outperform sparse and binary-preference baselines by 38 points on long-horizon manipulation.

  3. When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Supervised fine-tuning collapses LLM action diversity in board-game play beyond what the accuracy–diversity tradeoff requires; augmenting SFT data with all optimal actions per state partially prevents this.

  4. Generative Sequential Notification Optimization via Multi-Objective Decision Transformers

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A Decision Transformer with quantile-regression return prompts improved notification decisions at LinkedIn, boosting sessions by 0.72% over the deployed CQL baseline in a live A/B test.

  5. Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Doctor is a transformer-based offline RL method that samples candidate actions near a desired return and verifies them with a learned Q-function, improving target-return alignment.

  6. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  7. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools