Pith. sign in

REVIEW 6 cited by

Offline Reinforcement Learning as One Big Sequence Modeling Problem

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.02039 v4 pith:6DXSKQK6 submitted 2021-06-03 cs.LG cs.AI

classification cs.LGcs.AI
keywords sequencemodelingproblemlearningofflinealgorithmsapproachlong-horizon
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement learning (RL) is typically concerned with estimating stationary policies or single-step models, leveraging the Markov property to factorize problems in time. However, we can also view RL as a generic sequence modeling problem, with the goal being to produce a sequence of actions that leads to a sequence of high rewards. Viewed in this way, it is tempting to consider whether high-capacity sequence prediction models that work well in other domains, such as natural-language processing, can also provide effective solutions to the RL problem. To this end, we explore how RL can be tackled with the tools of sequence modeling, using a Transformer architecture to model distributions over trajectories and repurposing beam search as a planning algorithm. Framing RL as sequence modeling problem simplifies a range of design decisions, allowing us to dispense with many of the components common in offline RL algorithms. We demonstrate the flexibility of this approach across long-horizon dynamics prediction, imitation learning, goal-conditioned RL, and offline RL. Further, we show that this approach can be combined with existing model-free algorithms to yield a state-of-the-art planner in sparse-reward, long-horizon tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 41 citations worldwide. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. Agent-Centric Animal Pose Forecasting

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Agent-centric transformers trained through a composable library reproduce several marginal statistics of courting fly behavior, but discriminators still separate simulated from real flies.

  3. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.

  4. Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing

    cs.IT 2025-09 reject novelty 5.0 of 10

    An adaptive Decision Transformer policy, trained on heuristic-generated trajectories, is claimed to beat the non-adaptive query-count bound for quantitative group testing.

  5. Energy-Efficient Deep Reinforcement Learning with Spiking Transformers

    cs.LG 2025-05 reject novelty 5.0 of 10

    A spiking Transformer trained on A* maze demonstrations reaches 99.64% action accuracy on a custom 21x21 maze benchmark, but the energy-efficiency and baseline-comparison claims are not backed by proper experiments.

  6. AdaCred: Adaptive Causal Decision Transformers with Feature Crediting

    cs.LG 2024-12 reject novelty 4.0 of 10

    AdaCred trains decision transformers with learned GumbelSigmoid token masks and an efficiency loss, claiming improved offline RL and imitation learning with shorter, pruned sequences.

Pith tools