REVIEW 6 cited by
Offline Reinforcement Learning as One Big Sequence Modeling Problem
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement learning (RL) is typically concerned with estimating stationary policies or single-step models, leveraging the Markov property to factorize problems in time. However, we can also view RL as a generic sequence modeling problem, with the goal being to produce a sequence of actions that leads to a sequence of high rewards. Viewed in this way, it is tempting to consider whether high-capacity sequence prediction models that work well in other domains, such as natural-language processing, can also provide effective solutions to the RL problem. To this end, we explore how RL can be tackled with the tools of sequence modeling, using a Transformer architecture to model distributions over trajectories and repurposing beam search as a planning algorithm. Framing RL as sequence modeling problem simplifies a range of design decisions, allowing us to dispense with many of the components common in offline RL algorithms. We demonstrate the flexibility of this approach across long-horizon dynamics prediction, imitation learning, goal-conditioned RL, and offline RL. Further, we show that this approach can be combined with existing model-free algorithms to yield a state-of-the-art planner in sparse-reward, long-horizon tasks.
Forward citations
Cited by 6 Pith papers
-
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.
-
Agent-Centric Animal Pose Forecasting
Agent-centric transformers trained through a composable library reproduce several marginal statistics of courting fly behavior, but discriminators still separate simulated from real flies.
-
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text
A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.
-
Learning to Ask: Decision Transformers for Adaptive Quantitative Group Testing
An adaptive Decision Transformer policy, trained on heuristic-generated trajectories, is claimed to beat the non-adaptive query-count bound for quantitative group testing.
-
Energy-Efficient Deep Reinforcement Learning with Spiking Transformers
A spiking Transformer trained on A* maze demonstrations reaches 99.64% action accuracy on a custom 21x21 maze benchmark, but the energy-efficiency and baseline-comparison claims are not backed by proper experiments.
-
AdaCred: Adaptive Causal Decision Transformers with Feature Crediting
AdaCred trains decision transformers with learned GumbelSigmoid token masks and an efficiency loss, claiming improved offline RL and imitation learning with shorter, pruned sequences.
Discussion (0). Continue with ORCID to comment.