Filtered behavior cloning, an MLP baseline trained only on high-return trajectories, matches or outperforms Decision Transformer on the sparse-reward Robomimic and sparsified D4RL benchmarks tested.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
Filtered behavior cloning, an MLP baseline trained only on high-return trajectories, matches or outperforms Decision Transformer on the sparse-reward Robomimic and sparsified D4RL benchmarks tested.