REVIEW 4 cited by
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Decision Transformer, a promising approach that applies Transformer architectures to reinforcement learning, relies on causal self-attention to model sequences of states, actions, and rewards. While this method has shown competitive results, this paper investigates the integration of the Mamba framework, known for its advanced capabilities in efficient and effective sequence modeling, into the Decision Transformer architecture, focusing on the potential performance enhancements in sequential decision-making tasks. Our study systematically evaluates this integration by conducting a series of experiments across various decision-making environments, comparing the modified Decision Transformer, Decision Mamba, with its traditional counterpart. This work contributes to the advancement of sequential decision-making models, suggesting that the architecture and training methodology of neural networks can significantly impact their performance in complex tasks, and highlighting the potential of Mamba as a valuable tool for improving the efficacy of Transformer-based models in reinforcement learning scenarios.
Forward citations
Cited by 4 Pith papers
-
Success in Humanoid Reinforcement Learning under Partial Observation
This paper reports the first stable partial-observability training on Humanoid-v4, using a parallel history encoder that matches full-state TD3 performance in most tested state-removal settings.
-
Meta-Black-Box-Optimization through Offline Q-function Learning
Q-Mamba trains a Mamba-based Q-function controller for evolutionary algorithm configuration on an offline dataset and matches or slightly exceeds online baselines on BBOB benchmarks.
-
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
GCReinSL adds Q-conditioned maximization to supervised offline RL, using normalizing flows to estimate goal-reaching probabilities and expectile regression to condition actions on the best in-distribution value, impro...
-
Decision Transformer vs. Decision Mamba: Analysing the Complexity of Sequential Decision Making in Atari Games
Across 12 Atari games, Decision Transformer tends to outperform Decision Mamba in games with larger action spaces and more complex visuals, while Decision Mamba wins in simpler games, though the statistical support is weak.
Discussion (0). Continue with ORCID to comment.