Pith. sign in

REVIEW 7 cited by

MaIL: Improving Imitation Learning with Mamba

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08234 v2 pith:55FMH7CC submitted 2024-06-12 cs.LG cs.RO

classification cs.LGcs.RO
keywords maillearningmambadataimitationarchitectureavailableenhances
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work presents Mamba Imitation Learning (MaIL), a novel imitation learning (IL) architecture that provides an alternative to state-of-the-art (SoTA) Transformer-based policies. MaIL leverages Mamba, a state-space model designed to selectively focus on key features of the data. While Transformers are highly effective in data-rich environments due to their dense attention mechanisms, they can struggle with smaller datasets, often leading to overfitting or suboptimal representation learning. In contrast, Mamba's architecture enhances representation learning efficiency by focusing on key features and reducing model complexity. This approach mitigates overfitting and enhances generalization, even when working with limited data. Extensive evaluations on the LIBERO benchmark demonstrate that MaIL consistently outperforms Transformers on all LIBERO tasks with limited data and matches their performance when the full dataset is available. Additionally, MaIL's effectiveness is validated through its superior performance in three real robot experiments. Our code is available at https://github.com/ALRhub/MaIL.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DSSP: Diffusion State Space Policy with Full-History Encoding

    cs.RO 2026-05 conditional novelty 7.0 of 10

    DSSP is a history-conditioned diffusion state space policy that uses SSMs to encode full observation streams with an auxiliary dynamics objective and hierarchical fusion, achieving SOTA results with reduced model size...

  2. SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    SSI-Policy learns a robot-agnostic RGB-only scene interface from video to improve vision-language manipulation policies by 15% on LIBERO with only 10 demos per task.

  3. SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    SSI-Policy uses an RGB-only Structured Scene Interface to improve LIBERO benchmark performance by nearly 15% with only 10 demonstrations per task compared to prior methods.

  4. RoboSSM: Scalable In-context Imitation Learning via State-Space Models

    cs.RO 2025-09 conditional novelty 6.0 of 10

    RoboSSM shows that a state-space model backbone can extend in-context imitation learning to prompts much longer than those seen in training, where a Transformer-based baseline degrades.

  5. MUSE: Multimodal Uncertainty Quantification of State Estimation

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    MUSE applies Mamba sequential modeling to produce real-time uncertainty estimates for visual-inertial state estimation from asynchronous multimodal sensors.

  6. OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A generalist agent with shared shallow layers and task-separated deep experts outperforms single-domain GUI and embodied agents on AndroidControl, GUI-Odyssey, and LIBERO benchmarks.

  7. A Survey of Mamba

    cs.LG 2024-08 unverdicted novelty 2.0 of 10

    The paper consolidates existing research on Mamba models, their architecture variants, adaptations to different data modalities, and applications across domains.

Pith tools