REVIEW 5 cited by
X-IL: Exploring the Design Space of Imitation Learning Policies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Designing modern imitation learning (IL) policies requires making numerous decisions, including the selection of feature encoding, architecture, policy representation, and more. As the field rapidly advances, the range of available options continues to grow, creating a vast and largely unexplored design space for IL policies. In this work, we present X-IL, an accessible open-source framework designed to systematically explore this design space. The framework's modular design enables seamless swapping of policy components, such as backbones (e.g., Transformer, Mamba, xLSTM) and policy optimization techniques (e.g., Score-matching, Flow-matching). This flexibility facilitates comprehensive experimentation and has led to the discovery of novel policy configurations that outperform existing methods on recent robot learning benchmarks. Our experiments demonstrate not only significant performance gains but also provide valuable insights into the strengths and weaknesses of various design choices. This study serves as both a practical reference for practitioners and a foundation for guiding future research in imitation learning.
Forward citations
Cited by 5 Pith papers
-
RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning
Conditioning a pretrained ViT on per-pixel Plücker camera rays — via a gated-cross-attention class token and patch-level ray embeddings — makes imitation-learned manipulation policies substantially more robust to came...
-
Nautilus: From One Prompt to Plug-and-Play Robot Learning
NAUTILUS is a prompt-driven harness that automates plug-and-play adapters, typed contracts, and validation for policies, benchmarks, and robots in learning research.
-
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
A transformer-based compressor that turns old observations into fixed-size memory tokens improves non-Markovian imitation-learning robot policies, with large gains on memory-intensive simulated tasks.
-
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Dual-temporal VLM guidance injected into a JEPA predictor via multi-layer pyramid features improves hand-manipulation trajectory forecasting over VLM-only and JEPA-only baselines.
-
Expert Behavior Prior Reinforcement Learning
An online RL method that learns a generative behavior prior from the replay buffer via a Q-guided CVAE and uses adaptive gradient correction to combine Q-guidance with expert-action supervision.
Discussion (0). Continue with ORCID to comment.