A memory-guided LLM harness that calls a frozen VLA only for contact-rich phases lifts success to 82.4% on LIBERO-Pro, 55.4% on RoboCasa365, and 58.4% on RoboTwin C2R with no policy finetuning.
Atomvla: Scalable post-training for robotic manipulation via predictive latent world models
5 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.RO 5years
2026 5representative citing papers
AIM predicts aligned spatial value maps inside a shared video-generation transformer to produce reliable robot actions, reaching 94% success on RoboTwin 2.0 with larger gains on long-horizon and contact-rich tasks.
HiMem-WAM integrates hierarchical latent actions and boundary-aware memory gates into world action models to enhance robustness and performance on memory-dependent long-horizon robotic tasks.
Survey organizing world models for robotic manipulation into representation families, a functional taxonomy, and infrastructure roles across pretraining, post-training, and inference, while reviewing 34 datasets and evaluation protocols.
citing papers explorer
-
Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents
A memory-guided LLM harness that calls a frozen VLA only for contact-rich phases lifts success to 82.4% on LIBERO-Pro, 55.4% on RoboCasa365, and 58.4% on RoboTwin C2R with no policy finetuning.
-
AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
AIM predicts aligned spatial value maps inside a shared video-generation transformer to produce reliable robot actions, reaching 94% success on RoboTwin 2.0 with larger gains on long-horizon and contact-rich tasks.
-
HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation
HiMem-WAM integrates hierarchical latent actions and boundary-aware memory gates into world action models to enhance robustness and performance on memory-dependent long-horizon robotic tasks.
-
World Models for Robotic Manipulation: A Survey
Survey organizing world models for robotic manipulation into representation families, a functional taxonomy, and infrastructure roles across pretraining, post-training, and inference, while reviewing 34 datasets and evaluation protocols.
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?