{"work":{"id":"d0b2d257-524d-4d67-9daf-5fb43e5e977a","openalex_id":null,"doi":null,"arxiv_id":"2512.13030","raw_key":null,"title":"Motus: A Unified Latent Action World Model","authors":null,"authors_text":"Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu","year":2025,"venue":"cs.CV","abstract":"While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level \"delta action\" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.","external_url":"https://arxiv.org/abs/2512.13030","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T04:16:48.690751+00:00","pith_arxiv_id":"2512.13030","created_at":"2026-05-10T01:10:09.444541+00:00","updated_at":"2026-07-10T04:16:48.690751+00:00","title_quality_ok":true,"display_title":"Motus: A Unified Latent Action World Model","render_title":"Motus: A Unified Latent Action World Model"},"hub":{"state":{"work_id":"d0b2d257-524d-4d67-9daf-5fb43e5e977a","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":88,"external_cited_by_count":null,"distinct_field_count":4,"first_pith_cited_at":"2026-01-29T17:07:43+00:00","last_pith_cited_at":"2026-07-09T16:15:43+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T08:19:20.661048+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":20},{"context_role":"baseline","n":3},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":20},{"context_polarity":"baseline","n":3},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}