Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:04:11.999325Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 7 inbound Pith citation observations for arXiv:2607.15330.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T00:04:11.999325Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:17:24.951028Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T14:52:36.313517Z
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 11bb3548-07d5-4973-9d2d-123330fe7fc0 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605d0126-85e7-43ba-99c7-d5c89b8e032a · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos 3: Omnimodal World Models for Physical AI
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fdce7e2-72af-4e1a-9e39-af8369972acb · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a52982e-ab77-4977-b06b-2cd22399abde · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a11f851-67fc-4cb7-9c2d-616f550a2ab6 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06ac054-a993-41e9-8ea8-a60638044087 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RT-1: Robotics Transformer for Real-World Control at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3955dccf-29e0-4ce0-be99-4b7659549687 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Language models are few-shot learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94428611-91ed-4eab-9408-eac62345b40c · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8021b7-cd0e-472e-bb1b-de885e8a0e0b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa6b5db-92b7-4149-8ded-902f997da321 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-3 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b1a5b9-9e36-45b9-900f-ca28e94475f6 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a91aa2f1-0c5a-43bc-9af8-de7a746a5b36 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400c7946-e052-4d43-89aa-478c300bdf5e · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Strategies for Efficient Embodied Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71c2dc8-c7fa-425f-ba20-905e005dd208 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b72ce5ab-7741-470f-975e-0e8edfac7f1c · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1e7ee89-44c8-4c46-a6e3-27c88216a173 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88f2b9c8-6bb2-4ea0-868c-c7995c5825d6 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d05d05-5854-4299-a7f5-2f22d7fcd3a8 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0517684-f747-44da-b962-a8232099456e · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct2: Action Reasoning Models for Real-world Deployment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35698623-2e34-4d0e-9955-6d5a6314550b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea g0.5 technical report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e06fcd5-8fcf-4926-83b2-6d7765aa8b22 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52eb2fbd-0f66-4138-98fb-9e64b0d22fa6 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Compute-Optimal Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7378915d-97d3-498f-b507-32b18b84e27d · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fa531a-0a35-45ce-b0ba-8f1298fe14ec · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi^{*}_{0.6}$: a VLA That Learns From Experience
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a937599b-d17a-48bd-9544-644636988650 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c109cc-2363-4441-9bc7-5cf88537232a · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea Open-World Dataset and G0 Dual-System VLA Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce86acd5-91d1-4d46-9ab4-d3612851450b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scaling Laws for Neural Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8145c4-eb6d-4480-b9c1-d94e833b9b27 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b2b861-086a-422c-af9b-769064d5ff03 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RLDX-1 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e68fd07-b6ec-4942-850c-7059404b9e2a · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories OpenVLA: An Open-Source Vision-Language-Action Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3d625c-3303-4af3-a624-b3297f6c60d1 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2464ccfb-0554-4edd-90d8-cd32a7430533 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a13f53-493d-4e77-a9b1-3d46e3270576 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning to act from actionless videos through dense correspondences
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72bd1631-89b2-4849-8c71-2a2b582af427 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct: Action Reasoning Models that can Reason in Space
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d545d019-ecb6-415e-94db-c9ccff4e2e03 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3804a4ed-233c-4b10-9cda-bec18a3197da · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Causal World Modeling for Robot Control
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8149403-1b63-49f5-ad55-3c20cf0b6c71 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gr-mg: Leveraging partially- annotated data via multi-modal goal-conditioned policy.IEEE Robotics and Automation Letters, 10(2):1912–1919, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e33eb75-6072-430c-a5a5-df4184f466bc · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8e04ef-f768-449a-8b04-c065b67a56a0 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c26cbcde-3016-4a5e-85ad-d0bc8ccd912b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9275040c-f375-47c1-bde7-ebf1aaacd914 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified Video Action Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc5e2df-7a63-44a3-9226-42faad756336 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88dc05cd-c5fc-4fa2-9d21-a6d5acb375d6 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27963403-1a86-48a4-9baf-38ec2c91e2b7 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f7c628-ed11-4a50-ba6c-8172d8fdea8f · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DeepSeek-V3 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d605c88b-0a22-41c2-9d22-5eec95c27e62 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf314b0-bdca-44ef-99b7-2eac4c1d3ad9 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt-1b: a diffusion foundation model for bimanual manipulation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0da9d55-b206-4473-bd3b-509c8dabff88 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310, 2026
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae21ebe-a60d-4db7-aeda-c7554c421de5 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9301cf6-8c57-4712-b8bb-df456ef42c4e · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fbd697c-610d-4dc3-95ea-c91b01898eb0 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b7906b-00a3-40b6-9ede-df5d6eabe242 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d1673e-9e66-4725-88a3-376012e51931 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR00T N1: An open foundation model for generalist humanoid robots
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1eb89b-d28b-4701-bb23-c184692f9990 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adfdcd9a-76f3-4940-aaa1-6db85fa6e05b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3821579b-4c34-4dcd-b8ec-3a79210a916b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable diffusion models with transformers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b906f0-94ed-410b-b654-8cbc2c256d97 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4dad51-fdcd-4991-89a6-6f2bf3f0efbf · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Coordinated humanoid manipulation with choice policies.arXiv preprint arXiv:2512.25072, 2025
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ec629f-4625-414a-9ba2-4c4780ea2f4a · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ce2c8f-bff1-4071-9e9f-8baf8f6cca98 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526ff66a-1512-4c10-b841-31128f6c2545 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini: A Family of Highly Capable Multimodal Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fbf405-d6d4-4024-bcbc-18c48348837f · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d0cacf-9905-4812-9370-8b6541f3ad95 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini Robotics: Bringing AI into the Physical World
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea493c74-4e01-414c-85ec-437ffb62a7fa · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-0: Embodied foundation models that scale with physical interaction.Generalist AI Blog,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7129a0-07f2-4ea4-b51d-240ba3d2c36e · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd335565-bdce-4b16-94cf-dafbcb829c71 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gene-26.5: Advancing robotic manipulation to human level.Genesis AI Blog, May 2026
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6408f5d5-3515-41ab-979e-bda2709f60eb · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Motubrain: An Advanced World Action Model for Robot Control
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a07c5b-9198-4593-b574-79a2746497dd · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Octo: An Open-Source Generalist Robot Policy
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dc7a77b-da8f-4386-9f13-545297db9c9b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3.5: Accelerating productivity with native multimodal agents, February 2026
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37642b6d-bc80-4668-8e71-19881e850fcc · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c93c5e-71fb-43eb-a527-f39338861437 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories LLaMA: Open and Efficient Foundation Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb546a4-854a-4567-80b1-5f9eb81dba9b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World2Act: Latent Action Post-Training from World Model Dynamics
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10fc9c9-56a0-4063-b656-35fbd7d67500 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgedata v2: A dataset for robot learning at scale
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbb7262-ef3d-492a-be20-e2d399360cd3 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff7846f-3914-4da3-abbe-e85fa1fba95c · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories A Pragmatic VLA Foundation Model
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8847c4a0-c647-44d2-ba2c-5483b46002be · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571a8b58-02ec-4e07-8ad5-b643aa9fdfb4 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MemoryWAM: Efficient World Action Modeling with Persistent Memory
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeee5494-a31e-4895-9005-01475331ad58 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gigaworld-policy: An efficient action-centered world-action model.arXiv preprint arXiv:2603.17240, 2026
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7419d939-de71-4609-96b3-1baa5e0f54ab · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Starvla-α: Reducing complexity in vision-language-action systems
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88207e1d-9378-44e1-b8a2-f40c32a463f2 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World Action Models are Zero-shot Policies
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628babad-8bf5-49a2-b9e3-d20096c08eeb · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wall-OSS-0.5 Technical Report
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181cbaaa-febe-41b0-b66a-e702b18d0f09 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395ff754-9e63-40d0-a1d6-14444ec8c58b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e0fc5f-94c4-43fc-ae96-5be679b4161b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213f53be-0816-4ac8-8652-76f621b0e5f1 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Native Video-Action Pretraining for Generalizable Robot Control
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1a8114-70f9-4b9a-8a6b-382661e04863 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a64214-cc8c-46c0-8ab8-8c27ff66a29d · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8662e36b-e228-487d-864a-51bc0c18689b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd2b7e2-65a0-4fe7-9342-6116df8470fb · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fastumi: A scalable and hardware-independent universal manipulation interface with dataset
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7574aabc-55ee-4055-8bc9-af23deb11fc9 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories TesserAct: Learning 4D Embodied World Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe5f4eb-12bf-4b5c-bf0c-183af1e848f3 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17054e2-dfe6-4555-aa1a-2ee1aaba470b · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a015f2-28a9-445d-9777-5fd99f0a73e9 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404, 2026
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720928f3-040e-4442-b1ad-e10b739fc6e2 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDreamer: Learning Compositional World Models for Robot Imagination
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9c9bc2-b575-4938-a372-d6cf0c9f1af9 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb94489-b3e8-461a-985b-bc8690f8338e · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rt-2: Vision-language-action models transfer web knowledge to robotic control
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d79aef79-964d-4da3-96ab-35e32f332ef8 · outbound
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0aac408-4465-4ad5-a4a1-9ef4b06df91e · inbound
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4807410a-d1a0-4b83-b8e9-f36242eb9ed9 · inbound
$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93313b23-be72-440b-899f-91968f484b7d · inbound
OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a539bc-f8cc-403b-84fe-ecb4a2e103a7 · inbound
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35032d6b-ae97-4ea2-bd59-3487dbbc969a · inbound
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3a3086-a292-43c5-88be-b9a14fafaf44 · inbound
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11a420e-44e5-456c-94a5-f8161ea423b4 · inbound
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.