Pith. sign in

Paper Citation Record · LEDGER

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

As of 16 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 7 inbound Pith citation observations for arXiv:2607.15330.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15330 v2

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:04:11.999325Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:17:24.951028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:36.313517Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11bb3548-07d5-4973-9d2d-123330fe7fc0 · outbound

This paper cites GPT-4 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.117791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.117791Z digest=sha256:a723404ca2c20d8b1559e528e0ba4d343657c997872e0d3867978ead7b23a369

Observation 605d0126-85e7-43ba-99c7-d5c89b8e032a · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.225496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.225496Z digest=sha256:1129e03b08572e7bfc516da14ae1127ba58a42fbd92d418f026a6d0877adb742

Observation 7fdce7e2-72af-4e1a-9e39-af8369972acb · outbound

This paper cites Qwen3-VL Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.355811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.355811Z digest=sha256:118638258ae8493d8791526558f301f1703a87b20f105cd9eb184342ce511522

Observation 9a52982e-ab77-4977-b06b-2cd22399abde · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.504446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.504446Z digest=sha256:a335aaef1bbc5cc0bbd775b0340db7209371c80bd1936d16f8a3d1ea49b0a299

Observation 9a11f851-67fc-4cb7-9c2d-616f550a2ab6 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.709806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.709806Z digest=sha256:11db8671ff1dfd442ef7893ad0d98d0602b45a53226174bbfac947164ad9c746

Observation a06ac054-a993-41e9-8ea8-a60638044087 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.879682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.879682Z digest=sha256:1d4c92680abc82d81e0850df8dc39a30d12045d61261d52618dd1e703b112f1b

Observation 3955dccf-29e0-4ce0-be99-4b7659549687 · outbound

This paper cites Language models are few-shot learners.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.955958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.955958Z digest=sha256:bdae1888518b54f92f7bb2758bf039e6eaf738c8fcb0735b0d734b6b7dcb7175

Observation 94428611-91ed-4eab-9408-eac62345b40c · outbound

This paper cites Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution.arXiv preprint arXiv:2602.12684, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.063560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.063560Z digest=sha256:3a753ead63874cc524fef76c106fc5b34bcf869fefb08e80d509d96a54530cce

Observation de8021b7-cd0e-472e-bb1b-de885e8a0e0b · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.143342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.143342Z digest=sha256:3a6bac00e6b06baade126920d18ad13a72f6f6e44fbece10f2ca9348b13d7cea

Observation 5aa6b5db-92b7-4149-8ded-902f997da321 · outbound

This paper cites GR-3 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR-3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.224041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.224041Z digest=sha256:e275aaa385585f62a6bb5dbb8968b58dce7264adfa7081e36c109d7a5d7a8ac5

Observation 01b1a5b9-9e36-45b9-900f-ca28e94475f6 · outbound

This paper cites ABot-M0.5: Unified Mobility-and-Manipulation World Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.292824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.292824Z digest=sha256:51a64e6d5e4094b645b17fafcdb5bdb0c965e9afd9e67e464a8907c69013b8c3

Observation a91aa2f1-0c5a-43bc-9af8-de7a746a5b36 · outbound

This paper cites RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.390410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.390410Z digest=sha256:0a572ee7e713a86ab74031bd33a4287d89c11c88403ec558a2c0285c15a12282

Observation 400c7946-e052-4d43-89aa-478c300bdf5e · outbound

This paper cites Training Strategies for Efficient Embodied Reasoning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Strategies for Efficient Embodied Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.468016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.468016Z digest=sha256:1f0ca95363b599fcf6c7422cb7ea9d923d42d92eca9c56c383de929e2d59db46

Observation a71c2dc8-c7fa-425f-ba20-905e005dd208 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.537030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.537030Z digest=sha256:cf6276754af1a207d56fc67f023c46a3a74be6e76994f195f13aa0a212a12848

Observation b72ce5ab-7741-470f-975e-0e8edfac7f1c · outbound

This paper cites Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.633844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.633844Z digest=sha256:b94b3bd9ff4b91adf39b43262d8caac2c8aef2882e970bcb7a4e814f33210560

Observation d1e7ee89-44c8-4c46-a6e3-27c88216a173 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2024.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:00.778343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:00.778343Z digest=sha256:81453eb27738670bdafffeba596adcca4ce34fdf54cacda3d48d2901886d9081

Observation 88f2b9c8-6bb2-4ea0-868c-c7995c5825d6 · outbound

This paper cites Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.022781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.022781Z digest=sha256:4d75b536ac55ccea7b819dfcb1a0192517ddf6a9123ab0879ccde7a948a783f5

Observation 27d05d05-5854-4299-a7f5-2f22d7fcd3a8 · outbound

This paper cites Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning universal policies via text-guided video generation.Advances in neural information processing systems, 36:9156–9172, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.112802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.112802Z digest=sha256:6f10dd8db75bea727c9a169673347a42d2506eabd7e4c97f4b373cdc12e37ca9

Observation b0517684-f747-44da-b962-a8232099456e · outbound

This paper cites MolmoAct2: Action Reasoning Models for Real-world Deployment.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct2: Action Reasoning Models for Real-world Deployment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.196872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.196872Z digest=sha256:8f78b9e07ec8654c1ce48ec971e1ff2db4c1ab3777f8dde4ce5509181cb952f1

Observation 35698623-2e34-4d0e-9955-6d5a6314550b · outbound

This paper cites Galaxea g0.5 technical report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea g0.5 technical report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.316861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.316861Z digest=sha256:b9ece3772334a6351761f1505f0768a6991b719be8be49effbaa2319782b2801

Observation 3e06fcd5-8fcf-4926-83b2-6d7765aa8b22 · outbound

This paper cites Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.407672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.407672Z digest=sha256:13ed31bc7d81732c2a0ff9cb20da78e4b7418bcefaf38e247712d4e4832e5b2d

Observation 52eb2fbd-0f66-4138-98fb-9e64b0d22fa6 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Training Compute-Optimal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.478584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.478584Z digest=sha256:a758245b9cdc555d946b6cdbc959dabeea6a9085e2b4d4b5ab6ec7f61bc26c27

Observation 7378915d-97d3-498f-b507-32b18b84e27d · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.537382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.537382Z digest=sha256:218ad130babb0cd1067b0053da7079d4b356ea3f4a96f1a147cae0d2f2eb0c04

Observation 58fa531a-0a35-45ce-b0ba-8f1298fe14ec · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.650486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.650486Z digest=sha256:49e3aef867db6eb8feed662714ba1ce6d5e2111847e19e8e75acac809d1d70b1

Observation a937599b-d17a-48bd-9544-644636988650 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.761066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.761066Z digest=sha256:9addf59ac345ea615d22c0889568a68429f9b02a243a1c79a0769a05590620b3

Observation 32c109cc-2363-4441-9bc7-5cf88537232a · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.858008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.858008Z digest=sha256:4a5c73d88b8a315287b2957eecbd09fb8df22b8d506a8e4804d35447cbf58262

Observation ce86acd5-91d1-4d46-9ab4-d3612851450b · outbound

This paper cites Scaling Laws for Neural Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scaling Laws for Neural Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:01.934636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:01.934636Z digest=sha256:37a45c90f0936ef8049de7fa59b1ae82204e4efd688b08f30632b9c47c9e9614

Observation ec8145c4-eb6d-4480-b9c1-d94e833b9b27 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.006714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.006714Z digest=sha256:fbb1685326603380f5c4aff9bb57b524e26a5fe64c21867b48a9340812a8c7e0

Observation 81b2b861-086a-422c-af9b-769064d5ff03 · outbound

This paper cites RLDX-1 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RLDX-1 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.063691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.063691Z digest=sha256:68fc60a61b8f35eaa9fa66fedaf7e94ce426b266409d4b42da427e32dbad6d74

Observation 4e68fd07-b6ec-4942-850c-7059404b9e2a · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.186263Z digest=sha256:063ff8fd951523fccd268e6880531c714c7377e2e4063729ec033cf3bb105a60

Observation af3d625c-3303-4af3-a624-b3297f6c60d1 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.287670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.287670Z digest=sha256:4c40050c5e93bf09290d25ed5fc1e5876c0ac6a7ca16a2c68725fa117bfcc9f0

Observation 2464ccfb-0554-4edd-90d8-cd32a7430533 · outbound

This paper cites Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.398602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.398602Z digest=sha256:ac2f6c9ec8ef7547027b88c2f1a1517565de7dec568f0815208602c161cf10a3

Observation 85a13f53-493d-4e77-a9b1-3d46e3270576 · outbound

This paper cites Learning to act from actionless videos through dense correspondences.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Learning to act from actionless videos through dense correspondences

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.581358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.581358Z digest=sha256:bff793430d8b3b327e2b6db6c7ac891430e9bc33b0c5a3260d6c008fb006dbd4

Observation 72bd1631-89b2-4849-8c71-2a2b582af427 · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MolmoAct: Action Reasoning Models that can Reason in Space

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.790902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.790902Z digest=sha256:567333c90c1f9ff42879df6f36d465df3adcf0fc439822c8dab534d79c64bef7

Observation d545d019-ecb6-415e-94db-c9ccff4e2e03 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Spatial forcing: Implicit spatial representation alignment for vision-language-action model.arXiv preprint arXiv:2510.12276, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:02.943489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:02.943489Z digest=sha256:eaaa2e5a8961d774df460a5b678952fde16859974c65e0fe1516262d1df34616

Observation 3804a4ed-233c-4b10-9cda-bec18a3197da · outbound

This paper cites Causal World Modeling for Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Causal World Modeling for Robot Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.115676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.115676Z digest=sha256:19678c3de256d848c70dc2a5e02543697cc92bf1a9398dec4f3d413c01bad5a9

Observation c8149403-1b63-49f5-ad55-3c20cf0b6c71 · outbound

This paper cites Gr-mg: Leveraging partially- annotated data via multi-modal goal-conditioned policy.IEEE Robotics and Automation Letters, 10(2):1912–1919, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gr-mg: Leveraging partially- annotated data via multi-modal goal-conditioned policy.IEEE Robotics and Automation Letters, 10(2):1912–1919, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.322392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.322392Z digest=sha256:f9ee9f5e12aeb3983e7b179488dddfb4065c1cae6ab3fb4797b4e95b1b61a018

Observation 9e33eb75-6072-430c-a5a5-df4184f466bc · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.480239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.480239Z digest=sha256:31d68d9bb22d568446eaa5deecd11ef8f187ee3d1a63c6494deaafa918c5c58c

Observation da8e04ef-f768-449a-8b04-c065b67a56a0 · outbound

This paper cites SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.686284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.686284Z digest=sha256:a45e34a4b2903ef201df330f1ff77794faa52bccd81f32425ba9f6f07fe171fb

Observation c26cbcde-3016-4a5e-85ad-d0bc8ccd912b · outbound

This paper cites Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable vision-language-action model pretraining for robotic manipulation with real-life human activity videos.arXiv preprint arXiv:2510.21571, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.836866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.836866Z digest=sha256:66ca60b7457fd49058db9cbe2396b67c1ca820940d4297d3a374f01896700960

Observation 9275040c-f375-47c1-bde7-ebf1aaacd914 · outbound

This paper cites Unified Video Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified Video Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:03.991822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:03.991822Z digest=sha256:96cd799879143bc74532d1d3677c526735e664b0ac9da61d125bbd14dcb0a09d

Observation 9dc5e2df-7a63-44a3-9226-42faad756336 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.138071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.138071Z digest=sha256:24135df117a7dd9247b2e19f57c77e7a0e3c4f0dca1fcbdf578170ac1dbb88cf

Observation 88dc05cd-c5fc-4fa2-9d21-a6d5acb375d6 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.293371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.293371Z digest=sha256:21ef329e0e021e0ed13eb6f0502b20b56c0d345a1131bc318de95b0461084dbe

Observation 27963403-1a86-48a4-9baf-38ec2c91e2b7 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.419047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.419047Z digest=sha256:2a607ddf7f70b671788e3c0c56df481cfd0af050b81a492a9c220efdc1c28d9f

Observation 65f7c628-ed11-4a50-ba6c-8172d8fdea8f · outbound

This paper cites DeepSeek-V3 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories DeepSeek-V3 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.489716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.489716Z digest=sha256:14f0852e7451b8bb2686bab6192be8b6f83bbb411969b43354cac77acf45a506

Observation d605c88b-0a22-41c2-9d22-5eec95c27e62 · outbound

This paper cites ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.559029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.559029Z digest=sha256:2caad478765e3f88dd0e5006e0b69b9f2b1b83fedf9bbadad661ba852acabd09

Observation ddf314b0-bdca-44ef-99b7-2eac4c1d3ad9 · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.695444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.695444Z digest=sha256:40a44b87c39eef9de0708d1803c6f3661a9e4983c7514e4a33f2f23bf745779e

Observation d0da9d55-b206-4473-bd3b-509c8dabff88 · outbound

This paper cites Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rdt2: Exploring the scaling limit of umi data towards zero-shot cross-embodiment generalization.arXiv preprint arXiv:2602.03310, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.891865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.891865Z digest=sha256:6309badede549aaf7dd5b480c50263b65395b759faad33fbb8f16f9a705960a2

Observation 2ae21ebe-a60d-4db7-aeda-c7554c421de5 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.963758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.963758Z digest=sha256:8e9a87528fc007a638bf7d36b06a983899cd59dfa23572274dd3a7c115dc69b1

Observation 5efc2714-0f33-4327-81ef-d8033792d4ec · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.119524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.119524Z digest=sha256:688ac3cddd05f96c818e971a7801bd64adee2c5b5e3d52eda472dfc569c6b795

Observation d9301cf6-8c57-4712-b8bb-df456ef42c4e · outbound

This paper cites Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.287415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.287415Z digest=sha256:e0fe83e52091db537eaba55e9371270e41cfce49ce3fee408e28a2c8439a167f

Observation 4fbd697c-610d-4dc3-95ea-c91b01898eb0 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.391932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.391932Z digest=sha256:fbf3ec9b0a850d749e6655935c5525de7b4a7788bbab04c70ea7b76e2239e241

Observation b1b7906b-00a3-40b6-9ede-df5d6eabe242 · outbound

This paper cites Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.541501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.541501Z digest=sha256:df51f063d084e7f3b58ec986d2b4769957eb936e8f2153badf495fc332a60352

Observation 19d1673e-9e66-4725-88a3-376012e51931 · outbound

This paper cites GR00T N1: An open foundation model for generalist humanoid robots.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories GR00T N1: An open foundation model for generalist humanoid robots

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.652775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.652775Z digest=sha256:b515f5057ffdea71e7f7ce33dd3ae7db31e616a697c992662eb8e29674d7ce1c

Observation 4f1eb89b-d28b-4701-bb23-c184692f9990 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.791080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.791080Z digest=sha256:58c265fa668c1d69bddf838c3cf831d023d32f7bb6c06a3127da9eb4386a1242

Observation adfdcd9a-76f3-4940-aaa1-6db85fa6e05b · outbound

This paper cites mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:05.928505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:05.928505Z digest=sha256:28150f820c5128cc193c0fd16148198c99f3c8a59736d4129e67d108594c82b3

Observation 3821579b-4c34-4dcd-b8ec-3a79210a916b · outbound

This paper cites Scalable diffusion models with transformers.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Scalable diffusion models with transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.037779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.037779Z digest=sha256:bce2cf46f768c76fe33593e05a0d0465c6246ebb8e3aa95c64eff54dadc19f94

Observation b5b906f0-94ed-410b-b654-8cbc2c256d97 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.230214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.230214Z digest=sha256:85ce02c9512278c3aef6479091c34fdadc6c36b2a3e36d74fd7a98c955fbbf8b

Observation be4dad51-fdcd-4991-89a6-6f2bf3f0efbf · outbound

This paper cites Coordinated humanoid manipulation with choice policies.arXiv preprint arXiv:2512.25072, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Coordinated humanoid manipulation with choice policies.arXiv preprint arXiv:2512.25072, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.351663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.351663Z digest=sha256:6c7cbaf9d00d4e8e0fa027bf14162d6523a58bcf5db62682db0847691808fb76

Observation c7ec629f-4625-414a-9ba2-4c4780ea2f4a · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.479669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.479669Z digest=sha256:40cd12c79a3c0be4f742bef0efb175a6fa9d67c8e3f5da58245d37d3df2a08ad

Observation c3ce2c8f-bff1-4071-9e9f-8baf8f6cca98 · outbound

This paper cites Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.613967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.613967Z digest=sha256:0a363a071ecbe31d69420c05a6ccd5595499402403011a4e2b615d44f0bc9618

Observation 526ff66a-1512-4c10-b841-31128f6c2545 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini: A Family of Highly Capable Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.798973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.798973Z digest=sha256:ed3a5365bc9aad08d4e36ff9acfe5e652513807b78c9deba4df6e077643e167e

Observation 73fbf405-d6d4-4024-bcbc-18c48348837f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:06.931913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:06.931913Z digest=sha256:059ed24b30981eb3927792f681c50acf22a817570cb1977415ebc569de56c86a

Observation 54d0cacf-9905-4812-9370-8b6541f3ad95 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gemini Robotics: Bringing AI into the Physical World

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.086788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.086788Z digest=sha256:401bd8fb1f192b798b2f51da0f5a736cbc1fafe45ed7bcda35e8877fd662766c

Observation ea493c74-4e01-414c-85ec-437ffb62a7fa · outbound

This paper cites Gen-0: Embodied foundation models that scale with physical interaction.Generalist AI Blog,.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-0: Embodied foundation models that scale with physical interaction.Generalist AI Blog,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.181000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.181000Z digest=sha256:1cc843cd7d81e4f23e9d455f8405d1f7c2c81daeb8781a9864793abf9d5f6167

Observation 2e7129a0-07f2-4ea4-b51d-240ba3d2c36e · outbound

This paper cites Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.426810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.426810Z digest=sha256:82ee69559326b04995f666cf298f14a5ef5b9733d143b0dbb1e4deb2e0546616

Observation fd335565-bdce-4b16-94cf-dafbcb829c71 · outbound

This paper cites Gene-26.5: Advancing robotic manipulation to human level.Genesis AI Blog, May 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gene-26.5: Advancing robotic manipulation to human level.Genesis AI Blog, May 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.591352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.591352Z digest=sha256:29407c009396cf064d637de330ccc844db5863b9e4346194f0367d470c250848

Observation 6408f5d5-3515-41ab-979e-bda2709f60eb · outbound

This paper cites Motubrain: An Advanced World Action Model for Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Motubrain: An Advanced World Action Model for Robot Control

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.705109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.705109Z digest=sha256:ed519ce2afd591557522a2062ad442d57c148af20054a9b86863841eea23ca3d

Observation 05a07c5b-9198-4593-b574-79a2746497dd · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Octo: An Open-Source Generalist Robot Policy

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.846442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.846442Z digest=sha256:77e79d76d19c86cc3720a9b36b3af37a6edaf958d017abe7ec9941dab716c153

Observation 3dc7a77b-da8f-4386-9f13-545297db9c9b · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen3.5: Accelerating productivity with native multimodal agents, February 2026

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.970461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.970461Z digest=sha256:2193d9031ce87184897f88c776b8b9a8d80130399b31c7f160a96be920c4210a

Observation 37642b6d-bc80-4668-8e71-19881e850fcc · outbound

This paper cites Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.076726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.076726Z digest=sha256:2fc61eb98ddaddf5ddd451ebbd26c568edc88bfe42434915fc19384ba248a53a

Observation 44c93c5e-71fb-43eb-a527-f39338861437 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories LLaMA: Open and Efficient Foundation Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.222348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.222348Z digest=sha256:9efd3312dac345bd0c058a6bdfff8567f0cd2536394b20e91651c39be4396b99

Observation 1cb546a4-854a-4567-80b1-5f9eb81dba9b · outbound

This paper cites World2Act: Latent Action Post-Training from World Model Dynamics.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World2Act: Latent Action Post-Training from World Model Dynamics

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.361341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.361341Z digest=sha256:8cfc37d24247d896a81568c51b14d75ff445bfa2fe36bddf127e395f157db2bd

Observation e10fc9c9-56a0-4063-b656-35fbd7d67500 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Bridgedata v2: A dataset for robot learning at scale

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.546668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.546668Z digest=sha256:c0144d46384a0c8a32034607b789fb28c2b725536dfc380804d4ebafac304297

Observation 5bbb7262-ef3d-492a-be20-e2d399360cd3 · outbound

This paper cites WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.689523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.689523Z digest=sha256:802af35bd51a234f11890e987f04c8936f72afc4d6b3540bf7e6f134679bd6be

Observation 5ff7846f-3914-4da3-abbe-e85fa1fba95c · outbound

This paper cites A Pragmatic VLA Foundation Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories A Pragmatic VLA Foundation Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:08.883631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:08.883631Z digest=sha256:ac8303e802076eeec03064b8fd83f9f4dacf3d18ef6946cee8881f3b2c7f928a

Observation 8847c4a0-c647-44d2-ba2c-5483b46002be · outbound

This paper cites Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Dexumi: Using human hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.062173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.062173Z digest=sha256:f7de5820224dad4211640e9242a03fa36b5eeaf82e0907f1286082635a8fad76

Observation 571a8b58-02ec-4e07-8ad5-b643aa9fdfb4 · outbound

This paper cites MemoryWAM: Efficient World Action Modeling with Persistent Memory.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories MemoryWAM: Efficient World Action Modeling with Persistent Memory

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.236077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.236077Z digest=sha256:7b050162dc41b101e4f29e6f9c0ac3e0255e0a966fb0168a00047d58d5e58514

Observation aeee5494-a31e-4895-9005-01475331ad58 · outbound

This paper cites Gigaworld-policy: An efficient action-centered world-action model.arXiv preprint arXiv:2603.17240, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Gigaworld-policy: An efficient action-centered world-action model.arXiv preprint arXiv:2603.17240, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.365590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.365590Z digest=sha256:e90a2ad4dafb7ed02031b1b9835515b34fe99b68482e42225933e72272131451

Observation 7419d939-de71-4609-96b3-1baa5e0f54ab · outbound

This paper cites Starvla-α: Reducing complexity in vision-language-action systems.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Starvla-α: Reducing complexity in vision-language-action systems

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.530949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.530949Z digest=sha256:e619029381084292340e9d5669b06598f61489a908d6ee5c73ee5685d4b1125c

Observation 88207e1d-9378-44e1-b8a2-f40c32a463f2 · outbound

This paper cites World Action Models are Zero-shot Policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories World Action Models are Zero-shot Policies

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.676504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.676504Z digest=sha256:538526330eac92490baf19408599de4bd9e84802552ba011f7d44c86a6f5cdc3

Observation 628babad-8bf5-49a2-b9e3-d20096c08eeb · outbound

This paper cites Wall-OSS-0.5 Technical Report.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Wall-OSS-0.5 Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:09.858319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:09.858319Z digest=sha256:9b969fa2c6c8106ad171d6a5ef84fa49500210dc942cccf756a4da8749ec16b2

Observation 181cbaaa-febe-41b0-b66a-e702b18d0f09 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.006144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.006144Z digest=sha256:c71d04f2745e04d8b047798ac0b15bd60b4b4f8e11808c951bbf1f0b7d5f093e

Observation 395ff754-9e63-40d0-a1d6-14444ec8c58b · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.138148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.138148Z digest=sha256:4bef3665c403d8b90482a51a9fdf5bc708d073f8394fe7afee02d7e0987c65d3

Observation 20e0fc5f-94c4-43fc-ae96-5be679b4161b · outbound

This paper cites Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.301551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.301551Z digest=sha256:cbad140a14bf3a7bf1f7eaa9df3d92900a9a165571b78092b00074563a6b19a1

Observation 213f53be-0816-4ac8-8652-76f621b0e5f1 · outbound

This paper cites Native Video-Action Pretraining for Generalizable Robot Control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Native Video-Action Pretraining for Generalizable Robot Control

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.493281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.493281Z digest=sha256:b5af9e4e9eced2208bdb6cf6af61c3f04c9a37792f3bb0e13599a64c9db951b9

Observation 1a1a8114-70f9-4b9a-8a6b-382661e04863 · outbound

This paper cites Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.653923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.653923Z digest=sha256:fbc49d40bcd14004aadcc0451712e9565c3dab66092af80ddbeb0fa7682c1749

Observation 78a64214-cc8c-46c0-8ab8-8c27ff66a29d · outbound

This paper cites RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.767303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.767303Z digest=sha256:4f737e978fd67699711d50b920848376141f302da54ee611be5275c077b03325

Observation 8662e36b-e228-487d-864a-51bc0c18689b · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:10.975158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:10.975158Z digest=sha256:edfcdec94f462ae8085b16cf7c2043f63d7740640ac30d9748f439fc9e4d4f56

Observation abd2b7e2-65a0-4fe7-9342-6116df8470fb · outbound

This paper cites Fastumi: A scalable and hardware-independent universal manipulation interface with dataset.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Fastumi: A scalable and hardware-independent universal manipulation interface with dataset

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.125595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.125595Z digest=sha256:2ab93a7cef8189ee15d664e129821c1d270b1c07ce51fa1a699b72c276c27ce0

Observation 7574aabc-55ee-4055-8bc9-af23deb11fc9 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories TesserAct: Learning 4D Embodied World Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.217040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.217040Z digest=sha256:d2d1894f6388cb9de7520b7de410196d2cb3d485771137f0dc9ccfc517309d77

Observation 8fe5f4eb-12bf-4b5c-bf0c-183af1e848f3 · outbound

This paper cites X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.291258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.291258Z digest=sha256:2189a6739b9ffe052a8aac5b77da87d5b334f25a349fc95c410ab898a35c5e3e

Observation b17054e2-dfe6-4555-aa1a-2ee1aaba470b · outbound

This paper cites Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.427801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.427801Z digest=sha256:a13c086cfaf91f398688e08dd3e2c34a2a3c0ae86024fc4d96d77492ec8c2964

Observation b9a015f2-28a9-445d-9777-5fd99f0a73e9 · outbound

This paper cites Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404, 2026.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Acot-vla: Action chain-of-thought for vision-language-action models.arXiv preprint arXiv:2601.11404, 2026

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.570801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.570801Z digest=sha256:6580ace6af33166f31d52adb9abe522647cb96ac1704bb2e7664a359cdcede4d

Observation 720928f3-040e-4442-b1ad-e10b739fc6e2 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.705089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.705089Z digest=sha256:9c276a5e7fd77cbf7f1b820e8b47e21018319168eaf949f72eb40b4aae69fa15

Observation 6f9c9bc2-b575-4938-a372-d6cf0c9f1af9 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.845223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.845223Z digest=sha256:647dc516b94f6c221bcca2a1140d9c5a0b0b830c1220bfed779f7b5342117feb

Observation 9eb94489-b3e8-461a-985b-bc8690f8338e · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:11.999325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:11.999325Z digest=sha256:5722c796374dc4041cadd5447d580f7ea95e260de8598dc6163a9943d02f2ab9

Observation d79aef79-964d-4da3-96ab-35e32f332ef8 · outbound

This paper cites an unresolved cited work.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:07.331395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:07.331395Z digest=sha256:fde38382f7248baa871f9248669da4046d11b9dc9d4ffb01387fa592455c2895

Pith citing papers

Observation a0aac408-4465-4ad5-a4a1-9ef4b06df91e · inbound

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens cites this paper.

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T12:43:44.703075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:43:44.703075Z digest=sha256:5453334889ece2692509f82ed71ab1372ac9179facf8e3e9ac8869ce968fd916

Observation 4807410a-d1a0-4b83-b8e9-f36242eb9ed9 · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.603531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.603531Z digest=sha256:03672128bb52fd610dba8a06a2d7481af5864ba6d7ae11f827ee4659952e3f28

Observation 93313b23-be72-440b-899f-91968f484b7d · inbound

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation cites this paper.

OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:17:24.951028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:17:24.951028Z digest=sha256:fe7f3bdef0e6429e8fc717c1bf5b27acad0ff9846e7e57d20bd343adf2281ef0

Observation 26a539bc-f8cc-403b-84fe-ecb4a2e103a7 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:36.316445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.371109Z digest=sha256:7c15cfe24e1d5f3b34cedcdaf633b5a92220476a871fc099037cfbde68d7b267

Observation 35032d6b-ae97-4ea2-bd59-3487dbbc969a · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:52:22.178974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:52:22.178974Z digest=sha256:be12b39301754b1abb0891f9293ee82d7718b5346194582556ac56c1246981f7

Observation ba3a3086-a292-43c5-88be-b9a14fafaf44 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T04:54:24.134310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:54:24.134310Z digest=sha256:c6f2c25073c6e799f0742879fe0222d3f9b684bd0ef1f5b9b1aa483c5d4a0a83

Observation f11a420e-44e5-456c-94a5-f8161ea423b4 · inbound

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment cites this paper.

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:47.516551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:17:47.516551Z digest=sha256:ed06f51f38a9bf3f32fe39b14ad59f70f2eb53dd310388605e0a07fb8e1b73e1