Pith. sign in

Paper Citation Record · LEDGER

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

As of 21 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 4 inbound Pith citation observations for arXiv:2607.11643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11643 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T04:10:14.360463Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:52:22.573456Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:35.595553Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25d1cb50-f7ab-4983-8cba-96e8942e7e39 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:fd5c418fac905dd7975c9bfd82ffcbcea2a0462c85a4ea435ac652f78119cc42

Observation 9b703af7-e8db-496d-81b7-3c1d338916e4 · outbound

This paper cites EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model EgoWalk: A Multimodal Dataset for Robot Navigation in the Wild

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:2ee0e56ab014db54cb1af48302431eb0a9627b36a5e339de533324ef048abfa4

Observation 33f43ff5-8da2-437b-929b-e721582f7d0b · outbound

This paper cites Qwen3-VL Technical Report.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:6a48d9b4ff8ee788e77db03ebae38443a70031f263ff2a63f46897dbf2dc4be3

Observation b20ba17a-0b01-4b5e-9aeb-4ab1fa507b95 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:240cfd1df9c705163d519c16cbf7c261b96beaa22d963f47188981dd99db5dce

Observation 035d49b0-7836-48a4-b4aa-281da599d97d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:3be6404442de16f828a820d015cc4d592dab96c38ca1c0fbe5dd32085147f5aa

Observation a699d510-3ca3-4e4f-925f-0606dca85e9f · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Instructpix2pix: Learning to follow image editing instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:6f76e4fa8c5c8c87545177176155e9039196c749ebae6ab12cd1cd6e89fd23e0

Observation de833add-c929-4160-8baa-75a8d0e195a6 · outbound

This paper cites Genie: Generative interactive environments.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Genie: Generative interactive environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:288d6f94f5f74de88faa9250fddab6ca1b15c4f49880da6bf2d9b7eedd582876

Observation bf09b2d1-aff2-456b-8079-560c4bb064a1 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:8bffd28946b3c4b90ce4cbea440461abf71001cee4f8942c0b1b37c854840eb7

Observation e8f31ebe-1523-40c6-98f5-55c865a515be · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Sharegpt4v: Improving large multi-modal models with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:06277a1ad498df835dc7b0df5c5f4fba3df7351d123522b58e1b8d75fe5497aa

Observation e8569a7e-2de9-40f1-8e41-eb3990da8d6b · outbound

This paper cites Video depth anything: Consistent depth estimation for super-long videos.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Video depth anything: Consistent depth estimation for super-long videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:c5b6728dd5e24bfbcca271a96ac4b58f12dbbbae67bb26193190323d21ddb35a

Observation fa3e6eab-0b56-4ed5-b686-33805ab2dbc8 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:294bb763a62681380bf5bc39612b2a6aaf78b50006d0708f5f77c5e8be04a19b

Observation 1b260aef-6c6c-41b0-8cda-a66f0feda4be · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Anydoor: Zero-shot object-level image customization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:6ef6a840f718ce746813bb99e8bc111fae5cd73d39cf67c2b67a00fde9151d49

Observation ab5b3001-45ac-4c9f-8fa3-3dff5dd44c25 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:e13cf4b8d2729d95a0ffb5f730c883ba0a3ad2d14c27e0b727d2e22a38eb0970

Observation a1255016-3c8f-4d98-8479-9c32b3e2996b · outbound

This paper cites Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment.arXiv preprint arXiv:2603.23376, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Abot-physworld: Interactive world foundation model for robotic manipulation with physics alignment.arXiv preprint arXiv:2603.23376, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:d9f9310361ae5b6ada8e6e22fc6993023a0d4e988ac799b7ae2f5fe5aa7b3455

Observation d9c04fb4-e626-4a0b-a72e-39345b1d9364 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:08619850d492389453364f4085d86bffd44d8879bd035b9757eb1c5806416141

Observation 92e9c856-d691-4660-876c-8410059df1e1 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Emu3.5: Native Multimodal Models are World Learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:8fd249d66a7ac4dc1a71f871b3fcda85a4f83b304005cbdfaeaa6550ffe5ece3

Observation e2a933b3-dd95-4a6c-ac2f-318e67acd313 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Emerging Properties in Unified Multimodal Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:638eddefa6f24a9270873e5b366202d248f65766e4f9ea500f05888871bd53ce

Observation 8a226344-61f7-4e3b-9ad6-d64b8b3c88ee · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Datacomp: In search of the next generation of multimodal datasets

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:ac3751bfd9b9be742bd348370ed8085fd52077db09c95b49c0bac5a2642b9ac1

Observation 72c5905b-c8f3-4319-a756-6235e976ab95 · outbound

This paper cites CAT3D: Create Anything in 3D with Multi-View Diffusion Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model CAT3D: Create Anything in 3D with Multi-View Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:9c4bcd9f5b707dae0c835b1e8578467d23a3a2373566284a5b12139efa56486b

Observation 07d9e81a-d54f-453f-bd47-2ad3655233a8 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advancesin Neural Information Processing Systems, 36:52132–52152, 2023.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Geneval: An object-focused framework for evaluating text-to-image alignment.Advancesin Neural Information Processing Systems, 36:52132–52152, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:5bb110f8318b783d478d1f068441f43408230850bb219b81de18893f35a80c16

Observation 03af37ba-a092-426b-bb2c-cd10da259f62 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Ego4d: Around the world in 3,000 hours of egocentric video

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:f21c3a07e2d6a4b99c790a33ca44828092d7cf677310b9cf76697dba2d33e369

Observation e35addfd-360a-4cb4-a76d-6175336c4be9 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Dream to Control: Learning Behaviors by Latent Imagination

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:83bc7b1b1c53df8fe1060b152d684065be2953dba93d03e932bf5e980decaa12

Observation 9a52e92e-9286-49e8-b6c7-c703c6720758 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:c58a2f98ad6bd70a685133312ffcdccf52da4bfd8f86363857a26bf67e01dd7b

Observation 135de540-e5f2-41bd-972a-80662266ca6a · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:97bceac712a1054c7d4fdf8b06f82bb38d015f85dea49fc7df6aac4119531b6d

Observation 1f7e5050-230a-4779-ba7c-43cc16e35df0 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:ade322cc4ffd3215b07f9da0fcb7a8140613b2428b93639ad49fb88abfa15203

Observation a1371734-58d9-42c9-a963-cd1a3d975dc7 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Gonzalez, Hao Zhang, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:efdbd987e9d22278334ef82dec3334008cbe203b0005bbab57a5ec9f6a8dc691

Observation a121a23a-f707-4b60-b126-f9471ddd1c98 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:af7c1300393ffccb025d05fd06c1b19c850b7eaadbd6f1d386a693d8328ac886

Observation 4db0f5d5-6247-4c03-9472-e509b68c1613 · outbound

This paper cites Causal World Modeling for Robot Control.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Causal World Modeling for Robot Control

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:5c77a9dc80e2e3e7075c22c4862d5deb7331d38999de3429037724a893506a10

Observation 85502411-ea37-4d6a-b369-0909e560d945 · outbound

This paper cites Era3d: High-resolution multiview diffusion using efficient row-wise attention.Advancesin Neural Information Processing Systems, 37:55975–56000, 2024.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Era3d: High-resolution multiview diffusion using efficient row-wise attention.Advancesin Neural Information Processing Systems, 37:55975–56000, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:78c316ce0161e9cbe94d05f0b0c6218bce75ded95f0304bf5cc790a21471d989

Observation 955247b7-9473-4892-89f3-23a6332fb479 · outbound

This paper cites A Comprehensive Survey on World Models for Embodied AI.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model A Comprehensive Survey on World Models for Embodied AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:7ea460c86477a43cba81e9065ba95e7d83d9022a0fc77caab007bda7d1911cea

Observation d0fb85fc-7708-41b0-8ff5-271b3c8d9748 · outbound

This paper cites Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:cf397579450d2663a2d097c379b06e2c8d11efcc476f426d0593cd9baaa2687e

Observation 0be5d850-83a7-4e87-be62-76e42ccfc3cb · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:8d5b80e40f03a7302a61fa3ad7369db767c2efbedc37d04777b2509beec430b3

Observation 69fcc8ba-6883-4af7-9daa-ee0c12e034f1 · outbound

This paper cites Syncdreamer: Generating multiview-consistent images from a single-view image.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Syncdreamer: Generating multiview-consistent images from a single-view image

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:076b1dc009b845acf710eff483f1704569f96a58687c83070afb01a9867bf109

Observation 470bcd0c-8e53-4363-a6a0-954733f9581a · outbound

This paper cites Scaling world model for hierarchical manipulation policies.arXiv preprintarXiv:2602.10983, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Scaling world model for hierarchical manipulation policies.arXiv preprintarXiv:2602.10983, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:e57aa12e711988e02b5fa7b8d99397e53bd15973dc764b4e004a00df02c04071

Observation 383f7fdc-d685-4999-987d-e240a7e1c98e · outbound

This paper cites Wonder3d: Single image to 3d using cross-domain diffusion.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Wonder3d: Single image to 3d using cross-domain diffusion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:c6aee4408fc7b40b363acd9b440e5bf8e115dfea5a76eb4bea08e0a21bfc36ed

Observation ca8f010c-7c01-4108-b8b2-330e2babfa19 · outbound

This paper cites A Survey: Learning Embodied Intelligence from Physical Simulators and World Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:e01d608c7fbc1178560e7822415d2dd3ae2a94ddf00e51a0e81c0860557b8ada

Observation 2cfd7eff-2834-4579-ad67-7f2152fb55b2 · outbound

This paper cites an unresolved cited work.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:164873f81e757cfbc7383154aaea6709160ab4aa060a114df4735eb3b9c55be3

Observation 4409df06-45ec-416e-ae33-22c6eb28bc74 · outbound

This paper cites Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:f1c2ebe43acdd16eeb6d7b7a3130fa83075dceede285ad9d35c87229bc061e5c

Observation 453f219d-6670-4ee8-80d6-efd07f387f73 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:a1481ada6d74d3b3dc3706c9553d249ec9553e9ffbf7dc7263d17ea230d55b37

Observation cab8e87d-1bf4-4396-aa2f-186231b0206a · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:2716a5467b18a5e65ee6a1231f327e9ac3e2ab521e598b0666a1c47cf04cd615

Observation 3ac46b97-7aae-419f-a1b4-56b2d380b00f · outbound

This paper cites Video generation models as world simulators.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Video generation models as world simulators

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:8a0f84ab10bccac5ea801524a61d756a0613fd1b86cf4bae49e42cce38f5770f

Observation 2696929c-32f6-45d8-9165-8041f31334b9 · outbound

This paper cites Gpt image api.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Gpt image api

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:e4f35003af55be3763f8f60485deb786d08943b5e3a63597505fc481eed2b7da

Observation 3a1f785d-0ca8-4620-a1e5-47ce6260c783 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:d14a5474120c9be466934dd36f94b4e13e1305b82ca7a919fcee134cb001ab4a

Observation 32d835ba-d171-4bc2-8c5d-eaac82f9f34e · outbound

This paper cites Genie 2: A large-scale foundation world model.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Genie 2: A large-scale foundation world model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:459698c7c431fb243b6fa086985072e326a5a9d03950390beef235bac67030fc

Observation 9ba66005-3015-43db-96c3-45bffceb3259 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:5b1bf185887c085da99086d9dcb99d9dd2690d1540c42125f0bb87d1d48183cb

Observation 48887efc-c940-4fd8-b9d2-1e9935f3ec04 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:770fe37f3333402140107be60b34ed8ff338aca4947b7390bd80dbf9c674725f

Observation a164c70f-4fc0-4b92-8fb8-b1f11254c5f3 · outbound

This paper cites Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:7aaeb1a7ce69fc688ce6a1f92dfb2db6d734f34a376274024d48ec817ce6da98

Observation 8b82af18-daa0-4297-99a7-f29fe7b1e0a4 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:144a789d0b246156522ff75bc31cdd8ad2a7658dc09c4c969f7240b6f0af2170

Observation a3e54ae5-c7fb-4dd3-b500-3b0f118587a4 · outbound

This paper cites Worldarena: A unified benchmark for evaluating perception and functional utility of embodied world models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Worldarena: A unified benchmark for evaluating perception and functional utility of embodied world models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:79a01cda8c0c5ab25f48b0a73a47e71c22a04d3fb78ea3174f360a9e36571f7e

Observation 7b4eba0e-70fb-4e7a-83cc-82cd5264520a · outbound

This paper cites Roboscape: Physics-informed embodied world model.Advancesin Neural Information Processing Systems, 38:63674–63698, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Roboscape: Physics-informed embodied world model.Advancesin Neural Information Processing Systems, 38:63674–63698, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:2804c6d91ed337ab759ea26c1da2b34880cd3c17c1f7b07f4b9ec3c4cda1837c

Observation a914582d-f57a-4673-a7bd-fecb3df3bb63 · outbound

This paper cites Scalable image tokenization with index backpropagation quantization.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Scalable image tokenization with index backpropagation quantization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:165986c76d4fcfea4d9d21b90e147a754776db59a42e662bcdf22f425af62894

Observation a388b8ac-8e5f-49dd-a918-47c58ae4a629 · outbound

This paper cites Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:3da1abf03acb5d331cf47188b8e12b93a2e8138643741beb039e4c5c7c8c2355

Observation eb11014d-7839-4daf-b455-225cfb1a84d8 · outbound

This paper cites Mvdream: Multi-view diffusion for 3d generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Mvdream: Multi-view diffusion for 3d generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:67169b98fec58af57e78e387d8a5a6cf67514d699584c2e8de8dd0b84e984e84

Observation 91d17572-b55c-49c0-ac59-0f10e985d8ad · outbound

This paper cites Mvdiffusion++: A dense high-resolution multi-view diffusion model for single or sparse-view 3d object reconstruction.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Mvdiffusion++: A dense high-resolution multi-view diffusion model for single or sparse-view 3d object reconstruction

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:0192f1467756c941f54d814191c920c4fc376ff1e7b45d30d773f76f9e6f28a6

Observation ac68dd1a-6763-4cf3-a574-cd56f91116a1 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:c6b8dc8a131a2026aebdae1d30c070cc68ec494b97822f0d520f5e2f7f688d2e

Observation 395e0ef3-5d90-4275-976c-8414b6cdcfb2 · outbound

This paper cites Motubrain: An Advanced World Action Model for Robot Control.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Motubrain: An Advanced World Action Model for Robot Control

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:091ddafabc0badd6de4da5c1f83eb6c5cbfc3e7313e5bf3fc80f7743420ae7d9

Observation 3a860f55-7b1f-4cbd-a54e-27ebc8ac87ca · outbound

This paper cites Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:5318ee1963cf327ceb8d1c0ea68ce2af961ab465403cc7345a4fde67211b6b6d

Observation af4f344a-2ecc-4ca0-a6f4-90da90682ece · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Bridgedata v2: A dataset for robot learning at scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:4f6c1bba42e4aa241d7dd10a009f2ff68c8fd8236190dbf0773d7fc77e7ab676

Observation 0e83024b-49ed-4857-9cee-7161f90833f0 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Wan: Open and Advanced Large-Scale Video Generative Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:c80af9d024a3debb793f518bec5019708a9f597f9a24ef5f0e823ec6d228a4c3

Observation 12ab3c7f-c963-4256-bc8d-8683f3172a30 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:530da0ae60ad0480a8b0ab2eba335bc2fb48b2e4719c3790ea180af4efb25e5b

Observation 3d9ecfba-994c-42ed-b49c-70e0514c56b2 · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:74764d2867cce98b9a39558628de097464938256bdcfea30b3f4e9c1a7bcc7da

Observation 35487c09-54a4-41cb-a016-4cf73400acc3 · outbound

This paper cites RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:65e5862470309ada96a6ea368d709fc33b53400455bf9f62a42219eb04c2b97b

Observation 7b472a61-80de-4110-89e4-5be009d9d89d · outbound

This paper cites Omnigen: Unified image generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Omnigen: Unified image generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:28fc9991e6726dd110434ae8c4148cba3808b68441929b2ed17a731d25dd9262

Observation 17623436-a460-42a2-a002-4207ed47420a · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Show-o: One single transformer to unify multimodal understanding and generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:778c3150b21cbf3f41be32f0d71c200f6224941af26b65ecf719d753d81d6a71

Observation 91523548-53cd-4fba-bce7-be45c7218d8e · outbound

This paper cites Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:41a73535e9c952d4eda4f93698d3b223ddda9451aa55e84debea7c2dbd1a4b43

Observation 1cd938d0-8b01-4fec-b5fb-ddf04e83ae3b · outbound

This paper cites World Action Models are Zero-shot Policies.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model World Action Models are Zero-shot Policies

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:821f099ed75d671831b646da1bbf8d76ca40079ebbb367ca83b5f0d975d407c3

Observation 095b6f19-669c-40ab-9002-e10522d2af12 · outbound

This paper cites Imgedit: A unified image editing dataset and benchmark.Advancesin Neural Information Processing Systems, 38, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Imgedit: A unified image editing dataset and benchmark.Advancesin Neural Information Processing Systems, 38, 2026

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:b7bdc8b7cec55fef21ad5d541c47fa895d7b06671a71964cd8d84df226551e2f

Observation d30b35a0-19d2-46ba-91e1-3a2da8750ce7 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:ab3d127019416702afd457fc3b20b26a2deb1a3042cde1b9c42892f0ff0caea0

Observation 5baae18b-f968-442e-8278-df941a27b82e · outbound

This paper cites Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:4d004476e9e65209f413f6475e668ebf610023ac7b885f7196e6cc1968357418

Observation aa6b7954-e39c-4c9a-87cd-5413fa3c06c8 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:ab3d72aa774a613deca8a3302d6e5a81ae78327d83ee01dd185d5579905ff6cc

Observation 4955f32b-32ad-40b6-b654-5cf3c472f128 · outbound

This paper cites Scaling behavior cloning improves causal reasoning: An open model for real-time video game playing.arXiv preprint arXiv:2601.04575, 2026.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Scaling behavior cloning improves causal reasoning: An open model for real-time video game playing.arXiv preprint arXiv:2601.04575, 2026

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:f230f47616433856353bd04f35cb34644a32b46ff51dab94f133e19675ccff45

Observation 63bed55b-4d07-4791-969a-7ee1d39e4990 · outbound

This paper cites Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:05da8dd0327d17b13d4ae6dcb8e8885ddcda1f9f21e71cc6d4a15de1f9b99634

Observation 0414151a-9e34-412c-a634-0586e1233a1f · outbound

This paper cites FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:e43a4acbe5590a3576c7e5f3ba8d4ebae0b36ffde4285fc22aed4b70a43697ff

Pith citing papers

Observation b602de26-131f-44c4-a021-1052c3b0a0c6 · inbound

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens cites this paper.

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T12:43:44.545121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:43:44.545121Z digest=sha256:1f308981c9acab167978b0df9250a450ee1d7670c3e83183266ca913db5b5ea5

Observation c6fc42ca-b008-4b79-94ff-0eedc6c123fd · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Reference 219

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.753396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.753396Z digest=sha256:b7aba85f264719e8e37e4d749a22be863f690dfed2486f878d204b67d363a347

Observation 8027843d-d15f-41df-839d-1b1adf596ff4 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:35.600405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.554523Z digest=sha256:bf4907c22f2e15eaf3140e5583968f94fc67cbe712734b4cab64b48b668bb11d

Observation 89e2e64e-d393-45c9-a456-dde4bda4fe6c · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:52:22.573456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:52:22.573456Z digest=sha256:90fdcb718f989ade14e2a600f958e7f65f50386cebea1429a53e7c4b4eda8d1f