Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:00:48.172438Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 22 inbound Pith citation observations for arXiv:2510.27607.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:00:48.172438Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T22:22:57.515817Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T04:09:35.372388Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6df8e2d7-5b34-4efc-bbd4-3b7d601b151c · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cefecb-96fe-4570-833e-a39278f44d64 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Enerverse: Envisioning embodied future space for robotics manipulation.arXiv preprint arXiv:2501.01895,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b77faa-ffaf-447f-86f0-50c9ae4f7756 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ec8f49-49da-4f6e-8e9d-b268b49409eb · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1067cb-fee6-469e-a104-84c13e6da7a7 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f52d05e-4059-46fa-9a42-e8e09518940d · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINOv2: Learning Robust Visual Features without Supervision
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bbb80e-455e-4326-b6c8-5b277c9c80eb · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bddf48f-d85e-4084-a283-4d9b984e5ff2 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf338ef-eadf-4f53-90f2-7ef1640b41f8 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Genie centurion: Accelerating scalable real-world robot training with human rewind-and-refine guidance.arXiv preprint arXiv:2505.18793,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcaf861b-f407-475b-98a7-80f134dbf97c · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f89afb5-6bc5-496d-a056-6dc01b81b31c · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f87478-21fb-4e43-baed-96286917918a · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model FLARE: Robot Learning with Implicit World Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2975b3-4393-4e7f-bb87-a867b40fcd6c · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e5be1c8-6589-4dd6-835e-2f62be821ac1 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Image observations include 3 viewpoints from the left, right, and wrist
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67dcb501-5c7f-46b8-b48c-f4c26ca48b8d · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model The simulated robot is a GR-1 humanoid robot with Fourier dexterous hands, enabling fine-grained grasping and manipulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d37e046-644e-4868-8bf3-105848765b48 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea14750-e58d-4048-b56d-b8b74d8d3eb6 · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e2927a6-ed44-434f-9df3-2c89bc35768f · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model This&That: Language-Gesture Controlled Video Generation for Robot Planning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2874ea-6c8b-49ce-8efa-7104d79a790e · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90529829-446a-46ba-b581-2aed5cffbc7e · outbound
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b749c4-fc25-452f-abfa-6aecfe15734e · inbound
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88b916c2-e2cc-44d5-a27c-2c4b8992c4bb · inbound
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5c46c76-bf57-419e-a747-94e7625a8a7f · inbound
World Action Models are Zero-shot Policies Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 96b6dee1-172f-43e1-94f9-b09c71c70ee1 · inbound
Fast-WAM: Do World Action Models Need Test-time Future Imagination? Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0729ffa-837f-4672-bc4f-6d0e7cf66b5c · inbound
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f281ee1-c735-4eb8-bae8-d163e34277d2 · inbound
RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a164519-be3e-4db6-aed4-cf123fe18cc1 · inbound
RLDX-1 Technical Report Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7321da2-609e-46e7-91df-c5fbc8ddaa54 · inbound
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6f68adc5-ad17-4443-a28f-c243e88d430f · inbound
World Action Models: The Next Frontier in Embodied AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1756185-b4f1-4a41-8fdd-bea0b2f35d18 · inbound
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9a28dc9-a6ed-4b2a-a428-e0d1bd845a0d · inbound
Point Tracking Improves World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ec89e32-5612-46ea-bdf7-5cf57c296e4e · inbound
World Models for Robotic Manipulation: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aaaf0dfb-a2e0-4c18-9125-e261cee57598 · inbound
WALL-WM: Carving World Action Modeling at the Event Joints Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7a54e59b-80dc-4fd0-aae0-964e75e15e52 · inbound
Next Forcing: Causal World Modeling with Multi-Chunk Prediction Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 99f314d7-3ead-4642-aa27-b6946760aa0d · inbound
Making Foresight Actionable: Repurposing Representation Alignment in World Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d07178fc-7b55-4ba4-b038-3c950f608f1c · inbound
World Pilot: Steering Vision-Language-Action Models with World-Action Priors Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc3bc517-242b-425e-b212-a123e059d14d · inbound
MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d166564-449c-40cf-9737-b0f787255849 · inbound
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7edcb786-9da0-4025-b9f2-220881f800bd · inbound
World Action Models: A Survey Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 694347b4-43b9-4633-a04e-80a7d8ec6aeb · inbound
Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2c49e2-5f48-40c0-993c-b9314e9ac0ae · inbound
UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eecc954-fad2-4ecf-b563-8cbb4475a9dc · inbound
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.