Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:10:12.234233Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 8 inbound Pith citation observations for arXiv:2511.02776.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:10:12.234233Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:00:31.953138Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T20:16:29.417538Z
100 of 119 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 144983e2-5702-4ad3-9f44-d0b030581aea · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 753c30c1-7de9-45b7-8203-a5ac1cd3c7e2 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Affordances from human videos as a versatile representation for robotics
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e0108b-eac3-4dce-a6ff-c55ff0822cd8 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db6b1c64-be79-4b63-ab4d-746cf9ae12a4 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Katzschmann
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f600b06-495e-491c-b9d0-1d75ee74a2ea · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Hydra: Hybrid robot actions for imitation learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1cdcd1-0cc2-4e43-b7a6-c89bdcab0a4d · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations PaliGemma: A versatile 3B VLM for transfer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fd93ac-1f47-4625-8b33-eb70083efd76 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8106047e-be79-48de-a858-39188c005d9e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e9125e-3d14-4979-8d4d-8be679c1491a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb75da38-b5ec-46c3-851c-5fb5339c0625 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations RT-1: Robotics Transformer for Real-World Control at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04eef429-2e54-47c7-aecf-29342af00b36 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e71e394-ce41-4de4-b331-7baf1e0a7a8e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Univla: Learning to act anywhere with task-centric latent actions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c65b97c-59eb-4af6-a08d-e90cd4a2cb25 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mamba policy: Towards efficient 3d diffusion policy with hybrid selective state models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b38cc4-616c-4007-9599-7213a9cbc1d7 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f6a94b-1b35-437b-a9df-a8932f1fe9cf · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9567883d-a4c2-482c-9c37-5382cb4a54d1 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Berkeley UR5 demonstration dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3295f24a-8ad5-454f-b11f-46c65c8e7369 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Playfusion: Skill acquisition via diffusion from language-annotated play
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd6deab-5954-4d80-82e8-6a53268f3591 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Moto: Latent motion token as the bridging language for robot manipulation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f72c60d-595c-498c-a80e-14b4ad91529a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusion policy: Visuomotor policy learning via action diffusion
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79630e2c-e1e6-44b3-8fde-3bb04aaaa097 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations From play to policy: Conditional behavior generation from uncurated robot data
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4da7b4-5811-4065-8228-8b03bfa745b8 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdad1294-955c-4333-8c04-114a23788008 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04134676-4123-45f0-9d52-4d832ba0a1f8 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning universal policies via text-guided video generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cfa95a-1cb6-4a1c-82ba-b6d547c63b9e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bridge data: Boosting generalization of robotic skills with cross-domain datasets
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c0c563-8e68-481c-b63b-738e13460692 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusion trajectory-guided policy for long-horizon robot manipulation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b15971-601d-4791-9baa-676298fb8784 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Finetuning offline world models in the real world
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45a439b-ccd7-4c09-9838-b8abd4724564 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad50cffe-bb10-4229-9558-b830b5be44ae · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Llama-adapter v2: Parameter-efficient visual instruction model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f04a714-c3e4-4b15-9d6d-0826cfa7699e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Ego4d: Around the world in 3,000 hours of egocentric video
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d2c949-0b2a-466a-b74b-c3dce749f5cd · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Watch and match: Supercharging imitation with regularized optimal transport
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a524c1-1e7c-4f6f-86a5-f1a3281e72aa · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning an actionable discrete diffusion policy via large-scale actionless video pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5204f64-f07c-4c4a-8364-7af94f545343 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Masked autoencoders are scalable vision learners
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71a112a8-1baf-4317-a256-a0edf5bfd8a1 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db34190-8e2f-4bba-a705-4133b6d10262 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Sacson: Scalable autonomous control for social navigation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bd37ec-d88f-446b-8827-8ff991538b51 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8734c05c-1182-44e2-9a35-0fa48370b3ec · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9458080-7f50-4add-b098-02cbd520b0e8 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bc-z: Zero-shot task generalization with robotic imitation learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ed9004-e513-4d93-b29e-16648d32ab3e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b52e4d07-f359-490a-8300-c46515766fe6 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scalable deep reinforcement learning for vision-based robotic manipulation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e94f795-4a51-431f-82cf-55bf1fb36c1d · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f3b0ac-7986-46f0-a73d-bd727356852f · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Pre-and post-contact policy decomposition for non-prehensile manipulation with zero-shot sim-to-real transfer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c41af8c-e6fc-443a-b0e9-155dd9149553 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Openvla: An open-source vision-language-action model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca69b201-2f52-43d8-b5ba-17cc8db19fa2 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Molmoact: Action reasoning models that can reason in space
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9140fdce-3540-421c-b45a-f6981d7864ed · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Behavior generation with latent actions
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe9b4a1-4422-440b-bdb0-dd7bcc82efbc · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5284e8cb-b687-4602-b185-1511a45cc04a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6230c424-1809-4676-b3b9-9ae333c0935e · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unified video action model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 137ebdf4-0baa-4655-8017-cf376523f7cc · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Visual instruction tuning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e597d47-7142-4a7d-94e4-2257f2e5c1d3 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robot learning on the job: Human-in-the-loop autonomy and learning during deployment
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac21cc6c-f1ec-4bc4-91b7-a4036582c6bc · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40decb04-7570-4192-a8eb-534ffa2bed53 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Rdt-1b: a diffusion foundation model for bimanual manipulation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71ea9ee-6d4f-44da-9de7-5d36afc2ef8a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mla: A multisensory language-action model for multimodal understanding and forecasting in robotic manipulation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80e09eb-6a4c-470c-8664-5da4b8e13c5a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Multistage cable routing through hierarchical imitation learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d60139-4fbd-4752-95f8-0662f9078d03 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Fmb: a functional manipulation benchmark for generalizable robotic learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18ec6b9-1c8c-424b-995e-3b587543e078 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Interactive language: Talking to robots in real time
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb58e84-063f-42dc-8106-37bca937be25 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46874027-59bb-467f-a1b4-8e3a0cee3259 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Weblab xarm dataset, 2023
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c98143-8b0d-4e48-98ea-b5d444e02662 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Grounding language with visual affordances over unstructured data
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031cffd1-bd7f-4110-85ed-da9abdbfe50c · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Structured world models from human videos
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fe710df-da85-4924-ad5e-2f817067b96d · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Quest: Self-supervised skill abstractions for learning continuous control
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75dc6fd-c2ac-4cc1-86d2-a64d4bd57b81 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning and retrieval from prior data for skill-based imitation learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446a9553-f112-410e-88ea-df27b967d021 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations X-embodiment u-tokyo pr2 datasets
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e547628-86e6-4881-8229-76a0cab33be7 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Motion planning by learning the solution manifold in trajectory optimization
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b537d90d-2f9c-4d83-89e1-3e4b24a93e4b · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721772f7-7d4c-42bb-9ab0-cd457790e4a8 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations A guided reinforcement learning approach using shared control templates for learning manipulation skills in the real world
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26060803-b9b7-42e2-8a50-e309d76e1f15 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Guiding reinforcement learning with shared control templates
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e03fce4-c705-49dc-b7e1-2cdaee5b4a02 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations The surprising effectiveness of representation learning for visual imitation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119428c6-2dad-4f44-9c28-c2f66df48c30 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Supramodal and cross-modal representations of working memory in higher-order cortex
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc4743e-107b-49ed-8f56-38945ae5f088 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Embodied artificial intelligence: Trends and challenges
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e23507-f215-4a1e-83d9-44547ac2375d · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Spatialvla: Exploring spatial representations for visual-language-action model
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32460c80-671d-4a3e-9b6f-32ce07776ea1 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Shared control templates for assistive robotics
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d59140-da02-486c-bf2f-76d9bda11b0f · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robot learning with sensorimotor pre-training
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3cb414f-ae7d-4197-aeaf-8681cc03e89f · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Real-world robot learning with masked visual pre-training
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594eeb26-8057-40e5-9503-30ac30c1bb75 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent plans for task-agnostic offline reinforcement learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa28f33-5aa4-444c-9327-e4cda65d4a72 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Multi-resolution sensing for real-time control with vision-language models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95f86456-c115-4c7d-afb7-54fbf8a6e623 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Behavior transformers: Cloning k modes with one stone
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04dc17a6-6f36-47c0-b98f-f66ada632dec · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations On Bringing Robots Home
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a16ff2bf-65c2-4c28-8bdd-711f1dc82224 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Rapid exploration for open-world navigation with latent goal models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc10689-a07e-48dd-9ad7-15260452aa0b · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mutex: Learning unified policies from multimodal task specifications
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 355ce202-6c60-4ac3-931c-265353643d67 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Dense policy: Bidirectional autoregressive learning of actions
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb667ca3-93e9-4f27-a5e8-c68b890692eb · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Gemma: Open Models Based on Gemini Research and Technology
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36e2307-7873-4fc9-ae11-84f95fba208a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Octo: An open-source generalist robot policy
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1337f9f6-4bcb-46d6-aa06-d2a62a2b153b · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations WaveNet: A Generative Model for Raw Audio
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9585b2ab-8201-4936-9b02-bc02407b0284 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Neural discrete representation learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c37cfe8-27a7-4995-a3e1-cea88676ab5a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations o rn Vogel, Annette Hagengruber, Maged Iskandar, Gabriel Quere, Ulrike Leipscher, Samuel Bustamante, Alexander Dietrich, Hannes H \
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5c6a59-3c93-430c-8b80-aa370fc9cf32 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bridgedata v2: A dataset for robot learning at scale
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e48d89-d86b-4fcf-855e-d8e723f98961 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218c11ce-91c9-447b-a012-089f35b0b18c · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c611ed-4a5f-4560-a3b5-845630f0e112 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Dexvla: Vision-language model with plug-in diffusion expert for general robot control
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d67c732-7a4a-4e6f-939c-f55d671d6799 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73ce338-90dd-444f-aff9-77fc37c0a5ad · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusionvla: Scaling robot foundation models via unified diffusion and autoregression
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759101fa-7f2b-426a-988d-43d1bcc6909a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9fee724-7114-45e7-ae48-644538fc7b9f · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Discrete policy: Learning disentangled action space for multi-task robotic manipulation
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa98af1-187f-41ee-8774-769f5e8ed77a · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Florence-2: Advancing a unified representation for a variety of vision tasks
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d933d7bb-30bb-49b3-b79c-50967c8e9e05 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent Diffusion Planning for Imitation Learning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142850b9-fe0d-4e56-a3cd-16e5ed750074 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations ucsd kitchens Dataset
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b261577e-4f36-4937-9711-42cb1623578c · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Instructvla: Vision-language-action instruction tuning from understanding to manipulation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef1a4a8-df81-4eb6-bc88-5266618f1ba3 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent action pretraining from videos
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e37b523-ce46-4a6f-ba56-44be7cc1bb49 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4fcab4-40d5-4179-98f0-68524b9f79c6 · outbound
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917c4e4a-4397-44a5-b114-89a12a9de230 · inbound
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3a3171d-600d-4eab-99b5-c195e1c9bc09 · inbound
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8244850d-5b03-4c4e-8aed-2bc4a374edd2 · inbound
Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d20ff8bd-83b4-46dd-aad2-b09b0ff8d722 · inbound
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation caea0d14-8334-4eaa-b796-099b79db7b7b · inbound
Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e9949d2-8952-42b1-9ab0-353e7516db53 · inbound
GeoProp: Grounding Robot State in Vision for Generalist Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b92120f6-af3b-4f60-816f-d59befd8112a · inbound
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5949a4-a94d-45ae-b240-7ad55ab83bd6 · inbound
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.