Pith. sign in

Paper Citation Record · LEDGER

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

As of 19 August 2026, this Paper Citation Record lists 100 of 119 outbound references and 8 inbound Pith citation observations for arXiv:2511.02776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.02776 v3

Coverage vector

measured 100 of 119 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:10:12.234233Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:00:31.953138Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:16:29.417538Z

Reference resolution

100 of 119 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 144983e2-5702-4ad3-9f44-d0b030581aea · outbound

This paper cites write newline.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.646911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.646911Z digest=sha256:d2809c9aeff64d7a09411494394412ae601217ecd3cff3b4824a4d6058be80c1

Observation 753c30c1-7de9-45b7-8203-a5ac1cd3c7e2 · outbound

This paper cites Affordances from human videos as a versatile representation for robotics.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Affordances from human videos as a versatile representation for robotics

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.716220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.716220Z digest=sha256:5b926f37c2fb73572eaa10b2456dc122ab7ecbc43269c6f2c47aca5bd921d2b4

Observation c1e0108b-eac3-4dce-a6ff-c55ff0822cd8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.794099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.794099Z digest=sha256:b244d2931007eaca8fc24d8fae016a84764d0352113ac33d6e5f3e8b6c225c49

Observation db6b1c64-be79-4b63-ab4d-746cf9ae12a4 · outbound

This paper cites Katzschmann.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Katzschmann

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.857864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.857864Z digest=sha256:658f32084f5b2913409cd314741bf8cbc36d12f76887b27e70773dc9810bfe88

Observation 0f600b06-495e-491c-b9d0-1d75ee74a2ea · outbound

This paper cites Hydra: Hybrid robot actions for imitation learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Hydra: Hybrid robot actions for imitation learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.861375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.861375Z digest=sha256:f92ed49352007297e70f890df8333929fdb5e933a95581fae172d141aaa389fc

Observation 6c1cdcd1-0cc2-4e43-b7a6-c89bdcab0a4d · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.864723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.864723Z digest=sha256:64f4c7f42263dbcb204bd657a76fc79769692787dc9effdd8dde46f0d5e70c52

Observation 26fd93ac-1f47-4625-8b33-eb70083efd76 · outbound

This paper cites Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:00.932701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:00.932701Z digest=sha256:abc1b69067f7adc14e215582d36dc0212d6b7768da3d2c685d8402485a9598ff

Observation 8106047e-be79-48de-a858-39188c005d9e · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.100730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.100730Z digest=sha256:07c19bc91b4fb77f9b8253fc03f266c6a6c76e0d355b4968646069d7a9c32805

Observation 47e9125e-3d14-4979-8d4d-8be679c1491a · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.206655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.206655Z digest=sha256:c5f6dc7089cd7d34fc4b6d2aadedfe067ee035f78f909d9a8e46c58996b85955

Observation fb75da38-b5ec-46c3-851c-5fb5339c0625 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.371161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.371161Z digest=sha256:8ef45c1b69b488662fb0043fdd280e26f997d3b831605de487e90c461c1da940

Observation 04eef429-2e54-47c7-aecf-29342af00b36 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.535943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.535943Z digest=sha256:9053542123a7d90e0fd105195f480888356bb24b2d8fbf776b1b0fb1cdad6b58

Observation 9e71e394-ce41-4de4-b331-7baf1e0a7a8e · outbound

This paper cites Univla: Learning to act anywhere with task-centric latent actions.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Univla: Learning to act anywhere with task-centric latent actions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.704216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.704216Z digest=sha256:db3301237b05ce31713a6a9c46cfaa0fbe7edc2e5e162e0ade8b996ebc82bc9b

Observation 2c65b97c-59eb-4af6-a08d-e90cd4a2cb25 · outbound

This paper cites Mamba policy: Towards efficient 3d diffusion policy with hybrid selective state models.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mamba policy: Towards efficient 3d diffusion policy with hybrid selective state models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:01.871844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:01.871844Z digest=sha256:ae85e5c92ccecd9eb4d6f9c7a6f7332cae3a8e48aa9cdd890ad86ae0ef15b097

Observation a6b38cc4-616c-4007-9599-7213a9cbc1d7 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.038513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.038513Z digest=sha256:6fa32d0b70983162544b4f52fb3110c16f715be75ee089ffe048d0b2719e891b

Observation e9f6a94b-1b35-437b-a9df-a8932f1fe9cf · outbound

This paper cites GR-3 Technical Report.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations GR-3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.202806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.202806Z digest=sha256:1b31d35dba398b20e4de8b3092fcb156128a140db4962289c5b4fb5fd6a65e77

Observation 9567883d-a4c2-482c-9c37-5382cb4a54d1 · outbound

This paper cites Berkeley UR5 demonstration dataset.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Berkeley UR5 demonstration dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.371084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.371084Z digest=sha256:e8d520ecdb42d2211988f3c2f1ef02d7bc9504c9148e374e61a429b4268c20b9

Observation 3295f24a-8ad5-454f-b11f-46c65c8e7369 · outbound

This paper cites Playfusion: Skill acquisition via diffusion from language-annotated play.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Playfusion: Skill acquisition via diffusion from language-annotated play

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.544842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.544842Z digest=sha256:89b0c6dde2dc227782a825ab5d2eb3c7bd7fa528c51c2da8b520e1bb40394c66

Observation 3bd6deab-5954-4d80-82e8-6a53268f3591 · outbound

This paper cites Moto: Latent motion token as the bridging language for robot manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Moto: Latent motion token as the bridging language for robot manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.698063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.698063Z digest=sha256:aac976acf4c5f95ec1868246e620859df6137df6c8f8822692439490c4f051d3

Observation 3f72c60d-595c-498c-a80e-14b4ad91529a · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusion policy: Visuomotor policy learning via action diffusion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.806309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.806309Z digest=sha256:c9f03dc80796646f3d9c08e58a1f6f4ada3c28f222457aa3c1f2e66b348294e7

Observation 79630e2c-e1e6-44b3-8fde-3bb04aaaa097 · outbound

This paper cites From play to policy: Conditional behavior generation from uncurated robot data.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations From play to policy: Conditional behavior generation from uncurated robot data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:02.978234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:02.978234Z digest=sha256:6c2756922a4b967cf2be9181f840a14b861400838a3f0e41a0a705101fe6784c

Observation 2f4da7b4-5811-4065-8228-8b03bfa745b8 · outbound

This paper cites an unresolved cited work.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.136948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.136948Z digest=sha256:741affc89969a3dc436773e0c724ecf8e3383039b1ebaa15be5f560204871a7d

Observation bdad1294-955c-4333-8c04-114a23788008 · outbound

This paper cites an unresolved cited work.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.258438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.258438Z digest=sha256:9c97eb3425832989f697c61c08c0e3af972cdfb217d68d2574e1c9008478e750

Observation 04134676-4123-45f0-9d52-4d832ba0a1f8 · outbound

This paper cites Learning universal policies via text-guided video generation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning universal policies via text-guided video generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.386238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.386238Z digest=sha256:46bdd52e12ba2960e7647a2381909ec504e3261f5389b559a8a2eae574fb34e2

Observation 85cfa95a-1cb6-4a1c-82ba-b6d547c63b9e · outbound

This paper cites Bridge data: Boosting generalization of robotic skills with cross-domain datasets.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bridge data: Boosting generalization of robotic skills with cross-domain datasets

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.513094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.513094Z digest=sha256:1545b4cede2c575cc6c2898f237b959f661551d0f2958fa15437d3640fb914a6

Observation 51c0c563-8e68-481c-b63b-738e13460692 · outbound

This paper cites Diffusion trajectory-guided policy for long-horizon robot manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusion trajectory-guided policy for long-horizon robot manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.628278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.628278Z digest=sha256:679155ab656ba220c8c29bf8733002f6dddba5665bdb591188c308e92e048c26

Observation 33b15971-601d-4791-9baa-676298fb8784 · outbound

This paper cites Finetuning offline world models in the real world.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Finetuning offline world models in the real world

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.748331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.748331Z digest=sha256:35eee9cec6028334fa47b109b39f4956ad3d5d731fd5d4c16ef4f10275114ed4

Observation b45a439b-ccd7-4c09-9838-b8abd4724564 · outbound

This paper cites Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.869335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.869335Z digest=sha256:cea17396a10499f9d82525194ed8ee64b83d7ae55a7898ef3cde2ba3ea8fea6b

Observation ad50cffe-bb10-4229-9558-b830b5be44ae · outbound

This paper cites Llama-adapter v2: Parameter-efficient visual instruction model.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Llama-adapter v2: Parameter-efficient visual instruction model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:03.986449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:03.986449Z digest=sha256:5d34e592bec14ce0f927953e3aebe58db04aa8d7da307ad24a1d7d2d7c746296

Observation 9f04a714-c3e4-4b15-9d6d-0826cfa7699e · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Ego4d: Around the world in 3,000 hours of egocentric video

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.106057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.106057Z digest=sha256:c4192bec73b08cd3540bfae5e137bcdcc3eb3e67f8e29b6eb70f138db68e0691

Observation 22d2c949-0b2a-466a-b74b-c3dce749f5cd · outbound

This paper cites Watch and match: Supercharging imitation with regularized optimal transport.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Watch and match: Supercharging imitation with regularized optimal transport

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.226810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.226810Z digest=sha256:64632b6dcfcc4e42eae71eea9352bd8410f4eede0fa546ce15e8b25bf293c230

Observation 02a524c1-1e7c-4f6f-86a5-f1a3281e72aa · outbound

This paper cites Learning an actionable discrete diffusion policy via large-scale actionless video pre-training.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning an actionable discrete diffusion policy via large-scale actionless video pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.347887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.347887Z digest=sha256:48f4a949cb6113c320797e3c7a81987291f62dd3f788716eea698f4d35ca9e02

Observation e5204f64-f07c-4c4a-8364-7af94f545343 · outbound

This paper cites Masked autoencoders are scalable vision learners.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Masked autoencoders are scalable vision learners

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.504436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.504436Z digest=sha256:74367e16eccc3b07dc74f10f58ad7e3d6fe5394969f21a501dc4aa6fbcccb082

Observation 71a112a8-1baf-4317-a256-a0edf5bfd8a1 · outbound

This paper cites Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.615577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.615577Z digest=sha256:3d379abf4bdccc9d8b12c88192f8c381e47dfbb7e0c477ff35a3ff22efcac9a0

Observation 0db34190-8e2f-4bba-a705-4133b6d10262 · outbound

This paper cites Sacson: Scalable autonomous control for social navigation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Sacson: Scalable autonomous control for social navigation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.634366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.634366Z digest=sha256:c1b6304ffb54f5be7fbe7532a689cccd9868ff8fa4382500338735b6ed54507c

Observation 05bd37ec-d88f-446b-8827-8ff991538b51 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.640373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.640373Z digest=sha256:03612c193301a6ee6ced1c3378edc8fd03adfb5b0ddfacf687e9c70a139aa93d

Observation 8734c05c-1182-44e2-9a35-0fa48370b3ec · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.706491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.706491Z digest=sha256:e5a258443521248ffed8e27a13ad2fa16cf72cb51fef02926a63153e978ea9de

Observation e9458080-7f50-4add-b098-02cbd520b0e8 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:04.841290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:04.841290Z digest=sha256:5fd6584c51af6b41d1e22f17d0f9914594070bc28efd7afa9d40dea7ce897efa

Observation 00ed9004-e513-4d93-b29e-16648d32ab3e · outbound

This paper cites Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.004982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.004982Z digest=sha256:9bdd2080804ace013c89193f75a0dde843bfc6b375b2f16b73b4e5472befef4f

Observation b52e4d07-f359-490a-8300-c46515766fe6 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.127713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.127713Z digest=sha256:50a44c9972df6159347f08547832eaa761b373cb7542117e3c51880683c913ea

Observation 9e94f795-4a51-431f-82cf-55bf1fb36c1d · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.190823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.190823Z digest=sha256:e93a9d5792351f9fb0fab5c0354cb48920d47923c40e6ff9126eb3ef900b75ce

Observation 62f3b0ac-7986-46f0-a73d-bd727356852f · outbound

This paper cites Pre-and post-contact policy decomposition for non-prehensile manipulation with zero-shot sim-to-real transfer.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Pre-and post-contact policy decomposition for non-prehensile manipulation with zero-shot sim-to-real transfer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.253373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.253373Z digest=sha256:b5ca9891aa3e8e82cdcbdfdec301e69e4071b9c3aaa219d49010883e99f6cd00

Observation 6c41af8c-e6fc-443a-b0e9-155dd9149553 · outbound

This paper cites Openvla: An open-source vision-language-action model.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Openvla: An open-source vision-language-action model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.339723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.339723Z digest=sha256:62be49e1d60dfe8b786d2617e0b773a7a58a9a21c20cfa58dfd0a805417ab9c2

Observation ca69b201-2f52-43d8-b5ba-17cc8db19fa2 · outbound

This paper cites Molmoact: Action reasoning models that can reason in space.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Molmoact: Action reasoning models that can reason in space

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.470578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.470578Z digest=sha256:6fc892a7ca986cbc3ad6db3da19907a929eb73d3914dd3c87b1e753fa751d8b3

Observation 9140fdce-3540-421c-b45a-f6981d7864ed · outbound

This paper cites Behavior generation with latent actions.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Behavior generation with latent actions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.596613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.596613Z digest=sha256:1954ed2504d093052ff2e851ba8b03e647f7da8ff62167a4ef1cf9a7457b9675

Observation afe9b4a1-4422-440b-bdb0-dd7bcc82efbc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.731876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.731876Z digest=sha256:b3e9e04c46b56d9b98a82d982c904aae2ebe6e6b0d88dcad25a44e1de7d77758

Observation 5284e8cb-b687-4602-b185-1511a45cc04a · outbound

This paper cites SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.861028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.861028Z digest=sha256:4f858cfad428a4ee7abb0b609c49e17666f70c318fb7bae2f382937754879754

Observation 6230c424-1809-4676-b3b9-9ae333c0935e · outbound

This paper cites Unified video action model.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Unified video action model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:05.978753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:05.978753Z digest=sha256:bcc6850f6aacf42d6f83ecd79b826b84220c4c6163f530c5aafe57edfcfde723

Observation 137ebdf4-0baa-4655-8017-cf376523f7cc · outbound

This paper cites Visual instruction tuning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Visual instruction tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.132423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.132423Z digest=sha256:ebab00ec5c8f227fb05e3b69506875185a552126ad8023954564bc1b6d6cfe1f

Observation 5e597d47-7142-4a7d-94e4-2257f2e5c1d3 · outbound

This paper cites Robot learning on the job: Human-in-the-loop autonomy and learning during deployment.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robot learning on the job: Human-in-the-loop autonomy and learning during deployment

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.265230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.265230Z digest=sha256:cfa7bd886d0c4ca26a4cf44ddd6d91bb8985e20451162ba1639ffa71b0a4a934

Observation ac21cc6c-f1ec-4bc4-91b7-a4036582c6bc · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.438010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.438010Z digest=sha256:1e8a831386ba10de809edbabc5f695c89ed8528c71b5e0a9c6a083640887e317

Observation 40decb04-7570-4192-a8eb-534ffa2bed53 · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.563735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.563735Z digest=sha256:ce04e2d6c2a47d63f468bcdf593371f6820a4f1487b4228437001b1e626f448f

Observation e71ea9ee-6d4f-44da-9de7-5d36afc2ef8a · outbound

This paper cites Mla: A multisensory language-action model for multimodal understanding and forecasting in robotic manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mla: A multisensory language-action model for multimodal understanding and forecasting in robotic manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.687191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.687191Z digest=sha256:70f7ea1dba7d6e0c81c7d30b392e8a36f72fa645efbcb8714a06a1f2f00f4798

Observation c80e09eb-6a4c-470c-8664-5da4b8e13c5a · outbound

This paper cites Multistage cable routing through hierarchical imitation learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Multistage cable routing through hierarchical imitation learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.855578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.855578Z digest=sha256:71a2f39f18c06fa142616c05c6e02a1cc7e4ace82a5c3c24731d3fbc1becf09a

Observation 02d60139-4fbd-4752-95f8-0662f9078d03 · outbound

This paper cites Fmb: a functional manipulation benchmark for generalizable robotic learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Fmb: a functional manipulation benchmark for generalizable robotic learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:06.995250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:06.995250Z digest=sha256:7f0db75c84ce4d13b11f360ccb54505d819e73c4bf703dbb6d17b76de2faf331

Observation f18ec6b9-1c8c-424b-995e-3b587543e078 · outbound

This paper cites Interactive language: Talking to robots in real time.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Interactive language: Talking to robots in real time

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.112179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.112179Z digest=sha256:1db366614e0dc8fba2de8f6b49a8307f8ddfdfbccd8c578a668a56424cfe9848

Observation 1cb58e84-063f-42dc-8106-37bca937be25 · outbound

This paper cites Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.238888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.238888Z digest=sha256:66bc04bcfab22822f3b3e42e1877c129803da354a721414d1992f0f1581bab8b

Observation 46874027-59bb-467f-a1b4-8e3a0cee3259 · outbound

This paper cites Weblab xarm dataset, 2023.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Weblab xarm dataset, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.363956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.363956Z digest=sha256:c4733abda6e56ef160500da0380cce68f45faf636e4959e49c9b60e07cc6b659

Observation d7c98143-8b0d-4e48-98ea-b5d444e02662 · outbound

This paper cites Grounding language with visual affordances over unstructured data.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Grounding language with visual affordances over unstructured data

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.555878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.555878Z digest=sha256:c6e825d0a737d7652025a2ce0264c031798568983baea7a2348f7c43b7e9c6e6

Observation 031cffd1-bd7f-4110-85ed-da9abdbfe50c · outbound

This paper cites Structured world models from human videos.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Structured world models from human videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.671505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.671505Z digest=sha256:6cdaa5274d2c0c875f4a5e5890a9361387c473fceba06b9e2d2ac3e9057290b2

Observation 1fe710df-da85-4924-ad5e-2f817067b96d · outbound

This paper cites Quest: Self-supervised skill abstractions for learning continuous control.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Quest: Self-supervised skill abstractions for learning continuous control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.787613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.787613Z digest=sha256:69ac0ea369e6bad5db15a720d7549958330fb295c1b6b58af43e5ca4cb1b527e

Observation e75dc6fd-c2ac-4cc1-86d2-a64d4bd57b81 · outbound

This paper cites Learning and retrieval from prior data for skill-based imitation learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Learning and retrieval from prior data for skill-based imitation learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:07.914373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:07.914373Z digest=sha256:e26444e8e3b74c9b0764e335fde327fd32c23bf632d913339606587d02132346

Observation 446a9553-f112-410e-88ea-df27b967d021 · outbound

This paper cites X-embodiment u-tokyo pr2 datasets.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations X-embodiment u-tokyo pr2 datasets

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.045871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.045871Z digest=sha256:21809bcd21688f0961a3678deb35b1c737c1824bb2ac7a37864cb9d3ad1a6bc8

Observation 7e547628-86e6-4881-8229-76a0cab33be7 · outbound

This paper cites Motion planning by learning the solution manifold in trajectory optimization.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Motion planning by learning the solution manifold in trajectory optimization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.209916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.209916Z digest=sha256:cfea4bf0e6f352aef5e91d7be4b295d998cb19b9551001624490804661e8bd90

Observation b537d90d-2f9c-4d83-89e1-3e4b24a93e4b · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.344640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.344640Z digest=sha256:187296cd6d3fc55f6fefbc3f2e74c7b79eb6d51fba4bb4ff897030236b6e33d6

Observation 721772f7-7d4c-42bb-9ab0-cd457790e4a8 · outbound

This paper cites A guided reinforcement learning approach using shared control templates for learning manipulation skills in the real world.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations A guided reinforcement learning approach using shared control templates for learning manipulation skills in the real world

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.533185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.533185Z digest=sha256:8a1d177bbc0194fcd341ad83c5a4c1c5b548cd071da6ad4410bfa7d9c1d2bea7

Observation 26060803-b9b7-42e2-8a50-e309d76e1f15 · outbound

This paper cites Guiding reinforcement learning with shared control templates.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Guiding reinforcement learning with shared control templates

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.678685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.678685Z digest=sha256:7d79fe6f7b840ce81801f6fe9309577d70972241680d0fbe08b346a2f90ef662

Observation 9e03fce4-c705-49dc-b7e1-2cdaee5b4a02 · outbound

This paper cites The surprising effectiveness of representation learning for visual imitation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations The surprising effectiveness of representation learning for visual imitation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:08.813052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:08.813052Z digest=sha256:4983b44bec6643ae9c7012de8584968c6ebf467f09e2161fd171e21028a62626

Observation 119428c6-2dad-4f44-9c28-c2f66df48c30 · outbound

This paper cites Supramodal and cross-modal representations of working memory in higher-order cortex.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Supramodal and cross-modal representations of working memory in higher-order cortex

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.024235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.024235Z digest=sha256:8f572c2547497a56896c5d90473e8307bd4b99ab16d5e4735fb9daa889a6d0e6

Observation ccc4743e-107b-49ed-8f56-38945ae5f088 · outbound

This paper cites Embodied artificial intelligence: Trends and challenges.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Embodied artificial intelligence: Trends and challenges

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.208385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.208385Z digest=sha256:8454cd2a82e81c8e3f9a82898433f9978c6742f6f85e02d544457f4e9fa8b892

Observation e9e23507-f215-4a1e-83d9-44547ac2375d · outbound

This paper cites Spatialvla: Exploring spatial representations for visual-language-action model.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Spatialvla: Exploring spatial representations for visual-language-action model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.340867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.340867Z digest=sha256:ea375fac1d3c17f3418f4bd80d9849e2b5a997b23f2acb5793371966fb517e27

Observation 32460c80-671d-4a3e-9b6f-32ce07776ea1 · outbound

This paper cites Shared control templates for assistive robotics.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Shared control templates for assistive robotics

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.512222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.512222Z digest=sha256:2856f5822c91e39bc2da68e2d9d05c8abc16bd3bfd05125a6fc5698b114961b1

Observation 11d59140-da02-486c-bf2f-76d9bda11b0f · outbound

This paper cites Robot learning with sensorimotor pre-training.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robot learning with sensorimotor pre-training

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.571767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.571767Z digest=sha256:86997f6e195ee663f1d43405c502e86bfa7c6fc4956a8cd48fe7fd7ec0a57d04

Observation a3cb414f-ae7d-4197-aeaf-8681cc03e89f · outbound

This paper cites Real-world robot learning with masked visual pre-training.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Real-world robot learning with masked visual pre-training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.673256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.673256Z digest=sha256:707633e21de92d8ddc8c1c7b3a29d6acc6a7d3f374cc98d3ee10fc9f0422d0f2

Observation 594eeb26-8057-40e5-9503-30ac30c1bb75 · outbound

This paper cites Latent plans for task-agnostic offline reinforcement learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent plans for task-agnostic offline reinforcement learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.812719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.812719Z digest=sha256:6df92319c97c7a284b1e546213230832eda86ff7ef5da72dcd97aad9802a8aa6

Observation 9fa28f33-5aa4-444c-9327-e4cda65d4a72 · outbound

This paper cites Multi-resolution sensing for real-time control with vision-language models.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Multi-resolution sensing for real-time control with vision-language models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.886517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.886517Z digest=sha256:cfc3a0433db154e3548593d3574ac2f244b73597706b46391295aa0d4f316336

Observation 95f86456-c115-4c7d-afb7-54fbf8a6e623 · outbound

This paper cites Behavior transformers: Cloning k modes with one stone.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Behavior transformers: Cloning k modes with one stone

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:09.973268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:09.973268Z digest=sha256:288137626a0f68cd2d7e570abdde74539d780c32003a007c5b7e00c42c4c3561

Observation 04dc17a6-6f36-47c0-b98f-f66ada632dec · outbound

This paper cites On Bringing Robots Home.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations On Bringing Robots Home

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.046100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.046100Z digest=sha256:a29f53e440410a983c8da873c2b7ec9c06c092c003cca1e072fa18ce26b30cc0

Observation a16ff2bf-65c2-4c28-8bdd-711f1dc82224 · outbound

This paper cites Rapid exploration for open-world navigation with latent goal models.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Rapid exploration for open-world navigation with latent goal models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.105439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.105439Z digest=sha256:62896d2f58d449ee6a2381fc42eb8c7222d2ae1493b9542f8972caade809fe6f

Observation 9fc10689-a07e-48dd-9ad7-15260452aa0b · outbound

This paper cites Mutex: Learning unified policies from multimodal task specifications.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Mutex: Learning unified policies from multimodal task specifications

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.181018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.181018Z digest=sha256:dfa4a2e23a29ced653caa73bac7a9d11b5bd47891aac8e042ddbcf1a5e224809

Observation 355ce202-6c60-4ac3-931c-265353643d67 · outbound

This paper cites Dense policy: Bidirectional autoregressive learning of actions.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Dense policy: Bidirectional autoregressive learning of actions

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.253909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.253909Z digest=sha256:c76c0a28eddb08ddd179e095785f771cc834b29738f1c45c0bc95eefd40c9045

Observation fb667ca3-93e9-4f27-a5e8-c68b890692eb · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Gemma: Open Models Based on Gemini Research and Technology

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.403529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.403529Z digest=sha256:14901f40b9d09471ca727df51e418222fa72917e43ce61c75f9fb7a4dc39a5fb

Observation a36e2307-7873-4fc9-ae11-84f95fba208a · outbound

This paper cites Octo: An open-source generalist robot policy.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Octo: An open-source generalist robot policy

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.493855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.493855Z digest=sha256:60fd4ff26512d5521b4bcaf7cf26a016cd0b950de55ceea746bb4cc4fa30cf96

Observation 1337f9f6-4bcb-46d6-aa06-d2a62a2b153b · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations WaveNet: A Generative Model for Raw Audio

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.557966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.557966Z digest=sha256:b9abccd015b59fb85c55521a6e55825b01fe247492bb49629a286fc55a89ef3f

Observation 9585b2ab-8201-4936-9b02-bc02407b0284 · outbound

This paper cites Neural discrete representation learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Neural discrete representation learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.634298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.634298Z digest=sha256:2194149e5bcab0cc9cab98c2d9dfcda33eb21dfb857a87aa6d80c45638254e6a

Observation 8c37cfe8-27a7-4995-a3e1-cea88676ab5a · outbound

This paper cites o rn Vogel, Annette Hagengruber, Maged Iskandar, Gabriel Quere, Ulrike Leipscher, Samuel Bustamante, Alexander Dietrich, Hannes H \.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations o rn Vogel, Annette Hagengruber, Maged Iskandar, Gabriel Quere, Ulrike Leipscher, Samuel Bustamante, Alexander Dietrich, Hannes H \

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.710015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.710015Z digest=sha256:857638aa9632fa35166dbaf2e5edd456e9e581db22a9105958599afb7dcbe594

Observation 8c5c6a59-3c93-430c-8b80-aa370fc9cf32 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Bridgedata v2: A dataset for robot learning at scale

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.808096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.808096Z digest=sha256:e93f6174bc9155d60a9a5b903dcd0dddd995b9ab314de120eb0d838779ecede7

Observation 11e48d89-d86b-4fcf-855e-d8e723f98961 · outbound

This paper cites Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Scaling proprioceptive-visual learning with heterogeneous pre-trained transformers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:10.909954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:10.909954Z digest=sha256:89281f557d8b2becd1add994a579e247ac642025bcdee66c60bf547477db0baa

Observation 218c11ce-91c9-447b-a012-089f35b0b18c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.033236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.033236Z digest=sha256:92efa6fc955f4b09f775f82891c10fb195d3d5590bc67d032dde9cd6a4cea1b2

Observation 95c611ed-4a5f-4560-a3b5-845630f0e112 · outbound

This paper cites Dexvla: Vision-language model with plug-in diffusion expert for general robot control.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Dexvla: Vision-language model with plug-in diffusion expert for general robot control

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.139743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.139743Z digest=sha256:99981c5de038c792f0d9cac11b8dca2c6b5456ddbe01f6605493a05f784db144

Observation 6d67c732-7a4a-4e6f-939c-f55d671d6799 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.246567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.246567Z digest=sha256:057a3b416f798317e17b65f1017b0e60df20d1bba1d6db84a08f5dbd17f148e1

Observation d73ce338-90dd-444f-aff9-77fc37c0a5ad · outbound

This paper cites Diffusionvla: Scaling robot foundation models via unified diffusion and autoregression.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Diffusionvla: Scaling robot foundation models via unified diffusion and autoregression

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.372607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.372607Z digest=sha256:a8955639bd02460ee5b748813339f2e9e314567261406670afb0eaa805bce1b1

Observation 759101fa-7f2b-426a-988d-43d1bcc6909a · outbound

This paper cites Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.473471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.473471Z digest=sha256:91dbc0cd2f5154554f6b6ab613b9768ea90d79c9e5b7db86cc9bba4e409747ea

Observation e9fee724-7114-45e7-ae48-644538fc7b9f · outbound

This paper cites Discrete policy: Learning disentangled action space for multi-task robotic manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Discrete policy: Learning disentangled action space for multi-task robotic manipulation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.559465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.559465Z digest=sha256:07d0576aa44046d8b33c3d24b5573b8770a5077058bf7f89e9021bd96c0c9155

Observation dfa98af1-187f-41ee-8774-769f5e8ed77a · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.671428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.671428Z digest=sha256:5e83b98ae5722ddb11d84dc28f82b54c90db4216267aec120056a37c8b4346ac

Observation d933d7bb-30bb-49b3-b79c-50967c8e9e05 · outbound

This paper cites Latent Diffusion Planning for Imitation Learning.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent Diffusion Planning for Imitation Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.746390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.746390Z digest=sha256:4c182195b3a67b00187af9807f2f149a6b7f857cd85e72d9c005927ba08c39de

Observation 142850b9-fe0d-4e56-a3cd-16e5ed750074 · outbound

This paper cites ucsd kitchens Dataset.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations ucsd kitchens Dataset

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.861722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.861722Z digest=sha256:4504f6213e4c466375975f9e8edd68ba54d0f98b10c7100bbc6bd5a68ead91ba

Observation b261577e-4f36-4937-9711-42cb1623578c · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Instructvla: Vision-language-action instruction tuning from understanding to manipulation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:11.924637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:11.924637Z digest=sha256:6cc07b49c05937b908a3d7985a24fda735f159293fff8f9d8fb0c83872fc19a3

Observation 3ef1a4a8-df81-4eb6-bc88-5266618f1ba3 · outbound

This paper cites Latent action pretraining from videos.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Latent action pretraining from videos

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:12.041655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:12.041655Z digest=sha256:4da6cdff2a04ec3dd101db4ba2d0bbe37d35c8686a21b0e2cf0792933610674c

Observation 7e37b523-ce46-4a6f-ba56-44be7cc1bb49 · outbound

This paper cites From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:12.136184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:12.136184Z digest=sha256:008aaf63a87d3cb445f1a8aefcdf23fa5cf42b845aa41b1c659608206c715cba

Observation 5e4fcab4-40d5-4179-98f0-68524b9f79c6 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T00:10:12.234233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:10:12.234233Z digest=sha256:448018113a770eaab55b3483de3ccecb7f252f7bfb65da3f5dea3e2145cad9f3

Pith citing papers

Observation 917c4e4a-4397-44a5-b114-89a12a9de230 · inbound

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling cites this paper.

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:08.181865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:04:25.672157Z digest=sha256:e67a9b8f6f82ca6609bd1c276f9d24e852902cd05db1d9ffc25ad77a13626411

Observation f3a3171d-600d-4eab-99b5-c195e1c9bc09 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:08.181865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:ff1a75935c52427176db8f601b158055fb2df9715426c943be216221befa98c8

Observation 8244850d-5b03-4c4e-8aed-2bc4a374edd2 · inbound

Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation cites this paper.

Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:49:35.628764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T04:45:47.109272Z digest=sha256:0b75f47721c373be5a8f8a3840292ab821aba89807e7d8efcd7242a84f99495a

Observation d20ff8bd-83b4-46dd-aad2-b09b0ff8d722 · inbound

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models cites this paper.

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:24:04.228645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T00:18:33.150920Z digest=sha256:8a6f0f3bf34f81a7fd66a6b39700b62679028e952c68b6dfd5d5bda0cb8e219a

Observation caea0d14-8334-4eaa-b796-099b79db7b7b · inbound

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots cites this paper.

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T16:55:51.384865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T04:23:04.622902Z digest=sha256:4892f1a3d96efc7f7f0cb0beeeff301c9f8525399c69c5ed5c259f27898b5392

Observation 1e9949d2-8952-42b1-9ab0-353e7516db53 · inbound

GeoProp: Grounding Robot State in Vision for Generalist Manipulation cites this paper.

GeoProp: Grounding Robot State in Vision for Generalist Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:16:29.418832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T20:12:59.778086Z digest=sha256:c1d6ede0ffbfc07640c6316ed53f8ce034f2b66ba33c09836ff4a428631657c1

Observation b92120f6-af3b-4f60-816f-d59befd8112a · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:26.185245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:26.185245Z digest=sha256:4f8d2388fab67169ae58902aada42d27481c19b5f1adb9d800914ea3fd27a5f4

Observation ab5949a4-a94d-45ae-b240-7ad55ab83bd6 · inbound

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models cites this paper.

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:00:31.953138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:00:31.953138Z digest=sha256:29c55d099b354ec55fd7ffcf60f173ffaeb83209b815b51994ec86f81e127491