Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:42:06.146713Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 14 inbound Pith citation observations for arXiv:2506.09930.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:42:06.146713Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:39:40.204306Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
47 of 47 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation c27d9800-1868-4362-86fc-6cdad0e5b7b0 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b7dd7e-7714-4b6d-9c2a-fd9a277ec88b · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4bcf23-5d93-4151-9601-6a5045fa80a3 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b6f62db-9986-4bc5-904f-cb8758868817 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1c29cf-2742-466e-9b41-3ef9f52be5be · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Language models are few-shot learn- ers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91e20cde-cd8d-4722-8c14-d6fb38401d96 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341cf704-859f-40e6-a05b-41851bd41b86 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Pali: A jointly-scaled multilingual language-image model
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c747118b-51f4-45ef-9502-97fea4230882 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc69081c-67bc-45d0-8d79-3631fa2d1980 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d963e5b-7976-44ec-a70f-bf6c8bd322ad · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models BERT: Pre-training of deep bidirectional transformers for language understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7f056c45-23d6-4dce-b2af-3fa2239fa884 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba37c4f1-fd83-4a19-aea7-bc0a26fc0e0e · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Poor performance of openvla on bridge
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 44f349b3-9c3e-4c41-8ad6-84a9a674b104 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Gemini, 2023
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd51dde9-1c76-4a58-a908-dce00f64396b · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Maniskill2: A unified benchmark for generalizable manipulation skills
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 96adb65c-2cc8-4e41-8c87-0cedfbba00c6 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0532ff28-dd99-4a5c-8c30-6d2d8e0bca1c · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cbc4315-164b-4170-b50f-3d0196a42f86 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72427009-033b-452a-a48c-2ea4e9ea4ae3 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6077cad2-4627-44b2-84c5-80f9bd9a672f · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Open- vla: An open-source vision-language-action model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c25dea5-686f-4947-be56-be29b27d9155 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Evaluating Real-World Robot Manipulation Policies in Simulation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5f7fbd-95d5-4e82-9b39-1ffe23eaf7ce · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Flow Matching for Generative Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5366ed24-bd34-4772-91b0-b58102627c61 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd3e924-dcb6-4e59-84dd-32d9efb42b3f · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Visual instruction tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3a825f77-8c62-4827-a63d-fae94d00ac9d · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Rectified Flow: A Marginal Preserving Approach to Optimal Transport
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea21fac9-c7c1-4fc8-80f2-d937895f41d5 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d4a076-07f3-4d94-bb82-cfa8c353b254 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 156eb4df-4266-4f88-a898-01374b2dda86 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c66e6c8-1674-4f14-8461-7089d1f0e2c4 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models RoboManipBaselines, December
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6bf085e6-01f9-4da8-8d34-e10d1caae89b · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Octo: An open-source generalist robot policy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5a7a5c19-55b3-4716-97aa-d25409227988 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models gpt4o, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16c5c5a6-c9ea-457d-b02c-a96bcbe279f6 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models GPT-4o System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc60aae-10b8-4c33-9b05-dd9d98415e27 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 66c08d11-835d-4732-b7d2-f8eac72be357 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Transfer between Modalities with MetaQueries
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4f464ad-4aed-4261-9491-5e8dfbcfba1f · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7599869-dcfc-4e08-b4bc-b0bbdd366888 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63153e01-3e4f-4536-9664-d8ae44c94a4b · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Learning transferable visual models from natural language supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 928ae829-0606-4ba7-8e14-fcba1b06b291 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 952a6701-c7ec-479c-9ea3-1305e37ac407 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Gemini Robotics: Bringing AI into the Physical World
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a804a151-5955-4284-84e7-9cc7f52a29c2 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f7feb02-84ed-478e-b985-553503aef950 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Metamorph: Multimodal under- standing and generation via instruction tuning, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b2fa300-de27-470b-9a94-91a669883642 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Vlm see, robot do: Human demo video to robot action plan via vision language model.arXiv preprint arXiv:2410.08792, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581770ba-6bfe-444b-ad65-cdb3dfebdfae · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Decomposing the generalization gap in imitation learning for visual robotic manipulation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3fc7aefc-af35-4ad5-a4bf-727f9b6e9183 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Magma: A Foundation Model for Multimodal AI Agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95b2b9ed-2033-42bb-9e5c-a4c1aca5bdbe · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Sigmoid loss for language image pre-training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0479ad34-5cb8-45d6-b6a2-9d79d522827f · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74ebdc7-4d61-4f68-980e-a90cdf7e78fc · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models carrot on plate
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d042169b-3098-4228-bdd8-69403383f2f0 · outbound
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07cc7d3e-2ee1-483a-9247-4815e3e92ae3 · inbound
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bab77f63-a90c-4106-a0b7-b3e759dc711e · inbound
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e87e2d-283c-4f7d-bf50-7687c3e09ee7 · inbound
Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 60f0d20b-097b-459f-8ff6-a89a095e4604 · inbound
Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e144412d-e29f-4851-9400-08dafe29f882 · inbound
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 124dc2f3-cafd-47c4-be9a-8f050d0bbdf5 · inbound
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 19ff866d-7b68-434a-99ed-c057353da94d · inbound
AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation afcccc75-7d19-4c99-b990-e9cd5036660f · inbound
APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 42673a39-889b-48e5-933b-039500523b22 · inbound
MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 36726a43-7f1c-4f05-83f9-08bbb7e3535c · inbound
EmbodimentSemantic: A Spatial Scene-Graph Dataset and Benchmark for Vision-Language Models on Embodied Manipulation Trajectories From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7274a360-5f3c-4098-acd1-5d1578617171 · inbound
Robots Acquire Manipulation Skills in Seconds from a Single Human Video From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f466b2ab-6682-4105-9fb0-844986fa41c4 · inbound
RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82189999-d2a1-4aa9-8fbd-d1272d3f771b · inbound
RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78dbd86f-6df9-4060-87f0-f9c36e3b9296 · inbound
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.