Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:28.461260Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.11214.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:28.461260Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:56:22.911558Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T20:28:16.078786Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2d23df2-6129-4393-ac16-292debc961be · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Training language models to follow instructions with human feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cc614d-4ae6-4aea-a45b-8b0db30e8e83 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9837886e-a052-44b9-884f-cfd8efcfc023 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Gemini: A Family of Highly Capable Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b3bb3a-32b5-4421-9783-798e0704ef5a · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Visual Instruction Tuning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 75a71d25-1387-4dfc-9d95-8b5bb733ce80 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions BLIP -2: Bootstrapping Language - Image Pre -training with Frozen Image Encoders and Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 67836b1e-f353-4f6f-b6ee-27d5906336a7 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895c3502-6c74-4c1e-96ed-243edc68a864 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Prismatic VLMs : Investigating the Design Space of Visually - Conditioned Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5f536b50-b7c1-4291-a734-6928c560abe1 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LL a MA -adapter: Efficient fine-tuning of large language models with zero-initialized attention
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a2129233-259e-4803-8872-3cfcd773837a · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Sanketi, Grecia Salazar, Michael S
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b2297c5f-afdc-4902-97ae-1f19e6aa7f4c · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62778bcf-76fa-4b56-b216-7f24a0b45018 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions OpenVLA: An Open-Source Vision-Language-Action Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e2d3e5-449b-4db8-835b-5c6554cb432c · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RT -1: Robotics Transformer for Real - World Control at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59a5527-a746-42fc-a6ec-037925766b08 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Language Conditioned Imitation Learning Over Unstructured Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e5cd03-7cfa-4a63-bf53-636ac935f2f2 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Vision- Language Foundation Models as Effective Robot Imitators
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 74e367ab-0ddd-42bc-a9ba-b650cf217d50 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35913785-6f73-4528-b94a-e1c861a9f75f · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RDT -1b: a diffusion foundation model for bimanual manipulation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6d0cbb99-6dac-4f1f-a239-d73381e74d3c · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions What Matters in Building Vision-Language-Action Models for Generalist Robots
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fee3452-f08e-4d51-9183-496992f53f19 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a7a0a9-4637-46e2-9892-ee2983bb4cba · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705c1280-a39b-4398-a971-5b54c75c5ba9 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 663d3148-ceb7-48e1-8f76-4a48bfcdec63 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Diffusion Policy : Visuomotor Policy Learning via Action Diffusion
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91034c97-c35a-454e-82bb-a6b3caef9dd0 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38d7c17-ca60-473a-ab00-33601141c25b · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions 3D - VLA : A 3D Vision - Language - Action Generative World Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 504d44d2-dd21-4996-ad3c-eb1fd027f004 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Dream to manipulate: Compositional world models empowering robot imitation learning with imagination
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d92e8050-a612-454c-88e3-0b9330c8fa52 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140de4da-d8d0-4176-9040-ba8267deab69 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions RT-H: Action Hierarchies Using Language
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f229904-50ff-4ba9-a486-c6d94f54e345 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Robotic control via embodied chain-of-thought reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02fe49ba-6f9b-4a6c-a2af-a7bb1b3a62d2 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions DayDreamer : World Models for Physical Robot Learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 537c13e6-b982-4f62-b34a-a0fd35f8243b · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions VIMA : Robot Manipulation with Multimodal Prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e67f5d9-f7d3-4871-a214-97dd0b4c2fc2 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Interleave- VLA : Enhancing Robot Manipulation with Interleaved Image - Text Instructions , May 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d33075-eab2-4d16-8742-9deac8a80318 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464a12f3-9d26-428f-8762-c3b071891814 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Sigmoid Loss for Language Image Pre - Training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fb430f-7cde-4909-b37e-648acbc336c9 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions An image is worth 16x16 words: Transformers for image recognition at scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2355a54a-c018-4b16-8638-10958b256fa3 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Qwen Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e364c4d-9564-4082-bdab-640e6b79ba5c · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05ca3a6-9c4e-41f1-b1f9-5f3e2a272c95 · outbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2730981-4504-4ffa-9c78-9c5e3e9f836d · inbound
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e9948e22-77e3-4aba-b0a1-d6a0661eb9dd · inbound
CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904c3149-d11a-41bc-9d72-902213218bb8 · inbound
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.