Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:55:22.498055Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2605.18287.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:55:22.498055Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:36:25.569578Z
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 71955047-ede7-418f-956f-a69f5c026ee9 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Alemi, Ian Fischer, Joshua V
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2bd32c82-ab67-48d2-9d22-9f659fc72a3a · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Xcit: Cross-covariance image transformers.Advances in neural information processing systems, 34:20014–20027
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9f5a82e-6a2f-48b0-85f9-800bac36a3fa · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 990407d1-578f-47a0-b539-3a5d27f4b93c · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Are transformers more robust than cnns?Advances in neural information processing systems, 34:26831–26843
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b609cd65-b15f-420e-9a44-2b69ee6d4d8b · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7196ccb4-1309-43c6-aaeb-01500c19e9a0 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b062f4f4-2fce-4efa-b160-99257e6a74f8 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rt-1: Robotics transformer for real-world control at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83284948-60d4-42c9-afb4-05b04ec0c156 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation def13f74-7889-4b80-ab98-f879149070ef · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Diffusion policy: Visuomotor policy learning via action diffusion.Int
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b9209b9-5132-494c-a73d-edfbdafb69b2 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Diffusion policy: Visuomotor policy learning via action diffusion
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0940166a-f009-43ac-9f27-e08f3ed23227 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5ff1fa70-643a-49e0-970b-a18c124ef4ab · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Agibot world colosseum
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 55cbf250-217c-4f5e-a8bd-bf7bbea6cd76 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data HumanNet: Scaling Human-centric Video Learning to One Million Hours
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8817eccc-c6f2-434a-8879-e4e51139355b · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rethinking video generation model for the embodied world
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 89d890d8-b207-4627-b537-23a5b88b1dde · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a96f9eb6-5bc2-4696-b900-3453c0c8184b · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Towards Human-level Intelligence via Human-like Whole-Body Manipulation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 299f678f-ce8d-46fa-9016-45b0968d7cd7 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Benchmarking neural network robustness to common corruptions and perturbations.Proceedings of the International Conference on Learning Representations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 021fbe9e-c686-46ca-8a01-21c98422625a · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data The many faces of robustness: A critical analysis of out-of-distribution generalization.ICCV
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d4c71e2a-bc5d-467e-b756-89ca0afdf78e · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 614019a5-b697-4576-ae34-13d21013dd75 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Prismatic vlms: Investigating the design space of visually-conditioned language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b573aaef-d014-405b-a7f7-d410ffc5e92c · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Heewon Yang, Arefeh Yavary, Tony Z
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d6bb89dc-c2e5-43a6-83b9-8ac0c8c8c214 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data URLhttps://doi.org/10.15607/RSS.2024.XX.120
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 33ad01ef-49d9-4f57-9795-74514f67d4d5 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 797c3b15-6fc1-4ef6-9736-40169f9cef1d · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Fine-tuning vision-language-action models: Optimizing speed and success.CoRR
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7fcc94fd-64f8-4210-be26-44d7d2cb0665 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 190ca374-309c-4295-93ea-5713c30979c9 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Roboflamingo-plus: Fusion of depth and RGB perception with vision-language models for enhanced robotic manipulation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation afa9f423-e9a9-4b0f-b92c-148ece0f333c · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data doi: 10.1109/RCAR65431.2025.11139480
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2a1bb2f-64d7-4887-a4d1-2002365afd76 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e9e2ed33-0cb6-48a2-824c-2ea7970899d5 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 827d637a-0d04-46af-bae8-fcc3215e2e08 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data NVILA: Efficient Frontier Visual Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8155be90-5482-4f7f-81e0-1ac3dbb37368 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4be2b843-fac8-423e-b9ec-575d02854d37 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79fed709-a4b7-4c5e-91f6-2d306b465f78 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Robotwin: Dual-arm robot benchmark with generative digital twins (early version)
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation deebf55b-fe7f-4d3f-bcac-cad89a71ba21 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Open x-embodiment: Robotic learning datasets and rt-x models : Open x-embodiment collaboration
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ae8652a-4947-44bc-8b87-8e5007b23ff7 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data DINOv2: Learning Robust Visual Features without Supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 020750e1-701d-497e-87a9-439dadb69b65 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Vision transformers are robust learners
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d3e5c2c-bc1a-4027-aeb8-ce257eefd99b · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Octo: An Open-Source Generalist Robot Policy
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a766d65-3d94-4b58-846a-8adf51c5bbcc · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data The information bottleneck method
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8fdfb1a7-bafc-4a86-8e06-3b23f2bc4379 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Domain randomization for transferring deep neural networks from simulation to the real world
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0dc466fe-cffc-463a-96b4-5fd8c49b9121 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Augmax: Adversarial composition of random augmentations for robust training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce59cd29-1508-41e0-8a3e-a87b748766a5 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.CoRR
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5598384f-0e68-44ba-856f-3f83ebada8ea · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66ab4261-8476-4cc3-8f93-abdc1f35d5ed · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Robotic control via embodied chain-of-thought reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3768ab48-191b-4ac6-ab30-a3e3e200378e · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sigmoid loss for language image pre-training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d136d2bd-3abc-4304-ae0c-5f63586de58e · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ee54499-56bf-4f3c-9546-7c808168c652 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Understanding the robustness in vision transformers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3abfb6fc-90ee-4478-9f25-e4de7e526630 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 325e311d-d63c-49d2-8e8d-b8a322ccdeef · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rt-2: Vision-language-action models transfer web knowledge to robotic control
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 904c6979-8470-4e34-8079-cf713816dfbb · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b5363f7-b54c-487e-87fb-9c43a9713948 · outbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b093e008-a336-4fdf-a256-0eca4df13610 · inbound
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.