Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:00:20.896966Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 10 inbound Pith citation observations for arXiv:2512.21970.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:00:20.896966Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-07T12:28:21.318505Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T12:33:45.452965Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46c1f415-e2de-4090-84b6-4c84f4b9f9bc · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PaliGemma: A versatile 3B VLM for transfer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f57f7164-e780-41dd-b100-7eae2098dd6f · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc92c517-4865-4452-bc7b-95c19d52efc7 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision OpenVLA: An Open-Source Vision-Language-Action Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba07debc-8ef5-439a-8209-c1b2b6bdb2a3 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f730debf-8187-427d-967b-d0a5525fe6cd · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6825f712-3351-4e25-aa62-63d2296ebe0f · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03fd8ee7-6853-4824-ae1f-c8df13e08116 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rvt: Robotic view transformer for 3d object manipulation,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 662f0ae5-f798-498d-9708-7c14a6018916 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D-VLA: A 3D Vision-Language-Action Generative World Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0453c9e-7fcb-4835-9a80-0b48797b0afe · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c12eae7-0785-4216-8443-8ed0825365d4 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4c4c3a-4b9b-426e-968a-47eabbdc959b · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181423dd-04dc-410a-99c6-a4bc3b9677ba · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Foundationstereo: Zero-shot stereo matching,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554549d3-8d71-4939-8218-cd4de7077f43 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Prismatic vlms: Investigating the design space of visually- conditioned language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39029919-013e-4d46-853b-efa44f1a2299 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Palm-e: An embodied multimodal language model,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f73e351-04a5-481c-834b-d809dec0f86a · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43e69d0-a3c6-4f6e-aa8b-bb178e9ae143 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a851f77-53a9-486b-bb2e-7e3a89c5e85a · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Florence-2: Advancing a unified representation for a variety of vision tasks,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382014e7-f9dc-47f3-9831-1d1944bc0ead · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ed9fe0-4171-4711-b7d2-4fd74768c3e7 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea9d4a4-1c36-40ec-9b5f-a6275e93d73d · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca535a6-128e-4486-811c-d410d44d37d1 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rt-2: Vision-language-action models transfer web knowledge to robotic control,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f056f3fc-099d-46af-bbbc-0f95a73d027b · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8884fcb-8c89-4f58-b5b6-ee2dd654999b · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FAST: Efficient Action Tokenization for Vision-Language-Action Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbd01ba-5fc9-444c-8fa7-850c96c17140 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00390d9-18ec-48ce-a580-866cfea3d4fd · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97dccf5f-d256-472f-b89a-48a53538c6ff · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0450714-5eae-4d4f-8f1f-276e0e82d94a · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Internvla-m1: Latent spatial grounding for instruction-following robotic manipulation,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8ee129-bbb4-4602-bcb2-e0b2d8c276ad · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Tinyvla: Toward fast, data-efficient vision-language-action models for robotic manipulation,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c17244c-5f64-4dec-9b5d-a7bd080a0504 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe30e33-4bb9-4f8f-a2f5-80d8c2b5d870 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Generalizable Humanoid Manipulation with 3D Diffusion Policies
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ae4b1e-e783-49a8-8fe4-ca8fa6d8090e · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68888964-eab6-4441-904e-4ebba7a9d12e · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PointVLA: Injecting the 3D World into Vision-Language-Action Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f516bb1-7f71-487f-a83a-b3ad062311da · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b01426-c2c5-4119-a374-e7fc6457bbae · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FP3: A 3D Foundation Policy for Robotic Manipulation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e85bbd-a1e7-4674-b97b-17ba4355c0ab · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Evo-0: Vision- language-action model with implicit spatial understanding,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b14c6b1-069a-4c3f-8d63-d0608475797d · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gp3: A 3d geometry-aware policy with multi-view images for robotic manipulation,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647f1a5d-5aac-435e-a2c3-c5b1b0cdfba5 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning the distribution of er- rors in stereo matching for joint disparity and uncertainty estimation,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f295d6cf-0cf1-4c91-9e22-e1c45540f400 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Cfnet: Cascade and fused cost volume for robust stereo matching,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5d338b-0d2c-4dde-94ff-729c23ab1b40 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Pcw-net: Pyramid combination and warping cost volume for stereo matching,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be59683-b8ed-47d8-b656-0a7af496d967 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Aanet: Adaptive aggregation network for efficient stereo matching,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689ded80-bfb5-4e9a-83a1-2fbaaed4be91 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Arunet: Advancing real-time stereo matching for robotic perception on edge devices,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f029ab-af3e-4d83-b1b9-4943b8887eac · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gfanet: Group fusion aggregation network for real time stereo matching,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a17586-476b-42e5-9d95-d701b592fa3d · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Raft-stereo: Multilevel recurrent field transforms for stereo matching,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b657e44e-d0cd-407a-8b97-780df2ec3e1d · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Iterative geometry encoding volume for stereo matching,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 821d3cfa-00e1-47ac-85f8-27b550f0867b · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Practical stereo matching via cascaded recurrent network with adaptive correlation,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb26f67-0881-472b-849d-e25512166d71 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Uncertainty guided adaptive warping for robust and efficient stereo matching,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d67585-2b8e-409a-8f2f-255dc443fb49 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning intra- view and cross-view geometric knowledge for stereo matching,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536476e5-9dc8-4179-ac95-322f49b11667 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Efficient and hardware-friendly online adaptation for deep stereo depth estimation on embedded robots,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fd4d0a-e8fc-4911-b6fa-302c7204bdc8 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Stereo image- based visual servoing towards feature-based grasping,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c24e98b-43b4-43c6-a2d6-b31742e32627 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52159cb6-e386-4419-bbfb-962ed03dc739 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Simnet: Enabling robust unknown object manipulation from pure synthetic data via stereo,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183fd8e0-b25a-4fcc-a077-b286aba51e48 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision InternLM2 Technical Report
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a010d3a-800f-4e66-b2ae-51f5b95e2bb2 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Flow Matching for Generative Modeling
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60bafb7-2202-42ce-a5fd-18448b6c4d0f · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DINOv2: Learning Robust Visual Features without Supervision
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08702300-1375-4a36-89de-93277c83f09b · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7b73b3f-2007-4ada-b7e0-853868a86f79 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Mujoco: A physics engine for model-based control,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3143be9-2046-47ad-8b82-d46d4be7beb8 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Isaac Sim
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd27a75e-e7af-41b7-9a7c-8dc80099b7c9 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7fb575-bf4f-4a41-918c-8fbd10c5ec33 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Vggt: Visual geometry grounded transformer,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f44e8f0-4a64-4c6a-b38c-f743f13a6696 · outbound
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Depth map prediction from a single image using a multi-scale deep network,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29bcc586-fdd6-4570-a3f1-1810b5506305 · inbound
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff9f8454-ac41-4d72-a36f-9a1ebc423a2c · inbound
MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f5dc25b-ad35-4b5c-8263-6d52775a873b · inbound
MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f801e2ea-02e8-4226-a404-f3f954eb394a · inbound
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e57183a7-40e2-4b37-981a-ad9601eefce0 · inbound
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7561048-f502-4d09-aaf8-2992e5b5d62a · inbound
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d484f26f-42dd-4851-aad3-6662d8f9ab1b · inbound
Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7e3c69d-5d32-4bf0-8d84-b30f4260f15e · inbound
LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1ebb169-2504-4e9d-ba44-7e4680adbf17 · inbound
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation edf36261-3759-4659-977e-f093520ed7b4 · inbound
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.