Pith. sign in

Paper Citation Record · LEDGER

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

As of 5 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 10 inbound Pith citation observations for arXiv:2512.21970.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.21970 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:00:20.896966Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-07T12:28:21.318505Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:33:45.452965Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46c1f415-e2de-4090-84b6-4c84f4b9f9bc · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.737775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.737775Z digest=sha256:1d3d99a5cf6e2ccdbe06ca17aef5d7f847cdd5564d9263d1e9831f7601f42822

Observation f57f7164-e780-41dd-b100-7eae2098dd6f · outbound

This paper cites Qwen2.5-VL Technical Report.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.741580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.741580Z digest=sha256:0f6ffe2b6b6b25da993cb7612f2322474ef632563e833dcd367ab39d0bb38f78

Observation fc92c517-4865-4452-bc7b-95c19d52efc7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.744666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.744666Z digest=sha256:9733381a11f4ebccdf8618732de85b395ae0751082393241777887f8dcbfb962

Observation ba07debc-8ef5-439a-8209-c1b2b6bdb2a3 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.747597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.747597Z digest=sha256:a0ebab36e6f20bc90f6803c7c3528ea2a3b4ae2fd85cb2e40c6f2ed1579852e1

Observation f730debf-8187-427d-967b-d0a5525fe6cd · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.750637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.750637Z digest=sha256:90b2ad6f4e0856a4ea5030f6275290f0e1fa48ff3f1fbc77e66c8d71d1602988

Observation 6825f712-3351-4e25-aa62-63d2296ebe0f · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.753591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.753591Z digest=sha256:2f2e016a4fde56df3c44466dbc49cab760fb86b612bde246ca2c03ad2f3a0469

Observation 03fd8ee7-6853-4824-ae1f-c8df13e08116 · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rvt: Robotic view transformer for 3d object manipulation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.756756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.756756Z digest=sha256:30a801c069ec083575a3fe165af9e4fc134e06001f9c168adeb55a320f2498e9

Observation 662f0ae5-f798-498d-9708-7c14a6018916 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.759135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.759135Z digest=sha256:6c3f033b31a72e42fe85399d8208d174c30e8d89b705e10998f3c696b919a5a8

Observation d0453c9e-7fcb-4835-9a80-0b48797b0afe · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.761836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.761836Z digest=sha256:a99d61761a8477256303224874d52e07f20f74a8f7dacff617ac393c746925fa

Observation 6c12eae7-0785-4216-8443-8ed0825365d4 · outbound

This paper cites Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.764446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.764446Z digest=sha256:136efa1b7feda7e607698dc06de2bd7a04e81e47c68f67be88bc3a0415309171

Observation 6e4c4c3a-4b9b-426e-968a-47eabbdc959b · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.767550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.767550Z digest=sha256:a6f965ebb51fdfeb29a20c66683a10075db0b1a5e2320114556d3222768d6bca

Observation 181423dd-04dc-410a-99c6-a4bc3b9677ba · outbound

This paper cites Foundationstereo: Zero-shot stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Foundationstereo: Zero-shot stereo matching,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.770358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.770358Z digest=sha256:9208828a7b600dc73221edec6abd04e03a8a4c304dca70a1bff88a3d6399d80d

Observation 554549d3-8d71-4939-8218-cd4de7077f43 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually- conditioned language models,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Prismatic vlms: Investigating the design space of visually- conditioned language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.772653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.772653Z digest=sha256:34a7a9c3c2715138bbf40da6c2f531f161a0bfd5722b6690f987f8a0b0372ed3

Observation 39029919-013e-4d46-853b-efa44f1a2299 · outbound

This paper cites Palm-e: An embodied multimodal language model,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Palm-e: An embodied multimodal language model,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.775125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.775125Z digest=sha256:6a3ae3be4f8ca9126edb31215e509f8d338f6eefe57cb0490c0dadbcf2da10be

Observation 4f73e351-04a5-481c-834b-d809dec0f86a · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.777481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.777481Z digest=sha256:40232e6f541eac4a78b213416c593d0b589d9c50c0f25cfe7c9836249094af1f

Observation b43e69d0-a3c6-4f6e-aa8b-bb178e9ae143 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.780838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.780838Z digest=sha256:c41976c3a295a464e8055844b970d08e367cae2a49f1d2be92e305b94cb85df6

Observation 3a851f77-53a9-486b-bb2e-7e3a89c5e85a · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.784051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.784051Z digest=sha256:9454ae3361cd8ab756e48d7a50b7a72a217ee824258a2600efcb0c8fadd6bd49

Observation 382014e7-f9dc-47f3-9831-1d1944bc0ead · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.788875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.788875Z digest=sha256:670f3ddc254d2aeac96474d54709ec5f5520d75e4473911ca4b26b5834eddb91

Observation 44ed9fe0-4171-4711-b7d2-4fd74768c3e7 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.791509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.791509Z digest=sha256:e5871604b19a6e5b098ba367e9489823022e54988c183f99bcc1a690fed76e52

Observation fea9d4a4-1c36-40ec-9b5f-a6275e93d73d · outbound

This paper cites RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.794026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.794026Z digest=sha256:4bfb346d0b517a20f8146994fdd8bf1200f7fe94d4d17e649775ab8c36a9b9ce

Observation eca535a6-128e-4486-811c-d410d44d37d1 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.796749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.796749Z digest=sha256:1adf57707a96c45b833c93ee813a348e7fa07651349237aa97e7f0eb47dffeb8

Observation f056f3fc-099d-46af-bbbc-0f95a73d027b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.799400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.799400Z digest=sha256:191d62260b32837a8a4ffd3285ec2b3041ba55c63eaabda3930e1c650311698b

Observation a8884fcb-8c89-4f58-b5b6-ee2dd654999b · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.802389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.802389Z digest=sha256:449b6619a94a660f2c9e3f407426088a1948234acb240c99e278278a251e45dc

Observation 6fbd01ba-5fc9-444c-8fa7-850c96c17140 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.804995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.804995Z digest=sha256:af5e2fe530836e007087aac26038242c08fde21e1fc763d5a2e34e6fdbd5ff64

Observation b00390d9-18ec-48ce-a580-866cfea3d4fd · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.807558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.807558Z digest=sha256:13aea52680f0ab064b62680e4dd139f6bb67b548577a2d33c00afaa114941f1f

Observation 97dccf5f-d256-472f-b89a-48a53538c6ff · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.810263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.810263Z digest=sha256:9fef59191eb97762338d77a523092ce775e57d164231c99bfd006de053523f5e

Observation f0450714-5eae-4d4f-8f1f-276e0e82d94a · outbound

This paper cites Internvla-m1: Latent spatial grounding for instruction-following robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Internvla-m1: Latent spatial grounding for instruction-following robotic manipulation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.812858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.812858Z digest=sha256:51eed392f09b36f1fa656d3268a88f7916dd8a94334c850044a1c3582faf4ad0

Observation bd8ee129-bbb4-4602-bcb2-e0b2d8c276ad · outbound

This paper cites Tinyvla: Toward fast, data-efficient vision-language-action models for robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Tinyvla: Toward fast, data-efficient vision-language-action models for robotic manipulation,

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-03T14:00:20.815149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.815149Z digest=sha256:06eab5e992d7c16f0a6d9e0106c76e1fc053621d4cf25019580eafaae9cecac3

Observation 0c17244c-5f64-4dec-9b5d-a7bd080a0504 · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.817485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.817485Z digest=sha256:711fd4a319fd1e60c1cbb07b854ba4026e0f14cf3ec5159a528a7603d2a186fb

Observation fbe30e33-4bb9-4f8f-a2f5-80d8c2b5d870 · outbound

This paper cites Generalizable Humanoid Manipulation with 3D Diffusion Policies.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Generalizable Humanoid Manipulation with 3D Diffusion Policies

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.820085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.820085Z digest=sha256:2e5364ef2ec2961349f9f84fd19d807127f7157c2be49eca0c86afbb8e402d73

Observation 60ae4b1e-e783-49a8-8fe4-ca8fa6d8090e · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.822651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.822651Z digest=sha256:8ebe79a20039124baaf4f495e915ce852a41ac2961481c3477919e2eb3c347cd

Observation 68888964-eab6-4441-904e-4ebba7a9d12e · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.825243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.825243Z digest=sha256:6632af57d6abd197eeca17c2ae43cc2e4a0519a1349e5bbd8fd0a6db61bba814

Observation 3f516bb1-7f71-487f-a83a-b3ad062311da · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Bridgevla: Input-output alignment for efficient 3d manipulation learning with vision-language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.827780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.827780Z digest=sha256:9e0e056714762d7b8538a2880a4ce9d99ea490cf02965bae48ef494359b3b527

Observation 08b01426-c2c5-4119-a374-e7fc6457bbae · outbound

This paper cites FP3: A 3D Foundation Policy for Robotic Manipulation.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision FP3: A 3D Foundation Policy for Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.830245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.830245Z digest=sha256:253c81fed9ce392d538de79cb33d7de0bc23a1db460d1be7611df072b59f6158

Observation f3e85bbd-a1e7-4674-b97b-17ba4355c0ab · outbound

This paper cites Evo-0: Vision- language-action model with implicit spatial understanding,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Evo-0: Vision- language-action model with implicit spatial understanding,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.832873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.832873Z digest=sha256:39f9257490d307baee21292241d0c0718c26310a23b5c9892a7144c90652ec9c

Observation 4b14c6b1-069a-4c3f-8d63-d0608475797d · outbound

This paper cites Gp3: A 3d geometry-aware policy with multi-view images for robotic manipulation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gp3: A 3d geometry-aware policy with multi-view images for robotic manipulation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.835316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.835316Z digest=sha256:23454445bb9e2f26d5e7a62fe60f779768483c58579e0e93b302a58cebe1c2d5

Observation 647f1a5d-5aac-435e-a2c3-c5b1b0cdfba5 · outbound

This paper cites Learning the distribution of er- rors in stereo matching for joint disparity and uncertainty estimation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning the distribution of er- rors in stereo matching for joint disparity and uncertainty estimation,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.837678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.837678Z digest=sha256:f1db9ab9e8f3da32225c359e43a3a3792fe925fbdd8e81bf372daddd668eb817

Observation f295d6cf-0cf1-4c91-9e22-e1c45540f400 · outbound

This paper cites Cfnet: Cascade and fused cost volume for robust stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Cfnet: Cascade and fused cost volume for robust stereo matching,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.840251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.840251Z digest=sha256:2d3419207a23a70d19cac75d433de8b7bf8fa9bbfc4e80987582a28690171afa

Observation af5d338b-0d2c-4dde-94ff-729c23ab1b40 · outbound

This paper cites Pcw-net: Pyramid combination and warping cost volume for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Pcw-net: Pyramid combination and warping cost volume for stereo matching,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.842621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.842621Z digest=sha256:969f442a72c6f33f5f40d1424834403b364cd272eef0c0b5df74ec6754f3727c

Observation 0be59683-b8ed-47d8-b656-0a7af496d967 · outbound

This paper cites Aanet: Adaptive aggregation network for efficient stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Aanet: Adaptive aggregation network for efficient stereo matching,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.845047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.845047Z digest=sha256:c27a44d3c16d684c448b11ef25c642cc054b86da592f4972238a4c9f0d9f55b1

Observation 689ded80-bfb5-4e9a-83a1-2fbaaed4be91 · outbound

This paper cites Arunet: Advancing real-time stereo matching for robotic perception on edge devices,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Arunet: Advancing real-time stereo matching for robotic perception on edge devices,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.847397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.847397Z digest=sha256:825c457f8511b703251e66a0c4be11d0f8cbb485a492dc738a6f30f8452c8d26

Observation e4f029ab-af3e-4d83-b1b9-4943b8887eac · outbound

This paper cites Gfanet: Group fusion aggregation network for real time stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Gfanet: Group fusion aggregation network for real time stereo matching,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.849923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.849923Z digest=sha256:bbb2d7ed17691e8cf7760875299dde9a9c01084e4d232b83cadbf0bd41569c44

Observation 40a17586-476b-42e5-9d95-d701b592fa3d · outbound

This paper cites Raft-stereo: Multilevel recurrent field transforms for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Raft-stereo: Multilevel recurrent field transforms for stereo matching,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.852639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.852639Z digest=sha256:79b18e3ec3ae2e75b506d139b8440a0477ef83088c10556fd6f2bf3c5c09c57e

Observation b657e44e-d0cd-407a-8b97-780df2ec3e1d · outbound

This paper cites Iterative geometry encoding volume for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Iterative geometry encoding volume for stereo matching,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.855173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.855173Z digest=sha256:807c02ccbc8df7e4a78799fd78bac8a7d7b596f23e47d200fe05b9e879ef9495

Observation 821d3cfa-00e1-47ac-85f8-27b550f0867b · outbound

This paper cites Practical stereo matching via cascaded recurrent network with adaptive correlation,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Practical stereo matching via cascaded recurrent network with adaptive correlation,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.857433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.857433Z digest=sha256:35c18518cb7b6efc5ba5b747f1bdd9ce7de4772a542d7c272af7e26e85420354

Observation 9cb26f67-0881-472b-849d-e25512166d71 · outbound

This paper cites Uncertainty guided adaptive warping for robust and efficient stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Uncertainty guided adaptive warping for robust and efficient stereo matching,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.859742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.859742Z digest=sha256:74ada03b10988c527813486d6e94f9c2e6dbe02fc943f45662490fac1f21b553

Observation f7d67585-2b8e-409a-8f2f-255dc443fb49 · outbound

This paper cites Learning intra- view and cross-view geometric knowledge for stereo matching,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Learning intra- view and cross-view geometric knowledge for stereo matching,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.862334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.862334Z digest=sha256:4dc9695d105b729be192773587eb32e40a525c4c9637eef649c06eceeb1db101

Observation 536476e5-9dc8-4179-ac95-322f49b11667 · outbound

This paper cites Efficient and hardware-friendly online adaptation for deep stereo depth estimation on embedded robots,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Efficient and hardware-friendly online adaptation for deep stereo depth estimation on embedded robots,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.864728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.864728Z digest=sha256:a7b65096baebfdb87ee6cba39be2bcfaead37f3ff644a5f24a94d9c02d3bbb05

Observation 68fd4d0a-e8fc-4911-b6fa-302c7204bdc8 · outbound

This paper cites Stereo image- based visual servoing towards feature-based grasping,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Stereo image- based visual servoing towards feature-based grasping,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.867054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.867054Z digest=sha256:bfb96741c30b997be8b2a17fef928d6639c6a60b706f3caabbcec3fd7697bf24

Observation 6c24e98b-43b4-43c6-a2d6-b31742e32627 · outbound

This paper cites DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.869541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.869541Z digest=sha256:67911508c8fbbe7c5d45af904b4f01fc48f20a66ccd356771f09f78d7096bf73

Observation 52159cb6-e386-4419-bbfb-962ed03dc739 · outbound

This paper cites Simnet: Enabling robust unknown object manipulation from pure synthetic data via stereo,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Simnet: Enabling robust unknown object manipulation from pure synthetic data via stereo,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.872190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.872190Z digest=sha256:825015207cf506987bc565845563bc296a1d9b5ad900f60a088fc0529240af02

Observation 183fd8e0-b25a-4fcc-a077-b286aba51e48 · outbound

This paper cites InternLM2 Technical Report.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision InternLM2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.874567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.874567Z digest=sha256:205f116de2d4e91fe01eca31ab48546f5d5b6344e230bec06d41094ecb1c5fac

Observation 9a010d3a-800f-4e66-b2ae-51f5b95e2bb2 · outbound

This paper cites Flow Matching for Generative Modeling.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Flow Matching for Generative Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.877040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.877040Z digest=sha256:a3dabb6b30c9a621e0f824cc3ea8e14f54913b248293938cab83cf7394b33b39

Observation c60bafb7-2202-42ce-a5fd-18448b6c4d0f · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision DINOv2: Learning Robust Visual Features without Supervision

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.879627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.879627Z digest=sha256:aaed200645e93fea066c4aa055e1f5efb57d97e2c16d4ab26d4579f11def5609

Observation 08702300-1375-4a36-89de-93277c83f09b · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Open x-embodiment: Robotic learning datasets and rt-x models: Open x- embodiment collaboration 0,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.882762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.882762Z digest=sha256:7c93fc2582feac572525bca1ed06c780a649ef648cf543ca0830a1784c72e7a5

Observation c7b73b3f-2007-4ada-b7e0-853868a86f79 · outbound

This paper cites Mujoco: A physics engine for model-based control,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Mujoco: A physics engine for model-based control,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.885441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.885441Z digest=sha256:784b6ec27be47af6fcbc7f68e7a27f604d075f6e517e4adec8ce93a822409d70

Observation a3143be9-2046-47ad-8b82-d46d4be7beb8 · outbound

This paper cites Isaac Sim.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Isaac Sim

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.888280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.888280Z digest=sha256:e7ef5e2e76828914e8c006fb57a6bbfcc41daae8af87f8fbda183494e83c474a

Observation bd27a75e-e7af-41b7-9a7c-8dc80099b7c9 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.891559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.891559Z digest=sha256:be02868171523ffd3cfe85e1e32ba49d612886d7c2281b9f1521546c1460ac31

Observation bf7fb575-bf4f-4a41-918c-8fbd10c5ec33 · outbound

This paper cites Vggt: Visual geometry grounded transformer,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Vggt: Visual geometry grounded transformer,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.894490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.894490Z digest=sha256:2c3cc8ce207a3bd6cced5bc98208de486f2d0692ed7a7ede24a6601bdc4ccd5a

Observation 0f44e8f0-4a64-4c6a-b38c-f743f13a6696 · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep network,.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision Depth map prediction from a single image using a multi-scale deep network,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.896966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.896966Z digest=sha256:7b2de5860d3c70c0ad3b1517b247240196ed4fbf8a3b7363f1e07544693eedab

Pith citing papers

Observation 29bcc586-fdd6-4570-a3f1-1810b5506305 · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:21:45.365156Z digest=sha256:2e0c89093e3c7f78b4e45d31d2f63ad27444babe910c302505f399460e6f238c

Observation ff9f8454-ac41-4d72-a36f-9a1ebc423a2c · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T17:53:44.901684Z digest=sha256:9e89486fcd342de481ff9adc73c6523caf171ee19c615294163404e041484660

Observation 5f5dc25b-ad35-4b5c-8263-6d52775a873b · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:cbe6f7f1a6e9852e69e4c69d274f586d498ba29d7ff4f6d7c8d877cf8eb8503c

Observation f801e2ea-02e8-4226-a404-f3f954eb394a · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:23.857212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T03:57:10.338617Z digest=sha256:5a5f5aa4a028b6b9a0b5e7b6fc5ab4f2915a9e38227d635187c87a2444cf7f5e

Observation e57183a7-40e2-4b37-981a-ad9601eefce0 · inbound

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization cites this paper.

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:46.683779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T22:11:21.596611Z digest=sha256:c837abcbe50217c2a517eb274e4e146a50244c9d5e00fc7514270e7c7da80333

Observation a7561048-f502-4d09-aaf8-2992e5b5d62a · inbound

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model cites this paper.

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.196481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:39:15.662180Z digest=sha256:6b9e4f4ed759b98c6f404db7ef1e17432bcebe52193fea231321abca06e27989

Observation d484f26f-42dd-4851-aad3-6662d8f9ab1b · inbound

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning cites this paper.

Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:20.378897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T14:41:50.084254Z digest=sha256:9b78289f34f3e2836c89f2f8d58919c6a0f9a87cf6412afc142cd11f44259726

Observation a7e3c69d-5d32-4bf0-8d84-b30f4260f15e · inbound

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination cites this paper.

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:37:37.072871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:47:28.832078Z digest=sha256:eb44114c10ecaeeddaad5f97e198398ff9df6fd0de9a42789e1e848294988988

Observation b1ebb169-2504-4e9d-ba44-7e4680adbf17 · inbound

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model cites this paper.

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:04:21.054350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:01:48.330369Z digest=sha256:bee1c638a834cbe90cea3332f511d47069cfc8f7ffe89f6f1879af67ea64a3d2

Observation edf36261-3759-4659-977e-f093520ed7b4 · inbound

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model cites this paper.

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.454775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-07T12:28:21.318505Z digest=sha256:4751a6f16b57bf4a28849bf18360a1676cb281dfb6f6177ef974eaa8e0f4789d