Pith. sign in

Paper Citation Record · LEDGER

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2605.18287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18287 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T11:55:22.498055Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:36:25.569578Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact16
  • verified fuzzy26
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71955047-ede7-418f-956f-a69f5c026ee9 · outbound

This paper cites Alemi, Ian Fischer, Joshua V.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Alemi, Ian Fischer, Joshua V

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.728305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:1264335e4ac472c5b7a5d3f49606fa9839333bea8aee4df6a08ea790fe17546b

Observation 2bd32c82-ab67-48d2-9d22-9f659fc72a3a · outbound

This paper cites Xcit: Cross-covariance image transformers.Advances in neural information processing systems, 34:20014–20027.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Xcit: Cross-covariance image transformers.Advances in neural information processing systems, 34:20014–20027

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.731673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:41b9efc8c72fee734405d6e45c5d4854dadc3b24e0a32349508a3db1a3a2125f

Observation a9f5a82e-6a2f-48b0-85f9-800bac36a3fa · outbound

This paper cites Qwen2.5-VL Technical Report.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.173382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:bce7d64864d2a1f54e265854f1d3aeafb8da8fbc70d6b67b04e9458955681c83

Observation 990407d1-578f-47a0-b539-3a5d27f4b93c · outbound

This paper cites Are transformers more robust than cnns?Advances in neural information processing systems, 34:26831–26843.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Are transformers more robust than cnns?Advances in neural information processing systems, 34:26831–26843

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.724628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:17de4d014cbe195ed85614174678d9f41bd67123a83e7466e8089fd684726acd

Observation b609cd65-b15f-420e-9a44-2b69ee6d4d8b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.170126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:3a65ae3cbf64ad7bfc768140c9c93fb7b8e582058ffaeb2c884d60021c046918

Observation 7196ccb4-1309-43c6-aaeb-01500c19e9a0 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:58:15.167091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:3d5b699366ce6d3de11e851b7bc1b94c759be57cd9eae43bce3288da0ba31e38

Observation b062f4f4-2fce-4efa-b160-99257e6a74f8 · outbound

This paper cites Rt-1: Robotics transformer for real-world control at scale.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rt-1: Robotics transformer for real-world control at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.729946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:1483a0dc1cefa66a1ba5ce7c7048c8da2cac3ac35421dee760fbcf3d48a8a6d5

Observation 83284948-60d4-42c9-afb4-05b04ec0c156 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:14.767994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:dcad1b59094a32d8caf2ac16258c3772e2c702ebe65a02067da6cda5002894c7

Observation def13f74-7889-4b80-ab98-f879149070ef · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.Int.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Diffusion policy: Visuomotor policy learning via action diffusion.Int

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.702961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:4d02070c1506183c43bb9ac9da83a4d8f91761a6c9d4b2eb1cd47a4bc1f2ade7

Observation 8b9209b9-5132-494c-a73d-edfbdafb69b2 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Diffusion policy: Visuomotor policy learning via action diffusion

Reference 10

Resolution
metadata mismatch
doi, observed 2026-05-20T11:58:14.764574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:b55b8e4c8b0d3e3db750567ca3b2fb13dda1bdc3257c4205bec0cdbcaf18b58c

Observation 0940166a-f009-43ac-9f27-e08f3ed23227 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.136560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:4514b138a58ffbd5c5bfffa51aa41966cf771017bf8798493903fcc9c31f7dac

Observation 5ff1fa70-643a-49e0-970b-a18c124ef4ab · outbound

This paper cites Agibot world colosseum.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Agibot world colosseum

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.735442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:862db71c81b1f74b70f8bbf2ef333b33567c135516efd4da83a551affc229bf8

Observation 55cbf250-217c-4f5e-a8bd-bf7bbea6cd76 · outbound

This paper cites HumanNet: Scaling Human-centric Video Learning to One Million Hours.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data HumanNet: Scaling Human-centric Video Learning to One Million Hours

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.164341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:ad405e206599c31fda266e49184b0c7ca98c1a3fb0836b5500fd1e82be772b1d

Observation 8817eccc-c6f2-434a-8879-e4e51139355b · outbound

This paper cites Rethinking video generation model for the embodied world.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rethinking video generation model for the embodied world

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:15.161652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:38815154e5358359b5355fa7648bceae72bfa123fda27f11e493ed7d044c4e03

Observation 89d890d8-b207-4627-b537-23a5b88b1dde · outbound

This paper cites Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.733401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:9c3b5de40378a2dbf1d0141cab2a8ebd44086601060b391237f0027987a8a231

Observation a96f9eb6-5bc2-4696-b900-3453c0c8184b · outbound

This paper cites Towards Human-level Intelligence via Human-like Whole-Body Manipulation.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Towards Human-level Intelligence via Human-like Whole-Body Manipulation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:15.142464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:e0edf11262838453233d82e56bf21538de3846ae3ae5a37a8c03088981262163

Observation 299f678f-ce8d-46fa-9016-45b0968d7cd7 · outbound

This paper cites Benchmarking neural network robustness to common corruptions and perturbations.Proceedings of the International Conference on Learning Representations.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Benchmarking neural network robustness to common corruptions and perturbations.Proceedings of the International Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.722916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:56dc4abe49e6efc256724b66bd11735bc907bfcc4afa0a3305f08152d0c551b5

Observation 021fbe9e-c686-46ca-8a01-21c98422625a · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.ICCV.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data The many faces of robustness: A critical analysis of out-of-distribution generalization.ICCV

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.717307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:692477fb78995f96c2a7be9ac277874444b1897b34aba119d536659546141ded

Observation d4c71e2a-bc5d-467e-b756-89ca0afdf78e · outbound

This paper cites an unresolved cited work.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-20T12:03:27.711354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7edc65aa2d135ba039f81ce3d9e405790f19af9fc19a156c568c729f3e86ea19

Observation 614019a5-b697-4576-ae34-13d21013dd75 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.737307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:49a52a09e31b8a099c47b68348accbb27c6f84f5e27ea16654aa84a6a1cd92ab

Observation b573aaef-d014-405b-a7f7-d410ffc5e92c · outbound

This paper cites Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Heewon Yang, Arefeh Yavary, Tony Z.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Heewon Yang, Arefeh Yavary, Tony Z

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.685461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:53818999335022d6ed041315787ac99db0e8b43ee772e10f7468f33d59acd510

Observation d6bb89dc-c2e5-43a6-83b9-8ac0c8c8c214 · outbound

This paper cites URLhttps://doi.org/10.15607/RSS.2024.XX.120.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data URLhttps://doi.org/10.15607/RSS.2024.XX.120

Reference 22

Resolution
verified exact
doi, observed 2026-05-20T11:58:14.770581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:d859bb7dad47c2cc10ca3a10d1c19b5251e5cd2b7e92679028fa0a4f68f3ef49

Observation 33ad01ef-49d9-4f57-9795-74514f67d4d5 · outbound

This paper cites Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.701892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:1509c0dcf41b8669a862a12e10a406a3ed1968650bc38384508c385ee6ddb543

Observation 797c3b15-6fc1-4ef6-9736-40169f9cef1d · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success.CoRR.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Fine-tuning vision-language-action models: Optimizing speed and success.CoRR

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.704770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:40780f33c721956ca73196097b3ab38c4b953d205766f454e5fc7dc67e238d35

Observation 7fcc94fd-64f8-4210-be26-44d7d2cb0665 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.139418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:fda28b072f9bdfb26217adef0e24038c1960c375d8f6f54f51ec3b02fe1d0319

Observation 190ca374-309c-4295-93ea-5713c30979c9 · outbound

This paper cites Roboflamingo-plus: Fusion of depth and RGB perception with vision-language models for enhanced robotic manipulation.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Roboflamingo-plus: Fusion of depth and RGB perception with vision-language models for enhanced robotic manipulation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.719273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:bab270b2004c42644b47ba527045213907c5c669d2955f3067655fba974e7dbc

Observation afa9f423-e9a9-4b0f-b92c-148ece0f333c · outbound

This paper cites doi: 10.1109/RCAR65431.2025.11139480.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data doi: 10.1109/RCAR65431.2025.11139480

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:14.773798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:dfa8a6981808fd7e212e1a9ccf0d92e0c5c272888af48f65d03e5c38a7acbc2c

Observation a2a1bb2f-64d7-4887-a4d1-2002365afd76 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.693277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:965c280c49a3da20521da21e3a9a3dadee651748dc24583eae290040d3c5edfe

Observation e9e2ed33-0cb6-48a2-824c-2ea7970899d5 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.710680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7b5c64baeb9d3ce47583632188d70ec24dd6a8cb37d07632b5e17e41adea06ad

Observation 827d637a-0d04-46af-bae8-fcc3215e2e08 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data NVILA: Efficient Frontier Visual Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.155933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:4ff6bcf01e450ebd062d36cb9f001181448d7a93b2b20e2b1ef398e0b3060997

Observation 8155be90-5482-4f7f-81e0-1ac3dbb37368 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.713207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7003ad9054f4ca2b5277c5da9a325d693cc3cce1364a1f10e578ef89b48d0aa3

Observation 4be2b843-fac8-423e-b9ec-575d02854d37 · outbound

This paper cites Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:58:15.147995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:081fd09007d3a61e2cc0bb13aff6bbfd1746942dff8f95ce3023ffb63c767b21

Observation 79fed709-a4b7-4c5e-91f6-2d306b465f78 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version).

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Robotwin: Dual-arm robot benchmark with generative digital twins (early version)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.708820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:2daac19a0d345588d94577ae28559bc563830afcf086af658e315cc9b92dff65

Observation deebf55b-fe7f-4d3f-bcac-cad89a71ba21 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models : Open x-embodiment collaboration.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Open x-embodiment: Robotic learning datasets and rt-x models : Open x-embodiment collaboration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.712459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:a3ef23ff2a476a914622b5e5d06e113dd45227353d245f69a61d80abdd833656

Observation 8ae8652a-4947-44bc-8b87-8e5007b23ff7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.158759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:0c129499551e5010ab50d15c70ba6bd0f119b384996a51ff3eb67c2ca25fdaed

Observation 020750e1-701d-497e-87a9-439dadb69b65 · outbound

This paper cites Vision transformers are robust learners.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Vision transformers are robust learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.726701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:0445b93c83f163405e02cb578b222f83584b751a32ba451e258e70af2f98921c

Observation 8d3e5c2c-bc1a-4027-aeb8-ce257eefd99b · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Octo: An Open-Source Generalist Robot Policy

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.153259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7d2eec949a196bf69e148857f338dd6e3614e92b075fd40e99a673670ad5c357

Observation 2a766d65-3d94-4b58-846a-8adf51c5bbcc · outbound

This paper cites The information bottleneck method.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data The information bottleneck method

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.133169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:de956a111be00987c721c4b487c76088330af9ec2bbca2d87d9ae38145e98a45

Observation 8fdfb1a7-bafc-4a86-8e06-3b23f2bc4379 · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Domain randomization for transferring deep neural networks from simulation to the real world

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.720017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:6a736d7c486cfb508b91f5c80a5f2f4d45d0a33e11dad80431299f06484802e7

Observation 0dc466fe-cffc-463a-96b4-5fd8c49b9121 · outbound

This paper cites Augmax: Adversarial composition of random augmentations for robust training.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Augmax: Adversarial composition of random augmentations for robust training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.718175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:714613eddf532d545696b3ce55d5fbdc868a21b9d1981f7bb19c8cb73c8fb77b

Observation ce59cd29-1508-41e0-8a3e-a87b748766a5 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.CoRR.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.CoRR

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.694893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:fc9cfd0906f01075ec1d9d79ed4a7672a06686269f8bd3e5979effb1942c5ee1

Observation 5598384f-0e68-44ba-856f-3f83ebada8ea · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.145375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7ca486e6448c694eb6b72d63c082b9f636691350792d2a9670be917c14ffca59

Observation 66ab4261-8476-4cc3-8f93-abdc1f35d5ed · outbound

This paper cites Robotic control via embodied chain-of-thought reasoning.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Robotic control via embodied chain-of-thought reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.695629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:0873c080e00f59eac5ea335054e0dbb181073778a1519cc84610d04b93f97852

Observation 3768ab48-191b-4ac6-ab30-a3e3e200378e · outbound

This paper cites Sigmoid loss for language image pre-training.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Sigmoid loss for language image pre-training

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.721055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:acd392a54647caccffeff7d9d76afd9138de6407164c6e52335ec0d7cc7f9e3c

Observation d136d2bd-3abc-4304-ae0c-5f63586de58e · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 45

Resolution
metadata mismatch
doi, observed 2026-05-20T11:58:14.776025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:7dd8b1b9c946610e34847348aa8f3b57af1d15fd02cd4db542b0782a02cc25fc

Observation 8ee54499-56bf-4f3c-9546-7c808168c652 · outbound

This paper cites Understanding the robustness in vision transformers.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Understanding the robustness in vision transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T12:03:27.698339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:c76379d7781a4e96b84803fe15686762c847ceb584de3bbc7658853ec3ddf81c

Observation 3abfb6fc-90ee-4478-9f25-e4de7e526630 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.150672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:91bd2f09ba4933e07cf6de69c79db19729a8873f5c2f7fc9819f9f8c8f3540bc

Observation 325e311d-d63c-49d2-8e8d-b8a322ccdeef · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T12:03:27.707725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:40f259471d0258b9fe24f9ff9c085d596dc9e2d9d56f80c80a84fe0a3e97d0f7

Observation 904c6979-8470-4e34-8079-cf713816dfbb · outbound

This paper cites an unresolved cited work.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-20T12:03:27.705530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:6ee3f98aa11376448c756e6e07f4301873f541b5557b6fcb2a36eab21e8ffa80

Observation 3b5363f7-b54c-487e-87fb-9c43a9713948 · outbound

This paper cites an unresolved cited work.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-20T12:03:27.689625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:5496e9a999ce109ebcba75e0700b401da3d429b5182eaff556dd5af028cb22e7

Pith citing papers

Observation b093e008-a336-4fdf-a256-0eca4df13610 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:25.569578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:25.569578Z digest=sha256:1771f5d08134e3c3253468f8902260765dd5957ee0fd3740a294d8ba8371c384