Pith. sign in

Paper Citation Record · LEDGER

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 4 inbound Pith citation observations for arXiv:2605.19678.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19678 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:48:39.069675Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:27:58.462125Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact28
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e57640bf-08e9-4183-ad62-debaad09c2fb · outbound

This paper cites Robot manip- ulation based on embodied visual perception: A survey.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Robot manip- ulation based on embodied visual perception: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.505050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:57c119836819cecf4d5e5f9d246a2322a0d6105ca599dda853c8bb68f035c277

Observation d1acaf5d-ebe7-49fa-b3e8-ebc8a2b25beb · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.469762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4bf631c754a7252dddfd85e27947598a71ede442d98255ab836d7a50797badeb

Observation 18cd10a6-cb70-42fb-b6f3-6442a03c0f72 · outbound

This paper cites Pointnet++: Deep hierarchical feature learn- ing on point sets in a metric space.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Pointnet++: Deep hierarchical feature learn- ing on point sets in a metric space

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.500887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:2e507504c3b0f8ca5185326a1b2a25f9be2ed4cd27ef7379aa2ea9a9adb6f932

Observation af129ecb-b1ec-4f44-b6d0-53f56b6f953c · outbound

This paper cites Dspnet: Dual-vision scene perception for robust 3d question answering.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Dspnet: Dual-vision scene perception for robust 3d question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.483514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:845a9162a145bdf27b9f40d7082367e57236468d652440906e499ab351ff8ed6

Observation 99f3fa20-5457-4054-a13a-3bc8c71d2708 · outbound

This paper cites A Survey on Large Language Models for Automated Planning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models A Survey on Large Language Models for Automated Planning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:05.028904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:88376cc3124e23b850d6583b2c79562c31cdd40f878fffbb133a2076e254d49d

Observation b41d9bf2-095e-4f6a-9512-fe5ff9f4ad38 · outbound

This paper cites Qwen3 Technical Report.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Qwen3 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.098240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:e62870120f09cb7b1ed6ca3141d6d8e116fd6317b5eb5c761c27bebdeaf6844e

Observation 3db5227c-06f3-44f0-942b-b0599a6a8346 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.955065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:0d628d60ec72b99bf3c17fe27fdc1554c411d9894b8f78002e1c7d36b6b57950

Observation b4914052-5b62-4881-b68c-41af3e22c01a · outbound

This paper cites Scalable diffusion models with transformers.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Scalable diffusion models with transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.484971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:439fbe1781a3f2b69801c7cc00a15446427d860f6792aef02388ba1bee62a619

Observation f2708e3c-5574-48c3-94c3-0094d5ab062b · outbound

This paper cites Flow Matching for Generative Modeling.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Flow Matching for Generative Modeling

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.961336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:b2e5741f4e5644a4f081561f8e452da4268191afcf5230c741088da1bb7bd2c1

Observation f875bbee-ee14-4c57-9460-60ec2d321e7d · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.491411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:c4a710a756cf063c5d352cdaa85f4a7bbeeb15c1c177a10fa363e6a206ec8216

Observation 9999ba4d-16d6-4794-83c8-b0537da292a3 · outbound

This paper cites A Survey on Vision-Language-Action Models: An Action Tokenization Perspective.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.059092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4ca5e593008b34123b761375de8dafe464cf3a7c6e0d894dfb07f28753e43445

Observation 4a2abc5f-5600-481a-8793-8fb32cf11e7c · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.034940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:c9567ba5c4af5ed371adb3ac867cb4c96ade1367bd01e33b2592ee0cb2043298

Observation 8ae1af9a-4cc1-4cd8-b608-32cbbeb489f5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.980530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:e909d4d84a29e23db9ff6ce7aeb28eb9f490df1558b03de6732f22f07cdb89dc

Observation 8d9e74a5-ccef-4af3-9c75-182bc6dd3547 · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.948828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:895b398f3b01adf64161ed4e0168ba25ec857223280978e8c0122c950f762ce5

Observation 3804e7f5-f1ed-46b1-b6a0-4fca197f46a0 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.084084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:820578e423e0b7801f8f2b6ca6a785fff1a6cc210acd2c50deef5e1cb7f05b33

Observation 38d43cec-182e-4832-bbab-38dc98379317 · outbound

This paper cites 𝜋0.5: a vision-language-action model with open-world generalization.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models 𝜋0.5: a vision-language-action model with open-world generalization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.473722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:6e759069eb054aaa35c05a73e44ffd13c1e0013df45eaf85579b660423636040

Observation 2ea9bde2-168c-42eb-8016-4e8055baa41d · outbound

This paper cites GR00T N1.6: An Improved Open Foundation Model for Generalist Humanoid Robots.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models GR00T N1.6: An Improved Open Foundation Model for Generalist Humanoid Robots

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.486709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:ded8d7d7ba54a5208cfc11733bc47a226269cda49681c2c64956b003d8e28361

Observation 7f251330-dae6-420f-8dba-36fecf44ff78 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.065003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:f594784cff2031a673a86cfa56dfff3d8a1bc4fa86e57207932f60a59a388685

Observation a1cfc27f-2e73-4acd-8106-53cf64716e85 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.465903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:089688f8ac60da92b92a86c4c1111121f548c90983ad0508bfc38f956563f8d3

Observation c20cca0b-4517-466a-94cb-990df1101f8a · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.016329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:f97ce56977cc47921c3718aba15e79478beab643f6be0de6b95c7fbe9fee6c68

Observation 4eb40f09-a60f-4907-b16f-d8f321b61cea · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.942661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:1d745f824f634b10096f7184073373a57fe4e58e6e79c601b83389d9c1c1b954

Observation d31d1d5e-4cfa-4f65-99fa-ecff09788d4d · outbound

This paper cites Exploring the adversarial vulnerabilities of vision-language-action models in robotics.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Exploring the adversarial vulnerabilities of vision-language-action models in robotics

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.493662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:683d173d234d1a0b7253512e165ec100af3740aea1f76a6b2e66d61263898795

Observation 7fac4a71-de89-4aaf-bb46-0c121708859f · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Instructvla: Vision-language-action instruction tuning from understanding to manipulation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.968738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:8c6a22011c4189c818c396febebfbc597af98164353d4fce883b5ac16698e833

Observation 36728258-acf3-42fc-b438-0cc387b4ead3 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.344448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:a22388cc59b19d3353f20b7a88a1eaf1c3f967c37b706bbff5cb4adc32a62d4c

Observation bdfc9829-895f-48e9-9865-735dd8eb8ebd · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models WorldVLA: Towards Autoregressive Action World Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.936624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:7d67fe1abc375b59a9aff2ef27961671c14bfb39ef28bb229fa1ad3e586f7690

Observation 8d9aed89-8314-4974-815b-6ecdb9d97484 · outbound

This paper cites Unified Vision-Language-Action Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.974302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:10777b5dccb016658328bb41eb2494e48c6618761de6ccf9bddc87f4d2845d99

Observation d9def460-4881-4c81-a809-3ed73587fa4b · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Aligning cyber space with physical world: A comprehensive survey on embodied ai

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.444517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:3f2a6cd059a1a4e2f168703947a3d3cf58cff4c6bee27e95f9b34e8d224f63c5

Observation 5adfb2df-49cb-4ed2-8800-96e320924b09 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.466587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:caf4f7dd2081c89e371bed0706ef4784a528f1d25501e509fcd06410657f10f7

Observation 100cdfe6-13e6-4f98-a106-dacec8f07232 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Octo: An Open-Source Generalist Robot Policy

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.921971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:79ecb287640f5c421d43dc09f227cc13fce83a065f11d65b57c46a28e3df0d0e

Observation de6ad2ba-2729-45f1-b6b7-23bbae718b82 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.994309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:bdfc702008eb7cbbfc164c8a763075e7c7c55530be5f75f71c5675418bb93399

Observation 1bbed8e8-9328-4cbb-9cdf-cc0c8222db22 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.916497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:6025c65de69ea353a9d6d85abd5d2bc2cfd69851ec899f69b5c4a311231ec2e6

Observation aaae344f-c28c-4e9c-adf7-157f92a7a0d1 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.988204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:e00bf222439148a6fcc663a1a4ef78349f463f14167ef549171aabb2a70b9f5f

Observation 63d9bbee-b054-4230-8624-9a32f21dd066 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.053263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:f6d90b38f693cade4c862e2da6f44d19423005cbda7e54e7e38ca850826bb5ca

Observation 2973567a-d2b0-4726-94cb-f1b680de1706 · outbound

This paper cites Vlatest: Testing and evaluating vision-language-action models for robotic manipulation.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Vlatest: Testing and evaluating vision-language-action models for robotic manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.502640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:427284e53292d9106b0c4c2029f5b69b7a9a1b3df2e868cd01fc486ad3cc7993

Observation 430ae5d8-3844-4f86-887b-010754764c0b · outbound

This paper cites RynnVLA-002: A Unified Vision-Language-Action and World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models RynnVLA-002: A Unified Vision-Language-Action and World Model

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:36.254972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:1be49645537c27bf8d8e4676b5856c93be952689abfd7366b77f4f8d90995784

Observation e962c671-c9d7-4baa-b19e-ac5311ca4fd9 · outbound

This paper cites Motus: A Unified Latent Action World Model.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Motus: A Unified Latent Action World Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:04.909357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:3d8c1eb090a27efd20f121856dd94f13251367948c19f4f5bdfccd92d0d48e32

Observation 56dc571f-f0fb-4aed-8b64-3999e63c1f19 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.074152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:0c281cc6760dbd6e96155cc8f153e673faad33cb0ff4008ef51543605987e74b

Observation 7fd2a87a-9997-4dc3-9544-66ab4c0c7f21 · outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.009602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:1c03fcb6307eb99f41cfed73fdb3b184a1e46d2069f9b8dff199bc92e6042b87

Observation 1b34c62f-a977-4c36-b31c-df5317b9ebfd · outbound

This paper cites Robustvla: Robustness- aware reinforcement post-training for vision-language-action models.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Robustvla: Robustness- aware reinforcement post-training for vision-language-action models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.928533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:c9c0afe4276649b839ecb513b3df17dc4a29ae602cf283a8f1ead44a889d0b5f

Observation 967c2efc-9665-496c-bb1a-d3a15c81b749 · outbound

This paper cites Mean teachers are better role models: Weight- averaged consistency targets improve semi-supervised deep learning results.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Mean teachers are better role models: Weight- averaged consistency targets improve semi-supervised deep learning results

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.490131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:81181d9ccff026e0c3585095ea77603b438f180af27aeac97eb090f3247e1d91

Observation 9f6a548f-18fb-4bcb-8ec1-ea1abdd79833 · outbound

This paper cites Virtual adversarial training: a regularization method for supervised and semi-supervised learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Virtual adversarial training: a regularization method for supervised and semi-supervised learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.481359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:da5aec7f243d2cd839fa5aac757b8603a6872fd4f42aed42a7a956729d09d4c9

Observation 044070dc-4a57-4f35-84b8-249db13d47d0 · outbound

This paper cites Fixmatch: Simplifying semi-supervised learning with consistency and confidence.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Fixmatch: Simplifying semi-supervised learning with consistency and confidence

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.497406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:a22559caca16622a79a01d410e042b53d2a700c1e80119ad2de4339de825e4f2

Observation 62bfdd38-53b8-4dfe-b9e2-150272b6debf · outbound

This paper cites Image augmentation is all you need: Reg- ularizing deep reinforcement learning from pixels.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Image augmentation is all you need: Reg- ularizing deep reinforcement learning from pixels

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.488231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:dc0f67a77af8a27f92582dd5652746f6f823cc124c3a29a32d9a7c47cd5e359c

Observation 6be07c4e-cdab-407c-a83f-ea04d3444c97 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.090887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:4dffc963a69a48896f37223fcb5f2f98e989d25c99330df0109f5542cdcd722d

Observation 8ea2901a-c186-40e0-bcb5-984fc0e36e62 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Explaining and Harnessing Adversarial Examples

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:53:05.040820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:d5dcfc06e95b7af8e0b39a728a05dd976cb9536c6025aa4f4b518b057977555e

Observation 5e8f34f8-3642-4613-a558-7c3dacab4437 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:53:22.498790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:ac4811e378c9f4479d8c635a28f1d6f9698467d04b22001d13c932cf99493a5f

Observation bc487d13-178a-4d5a-b86a-e50d35ff3eed · outbound

This paper cites Decoupled Weight Decay Regularization.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Decoupled Weight Decay Regularization

Reference 47

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T04:53:05.021961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:fbb869585fcb70e46a0458db12854594f81d211bc11879c114333e8e79a0fc82

Pith citing papers

Observation 5e01aef6-0534-4d92-b406-61603fcd4d26 · inbound

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution cites this paper.

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-01T20:33:34.710891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-01T20:28:18.919704Z digest=sha256:cc85e4a852816562ba6649422500d8f58eef1e23affbcb439d7b7db8a6f7750e

Observation 70fd44b0-e1af-4694-89c6-37006110f8a1 · inbound

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models cites this paper.

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T06:10:33.302184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:10:33.302184Z digest=sha256:304b8c37f74e1c0e5516cd3c0bbf3adafda4f7728413c94a442f436b1a83d94f

Observation 75fb5270-f809-4809-ac2d-dcdb8efcabcd · inbound

Probabilistic Reachable-Action Verification of Visuomotor Policies via Set-Based Training cites this paper.

Probabilistic Reachable-Action Verification of Visuomotor Policies via Set-Based Training RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:07:02.883441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:07:02.883441Z digest=sha256:29abc969ffbc579ecfbe938f452d286550500622f52160664e8ad22929cfea04

Observation 926b2cab-789c-4ed3-8301-018a6cbbb54f · inbound

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling cites this paper.

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:58.462125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:58.462125Z digest=sha256:22c02063e90d3da48e87b3e011d220754183421c1230bcf78d720d50daebc639