Pith. sign in

Paper Citation Record · LEDGER

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 11 inbound Pith citation observations for arXiv:2510.00037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.00037 v6

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554805Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:33.899884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:19:48.867283Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93d059bf-220b-455a-8068-86b1ca67a759 · outbound

This paper cites Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.376470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.376470Z digest=sha256:8155b5f7c0f0ec2acb060acd19b278f43a75711bd47914d986ca471ec186a1e8

Observation 13b5c165-cb7a-43c6-9a7c-9599299a8a2e · outbound

This paper cites Flow Matching for Generative Modeling.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Flow Matching for Generative Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.024872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.024872Z digest=sha256:1c0f570feaff0a46493fba19256c055e911934b646e5bb86f183da0992d1eb83

Observation aed448bd-e510-4354-be6d-baa4945462da · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.250038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.250038Z digest=sha256:3f1deeb2f650e8b7e5f1ddba08c33e8b6c4946b6527847c656502a6e3b55d80a

Observation 0de6fb67-aa8b-45f4-8da8-a62bb926f755 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.522955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.522955Z digest=sha256:e3d8d3b8fe1134c139dfc1b572ffa75bd6c5e282fc7192a16bca606a50dfe193

Observation 2d5cdbe7-f040-4892-af65-3a7cf7e363ea · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.704747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.704747Z digest=sha256:3beee3cb25fa9d1c67ca83c07cd380e3b08fd9e78f49cd41753e2396121b3e90

Observation 7672aea4-10b7-4194-b65f-2e08c9250624 · outbound

This paper cites Vision-language-action models: Concepts, progress, applications and challenges.arXiv preprint arXiv:2505.04769,.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Vision-language-action models: Concepts, progress, applications and challenges.arXiv preprint arXiv:2505.04769,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.834746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.834746Z digest=sha256:412b0f46af9d686cb810026fbd2312422ebe782f9fbf009c68160be8d8fc7817

Observation 528b7f4b-f742-4a42-be34-deed34c3db36 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Octo: An Open-Source Generalist Robot Policy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.974007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.974007Z digest=sha256:5a4bc51f66e320ac422e3e80b6de196c9a7d5e82176329efc72ac8a40ab1bdc4

Observation 4a1e4bdd-335c-4b95-8e03-7a2c428acd87 · outbound

This paper cites Robust Reinforcement Learning on State Observations with Learned Optimal Adversary.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robust Reinforcement Learning on State Observations with Learned Optimal Adversary

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.134753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.134753Z digest=sha256:2f1747e2304f7e9727a54487da53354c3bbce3f0fd14b7225b37777509a6695a

Observation 07df484b-414d-4dc7-ab4c-71b1074eeebf · outbound

This paper cites In robustness against VLA input, we additionally use UCB algorithm to select the best perturbation.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations In robustness against VLA input, we additionally use UCB algorithm to select the best perturbation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.554805Z digest=sha256:00fd9ac66d0865f399634ed3176b52cb2698504fdfb497b7c49ecf2124d48b2f

Observation 623bb6f6-37c1-47b5-8afd-102391bce1c6 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.452731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.452731Z digest=sha256:2abf9c8d237fc9d8d51d38159dbaf9ea34fbbc0b7d51fbe496ac7a97acab0e0b

Observation a1ca3ec4-4b91-4633-a0fa-9f62958c0589 · outbound

This paper cites Robust Reinforcement Learning for Continuous Control with Model Misspecification.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Robust Reinforcement Learning for Continuous Control with Model Misspecification

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.384743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.384743Z digest=sha256:0240c4dc1daded88aa45f1545b36a772f1f8c0710f6bc95807156b4d24b82978

Observation 4bf2ccd7-86ba-40a2-9672-9a0f50fc6633 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.912617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.912617Z digest=sha256:c30de479964359ad277820ea0ae669e340198ab0056bbf69f903da0718963b0a

Observation a8290e40-00db-4c03-a400-30c712074d45 · outbound

This paper cites A Survey on Vision-Language-Action Models: An Action Tokenization Perspective.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.253146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:38.253146Z digest=sha256:d35e60ec535980e4950a66bb9ec34525bddf52c33defe1fc1e8ae46a5f0f6055

Observation 4a3fec9a-378b-4b3f-a59a-966e6d35fca0 · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.682378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.682378Z digest=sha256:e7f2b29ef08baadb66c7f7b254b58be6641364476c1062d1a2d9eb7165f8a166

Observation 43db3ebd-a31d-4ac1-b4c1-fd96d6a97837 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.144747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:37.144747Z digest=sha256:1b81a2dec0702a70288a03f23a9c2ad75acb346e572299e7ec0f9d8bd0e86f0e

Observation 368931d7-7fb9-48d5-a52f-578c8297282d · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.544334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.544334Z digest=sha256:10d2faa06769587ef12a6d2688e4fb9329bf066e9c932a706142eba96b62de98

Observation 54cacf35-c3c2-465d-a4e0-e4b79061e6bc · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.786424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.786424Z digest=sha256:b75930242c957fc4be51253a5c5fdc8061c08bc491d66b74af7757ac8ec6d83d

Pith citing papers

Observation 41a02e7a-a149-45f3-9f0c-e1abf2527fb8 · inbound

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations cites this paper.

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:33:16.554905Z digest=sha256:912785e733ee9a3778f49e499662c93b6e95de9d3233ddb11f81104c14a4a549

Observation a6ec3215-5d58-4e78-96af-43ee881f117a · inbound

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation cites this paper.

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T08:52:07.545191Z digest=sha256:78b34d0dc6d89f3439f8c8570546ebb610d48c0bd26e4c12a0b0bb77b8353b91

Observation 0f2b45cd-7d25-4d2a-9543-c7061f181417 · inbound

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty cites this paper.

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T16:08:23.588699Z digest=sha256:bfa80333466305068f52a7d3b5392e0f58967a7d3a0b78057e6dca9e73940ed5

Observation 3d81e558-2497-4c19-9d29-0fa4eff0c2eb · inbound

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty cites this paper.

Breaking the Epistemic Trap: Active Perception Under Compound Uncertainty RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T11:18:08.362562Z digest=sha256:44e65f9e7e6bd21af371c7abab30ff705ae0872074a562af672b9ad2f5e0605c

Observation 94ea7ac5-e1cd-45b5-869b-65289d836159 · inbound

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures cites this paper.

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T11:40:15.064339Z digest=sha256:34bf936bf6d28d00d36cb822513c165f821eb9805cdd28b2f15179646131e0c4

Observation 5a1d6dc5-1b70-4d21-a5e9-bf1e01fe7688 · inbound

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies cites this paper.

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T09:49:56.894300Z digest=sha256:2b28862b634c32712dcbb9f3d5d744f684e2f6db5ed61de5e4d9c1f102507e7d

Observation 321c6bde-b5df-4f2e-bc80-b8572ddd5ba8 · inbound

Flatness Preserves Instruction Following in Vision-Language-Action Models cites this paper.

Flatness Preserves Instruction Following in Vision-Language-Action Models RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T08:06:36.393665Z digest=sha256:34d78736302cf2f7fb5b14052a5aa04754ed7464ea8b36af503300ce5ec0a046

Observation bcfa979a-1433-48c1-ac78-9dade9708810 · inbound

Sequential Planning via Anchored Robotic Keypoints cites this paper.

Sequential Planning via Anchored Robotic Keypoints RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-21T02:20:31.091014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T04:59:12.363425Z digest=sha256:fca4897581ed73dd735196b2679f214973ea76efda1105139e3956160c2b6f20

Observation dd5d62c6-1358-4f73-8161-a2f3e8dbf3ff · inbound

Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models cites this paper.

Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:41:31.531782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:41:31.531782Z digest=sha256:d99f69413de6f61002e3fcd3a4679b45ae71fe47e7d50eb9a770fa81854a8926

Observation 463ffd35-cc7a-4d8a-9694-0c890eb33c4d · inbound

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning cites this paper.

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:33.899884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:33.899884Z digest=sha256:a9901b5367fc4fbbde04f9189f42d8d6478e15d948e241631f7b14dd8fd8b7cf

Observation ab5e601c-7f8a-4286-855b-52ab87d66ff3 · inbound

UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models cites this paper.

UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T10:45:44.781272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:45:44.781272Z digest=sha256:d9e4fde3d58dccb54c9783c936856bdd6b1d7def5883220615c787e6a85e9ee2