Pith. sign in

Paper Citation Record · LEDGER

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2502.14420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14420 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:17.808710Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:56:35.270996Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ce718d6f-8e18-4186-b3aa-289210808907 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.648354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:5d34cfab68091ff4d3cc96ee80d8339d99f27f16e58cca5976bbc519724055c4

Observation 9dc6f745-b6fe-4d6f-ac3e-d374ec8ffdc1 · inbound

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning cites this paper.

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:17.808710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:17.808710Z digest=sha256:edc4eb62c5d73807c7df3d2e5a2fa2cc67a73c09a90749f7dff5c8c19dba0c56

Observation 226371c4-8467-4e3d-a7c1-017f7d6f1c8c · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:57.438562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:57.438562Z digest=sha256:689e628afc8343dad0f5c9bfc625024c85c4b8c66507238171c4f3881b533f8c

Observation 15ee6829-a0a9-4ab3-a46f-971d4b4f73f7 · inbound

Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents cites this paper.

Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:35.637019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:49:35.637019Z digest=sha256:f30c226dce6a968a2f883ff01e0c4e5eaa7ee12de8bc9fd8adbdea2e6ce8aa53

Observation 6082cc5c-280e-499e-852e-b1f654eeb5af · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:16.632623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:16.632623Z digest=sha256:712828ab64a41cc0bddf34f0e10ec61cd3e3ddbf52c4a30d802795af6d2db453

Observation 31675d45-1f75-4d09-ad2e-e1456f0170f9 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 271

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:38.147800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:38.147800Z digest=sha256:733f74b01104bad10a88550ae0e991c82bf98f37924599ddbf64441945d2e6af

Observation 165a2863-afd9-46ef-9a11-eeb195aa3bcd · inbound

ROSA: Harnessing Robot States for Vision-Language and Action Alignment cites this paper.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.745215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.745215Z digest=sha256:1755c6f0d6e17991d3fd9b23b5fda9f4367b64c6cef8aa5c7b529fde02b5013b

Observation 7750dc80-a3b8-4e5e-b028-6508e0dda973 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:42:41.600535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:55added44c5bbddfff401d9d937e2f74668c536128c5e8831c587fa2465e5b0f

Observation 7e44a777-c2af-4e07-a7c5-f5b07db981bc · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.340673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:ab0f7f3716741cdda45f2a3d34ba95a10111736c22f4a2b3f665f7e1372b59ce

Observation 9a1b0fe7-1d5d-4f37-8fab-cf43743bbc1a · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.091468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.091468Z digest=sha256:7be0619d6f29d9db1de27173fd5249175a886cc6621613ce770a3326148e0c56

Observation 5590b66d-51b2-45ab-8e61-386d196ec944 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.260522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.260522Z digest=sha256:390074edeffdec4e0a24b0b5d59de10a2a73ad22238cd35aec619557b28ca6e4

Observation ee066cac-abb4-4331-9d3d-4ca7f09bc4e9 · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.078793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.078793Z digest=sha256:34304492f4e2693694c567d05c6166dd5ebc2e343f142bc7e63149c9093757fc

Observation 95a960aa-0b45-45b7-816e-f31416027001 · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:50.518362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:50.518362Z digest=sha256:35782b8aaf2d6cbede2d98e8b4a08e72009f457e54872e2e03e735be4ffc227e

Observation 81a5e571-eef7-4a7d-8bfe-cc1b12e59f34 · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.357317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.357317Z digest=sha256:e16df2d19367be63392bda5bb25861831dfca35a37024b9053bc490018aaf49d

Observation d35bf74b-9c79-41c5-a2b3-3b570a5605dc · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:53.480873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:53.480873Z digest=sha256:537c52382afc06089409459fc984f73ac1bdb0c240c0f686510239cb41a853de

Observation 714b5cbe-822d-4707-80fa-1c8fb1481890 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:24.338053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:24.338053Z digest=sha256:cabf105001f0202a526ed4dd9f4a8eb129a0f8823f5ef6fe640ae4add4f865b6

Observation 9970abb1-2bbe-41cf-9bd2-0c321d4a3f31 · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:48.142773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:48.142773Z digest=sha256:b50ca5db1a6ec3c7cf00de437ca66b82d9125e31361cdd9aa6cc67de26f902f2

Observation 52e7d8cf-2367-4189-8a01-9161a1456139 · inbound

Causal World Modeling for Robot Control cites this paper.

Causal World Modeling for Robot Control ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:53:52.273673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T13:53:52.188890Z digest=sha256:33039be548ee8cde5ed06e898245a373809ee0c442a31b90a7896d0d17f948cd

Observation 4603c4c1-bc1c-4a54-ae48-c84a8ea3e876 · inbound

Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation cites this paper.

Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:21:31.938466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:21:31.938466Z digest=sha256:f77270ddf57d61fea6c48fc2d153aaefb04f0ca5b63ded5cf1cce4f401705257

Observation 25a6971c-98df-42d9-8f0a-5688ce22e5dd · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:20:17.660663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:1426ff077c925b960ed6ba02fc46b48d4d386b5ec42b96c59173d7dd39ee8119

Observation a3ede2b6-4afd-416e-8adc-8a04621e0239 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:45:21.885175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:c1153d31cd9d34ed096a0adc9e0479035c64a38c3d51cc625036d9d93c19a37d

Observation c89e4ba4-0911-4812-8646-87c3cad1273c · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:48:48.378335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:289e255a2d86b95dfb9c1d01e4aa2fecd3027289ff3e9e0e6fce9fc450a06728

Observation b41e6278-0791-4b48-a7c1-92779fb0379b · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.599144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:05c740a3d9d406ad1ac0dc254c5a5818f98b8ff4777d5c5c7d17a5eb694520f3

Observation 2601a711-a270-4ee6-a3b8-75518bd1020b · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T14:25:59.915385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:25:59.915385Z digest=sha256:b8252cac1e70ebd88988d777eb8356ce0575b793f86059b9ce5dcb6509bc98f3

Observation fa7d8c17-1731-43a7-b7fe-455644114272 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.091658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:09870daed63d07f44dfa653559b368226d2f60dbbe934d1b86e6b2de49cc66c8

Observation 94d3015b-a23f-4590-b598-68fbad6a88d0 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:53.160490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:53.160490Z digest=sha256:94b8eeb80d3de7ed6b206249ddb18e3b5fbad2231c34ac980a8ffe29820feae3

Observation e17c54e4-9524-42c0-83ab-f12a4be8cdd6 · inbound

VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving cites this paper.

VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:43:06.012813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:39:04.567741Z digest=sha256:0b0b9ea45b849793a4ca5c4791c43e481564c028b05e973d05abed89c1cac212

Observation 2ef5d5d5-c09b-4a28-9d18-af9f70033dd6 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.764526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:d6c4fc0f60ae1b4d8e4f0784927235540b0b4a960a5ef3578fb31f0abdc75bf9

Observation dc42698d-7344-40de-8ccc-e74059df9d02 · inbound

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform cites this paper.

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:56:35.273554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T09:31:56.289129Z digest=sha256:30ffd4055500a8f1bbe29f0c91ff95e93b2aff8048e50e700f9e7975ef904a26

Observation 5d4e160d-c8c2-441f-ad35-0690867a9bb5 · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:43.132982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:43.132982Z digest=sha256:9e137a76458d87d6ba8f40d6a52f2ff02160087d7791169fb20cac8fd6aecc18

Observation 084a202f-25d2-4bdf-8987-b7df96f0cfc3 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:43.165360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:43.165360Z digest=sha256:be1b658b27e131e06c709d12b3f2a76c24a0a68bc438e7119af49538cf7a113e