Pith. sign in

Paper Citation Record · LEDGER

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2505.21906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21906 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:46.040172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.316884Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:623ff97560d4e22586cb4aaf1cfb8b2a45bd6b9a2e2bc23b42023e8db86aa40b

Observation a0e2b8c1-28e1-46ef-bed6-456df92c9d11 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.586386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:026cc2fa61f471f9e777dc0754959d924b0069736623ab1a41be651318e17639

Observation 3d798a09-016e-4b4a-badf-e51fd4ed9225 · inbound

RationalVLA: A Rational Vision-Language-Action Model with Dual System cites this paper.

RationalVLA: A Rational Vision-Language-Action Model with Dual System ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:46.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:24:46.040172Z digest=sha256:70365382180d6ab77dc549f1959ab1c750312523a28a1dd10b0ba8e2c2bf3675

Observation 91bd68e6-817b-4383-90fe-56f24360044f · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.346015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:aa8748a0e149ce73621f827c39c9f233f245461a48d9c4f61b3303482a8acf31

Observation 6cb31c1a-69ee-47cd-914e-be08a3b32bd5 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:28.034660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:28.034660Z digest=sha256:4961b5a59a36339676e16a7722df5acd3d8b643080c522ee3e40426efca5ac2d

Observation 6a9e238f-c452-44a9-ac56-732d17bd6da2 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.693054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:6726d01b3e864c786fb4e30deac4a128955b8a7c125345f68dc16b05611d339b

Observation 4cfccbd2-71cf-4a53-9f38-40d77c92705b · inbound

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning cites this paper.

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:17.365848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:12:45.211236Z digest=sha256:d70ca3fa19e04c4d6b1b1890a1e7f062c5460a682c6a1e3d5c9ab0b946152126

Observation 51b4255c-f5d1-4777-b590-5064be053c6e · inbound

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models cites this paper.

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:21.135705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:45:10.232982Z digest=sha256:d99a55dba05ce4d2f14c2d863bb99ad39b64f982ba149a75d0b4b4c7dfb03c61

Observation 20f8ca47-5f32-4d55-89c2-651193f0b6ff · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:46:04.954056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:ce3c5c82014a958bb33bd7f497a1ebb1d3024ba92ece70dda279492a6067fb0b

Observation 1eb1fddc-ed36-443c-8628-0548704f57c9 · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.083383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:06b316977565e57cf2407065f8ddcd6f80e00242d7887827f7ae924e34ac9959

Observation 999a750d-1f10-4cbf-95c6-c0e76e5ec01a · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.282356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:97250720510964316d93f9efdccb3a0dcc4d5e3decca0b3c2df65e1df7391731

Observation c4fedb81-ba07-4d9a-9aef-731750d8ad25 · inbound

PhysBrain 1.0 Technical Report cites this paper.

PhysBrain 1.0 Technical Report ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.958760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:34:44.204055Z digest=sha256:3a6082a058efd380aad9205cda5fc502a51351d885adbbdc2b802dbc280039a4

Observation 39aa038f-dc04-48fb-a73a-f3a8e39bc85b · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.242721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:0d8c9a96466cdce752142c890a978b37154f30e052ff080a4dfc0c611c34c62d

Observation e79f361a-6613-4805-b7cc-5660a41a7c0e · inbound

Continuous Reasoning for Vision-Language-Action cites this paper.

Continuous Reasoning for Vision-Language-Action ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.541501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T22:07:07.401335Z digest=sha256:520753210fb7f23e5cb43f67f87b44356e1a4b28581a7d282abb586e13b7892e

Observation 13aab588-7370-48d8-8881-0bc4931c2e1e · inbound

Policy-based Foveated Imaging and Perception cites this paper.

Policy-based Foveated Imaging and Perception ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:17.704616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T15:23:48.688620Z digest=sha256:5015f25026ed28fa5740c041b0f7a8717876722f5c3e35021bca31e334f11227

Observation 12d49e0d-37ca-4ab3-97e5-d7c91875184a · inbound

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform cites this paper.

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.237371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:31:56.289129Z digest=sha256:609e8ba592d426d4f566b8380aefd842960c50c2b5abd2b0c4a762014346ea45

Observation 668abe37-031f-4b57-9a78-5655687e2540 · inbound

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation cites this paper.

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.318532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:28:33.173481Z digest=sha256:d2f0bf0d298086f5fcb3eed0fb9433cef08cd618e02196243bc7c988a2a73a3b

Observation 89823bb4-2bfc-4161-b2cc-78d9df11fd8a · inbound

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach cites this paper.

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T19:52:24.975588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:52:24.975588Z digest=sha256:1d800724956dc6625424eff9c8f6581d38713d591d4d65950ff57692d9b42781

Observation 731e90b9-b286-469b-b613-037f883a850c · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:43.005296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:43.005296Z digest=sha256:a809f786d96068a816fa338c7e86d9a5e1fc2df71aee44a0e0cb8cfd20afea87