Pith. sign in

Paper Citation Record · LEDGER

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2505.21906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21906 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:24:46.040172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.316884Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:4f657604dbd37824568aa19f8be7d107a569d0d403246cbad7f3628fd9ab3a0a

Observation a0e2b8c1-28e1-46ef-bed6-456df92c9d11 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.586386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:3f3c8f540ef9c8798f1b43cd08cee7c259f217bd52c4d0de4cbdf018a291cd7c

Observation 3d798a09-016e-4b4a-badf-e51fd4ed9225 · inbound

RationalVLA: A Rational Vision-Language-Action Model with Dual System cites this paper.

RationalVLA: A Rational Vision-Language-Action Model with Dual System ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:46.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:24:46.040172Z digest=sha256:70365382180d6ab77dc549f1959ab1c750312523a28a1dd10b0ba8e2c2bf3675

Observation 91bd68e6-817b-4383-90fe-56f24360044f · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.346015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:9470cc1ab0be35f7768e5f48ff57433963617557b601f06a3ba316e10816ce67

Observation 6cb31c1a-69ee-47cd-914e-be08a3b32bd5 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:28.034660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:28.034660Z digest=sha256:4961b5a59a36339676e16a7722df5acd3d8b643080c522ee3e40426efca5ac2d

Observation 6a9e238f-c452-44a9-ac56-732d17bd6da2 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.693054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:34f6c2d5dc37c137d7fb0b638b5d047689846620ba951c2994ccc46536103223

Observation 4cfccbd2-71cf-4a53-9f38-40d77c92705b · inbound

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning cites this paper.

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:51:17.365848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T16:12:45.211236Z digest=sha256:9342a453d12c64b2d2138f0c5f4a3bf42f68871793a80c2505111963d50654ff

Observation 51b4255c-f5d1-4777-b590-5064be053c6e · inbound

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models cites this paper.

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:21.135705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T19:45:10.232982Z digest=sha256:fb4d980b2daad1f075009fc439ab07f946c5f970f30aa4983491ca220e03d542

Observation 20f8ca47-5f32-4d55-89c2-651193f0b6ff · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:46:04.954056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:3f9a9857076367cf5595663e3b01790264bfa2ddd9c1885cee46fbbaf2fd0410

Observation 1eb1fddc-ed36-443c-8628-0548704f57c9 · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.083383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:b3c340dda13ef8d9411ccb3cf3d052d922fba9dac2b4bfa8c749433dc2eb6a14

Observation 999a750d-1f10-4cbf-95c6-c0e76e5ec01a · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.282356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:33635df89158b0fa074caac58395c6647b208190f56551f8304c627990dc3499

Observation c4fedb81-ba07-4d9a-9aef-731750d8ad25 · inbound

PhysBrain 1.0 Technical Report cites this paper.

PhysBrain 1.0 Technical Report ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.958760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:34:44.204055Z digest=sha256:8f13834455b6fcac2ea788c3cbc4f781be0743254dd6ec1e2298c23d6c34d9ba

Observation 39aa038f-dc04-48fb-a73a-f3a8e39bc85b · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.242721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:8058cf2038ff9c52d6516a9bc3a9e64bd897706f0e1067e250764fe9effdcdd8

Observation e79f361a-6613-4805-b7cc-5660a41a7c0e · inbound

Continuous Reasoning for Vision-Language-Action cites this paper.

Continuous Reasoning for Vision-Language-Action ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:46:10.541501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T22:07:07.401335Z digest=sha256:61cb12ab91b54aa8b749cfac977933c8bc992c2b8174c3f28c28afc06d20d4cb

Observation 13aab588-7370-48d8-8881-0bc4931c2e1e · inbound

Policy-based Foveated Imaging and Perception cites this paper.

Policy-based Foveated Imaging and Perception ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:17.704616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T15:23:48.688620Z digest=sha256:2d61c0bf2f59d214f4e86731f6eff4023666fe8f5accb1dbef73f3f51f8ce196

Observation 12d49e0d-37ca-4ab3-97e5-d7c91875184a · inbound

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform cites this paper.

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.237371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T09:31:56.289129Z digest=sha256:b0832c7f3f8c2a61583f9b563e21eadf63cb2e3d50a5ea332e5478371ace3fd6

Observation 668abe37-031f-4b57-9a78-5655687e2540 · inbound

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation cites this paper.

Bridging Semantics and Kinematics: A Modular Framework for Zero-Shot Robotic Manipulation ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.318532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:28:33.173481Z digest=sha256:9211447e22f41779fe64abd6de9356e247c4a7f9c4b0672865a47afcd5c08b52

Observation 89823bb4-2bfc-4161-b2cc-78d9df11fd8a · inbound

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach cites this paper.

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T19:52:24.975588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T19:52:24.975588Z digest=sha256:1d800724956dc6625424eff9c8f6581d38713d591d4d65950ff57692d9b42781

Observation 731e90b9-b286-469b-b613-037f883a850c · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:43.005296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:43.005296Z digest=sha256:a809f786d96068a816fa338c7e86d9a5e1fc2df71aee44a0e0cb8cfd20afea87