Pith. sign in

Paper Citation Record · LEDGER

Action with Visual Primitives

As of 19 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2605.22183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22183 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T05:26:55.722814Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:36:28.020890Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact38
  • verified fuzzy2
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c2219a6-8813-4352-87ac-15de83e743e6 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Action with Visual Primitives OpenVLA: An Open-Source Vision-Language-Action Model

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.107209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:a814771a1d387772521ee7a4d3f7ab71d93845275d92e70eba3614b31bf4dd5d

Observation c80375f3-6376-4c18-bb32-882db7a75901 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Action with Visual Primitives $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.100163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:145ddf8ee72112b48beb6d03d6c3961188007c04dfb1625fb4b75ee9e632e2db

Observation ac06387b-904e-4b19-a2df-5e9ab0659af7 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Action with Visual Primitives $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.035463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:a173252188c6b0a7e80e51738e8fd202adf55215745002da6f3908517f297152

Observation 201d4021-2e41-424d-a8d1-e8799982bbe6 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Action with Visual Primitives RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.040865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:a670bf5ea94297c838a491081825e2539f948393af40f9f33eb6fa83b8fc8ba8

Observation 8ca54361-1a53-46d2-acfe-1014a844212f · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Action with Visual Primitives Octo: An Open-Source Generalist Robot Policy

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.024464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:35b2ae8670cca521511bd18f8735d742dee84399486e249f4eebb531d63bb136

Observation b341e1b7-c2a2-46c8-898b-bbb6a2b1e15c · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Action with Visual Primitives Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.079597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:17e7f40eb08da60cf920dba2c846fc236c08bccf613133071635a85283e22907

Observation ac534265-0617-44f7-9020-9dfbd4aaef98 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Action with Visual Primitives DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.046556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:03094f86fb59865ff5344fcb5d48bd92de45bf16e9fb68d54b6d5a7b753586d7

Observation 679df4ea-e655-48b8-a82f-67a64becc654 · outbound

This paper cites Walke, K.

Action with Visual Primitives Walke, K

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:36:08.408336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:79ad2cc0e379b9b4f5629eaacd219b622cf05f754d56831baa7617db6ca94457

Observation ef426bcf-6ff1-48ea-bfb2-639ff7a7706b · outbound

This paper cites RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning.

Action with Visual Primitives RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:08.064923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:6ea1ad4473d8e6e8af65178764c9e8d3534cd25e277a366a4455fbede738002b

Observation ecda8d99-0249-4c68-ae19-73f5f4f50414 · outbound

This paper cites Zitkovich, T.

Action with Visual Primitives Zitkovich, T

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T05:36:08.404963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:f3e750bd0d239a6e2cdcf43ceca6bb8ea5f88d5eb75ef571c23f72a08697c905

Observation ecfafcda-8e8c-43b1-a90e-81137c971f44 · outbound

This paper cites VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models.

Action with Visual Primitives VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:37.239527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:66864d44630772fd382ca63915edb7c5244591af36715e71a8b604c30a483c7e

Observation 865aa20a-7832-4a17-9987-8e0f3da63a87 · outbound

This paper cites Kachaev, M.

Action with Visual Primitives Kachaev, M

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:08.051621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:4b4658efbfeca399de1c0a17b042a173ad3070dfdaf6fd20924563be59e64326

Observation e5d81aae-4026-45a3-a71b-61037fcd3fb2 · outbound

This paper cites Ac- tions as language: Fine-tuning vlms into vlas without catastrophic forgetting.

Action with Visual Primitives Ac- tions as language: Fine-tuning vlms into vlas without catastrophic forgetting

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:08.094132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:1ec64193ef6dac74e055111d9acf05b65aa11eaef0a4af610d403f22d0a5b84e

Observation 7d750c11-1fcc-46f2-9bb0-f27002834235 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Action with Visual Primitives RT-H: Action Hierarchies Using Language

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.073565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:cb3bc92b92d60f248669f2b5f931f7bd5e4077f9d2a3aab3f0a855f3ea17f875

Observation a4f803d9-cfcf-4bf1-939a-a1b638b79833 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Action with Visual Primitives Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.115243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:4aaa59a3a5ae238e2c3f714a3e6a421b7e2e20822b38737745907b15bc680574

Observation d40a8543-9797-45e5-a69a-ee0055a5ddb2 · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Action with Visual Primitives ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.057225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:f06c8cf9413103e0ec4c84268e195a97e57c2ebd8021af05ceee1782ffdb388d

Observation 21dab57a-3b9d-45bc-8b77-312b3702150b · outbound

This paper cites Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability.

Action with Visual Primitives Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:08.030793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:f4fef12c1fe017eab730d5c4a74837b4afe5d55ab57493c4bf88aa94a58e3332

Observation f193b8b2-3c19-4e9f-ba7c-f969b89d8965 · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

Action with Visual Primitives Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.019015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:a8a9a58c263f836d087e4e904d81f350d917eb9e1912af5708098489523d3404

Observation 78e2ca3d-685c-46f5-9fb7-9f06cb731cd3 · outbound

This paper cites Point what you mean: Visually grounded instruction policy.

Action with Visual Primitives Point what you mean: Visually grounded instruction policy

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.976747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:ff3e81d603c965a3a65028011eed2ed21e2128f86439b13d7c70262e1185d8d2

Observation e6cb6016-594c-42ad-9e92-3c5cf079da9c · outbound

This paper cites VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models.

Action with Visual Primitives VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.991337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:b075572d477fc60d79a38c8e66cef0f5efce039ca1f709041f9082d6832514c1

Observation c03c2516-f793-4f35-93f1-639370ad4df9 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Action with Visual Primitives TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.008286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:71b6cbe91b640d4fe84f26b9f8e28efdd1672dda9f1efae591d593cbcf65f689

Observation 285f02cd-8199-483b-92c3-622abfe45b1c · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.401845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:3d9f06d84a20fdcdcd07e97a10cd1e6ba27d6fefed451aa462e06e6682af9061

Observation 802a53c7-6813-4fa8-b47f-ba70079a3363 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Action with Visual Primitives GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.956335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:f9ba4da766c7aa48454d185891d048a0c34ccb8a4a8d00a4ea0b650ae8cb6d04

Observation 039a26b0-f02a-4136-af11-5ea24230d9ec · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Action with Visual Primitives RT-1: Robotics Transformer for Real-World Control at Scale

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.981436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:68a48d5e7e0008162fad6257b8cef602c4c8e68e5918be2ab9f9578ed4c0d3c5

Observation 6b265a3d-85be-43b4-bc1c-319c77609ecc · outbound

This paper cites GR-3 Technical Report.

Action with Visual Primitives GR-3 Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.013331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:f09f3f5ef5106f3ef370955c9c831bcc37cec294cfb86bc150126a16fc664a46

Observation 0e1a9a18-baab-4874-9c45-2af5e57cdd28 · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.398742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:c58895f592e703cf71153b05653452e9a4ec085a23754713da95d4fbd8e5a38e

Observation ccdb12fd-1457-49e2-96f6-9e95b40694a7 · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.394347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:3a858d7923cbae9be16e1d3722fb1073a4122daf88d2a4eff3715bd7728aa94c

Observation 9065a5bc-694a-4060-9b51-5cd42a0f5bd3 · outbound

This paper cites Yu et al.

Action with Visual Primitives Yu et al

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.946010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:23c1033275df7ad2cb99885108eb900b7908082317b014d565a805717ea0a0ad

Observation 8b5e1a92-777a-48d7-884b-afc4144490e1 · outbound

This paper cites LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion.

Action with Visual Primitives LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:08.962760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:473105d4726bdcd3ba319278e1f295bcc84ee82084d77842f296721fa5d02707

Observation f00af313-b15e-46b7-8453-dd350a8d66ef · outbound

This paper cites Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416.

Action with Visual Primitives Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.966330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:2980baaecd479f09651675fd1cf9e6e1cf1c5d0ec3678e1910ec091f5c466aa0

Observation cef54000-5c93-4d58-af0b-990bc5452eec · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

Action with Visual Primitives LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.916410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:1b37409f56986aef88dc3dfb647918656ed7a234896f810e6dc3dc8fba72cc97

Observation d2a6fc74-dcc9-444f-9532-58f6e9589845 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Action with Visual Primitives Gemini Robotics: Bringing AI into the Physical World

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.951707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:d19eeff91e6238eb15980339c94abe9a2617478b787e30325d652abd9f265313

Observation 33e0a215-f637-4466-a354-1dc414125d1b · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Action with Visual Primitives SAM 2: Segment Anything in Images and Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.906937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:8c39467c18bb0d1858643ee804cbf2bf651a0a92635748145a6e3edba0a9a637

Observation 56fbf6c2-eae6-4bad-a7a8-8d7d541bc1d2 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Action with Visual Primitives SAM 3: Segment Anything with Concepts

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.911468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:da3f9525d6ea290d583345c32f83bb49659745498d111299338966d5fdee899e

Observation fb546123-39ab-4148-8813-174612dd32e7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Action with Visual Primitives Gemini: A Family of Highly Capable Multimodal Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.940275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:2e457ed445fe8ed0d8587bc98fd13f3d628d6936f3d2964ae0267994145e0220

Observation f1232b0b-c5c7-49c0-8c88-c3a558009356 · outbound

This paper cites GPT-4 Technical Report.

Action with Visual Primitives GPT-4 Technical Report

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.960837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:6cdae6524ece0338f1e1f67e32f6ed33f4733ee30cc2e81ae580cfff0fe3dc3a

Observation 1adf1113-343f-495f-b7b2-54c6f446b331 · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.390750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:3f3589ad6bdf8a41e96bdcd60f1583da571a331a99ac693bb19801856a998c8d

Observation 4420f193-103f-4d8c-af71-be52ecf4920b · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.387116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:d24563fb7e3d5a298a041d6524a818d2970cc207f102c9d1547b93d783c6bb34

Observation 5724532d-63a9-4b65-81f3-4873ec79ff14 · outbound

This paper cites CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation.

Action with Visual Primitives CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.996827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:e107770da820e9f9c8b12750da77cf97e50614042bc0225a7117374f97daec8b

Observation 6b1b35f4-0e36-4d1b-ade6-6ba1fa3b910c · outbound

This paper cites ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation.

Action with Visual Primitives ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.971051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:ce57a2470da5725c1f8b16f2c83e65ea09c215b0259c28a2ccaeba35c1adfb39

Observation ad897b99-3318-4711-8225-62020b69ddb8 · outbound

This paper cites A3VLM: Actionable Articulation-Aware Vision Language Model.

Action with Visual Primitives A3VLM: Actionable Articulation-Aware Vision Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.921165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:d831ee45fddf6a97eae9da5f6afb18b967d6ea3e86463e247fc6fab122c36d84

Observation 7bf4e6b9-f7ff-40d4-b673-ba1f8a125b34 · outbound

This paper cites Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation.

Action with Visual Primitives Sofar: Language-grounded orientation bridges spatial reasoning and object manipulation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:08.002642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:02f5fd03643869e4b4e7aeab953663f5222913899611a579611553d15e2a20e3

Observation 9b7624af-d98f-4ede-99dd-74b257996845 · outbound

This paper cites an unresolved cited work.

Action with Visual Primitives Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-22T05:36:08.383670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:c4275190d4991201c28a64df4c761c0d707bf7ca7842bf0790941d1390b7afe5

Observation 6de0b034-d246-48b8-aaa4-5591af29bbb8 · outbound

This paper cites See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation.

Action with Visual Primitives See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:42.139620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:4e3ce1df557d680a8c2abe788129c230637c7925dd9ed9562963ca1ddc9adae3

Observation d6c68e6f-255c-4875-8dff-e53464f1ccd1 · outbound

This paper cites Robotic Visual Instruction.

Action with Visual Primitives Robotic Visual Instruction

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:31:07.986405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:96f3037634808cde7aed81038fc52e55bf7ef06e2c885d373c7d1f24bbc8c878

Observation 32a3996b-44e4-4cc7-8abd-4e469a53bff8 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Action with Visual Primitives Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:07.930649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:d0eaee07622fc9cca2646a4150c0b27ef5d326980b1b179fed64af90ff6dd4a6

Pith citing papers

Observation 374d8661-814f-40ec-b1c2-2dd2de87bcf3 · inbound

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation cites this paper.

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation Action with Visual Primitives

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:28.020890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:28.020890Z digest=sha256:f7ee9b4ef3b0a5b02f18566b3172b6d494c323b23c377b16549982c91c5edf9b