Pith. sign in

Paper Citation Record · LEDGER

Open-World Object Manipulation using Pre-trained Vision-Language Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2303.00905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.00905 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:51:57.719892Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6e6e8e92-5167-437a-8f8e-3ba29a7018f1 · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.570726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:e5e799fdbf5cb349dd012810944415897a05aadb473618b6d5ff38eba268b739

Observation ab904600-e585-4a6e-9663-7038b8ae26e5 · inbound

Open X-Embodiment: Robotic Learning Datasets and RT-X Models cites this paper.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:23:24.351434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T17:23:24.255829Z digest=sha256:e3e2cde37cedd7f905c91435cdab8f69d3d1fc3b7e48302d1bb57a965931549a

Observation 9e46c221-234c-4ae0-acd1-72c5d6fb3d78 · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:54:59.132287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:ea02fa17a8568af77ee485ae41962802a4bd8cc67ee6b2d9e5165df1d7d69084

Observation 9435dbc9-57ce-45f9-9a79-2a792a253746 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.465586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:828ee02edc517197274350eb6f5ad30a0699c3f68b0ff60d06f8cda8d36789d6

Observation b8513001-5337-4b36-9899-284f6eaff38c · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:46:36.281682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:4f475b98faf5fb32c86d26c73c5e413807754400698bd4a3eb65f3de1bfe45b7

Observation 94e238f7-627f-4b2d-884d-609a3aaea92d · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:27:22.835880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:98f67c83504c4f1f09d7090f0bc252fcc88a9884f626c530d1e44b34046fa130

Observation ca32e68d-fc72-41e1-957f-8f24fc0e225e · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.275383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:33f9f6291c98847b3a549cea9fddf1c22e7308cb2589e532e4d37352ead6d22b

Observation c1c0d234-863a-4774-98df-cc50b82325de · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:32.851181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:43de5bcfdb53143c1cb6e2f41f5fc9dd5923630bf8b96c21461412719aa58d10

Observation 8a97b39f-b308-47ce-a50e-cf8997c5a9b2 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.899383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:0060817d01e7686a7073744d7abaf9d7ca4240b7f47bf421d41fcee5ba16f3b9

Observation 3f51432a-c1c2-4e04-9d52-b14534762b3a · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:52.251612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:f286f81747327c46af32165835787faac53b74c0947320f8e8a1da24d7161aa6

Observation f27ff8c3-ca20-4a21-aa1e-46a9ded69031 · inbound

Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors cites this paper.

Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T22:51:57.719892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:51:57.719892Z digest=sha256:d9dfff30ecaf531f27cd429c90c08fa9b1ccff5c400d2aa81e3702044778f34b

Observation 56f02102-3db8-4d18-9986-01ad136caf58 · inbound

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception cites this paper.

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:05:15.672460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:03:28.341042Z digest=sha256:c94bee1722a6414ea00346bf88f177c6377ab604a2d4b6a35bc9474682cf41de

Observation ed491bfc-b2bb-4772-9e84-699379fa0973 · inbound

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics cites this paper.

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T21:43:09.897136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:43:09.897136Z digest=sha256:dd90faad854cb8537948c639225b5e9882e5c09f3a9352a5aed0bf0c1b7d06b5

Observation c80735aa-51fb-49ea-b4f0-a69e76cfb67e · inbound

Tactile Modality Fusion for Vision-Language-Action Models cites this paper.

Tactile Modality Fusion for Vision-Language-Action Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T18:13:10.246106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:13:10.246106Z digest=sha256:5c0b398f7aa4b7012830bf152f5e21729d570ce17a391f695395d26c2c0d5bcf

Observation 35314711-93ca-455e-b754-e5bb313e6b82 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.565939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:871ab5372e0e8c9aa1b91b79c0125dfc8e6c7a24cee18f9ceedc5c88eac09757

Observation 74c964fd-9da7-453c-99a8-06e4da12d3dd · inbound

Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation cites this paper.

Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.233360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:51:05.220627Z digest=sha256:051fec4867359d7f99f3e5b8ac33628c4e1ec068299d79ded30185f4ad458f5d

Observation 9294b453-f12e-4635-a6e4-bc7a40285662 · inbound

SEVO: Semantic-Enhanced Virtual Observation for Robust VLA Manipulation via Active Illumination and Data-Centric Collection cites this paper.

SEVO: Semantic-Enhanced Virtual Observation for Robust VLA Manipulation via Active Illumination and Data-Centric Collection Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.744611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:50:35.117038Z digest=sha256:17d84ea050a8a238721a9307df8fe41f6ac0dc05618e9a4cc620e5dc9bc91061

Observation 3c615202-ee07-43c2-bd21-08cd1dc63ac1 · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.889350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:45061c05d1012f0afb33306de882bc18dcbcc0b11b7e1f3a339f7874fee4eb19

Observation 9be766c6-d92c-41cf-a173-1d58118890eb · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.300839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:9d4a39354c65bdb5461adb247a5139128a3d1d41ac744aa5acaa9ab2b0651431

Observation 3da45ba8-22cc-43d9-b07f-ef683d0fb025 · inbound

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning cites this paper.

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:33.491536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:46:45.209053Z digest=sha256:7326ec39f8983ef0b9209abb63245f3d02d5f8da29c39cd9dcb11da03d30426e

Observation 2740f7c6-48d1-4492-b47b-64ec4d8d0694 · inbound

FeVOS: Foresight Expression Video Object Segmentation cites this paper.

FeVOS: Foresight Expression Video Object Segmentation Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:30:08.043849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:13:21.500674Z digest=sha256:0658b380f6f2f0aab652842314fe7330d33a065a9704cd1513aa54ac419db2a8

Observation b1813dd7-829b-4727-bde7-d7c5c2cd8a35 · inbound

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure cites this paper.

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.140694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:32:52.472573Z digest=sha256:4ccddd0b67907cb350a64c7c39625c831e989ba3e6f5f9454fe44a86d14a7717

Observation 7382167f-f5a6-4da5-91ab-de49c429aa6d · inbound

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution cites this paper.

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T02:25:55.766741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:25:55.766741Z digest=sha256:b29f735ddd2eb9d94f330bac0448b3a1d4f506e0a551cdbb37c7e6aa5fb4a237

Observation 6ca2a3f3-8f35-440b-9e23-158707323d9f · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.346741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.346741Z digest=sha256:a311ad591eca1ac0059411aff550e84efb68e897acbb314a6985cfbbb0582ac0

Observation 64c74535-d0a0-461b-8f49-1c20b198fbf9 · inbound

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models cites this paper.

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:10:34.688795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:10:34.688795Z digest=sha256:b6b9e730fc3b7d1882f8de3e8ecb165dddf599d0debb925d02ae78b3f2c85db9