Pith. sign in

Paper Citation Record · LEDGER

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2402.07872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07872 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:44:58.996315Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.793749Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2991d99c-93e6-4bb8-a16b-cb2a4873dcc4 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:53:27.756324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:ec1223a2057c32b005dab02fe6bf9190eb4ed58157d42d46b6867edfd0982588

Observation 8658726a-81ab-4629-9493-68cb57c47fea · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:17.946537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:9ce86a3c435bab0eea6ba00e8a0ec2425f8765d19f183379ef432ce53e51da73

Observation 9712e9ff-b380-4a6d-aeb8-f7cbc3a4af07 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:22.974511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:4a8f795a8266c8ee7d73082be08cfb22e0bf49ea3ff7b5096870141deff41c0e

Observation bac584d5-64b5-4b8d-85e2-2279e245f1fa · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.226372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:787ef06aaf39aac00dd05e01194b6a2343eb64dec716a3d10da9f23b7f3e1d5b

Observation 395c7c37-590c-48b4-bf69-aaf3dd4bf8a8 · inbound

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning cites this paper.

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:27:17.800649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T00:26:58.273861Z digest=sha256:9f419654a2c5dfc233abec448e6872d4cf781e227b337d57b54dfa7a72c67a73

Observation 753df86c-10a5-4030-87b2-d0e9b3d2b2d4 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:15361e927dcd8567439743578772fd029fe3f7720e111d7260f152106f562f91

Observation 844b8e2d-5cc2-4da6-a348-74db30750c5b · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:58.996315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:58.996315Z digest=sha256:2a79b33cbb792569b5f01105cbd32b6e6c549b7609b82b895c7b629b1fe934cf

Observation 74a96285-36ca-470f-9429-f7ac7f3740d6 · inbound

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT cites this paper.

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:23:18.991949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:23:18.991949Z digest=sha256:7d9e3a26a520f96fac83f9f933cc6b1dae5f4c4328bf7b50fdbc7f53b7e6d0cc

Observation 2494ea49-5bac-4ac1-8540-65685f94adef · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:28.691845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:28.691845Z digest=sha256:b01cd78a5769300c7874a69cccfb7f87c134efd777d34934e369a39a5c673c1b

Observation 3052474a-8db0-4b40-ab2a-d602a08bbd40 · inbound

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals cites this paper.

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:06.276723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:06.276723Z digest=sha256:6351a40a618b3e17dbbeea18a34852386106a594a633c3b2cefe075777c24fb3

Observation 2ace252f-c4e0-4843-b1d3-24b59dcc3d43 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:52.414897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:52.414897Z digest=sha256:ee551a5aa42a0e4b9fc833fb420bcf4aeb479e34ab4cd3738079328d89663db1

Observation dd755236-743f-4464-9e78-88eb3b1ba4a9 · inbound

EVE: A Generator-Verifier System for Generative Policies cites this paper.

EVE: A Generator-Verifier System for Generative Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:36.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:36.225511Z digest=sha256:00b64edd5cfa2c0a02a2583e6418e4a5521536448ad301d307a3c875366c6d0a

Observation 6693fff6-e3fc-4240-9e81-4f5b1e23c063 · inbound

Visual-Language-Guided Task Planning for Horticultural Robots cites this paper.

Visual-Language-Guided Task Planning for Horticultural Robots PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:58:53.900823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:58:53.900823Z digest=sha256:ccd03f976d6dd66ddc0802a64af7c4955e789b20f8ad3f03d84eba1d305b4cca

Observation 3047f114-df0f-444d-9136-f0d4615c164b · inbound

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors cites this paper.

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:21.189276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:43:21.189276Z digest=sha256:be852ffc2e5fc60fcac06525f4271b940bebf05368ab74a3acbf315f45fac614

Observation b07d05f7-0c8d-40d3-97d8-29f0648ac513 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:2e5babb2244c8dd891692107642cd273b2cf1a06bf2c829b95a72b9a0000e136

Observation 33c2c122-6f5b-4a7f-815d-bf60d59310c0 · inbound

JailWAM: Jailbreaking World Action Models in Robot Control cites this paper.

JailWAM: Jailbreaking World Action Models in Robot Control PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.143404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:22:13.890495Z digest=sha256:e492f3b3d1e529482445527b65d195ff24c35339303e6ca93056bbbfa47adb6e

Observation a876af4c-f548-405c-9b3d-d3731957eca7 · inbound

Visually-grounded Humanoid Agents cites this paper.

Visually-grounded Humanoid Agents PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:04.606662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:10:32.710853Z digest=sha256:06a1e6e3d2bda49e47ec75c6a4854de2d288d324e4f40f55594683219cacdcf5

Observation 4942f695-e5c0-4837-bbda-d22b9cf0164e · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.824129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:d0f889eca5849df309cabe902c18aa1574338084b76a3cdd4244b71ba4942c8b

Observation 6b6447d3-9169-4cee-89e4-3107a2ef1ef0 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.256553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:f1085e75186975fe62c3fce08b345f6b5adcd640804be10391e4db58d095b822

Observation 656704d4-258c-4f0b-99f6-d44d53412430 · inbound

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning cites this paper.

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.795525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T20:48:49.071281Z digest=sha256:941c265479d8f07f1cfe056ac404401a537f5f168e0b3a95f35e2766ab9c383f

Observation bec7434e-6f77-4818-ba85-a7f72b2950db · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:56.957400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:25:21.796778Z digest=sha256:d7d31a6bfe84c57b8b0f59e99530d17272cb9127bbd2171c966d651a066e43ce

Observation c2a296ef-5bae-432c-a4be-7fbb039a02d8 · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.671107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:54:50.393828Z digest=sha256:31638a2eb3bc58906bb0bc5bf40dec5c070cecc29957ccef18b7786bdf741734

Observation b385d68e-4993-401d-b552-27f589b8085f · inbound

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models cites this paper.

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.638351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T09:10:55.016700Z digest=sha256:b5691a929dfaeeee534d8eaa9e4bb8c60c2e65997c0a8b3b1c3a31caf3fc32d8

Observation de5719f8-13bf-4504-889a-2a8122df22d1 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:41.952355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:17:25.534715Z digest=sha256:23ae268249a0bffec469a50acc3d4e142efcbb7da436a432b5e3c755727ca588

Observation 5f1c91ea-0129-49e2-9af2-8b921cd5b863 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.682933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T19:15:49.118807Z digest=sha256:1a16a79a3b0aa792af58292da9f7ce9c3c8c36f656986ac8e2b9a5874f253c68

Observation cd4561a9-4371-447c-8fda-5928d69a8e9c · inbound

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies cites this paper.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:fc449900ae7e05b6c215da51fac78c45ed9472c151d51af6f4a6bee3b07ecdbc

Observation b67453e3-438b-4aeb-9b8c-403755cc4fee · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:13.878543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:13.878543Z digest=sha256:5c61ac880c416835ec5c98cfc755c280753f148c36ab45be0202844b36a69519

Observation b2cec11d-36f2-49d3-9adf-56765bcaf9fe · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:50.899469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:50.899469Z digest=sha256:891d673f257dfa314c3ade7f4b13522d2a24330927c8a9622ca996f023574f17

Observation 9bb385f8-e7c9-409c-9286-d8a5d9ebee03 · inbound

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models cites this paper.

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:59:36.361796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:59:36.361796Z digest=sha256:8d1f7105856716185e7b5f5a16135c7f963b4cced52bebdcc05e11fe680fba4b

Observation ce1420cd-c1cb-421e-9729-6b2f8b0ce335 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.094779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.094779Z digest=sha256:d0f6accfa5632e3fedc4648f5a943f1c661af4218b151456d4da37556b205412