Pith. sign in

Paper Citation Record · LEDGER

VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2410.13860.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.13860 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:47.340374Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:59:20.317233Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0179f253-307d-416e-9eac-3bcda0740004 · inbound

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding cites this paper.

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:49.880892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:49.880892Z digest=sha256:a1917c594f52097fbde8f0aca3a814a6fb52aeb5ef386ce5abf43e8988191217

Observation a12b5421-9868-4975-bd19-3f873b99747a · inbound

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues cites this paper.

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:35.395155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:36:35.395155Z digest=sha256:0cf359942ea2520da4136dcec3d8a1eda8ef252279597548af4b0e8e73e7dc20

Observation 9e741549-ac26-43a8-b61d-56ee16b6232a · inbound

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models cites this paper.

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:47.340374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:47.340374Z digest=sha256:d1acf71532f58ad22278f23698d659d7626e70da528b37c66252e8dff3ae7b2c

Observation e280d0e7-be15-4340-a6c4-353c5aaa6c8e · inbound

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models cites this paper.

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:16:31.258432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:16:31.258432Z digest=sha256:1809f9376b47391d5e41a94dcc6dd8162dbde18675d67682e0e0045e35836444

Observation edcfabd1-07f9-4cd5-b156-2823a5aa0bc2 · inbound

Zero-Shot 3D Visual Grounding from Vision-Language Models cites this paper.

Zero-Shot 3D Visual Grounding from Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:19.231195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:19.231195Z digest=sha256:1897634831e57681bbc46512deac5d1786cb0a448bb7cf6dc77f9086818e6dc5

Observation 496689e8-17f6-4ea9-a748-b07daff7a766 · inbound

FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding cites this paper.

FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:02:00.869430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:02:00.869430Z digest=sha256:1f3598298a97a5d33e7cd1903a21beb4e97c72eb5c1a8fdcde88f760e8a5b58f

Observation 6b33626f-3e6d-49ac-a012-b6b4d3b5ab12 · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:08.430220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:08.430220Z digest=sha256:c4c62792ec2a499ad5f7333d14a8a32665720b499297729b82426f15cfce8258

Observation e338d33b-6e96-42b4-afde-065288427dff · inbound

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding cites this paper.

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:56.561970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:52:56.561970Z digest=sha256:5369bab634ae84ce2835f05a16d2db1992b1371cb90782ae88137fa74f28628c

Observation f4b65a09-a4b9-479e-9008-e80826c52478 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:05:10.065051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:7d66978f2688d197a45acb0c0d4e686bab3c428050199be9009db355f1d939bd

Observation 3354e406-cad3-44e4-a168-aa6be5881022 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:15.172508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:15.172508Z digest=sha256:a62c880c881b5a0e4595298276801f89dfad8b5469b05b57d5a2c0e33696ee15

Observation 85c84e12-c257-4e9a-9a62-ff8967061ac8 · inbound

QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight cites this paper.

QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:13.144096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:16:36.843013Z digest=sha256:21613b30f403633df6eb41d63dd7d45d374c395fe918ec3aee74309c774165c4

Observation a173309e-a90b-4a2b-81a9-2be06ae79344 · inbound

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding cites this paper.

Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.479613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T13:48:28.695094Z digest=sha256:9dda33f1605cfbe0e165ece42cea285457c6f70f5aefc4f5cafa5f782f46b60e

Observation a42499ad-265c-4a59-a797-7260600de9e4 · inbound

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching cites this paper.

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.439609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T08:47:25.712575Z digest=sha256:1895d33da570ce8e64277f89b674b21e7f0d4788eba5f2e2d6849657061ab755

Observation 2d930105-f2a8-4667-b03d-01ac17814941 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.672837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:d8f9f8cd43020296c80494e9816eb2cd518b585ca6690b3e601fc6c6e79a7fec

Observation c656486a-9678-4d86-8d69-55b8f97fdd94 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.320423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:85c0c686ccdf3c0dd584767cd3359eeec9b82dc82f814e89414b6183d1804ba6

Observation 75455b46-ea1f-4277-91c1-37e3a05e8699 · inbound

G$^2$TAM: Geometry Grounded Track Anything Model cites this paper.

G$^2$TAM: Geometry Grounded Track Anything Model VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T23:56:52.009530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T23:56:52.009530Z digest=sha256:2ced557783ec8fb30ac79db04a3edf356e14fb7e6b57c765f8b4921a71f25135

Observation 3adfee7b-12a5-4b11-9ef3-89a8da7c8d6c · inbound

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding cites this paper.

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:12.444281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:12.444281Z digest=sha256:b3930ecac9c7b5efbf8f2c7291560e87b5834c44d02784b1059ad375987b027e

Observation 174dfef7-f6eb-4c34-9be9-a3e11ae3c7dd · inbound

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching cites this paper.

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:45.060750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:45.060750Z digest=sha256:576b7896c380483c03b72dcb5b0dee8046d9fa1d9e6ac0355d28a36ae7017a1c