Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot 3D Visual Grounding from Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2505.22429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22429 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:12:19.747542Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:41:38.649424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T06:45:29.442037Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact5
  • verified fuzzy57
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 960747c0-afc7-4659-b22b-4e9984151d05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scanrefer: 3d object localization in rgb-d scans using natural language,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:33.112862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.394174Z digest=sha256:2df58d42a439eb52ba99d50740053a4bf41b2b543f2418654b795c6012e8950a

Observation 8b9a2380-ea31-4728-a03e-027b0f870747 · outbound

This paper cites RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.282429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.502388Z digest=sha256:00e0da677f2012088960647d36bcd2fde0fc5adcb1cfbd5b2577aac74ae14715

Observation dbfea0d5-1991-4263-ac19-3d58190694cb · outbound

This paper cites Deep view synthesis via self-consistent gen- erative network,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Deep view synthesis via self-consistent gen- erative network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.596606Z digest=sha256:dc2dc6040b8bdbe6a811431d8f122b99e6b9217525c1ad17c1587b9e9f9ce02b

Observation e3b704f0-aab2-4e93-86ab-32d1e3515bac · outbound

This paper cites PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation.

Zero-Shot 3D Visual Grounding from Vision-Language Models PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:13.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:13.660944Z digest=sha256:55fdcb47d9abe1cd0953be6956f46579d3764182b49911fe73f6eec67a8cb3ef

Observation 57a6e328-f12f-4ab9-af53-7caa8174205f · outbound

This paper cites SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting.

Zero-Shot 3D Visual Grounding from Vision-Language Models SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.087050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.744633Z digest=sha256:e7ebc23beecfb353bdad34c87d74c34f981789ebbb460891a556454bf7fbebd2

Observation 4f6dd310-9524-424d-8221-e830ebea8546 · outbound

This paper cites An Examination of the Compositionality of Large Generative Vision-Language Models.

Zero-Shot 3D Visual Grounding from Vision-Language Models An Examination of the Compositionality of Large Generative Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.851836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.802829Z digest=sha256:b0b7e55e8cd328a440d42b66039a08bcb02b547d2adf11915564abc31282365f

Observation aa17f378-4420-4b33-b61d-de83946f313d · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Think global, act local: Dual-scale graph transformer for vision-and-language navigation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.731576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:13.940140Z digest=sha256:4e75fd3882ff7652aebf9330d5ae1d54fb93f95256a5e85c74d5df110b3713b7

Observation 330b875b-8571-4c73-ae2c-62fa7c4aaeb1 · outbound

This paper cites Assister: Assistive navigation via condi- tional instruction generation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Assister: Assistive navigation via condi- tional instruction generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.568017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.029025Z digest=sha256:5abcbf790c4f9a201c971a14695672e79d1bc4b95dc1db92c97c0151e2faf4b2

Observation 9f53beb9-7803-4995-a8ac-6306a5f28797 · outbound

This paper cites From Cognition to Precognition: A Future-Aware Framework for Social Navigation.

Zero-Shot 3D Visual Grounding from Vision-Language Models From Cognition to Precognition: A Future-Aware Framework for Social Navigation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.654751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.101940Z digest=sha256:96b62419a43fee04deb47ca0e114033ea41451a9d0c74297a030ab20deda6268

Observation dd10ce10-0d22-4d89-9acd-dcdaeb9deac4 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.413691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.175634Z digest=sha256:bbe9aa217549e2dd1c21ce6fd0bc3ef1981ed57aa50b0d97c2c07683ebc7097f

Observation 31d13b7b-2b84-4766-85e6-5592b983dad8 · outbound

This paper cites Robo3d: Towards robust and reliable 3d perception against corruptions,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robo3d: Towards robust and reliable 3d perception against corruptions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.257194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.211674Z digest=sha256:db2075f1ae387cd94ec99f0df1d227663c53b29b6b717641da374c2946403421

Observation c1d22c9f-0892-4398-8716-e2d9ccbc6a59 · outbound

This paper cites Rethinking range view representation for lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Rethinking range view representation for lidar segmentation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.093229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.269354Z digest=sha256:f731e2979db082b6425a8ff9af2005d425fc9f95f20646a3fd4ebb95d56e3614

Observation 2bedc6dc-5cfb-46f0-91a9-202261a576b1 · outbound

This paper cites Xvo: Generalized visual odometry via cross- modal self-training,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Xvo: Generalized visual odometry via cross- modal self-training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.912696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.340641Z digest=sha256:0d97d58269628705f4bf8414cc961e15b14207c743a9665d0bf46835bfa25368

Observation 03bfdf04-2234-4fe5-a716-86bba0b57d4b · outbound

This paper cites COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.416591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.416591Z digest=sha256:56f2e1a59eadc4324e37ec36e27129a6ab69cd814f58ab7f8bc29b2b7f8441b0

Observation fa261fc4-6dda-44f3-a79d-88fba013e8b5 · outbound

This paper cites Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.722848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.467600Z digest=sha256:7ebdcb0ed9ebb04552be0f698a3f652fba4b4864cc654348805c2ce36aa6430c

Observation 6b23063d-3182-489e-ae81-83d3d6d434eb · outbound

This paper cites Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.545958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.521529Z digest=sha256:55a14cc2a536a99a47d7d8004a202806690f7f6370e912f8b976280f39527674

Observation 1040d8e3-f317-4bd5-8225-322ce41fcea5 · outbound

This paper cites Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.385278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.602509Z digest=sha256:33f3946d64acab76ff044dbea94e67d9bba6243e6f22f3c0aec2c05ba6dba188

Observation dd2fa089-3b28-4518-a393-274e9128ad3c · outbound

This paper cites Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.678694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.678694Z digest=sha256:d6bebd41b4cf95effacb4c88c0f804deceeef331230d05f5aadcc61806a3d6fe

Observation 6adcb6dd-1a21-4311-bcf5-680846d8444c · outbound

This paper cites Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.170166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.771248Z digest=sha256:4adfbae93bf82178260c50293785b96ef8f6bddfe219b07a85e5c041b0116bfd

Observation dd1c1904-35de-4441-b3eb-dd532efb8937 · outbound

This paper cites Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.987688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:14.854042Z digest=sha256:9881843734a16dccc55cac141831dc11ec0c9f31a65e237be5f80644edd732c3

Observation 77efeec2-affe-46a1-8e6c-5590ca719030 · outbound

This paper cites Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.916380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.916380Z digest=sha256:2b37a9263a7ed0c6556e75bf8c0e8e2817a251ec4e6b6d8612efd1badec2ab49

Observation 21947037-35b4-4109-83a4-c704eadff963 · outbound

This paper cites Calib3d: Calibrating model preferences for reliable 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Calib3d: Calibrating model preferences for reliable 3d scene understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.705023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.013655Z digest=sha256:52a4fb4f986b7cdc0fd0989df81d1fdce6c609a6c3a05b7d2c45430b49626649

Observation 5c6df824-f55a-43df-b72b-d33033b5e4f1 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Bottom up top down detection transform- ers for language grounding in images and point clouds,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.534444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.130922Z digest=sha256:935215e39cb2539123d17bb81e8b20a365ea2f819a7089f7c4b9e65e1bf79e8a

Observation bc2adfb0-b5bc-4cec-8761-c20b7c790c46 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-vista: Pre-trained transformer for 3d vision and text alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.315718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.204457Z digest=sha256:d9c37b84f6dd5e1b82c36c5a254686a7528b2e51f09df7671fa07c8491ef4c3c

Observation 7c0380bf-1754-4b5b-91dd-f9bc7301825e · outbound

This paper cites Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.141405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.273293Z digest=sha256:973ee86d551c0cda0b570cdcb5b49d06b35eaa95fcf4113e83abd41a94ff3256

Observation d50e43f0-cf9a-41c7-b7b3-64936739d156 · outbound

This paper cites 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.002253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.322438Z digest=sha256:c2d7668afa8989281c4902b17dc074fafc1a05eecbf40792857c0b0afa94eb27

Observation bd712acb-4ec8-4fd0-bc3f-3c2a28ce6d8f · outbound

This paper cites Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.817183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.400260Z digest=sha256:2ec500c39b20c5f119dcbbeaf1443f9acf400c29ac498fa010156c9478502a91

Observation f3b330c4-f015-42f9-86e6-d02a4f2ff0f9 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-branch collaborative learning network for 3d visual grounding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.648448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.451650Z digest=sha256:1d25b92a6f6972e58ff4c78f7277bdaaf9f91a4c459a4317b9ebbbf43d1b6232

Observation 5f07f699-f8bc-441a-8085-dcc8ec68e53d · outbound

This paper cites Semantickitti: A dataset for semantic scene understanding of lidar sequences,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Semantickitti: A dataset for semantic scene understanding of lidar sequences,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.457192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.491354Z digest=sha256:95ad916ad19f636b42146d180b36acecff9f5cc3611fa27e3e524f5712034c5f

Observation 2f3a2f58-fa72-4e45-9ff5-3861b4926bab · outbound

This paper cites Scalability in perception for autonomous driv- ing: Waymo open dataset,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scalability in perception for autonomous driv- ing: Waymo open dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.255587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.564839Z digest=sha256:8f6aedaf709f2e4838694aaed1716bdaa5cc25645b2d9e207028dac69368c698

Observation 7354791c-6c9e-4f60-956a-af68e73e773c · outbound

This paper cites Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.002623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.623601Z digest=sha256:cc231625b74b1f4d97bcbf97e34133a088500ccf514a0e7ff4f1b8342edf2ac2

Observation 11375c3b-8862-4593-bad2-7761dfc1ac0e · outbound

This paper cites Visual programming for zero-shot open- vocabulary 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Visual programming for zero-shot open- vocabulary 3d visual grounding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.773037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.701053Z digest=sha256:facfde89a16e70cb6bf512fb6aa554facdc259c1720121401993334e62530910

Observation a012d746-72fe-44e3-bdeb-0d1aee487572 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.545584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.776796Z digest=sha256:dcbb165a569c274a593aa8f0f53aaad4bfcda7ebbdba5b9328bfc6107788296a

Observation c155db64-fdc3-4454-8b88-6e86c1baba4a · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Training language models to fol- low instructions with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.325098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:15.878975Z digest=sha256:cb29e1aef7de02fc1a2f5b31d451ed23fa671f3cf54f126718ab896bd600774d

Observation 1e1b7fb4-f508-42b6-8fab-a9bf6ea0c806 · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot 3D Visual Grounding from Vision-Language Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:15.966014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:15.966014Z digest=sha256:bd6508502632a7cb4b8cd8936ae60cd13bb40ef83cc7bf5283841b3604fcf32d

Observation 3579b74e-7593-4f22-8ebb-e28f7c17fb6a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Zero-Shot 3D Visual Grounding from Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.033226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.033226Z digest=sha256:12cda3d8c773fc4657c8b981862bf224bc8655c144c716a92809333b29de49d7

Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.099070Z digest=sha256:354e1754f0ffc9c2844aa209db577c163c61aea3e905b309b3bc8e827458730b

Observation 63c38e4a-105a-4267-8616-ce139f488a22 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.062525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.172492Z digest=sha256:b7bc3c31ddfc72e5e6fee2376bfa32dd898662f974187b923ea1cc5303f9b2ad

Observation 37219787-af06-4101-8af5-57e158d3730c · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.243728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.243728Z digest=sha256:0032343873512638c916f8d61669e97433203d842426111c733f55f2416b07ef

Observation dd2be870-d9b6-4bae-b676-1b4ff513377a · outbound

This paper cites Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.868411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.327105Z digest=sha256:674934a6fd551cad666eef5389500d48cf3ee33d80169f529345a916c789a38b

Observation 1c3d62e2-a284-4093-9485-c07bfd12986b · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.642877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.383282Z digest=sha256:a6c5a5a651862bcfbb4a43993ecc97c6c9c82a209ea931c45ad27bc953233c53

Observation a70094c3-e7ce-4b6c-8fc8-cb4d1670b286 · outbound

This paper cites Multi-view transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-view transformer for 3d visual grounding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.299693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.457874Z digest=sha256:c2d9755f2071a2f23b291a38ab262d4e2a0bcf1d9c06823276c82cc1efc88183

Observation 9229b3cc-9dcb-4273-9701-4f45767296f1 · outbound

This paper cites Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.103416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.537325Z digest=sha256:f8d0bcf82db8e4a876a79e27773e4fd091a33809d31b0c0dd4c08567b75717f5

Observation d6ce677b-b1b6-4488-b020-89df092fe040 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sat: 2d semantics assisted training for 3d visual grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.794171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.586571Z digest=sha256:f5d52ad5dc202e81ea973e0077e126ae9eb9f880f3b85ed40155e5f1bd7e73af

Observation 1ca3c590-6dea-4a27-874d-747b1befea47 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.612752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.657203Z digest=sha256:93de8d90e1ae67f21549c03946d4576b5add58258a42e1454be2ba8bb1319e2f

Observation 8b9e7394-52c4-4635-90bc-357d17e6ace1 · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.393051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.723540Z digest=sha256:b6d15674e6491b2a15b1cbac65142f91c7b4f5cdc09ebf8f3bf27fd861378c5c

Observation 8ad91c79-521c-4f81-8bfc-9e2577bc4909 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Unifying 3d vision-language understanding via promptable queries,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.169590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.772423Z digest=sha256:f875cac0c56cdba9e22485fa1db08fd3f68a78c71d874f873e718a1a4f95366b

Observation f858176d-4543-448f-94e2-5417467643c2 · outbound

This paper cites OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies.

Zero-Shot 3D Visual Grounding from Vision-Language Models OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.828825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.828825Z digest=sha256:abc03d10114425860dc97319f7002fd3004140b7d60f6f13a8b277523e9b8dd9

Observation b52141f5-02d0-4ff5-8e29-8998abfa1c46 · outbound

This paper cites Multi-space alignments towards universal lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-space alignments towards universal lidar segmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.960756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.895168Z digest=sha256:01c2bff0a0711202455311611375310fad1693a68f6505dc0cb92f4d25419609

Observation 02ec2c3f-ff86-42ff-8ffb-34b28baecacf · outbound

This paper cites Towards label-free scene understanding by vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Towards label-free scene understanding by vision foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.756569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:16.967579Z digest=sha256:003c0a3c14bd4b282339e3de9c492a4f67f65265638a6d6c2d312d95f1c85e52

Observation 72357686-fe9d-482c-a614-4c8e5f09b4b6 · outbound

This paper cites LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes.

Zero-Shot 3D Visual Grounding from Vision-Language Models LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.070727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.070727Z digest=sha256:6129f882b3288512bbb6905269c7f0f0be7013ae7dfb651ee8fd4b6dec513a59

Observation 197f289d-d06e-4527-8ea0-b7333f8cc82a · outbound

This paper cites GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.150193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.150193Z digest=sha256:67645e5f6bdacfa1ec3f82955b7d8ff73ae9e59b1531c97d123739339a616dee

Observation 030e2204-a25d-4b24-a0b4-3397826df452 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openscene: 3d scene understanding with open vocabularies,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.487986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.213896Z digest=sha256:81be62f2c77c78d3b5f75179d42380c1464b8097b0f94d1ed11f22df4e3514e1

Observation 423967c6-2e0e-4c80-bcf4-565da2f9c4d3 · outbound

This paper cites Lerf: Language embedded radiance fields,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lerf: Language embedded radiance fields,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.225323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.283351Z digest=sha256:27e7f47575c14480bc6711ce9ca5250dc6925fafa706ec6d56e7eb102af30c43

Observation 94103a13-4e1b-43a6-bbc2-ed6e9e18a290 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.019581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.343736Z digest=sha256:7bec8dba032fa5eedc699c09d7297d39c721cca9a73aa1a8e7dfab3cd5a4cf4a

Observation 1f8f6e63-38ef-47e1-98da-a0ea5409aeeb · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.411309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.411309Z digest=sha256:a0a99cf5c797949fbc52723cb88d647a725a32a0ad63da76fa53759a59b4b6bd

Observation 1d2c5493-8c26-42f6-90f4-97d8c35ede01 · outbound

This paper cites Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.776061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.470353Z digest=sha256:c025636ac3103ad75c8252f795a4265c950922f7e058b7a96e9bd6cfb331ef40

Observation 739b960d-0787-420d-8179-33da110aa4cf · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.523576Z digest=sha256:909a753ded8934e64614908fd3f315b9d3d73141bda6d001a4b0f91f3c03b7b5

Observation 9fff50e1-dbfb-428b-bd79-ed7ded1e7bfd · outbound

This paper cites Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.563327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.579259Z digest=sha256:330cc9e294eb76f75647642a5117d5f7550e55cba68235c7b877f58152450886

Observation 6b063554-7f08-45ac-8735-ec16104d3204 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sai3d: Segment any instance in 3d scenes,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.321836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.637874Z digest=sha256:b9c056a0ad35e27aad76e2548f0aab74a6bccb361a9a639e25e12a119bbc99df

Observation b8e8a98d-4d75-4d7e-859e-cbf1efa418be · outbound

This paper cites Lasermix for semi-supervised lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lasermix for semi-supervised lidar semantic segmentation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.131536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.710770Z digest=sha256:708520cf1de49e0961c14ee63871e33876d65ab5cb2276b0c05597a0d1376aad

Observation 88a4132e-37ac-4dc3-aec5-a64dc8004e8f · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Segment any point cloud sequences by distilling vision foundation models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.906258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.807087Z digest=sha256:cacd873301b35b7e26135e59f419a03eaa8279207d1319dd251e75a6023fd74d

Observation c44806ef-9311-4844-9525-3cd2395c8625 · outbound

This paper cites 4d contrastive superflows are dense 3d repre- sentation learners,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 4d contrastive superflows are dense 3d repre- sentation learners,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.683419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:17.935512Z digest=sha256:4ddf89f40bd4f1ad8fdfe40c6f86387a83910c963d98aeebca28bef2f66e4ae4

Observation 1c3f78e0-9855-4610-a55d-3cdbde6b628c · outbound

This paper cites Frnet: Frustum-range networks for scalable lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Frnet: Frustum-range networks for scalable lidar segmentation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.484256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.034869Z digest=sha256:ffa02deed34b020fc87489990c15c62b5fc4434f92ae5a2040990718f38134f8

Observation c4b43a23-36bf-481c-ae1a-d60fab1e0d84 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.091406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.091406Z digest=sha256:13684a4d1e09dff08aa45b58a0a276d9a19ade599b43505543740dc612684a23

Observation b469ac04-b778-4bfe-b0ae-85bc252010b2 · outbound

This paper cites Uni3DL: Unified Model for 3D and Language Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Uni3DL: Unified Model for 3D and Language Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.164538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.156553Z digest=sha256:815f1f0c42bbd073626f798e3ae9ec7304081e4d5e5b42034b1f8d02a26a3f12

Observation 6171bf57-8155-4ef5-bc3d-97e28edb525c · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Conceptfusion: Open-set multi- modal 3d mapping,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.254570Z digest=sha256:576c1c646e095241d600b3df97fd4ea35f27a201e74a0d423db1d6c0c14399e8

Observation b1e0260d-814e-4607-846a-a2b4f28ac11a · outbound

This paper cites GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping.

Zero-Shot 3D Visual Grounding from Vision-Language Models GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.478933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.478933Z digest=sha256:49ccffd2a02070ff9f971e80e0251c9d5416b8944577634473165abf89668518

Observation 11fb35db-9415-452a-8a43-1742b926fdb6 · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Interactive planning using large language models for partially observable robotic tasks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.161734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.706597Z digest=sha256:2b48b415e5302c2b28a164d67dc1a295c1db81d87e6c4922a1d92904baeac749

Observation fce09665-8e57-4ecf-ac78-7d026fcc53c8 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-llm: Injecting the 3d world into large language models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.964761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.817427Z digest=sha256:44af36cb007cf0ff40d62f8c6ad554b265093574259b6c69306cf1b579397da7

Observation 5fcd2776-fd95-49b2-b879-6e63ed38e8fd · outbound

This paper cites Is your lidar placement optimized for 3d scene understanding?,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Is your lidar placement optimized for 3d scene understanding?,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.736693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.907550Z digest=sha256:37bae24201afc3eb95c52dfe2ade43168cb57893d5e26f12c100327167d122e3

Observation 7343377b-7b76-4d17-b048-d0c15d351101 · outbound

This paper cites G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,.

Zero-Shot 3D Visual Grounding from Vision-Language Models G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.557138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:18.980280Z digest=sha256:3e44106bf38455912ef4656d934c15df69377b50f59ea6b59a4677e917786fe6

Observation 3f863d07-94ef-4de9-ad28-d001bf53d84e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.328649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:19.127295Z digest=sha256:8d9bb00d9a522d413cb2551cb37111ae45af8eb82cb9630df87fe9b3316fab9b

Observation edcfabd1-07f9-4cd5-b156-2823a5aa0bc2 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:19.231195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:19.231195Z digest=sha256:af7ed8ea4694a10245cd1cf6ee8358e8bad378dfeb56a2a55481e61bc2166e96

Observation 6bb7d6ca-a924-4237-8f5e-d5f833206cbb · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Text-guided graph neural networks for referring 3d instance segmentation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.080070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:19.353586Z digest=sha256:d5603e429de3a1df25a6dfa0ee636ff5ff123c227b2d8449ac20e852acd8860a

Observation 4761a55c-7a2e-4aa9-9ef1-8cf0368ce164 · outbound

This paper cites Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.837968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:19.499102Z digest=sha256:75beaf4b34f2278a7c82ecb23406226ec487abc48f330754f486f38da536d8e5

Observation 0de609bf-af3c-4a3d-b24c-fc6057d60b72 · outbound

This paper cites Language conditioned spatial relation rea- soning for 3d object grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Language conditioned spatial relation rea- soning for 3d object grounding,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.636127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:19.626429Z digest=sha256:1de1248a0b259a997f3466d7aaacd241ab41f6f4f82afbbd61f1157336cc4bdd

Observation b9c59bea-69c3-4075-8a2c-08946499f3a6 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.510030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:12:19.747542Z digest=sha256:101b400061b11164a07be1b9dedc8332cd8c1802f55971be1a3df8e8032c137d

Pith citing papers

Observation 7e135835-ad52-43f6-9e78-62ef14a3da16 · inbound

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding cites this paper.

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding Zero-Shot 3D Visual Grounding from Vision-Language Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.444056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T06:41:38.649424Z digest=sha256:9c629bdceca168a26d0a2c7fe63febd652f40d487888797cfc3ccd379a1ef875