Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot 3D Visual Grounding from Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 1 inbound Pith citation observation for arXiv:2505.22429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22429 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:12:19.747542Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:41:38.649424Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T06:45:29.442037Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact5
  • verified fuzzy57
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 960747c0-afc7-4659-b22b-4e9984151d05 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scanrefer: 3d object localization in rgb-d scans using natural language,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:33.112862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.394174Z digest=sha256:26b6152cd003222d4bc32c2c7c05713d1d1d492db4ac09c7c8793349af71dbab

Observation 8b9a2380-ea31-4728-a03e-027b0f870747 · outbound

This paper cites RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models RayDF: Neural Ray-surface Distance Fields with Multi-view Consistency

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.282429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.502388Z digest=sha256:48c8987976182feafaad19834a2ad2e34e15bd1ec58403a0155261d830354a4d

Observation dbfea0d5-1991-4263-ac19-3d58190694cb · outbound

This paper cites Deep view synthesis via self-consistent gen- erative network,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Deep view synthesis via self-consistent gen- erative network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.931597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.596606Z digest=sha256:abb3029dd4dcea4e3683b96236b224c310e145c3c5ce2dea67fba7e970651354

Observation e3b704f0-aab2-4e93-86ab-32d1e3515bac · outbound

This paper cites PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation.

Zero-Shot 3D Visual Grounding from Vision-Language Models PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:13.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:13.660944Z digest=sha256:55fdcb47d9abe1cd0953be6956f46579d3764182b49911fe73f6eec67a8cb3ef

Observation 57a6e328-f12f-4ab9-af53-7caa8174205f · outbound

This paper cites SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting.

Zero-Shot 3D Visual Grounding from Vision-Language Models SIR: Multi-view Inverse Rendering with Decomposable Shadow Under Indoor Intense Lighting

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:21.087050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.744633Z digest=sha256:7eab82b8630021a8d9964a2974166470897de94c32c18e615e70005b7bee4c4a

Observation 4f6dd310-9524-424d-8221-e830ebea8546 · outbound

This paper cites An Examination of the Compositionality of Large Generative Vision-Language Models.

Zero-Shot 3D Visual Grounding from Vision-Language Models An Examination of the Compositionality of Large Generative Vision-Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.851836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.802829Z digest=sha256:c29edc49d11992bcc962234411d69318331b814d95eccd058441bd0fbfa4a435

Observation aa17f378-4420-4b33-b61d-de83946f313d · outbound

This paper cites Think global, act local: Dual-scale graph transformer for vision-and-language navigation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Think global, act local: Dual-scale graph transformer for vision-and-language navigation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.731576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:13.940140Z digest=sha256:e13dc841f214451622351e699613746ae335bf3cea00c67ddf9350cc876778c3

Observation 330b875b-8571-4c73-ae2c-62fa7c4aaeb1 · outbound

This paper cites Assister: Assistive navigation via condi- tional instruction generation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Assister: Assistive navigation via condi- tional instruction generation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.568017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.029025Z digest=sha256:ce471aa5519c0dfdffcbd569d881fb1277aec79367a36f1e86205ac6dfe5f6ae

Observation 9f53beb9-7803-4995-a8ac-6306a5f28797 · outbound

This paper cites From Cognition to Precognition: A Future-Aware Framework for Social Navigation.

Zero-Shot 3D Visual Grounding from Vision-Language Models From Cognition to Precognition: A Future-Aware Framework for Social Navigation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.654751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.101940Z digest=sha256:491e7d2351fc0b539d36815040dd36dfd5f266da98448db8915f0255260470a0

Observation dd10ce10-0d22-4d89-9acd-dcdaeb9deac4 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Clip2scene: Towards label-efficient 3d scene understanding by clip,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.413691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.175634Z digest=sha256:6aa2aa0082f953774632f98334752abf9bb33d7ad8151ea2ab5789d52712ec05

Observation 31d13b7b-2b84-4766-85e6-5592b983dad8 · outbound

This paper cites Robo3d: Towards robust and reliable 3d perception against corruptions,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robo3d: Towards robust and reliable 3d perception against corruptions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.257194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.211674Z digest=sha256:7208afd9645fc24deb984ee89eb490b1a6fb98062f5bd9dbbfa456faac6d78fc

Observation c1d22c9f-0892-4398-8716-e2d9ccbc6a59 · outbound

This paper cites Rethinking range view representation for lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Rethinking range view representation for lidar segmentation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:32.093229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.269354Z digest=sha256:ea40ac6879e1847aacad71bc1595d199401d91bc3daf1b6981cccdb3cd0bed21

Observation 2bedc6dc-5cfb-46f0-91a9-202261a576b1 · outbound

This paper cites Xvo: Generalized visual odometry via cross- modal self-training,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Xvo: Generalized visual odometry via cross- modal self-training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.912696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.340641Z digest=sha256:3ec980fd45e36767db6f92dedf70c7e97da4f62cdf288785d8f43cb4df0bf313

Observation 03bfdf04-2234-4fe5-a716-86bba0b57d4b · outbound

This paper cites COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models COARSE3D: Class-Prototypes for Contrastive Learning in Weakly-Supervised 3D Point Cloud Segmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.416591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.416591Z digest=sha256:56f2e1a59eadc4324e37ec36e27129a6ab69cd814f58ab7f8bc29b2b7f8441b0

Observation fa261fc4-6dda-44f3-a79d-88fba013e8b5 · outbound

This paper cites Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Perception-aware multi-sensor fusion for 3d lidar semantic segmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.722848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.467600Z digest=sha256:1a0f8f1002d71bb2f46b3343eedadf6289d145a7e6b81aae40ee21f7683ef231

Observation 6b23063d-3182-489e-ae81-83d3d6d434eb · outbound

This paper cites Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Epmf: Efficient perception-aware multi- sensor fusion for 3d semantic segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.545958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.521529Z digest=sha256:c7395bbaf5d6a3a4addcdf5d839d5a7fe035ac210c14b296bf4caec48f217853

Observation 1040d8e3-f317-4bd5-8225-322ce41fcea5 · outbound

This paper cites Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Tfnet: Exploiting temporal cues for fast and accurate lidar semantic segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.385278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.602509Z digest=sha256:8a105a73971be7e9f2da4155810d876482e2dca3005e6f1e86f7e4957f1dc337

Observation dd2fa089-3b28-4518-a393-274e9128ad3c · outbound

This paper cites Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation.

Zero-Shot 3D Visual Grounding from Vision-Language Models Robust 3D Semantic Occupancy Prediction with Calibration-free Spatial Transformation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.678694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.678694Z digest=sha256:d6bebd41b4cf95effacb4c88c0f804deceeef331230d05f5aadcc61806a3d6fe

Observation 6adcb6dd-1a21-4311-bcf5-680846d8444c · outbound

This paper cites Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dhp-mapping: A dense panoptic mapping sys- tem with hierarchical world representation and label opti- mization techniques,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:31.170166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.771248Z digest=sha256:80b93e80ae890af5208ce2e16788b8ec66439c82603b7df53a30384fc3972170

Observation dd1c1904-35de-4441-b3eb-dd532efb8937 · outbound

This paper cites Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-modal data-efficient 3d scene un- derstanding for autonomous drivin,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.987688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:14.854042Z digest=sha256:fb02a79d37582b8f46fc62e81b317dc10918847e3cf6cbd39bc38f6d6040a89a

Observation 77efeec2-affe-46a1-8e6c-5590ca719030 · outbound

This paper cites Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Dynamiccity: Large-scale 4d occu- pancy generation from dynamic scenes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:14.916380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:14.916380Z digest=sha256:2b37a9263a7ed0c6556e75bf8c0e8e2817a251ec4e6b6d8612efd1badec2ab49

Observation 21947037-35b4-4109-83a4-c704eadff963 · outbound

This paper cites Calib3d: Calibrating model preferences for reliable 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Calib3d: Calibrating model preferences for reliable 3d scene understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.705023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.013655Z digest=sha256:b681f9ca5a7102e9162448e3383733e8deaefa0ae14732d755a230e7920a7d3b

Observation 5c6df824-f55a-43df-b72b-d33033b5e4f1 · outbound

This paper cites Bottom up top down detection transform- ers for language grounding in images and point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Bottom up top down detection transform- ers for language grounding in images and point clouds,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.534444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.130922Z digest=sha256:97c4fd3756eb6172abee9db3fdde6f9dd07a30fa9b115c6650793bd839eff356

Observation bc2adfb0-b5bc-4cec-8761-c20b7c790c46 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-vista: Pre-trained transformer for 3d vision and text alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.315718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.204457Z digest=sha256:8b995e747c9771d5c4b245abd7f566ef041e84cca7dbfb830f07de49335fddc7

Observation 7c0380bf-1754-4b5b-91dd-f9bc7301825e · outbound

This paper cites Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Eda: Explicit text-decoupling and dense align- ment for 3d visual grounding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.141405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.273293Z digest=sha256:622a47ca93b9024f4b4142437bbe46e9c92dfe51915668c0942e14d279d6e3cc

Observation d50e43f0-cf9a-41c7-b7b3-64936739d156 · outbound

This paper cites 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3dvg-transformer: Relation modeling for vi- sual grounding on point clouds,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:30.002253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.322438Z digest=sha256:31e058bb9cb95f01ed3ef8ba44d29d165a550df58418136b9ad11c02d189b775

Observation bd712acb-4ec8-4fd0-bc3f-3c2a28ce6d8f · outbound

This paper cites Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Instancerefer: Cooperative holistic under- standing for visual grounding on point clouds through in- stance multi-level contextual referring,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.817183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.400260Z digest=sha256:132aa6f596a19a8c23e33b140178b97661793815a6921e41088f9f8962493558

Observation f3b330c4-f015-42f9-86e6-d02a4f2ff0f9 · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-branch collaborative learning network for 3d visual grounding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.648448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.451650Z digest=sha256:fa309b4a1b0a3ba228b18e29256bc61ff557bd29b7cef7216ec9eaf181f28c1e

Observation 5f07f699-f8bc-441a-8085-dcc8ec68e53d · outbound

This paper cites Semantickitti: A dataset for semantic scene understanding of lidar sequences,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Semantickitti: A dataset for semantic scene understanding of lidar sequences,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.457192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.491354Z digest=sha256:ebcc6893cc1d3ce99a3dc2edcfc19a0607e1842c55361ccb26354a646c6cd54a

Observation 2f3a2f58-fa72-4e45-9ff5-3861b4926bab · outbound

This paper cites Scalability in perception for autonomous driv- ing: Waymo open dataset,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scalability in perception for autonomous driv- ing: Waymo open dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.255587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.564839Z digest=sha256:c1afe1fa4c35534df668fdfde12c99044dbb0d2e28736ef4109c2bd106a69b40

Observation 7354791c-6c9e-4f60-956a-af68e73e773c · outbound

This paper cites Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Panoptic nuscenes: A large-scale bench- mark for lidar panoptic segmentation and tracking,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:29.002623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.623601Z digest=sha256:36c09ed990c1aa0f748f3c6c1334aa7297e8675468c5419db4c69858cc8cd454

Observation 11375c3b-8862-4593-bad2-7761dfc1ac0e · outbound

This paper cites Visual programming for zero-shot open- vocabulary 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Visual programming for zero-shot open- vocabulary 3d visual grounding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.773037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.701053Z digest=sha256:7796bf7d37c84df7e0d3fbfce19049bc3f15dfbfdc2d38182cfd947bdfb44c0a

Observation a012d746-72fe-44e3-bdeb-0d1aee487572 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.545584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.776796Z digest=sha256:e0fd92cd0bd0aec879fe78f16ee0f533a6886d42fce499fc623321816421bcdd

Observation c155db64-fdc3-4454-8b88-6e86c1baba4a · outbound

This paper cites Training language models to fol- low instructions with human feedback,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Training language models to fol- low instructions with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.325098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:15.878975Z digest=sha256:dc7a35c22ab7e984204e5c91742ef759871c70a1069deb8e2c5714bebb947f51

Observation 1e1b7fb4-f508-42b6-8fab-a9bf6ea0c806 · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot 3D Visual Grounding from Vision-Language Models GPT-4 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:15.966014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:15.966014Z digest=sha256:bd6508502632a7cb4b8cd8936ae60cd13bb40ef83cc7bf5283841b3604fcf32d

Observation 3579b74e-7593-4f22-8ebb-e28f7c17fb6a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Zero-Shot 3D Visual Grounding from Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.033226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.033226Z digest=sha256:12cda3d8c773fc4657c8b981862bf224bc8655c144c716a92809333b29de49d7

Observation 5426ae40-e7e6-4ba6-bb60-6b115f0d2f4a · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.099070Z digest=sha256:354e1754f0ffc9c2844aa209db577c163c61aea3e905b309b3bc8e827458730b

Observation 63c38e4a-105a-4267-8616-ce139f488a22 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sceneverse: Scaling 3d vision-language learn- ing for grounded scene understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:28.062525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.172492Z digest=sha256:c3773ef90f0c2458126d855df5ad03f19ff62a3e2b714f7b271a8edc4696ecf3

Observation 37219787-af06-4101-8af5-57e158d3730c · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.243728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.243728Z digest=sha256:0032343873512638c916f8d61669e97433203d842426111c733f55f2416b07ef

Observation dd2be870-d9b6-4bae-b676-1b4ff513377a · outbound

This paper cites Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.868411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.327105Z digest=sha256:16b310ab4bf2435557e823791bb7c2480b7a35d2d1b5d5a56d6193e8f4789fa8

Observation 1c3d62e2-a284-4093-9485-c07bfd12986b · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Viewrefer: Grasp the multi-view knowledge for 3d visual grounding,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.642877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.383282Z digest=sha256:0de2fde57e6447edce4a2232ed3f267f5e3dc3597ce24f475b3a1042c1d79ddb

Observation a70094c3-e7ce-4b6c-8fc8-cb4d1670b286 · outbound

This paper cites Multi-view transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-view transformer for 3d visual grounding,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.299693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.457874Z digest=sha256:199dddbbd1f15cbea8a58911019cfda42b889ec01303e6bb143863ac0e59a431

Observation 9229b3cc-9dcb-4273-9701-4f45767296f1 · outbound

This paper cites Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Look around and refer: 2d synthetic se- mantics knowledge distillation for 3d visual grounding,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:27.103416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.537325Z digest=sha256:dfb3fefb1724f73afa69ddb27de105e5ad32dc4c348afe8d4abeeca0dd055528

Observation d6ce677b-b1b6-4488-b020-89df092fe040 · outbound

This paper cites Sat: 2d semantics assisted training for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sat: 2d semantics assisted training for 3d visual grounding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.794171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.586571Z digest=sha256:7d87385bca922d747c519d3274a6f8cf4f33b70d555eed4fd6539b8ac0e377c5

Observation 1ca3c590-6dea-4a27-874d-747b1befea47 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.612752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.657203Z digest=sha256:fafd001731259cc0f12fe08b1b3d09601888227c8ddcbdc790867db6182c879f

Observation 8b9e7394-52c4-4635-90bc-357d17e6ace1 · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.393051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.723540Z digest=sha256:8d8e00157a51ebdc3028a24dabb7be3ac1afe60cd1fc853b7aef7add31949a14

Observation 8ad91c79-521c-4f81-8bfc-9e2577bc4909 · outbound

This paper cites Unifying 3d vision-language understanding via promptable queries,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Unifying 3d vision-language understanding via promptable queries,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:26.169590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.772423Z digest=sha256:3ab51b1fd904321a48a00e6b31ea09839390ef24fd871f930d46afd489bddffb

Observation f858176d-4543-448f-94e2-5417467643c2 · outbound

This paper cites OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies.

Zero-Shot 3D Visual Grounding from Vision-Language Models OVGaussian: Generalizable 3D Gaussian Segmentation with Open Vocabularies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:16.828825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:16.828825Z digest=sha256:abc03d10114425860dc97319f7002fd3004140b7d60f6f13a8b277523e9b8dd9

Observation b52141f5-02d0-4ff5-8e29-8998abfa1c46 · outbound

This paper cites Multi-space alignments towards universal lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Multi-space alignments towards universal lidar segmentation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.960756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.895168Z digest=sha256:40328a63bc9f850147aa50a20217016377496cfa6fe40043c77efd3920147aac

Observation 02ec2c3f-ff86-42ff-8ffb-34b28baecacf · outbound

This paper cites Towards label-free scene understanding by vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Towards label-free scene understanding by vision foundation models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.756569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:16.967579Z digest=sha256:66ae85a9c88cc128ba9ff0eb4dedae632c6b06db1475ed8d4ec010a5bd0b5d50

Observation 72357686-fe9d-482c-a614-4c8e5f09b4b6 · outbound

This paper cites LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes.

Zero-Shot 3D Visual Grounding from Vision-Language Models LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.070727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.070727Z digest=sha256:6129f882b3288512bbb6905269c7f0f0be7013ae7dfb651ee8fd4b6dec513a59

Observation 197f289d-d06e-4527-8ea0-b7333f8cc82a · outbound

This paper cites GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency.

Zero-Shot 3D Visual Grounding from Vision-Language Models GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.150193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.150193Z digest=sha256:67645e5f6bdacfa1ec3f82955b7d8ff73ae9e59b1531c97d123739339a616dee

Observation 030e2204-a25d-4b24-a0b4-3397826df452 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openscene: 3d scene understanding with open vocabularies,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.487986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.213896Z digest=sha256:0172e41d340dde623b49dd1ef48c3e004fddb09081befd43b8864a8f9e86bcf0

Observation 423967c6-2e0e-4c80-bcf4-565da2f9c4d3 · outbound

This paper cites Lerf: Language embedded radiance fields,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lerf: Language embedded radiance fields,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.225323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.283351Z digest=sha256:a5f0067b9947356b430d32ef09df954fd8eedfe9fff0d90cb2b575318d28b93f

Observation 94103a13-4e1b-43a6-bbc2-ed6e9e18a290 · outbound

This paper cites Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:25.019581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.343736Z digest=sha256:5a008ac46371cfa4338f3202d2bea9b8f2211320110c9b16a2e2c7f0b6463220

Observation 1f8f6e63-38ef-47e1-98da-a0ea5409aeeb · outbound

This paper cites Agent3D-Zero: An Agent for Zero-shot 3D Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Agent3D-Zero: An Agent for Zero-shot 3D Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.411309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.411309Z digest=sha256:a0a99cf5c797949fbc52723cb88d647a725a32a0ad63da76fa53759a59b4b6bd

Observation 1d2c5493-8c26-42f6-90f4-97d8c35ede01 · outbound

This paper cites Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Regionplc: Regional point-language con- trastive learning for open-world 3d scene understanding,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.776061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.470353Z digest=sha256:cb6e8f2970031fd921841e540832dacb6acdc6d1479c5b06333a61a2f0e5c6f6

Observation 739b960d-0787-420d-8179-33da110aa4cf · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

Zero-Shot 3D Visual Grounding from Vision-Language Models OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:17.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:17.523576Z digest=sha256:909a753ded8934e64614908fd3f315b9d3d73141bda6d001a4b0f91f3c03b7b5

Observation 9fff50e1-dbfb-428b-bd79-ed7ded1e7bfd · outbound

This paper cites Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Openins3d: Snap and lookup for 3d open- vocabulary instance segmentation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.563327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.579259Z digest=sha256:6e70b7b76df6c2d783fcc91543acc3ace2df53cbb0cdf09aaf245ef70f7a9fe6

Observation 6b063554-7f08-45ac-8735-ec16104d3204 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Sai3d: Segment any instance in 3d scenes,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.321836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.637874Z digest=sha256:d59c0cf9e7a2cef759d7801f907bb51d0733b009a2c9bd9ecf8c58e6680e3cba

Observation b8e8a98d-4d75-4d7e-859e-cbf1efa418be · outbound

This paper cites Lasermix for semi-supervised lidar semantic segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Lasermix for semi-supervised lidar semantic segmentation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:24.131536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.710770Z digest=sha256:caf4edd84861ffe9088838f0b03a023e00cae17eb24eda746e21adabfce79dff

Observation 88a4132e-37ac-4dc3-aec5-a64dc8004e8f · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Segment any point cloud sequences by distilling vision foundation models,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.906258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.807087Z digest=sha256:ce85e35814bf8373c2674f33c06aaedfce4fcd97ad6686ccc47948c2f0001418

Observation c44806ef-9311-4844-9525-3cd2395c8625 · outbound

This paper cites 4d contrastive superflows are dense 3d repre- sentation learners,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 4d contrastive superflows are dense 3d repre- sentation learners,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.683419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:17.935512Z digest=sha256:0e9297d3fc0c1486f6fa532b6d2cf7e92eac4a6fe9de48cf519716abc22afd41

Observation 1c3f78e0-9855-4610-a55d-3cdbde6b628c · outbound

This paper cites Frnet: Frustum-range networks for scalable lidar segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Frnet: Frustum-range networks for scalable lidar segmentation,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.484256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.034869Z digest=sha256:af30bde0c5ab7c3c0745953e47dff0e88f0c22f3507719b73219e414591e5220

Observation c4b43a23-36bf-481c-ae1a-d60fab1e0d84 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Zero-Shot 3D Visual Grounding from Vision-Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.091406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.091406Z digest=sha256:13684a4d1e09dff08aa45b58a0a276d9a19ade599b43505543740dc612684a23

Observation b469ac04-b778-4bfe-b0ae-85bc252010b2 · outbound

This paper cites Uni3DL: Unified Model for 3D and Language Understanding.

Zero-Shot 3D Visual Grounding from Vision-Language Models Uni3DL: Unified Model for 3D and Language Understanding

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:12:20.164538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.156553Z digest=sha256:7f21fb4b5022b2690aacd2317e543d3e5041cf476b6d5d03bb2ed559f3235c64

Observation 6171bf57-8155-4ef5-bc3d-97e28edb525c · outbound

This paper cites Conceptfusion: Open-set multi- modal 3d mapping,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Conceptfusion: Open-set multi- modal 3d mapping,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.283686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.254570Z digest=sha256:f4c0ffa353be69ab391b8f388482b41f3030ed2aea34623f10a110058a73a07d

Observation b1e0260d-814e-4607-846a-a2b4f28ac11a · outbound

This paper cites GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping.

Zero-Shot 3D Visual Grounding from Vision-Language Models GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:18.478933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:18.478933Z digest=sha256:49ccffd2a02070ff9f971e80e0251c9d5416b8944577634473165abf89668518

Observation 11fb35db-9415-452a-8a43-1742b926fdb6 · outbound

This paper cites Interactive planning using large language models for partially observable robotic tasks,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Interactive planning using large language models for partially observable robotic tasks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:23.161734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.706597Z digest=sha256:9aa5fdc9a57fc54903a9821240fc04b9aa71bbd9c106559dcbe3d331e7e59dea

Observation fce09665-8e57-4ecf-ac78-7d026fcc53c8 · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models,.

Zero-Shot 3D Visual Grounding from Vision-Language Models 3d-llm: Injecting the 3d world into large language models,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.964761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.817427Z digest=sha256:6f97d4805afea221a3ac9f840300f630fa24b8ad93766b71c8ec854e32ca95c3

Observation 5fcd2776-fd95-49b2-b879-6e63ed38e8fd · outbound

This paper cites Is your lidar placement optimized for 3d scene understanding?,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Is your lidar placement optimized for 3d scene understanding?,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.736693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.907550Z digest=sha256:7ebe0897c390c042e0d365c4c174117f93db4d2f9e8b3f2f91e1c3f44a7c176d

Observation 7343377b-7b76-4d17-b048-d0c15d351101 · outbound

This paper cites G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,.

Zero-Shot 3D Visual Grounding from Vision-Language Models G3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual ground- ing,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.557138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:18.980280Z digest=sha256:a5da25738b91e16321d632ee3f3502c57cd07761222a528da97eeacb506537f6

Observation 3f863d07-94ef-4de9-ad28-d001bf53d84e · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Learning transferable visual models from natural language supervision,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.328649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:19.127295Z digest=sha256:178eab6d40ed57bf7d876e52cbbfecab275fc85d1a2df8ac4d88b6d255da2708

Observation edcfabd1-07f9-4cd5-b156-2823a5aa0bc2 · outbound

This paper cites VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding.

Zero-Shot 3D Visual Grounding from Vision-Language Models VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:19.231195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:19.231195Z digest=sha256:af7ed8ea4694a10245cd1cf6ee8358e8bad378dfeb56a2a55481e61bc2166e96

Observation 6bb7d6ca-a924-4237-8f5e-d5f833206cbb · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Text-guided graph neural networks for referring 3d instance segmentation,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:22.080070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:19.353586Z digest=sha256:ed4d3fa90c87d1924cc6d69e306baceaff72a7aa30e2e7130f26cb128e9500b2

Observation 4761a55c-7a2e-4aa9-9ef1-8cf0368ce164 · outbound

This paper cites Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mikasa: Multi-key-anchor & scene- aware transformer for 3d visual grounding,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.837968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:19.499102Z digest=sha256:526c04970d95d8ffe143cd012db2c65357e72bdd3ef83dd23a61d90be4271946

Observation 0de609bf-af3c-4a3d-b24c-fc6057d60b72 · outbound

This paper cites Language conditioned spatial relation rea- soning for 3d object grounding,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Language conditioned spatial relation rea- soning for 3d object grounding,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.636127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:19.626429Z digest=sha256:52f048657294c6bbfda7bae77899eaa558664c43a105b1dbedee3630fa5e8995

Observation b9c59bea-69c3-4075-8a2c-08946499f3a6 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

Zero-Shot 3D Visual Grounding from Vision-Language Models Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:12:21.510030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:12:19.747542Z digest=sha256:3b58e59089911a17b1e6283dfb872a9348ba88e56151609a07f884fa21134d60

Pith citing papers

Observation 7e135835-ad52-43f6-9e78-62ef14a3da16 · inbound

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding cites this paper.

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding Zero-Shot 3D Visual Grounding from Vision-Language Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.444056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T06:41:38.649424Z digest=sha256:04dd59f5c24cb9b412728a289941300997c053e67bb18df1786b6478e90c5d1a