Pith. sign in

Paper Citation Record · LEDGER

A Review of 3D Object Detection with Vision-Language Models

As of 18 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 1 inbound Pith citation observation for arXiv:2504.18738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18738 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:13:25.257063Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:34:07.995899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:16:17.039197Z

Reference resolution

7 of 7 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98957a5e-5342-4994-8ea7-0436f0597901 · outbound

This paper cites Leveraging VLM-Based Pipelines to Annotate 3D Objects.

A Review of 3D Object Detection with Vision-Language Models Leveraging VLM-Based Pipelines to Annotate 3D Objects

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.741575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.139862Z digest=sha256:ac9b881747e93dfdf77ab97368b9cad108426a9b147eba1058b96fddae3ae308

Observation 33ac26f8-7c1c-4483-ba56-11a9eb3bf281 · outbound

This paper cites SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.627225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.173316Z digest=sha256:1242e77b0d15b3b83c62970ae7c65a917a9a162c8fc7ec053b790be3b561078a

Observation 7018c0b3-d611-47e9-8421-b7f103a98985 · outbound

This paper cites Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks.

A Review of 3D Object Detection with Vision-Language Models Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.257063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.257063Z digest=sha256:8133e5de2bdcd79dd4c3c087a553613dfc34d5441292ac7751e075f7783bb4cd

Observation 8febbffd-7988-46e7-b6ce-d4a41c792e9f · outbound

This paper cites OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference.

A Review of 3D Object Detection with Vision-Language Models OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:26.010057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:13:25.113330Z digest=sha256:85148d404ef963afcbd14239128aceb5318c537bc68d300c9881b05c012b592b

Observation f2f5b1d2-5aff-40db-a7c9-fc4a66b70aa8 · outbound

This paper cites Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion.

A Review of 3D Object Detection with Vision-Language Models Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.147697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.147697Z digest=sha256:86825fdc10af40605c97f53b47ed84ea1374fd3063f08b204414d0536c880f2e

Observation 0abec588-bf00-4e50-ae31-19e61f2e006e · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

A Review of 3D Object Detection with Vision-Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.100455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.100455Z digest=sha256:506bf1bf1c4e8ddf20e15788f9ce3e837f3b2c1ae89cd64201f56236affbddbc

Observation ba40995c-8a85-4c7d-bded-77e27378dee3 · outbound

This paper cites OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.128583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.128583Z digest=sha256:6c2b4061ccd5c87f6de0533eb5773321ddc9cb029e4fb91c7c015102dfc9f87b

Pith citing papers

Observation f52befdf-3477-40ff-8c9b-f9d2922e3eb1 · inbound

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks cites this paper.

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks A Review of 3D Object Detection with Vision-Language Models

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:34:08.064561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:34:07.995899Z digest=sha256:6defe0ad0ba7982b0ec99049589bb67c13a761dc6a12aaeed7bb85a7dd6d126e