Pith. sign in

Paper Citation Record · LEDGER

A Review of 3D Object Detection with Vision-Language Models

As of 22 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 1 inbound Pith citation observation for arXiv:2504.18738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18738 v1

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:13:25.257063Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:34:07.995899Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:16:17.039197Z

Reference resolution

7 of 7 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98957a5e-5342-4994-8ea7-0436f0597901 · outbound

This paper cites Leveraging VLM-Based Pipelines to Annotate 3D Objects.

A Review of 3D Object Detection with Vision-Language Models Leveraging VLM-Based Pipelines to Annotate 3D Objects

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.741575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:13:25.139862Z digest=sha256:d02c8388b752f95164e2a584212fce78447ec400ae0561c77b54851baab9132a

Observation 33ac26f8-7c1c-4483-ba56-11a9eb3bf281 · outbound

This paper cites SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:25.627225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:13:25.173316Z digest=sha256:33688e9cee1ab9abc13d973628d245e6c65d945590f77cc1cb364f6df7a3541f

Observation 7018c0b3-d611-47e9-8421-b7f103a98985 · outbound

This paper cites Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks.

A Review of 3D Object Detection with Vision-Language Models Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.257063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.257063Z digest=sha256:5398160af0aad358c4bc51f3413206c8174489e1689f278c6dc3e2c411335eda

Observation 8febbffd-7988-46e7-b6ce-d4a41c792e9f · outbound

This paper cites OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference.

A Review of 3D Object Detection with Vision-Language Models OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:26.010057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T10:13:25.113330Z digest=sha256:714f72c69b0562731668cd54b8caf9ae9331eb2088f784975d029541c5493cea

Observation f2f5b1d2-5aff-40db-a7c9-fc4a66b70aa8 · outbound

This paper cites Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion.

A Review of 3D Object Detection with Vision-Language Models Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.147697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.147697Z digest=sha256:a305f4f097934b9cf745ad1b32a0023ead2d1da4e4e99303c4333aa1af143eec

Observation 0abec588-bf00-4e50-ae31-19e61f2e006e · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

A Review of 3D Object Detection with Vision-Language Models M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.100455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.100455Z digest=sha256:c412c09837b96afbcb10800efb1394251e9176bd6923c9f55fb6eb81fcfba096

Observation ba40995c-8a85-4c7d-bded-77e27378dee3 · outbound

This paper cites OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection.

A Review of 3D Object Detection with Vision-Language Models OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T10:13:25.128583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:13:25.128583Z digest=sha256:d264dd0597e3ab26c93db3235c90870561a418f0375e8f504f5076ded04759a4

Pith citing papers

Observation f52befdf-3477-40ff-8c9b-f9d2922e3eb1 · inbound

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks cites this paper.

Plant Disease Detection through Multimodal Large Language Models and Convolutional Neural Networks A Review of 3D Object Detection with Vision-Language Models

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:34:08.064561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:34:07.995899Z digest=sha256:f5804a05ef313fee05bd97913d8ef66d3008ffa2d4e68b857a3acdf25de4ac9e