Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2508.16812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16812 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:06.935892Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact11
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 840f8725-d2b5-4625-9829-401e7c8370a8 · outbound

This paper cites TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.347860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.437631Z digest=sha256:f2c6b1076bdad51fe524525c444021640c7ae63dd3ad60d062f74c57fb53558c

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:b939b5c0515c2a4597f681cf658db761f1fcf0511181824fe8b631ef5a7275d0

Observation 0323529f-d181-4656-9e98-f1df76763a63 · outbound

This paper cites Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.331370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.550840Z digest=sha256:1dbe99409143b9817cb009f7c096d72e2cd2361fb8283dfc4d512def21f87a54

Observation 92bfde35-7860-4701-81bd-53daddc5e63f · outbound

This paper cites CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.315917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.619614Z digest=sha256:b4b4d1e40e019e4fd7e3393a67a3617ad9ce1b6dad1eea42a4760afc12e175bc

Observation 5e6d3ffb-ca94-414c-9c2d-5ae815fbcfda · outbound

This paper cites Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.748655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.666088Z digest=sha256:44a10e3dcb074423a80d4edd64a208125a0bed29d2cce75a359634a033ef5291

Observation b85c14d1-2749-4e90-94ee-c905dc863425 · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.758342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.758342Z digest=sha256:0db3a96093a179506098d087a3176a2c503f4077b78edca57a0ea00fd6f4e90e

Observation 103b1be2-881b-4092-a9de-cdf760126ccd · outbound

This paper cites Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.568522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.824856Z digest=sha256:aa6a15d31eed6f8b59e556661c77faa95bdea46e932a8b340f821f262c3a899b

Observation fcc17589-f298-479f-9021-265f579e2956 · outbound

This paper cites Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.337877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.870710Z digest=sha256:1b3015fd690fa5734012ac099f1b76fbee8977e971f32df4e69872b85a8d613b

Observation b9c0e4d5-992e-4efb-9ce4-80e17e4a012e · outbound

This paper cites Fully sparse 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Fully sparse 3D object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.297804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.958733Z digest=sha256:82f91bd68c316a120fe309e8f4f86b39d85954f20959a753c3fd75cdd3a5c36c

Observation 65a7e81d-22fa-4f21-9b72-55d7e8009575 · outbound

This paper cites Multi-modal transformer for video retrieval.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Multi-modal transformer for video retrieval

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.281744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.024764Z digest=sha256:24b07039921b0d5134079b28ef1e508e04b0e51ae1eefef5b9a12ec879ae34f8

Observation ff00328c-a284-450d-a55a-0bed8daf3e52 · outbound

This paper cites Are we ready for autonomous driving? the KITTI vision benchmark suite.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Are we ready for autonomous driving? the KITTI vision benchmark suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.093300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.093300Z digest=sha256:e14f896b62835f7fe699e979fe84f67d0866f90d916e6805d2a5bbcdd79b6528

Observation 5d3e40f7-cc5b-4a84-9b7c-20e5284fab6a · outbound

This paper cites ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.141605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.141605Z digest=sha256:c7caf0ab554dd5a00aec97fe3e04e958077606accf97586b8bfbd6a0294b399d

Observation e667cdb0-2ac7-45a1-b463-6f4ee791ac75 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary object detection via vision and language knowledge distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.265554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.230640Z digest=sha256:9b9994a02843a90ca4408f554278e5fcea0f76be7f103689bca2171c9ac8f81d

Observation 4daed349-6ad7-46ce-8ed6-7562da3577c9 · outbound

This paper cites OneLLM: One framework to align all modalities with language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One framework to align all modalities with language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.250070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.295762Z digest=sha256:0d13313a42fe64fbe468bda86819525c318b8b8a03dd9a07d5d3e25cd315b01b

Observation 1feac009-2953-4277-953f-eb5175876ced · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.231011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.472014Z digest=sha256:eb32e52054c0450065880f4022ae0d18ca52798c6c16cf2a1a5c2892e8a36e7b

Observation a42b1dbc-f053-4791-b64e-845761745580 · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.214938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.563889Z digest=sha256:93a65ffb194b16eecb2f98437c492ced7e57b1583d193a48609155d991fd257a

Observation eb6ac23e-a5bd-4605-8d95-8fc262d9962e · outbound

This paper cites Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.655417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.655417Z digest=sha256:678a1331b903b5f952e00a5bbd2c3b507b7077c41acf5b7f1381558c7e7dde01

Observation a27b0254-1bb4-49a8-990b-1cdea1a248cb · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptFusion: Open-set Multimodal 3D Mapping

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.731888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.731888Z digest=sha256:56a3c46bc37b35047026bb3e685880ce32845c047a228af78837c98a2535b9df

Observation 1e0ae063-aee7-4e35-80d7-5e7f021646cb · outbound

This paper cites Action genome: Actions as composition of spatio-temporal scene graphs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Action genome: Actions as composition of spatio-temporal scene graphs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.198930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.811170Z digest=sha256:9671f5044525be07e118c05d72b5b527ec4725487537deb21e0121d85f3cdf39

Observation 81b7e8c8-d47a-4b5b-980c-41e71feb883e · outbound

This paper cites PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.182991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.874237Z digest=sha256:1b21c6a30f8776af04b9e76d45acd3751d3ae6be8d303d15f8f8c41aa6b936cb

Observation c734e3c7-3a44-4ebe-b8d5-b10c179ca67a · outbound

This paper cites Grounded language-image pre-training.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounded language-image pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.165658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.980304Z digest=sha256:e5bfe6300786432d538ecb9cc3e0106198f148be7c2abb386ac975ca323bea06

Observation f5844c12-2d4d-467f-b559-f74bc885388d · outbound

This paper cites OpenShape: Scaling up 3D shape representation towards open-world understanding.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenShape: Scaling up 3D shape representation towards open-world understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.146762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.053531Z digest=sha256:d871832acaf3e56fecedaa7a6eaffd8fdb4132d9e3551d39e664acfc16afa3cf

Observation 47339047-0ceb-48aa-be97-d11fc8eb48e5 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.127565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.160071Z digest=sha256:27299411761b67c5e8b77f1bbe167e8dc314913c11e280955d825670b7d02237

Observation a1648445-19ac-4617-ac79-9ebce9000c08 · outbound

This paper cites Open-vocabulary point-cloud object detection without 3D annotation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary point-cloud object detection without 3D annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.111280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.291623Z digest=sha256:e54f3c983f050212f74f390359023c3995ba0e090627bc7f438f172ddfc3e934

Observation 7f458115-3bbd-44f3-ab37-2168437dc4ee · outbound

This paper cites An End-to-End Transformer Model for 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes An End-to-End Transformer Model for 3D Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.369392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.369392Z digest=sha256:43f7ddec78804df30596db4fc47f8c0fca5477ac12a647cc7055d21bbd2e2738

Observation b0c79d08-7717-4a79-a1ee-1dd63bd53f09 · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Modeling temporal structure of decomposable motion segments for activity classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.094901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.471688Z digest=sha256:c8e6e0595f7d6c310b300f0198291c62c131fe8c07fcccc3e73850d0c9adb3a1

Observation dabae4fa-8546-48ee-8e6e-dc06fcded168 · outbound

This paper cites PyT orch: An imperative style, high-performance deep learning library.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PyT orch: An imperative style, high-performance deep learning library

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.076037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.583519Z digest=sha256:8f61645864835c8f7282a763f206c79a9b2d311d42e9a01b2419d9294cca59cf

Observation 1bd61506-50b3-4391-9cd5-498424ec4d70 · outbound

This paper cites OpenScene: 3D scene understanding with open vocabularies.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenScene: 3D scene understanding with open vocabularies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.058826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.686010Z digest=sha256:6634b28ce033736f0c4942abbb3ae08ac295f5af075a180022aa9805d9da2783

Observation 166f1343-26ab-447c-9b6d-14a5b9b70c69 · outbound

This paper cites Qi, Hao Su, Kaichun Mo, and Leonidas J.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Qi, Hao Su, Kaichun Mo, and Leonidas J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.788184Z digest=sha256:89595c7569ac5fb4fe4fd4f28c48e669484e70cefdfe2f5cc2a71881dddf49ef

Observation 2f536505-b225-487c-a519-1c9ef552fcc6 · outbound

This paper cites Frustum PointNets for 3D Object Detection from RGB-D Data.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Frustum PointNets for 3D Object Detection from RGB-D Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.886207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.886207Z digest=sha256:427ae90cec335fd55826d83d35c83fa27f27bf71cb8832825c1a98be7cd87322

Observation 0303b28b-8ef5-4362-897b-8707f7336768 · outbound

This paper cites Deep Hough Voting for 3D Object Detection in Point Clouds.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Deep Hough Voting for 3D Object Detection in Point Clouds

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.987522Z digest=sha256:d214ff73cdf1a30ee1bd21bc2e72555fdce9f6f2b189a0b63eb77520aaf2a4ac

Observation 6d60b02a-280e-4cea-8201-22b4b411d242 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:04.056417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:04.056417Z digest=sha256:cc516fbfe392707e61f5b8b670990f5000c7107736227ed8da7b74fa0605abc9

Observation 70bcaecc-0c36-41b1-9450-c0d54195480f · outbound

This paper cites PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.034757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.157255Z digest=sha256:7ccb9e3c882ec8ee9ca25756b32f875b21c922c93d41517226a32ddc08590dee

Observation cd9525d2-b4da-4dea-9189-e38ee5217e14 · outbound

This paper cites PV -RCNN: Point-voxel feature set abstraction for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PV -RCNN: Point-voxel feature set abstraction for 3D object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.021635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.256401Z digest=sha256:af91139a63a80c5973d260dacbdc4e1aa2fe8a03a440f011a8d3ba22f43a9708

Observation 75bf4426-3b0a-45c4-98d4-8e737839456c · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes VideoBERT: A joint model for video and language representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.005712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.338428Z digest=sha256:b4795125cab0a7a5aae9f22c6d6a69369139271b5c6ecd2681c5f63ee41e7c02

Observation 9536a125-4495-40da-9be9-6671aa89b29a · outbound

This paper cites Scalability in perception for autonomous driving: W aymo open dataset, 2020.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Scalability in perception for autonomous driving: W aymo open dataset, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.986709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.558032Z digest=sha256:6fb39bd1e60da6c2faf9d455ecfcd2d1f4621c5234d0e07a3a57feb748597876

Observation ca3ff166-92f1-495c-ac5c-aa2b4d7e419f · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning spatiotemporal features with 3D convolutional networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.956623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.735453Z digest=sha256:96279303a5b0f35db9afb8509f0c87b7fc2308513e2be632419c6593b6edd7a2

Observation a3914318-82fa-4337-b98e-878381f55c9f · outbound

This paper cites DSVT: Dynamic sparse voxel transformer with rotated sets.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes DSVT: Dynamic sparse voxel transformer with rotated sets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.935947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.934308Z digest=sha256:4ebb0d56a219a7cf146f7268cb06d0768d5f0817046a03a80a210dd7376301c5

Observation 76edf698-e339-4015-b383-1b7c9cfd170c · outbound

This paper cites Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.358512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.358512Z digest=sha256:437fe946fcd2bb5b88a7c55a30426246d6256989bc64d6b110e158f86880a060

Observation 5f9d2efa-63c2-4004-8447-bccd567f0af8 · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Argoverse 2: Next generation datasets for self-driving perception and forecasting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.914171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.456330Z digest=sha256:28375ae8960969065381474a7dd8e2401009d2bac96550d8d1d22f4894d1b823

Observation 6e6596de-4638-4dfd-90ba-f7288e8a3ee8 · outbound

This paper cites Transformation- equivariant 3D object detection for autonomous driving.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Transformation- equivariant 3D object detection for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.895793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.555846Z digest=sha256:4cf17c305e48586a471ad0f3caa2aafb55a6196acf9b0c247a3962c77bb614a5

Observation 74593d27-35fd-4cd5-be04-5504b25e89aa · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Towards Open Vocabulary Learning: A Survey

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.006377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.638687Z digest=sha256:fbda12b9350634cf849c4d5233fd4294e8d6b35a086f2098f7de9e79c43207d9

Observation 7e732a7a-114b-4d1f-b7be-a104fda0f275 · outbound

This paper cites FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.763277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.763277Z digest=sha256:36956811bd415827fa5aff773b1757f99576ea55a02f2e4661d3792358530075

Observation d4098634-ae30-4531-82c2-b551b1eddb60 · outbound

This paper cites 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:13:07.953455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.930584Z digest=sha256:864460deaf7cce127c5bbd1f9cd53eb5983c4dcb8a830409964c89fa50f4c25b

Observation 6c5b56a5-94f0-4a35-af10-38ab6e6203b0 · outbound

This paper cites EffiPerception: an Efficient Framework for Various Perception Tasks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes EffiPerception: an Efficient Framework for Various Perception Tasks

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.702417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.036014Z digest=sha256:36b0c7081c533140fd39a6fcfe0a91acf198aac031ab8eb782056c1bfcff31c1

Observation 91b533a6-5f6e-4b49-86d4-4fb4a3360e09 · outbound

This paper cites Graph R-CNN for Scene Graph Generation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Graph R-CNN for Scene Graph Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.519380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.116837Z digest=sha256:2ed2cbd3eaa4384225d7eda8c55e3f9f649b52b5b0471552306a8499b442df01

Observation d37adc28-e3b3-44ff-ba67-1b6365d0cabd · outbound

This paper cites Open-vocabulary DETR with conditional matching.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary DETR with conditional matching

Reference 49

Resolution
verified exact
doi, observed 2026-08-05T17:13:07.148949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.198543Z digest=sha256:7dfdb7fe56e445df17042be92490e9302a594d56ffdad5eaf09d120faca6b1cb

Observation 32c79f97-0f4c-4abb-9c36-4cfb103414ca · outbound

This paper cites Open-Vocabulary Object Detection Using Captions.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-Vocabulary Object Detection Using Captions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.337789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.300876Z digest=sha256:3ba039909a8424baae96f72c60e270e603b1a965ae7d38866d38282bcf5458f8

Observation f02b52ce-72b7-4aed-9ed4-519b2f9548b2 · outbound

This paper cites FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.878591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.391908Z digest=sha256:cdaf2d6f2a5dea8d420bce86e749995d01baf6f4b3cfb75b384f661ce4f16a08

Observation 0ffd2678-aa19-4326-be4b-169e30a6c3a8 · outbound

This paper cites OpenSight: A simple open-vocabulary framework for LiDAR-based object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenSight: A simple open-vocabulary framework for LiDAR-based object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.860512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.478801Z digest=sha256:c7011841fe3c8b3bd3c7636b7644ec72adf64d4cece6c0dc23335ad619661325

Observation c38fb163-4d80-43bf-ba6d-f93513cef3f2 · outbound

This paper cites G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.839413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.561654Z digest=sha256:d4b7e013e15a0e112dcb7be8536b8d228c266ff7f62791c7d04e60d27991bbc2

Observation 8fd75d1e-f795-4927-90fd-bd099936e1de · outbound

This paper cites OcTr: Octree-based transformer for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OcTr: Octree-based transformer for 3D object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.816793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.650142Z digest=sha256:267ebe53250f0fa43124be1b2cc4117eae5ce988a5174bd6caea6d50a3b9f8b8

Observation e84cafae-c49f-4a74-b668-7d794db8d9da · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:06.743796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:06.743796Z digest=sha256:e3d92c467629c5a0af3b077fddc593cc29a2972261a5f380faba06c773209d98

Observation af50caab-0370-43a9-b0b5-631aea67a62b · outbound

This paper cites PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.799818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.857653Z digest=sha256:daf3841eff2d2582127de0bafc5211ae36c3294527884e504dc3e328176a7b2c

Observation f4e88895-2b6d-422c-95bc-46b5a19585ec · outbound

This paper cites Then, the model continues to be trained for 20 epochs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Then, the model continues to be trained for 20 epochs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.784314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.935892Z digest=sha256:7ce6af69da6204ac9d2ada3917cb22f29bcc8262b04fa72edc1b497c74440dea

Observation 3a331ea2-4e45-46c3-b90b-5d7b8e65f565 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One Framework to Align All Modalities with Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.378508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.378508Z digest=sha256:fd01ccf1e5b95c348abb9e1c7a640db6656da784939413a6958a4da970c5408e

Pith citing papers

No inbound Pith citation observations are available.