Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

As of 14 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2508.16812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16812 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:06.935892Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact11
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 840f8725-d2b5-4625-9829-401e7c8370a8 · outbound

This paper cites TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.347860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.437631Z digest=sha256:5a1da1f69d2e0708147e44fabd79f3d2835fc7b914ef8ce72f6fddc2cd90fda4

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:8ac32b10fa9f8f7f1caf37d88d970ad88042a6b73d844f52be9aec0fd0978150

Observation 0323529f-d181-4656-9e98-f1df76763a63 · outbound

This paper cites Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.331370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.550840Z digest=sha256:943cf7307447c87f2dc8bfef60001d8be4eeada8442818feda684432f6ecc406

Observation 92bfde35-7860-4701-81bd-53daddc5e63f · outbound

This paper cites CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.315917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.619614Z digest=sha256:8d3b35bb57129e641dfb8436e292a0ba3b934e92e73573324296db8cff5778e4

Observation 5e6d3ffb-ca94-414c-9c2d-5ae815fbcfda · outbound

This paper cites Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.748655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.666088Z digest=sha256:7789ce61d86c09d118481c30e9339b6741758c48d52c73e9504497db85df49f1

Observation b85c14d1-2749-4e90-94ee-c905dc863425 · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.758342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.758342Z digest=sha256:0db3a96093a179506098d087a3176a2c503f4077b78edca57a0ea00fd6f4e90e

Observation 103b1be2-881b-4092-a9de-cdf760126ccd · outbound

This paper cites Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.568522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.824856Z digest=sha256:27e32f28f296ee1db9fc57a2513d060c09de2b825e8293af1b399f76359e5e84

Observation fcc17589-f298-479f-9021-265f579e2956 · outbound

This paper cites Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.337877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.870710Z digest=sha256:9f7acdb3c233d5930a2c715bbd97435301c2c76c4e583d1d1487c7e269db0737

Observation b9c0e4d5-992e-4efb-9ce4-80e17e4a012e · outbound

This paper cites Fully sparse 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Fully sparse 3D object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.297804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:01.958733Z digest=sha256:c4575f19bfebb225085f7b82906159137b560556d28425b140b188a20dcaab50

Observation 65a7e81d-22fa-4f21-9b72-55d7e8009575 · outbound

This paper cites Multi-modal transformer for video retrieval.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Multi-modal transformer for video retrieval

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.281744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.024764Z digest=sha256:97b545d23598d135fc4fcbf8a1b50047c5724b20ddf83a9e9d61a4a4349ef8a9

Observation ff00328c-a284-450d-a55a-0bed8daf3e52 · outbound

This paper cites Are we ready for autonomous driving? the KITTI vision benchmark suite.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Are we ready for autonomous driving? the KITTI vision benchmark suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.093300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.093300Z digest=sha256:e14f896b62835f7fe699e979fe84f67d0866f90d916e6805d2a5bbcdd79b6528

Observation 5d3e40f7-cc5b-4a84-9b7c-20e5284fab6a · outbound

This paper cites ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.141605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.141605Z digest=sha256:a7935e81cade3d7ec9a65a62fd69c09792ab2aaac191a986c2faf3cfb8c4c49d

Observation e667cdb0-2ac7-45a1-b463-6f4ee791ac75 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary object detection via vision and language knowledge distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.265554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.230640Z digest=sha256:3cc3e912f1e4672c3d46011c3b8f51ce9fb578c94c7da2dc636beab476fb8bed

Observation 4daed349-6ad7-46ce-8ed6-7562da3577c9 · outbound

This paper cites OneLLM: One framework to align all modalities with language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One framework to align all modalities with language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.250070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.295762Z digest=sha256:9b31e4a18d53fb2867b3f1a8c25d56a621a35d3810696be848ae91c2c6d10262

Observation 1feac009-2953-4277-953f-eb5175876ced · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.231011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.472014Z digest=sha256:4c0fc177a67cb5fef17c659df7c3cc733741e327f70d02fda89d5658bf52cdc3

Observation a42b1dbc-f053-4791-b64e-845761745580 · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.214938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.563889Z digest=sha256:d2cee1e4b4512a7d24decbb6164eec7b5f6f0626de88f6e408e7718806ce2d1b

Observation eb6ac23e-a5bd-4605-8d95-8fc262d9962e · outbound

This paper cites Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.655417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.655417Z digest=sha256:678a1331b903b5f952e00a5bbd2c3b507b7077c41acf5b7f1381558c7e7dde01

Observation a27b0254-1bb4-49a8-990b-1cdea1a248cb · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptFusion: Open-set Multimodal 3D Mapping

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.731888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.731888Z digest=sha256:f0f6349bf33588a9013f650b4b9eab6d3f2af622473951e44c1cc437a86a0eac

Observation 1e0ae063-aee7-4e35-80d7-5e7f021646cb · outbound

This paper cites Action genome: Actions as composition of spatio-temporal scene graphs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Action genome: Actions as composition of spatio-temporal scene graphs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.198930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.811170Z digest=sha256:85ca089fc00b60b4b722374336cedcafeb2a613ab31c44955a4384afe988d6e5

Observation 81b7e8c8-d47a-4b5b-980c-41e71feb883e · outbound

This paper cites PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.182991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.874237Z digest=sha256:bad49f34ee41a06dc2831a81f2b5c038b6c3609a40c2b58272853337cae94c45

Observation c734e3c7-3a44-4ebe-b8d5-b10c179ca67a · outbound

This paper cites Grounded language-image pre-training.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounded language-image pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.165658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:02.980304Z digest=sha256:703c9ce5bbf37ed64eed34a670683826d37fe4e0c337a202a5863b02d6f2dbd7

Observation f5844c12-2d4d-467f-b559-f74bc885388d · outbound

This paper cites OpenShape: Scaling up 3D shape representation towards open-world understanding.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenShape: Scaling up 3D shape representation towards open-world understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.146762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.053531Z digest=sha256:03676a889e706126b6bce57663ae5a5a224a69dd4ec169d88c9cf318fad6a601

Observation 47339047-0ceb-48aa-be97-d11fc8eb48e5 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.127565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.160071Z digest=sha256:cd34b8c3c2deea2e24514a4e428f8058f75b8732371f13e00ccdef48e708304f

Observation a1648445-19ac-4617-ac79-9ebce9000c08 · outbound

This paper cites Open-vocabulary point-cloud object detection without 3D annotation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary point-cloud object detection without 3D annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.111280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.291623Z digest=sha256:eea34845bbc54313217c116f24b9e812a2d167e3731f9f683efd5b2b5f9d144d

Observation 7f458115-3bbd-44f3-ab37-2168437dc4ee · outbound

This paper cites An End-to-End Transformer Model for 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes An End-to-End Transformer Model for 3D Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.369392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.369392Z digest=sha256:889ca63341401798e0324d23eaadfe7f629ae836ea0d946ecc6b7418d6512fca

Observation b0c79d08-7717-4a79-a1ee-1dd63bd53f09 · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Modeling temporal structure of decomposable motion segments for activity classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.094901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.471688Z digest=sha256:e15e5fe33221d058e7e523d52bd5fb99d0a05116ccb82c9e0938b13039d427cf

Observation dabae4fa-8546-48ee-8e6e-dc06fcded168 · outbound

This paper cites PyT orch: An imperative style, high-performance deep learning library.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PyT orch: An imperative style, high-performance deep learning library

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.076037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.583519Z digest=sha256:8d9a2447aaee7cceb108d691e2634bfcefcc100f46766fb2b8a25dcfcecfdd6b

Observation 1bd61506-50b3-4391-9cd5-498424ec4d70 · outbound

This paper cites OpenScene: 3D scene understanding with open vocabularies.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenScene: 3D scene understanding with open vocabularies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.058826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.686010Z digest=sha256:60d7f76a931078cfda887c8f7f89f46e185b515c01b53900afd22a802159652f

Observation 166f1343-26ab-447c-9b6d-14a5b9b70c69 · outbound

This paper cites Qi, Hao Su, Kaichun Mo, and Leonidas J.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Qi, Hao Su, Kaichun Mo, and Leonidas J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.788184Z digest=sha256:c309ad80dd52d5fa43c255d1c902f9a15f37b430473045bd61cb31dd43bb84a5

Observation 2f536505-b225-487c-a519-1c9ef552fcc6 · outbound

This paper cites Frustum PointNets for 3D Object Detection from RGB-D Data.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Frustum PointNets for 3D Object Detection from RGB-D Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.886207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.886207Z digest=sha256:521c9d15bbae2bed3b7ba709a87c126de4fd49bd4450026b11870880e7af54fa

Observation 0303b28b-8ef5-4362-897b-8707f7336768 · outbound

This paper cites Deep Hough Voting for 3D Object Detection in Point Clouds.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Deep Hough Voting for 3D Object Detection in Point Clouds

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:03.987522Z digest=sha256:46850c51c4a6c7ed160b4670fafb73497e3ddfe77d076ec22d5f4b169a82c150

Observation 6d60b02a-280e-4cea-8201-22b4b411d242 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:04.056417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:04.056417Z digest=sha256:cc516fbfe392707e61f5b8b670990f5000c7107736227ed8da7b74fa0605abc9

Observation 70bcaecc-0c36-41b1-9450-c0d54195480f · outbound

This paper cites PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.034757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.157255Z digest=sha256:05cb5f48a6ac3d44c94b20d53fa82d0e0ecb5620f5673a1350a635cb33f37373

Observation cd9525d2-b4da-4dea-9189-e38ee5217e14 · outbound

This paper cites PV -RCNN: Point-voxel feature set abstraction for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PV -RCNN: Point-voxel feature set abstraction for 3D object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.021635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.256401Z digest=sha256:dbdb922a9f7e7285cdbe483c9a46f9f1ea95faffabe903e32a38fc16addd8098

Observation 75bf4426-3b0a-45c4-98d4-8e737839456c · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes VideoBERT: A joint model for video and language representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.005712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.338428Z digest=sha256:56a1c67d1cc44629e9170437e57833731b8ef6378ba44f8a5d2df9acb28db680

Observation 9536a125-4495-40da-9be9-6671aa89b29a · outbound

This paper cites Scalability in perception for autonomous driving: W aymo open dataset, 2020.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Scalability in perception for autonomous driving: W aymo open dataset, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.986709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.558032Z digest=sha256:0f960d19236a02fc7efa99ef541c176c03b42b26345297ac738429a86d0a6d1f

Observation ca3ff166-92f1-495c-ac5c-aa2b4d7e419f · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning spatiotemporal features with 3D convolutional networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.956623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.735453Z digest=sha256:98721966e9a65dcc89ac77058ef7cf21bc8c4ef672f83a1b7f9f9b7ca7be89d4

Observation a3914318-82fa-4337-b98e-878381f55c9f · outbound

This paper cites DSVT: Dynamic sparse voxel transformer with rotated sets.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes DSVT: Dynamic sparse voxel transformer with rotated sets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.935947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:04.934308Z digest=sha256:f57249cad7a5acc1ba209a7ee5a71a0bf688c70c2dbe0b05b4fda9783dbb8f86

Observation 76edf698-e339-4015-b383-1b7c9cfd170c · outbound

This paper cites Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.358512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.358512Z digest=sha256:437fe946fcd2bb5b88a7c55a30426246d6256989bc64d6b110e158f86880a060

Observation 5f9d2efa-63c2-4004-8447-bccd567f0af8 · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Argoverse 2: Next generation datasets for self-driving perception and forecasting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.914171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:05.456330Z digest=sha256:94f0c5a863316a35f81601d993d88fb54974bee378a170ec7aca53bfde7cdde8

Observation 6e6596de-4638-4dfd-90ba-f7288e8a3ee8 · outbound

This paper cites Transformation- equivariant 3D object detection for autonomous driving.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Transformation- equivariant 3D object detection for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.895793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:05.555846Z digest=sha256:d3c0401dfa25ebd9fa3b3cf7c1401460848cfc9df732e9f00a216ee7b659e3ae

Observation 74593d27-35fd-4cd5-be04-5504b25e89aa · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Towards Open Vocabulary Learning: A Survey

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.006377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:05.638687Z digest=sha256:49eb23f216df91644b921c9a2778b080bf9316f7cc5bc07c50ff60d16c1b0f6b

Observation 7e732a7a-114b-4d1f-b7be-a104fda0f275 · outbound

This paper cites FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.763277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.763277Z digest=sha256:3acfe7c098452c055156f08507dd6a77c47da5621ab8e205a5f1a3ec39bc5036

Observation d4098634-ae30-4531-82c2-b551b1eddb60 · outbound

This paper cites 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:13:07.953455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:05.930584Z digest=sha256:07bb8fabe1201ae25cf74cc1fe60f6480b3ce98753d052388ef2886d226265e0

Observation 6c5b56a5-94f0-4a35-af10-38ab6e6203b0 · outbound

This paper cites EffiPerception: an Efficient Framework for Various Perception Tasks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes EffiPerception: an Efficient Framework for Various Perception Tasks

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.702417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.036014Z digest=sha256:67a3d13f958e522e310878ec8859b5bb123694b720c9f7ef31b0e40b0a275908

Observation 91b533a6-5f6e-4b49-86d4-4fb4a3360e09 · outbound

This paper cites Graph R-CNN for Scene Graph Generation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Graph R-CNN for Scene Graph Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.519380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.116837Z digest=sha256:6c0619520939bcc7aefefbac57547419a2bdcec9a1d103ff9022619bf12e7c9f

Observation d37adc28-e3b3-44ff-ba67-1b6365d0cabd · outbound

This paper cites Open-vocabulary DETR with conditional matching.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary DETR with conditional matching

Reference 49

Resolution
verified exact
doi, observed 2026-08-05T17:13:07.148949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.198543Z digest=sha256:ea70613a29698fba8e14e5dd392f97e406f008aada3a85d48e5f9b7225fef54e

Observation 32c79f97-0f4c-4abb-9c36-4cfb103414ca · outbound

This paper cites Open-Vocabulary Object Detection Using Captions.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-Vocabulary Object Detection Using Captions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.337789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.300876Z digest=sha256:73569962e7db8e3496558b47491248cd5c32a874690f465a86ab122d2b89a780

Observation f02b52ce-72b7-4aed-9ed4-519b2f9548b2 · outbound

This paper cites FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.878591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.391908Z digest=sha256:57c55d7e530fce9141287b7da24d198422c049d524553fd93d7a0119a2bfe50e

Observation 0ffd2678-aa19-4326-be4b-169e30a6c3a8 · outbound

This paper cites OpenSight: A simple open-vocabulary framework for LiDAR-based object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenSight: A simple open-vocabulary framework for LiDAR-based object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.860512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.478801Z digest=sha256:a1e59f1667fd4b90c77549d6ece05888fdeba24b13c007b2f4ae6dd815299a94

Observation c38fb163-4d80-43bf-ba6d-f93513cef3f2 · outbound

This paper cites G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.839413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.561654Z digest=sha256:893fa5d0f06b2ff56b88f3da145623f4b82ce62f3303641804fcede53ee84724

Observation 8fd75d1e-f795-4927-90fd-bd099936e1de · outbound

This paper cites OcTr: Octree-based transformer for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OcTr: Octree-based transformer for 3D object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.816793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.650142Z digest=sha256:582479eb3d90a6e7208f754c8ead01501811969a5d362cb753909e34e71f0751

Observation e84cafae-c49f-4a74-b668-7d794db8d9da · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:06.743796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:06.743796Z digest=sha256:e3d92c467629c5a0af3b077fddc593cc29a2972261a5f380faba06c773209d98

Observation af50caab-0370-43a9-b0b5-631aea67a62b · outbound

This paper cites PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.799818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.857653Z digest=sha256:cab218bb9f3c92ec4449f6f379a1f9b484db2d4c6010a3a605a73d31471e4c0a

Observation f4e88895-2b6d-422c-95bc-46b5a19585ec · outbound

This paper cites Then, the model continues to be trained for 20 epochs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Then, the model continues to be trained for 20 epochs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.784314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T17:13:06.935892Z digest=sha256:09862c8426ee79a0292749a0ea6631d6b5a8c671c7121dbd21ccae502ac907c2

Observation 3a331ea2-4e45-46c3-b90b-5d7b8e65f565 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One Framework to Align All Modalities with Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.378508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.378508Z digest=sha256:d4ba20bab486fb3113dcd130bc2f6a02f19844fcfb7a71caa0118da954d97c31

Pith citing papers

No inbound Pith citation observations are available.