Pith. sign in

Paper Citation Record · LEDGER

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2508.16812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16812 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:13:06.935892Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact11
  • verified fuzzy32
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 840f8725-d2b5-4625-9829-401e7c8370a8 · outbound

This paper cites TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.347860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.437631Z digest=sha256:3838de3fb850a810f577e7d10cc27ae35f01a70b7ef7a62f2c927140367527dd

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:297b01a886a648b3e90000058a2496fd45336ef302470c3a36a715fd883dd13a

Observation 0323529f-d181-4656-9e98-f1df76763a63 · outbound

This paper cites Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Lang, Sourabh V ora, V enice Erin Liong, Qiang Xu, Anush Krishnan, Y u Pan, Giancarlo Baldan, and Oscar Beijbom

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.331370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.550840Z digest=sha256:5afb4c0fe7094cd71c5f2190fbfd847eb441a3ecb6a9855f61de97e0cb904295

Observation 92bfde35-7860-4701-81bd-53daddc5e63f · outbound

This paper cites CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CoDA: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3D object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.315917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.619614Z digest=sha256:20e8d4d4672ad139a4c76228a131978b7873911e243327bf9714b04ac3ebe96f

Observation 5e6d3ffb-ca94-414c-9c2d-5ae815fbcfda · outbound

This paper cites Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.748655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.666088Z digest=sha256:b7ceb669046547e0c02603e712db3a4ab0aa64fedfc7e9242b4b06eba9ef11b0

Observation b85c14d1-2749-4e90-94ee-c905dc863425 · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.758342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.758342Z digest=sha256:531530b9b5ce9499a519b602d0ae2b11423ace0b51cd1ebb51dbfa302f3bfc4d

Observation 103b1be2-881b-4092-a9de-cdf760126ccd · outbound

This paper cites Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.568522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.824856Z digest=sha256:8215377bbb1de38638d83b460d552a054992fad35eabd4f75aee5c39a0032266

Observation fcc17589-f298-479f-9021-265f579e2956 · outbound

This paper cites Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Find n' Propagate: Open-Vocabulary 3D Object Detection in Urban Environments

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.337877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.870710Z digest=sha256:209ac98f962e9c65d8674b3af8f231fa5ccc903d06e5c57890dfea1a11a7557c

Observation b9c0e4d5-992e-4efb-9ce4-80e17e4a012e · outbound

This paper cites Fully sparse 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Fully sparse 3D object detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.297804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:01.958733Z digest=sha256:78879d3516d2900cce6fa8ba0742f9f5eefdb38a7763a91e59982d4eedd4a758

Observation 65a7e81d-22fa-4f21-9b72-55d7e8009575 · outbound

This paper cites Multi-modal transformer for video retrieval.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Multi-modal transformer for video retrieval

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.281744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.024764Z digest=sha256:844e35b5ee41f47811ba4eac425dec0c290ebe4d88335d89f62376fa83068664

Observation ff00328c-a284-450d-a55a-0bed8daf3e52 · outbound

This paper cites Are we ready for autonomous driving? the KITTI vision benchmark suite.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Are we ready for autonomous driving? the KITTI vision benchmark suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.093300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.093300Z digest=sha256:28aa010dd41a3f0a5f16dfa9e00bdb8a7f17b6ba50c51c7eac63d99a4cd1eda3

Observation 5d3e40f7-cc5b-4a84-9b7c-20e5284fab6a · outbound

This paper cites ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.141605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.141605Z digest=sha256:96db4ce1d2182536e277068d2fab4cee0d4be3e370d708bc02e861c8cb76d444

Observation e667cdb0-2ac7-45a1-b463-6f4ee791ac75 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary object detection via vision and language knowledge distillation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.265554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.230640Z digest=sha256:8c278684ead88aecf7df0cbf83d5e1adf7dd346190d0dfd863a36f6694b5a5c6

Observation 4daed349-6ad7-46ce-8ed6-7562da3577c9 · outbound

This paper cites OneLLM: One framework to align all modalities with language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One framework to align all modalities with language

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.250070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.295762Z digest=sha256:cfa3a3dc67d86b339d2c1ad245c465539439a6e6c451793b8ac20c94270bfcbc

Observation 1feac009-2953-4277-953f-eb5175876ced · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.231011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.472014Z digest=sha256:b82090f03419d5dcc883b50951e7fb7de9e17608ef3f8d0fcbd7cbe20ca843cf

Observation a42b1dbc-f053-4791-b64e-845761745580 · outbound

This paper cites Jones, and Vishal M.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Jones, and Vishal M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.214938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.563889Z digest=sha256:9d9fc620ab19efc9f0317991143620d452b2675ef7a879bd16555052f3dbafd2

Observation eb6ac23e-a5bd-4605-8d95-8fc262d9962e · outbound

This paper cites Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Long short-term memory.Neural Comput., 9 (8):1735–1780, November 1997

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.655417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.655417Z digest=sha256:4d727e49c5a5e3ee5a897ee9f0d7efefb53b76c66cc5b0b606a89d15e79a3bab

Observation a27b0254-1bb4-49a8-990b-1cdea1a248cb · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes ConceptFusion: Open-set Multimodal 3D Mapping

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.731888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.731888Z digest=sha256:6034da1d9aa792eaf2b71c1a004fbe2fabe5bdae21d1fbcbfc281b5eee167bd0

Observation 1e0ae063-aee7-4e35-80d7-5e7f021646cb · outbound

This paper cites Action genome: Actions as composition of spatio-temporal scene graphs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Action genome: Actions as composition of spatio-temporal scene graphs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.198930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.811170Z digest=sha256:43621b2c7583acbd93c27d035ca38bf815a8b9b55a5c2f2ee3db481d19d1c4a8

Observation 81b7e8c8-d47a-4b5b-980c-41e71feb883e · outbound

This paper cites PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PF3Det: A prompted foundation feature assisted visual LiDAR 3D detector

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.182991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.874237Z digest=sha256:5baadb59c0488b009ebc3e814ecb82981ecdaa7cadbf2dce789295ed2687983b

Observation c734e3c7-3a44-4ebe-b8d5-b10c179ca67a · outbound

This paper cites Grounded language-image pre-training.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounded language-image pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.165658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:02.980304Z digest=sha256:370ab0a26fd9f30f7ce0880318ef7472e88ba830cdd4788c0b4cbeaacf3092d2

Observation f5844c12-2d4d-467f-b559-f74bc885388d · outbound

This paper cites OpenShape: Scaling up 3D shape representation towards open-world understanding.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenShape: Scaling up 3D shape representation towards open-world understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.146762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.053531Z digest=sha256:81377d7a08da798f6acf5cf341b4fd5df4daddaf68cb16e649d8e077349af82c

Observation 47339047-0ceb-48aa-be97-d11fc8eb48e5 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.127565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.160071Z digest=sha256:e26d97e37fdae2da8dc628d4186a3df1ece4385efcd260f36f4e3dabcfb59348

Observation a1648445-19ac-4617-ac79-9ebce9000c08 · outbound

This paper cites Open-vocabulary point-cloud object detection without 3D annotation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary point-cloud object detection without 3D annotation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.111280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.291623Z digest=sha256:2d4bf1bcbbf920354113239d745911ba7e2a448b9e0aeeab0cdcd4bbed348a5a

Observation 7f458115-3bbd-44f3-ab37-2168437dc4ee · outbound

This paper cites An End-to-End Transformer Model for 3D Object Detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes An End-to-End Transformer Model for 3D Object Detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.369392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.369392Z digest=sha256:ce9a9b03d8255f82352bdf958aca425ccc737c2cca33708b85f4eb20745fda92

Observation b0c79d08-7717-4a79-a1ee-1dd63bd53f09 · outbound

This paper cites Modeling temporal structure of decomposable motion segments for activity classification.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Modeling temporal structure of decomposable motion segments for activity classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.094901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.471688Z digest=sha256:bf9abb3aa7e572a2c71e911ca86cd9b834b1da7efa137466663cd39793e1fc77

Observation dabae4fa-8546-48ee-8e6e-dc06fcded168 · outbound

This paper cites PyT orch: An imperative style, high-performance deep learning library.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PyT orch: An imperative style, high-performance deep learning library

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.076037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.583519Z digest=sha256:89df4ee264204e89af2731a321fae7dfda944edac97cb1166c3e943cc8ad54c0

Observation 1bd61506-50b3-4391-9cd5-498424ec4d70 · outbound

This paper cites OpenScene: 3D scene understanding with open vocabularies.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenScene: 3D scene understanding with open vocabularies

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.058826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.686010Z digest=sha256:7abd25b6c63482853fed61917ed21c88209eb2d1031d4db08a9f4fd2306b4437

Observation 166f1343-26ab-447c-9b6d-14a5b9b70c69 · outbound

This paper cites Qi, Hao Su, Kaichun Mo, and Leonidas J.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Qi, Hao Su, Kaichun Mo, and Leonidas J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.788184Z digest=sha256:fcbc437d404427799c45c889fdb4564facb4f530338ebc9631913782c89287d4

Observation 2f536505-b225-487c-a519-1c9ef552fcc6 · outbound

This paper cites Frustum PointNets for 3D Object Detection from RGB-D Data.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Frustum PointNets for 3D Object Detection from RGB-D Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:03.886207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:03.886207Z digest=sha256:7462e9d331392c6cee77a74bccba5f0175bcc49dec179ce0b7317c38cf645016

Observation 0303b28b-8ef5-4362-897b-8707f7336768 · outbound

This paper cites Deep Hough Voting for 3D Object Detection in Point Clouds.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Deep Hough Voting for 3D Object Detection in Point Clouds

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.084109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:03.987522Z digest=sha256:50f0adf63294071d2d6c80078d636ba6ac8bac2291d7d9225312865d1d730443

Observation 6d60b02a-280e-4cea-8201-22b4b411d242 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:04.056417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:04.056417Z digest=sha256:bf5ae73d39076348dfa1b233bbc7bec91c1b66c1422c126d497533cb7b826971

Observation 70bcaecc-0c36-41b1-9450-c0d54195480f · outbound

This paper cites PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.034757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.157255Z digest=sha256:37d7512d76b3bf7d4bff7539702aa6b3aef8ae97d4e159664435f055f53dd69f

Observation cd9525d2-b4da-4dea-9189-e38ee5217e14 · outbound

This paper cites PV -RCNN: Point-voxel feature set abstraction for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PV -RCNN: Point-voxel feature set abstraction for 3D object detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.021635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.256401Z digest=sha256:9f2a7d26772e62bcc5c7ef56950ce117c526032e4835a39678339ac49a6c8d35

Observation 75bf4426-3b0a-45c4-98d4-8e737839456c · outbound

This paper cites VideoBERT: A joint model for video and language representation learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes VideoBERT: A joint model for video and language representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:09.005712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.338428Z digest=sha256:a3c77bc788aa9950aea991b3de26fbdb9938c25ec4858a3d97e8a1d029dad91d

Observation 9536a125-4495-40da-9be9-6671aa89b29a · outbound

This paper cites Scalability in perception for autonomous driving: W aymo open dataset, 2020.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Scalability in perception for autonomous driving: W aymo open dataset, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.986709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.558032Z digest=sha256:7d2ab9edd3edb5b7bd65c3912a6e9c7383d9d9e1e7cf817deb9131e08dd20507

Observation ca3ff166-92f1-495c-ac5c-aa2b4d7e419f · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Learning spatiotemporal features with 3D convolutional networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.956623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.735453Z digest=sha256:ac4106184624b1dc183e3b9758cbf9add68f1000be5eaf8e4fc5ef82af36eaf3

Observation a3914318-82fa-4337-b98e-878381f55c9f · outbound

This paper cites DSVT: Dynamic sparse voxel transformer with rotated sets.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes DSVT: Dynamic sparse voxel transformer with rotated sets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.935947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:04.934308Z digest=sha256:b616feaa40af88c18f838e9588f16d94e28d4c28124f43342887a317c3c2c59b

Observation 76edf698-e339-4015-b383-1b7c9cfd170c · outbound

This paper cites Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Hierarchical open-vocabulary 3D scene graphs for language-grounded robot navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.358512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.358512Z digest=sha256:40c14c3e1e78b2c38ad3867d56c9fad8ac7f24742a939896b7a2faecb2875c39

Observation 5f9d2efa-63c2-4004-8447-bccd567f0af8 · outbound

This paper cites Argoverse 2: Next generation datasets for self-driving perception and forecasting.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Argoverse 2: Next generation datasets for self-driving perception and forecasting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.914171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.456330Z digest=sha256:49e221e4a8d90f44e87d8b5b80686e3a14a095b5fd7cc871ab8f0fc11c990cac

Observation 6e6596de-4638-4dfd-90ba-f7288e8a3ee8 · outbound

This paper cites Transformation- equivariant 3D object detection for autonomous driving.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Transformation- equivariant 3D object detection for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.895793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.555846Z digest=sha256:e98037bc1d3cd660bd68d73111546731655d7c154b9f8e6cb488767b5aa2620f

Observation 74593d27-35fd-4cd5-be04-5504b25e89aa · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Towards Open Vocabulary Learning: A Survey

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:08.006377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.638687Z digest=sha256:79879a1bdacd272de537679e8b00b3722600d0c237590e4e17b387e6147a0c5f

Observation 7e732a7a-114b-4d1f-b7be-a104fda0f275 · outbound

This paper cites FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:05.763277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:05.763277Z digest=sha256:090268bb221278f19f67ada99f3069bfc15cc5a30a341dd3f851b9f8bcfae0f7

Observation d4098634-ae30-4531-82c2-b551b1eddb60 · outbound

This paper cites 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes 3DifFusionDet: Diffusion model for 3D object detection with robust LiDAR-camera fusion.arXiv preprint arXiv:2311.0374, 2023

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-05T17:13:07.953455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:05.930584Z digest=sha256:647b248d398f9f3fe11d0bc93745017a3f6504f5d9212efd589ededaf838d7ec

Observation 6c5b56a5-94f0-4a35-af10-38ab6e6203b0 · outbound

This paper cites EffiPerception: an Efficient Framework for Various Perception Tasks.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes EffiPerception: an Efficient Framework for Various Perception Tasks

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.702417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.036014Z digest=sha256:dbf65710534591ed45f702f786c886ed2bf16341e4ed816d5b3c17acb3357a39

Observation 91b533a6-5f6e-4b49-86d4-4fb4a3360e09 · outbound

This paper cites Graph R-CNN for Scene Graph Generation.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Graph R-CNN for Scene Graph Generation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.519380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.116837Z digest=sha256:859652c05cb57e7df556c649fa30e03bee13a9e52de7f8256223cf105b38f6a6

Observation d37adc28-e3b3-44ff-ba67-1b6365d0cabd · outbound

This paper cites Open-vocabulary DETR with conditional matching.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-vocabulary DETR with conditional matching

Reference 49

Resolution
verified exact
doi, observed 2026-08-05T17:13:07.148949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.198543Z digest=sha256:af2c08ab457df18415ca10b14d1bf76cfa1937200c5c338deaf869f3673298c1

Observation 32c79f97-0f4c-4abb-9c36-4cfb103414ca · outbound

This paper cites Open-Vocabulary Object Detection Using Captions.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Open-Vocabulary Object Detection Using Captions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:13:07.337789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.300876Z digest=sha256:38966d2c17b787c8533dc0499f659538dcc4b717b2f014f64865e6718080b248

Observation f02b52ce-72b7-4aed-9ed4-519b2f9548b2 · outbound

This paper cites FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes FM-OV3D: Foundation model-based cross-modal knowledge blending for open-vocabulary 3D detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.878591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.391908Z digest=sha256:46af47823ccf8e4723e1a68e4b6f04ffe36f27c74d1e652f844d7cd973c13c6c

Observation 0ffd2678-aa19-4326-be4b-169e30a6c3a8 · outbound

This paper cites OpenSight: A simple open-vocabulary framework for LiDAR-based object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OpenSight: A simple open-vocabulary framework for LiDAR-based object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.860512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.478801Z digest=sha256:7a5d7040301f0b3b1eec1955dd26fb416508249aba0d0fcd6a8540f6a85aacad

Observation c38fb163-4d80-43bf-ba6d-f93513cef3f2 · outbound

This paper cites G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes G, Anastasis Stathopoulos, Manmohan Chandraker, and Dimitris Metaxas

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.839413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.561654Z digest=sha256:01d88377b28f6eb7090082d8cb9b6246540322f19fe33933daf351a7b570b68d

Observation 8fd75d1e-f795-4927-90fd-bd099936e1de · outbound

This paper cites OcTr: Octree-based transformer for 3D object detection.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OcTr: Octree-based transformer for 3D object detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.816793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.650142Z digest=sha256:9004fd1f5e4564932496ddcfc1007fc9974d71fae3640ccff0e95cac773fc2a7

Observation e84cafae-c49f-4a74-b668-7d794db8d9da · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes CogVLM: Visual Expert for Pretrained Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:06.743796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:06.743796Z digest=sha256:8300044cb1a50af57b53ce469b8abdea5058616c867eff7a8ac1d1eb8593a474

Observation af50caab-0370-43a9-b0b5-631aea67a62b · outbound

This paper cites PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes PointCLIP V2: Prompting CLIP and GPT for powerful 3D open-world learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.799818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.857653Z digest=sha256:5d8632767f0f72ffa6b3bdc4fd032b7e4e15a3b9843ae50578b97e6a0d06a6e9

Observation f4e88895-2b6d-422c-95bc-46b5a19585ec · outbound

This paper cites Then, the model continues to be trained for 20 epochs.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Then, the model continues to be trained for 20 epochs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:13:08.784314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T17:13:06.935892Z digest=sha256:3678a1f886ea9543561a25159fc1c1c6240ef1bff7a30403b316a547f4c1cb68

Observation 3a331ea2-4e45-46c3-b90b-5d7b8e65f565 · outbound

This paper cites OneLLM: One Framework to Align All Modalities with Language.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes OneLLM: One Framework to Align All Modalities with Language

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:02.378508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:02.378508Z digest=sha256:51c4e687bf4ae4155a5c6299a815732b0769872d47077c39cc2430ac44bef8bf

Pith citing papers

No inbound Pith citation observations are available.