Pith. sign in

REVIEW 2 cited by

Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09456 v1 pith:VCFY375R submitted 2023-09-18 cs.CV

classification cs.CV
keywords datasetsdetectionobject2sceneopen-vocabularyobjectobjectscategoriesexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Point cloud-based open-vocabulary 3D object detection aims to detect 3D categories that do not have ground-truth annotations in the training set. It is extremely challenging because of the limited data and annotations (bounding boxes with class labels or text descriptions) of 3D scenes. Previous approaches leverage large-scale richly-annotated image datasets as a bridge between 3D and category semantics but require an extra alignment process between 2D images and 3D points, limiting the open-vocabulary ability of 3D detectors. Instead of leveraging 2D images, we propose Object2Scene, the first approach that leverages large-scale large-vocabulary 3D object datasets to augment existing 3D scene datasets for open-vocabulary 3D object detection. Object2Scene inserts objects from different sources into 3D scenes to enrich the vocabulary of 3D scene datasets and generates text descriptions for the newly inserted objects. We further introduce a framework that unifies 3D detection and visual grounding, named L3Det, and propose a cross-domain category-level contrastive learning approach to mitigate the domain gap between 3D objects from different datasets. Extensive experiments on existing open-vocabulary 3D object detection benchmarks show that Object2Scene obtains superior performance over existing methods. We further verify the effectiveness of Object2Scene on a new benchmark OV-ScanNet-200, by holding out all rare categories as novel categories not seen during training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    DDStereo uses two lightweight decoder branches sharing object queries plus a compact disparity extractor to deliver SOTA closed- and open-set accuracy with real-time inference on stereo 3D benchmarks.

  2. TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP

    cs.CV 2025-07 conditional novelty 5.0 of 10

    TriCLIP-3D encodes point clouds, images, and text with one frozen CLIP model plus adapters, reporting 6.5-point AP25 gains on EmbodiedScan 3D detection and grounding while cutting trainable parameters by 58%.

Pith tools