Pith. sign in

REVIEW 3 cited by

RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00962 v4 pith:TLKSLAYA submitted 2023-04-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords learningcontrastivelanguageopen-worldregionalsceneunderstandingdense
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a lightweight and scalable Regional Point-Language Contrastive learning framework, namely \textbf{RegionPLC}, for open-world 3D scene understanding, aiming to identify and recognize open-set objects and categories. Specifically, based on our empirical studies, we introduce a 3D-aware SFusion strategy that fuses 3D vision-language pairs derived from multiple 2D foundation models, yielding high-quality, dense region-level language descriptions without human 3D annotations. Subsequently, we devise a region-aware point-discriminative contrastive learning objective to enable robust and effective 3D learning from dense regional language supervision. We carry out extensive experiments on ScanNet, ScanNet200, and nuScenes datasets, and our model outperforms prior 3D open-world scene understanding approaches by an average of 17.2\% and 9.1\% for semantic and instance segmentation, respectively, while maintaining greater scalability and lower resource demands. Furthermore, our method has the flexibility to be effortlessly integrated with language models to enable open-ended grounded 3D reasoning without extra task-specific training. Code is available at https://github.com/CVMI-Lab/PLA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  2. Geospatial-Prior Guidance for 3D Semantic Scene Completion

    cs.CV 2026-08 conditional novelty 6.0 of 10

    GeoScene uses weighted fusion of satellite imagery and OpenStreetMap priors to improve camera-based 3D semantic scene completion on SemanticKITTI and SSCBench-KITTI-360.

  3. OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model

    cs.CV 2025-06 reject novelty 3.0 of 10

    OpenMaskDINO3D reports state-of-the-art 3D reasoning segmentation with a LISA-style SEG token and object identifiers, but uses Mask3D pseudo-labels as ground truth and lacks released code.

Pith tools