Pith. sign in

Uni3DL: Unified Model for 3D and Language Understanding

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In this work, we present Uni3DL, a unified model for 3D and Language understanding. Distinct from existing unified vision-language models in 3D which are limited in task variety and predominantly dependent on projected multi-view images, Uni3DL operates directly on point clouds. This approach significantly expands the range of supported tasks in 3D, encompassing both vision and vision-language tasks in 3D. At the core of Uni3DL, a query transformer is designed to learn task-agnostic semantic and mask outputs by attending to 3D visual features, and a task router is employed to selectively generate task-specific outputs required for diverse tasks. With a unified architecture, our Uni3DL model enjoys seamless task decomposition and substantial parameter sharing across tasks. Uni3DL has been rigorously evaluated across diverse 3D vision-language understanding tasks, including semantic segmentation, object detection, instance segmentation, visual grounding, 3D captioning, and text-3D cross-modal retrieval. It demonstrates performance on par with or surpassing state-of-the-art (SOTA) task-specific models. We hope our benchmark and Uni3DL model will serve as a solid step to ease future research in unified models in the realm of 3D and language understanding. Project page: https://uni3dl.github.io.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Zero-Shot 3D Visual Grounding from Vision-Language Models

cs.CV · 2025-05-28 · conditional · novelty 5.0

SeeGround localizes objects in 3D scenes from natural language without 3D-specific training, using query-aligned rendered views and spatially enriched text fed to a 2D vision-language model.

citing papers explorer

Showing 1 of 1 citing paper.

  • Zero-Shot 3D Visual Grounding from Vision-Language Models cs.CV · 2025-05-28 · conditional · none · ref 66 · internal anchor

    SeeGround localizes objects in 3D scenes from natural language without 3D-specific training, using query-aligned rendered views and spatially enriched text fed to a 2D vision-language model.