Pith. sign in

REVIEW 3 cited by

Unifying 3D Vision-Language Understanding via Promptable Queries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11442 v2 pith:YEPSRUJQ submitted 2024-05-19 cs.CV

classification cs.CV
keywords pq3dtasksd-vlmodelrepresentationssceneunifiedmulti-task
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene. However, a considerable gap exists between existing methods and such a unified model, due to the independent application of representation and insufficient exploration of 3D multi-task training. In this paper, we introduce PQ3D, a unified model capable of using Promptable Queries to tackle a wide range of 3D-VL tasks, from low-level instance segmentation to high-level reasoning and planning. This is achieved through three key innovations: (1) unifying various 3D scene representations (i.e., voxels, point clouds, multi-view images) into a shared 3D coordinate space by segment-level grouping, (2) an attention-based query decoder for task-specific information retrieval guided by prompts, and (3) universal output heads for different tasks to support multi-task training. Tested across ten diverse 3D-VL datasets, PQ3D demonstrates impressive performance on these tasks, setting new records on most benchmarks. Particularly, PQ3D improves the state-of-the-art on ScanNet200 by 4.9% (AP25), ScanRefer by 5.4% (acc@0.5), Multi3DRefer by 11.7% (F1@0.5), and Scan2Cap by 13.4% (CIDEr@0.5). Moreover, PQ3D supports flexible inference with individual or combined forms of available 3D representations, e.g., solely voxel input.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Multi-view relational distillation transfers geometric knowledge to vision-language models by matching cross-view patch similarity matrices, improving spatial reasoning with minimal overhead.

  2. Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Qwen-3D feeds the full visual-language representation of a 3D scene into a mask decoder, beating prior 3D LMMs on grounding and segmentation benchmarks.

  3. ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ViGiL3D is a 350-prompt diagnostic dataset showing that existing 3D visual grounding models lose 20 or more points on linguistically diverse prompts compared to ScanRefer.

Pith tools