Pith. sign in

REVIEW 6 cited by

Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.02329 v1 pith:2XW44G3Q submitted 2023-06-04 cs.CV

classification cs.CV
keywords answeringquestionscenetasksdownstreammulti-clippre-trainingvision-language
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Training models to apply common-sense linguistic knowledge and visual concepts from 2D images to 3D scene understanding is a promising direction that researchers have only recently started to explore. However, it still remains understudied whether 2D distilled knowledge can provide useful representations for downstream 3D vision-language tasks such as 3D question answering. In this paper, we propose a novel 3D pre-training Vision-Language method, namely Multi-CLIP, that enables a model to learn language-grounded and transferable 3D scene point cloud representations. We leverage the representational power of the CLIP model by maximizing the agreement between the encoded 3D scene features and the corresponding 2D multi-view image and text embeddings in the CLIP space via a contrastive objective. To validate our approach, we consider the challenging downstream tasks of 3D Visual Question Answering (3D-VQA) and 3D Situated Question Answering (3D-SQA). To this end, we develop novel multi-modal transformer-based architectures and we demonstrate how our pre-training method can benefit their performance. Quantitative and qualitative experimental results show that Multi-CLIP outperforms state-of-the-art works across the downstream tasks of 3D-VQA and 3D-SQA and leads to a well-structured 3D scene feature space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ego-Human Motion Prediction with 3D-Aware LLM

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Ego3DLM jointly predicts past and future 3D body pose and motion descriptions in a single autoregressive pass, conditioned on egocentric video, 3D scene features, and three-point tracking, achieving state-of-the-art o...

  2. HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HCNQA supervises three intermediate reasoning phases (BoI, OoI, OoT) derived from ScanQA annotations, improving EM@1 from 25.94 to 27.01 over 3D-VisTA.

  3. 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A learnable 3D scene graph representation that injects semantic relation embeddings between objects improves LLM performance on 3D grounding, captioning, and question answering.

  4. PerLA: Perceptive 3D Language Assistant

    cs.CV 2024-11 conditional novelty 5.0 of 10

    PerLA improves 3D question answering and dense captioning by combining Hilbert-curve partitioned local point-cloud details with global context through cross-attention and graph convolution.

  5. Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.

  6. 3D Scene Graph Guided Vision-Language Pre-training

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A scene-graph-guided contrastive and masked-modality pre-training scheme improves performance on three 3D vision-language benchmarks, but the pre-training uses the same dataset as downstream fine-tuning.

Pith tools