Pith. sign in

REVIEW 4 cited by

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08275 v4 pith:54TFT3NP submitted 2023-05-14 cs.CV

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

classification cs.CV
keywords multimodalulip-2languagedatasetsdescriptionsscalableclassificationlearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advancements in multimodal pre-training have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes, their 2D counterparts, and language descriptions. However, the methods used by existing frameworks to curate such multimodal data, in particular language descriptions for 3D shapes, are not scalable, and the collected language descriptions are not diverse. To address this, we introduce ULIP-2, a simple yet effective tri-modal pre-training framework that leverages large multimodal models to automatically generate holistic language descriptions for 3D shapes. It only needs 3D data as input, eliminating the need for any manual 3D annotations, and is therefore scalable to large datasets. ULIP-2 is also equipped with scaled-up backbones for better multimodal representation learning. We conduct experiments on two large-scale 3D datasets, Objaverse and ShapeNet, and augment them with tri-modal datasets of 3D point clouds, images, and language for training ULIP-2. Experiments show that ULIP-2 demonstrates substantial benefits in three downstream tasks: zero-shot 3D classification, standard 3D classification with fine-tuning, and 3D captioning (3D-to-language generation). It achieves a new SOTA of 50.6% (top-1) on Objaverse-LVIS and 84.7% (top-1) on ModelNet40 in zero-shot classification. In the ScanObjectNN benchmark for standard fine-tuning, ULIP-2 reaches an overall accuracy of 91.5% with a compact model of only 1.4 million parameters. ULIP-2 sheds light on a new paradigm for scalable multimodal 3D representation learning without human annotations and shows significant improvements over existing baselines. The code and datasets are released at https://github.com/salesforce/ULIP.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

    cs.CV 2026-05 conditional novelty 6.0

    Adaptive saliency-guided supervoxel tokenization cuts 3D AR token length to 12.8% of uniform voxels while claiming SOTA quality and ~10× speedup on Trellis-500K.

  2. PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

    cs.CV 2026-01 conditional novelty 6.0

    A feed-forward 3D encoder aligning patch-level point-cloud features with part-name text embeddings achieves state-of-the-art zero-shot 3D part segmentation, surpassing multi-view rendering pipelines by large margins o...

  3. Transolver: A Fast Transformer Solver for PDEs on General Geometries

    cs.LG 2024-02 conditional novelty 6.0

    Transolver learns intrinsic physical states from discretized meshes by adaptively splitting domains into flexible learnable slices and computing attention over physics-aware tokens, achieving state-of-the-art PDE solv...

  4. SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    SuperVoxelGPT creates shape-adaptive, deterministically ordered supervoxel tokens via saliency-guided CVT, cutting sequence length to 12.8% of uniform voxels while claiming SOTA quality and 10x speedup on Trellis-500K.