Pith. sign in

REVIEW 2 cited by

PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.08129 v1 pith:YU5QPYMV submitted 2023-03-14 cs.CV cs.AI

classification cs.CVcs.AI
keywords modalitiespimaeautoencoderscloudcross-modaldetectorsimageimprove
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked Autoencoders learn strong visual representations and achieve state-of-the-art results in several independent modalities, yet very few works have addressed their capabilities in multi-modality settings. In this work, we focus on point cloud and RGB image data, two modalities that are often presented together in the real world, and explore their meaningful interactions. To improve upon the cross-modal synergy in existing works, we propose PiMAE, a self-supervised pre-training framework that promotes 3D and 2D interaction through three aspects. Specifically, we first notice the importance of masking strategies between the two sources and utilize a projection module to complementarily align the mask and visible tokens of the two modalities. Then, we utilize a well-crafted two-branch MAE pipeline with a novel shared decoder to promote cross-modality interaction in the mask tokens. Finally, we design a unique cross-modal reconstruction module to enhance representation learning for both modalities. Through extensive experiments performed on large-scale RGB-D scene understanding benchmarks (SUN RGB-D and ScannetV2), we discover it is nontrivial to interactively learn point-image features, where we greatly improve multiple 3D detectors, 2D detectors, and few-shot classifiers by 2.9%, 6.7%, and 2.4%, respectively. Code is available at https://github.com/BLVLab/PiMAE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FOLIAGE: Towards Physical Intelligence World Models Via Unbounded Surface Evolution

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FOLIAGE combines image, point-cloud, and mesh encoders with an action-conditioned latent predictor to forecast accretive surface growth, outperforming baselines on the new synthetic SURF-BENCH benchmark.

  2. IndoorBEV: Joint Detection and Footprint Completion of Objects via Mask-based Prediction in Indoor Scenarios for Bird's-Eye View Perception

    cs.RO 2025-07 conditional novelty 4.0 of 10

    IndoorBEV uses a query-based transformer decoder on a bird's-eye view lidar grid to jointly detect objects and predict footprint masks in indoor scenes.

Pith tools