Pith. sign in

REVIEW 1 cited by

Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15676 v2 pith:WGQGDQNG submitted 2023-10-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords multi-modalrecentcomprehensiveespeciallyintelligencemethodsseveralsurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal 3D Intelligence has gained considerable attention due to its wide applications in autonomous driving and world simulation, etc. Compared to conventional single-modal 3D understanding, introducing an additional modality not only elevates the richness and precision of scene interpretation but also provides a foundation for higher-level physical world interaction. This becomes especially crucial in varied and challenging environments where solely relying on 3D data might be inadequate. While there has been a surge in the development of multi-modal 3D methods over the past six years, especially those integrating multi-camera images (3D+2D) and textual descriptions (3D+language), a comprehensive and in-depth review is notably absent. In this paper, we present a systematic survey of recent progress to bridge this gap. We begin by briefly summarizing the unique challenges among various 3D multi-modal tasks. After that, we present a novel taxonomy that delivers a thorough categorization of existing methods according to modalities and tasks, exploring their respective strengths and limitations. Furthermore, comparative results of recent approaches on several benchmark datasets, together with insightful analysis, are offered. Finally, we discuss the unresolved issues and provide several potential avenues for future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning from Silence and Noise for Visual Sound Source Localization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Adding silence and Gaussian noise as negative training pairs improves self-supervised visual sound source localization, and the authors provide IS3+ and a separability metric.

Pith tools