CM3AE combines masked autoencoding, cross-modal fusion reconstruction, and contrastive learning to pre-train a ViT on 2.53 million RGB-event pairs.
Weakly misalignment-free adaptive feature alignment for uavs- based multimodal object detection
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
CM3AE combines masked autoencoding, cross-modal fusion reconstruction, and contrastive learning to pre-train a ViT on 2.53 million RGB-event pairs.