REVIEW 5 cited by
What to align in multimodal contrastive learning?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive MultiModal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on the seven multimodal benchmarks. Code is available at https://github.com/Duplums/CoMM
Forward citations
Cited by 5 Pith papers
-
Information-Theoretic Decomposition for Multimodal Interaction Learning
DMIL is a multimodal learning framework that decomposes sample-specific interactions into redundant, unique, and synergistic components via variational architecture and uses them for adaptive fine-tuning.
-
The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment
Contrastive Fusion (ConFu) adds a fused-modality contrastive term to jointly align individual modalities and their combinations, enabling capture of higher-order dependencies like XOR relations while preserving pairwi...
-
Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
A regularization method enforces diverse intra-modal embeddings and bounded inter-modal drift to improve both multimodal fusion and unimodal robustness.
-
Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
Adding a dispersion loss plus a bounded cross-modal drift penalty to intermediate embeddings improves unimodal and multimodal accuracy across audio-visual, image-text, and RF benchmarks.
-
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
A dual-path ResNet/DenseNet framework with multi-stage contrastive learning and confidence-driven gradient modulation is presented for multimodal human activity recognition, with reported improvements on four public datasets.
Discussion (0). Sign in to comment.