Pith. sign in

REVIEW 2 cited by

What to align in multimodal contrastive learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07402 v2 pith:D4YB7HSD submitted 2024-09-11 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords multimodalcomminformationmodalitieslearningaligncontrastivedifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive MultiModal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on the seven multimodal benchmarks. Code is available at https://github.com/Duplums/CoMM

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

    cs.CV 2026-01 unverdicted novelty 5.0 of 10

    A regularization method enforces diverse intra-modal embeddings and bounded inter-modal drift to improve both multimodal fusion and unimodal robustness.

  2. Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach

    cs.CV 2025-07 reject novelty 4.0 of 10

    A dual-path ResNet/DenseNet framework with multi-stage contrastive learning and confidence-driven gradient modulation is presented for multimodal human activity recognition, with reported improvements on four public datasets.

Pith tools