Pith. sign in

REVIEW 5 cited by

What to align in multimodal contrastive learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07402 v2 pith:D4YB7HSD submitted 2024-09-11 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords multimodalcomminformationmodalitieslearningaligncontrastivedifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive MultiModal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on the seven multimodal benchmarks. Code is available at https://github.com/Duplums/CoMM

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Information-Theoretic Decomposition for Multimodal Interaction Learning

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    DMIL is a multimodal learning framework that decomposes sample-specific interactions into redundant, unique, and synergistic components via variational architecture and uses them for adaptive fine-tuning.

  2. The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment

    cs.CV 2025-11 unverdicted novelty 6.0 of 10

    Contrastive Fusion (ConFu) adds a fused-modality contrastive term to jointly align individual modalities and their combinations, enabling capture of higher-order dependencies like XOR relations while preserving pairwi...

  3. Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

    cs.CV 2026-01 unverdicted novelty 5.0 of 10

    A regularization method enforces diverse intra-modal embeddings and bounded inter-modal drift to improve both multimodal fusion and unimodal robustness.

  4. Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Adding a dispersion loss plus a bounded cross-modal drift penalty to intermediate embeddings improves unimodal and multimodal accuracy across audio-visual, image-text, and RF benchmarks.

  5. Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach

    cs.CV 2025-07 reject novelty 4.0 of 10

    A dual-path ResNet/DenseNet framework with multi-stage contrastive learning and confidence-driven gradient modulation is presented for multimodal human activity recognition, with reported improvements on four public datasets.

Pith tools