Pith. sign in

REVIEW 4 cited by

Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.12221 v1 pith:O7GCJ6XU submitted 2022-03-23 cs.LG

classification cs.LG
keywords multi-modaljointnetworktrainingcompetitionmodalitiesmodalitybeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network, which is counter-intuitive since multiple signals generally bring more information. This work provides a theoretical explanation for the emergence of such performance gap in neural networks for the prevalent joint training framework. Based on a simplified data distribution that captures the realistic property of multi-modal data, we prove that for the multi-modal late-fusion network with (smoothed) ReLU activation trained jointly by gradient descent, different modalities will compete with each other. The encoder networks will learn only a subset of modalities. We refer to this phenomenon as modality competition. The losing modalities, which fail to be discovered, are the origins where the sub-optimality of joint training comes from. Experimentally, we illustrate that modality competition matches the intrinsic behavior of late-fusion joint training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Information-Theoretic Decomposition for Multimodal Interaction Learning

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    DMIL is a multimodal learning framework that decomposes sample-specific interactions into redundant, unique, and synergistic components via variational architecture and uses them for adaptive fine-tuning.

  2. Boosting Multimodal Federated Learning via Chained Modality Optimization

    cs.DC 2026-06 unverdicted novelty 6.0 of 10

    FedMChain improves multimodal federated learning by chaining modality-wise optimization phases with error-compensated regularization and sparse sign-guided aggregation to mitigate modality competition and cut communic...

  3. Optimizing Deep Learning Photometric Redshifts for the Roman Space Telescope with HST/CANDELS

    astro-ph.IM 2026-02 unverdicted novelty 6.0 of 10

    PITA, a new semi-supervised deep learning algorithm, outperforms prior photo-z methods by using a triple-task loss on images, colors, and available redshifts to produce a smooth latent space.

  4. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

Pith tools