Pith. sign in

REVIEW 6 cited by

MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17730 v1 pith:H4AS6ZNF submitted 2024-05-28 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords multimodalunimodallearninggradientlossmmparetoobjectivesassistance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal learning methods with targeted unimodal learning objectives have exhibited their superior efficacy in alleviating the imbalanced multimodal learning problem. However, in this paper, we identify the previously ignored gradient conflict between multimodal and unimodal learning objectives, potentially misleading the unimodal encoder optimization. To well diminish these conflicts, we observe the discrepancy between multimodal loss and unimodal loss, where both gradient magnitude and covariance of the easier-to-learn multimodal loss are smaller than the unimodal one. With this property, we analyze Pareto integration under our multimodal scenario and propose MMPareto algorithm, which could ensure a final gradient with direction that is common to all learning objectives and enhanced magnitude to improve generalization, providing innocent unimodal assistance. Finally, experiments across multiple types of modalities and frameworks with dense cross-modal interaction indicate our superior and extendable method performance. Our method is also expected to facilitate multi-task cases with a clear discrepancy in task difficulty, demonstrating its ideal scalability. The source code and dataset are available at https://github.com/GeWu-Lab/MMPareto_ICML2024.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Pareto LoRA applies Pareto-optimal gradient integration to balance text and image objectives in LoRA-based fine-tuning of unified multimodal models, reporting up to 44.9% gains in image quality on the CoMM benchmark w...

  2. Boosting Multimodal Federated Learning via Chained Modality Optimization

    cs.DC 2026-06 unverdicted novelty 6.0 of 10

    FedMChain improves multimodal federated learning by chaining modality-wise optimization phases with error-compensated regularization and sparse sign-guided aggregation to mitigate modality competition and cut communic...

  3. PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Imbalanced multimodal learning that prioritizes the performance-dominant modality via unimodal ranking and asymmetric gradient modulation outperforms balanced approaches.

  4. Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    Random label bridge training aligns LLM parameters with vision tasks, and partial training of certain layers often suffices due to their foundational properties.

  5. Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

    cs.CV 2026-01 unverdicted novelty 5.0 of 10

    A regularization method enforces diverse intra-modal embeddings and bounded inter-modal drift to improve both multimodal fusion and unimodal robustness.

  6. Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Adding a dispersion loss plus a bounded cross-modal drift penalty to intermediate embeddings improves unimodal and multimodal accuracy across audio-visual, image-text, and RF benchmarks.

Pith tools