Pith. sign in

REVIEW 1 cited by

SHAPE: An Unified Approach to Evaluate the Contribution and Cooperation of Individual Modalities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.00302 v1 pith:KXBBAR7T submitted 2022-04-30 cs.LG

classification cs.LG
keywords modalitiesdifferentmulti-modalmodelscooperationscorescontributionmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As deep learning advances, there is an ever-growing demand for models capable of synthesizing information from multi-modal resources to address the complex tasks raised from real-life applications. Recently, many large multi-modal datasets have been collected, on which researchers actively explore different methods of fusing multi-modal information. However, little attention has been paid to quantifying the contribution of different modalities within the proposed models. In this paper, we propose the {\bf SH}apley v{\bf A}lue-based {\bf PE}rceptual (SHAPE) scores that measure the marginal contribution of individual modalities and the degree of cooperation across modalities. Using these scores, we systematically evaluate different fusion methods on different multi-modal datasets for different tasks. Our experiments suggest that for some tasks where different modalities are complementary, the multi-modal models still tend to use the dominant modality alone and ignore the cooperation across modalities. On the other hand, models learn to exploit cross-modal cooperation when different modalities are indispensable for the task. In this case, the scores indicate it is better to fuse different modalities at relatively early stages. We hope our scores can help improve the understanding of how the present multi-modal models operate on different modalities and encourage more sophisticated methods of integrating multiple modalities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asymmetric Reinforcing against Multi-modal Representation Bias

    cs.CV 2025-01 reject novelty 6.0 of 10

    ARM, a mutual-information-based asymmetric reinforcement method, narrows modality contribution gaps and reports improved accuracy on three multimodal classification datasets.

Pith tools