Pith. sign in

REVIEW 3 cited by

MSAF: Multimodal Split Attention Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.07175 v2 pith:FF5YDVY7 submitted 2020-12-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords multimodalfusionmsafmoduleattentionfeaturesnetworksacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal learning mimics the reasoning process of the human multi-sensory system, which is used to perceive the surrounding world. While making a prediction, the human brain tends to relate crucial cues from multiple sources of information. In this work, we propose a novel multimodal fusion module that learns to emphasize more contributive features across all modalities. Specifically, the proposed Multimodal Split Attention Fusion (MSAF) module splits each modality into channel-wise equal feature blocks and creates a joint representation that is used to generate soft attention for each channel across the feature blocks. Further, the MSAF module is designed to be compatible with features of various spatial dimensions and sequence lengths, suitable for both CNNs and RNNs. Thus, MSAF can be easily added to fuse features of any unimodal networks and utilize existing pretrained unimodal model weights. To demonstrate the effectiveness of our fusion module, we design three multimodal networks with MSAF for emotion recognition, sentiment analysis, and action recognition tasks. Our approach achieves competitive results in each task and outperforms other application-specific networks and multimodal fusion benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PIDNet uses progressive implicit decoupling with iMambaWave and Group3M blocks to fuse multimodal cues for improved action quality assessment on gymnastics datasets.

  2. MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    MLCR organizes quality cues at intra-modal, cross-modal, and stage-wise levels to improve long-term multimodal action quality assessment, achieving top results on gymnastics datasets.

  3. Attention-Driven Multimodal Alignment for Long-term Action Quality Assessment

    cs.CV 2025-07 reject novelty 5.0 of 10

    LMAC-Net reports state-of-the-art Spearman correlations on the RG and Fis-V benchmarks by aligning attention centers across RGB, optical flow, and audio branches.

Pith tools