Pith. sign in

REVIEW 3 cited by

Adaptive Fusion Techniques for Multimodal Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03821 v2 pith:LSR4PQC6 submitted 2019-11-10 cs.CL cs.CVcs.LGeess.AS

classification cs.CLcs.CVcs.LGeess.AS
keywords modalitiescontextfusionmultimodaladaptivedatanetworksdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from different modalities effectively. Instead of defining a deterministic fusion operation, such as concatenation, for the network, we let the network decide "how" to combine a given set of multimodal features more effectively. We propose two networks: 1) Auto-Fusion, which learns to compress information from different modalities while preserving the context, and 2) GAN-Fusion, which regularizes the learned latent space given context from complementing modalities. A quantitative evaluation on the tasks of multimodal machine translation and emotion recognition suggests that our lightweight, adaptive networks can better model context from other modalities than existing methods, many of which employ massive transformer-based networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A shared-compression MLP fusion with 32-head ensemble learning achieves MSE 0.1824, the top score in the AVI 2025 interview assessment track.

  2. Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

    cs.MM 2025-07 conditional novelty 4.0 of 10

    Sync-TVA reports modest accuracy and weighted-F1 improvements over prior graph-based models on MELD and IEMOCAP, using modality-specific enhancement and cross-modal graph fusion.

  3. ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection

    cs.MM 2025-05 conditional novelty 4.0 of 10

    ISMAF reports state-of-the-art rumor detection accuracy on Weibo and PHEME by aligning text-image intrinsic features with social graph features and fusing them adaptively.

Pith tools