A shared-compression MLP fusion with 32-head ensemble learning achieves MSE 0.1824, the top score in the AVI 2025 interview assessment track.
Adaptive Fusion Techniques for Multimodal Data
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from different modalities effectively. Instead of defining a deterministic fusion operation, such as concatenation, for the network, we let the network decide "how" to combine a given set of multimodal features more effectively. We propose two networks: 1) Auto-Fusion, which learns to compress information from different modalities while preserving the context, and 2) GAN-Fusion, which regularizes the learned latent space given context from complementing modalities. A quantitative evaluation on the tasks of multimodal machine translation and emotion recognition suggests that our lightweight, adaptive networks can better model context from other modalities than existing methods, many of which employ massive transformer-based networks.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment
A shared-compression MLP fusion with 32-head ensemble learning achieves MSE 0.1824, the top score in the AVI 2025 interview assessment track.