AlignMamba fuses audio, video, and language by matching tokens to a language anchor and enforcing distribution similarity, reporting small accuracy gains with large efficiency gains on MOSI and MOSEI.
On the parameterization and initialization of diagonal state space models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
AlignMamba fuses audio, video, and language by matching tokens to a language anchor and enforcing distribution similarity, reporting small accuracy gains with large efficiency gains on MOSI and MOSEI.