Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not replace, attribution methods.
Multimodal speech emotion recognition using cross attention with aligned audio and text
1 Pith paper cite this work, alongside 24 external citations. Polarity classification is still indexing.
1
Pith paper citing it
24
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cross-Attention is Half Explanation in Speech-to-Text Models
Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not replace, attribution methods.