Generated multilingual ASR transcripts fused by cascaded cross-modal transformers raise audio sentiment accuracy, and the multimodal knowledge distills into a stronger audio-only WavLM student with no inference cost.
Optimized Sentiment Analysis in Tagalog Speech Using PCA and BRNN on Prosodic Suprasegmental and MFCC Features, in: Proceedings of ICTC, pp
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
Generated multilingual ASR transcripts fused by cascaded cross-modal transformers raise audio sentiment accuracy, and the multimodal knowledge distills into a stronger audio-only WavLM student with no inference cost.