Pith. sign in

REVIEW 5 cited by

Multi-modal Attention for Speech Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.04107 v1 pith:43QWM66N submitted 2020-09-09 eess.AS cs.MMcs.SDeess.IV

classification eess.AScs.MMcs.SDeess.IV
keywords speechemotionattentionrecognitionclstm-mmacuesfusionmulti-modal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotion represents an essential aspect of human speech that is manifested in speech prosody. Speech, visual, and textual cues are complementary in human communication. In this paper, we study a hybrid fusion method, referred to as multi-modal attention network (MMAN) to make use of visual and textual cues in speech emotion recognition. We propose a novel multi-modal attention mechanism, cLSTM-MMA, which facilitates the attention across three modalities and selectively fuse the information. cLSTM-MMA is fused with other uni-modal sub-networks in the late fusion. The experiments show that speech emotion recognition benefits significantly from visual and textual cues, and the proposed cLSTM-MMA alone is as competitive as other fusion methods in terms of accuracy, but with a much more compact network structure. The proposed hybrid network MMAN achieves state-of-the-art performance on IEMOCAP database for emotion recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Image Editing As Programs with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.

  2. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.

  3. WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A text-to-image pipeline that lets users restyle individual characters or regions of artistic typography and iteratively refine them with region-specific prompts.

  4. RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.

  5. Learning Annotation Consensus for Continuous Emotion Recognition

    cs.HC 2025-05 conditional novelty 4.0 of 10

    A consensus network over multiple annotator labels improves continuous emotion prediction on RECOLA, but the claimed COGNIMUSE results are absent from the paper.

Pith tools