Pith. sign in

REVIEW

Temporal aggregation of audio-visual modalities for emotion recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.04364 v1 pith:MNJ4ZZUL submitted 2020-07-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords emotionrecognitiontemporalaudio-visualcurrenthumaninformationinteraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotion recognition has a pivotal role in affective computing and in human-computer interaction. The current technological developments lead to increased possibilities of collecting data about the emotional state of a person. In general, human perception regarding the emotion transmitted by a subject is based on vocal and visual information collected in the first seconds of interaction with the subject. As a consequence, the integration of verbal (i.e., speech) and non-verbal (i.e., image) information seems to be the preferred choice in most of the current approaches towards emotion recognition. In this paper, we propose a multimodal fusion technique for emotion recognition based on combining audio-visual modalities from a temporal window with different temporal offsets for each modality. We show that our proposed method outperforms other methods from the literature and human accuracy rating. The experiments are conducted over the open-access multimodal dataset CREMA-D.

Discussion (0). Continue with ORCID to comment.

Pith tools