Pith. sign in

REVIEW 1 cited by

Detecting expressions with multimodal transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.00063 v1 pith:FSQBATQU submitted 2020-11-30 eess.AS cs.SDeess.IV

classification eess.AScs.SDeess.IV
keywords audio-visualexpressionlayersaff-wild2algorithmsarchitecturebaselinebetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing machine learning algorithms to understand person-to-person engagement can result in natural user experiences for communal devices such as Amazon Alexa. Among other cues such as voice activity and gaze, a person's audio-visual expression that includes tone of the voice and facial expression serves as an implicit signal of engagement between parties in a dialog. This study investigates deep-learning algorithms for audio-visual detection of user's expression. We first implement an audio-visual baseline model with recurrent layers that shows competitive results compared to current state of the art. Next, we propose the transformer architecture with encoder layers that better integrate audio-visual features for expressions tracking. Performance on the Aff-Wild2 database shows that the proposed methods perform better than baseline architecture with recurrent layers with absolute gains approximately 2% for arousal and valence descriptors. Further, multimodal architectures show significant improvements over models trained on single modalities with gains of up to 3.6%. Ablation studies show the significance of the visual modality for the expression detection on the Aff-Wild2 database.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects

    cs.DL 2025-06 conditional novelty 4.0 of 10

    A review arguing LLM-based quality scores could surpass bibliometrics in accuracy and coverage, but with unknown biases and gaming risks that currently block real-world use.

Pith tools