Pith. sign in

REVIEW

Spatial and Temporal Networks for Facial Expression Recognition in the Wild Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.05160 v1 pith:CQSPQFGT submitted 2021-07-12 cs.CV eess.IV

classification cs.CVeess.IV
keywords expressionmodelrecognitioncnn-transformerensemblefacialin-the-wildnetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The paper describes our proposed methodology for the seven basic expression classification track of Affective Behavior Analysis in-the-wild (ABAW) Competition 2021. In this task, facial expression recognition (FER) methods aim to classify the correct expression category from a diverse background, but there are several challenges. First, to adapt the model to in-the-wild scenarios, we use the knowledge from pre-trained large-scale face recognition data. Second, we propose an ensemble model with a convolution neural network (CNN), a CNN-recurrent neural network (CNN-RNN), and a CNN-Transformer (CNN-Transformer), to incorporate both spatial and temporal information. Our ensemble model achieved F1 as 0.4133, accuracy as 0.6216 and final metric as 0.4821 on the validation set.

Discussion (0). Continue with ORCID to comment.

Pith tools