Pith. sign in

REVIEW 9 cited by

Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04855 v1 pith:BWFS4AKR submitted 2019-09-25 cs.CV cs.HCcs.LGeess.IV

classification cs.CVcs.HCcs.LGeess.IV
keywords aff-wild2recognitionavailabledatabaseemotionnetworksactiondatabases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Affective computing has been largely limited in terms of available data resources. The need to collect and annotate diverse in-the-wild datasets has become apparent with the rise of deep learning models, as the default approach to address any computer vision task. Some in-the-wild databases have been recently proposed. However: i) their size is small, ii) they are not audiovisual, iii) only a small part is manually annotated, iv) they contain a small number of subjects, or v) they are not annotated for all main behavior tasks (valence-arousal estimation, action unit detection and basic expression classification). To address these, we substantially extend the largest available in-the-wild database (Aff-Wild) to study continuous emotions such as valence and arousal. Furthermore, we annotate parts of the database with basic expressions and action units. As a consequence, for the first time, this allows the joint study of all three types of behavior states. We call this database Aff-Wild2. We conduct extensive experiments with CNN and CNN-RNN architectures that use visual and audio modalities; these networks are trained on Aff-Wild2 and their performance is then evaluated on 10 publicly available emotion databases. We show that the networks achieve state-of-the-art performance for the emotion recognition tasks. Additionally, we adapt the ArcFace loss function in the emotion recognition context and use it for training two new networks on Aff-Wild2 and then re-train them in a variety of diverse expression recognition databases. The networks are shown to improve the existing state-of-the-art. The database, emotion recognition models and source code are available at http://ibug.doc.ic.ac.uk/resources/aff-wild2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors propose a new frame-level annotated violence detection dataset, DVD, with 500 videos and rich metadata, but it is not yet available and lacks validation experiments.

  2. AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adding a conditional rectified-flow head to a DINOv3 multi-task affect model improves valence-arousal CCC by +0.058 when the backbone is frozen and, with fine-tuning and validation-tuned calibration, reaches P_MTL=1.1...

  3. Causal Supervision of Attention for Affective Behaviour Analysis

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Causal supervision plus K-V independence and SwiGLU attention pooling yields a multi-task P-score of 1.2214 on s-Aff-Wild2 validation for valence-arousal, expression, and action-unit prediction.

  4. Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Parameter-isolated LoRA experts on one face backbone stay decorrelated (0.91 vs 0.98 for full fine-tuning) and improve an ABAW affect ensemble from 1.6669 to 1.6949–1.7259 validation score.

  5. A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition

    cs.CV 2026-07 accept novelty 5.0 of 10

    A shared variational affect latent that marginalizes missing labels lifts rare expression and action-unit recognition on s-Aff-Wild2 beyond masked-loss training.

  6. Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

    cs.CV 2026-03 reject novelty 5.0 of 10

    Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.

  7. AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A multi-task affective system combining frozen AffectNet backbones, LoRA-adapted MAE for action units, temporal heads, fusion, and ensembling attains P=1.7302 on the s-Aff-Wild2 validation split.

  8. Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

    cs.CV 2025-07 reject novelty 4.0 of 10

    A GRU-based cross-attention fusion of frozen vision-language encoders is claimed to achieve strong results on DVD and Aff-Wild2, but the supporting experiments are missing from the paper.

  9. TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

    cs.MM 2025-07 reject novelty 4.0 of 10

    TAGF adds a BiLSTM-based gate that reweights recursive cross-attention outputs for valence-arousal prediction, with results slightly below several existing methods on Aff-Wild2.

Pith tools