REVIEW 10 cited by
Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Affect recognition based on subjects' facial expressions has been a topic of major research in the attempt to generate machines that can understand the way subjects feel, act and react. In the past, due to the unavailability of large amounts of data captured in real-life situations, research has mainly focused on controlled environments. However, recently, social media and platforms have been widely used. Moreover, deep learning has emerged as a means to solve visual analysis and recognition problems. This paper exploits these advances and presents significant contributions for affect analysis and recognition in-the-wild. Affect analysis and recognition can be seen as a dual knowledge generation problem, involving: i) creation of new, large and rich in-the-wild databases and ii) design and training of novel deep neural architectures that are able to analyse affect over these databases and to successfully generalise their performance on other datasets. The paper focuses on large in-the-wild databases, i.e., Aff-Wild and Aff-Wild2 and presents the design of two classes of deep neural networks trained with these databases. The first class refers to uni-task affect recognition, focusing on prediction of the valence and arousal dimensional variables. The second class refers to estimation of all main behavior tasks, i.e. valence-arousal prediction; categorical emotion classification in seven basic facial expressions; facial Action Unit detection. A novel multi-task and holistic framework is presented which is able to jointly learn and effectively generalize and perform affect recognition over all existing in-the-wild databases. Large experimental studies illustrate the achieved performance improvement over the existing state-of-the-art in affect recognition.
Forward citations
Cited by 10 Pith papers
-
DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios
The authors propose a new frame-level annotated violence detection dataset, DVD, with 500 videos and rich metadata, but it is not yet available and lacks validation experiments.
-
AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
Adding a conditional rectified-flow head to a DINOv3 multi-task affect model improves valence-arousal CCC by +0.058 when the backbone is frozen and, with fine-tuning and validation-tuned calibration, reaches P_MTL=1.1...
-
Causal Supervision of Attention for Affective Behaviour Analysis
Causal supervision plus K-V independence and SwiGLU attention pooling yields a multi-task P-score of 1.2214 on s-Aff-Wild2 validation for valence-arousal, expression, and action-unit prediction.
-
Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition
Parameter-isolated LoRA experts on one face backbone stay decorrelated (0.91 vs 0.98 for full fine-tuning) and improve an ABAW affect ensemble from 1.6669 to 1.6949–1.7259 validation score.
-
A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition
A shared variational affect latent that marginalizes missing labels lifts rare expression and action-unit recognition on s-Aff-Wild2 beyond masked-loss training.
-
Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation
Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.
-
Team RAS in 9th ABAW Competition: Multimodal Compound Expression Recognition Approach
A six-modality zero-shot pipeline with CLIP, Qwen-VL, WavLM, Mamba, and new fusion/aggregation modules reports F1 scores of 46.95 (AffWild2), 49.02 (AFEW), and 34.85 (C-EXPR-DB) without target-domain fine-tuning.
-
AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis
A multi-task affective system combining frozen AffectNet backbones, LoRA-adapted MAE for action units, temporal heads, fusion, and ensembling attains P=1.7302 on the s-Aff-Wild2 validation split.
-
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
A GRU-based cross-attention fusion of frozen vision-language encoders is claimed to achieve strong results on DVD and Aff-Wild2, but the supporting experiments are missing from the paper.
-
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
TAGF adds a BiLSTM-based gate that reweights recursive cross-attention outputs for valence-arousal prediction, with results slightly below several existing methods on Aff-Wild2.
Discussion (0). Sign in to comment.