REVIEW 12 cited by
Aff-Wild2: Extending the Aff-Wild Database for Affect Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Aff-Wild2: Extending the Aff-Wild Database for Affect Recognition
read the original abstract
Automatic understanding of human affect using visual signals is a problem that has attracted significant interest over the past 20 years. However, human emotional states are quite complex. To appraise such states displayed in real-world settings, we need expressive emotional descriptors that are capable of capturing and describing this complexity. The circumplex model of affect, which is described in terms of valence (i.e., how positive or negative is an emotion) and arousal (i.e., power of the activation of the emotion), can be used for this purpose. Recent progress in the emotion recognition domain has been achieved through the development of deep neural architectures and the availability of very large training databases. To this end, Aff-Wild has been the first large-scale "in-the-wild" database, containing around 1,200,000 frames. In this paper, we build upon this database, extending it with 260 more subjects and 1,413,000 new video frames. We call the union of Aff-Wild with the additional data, Aff-Wild2. The videos are downloaded from Youtube and have large variations in pose, age, illumination conditions, ethnicity and profession. Both database-specific as well as cross-database experiments are performed in this paper, by utilizing the Aff-Wild2, along with the RECOLA database. The developed deep neural architectures are based on the joint training of state-of-the-art convolutional and recurrent neural networks with attention mechanism; thus exploiting both the invariant properties of convolutional features, while modeling temporal dynamics that arise in human behaviour via the recurrent layers. The obtained results show premise for utilization of the extended Aff-Wild, as well as of the developed deep neural architectures for visual analysis of human behaviour in terms of continuous emotion dimensions.
Forward citations
Cited by 12 Pith papers
-
Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition
Chehre introduces a new emoji-prompted video dataset with multi-annotator labels to benchmark models on dominant and distributional facial expression recognition tasks.
-
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
DAR provides 15k videos with 37k event-aligned viewer-emotion segments and causal chains, and DAR-R1 (SFT+GRPO) leads 10+ MLLMs on segmentation, emotion accuracy, and reasoning quality.
-
LaScA: Language-Conditioned Scalable Modelling of Affective Dynamics
A framework converts interpretable facial and acoustic features into language descriptions, feeds them to a pretrained LM for semantic embeddings, and uses those embeddings as priors to improve valence and arousal cha...
-
Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
Facial expression signals via prompt integration improve empathetic responsiveness in LLM-based tutoring systems.
-
Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony
Group-level EEG dynamic neural synchrony (CorrCA) preferentially tracks the rate of change of continuous arousal and shows valence-dependent structure across four datasets.
-
Causal Supervision of Attention for Affective Behaviour Analysis
Causal supervision plus K-V independence and SwiGLU attention pooling yields a multi-task P-score of 1.2214 on s-Aff-Wild2 validation for valence-arousal, expression, and action-unit prediction.
-
FIDAC: An Easy-to-use Pipeline to Extract and Interpret Interpersonal Distance From Video
FIDAC merges multi-model face detection, human coding, and plane benchmarking to extract interpersonal distance from ordinary 2D video.
-
AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis
A multi-task affective system combining frozen AffectNet backbones, LoRA-adapted MAE for action units, temporal heads, fusion, and ensembling attains P=1.7302 on the s-Aff-Wild2 validation split.
-
Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis
A task-adaptive system that mixes two frozen visual features with per-task fusion and temporal strategies scores 1.6341 on the ABAW11 validation set, beating its own shared multi-task baselines.
-
Causal Supervision of Attention for Affective Behaviour Analysis
Causal supervision of attention plus K-V cross-covariance regularization yields composite P=1.2214 on ABAW 11 MTL validation, up from the official baseline's 0.45.
-
Facial Expression Recognition in the Deep Learning Era: A Systematic Multi-Criteria Review of Methods, Models, Datasets, Performance, Challenges, and Future Research Directions
This survey organizes deep learning FER literature into five evolutionary phases and a seven-criteria taxonomy, compares datasets and performance, and outlines challenges.
-
Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
Feeding LLM tutors a text description or AUM-selected frame of a student's facial expression improves rated empathetic responsiveness across three backbones.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.