Pith. sign in

REVIEW 9 cited by

A Survey on Facial Expression Recognition of Static and Dynamic Emotions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15777 v1 pith:MVBJCKWR submitted 2024-08-28 cs.CV

classification cs.CV
keywords expressionchallengesdynamicstaticanalyzedevelopmentdferemotions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Facial expression recognition (FER) aims to analyze emotional states from static images and dynamic sequences, which is pivotal in enhancing anthropomorphic communication among humans, robots, and digital avatars by leveraging AI technologies. As the FER field evolves from controlled laboratory environments to more complex in-the-wild scenarios, advanced methods have been rapidly developed and new challenges and apporaches are encounted, which are not well addressed in existing reviews of FER. This paper offers a comprehensive survey of both image-based static FER (SFER) and video-based dynamic FER (DFER) methods, analyzing from model-oriented development to challenge-focused categorization. We begin with a critical comparison of recent reviews, an introduction to common datasets and evaluation criteria, and an in-depth workflow on FER to establish a robust research foundation. We then systematically review representative approaches addressing eight main challenges in SFER (such as expression disturbance, uncertainties, compound emotions, and cross-domain inconsistency) as well as seven main challenges in DFER (such as key frame sampling, expression intensity variations, and cross-modal alignment). Additionally, we analyze recent advancements, benchmark performances, major applications, and ethical considerations. Finally, we propose five promising future directions and development trends to guide ongoing research. The project page for this paper can be found at https://github.com/wangyanckxx/SurveyFER.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change

    cs.CV 2025-05 accept novelty 8.0 of 10

    Introduces the BAH dataset with 1,427 annotated videos for multimodal recognition of ambivalence/hesitancy in digital behavior change contexts.

  2. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    ARGen generates high-fidelity dynamic facial expression videos using affective semantic injection and adaptive reinforcement diffusion to improve emotion recognition models facing data scarcity and long-tail distributions.

  3. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

    cs.CV 2026-04 conditional novelty 6.0 of 10

    ARGen uses AU-guided prompts and a reinforcement-learned diffusion strategy to synthesize scarce-class facial expression videos that improve dynamic emotion recognition.

  4. FIELDS: Face reconstruction with accurate Inference of Expression using Learning with Direct Supervision

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Adding direct FLAME expression supervision from 3D scans plus an emotion-recognition head improves AffectNet facial-expression prediction from reconstructed faces without hurting geometric accuracy.

  5. Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    DuSE is a new dual-stream model for dynamic facial expression recognition that explicitly models cognitive priming and conceptual knowledge integration to reach state-of-the-art accuracy on in-the-wild benchmarks.

  6. Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    Standard deep multimodal models and LLM zero-shot inference achieve only limited performance on video ambivalence/hesitancy recognition for digital health personalization.

  7. Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    Multimodal deep learning for ambivalence/hesitancy recognition in videos yields limited results on the BAH dataset, highlighting the need for improved spatio-temporal and cross-modal fusion methods.

  8. Facial Expression Recognition in the Deep Learning Era: A Systematic Multi-Criteria Review of Methods, Models, Datasets, Performance, Challenges, and Future Research Directions

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    This survey organizes deep learning FER literature into five evolutionary phases and a seven-criteria taxonomy, compares datasets and performance, and outlines challenges.

  9. Evaluating multimodal emotion recognition in proactive conversational agents: A user study

    cs.HC 2026-04 unverdicted novelty 3.0 of 10

    A user study with 20 participants found that linguistic analysis is more reliable than facial recognition for detecting emotions in proactive AI agents due to users displaying neutral 'poker faces,' while also showing...

Pith tools