REVIEW 5 cited by
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents a novel visual-language model called DFER-CLIP, which is based on the CLIP model and designed for in-the-wild Dynamic Facial Expression Recognition (DFER). Specifically, the proposed DFER-CLIP consists of a visual part and a textual part. For the visual part, based on the CLIP image encoder, a temporal model consisting of several Transformer encoders is introduced for extracting temporal facial expression features, and the final feature embedding is obtained as a learnable "class" token. For the textual part, we use as inputs textual descriptions of the facial behaviour that is related to the classes (facial expressions) that we are interested in recognising -- those descriptions are generated using large language models, like ChatGPT. This, in contrast to works that use only the class names and more accurately captures the relationship between them. Alongside the textual description, we introduce a learnable token which helps the model learn relevant context information for each expression during training. Extensive experiments demonstrate the effectiveness of the proposed method and show that our DFER-CLIP also achieves state-of-the-art results compared with the current supervised DFER methods on the DFEW, FERV39k, and MAFW benchmarks. Code is publicly available at https://github.com/zengqunzhao/DFER-CLIP.
Forward citations
Cited by 5 Pith papers
-
Reweighting Framewise Attention in Video Transformers for Facial Expression Understanding
MiRA is a parameter-free frame-marginal attention redistribution technique for ViT video models that improves sensitivity to localized facial cues on FER benchmarks.
-
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
VEGA aligns multimodal emotion features with CLIP-derived visual emotion prototypes and reports SOTA on IEMOCAP and MELD.
-
Trade-offs in Image Generation: How Do Different Dimensions Interact?
A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.
-
From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition
GRACE pairs motion-weighted video tokens with AI-refined emotion text tokens using optimal transport, reporting new UAR and WAR records on DFEW, FERV39k, and MAFW.
-
Multimodal Prompt Alignment for Facial Expression Recognition
A frozen-CLIP facial expression recognition framework with LLM-guided prompts, prototype regularization, and sparse global-local alignment claims new state-of-the-art accuracy across RAF-DB, FERPlus, and AffectNet.
Discussion (0). Sign in to comment.