REVIEW 1 cited by
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With wearing masks becoming a new cultural norm, facial expression recognition (FER) while taking masks into account has become a significant challenge. In this paper, we propose a unified multi-branch vision transformer for facial expression recognition and mask wearing classification tasks. Our approach extracts shared features for both tasks using a dual-branch architecture that obtains multi-scale feature representations. Furthermore, we propose a cross-task fusion phase that processes tokens for each task with separate branches, while exchanging information using a cross attention module. Our proposed framework reduces the overall complexity compared with using separate networks for both tasks by the simple yet effective cross-task fusion phase. Extensive experiments demonstrate that our proposed model performs better than or on par with different state-of-the-art methods on both facial expression recognition and facial mask wearing classification task.
Forward citations
Cited by 1 Pith paper
-
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
OpenFace 3.0 shows a single lightweight multi-task model can handle four facial behavior tasks at speeds competitive with specialized toolkits, though the 'rivals SOTA' claim is not equally supported across all four tasks.
Discussion (0). Sign in to comment.