FaceInsight, an MLLM with segmentation inputs, correlation priors, and logic rules, reports higher face-attribute, age/gender/race, and expression accuracy than nine general MLLMs across six benchmarks.
Task-adaptive Q-Face
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Although face analysis has achieved remarkable improvements in the past few years, designing a multi-task face analysis model is still challenging. Most face analysis tasks are studied as separate problems and do not benefit from the synergy among related tasks. In this work, we propose a novel task-adaptive multi-task face analysis method named as Q-Face, which simultaneously performs multiple face analysis tasks with a unified model. We fuse the features from multiple layers of a large-scale pre-trained model so that the whole model can use both local and global facial information to support multiple tasks. Furthermore, we design a task-adaptive module that performs cross-attention between a set of query vectors and the fused multi-stage features and finally adaptively extracts desired features for each face analysis task. Extensive experiments show that our method can perform multiple tasks simultaneously and achieves state-of-the-art performance on face expression recognition, action unit detection, face attribute analysis, age estimation, and face pose estimation. Compared to conventional methods, our method opens up new possibilities for multi-task face analysis and shows the potential for both accuracy and efficiency.
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FaceInsight: A Multimodal Large Language Model for Face Perception
FaceInsight, an MLLM with segmentation inputs, correlation priors, and logic rules, reports higher face-attribute, age/gender/race, and expression accuracy than nine general MLLMs across six benchmarks.