Pith. sign in

REVIEW 3 cited by

AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.23450 v1 pith:5P2O44RU submitted 2025-03-30 cs.CV

classification cs.CV
keywords detectionfacialperformancetrainingactionadditionallyau-tttchallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Facial Action Units (AUs) detection is a cornerstone of objective facial expression analysis and a critical focus in affective computing. Despite its importance, AU detection faces significant challenges, such as the high cost of AU annotation and the limited availability of datasets. These constraints often lead to overfitting in existing methods, resulting in substantial performance degradation when applied across diverse datasets. Addressing these issues is essential for improving the reliability and generalizability of AU detection methods. Moreover, many current approaches leverage Transformers for their effectiveness in long-context modeling, but they are hindered by the quadratic complexity of self-attention. Recently, Test-Time Training (TTT) layers have emerged as a promising solution for long-sequence modeling. Additionally, TTT applies self-supervised learning for iterative updates during both training and inference, offering a potential pathway to mitigate the generalization challenges inherent in AU detection tasks. In this paper, we propose a novel vision backbone tailored for AU detection, incorporating bidirectional TTT blocks, named AU-TTT. Our approach introduces TTT Linear to the AU detection task and optimizes image scanning mechanisms for enhanced performance. Additionally, we design an AU-specific Region of Interest (RoI) scanning mechanism to capture fine-grained facial features critical for AU detection. Experimental results demonstrate that our method achieves competitive performance in both within-domain and cross-domain scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    CoEmoGen generates emotionally faithful images from emotion categories using MLLM-crafted captions and a hierarchical LoRA module, validated on a new EmoArt dataset.

  2. AU-LLM: Micro-Expression Action Unit Detection via Enhanced LLM-Based Feature Fusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AU-LLM uses a large language model and a feature-fusion projector to detect micro-expression action units, reporting the best average F1 scores on CASME II and SAMM.

  3. Distribution-Specific Learning for Joint Salient and Camouflaged Object Detection

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A shared network with about 2,000 decoder-specific parameters and a saliency-filtered, size-balanced training set reaches state-of-the-art accuracy on both salient and camouflaged object detection simultaneously.

Pith tools