Pith. sign in

REVIEW 8 cited by

Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.03790 v1 pith:PD72TRTJ submitted 2021-05-08 cs.CV

classification cs.CV
keywords taskslearningapproachdetectionfaceknowledgematchingannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-Task Learning has emerged as a methodology in which multiple tasks are jointly learned by a shared learning algorithm, such as a DNN. MTL is based on the assumption that the tasks under consideration are related; therefore it exploits shared knowledge for improving performance on each individual task. Tasks are generally considered to be homogeneous, i.e., to refer to the same type of problem. Moreover, MTL is usually based on ground truth annotations with full, or partial overlap across tasks. In this work, we deal with heterogeneous MTL, simultaneously addressing detection, classification & regression problems. We explore task-relatedness as a means for co-training, in a weakly-supervised way, tasks that contain little, or even non-overlapping annotations. Task-relatedness is introduced in MTL, either explicitly through prior expert knowledge, or through data-driven studies. We propose a novel distribution matching approach, in which knowledge exchange is enabled between tasks, via matching of their predictions' distributions. Based on this approach, we build FaceBehaviorNet, the first framework for large-scale face analysis, by jointly learning all facial behavior tasks. We develop case studies for: i) continuous affect estimation, action unit detection, basic emotion recognition; ii) attribute detection, face identification. We illustrate that co-training via task relatedness alleviates negative transfer. Since FaceBehaviorNet learns features that encapsulate all aspects of facial behavior, we conduct zero-/few-shot learning to perform tasks beyond the ones that it has been trained for, such as compound emotion recognition. By conducting a very large experimental study, utilizing 10 databases, we illustrate that our approach outperforms, by large margins, the state-of-the-art in all tasks and in all databases, even in these which have not been used in its training.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adding a conditional rectified-flow head to a DINOv3 multi-task affect model improves valence-arousal CCC by +0.058 when the backbone is frozen and, with fine-tuning and validation-tuned calibration, reaches P_MTL=1.1...

  2. Causal Supervision of Attention for Affective Behaviour Analysis

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Causal supervision plus K-V independence and SwiGLU attention pooling yields a multi-task P-score of 1.2214 on s-Aff-Wild2 validation for valence-arousal, expression, and action-unit prediction.

  3. Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Parameter-isolated LoRA experts on one face backbone stay decorrelated (0.91 vs 0.98 for full fine-tuning) and improve an ABAW affect ensemble from 1.6669 to 1.6949–1.7259 validation score.

  4. Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

    cs.CV 2026-03 reject novelty 5.0 of 10

    Distance-aware soft prompts over a 3×3 emotion grid with CLIP text prototypes and audio-visual GRU fusion achieve CCC_mean 0.5361 on Aff-Wild2, beating only the paper's self-defined baselines.

  5. Team RAS in 9th ABAW Competition: Multimodal Compound Expression Recognition Approach

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A six-modality zero-shot pipeline with CLIP, Qwen-VL, WavLM, Mamba, and new fusion/aggregation modules reports F1 scores of 46.95 (AffWild2), 49.02 (AFEW), and 34.85 (C-EXPR-DB) without target-domain fine-tuning.

  6. AffectFuse: Cross-Task Feature Fusion with Temporal Modeling for Multi-Task Affective Behavior Analysis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A multi-task affective system combining frozen AffectNet backbones, LoRA-adapted MAE for action units, temporal heads, fusion, and ensembling attains P=1.7302 on the s-Aff-Wild2 validation split.

  7. Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding

    cs.CV 2025-07 reject novelty 4.0 of 10

    A GRU-based cross-attention fusion of frozen vision-language encoders is claimed to achieve strong results on DVD and Aff-Wild2, but the supporting experiments are missing from the paper.

  8. TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

    cs.MM 2025-07 reject novelty 4.0 of 10

    TAGF adds a BiLSTM-based gate that reweights recursive cross-attention outputs for valence-arousal prediction, with results slightly below several existing methods on Aff-Wild2.

Pith tools