Audit of depression detection benchmarks finds that official splits yield unstable model rankings, zero-shot transfer across datasets is weak, and text models but not audio models improve on symptom-dense interview segments.
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
This study investigates whether speech-based depression detection models learn depression-related acoustic biomarkers or instead rely on speaker identity cues. Using the DAIC-WOZ dataset, we propose a data-splitting strategy that controls speaker overlap between training and test sets while keeping the training size constant, and evaluate three models of varying complexity. Results show that speaker overlap significantly boosts performance, whereas accuracy drops sharply on unseen speakers. Even with a Domain-Adversarial Neural Network, a substantial performance gap remains. These findings indicate that depression-related features extracted by current speech models are highly entangled with speaker identity. Conventional evaluation protocols may therefore overestimate generalization and clinical utility, highlighting the need for strictly speaker-independent evaluation.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Speech-based depression detection models primarily learn speaker identity rather than depression biomarkers, with performance dropping sharply on unseen speakers even under adversarial training.
Framework applies XAI feature selection to low-complexity ML models for interpretable, fair speech-based depression detection on DAIC-WOZ, claiming 82% accuracy as state-of-the-art.
citing papers explorer
-
A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks
Audit of depression detection benchmarks finds that official splits yield unstable model rankings, zero-shot transfer across datasets is weak, and text models but not audio models improve on symptom-dense interview segments.
-
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
Speech-based depression detection models primarily learn speaker identity rather than depression biomarkers, with performance dropping sharply on unseen speakers even under adversarial training.
-
A Fair and Transparent Framework for Speech-Based Depression Detection: Balancing Interpretability and Performance
Framework applies XAI feature selection to low-complexity ML models for interpretable, fair speech-based depression detection on DAIC-WOZ, claiming 82% accuracy as state-of-the-art.