Audit of depression detection benchmarks finds that official splits yield unstable model rankings, zero-shot transfer across datasets is weak, and text models but not audio models improve on symptom-dense interview segments.
Global, regional and national burden of depressive disorders and attributable risk factors, from 1990 to 2021: results from the 2021 Global Burden of Disease study,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks
Audit of depression detection benchmarks finds that official splits yield unstable model rankings, zero-shot transfer across datasets is weak, and text models but not audio models improve on symptom-dense interview segments.