Pith. sign in

REVIEW 4 major objections 6 minor 48 references

Multi-level Attention network using text, audio and video for Depression Prediction

T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Multi-level attention over text, audio, and video predicts PHQ-8 depression scores with RMSE 4.28 on E-DAIC development data, beating the AVEC 2019 baseline of 5.03.

desk verdict Useful but modest AVEC 2019 fusion work whose headline improvement is dev-only and arithmetically overstated; the text-only test numbers are the most solid part. read the letter →

arxiv 1909.01417 v1 pith:4XWJ5YND submitted 2019-09-03 cs.CV eess.AS

classification cs.CVeess.AS
keywords depressionpredictionmulti-modalfusionattentionnetworksbidirectionalLSTME-DAICPHQ-8affectivecomputingaudio-video-text
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Depression severity is usually scored through subjective questionnaires, and this paper argues that behavioural cues can do the scoring automatically. It proposes a multi-level attention network that fuses features from text transcripts, audio, and video, with attention applied both inside each modality and again when the modalities are combined. The network learns which features and which modalities matter for predicting PHQ-8 scores, and its best all-feature fusion configuration reaches an RMSE of 4.28 on the E-DAIC development set, against the AVEC 2019 baseline's 5.03. If the result holds on unseen data, automated screening from an interview recording becomes plausible without expensive clinical assessment.

What carries the argument

The central object is the multi-level attention fusion network: stacked bidirectional LSTMs over each modality produce hidden states, an attention layer over each stream builds a context vector, feedforward layers compress each context, and a second attention layer over the compressed modalities produces the final weighted representation that is multiplied into a stacked-BLSTM output and regressed to the PHQ-8 score. The intra-modality attention selects the informative timesteps or features within a stream, while the inter-modality attention learns the contribution ratios of text, audio and video. To stabilise training of the three-way fusion, the paper initialises a multiplicative 'nudge' vector with the reciprocal ratios of each single-modality RMSE, which biases the optimisation initially toward the text pathway; the final attention ratios after convergence are approximately 0.21 (video), 0.21 (audio) and 0.57 (text).

What would settle it

Run the proposed all-feature fusion model on the E-DAIC test partition and compare its RMSE with the AVEC 2019 baseline; if the test RMSE does not beat 5.03 by a comparable margin, the central improvement claim fails. A simpler check is arithmetic: (5.03 - 4.28)/5.03 is about 14.9%, which already differs from the reported 17.52% under the usual relative-error definition.

Watch

Extended reading notes

Core claim

The paper's central claim is that applying attention at multiple levels—once over the sequence outputs within each modality and again over the fused modality representations—lets a regression network predict depression severity better than the challenge baseline. The all-feature fusion model, built from text sentence embeddings, audio descriptors (MFCC, eGeMAPS, BoAW, deep densenet features) and video descriptors (pose, gaze, facial action units, bag-of-visual-words), achieves the best development-set RMSE of 4.28 compared with the baseline's 5.03, which the authors report as a 17.52% improvement. The learned attention ratios give the text modality roughly 57% of the weight, with audio and video near 21% each, and the authors interpret this as the network discovering that verbal content is the strongest marker. A text-only version of the model also performs well on the withheld test set, with a concordance correlation coefficient of 0.67, and the authors report that it outperforms the closest prior attention-based work by 8.95%.

Load-bearing premise

The paper assumes that the development partition's labels, which were used both to choose the best configurations and to initialize the fusion weights, are a reliable stand-in for the withheld test partition, so the reported improvement will survive on unseen data.

Editorial extensions

If this is right

  • The all-feature fusion network reaches RMSE 4.28 on the E-DAIC development set, beating the AVEC 2019 baseline's 5.03.
  • A text-only network scores 4.37 RMSE on dev and, on the challenge test set, MAE 4.02, RMSE 4.73 and CCC 0.67, which the authors report as stronger than the closest prior attention-based model.
  • Learned attention weights rank text first (about 0.57), with audio and video nearly equal (about 0.21 each), implying that verbal content dominates automated depression scoring in this corpus.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the development partition was used both to select the best configurations and to initialize the fusion nudge vector, the reported 17.52% improvement should be read as a dev-set estimate rather than proven generalization to the withheld test set.
  • The nudge initialization biases the network toward the text modality from the start, so the final attention ratios may partly reflect that initialization rather than an unbiased discovery of modality importance; ablating the nudge vector would clarify this.
  • The same multi-level attention architecture could be applied to other PHQ-based assessments or to anxiety and PTSD scores in the E-DAIC corpus, where similar behavioural markers are recorded.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a multi-level attention network that fuses audio, video, and text features to predict PHQ-8 depression severity scores on the E-DAIC corpus. The authors build per-modality regression models using BLSTM networks, then experiment with several fusion configurations that apply attention within and across modalities. Their best model, an all-feature fusion network, is reported to achieve RMSE 4.28 on the development partition, which they describe as outperforming the AVEC 2019 baseline (RMSE 5.03) by 17.52%. The text-only model is also evaluated on the withheld test partition, achieving CCC 0.67. The manuscript includes extensive ablations across individual features and fusion variants.

Significance. If the central claim were fully supported, the paper would contribute a reasonably investigated multimodal attention architecture for depression prediction, with a thorough ablation over audio, video, and text features. The text-only model's test-set CCC of 0.67 and the systematic per-feature comparisons against the AVEC 2019 baseline are genuine strengths. The proposed multi-level attention idea is plausible and the empirical exploration is fairly detailed. However, the headline improvement claim is numerically inconsistent with the reported table, the fusion results are limited to the development partition with no held-out evaluation, and the attention-ratio interpretation is partly circular due to the initialization described in Section 4.4. These issues currently weaken the otherwise valuable empirical study.

major comments (4)
  1. [Abstract and Section 1, with Table 2] The claim of outperforming the baseline by 17.52% is not consistent with the numbers in Table 2 under the standard relative-error definition: (5.03 - 4.28) / 5.03 = 14.9%. The value 17.52% corresponds to dividing by the proposed model's RMSE, i.e., (5.03 - 4.28) / 4.28, a convention that is never stated. This arithmetic inconsistency is load-bearing because the abstract and Section 1 both present the 17.52% figure as the paper's headline result. Please correct the percentage or explicitly define the relative-improvement formula used, and ensure the same convention is applied consistently across all reported improvements.
  2. [Section 5 and Table 2] The all-feature fusion result (RMSE 4.28) is reported only on the development partition. Section 5 states that test labels are unavailable and that most results are on the dev partition; the only test-set scores reported are for the text-based model (Section 5.1). Therefore the central claim 'outperforms the baseline by 17.52%' is not established on unseen data. The paper should clearly restrict the headline claim to the development partition, or, preferably, provide a test-set evaluation of the fusion model. As written, a reader could reasonably infer a held-out performance claim that the manuscript does not actually support.
  3. [Section 4.4, fusion initialization and attention ratios] The fusion nudge vector is initialized using reciprocal RMSE losses computed on the same development labels that are later used to report the final RMSE and attention ratios. This introduces circularity into the reported attention ratios [0.21262352, 0.21262285, 0.57475364]: the initialization explicitly prioritizes the text modality, and the final ratios still reflect that preference. The paper presents these ratios in Section 7 as learned 'importance' weights without acknowledging this bias. Please provide an ablation without this initialization, or initialize from training-fold-only losses and report the resulting attention ratios, so that the claimed modality-importance finding is not an artifact of the chosen nudge.
  4. [Section 4.4 and Table 2] Two different models are both described as 'Video-Text fused': the third fusion model (attention vector output from the video modality combined with text) and the fifth fusion model (video sub-modalities combined with text through a Bi-LSTM and attention). It is therefore unclear which configuration is reported in Table 2 under the name 'Video-Text fused.' This ambiguity harms reproducibility and should be resolved by giving each fusion variant a distinct name and consistent description in both the text and the table.
minor comments (6)
  1. [Section 1.1] The second contribution bullet is incomplete: 'The proposed approach outperforms the baseline fusion network by - on root mean square error.' Please complete or remove this bullet.
  2. [Abstract and Section 5] The abstract states the 17.52% improvement without any qualifier, while Section 5 explicitly says most results are on the dev partition. Please make the abstract and body consistent by stating that the fusion result is on the development partition.
  3. [Table 2 caption and Section 6] The third row of Table 2 (Qureshi et al.) is based on the test partition of DAIC-WOZ, not the E-DAIC dev partition used in this paper. The text acknowledges this caveat, but the table caption should state it as well to prevent misinterpretation.
  4. [Section 7] The statement that the baseline achieves a CCC of 0.1 on the test set is not supported by any cited source in the manuscript. Please add a citation to the AVEC 2019 baseline paper or remove the unsubstantiated number.
  5. [Section 3, Figure 1] The paper references Figure 1 in Section 3, but the figure is not placed within the text of this version. In the final version, ensure the block diagram is included and properly captioned.
  6. [General] There are several typographical and stylistic errors, including 'questionnaries' in the Introduction, inconsistent hyphenation of 'state-of-art' versus 'state-of-the-art', and the phrase 'DS-DNet features is 0.13 seconds' in Section 5.2. A careful proofread is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central fusion result is an external benchmark comparison, and the attention-ratio initialization is a training heuristic that does not force the reported outcome by construction.

full rationale

The paper's central claim is an empirical RMSE comparison against the AVEC 2019 baseline on the E-DAIC development partition (Table 2: 4.28 vs. 5.03). The baseline is an external result from the challenge organizers, not an input to the proposed network, so the comparison is self-contained and falsifiable. The only candidate circular step is the fusion nudge vector in Section 4.4, initialized with reciprocal per-modality RMSEs to prioritize text. However, the final attention ratios ([0.2126, 0.2126, 0.5748]) are learned weights that differ from the normalized reciprocal-RMSE initialization, and the reported fusion RMSE is a trained model output, not an algebraic consequence of the initialization. The paper explicitly acknowledges that test labels are unavailable and that most results are on the dev partition; this is an evaluation limitation, not circularity. The 17.52% improvement figure is arithmetically inconsistent with the standard relative-error convention, but a miscomputed percentage is a correctness issue, not a circular derivation. There is no load-bearing self-citation, no imported uniqueness theorem, and no renamed known result. Therefore no step of the claimed derivation reduces to its own input by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the representativeness of the E-DAIC dev split, the validity of PHQ-8 labels, and an ad hoc fusion initialization built from per-modality dev RMSEs. The nudge vector is the clearest free parameter: it encodes dev performance into the fusion and biases the attention ratios, while no new physical or conceptual entities are introduced.

free parameters (3)
  • Nudge initialization vector for fusion attention = Reciprocal RMSE ratios from individual modality models; final attention ratios [0.2126, 0.2126, 0.5748]
    Section 4.4 initializes the fusion output with a vector built from reciprocal RMSE losses of the individual modalities to prioritize text. This injects dev-set performance into the fusion and biases the learned attention ratios.
  • Text sequence length cutoff = 400 timesteps, zero-padded
    Section 4.1 fixes the tensor size by zero padding to 400 timesteps; truncation and padding choices affect what the text model sees and are chosen by hand.
  • Network hyperparameters = 200 hidden units per BLSTM layer, MLP widths 500-100-60-1, batch size 10, 15 epochs
    Section 4 and Section 5.1 state these values were chosen empirically; the final results depend on them, but no sensitivity analysis is provided.
assumptions (4)
  • domain assumption E-DAIC challenge-provided features and OpenFace outputs are reliable representations of interview behavior
    The audio, video, and text inputs are taken as given from the challenge and OpenFace; no feature quality verification is provided.
  • domain assumption PHQ-8 scores in E-DAIC are valid ground truth labels for depression severity
    The regression target is the self-report PHQ-8 score; the introduction acknowledges subjectivity in assessment, but the model treats these labels as noise-free ground truth.
  • domain assumption The development partition is representative of the withheld test partition
    Section 5 reports fusion results only on dev because test labels are unavailable; the headline claim assumes dev RMSE transfers to unseen interviews.
  • standard math Standard backpropagation training of the attention and LSTM networks converges to a stable solution
    The paper relies on standard deep learning training assumptions and provides no convergence guarantees or repeated-seed analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-level Attention network using text, audio and video for Depression Prediction." pith.science (2026). https://pith.science/paper/4XWJ5YND

@misc{pith2026190901417,
  author       = {Pith},
  title        = {Pith review of: Multi-level Attention network using text, audio and video for Depression Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4XWJ5YND}},
  note         = {Machine review of arXiv:1909.01417}
}
read the original abstract

Depression has been the leading cause of mental-health illness worldwide. Major depressive disorder (MDD), is a common mental health disorder that affects both psychologically as well as physically which could lead to loss of lives. Due to the lack of diagnostic tests and subjectivity involved in detecting depression, there is a growing interest in using behavioural cues to automate depression diagnosis and stage prediction. The absence of labelled behavioural datasets for such problems and the huge amount of variations possible in behaviour makes the problem more challenging. This paper presents a novel multi-level attention based network for multi-modal depression prediction that fuses features from audio, video and text modalities while learning the intra and inter modality relevance. The multi-level attention reinforces overall learning by selecting the most influential features within each modality for the decision making. We perform exhaustive experimentation to create different regression models for audio, video and text modalities. Several fusions models with different configurations are constructed to understand the impact of each feature and modality. We outperform the current baseline by 17.52% in terms of root mean squared error.

Figures

Figures reproduced from arXiv: 1909.01417 by the authors.

Figure 1
Figure 1. Block diagram of proposed multi-layer attention network on multi-modality input features [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 43 canonical work pages

  1. [1]

    Sharifa Alghowinem, Roland Goecke, Julien Epps, Michael Wagner, and Jeffrey F Cohn. 2016. Cross-Cultural Depression Recognition from Vocal Biomarkers.. In INTERSPEECH. 1943–1947

  2. [2]

    Sharifa Alghowinem, Roland Goecke, Michael Wagner, Julien Epps, Matthew Hyett, Gordon Parker, and Michael Breakspear. 2016. Multimodal depression detection: fusion analysis of paralinguistic, head pose and eye gaze behaviors. IEEE Transactions on Affective Computing 9, 4 (2016), 478–490

  3. [3]

    Haroon Ansari, Aditya Vijayvergia, and Krishan Kumar. 2018. DCR-HMM: Depression detection based on Content Rating using Hidden Markov Model. In 2018 Conference on Information and Communication Technology (CICT) . IEEE, 1–6

  4. [4]

    Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In 2016 IEEE Winter Conference on Applications of Computer Vision (W ACV). IEEE, 1–10

  5. [5]

    Daniel Bone, Chi-Chun Lee, and Shrikanth Narayanan. 2014. Robust unsupervised arousal rating: A rule-based framework withknowledge-inspired vocal features. IEEE transactions on affective computing 5, 2 (2014), 201–213

  6. [6]

    Meadows G Carey M, Jones K. 2014. Accuracy of general practitioner unassisted detection of depression. Aust N Z J Psychiatry. 48(6) (4 2014), 571–578

  7. [7]

    Patricia A Cavazos-Rehg, Melissa J Krauss, Shaina Sowles, Sarah Connolly, Carlos Rosas, Meghana Bharadwaj, and Laura J Bierut. 2016. A content analysis of depression-related tweets. Computers in human behavior 54 (2016), 351–357

  8. [8]

    John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun- Hsuan Sung, Brian Strope, and Ray Kurzweil

    Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun- Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Universal Sentence Encoder. CoRR abs/1803.11175 (2018). http://arxiv.org/abs/1803.11175

Show all 48 references
  1. [9]

    J. F. Cohn, T. S. Kruez, I. Matthews, Y. Yang, M. H. Nguyen, M. T. Padilla, F. Zhou, and F. De la Torre. 2009. Detecting depression from facial actions and vocal prosody. In 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops

  2. [10]

    Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, and Thomas F Quatieri. 2015. A review of depression and suicide risk assessment using speech analysis. Speech Communication 71 (2015), 10–49

  3. [11]

    Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Sebastian Schnieder, and Jarek Krajewski. 2015. Analysis of acoustic space variability in speech affected by depression. Speech Communication 75 (2015), 27–49

  4. [12]

    MN Dalili, IS Penton-Voak, CJ Harmer, and MR Munafò. 2015. Meta-analysis of emotion recognition deficits in major depressive disorder. Psychological medicine 45, 6 (2015), 1135–1144

  5. [13]

    David DeVault, Ron Artstein, Grace Benn, Teresa Dey, Ed Fast, Alesia Gainer, Kallirroi Georgila, Jon Gratch, Arno Hartholt, Margaux Lhommet, Gale Lucas, Stacy Marsella, Fabrizio Morbini, Angela Nazarian, Stefan Scherer, Giota Stratou, Apar Suri, David Traum, Rachel Wood, Yuyu ...

  6. [14]

    Samira Ebrahimi Kahou, Vincent Michalski, Kishore Konda, Roland Memisevic, and Christopher Pal. 2015. Recurrent neural networks for emotion recognition in video. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 467–474

  7. [15]

    Florian Eyben, Klaus Scherer, Björn Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shrikanth Narayanan, and Khiet Phuong Truong. 2016. The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affect...

  8. [16]

    Florian Eyben, F Weninger, and BjÃűrn Schuller. 2013. Affect recognition in real-life acoustic conditions - A new perspective on feature selection.Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH (2013), 2044–2048

  9. [17]

    Fabien Ringeval and Björn Schuller and Michel Valstar and Nicholas Cummins and Roddy Cowie and Leili Tavabi and Maximilian Schmitt and Sina Alisamir and Shahin Amiriparian and Eva-Maria Messner and Siyang Song and Shuo Lui and Ziping Zhao and Adria Mallol-Ragolta and Zhao Ren,...

  10. [18]

    Jonathan Gratch, Ron Arstein, Gale Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, David Traum, Albert Rizzo, and L P. Morency. 2014. The Distress Analysis Interview Corpus of human and computer interviews

  11. [19]

    Jonathan Gratch, Ron Artstein, Gale M Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, et al

  12. [20]

    Paulo A Graziano and Alexis Garcia. 2016. Attention-deficit hyperactivity disor- der and children’s emotion dysregulation: A meta-analysis. Clinical psychology review 46 (2016), 106–123

  13. [21]

    Fei Hao, Guangyao Pang, Yulei Wu, Zhongling Pi, Lirong Xia, and Geyong Min

  14. [22]

    Anees Ul Hassan, Jamil Hussain, Musarrat Hussain, Muhammad Sadiq, and Sungy- oung Lee. 2017. Sentiment analysis of social networking sites (SNS) data using machine learning approach for the measurement of depression. In 2017 Interna- tional Conference on Information and Commun...

  15. [23]

    Jia Jia. 2018. Mental Health Computing via Harvesting Social Media Data.. In IJCAI. 5677–5681

  16. [24]

    IEEE Transactions on Computational Social Systems (2019)

    Providing Appropriate Social Support to Prevention of Depression for Highly Anxious Sufferers. IEEE Transactions on Computational Social Systems (2019)

  17. [25]

    Genevieve Lam, Huang Dongyan, and Weisi Lin. 2019. Context-aware Deep Learning for Multi-modal Depression Detection. In ICASSP 2019-2019 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3946–3950

  18. [26]

    Amanda McCleery, Junghee Lee, Aditi Joshi, Jonathan K Wynn, Gerhard S Helle- mann, and Michael F Green. 2015. Meta-analysis of face processing event-related potentials in schizophrenia. Biological psychiatry 77, 2 (2015), 116–126

  19. [27]

    Kurt Kroenke, Tara Strine, Robert L Spitzer, Janet Williams, Joyce T Berry, and Ali Mokdad. 2008. The PHQ-8 as a Measure of Current Depression in the General Population. Journal of affective disorders 114 (09 2008), 163–73

  20. [28]

    Michelle Morales, Stefan Scherer, and Rivka Levitan. 2018. A linguistically- informed fusion approach for multimodal depression detection. In Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic. 13–24

  21. [29]

    Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Context-dependent sentiment anal- ysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume...

  22. [30]

    Vikramjit Mitra, Andreas Tsiartas, and Elizabeth Shriberg. 2016. Noise and rever- beration effects on depression detection from speech. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5795–5799

  23. [31]

    Soujanya Poria, Iti Chaturvedi, Erik Cambria, and Amir Hussain. 2016. Convolu- tional MKL based multimodal emotion recognition and sentiment analysis. In 2016 IEEE 16th international conference on data mining (ICDM) . IEEE, 439–448

  24. [32]

    Syed Arbaaz Qureshi, Mohammed Hasanuzzaman, Sriparna Saha, and Gaël Dias

  25. [33]

    Soujanya Poria, Erik Cambria, Newton Howard, Guang-Bin Huang, and Amir Hussain. 2016. Fusing audio, visual and textual clues for sentiment analysis from multimodal content. Neurocomputing 174 (2016), 50–59

  26. [34]

    Stefan Scherer, Giota Stratou, Gale Lucas, Marwa Mahmoud, Jill Boberg, Jonathan Gratch, Albert (Skip) Rizzo, and Louis-Philippe Morency. 2014. Automatic audio- visual behavior descriptors for psychological disorder analysis. Image and Vision Computing Journal 32 (Oct. 2014)

  27. [35]

    Schuller

    Maximilian Schmitt and Björn W. Schuller. 2016. openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit. CoRR abs/1605.06778 (2016). arXiv:1605.06778 http://arxiv.org/abs/1605.06778

  28. [36]

    The Verbal and Non Verbal Signals of Depression–Combining Acoustics, Text and Visuals for Estimating Depression Level.arXiv preprint arXiv:1904.07656 (2019)

  29. [37]

    Fabien Ringeval, Björn Schuller, Michel Valstar, Shashank Jaiswal, Erik Marchi, Denis Lalanne, Roddy Cowie, and Maja Pantic. 2015. AV+EC 2015: The First Affect Recognition Challenge Bridging Across Audio, Video, and Physiological Data. In Proceedings of the 5th International W...

  30. [38]

    Brian Stasak, Julien Epps, Nicholas Cummins, and Roland Goecke. 2016. An Investigation of Emotional Speech in Depression Classification.. In Interspeech. 485–489

  31. [39]

    Lei Tong, Qianni Zhang, Abdul Sadka, Ling Li, Huiyu Zhou, et al . 2019. In- verse boosting pruning trees for depression detection on Twitter. arXiv preprint arXiv:1906.00398 (2019)

  32. [40]

    Schuller, Anton Batliner, Dino Seppi, Stefan Steidl, Thurid Vogt, Jo- hannes Wagner, Laurence Devillers, Laurence Vidrascu, Noam Amir, Loïc Kessous, and Vered Aharonson

    Björn W. Schuller, Anton Batliner, Dino Seppi, Stefan Steidl, Thurid Vogt, Jo- hannes Wagner, Laurence Devillers, Laurence Vidrascu, Noam Amir, Loïc Kessous, and Vered Aharonson. 2007. The relevance of feature type for the automatic classification of emotional user states: low...

  33. [41]

    Guangyao Shen, Jia Jia, Liqiang Nie, Fuli Feng, Cunjun Zhang, Tianrui Hu, Tat- Seng Chua, and Wenwu Zhu. 2017. Depression Detection via Harvesting Social Media: A Multimodal Dictionary Learning Solution.. In IJCAI. 3838–3844

  34. [42]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. In Proceedings of the 31st International Conference on Neural Informa- tion Processing Systems (NIPS’17)

  35. [43]

    James R Williamson, Elizabeth Godoy, Miriam Cha, Adrianne Schwarzentru- ber, Pooya Khorrami, Youngjune Gwon, Hsiang-Tsung Kung, Charlie Dagli, and Thomas F Quatieri. 2016. Detecting depression using vocal, facial and seman- tic communication cues. In Proceedings of the 6th Int...

  36. [44]

    Valstar, Jonathan Gratch, Björn W

    Michel F. Valstar, Jonathan Gratch, Björn W. Schuller, Fabien Ringeval, Denis Lalanne, Mercedes Torres, Stefan Scherer, Giota Stratou, Roddy Cowie, and Maja Pantic. 2016. AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge. CoRR abs/1605.01600 (2016). ...

  37. [45]

    Michel F Valstar, Enrique Sánchez-Lozano, Jeffrey F Cohn, László A Jeni, Jeffrey M Girard, Zheng Zhang, Lijun Yin, and Maja Pantic. 2017. Fera 2017-addressing head pose in the third facial expression recognition and analysis challenge. In 2017 12th IEEE International Conferenc...

  38. [48]

    Zheng Zhang, Jeff M Girard, Yue Wu, Xing Zhang, Peng Liu, Umur Ciftci, Shaun Canavan, Michael Reale, Andy Horowitz, Huiyuan Yang, et al. 2016. Multimodal spontaneous emotion corpus for human behavior analysis. In Proceedings of the IEEE Conference on Computer Vision and Patter...

  39. [2014]

    The distress analysis interview corpus of human and computer interviews.. In LREC. 3123–3128

  40. [2019]

    AVEC 2019 Workshop and Challenge: State-of-Mind, Depression with AI, and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio/Visual Emotion Challenge, A VEC’19, co-located with the 27th ACM International Conference on Multimedia, MM 201...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.