REVIEW 4 major objections 6 minor 48 references
Multi-level Attention network using text, audio and video for Depression Prediction
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Multi-level attention over text, audio, and video predicts PHQ-8 depression scores with RMSE 4.28 on E-DAIC development data, beating the AVEC 2019 baseline of 5.03.
desk verdict Useful but modest AVEC 2019 fusion work whose headline improvement is dev-only and arithmetically overstated; the text-only test numbers are the most solid part. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multi-level attention fusion network: stacked bidirectional LSTMs over each modality produce hidden states, an attention layer over each stream builds a context vector, feedforward layers compress each context, and a second attention layer over the compressed modalities produces the final weighted representation that is multiplied into a stacked-BLSTM output and regressed to the PHQ-8 score. The intra-modality attention selects the informative timesteps or features within a stream, while the inter-modality attention learns the contribution ratios of text, audio and video. To stabilise training of the three-way fusion, the paper initialises a multiplicative 'nudge' vector with the reciprocal ratios of each single-modality RMSE, which biases the optimisation initially toward the text pathway; the final attention ratios after convergence are approximately 0.21 (video), 0.21 (audio) and 0.57 (text).
What would settle it
Run the proposed all-feature fusion model on the E-DAIC test partition and compare its RMSE with the AVEC 2019 baseline; if the test RMSE does not beat 5.03 by a comparable margin, the central improvement claim fails. A simpler check is arithmetic: (5.03 - 4.28)/5.03 is about 14.9%, which already differs from the reported 17.52% under the usual relative-error definition.
Extended reading notes
Core claim
The paper's central claim is that applying attention at multiple levels—once over the sequence outputs within each modality and again over the fused modality representations—lets a regression network predict depression severity better than the challenge baseline. The all-feature fusion model, built from text sentence embeddings, audio descriptors (MFCC, eGeMAPS, BoAW, deep densenet features) and video descriptors (pose, gaze, facial action units, bag-of-visual-words), achieves the best development-set RMSE of 4.28 compared with the baseline's 5.03, which the authors report as a 17.52% improvement. The learned attention ratios give the text modality roughly 57% of the weight, with audio and video near 21% each, and the authors interpret this as the network discovering that verbal content is the strongest marker. A text-only version of the model also performs well on the withheld test set, with a concordance correlation coefficient of 0.67, and the authors report that it outperforms the closest prior attention-based work by 8.95%.
Load-bearing premise
The paper assumes that the development partition's labels, which were used both to choose the best configurations and to initialize the fusion weights, are a reliable stand-in for the withheld test partition, so the reported improvement will survive on unseen data.
Editorial extensions
If this is right
- The all-feature fusion network reaches RMSE 4.28 on the E-DAIC development set, beating the AVEC 2019 baseline's 5.03.
- A text-only network scores 4.37 RMSE on dev and, on the challenge test set, MAE 4.02, RMSE 4.73 and CCC 0.67, which the authors report as stronger than the closest prior attention-based model.
- Learned attention weights rank text first (about 0.57), with audio and video nearly equal (about 0.21 each), implying that verbal content dominates automated depression scoring in this corpus.
Reading between the lines
- Because the development partition was used both to select the best configurations and to initialize the fusion nudge vector, the reported 17.52% improvement should be read as a dev-set estimate rather than proven generalization to the withheld test set.
- The nudge initialization biases the network toward the text modality from the start, so the final attention ratios may partly reflect that initialization rather than an unbiased discovery of modality importance; ablating the nudge vector would clarify this.
- The same multi-level attention architecture could be applied to other PHQ-based assessments or to anxiety and PTSD scores in the E-DAIC corpus, where similar behavioural markers are recorded.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-level attention network that fuses audio, video, and text features to predict PHQ-8 depression severity scores on the E-DAIC corpus. The authors build per-modality regression models using BLSTM networks, then experiment with several fusion configurations that apply attention within and across modalities. Their best model, an all-feature fusion network, is reported to achieve RMSE 4.28 on the development partition, which they describe as outperforming the AVEC 2019 baseline (RMSE 5.03) by 17.52%. The text-only model is also evaluated on the withheld test partition, achieving CCC 0.67. The manuscript includes extensive ablations across individual features and fusion variants.
Significance. If the central claim were fully supported, the paper would contribute a reasonably investigated multimodal attention architecture for depression prediction, with a thorough ablation over audio, video, and text features. The text-only model's test-set CCC of 0.67 and the systematic per-feature comparisons against the AVEC 2019 baseline are genuine strengths. The proposed multi-level attention idea is plausible and the empirical exploration is fairly detailed. However, the headline improvement claim is numerically inconsistent with the reported table, the fusion results are limited to the development partition with no held-out evaluation, and the attention-ratio interpretation is partly circular due to the initialization described in Section 4.4. These issues currently weaken the otherwise valuable empirical study.
major comments (4)
- [Abstract and Section 1, with Table 2] The claim of outperforming the baseline by 17.52% is not consistent with the numbers in Table 2 under the standard relative-error definition: (5.03 - 4.28) / 5.03 = 14.9%. The value 17.52% corresponds to dividing by the proposed model's RMSE, i.e., (5.03 - 4.28) / 4.28, a convention that is never stated. This arithmetic inconsistency is load-bearing because the abstract and Section 1 both present the 17.52% figure as the paper's headline result. Please correct the percentage or explicitly define the relative-improvement formula used, and ensure the same convention is applied consistently across all reported improvements.
- [Section 5 and Table 2] The all-feature fusion result (RMSE 4.28) is reported only on the development partition. Section 5 states that test labels are unavailable and that most results are on the dev partition; the only test-set scores reported are for the text-based model (Section 5.1). Therefore the central claim 'outperforms the baseline by 17.52%' is not established on unseen data. The paper should clearly restrict the headline claim to the development partition, or, preferably, provide a test-set evaluation of the fusion model. As written, a reader could reasonably infer a held-out performance claim that the manuscript does not actually support.
- [Section 4.4, fusion initialization and attention ratios] The fusion nudge vector is initialized using reciprocal RMSE losses computed on the same development labels that are later used to report the final RMSE and attention ratios. This introduces circularity into the reported attention ratios [0.21262352, 0.21262285, 0.57475364]: the initialization explicitly prioritizes the text modality, and the final ratios still reflect that preference. The paper presents these ratios in Section 7 as learned 'importance' weights without acknowledging this bias. Please provide an ablation without this initialization, or initialize from training-fold-only losses and report the resulting attention ratios, so that the claimed modality-importance finding is not an artifact of the chosen nudge.
- [Section 4.4 and Table 2] Two different models are both described as 'Video-Text fused': the third fusion model (attention vector output from the video modality combined with text) and the fifth fusion model (video sub-modalities combined with text through a Bi-LSTM and attention). It is therefore unclear which configuration is reported in Table 2 under the name 'Video-Text fused.' This ambiguity harms reproducibility and should be resolved by giving each fusion variant a distinct name and consistent description in both the text and the table.
minor comments (6)
- [Section 1.1] The second contribution bullet is incomplete: 'The proposed approach outperforms the baseline fusion network by - on root mean square error.' Please complete or remove this bullet.
- [Abstract and Section 5] The abstract states the 17.52% improvement without any qualifier, while Section 5 explicitly says most results are on the dev partition. Please make the abstract and body consistent by stating that the fusion result is on the development partition.
- [Table 2 caption and Section 6] The third row of Table 2 (Qureshi et al.) is based on the test partition of DAIC-WOZ, not the E-DAIC dev partition used in this paper. The text acknowledges this caveat, but the table caption should state it as well to prevent misinterpretation.
- [Section 7] The statement that the baseline achieves a CCC of 0.1 on the test set is not supported by any cited source in the manuscript. Please add a citation to the AVEC 2019 baseline paper or remove the unsubstantiated number.
- [Section 3, Figure 1] The paper references Figure 1 in Section 3, but the figure is not placed within the text of this version. In the final version, ensure the block diagram is included and properly captioned.
- [General] There are several typographical and stylistic errors, including 'questionnaries' in the Introduction, inconsistent hyphenation of 'state-of-art' versus 'state-of-the-art', and the phrase 'DS-DNet features is 0.13 seconds' in Section 5.2. A careful proofread is recommended.
Circularity Check
No significant circularity: the central fusion result is an external benchmark comparison, and the attention-ratio initialization is a training heuristic that does not force the reported outcome by construction.
full rationale
The paper's central claim is an empirical RMSE comparison against the AVEC 2019 baseline on the E-DAIC development partition (Table 2: 4.28 vs. 5.03). The baseline is an external result from the challenge organizers, not an input to the proposed network, so the comparison is self-contained and falsifiable. The only candidate circular step is the fusion nudge vector in Section 4.4, initialized with reciprocal per-modality RMSEs to prioritize text. However, the final attention ratios ([0.2126, 0.2126, 0.5748]) are learned weights that differ from the normalized reciprocal-RMSE initialization, and the reported fusion RMSE is a trained model output, not an algebraic consequence of the initialization. The paper explicitly acknowledges that test labels are unavailable and that most results are on the dev partition; this is an evaluation limitation, not circularity. The 17.52% improvement figure is arithmetically inconsistent with the standard relative-error convention, but a miscomputed percentage is a correctness issue, not a circular derivation. There is no load-bearing self-citation, no imported uniqueness theorem, and no renamed known result. Therefore no step of the claimed derivation reduces to its own input by construction.
Assumptions & free parameters
free parameters (3)
- Nudge initialization vector for fusion attention =
Reciprocal RMSE ratios from individual modality models; final attention ratios [0.2126, 0.2126, 0.5748]
- Text sequence length cutoff =
400 timesteps, zero-padded
- Network hyperparameters =
200 hidden units per BLSTM layer, MLP widths 500-100-60-1, batch size 10, 15 epochs
assumptions (4)
- domain assumption E-DAIC challenge-provided features and OpenFace outputs are reliable representations of interview behavior
- domain assumption PHQ-8 scores in E-DAIC are valid ground truth labels for depression severity
- domain assumption The development partition is representative of the withheld test partition
- standard math Standard backpropagation training of the attention and LSTM networks converges to a stable solution
Cite this review
Pith. "Pith review of Multi-level Attention network using text, audio and video for Depression Prediction." pith.science (2026). https://pith.science/paper/4XWJ5YND
@misc{pith2026190901417,
author = {Pith},
title = {Pith review of: Multi-level Attention network using text, audio and video for Depression Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/4XWJ5YND}},
note = {Machine review of arXiv:1909.01417}
}
read the original abstract
Depression has been the leading cause of mental-health illness worldwide. Major depressive disorder (MDD), is a common mental health disorder that affects both psychologically as well as physically which could lead to loss of lives. Due to the lack of diagnostic tests and subjectivity involved in detecting depression, there is a growing interest in using behavioural cues to automate depression diagnosis and stage prediction. The absence of labelled behavioural datasets for such problems and the huge amount of variations possible in behaviour makes the problem more challenging. This paper presents a novel multi-level attention based network for multi-modal depression prediction that fuses features from audio, video and text modalities while learning the intra and inter modality relevance. The multi-level attention reinforces overall learning by selecting the most influential features within each modality for the decision making. We perform exhaustive experimentation to create different regression models for audio, video and text modalities. Several fusions models with different configurations are constructed to understand the impact of each feature and modality. We outperform the current baseline by 17.52% in terms of root mean squared error.
Figures
Reference graph
Works this paper leans on
-
[1]
Sharifa Alghowinem, Roland Goecke, Julien Epps, Michael Wagner, and Jeffrey F Cohn. 2016. Cross-Cultural Depression Recognition from Vocal Biomarkers.. In INTERSPEECH. 1943–1947
work page 2016
-
[2]
Sharifa Alghowinem, Roland Goecke, Michael Wagner, Julien Epps, Matthew Hyett, Gordon Parker, and Michael Breakspear. 2016. Multimodal depression detection: fusion analysis of paralinguistic, head pose and eye gaze behaviors. IEEE Transactions on Affective Computing 9, 4 (2016), 478–490
work page 2016
-
[3]
Haroon Ansari, Aditya Vijayvergia, and Krishan Kumar. 2018. DCR-HMM: Depression detection based on Content Rating using Hidden Markov Model. In 2018 Conference on Information and Communication Technology (CICT) . IEEE, 1–6
work page 2018
-
[4]
Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In 2016 IEEE Winter Conference on Applications of Computer Vision (W ACV). IEEE, 1–10
work page 2016
-
[5]
Daniel Bone, Chi-Chun Lee, and Shrikanth Narayanan. 2014. Robust unsupervised arousal rating: A rule-based framework withknowledge-inspired vocal features. IEEE transactions on affective computing 5, 2 (2014), 201–213
work page 2014
-
[6]
Meadows G Carey M, Jones K. 2014. Accuracy of general practitioner unassisted detection of depression. Aust N Z J Psychiatry. 48(6) (4 2014), 571–578
work page 2014
-
[7]
Patricia A Cavazos-Rehg, Melissa J Krauss, Shaina Sowles, Sarah Connolly, Carlos Rosas, Meghana Bharadwaj, and Laura J Bierut. 2016. A content analysis of depression-related tweets. Computers in human behavior 54 (2016), 351–357
work page 2016
-
[8]
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun- Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Universal Sentence Encoder. CoRR abs/1803.11175 (2018). http://arxiv.org/abs/1803.11175
arXiv 2018
Show all 48 references
-
[9]
J. F. Cohn, T. S. Kruez, I. Matthews, Y. Yang, M. H. Nguyen, M. T. Padilla, F. Zhou, and F. De la Torre. 2009. Detecting depression from facial actions and vocal prosody. In 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops
2009
-
[10]
Nicholas Cummins, Stefan Scherer, Jarek Krajewski, Sebastian Schnieder, Julien Epps, and Thomas F Quatieri. 2015. A review of depression and suicide risk assessment using speech analysis. Speech Communication 71 (2015), 10–49
2015
-
[11]
Nicholas Cummins, Vidhyasaharan Sethu, Julien Epps, Sebastian Schnieder, and Jarek Krajewski. 2015. Analysis of acoustic space variability in speech affected by depression. Speech Communication 75 (2015), 27–49
2015
-
[12]
MN Dalili, IS Penton-Voak, CJ Harmer, and MR Munafò. 2015. Meta-analysis of emotion recognition deficits in major depressive disorder. Psychological medicine 45, 6 (2015), 1135–1144
2015
-
[13]
David DeVault, Ron Artstein, Grace Benn, Teresa Dey, Ed Fast, Alesia Gainer, Kallirroi Georgila, Jon Gratch, Arno Hartholt, Margaux Lhommet, Gale Lucas, Stacy Marsella, Fabrizio Morbini, Angela Nazarian, Stefan Scherer, Giota Stratou, Apar Suri, David Traum, Rachel Wood, Yuyu ...
2014
-
[14]
Samira Ebrahimi Kahou, Vincent Michalski, Kishore Konda, Roland Memisevic, and Christopher Pal. 2015. Recurrent neural networks for emotion recognition in video. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 467–474
2015
-
[15]
Florian Eyben, Klaus Scherer, Björn Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shrikanth Narayanan, and Khiet Phuong Truong. 2016. The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affect...
2016
-
[16]
Florian Eyben, F Weninger, and BjÃűrn Schuller. 2013. Affect recognition in real-life acoustic conditions - A new perspective on feature selection.Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH (2013), 2044–2048
2013
-
[17]
Fabien Ringeval and Björn Schuller and Michel Valstar and Nicholas Cummins and Roddy Cowie and Leili Tavabi and Maximilian Schmitt and Sina Alisamir and Shahin Amiriparian and Eva-Maria Messner and Siyang Song and Shuo Lui and Ziping Zhao and Adria Mallol-Ragolta and Zhao Ren,...
-
[18]
Jonathan Gratch, Ron Arstein, Gale Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, David Traum, Albert Rizzo, and L P. Morency. 2014. The Distress Analysis Interview Corpus of human and computer interviews
2014
-
[19]
Jonathan Gratch, Ron Artstein, Gale M Lucas, Giota Stratou, Stefan Scherer, Angela Nazarian, Rachel Wood, Jill Boberg, David DeVault, Stacy Marsella, et al
-
[20]
Paulo A Graziano and Alexis Garcia. 2016. Attention-deficit hyperactivity disor- der and children’s emotion dysregulation: A meta-analysis. Clinical psychology review 46 (2016), 106–123
2016
-
[21]
Fei Hao, Guangyao Pang, Yulei Wu, Zhongling Pi, Lirong Xia, and Geyong Min
-
[22]
Anees Ul Hassan, Jamil Hussain, Musarrat Hussain, Muhammad Sadiq, and Sungy- oung Lee. 2017. Sentiment analysis of social networking sites (SNS) data using machine learning approach for the measurement of depression. In 2017 Interna- tional Conference on Information and Commun...
2017
-
[23]
Jia Jia. 2018. Mental Health Computing via Harvesting Social Media Data.. In IJCAI. 5677–5681
2018
-
[24]
IEEE Transactions on Computational Social Systems (2019)
Providing Appropriate Social Support to Prevention of Depression for Highly Anxious Sufferers. IEEE Transactions on Computational Social Systems (2019)
2019
-
[25]
Genevieve Lam, Huang Dongyan, and Weisi Lin. 2019. Context-aware Deep Learning for Multi-modal Depression Detection. In ICASSP 2019-2019 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3946–3950
2019
-
[26]
Amanda McCleery, Junghee Lee, Aditi Joshi, Jonathan K Wynn, Gerhard S Helle- mann, and Michael F Green. 2015. Meta-analysis of face processing event-related potentials in schizophrenia. Biological psychiatry 77, 2 (2015), 116–126
2015
-
[27]
Kurt Kroenke, Tara Strine, Robert L Spitzer, Janet Williams, Joyce T Berry, and Ali Mokdad. 2008. The PHQ-8 as a Measure of Current Depression in the General Population. Journal of affective disorders 114 (09 2008), 163–73
2008
-
[28]
Michelle Morales, Stefan Scherer, and Rivka Levitan. 2018. A linguistically- informed fusion approach for multimodal depression detection. In Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic. 13–24
2018
-
[29]
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Context-dependent sentiment anal- ysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume...
2017
-
[30]
Vikramjit Mitra, Andreas Tsiartas, and Elizabeth Shriberg. 2016. Noise and rever- beration effects on depression detection from speech. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5795–5799
2016
-
[31]
Soujanya Poria, Iti Chaturvedi, Erik Cambria, and Amir Hussain. 2016. Convolu- tional MKL based multimodal emotion recognition and sentiment analysis. In 2016 IEEE 16th international conference on data mining (ICDM) . IEEE, 439–448
2016
-
[32]
Syed Arbaaz Qureshi, Mohammed Hasanuzzaman, Sriparna Saha, and Gaël Dias
-
[33]
Soujanya Poria, Erik Cambria, Newton Howard, Guang-Bin Huang, and Amir Hussain. 2016. Fusing audio, visual and textual clues for sentiment analysis from multimodal content. Neurocomputing 174 (2016), 50–59
2016
-
[34]
Stefan Scherer, Giota Stratou, Gale Lucas, Marwa Mahmoud, Jill Boberg, Jonathan Gratch, Albert (Skip) Rizzo, and Louis-Philippe Morency. 2014. Automatic audio- visual behavior descriptors for psychological disorder analysis. Image and Vision Computing Journal 32 (Oct. 2014)
2014
-
[35]
Schuller
Maximilian Schmitt and Björn W. Schuller. 2016. openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit. CoRR abs/1605.06778 (2016). arXiv:1605.06778 http://arxiv.org/abs/1605.06778
2016 arXiv
-
[36]
The Verbal and Non Verbal Signals of Depression–Combining Acoustics, Text and Visuals for Estimating Depression Level.arXiv preprint arXiv:1904.07656 (2019)
2019 arXiv
-
[37]
Fabien Ringeval, Björn Schuller, Michel Valstar, Shashank Jaiswal, Erik Marchi, Denis Lalanne, Roddy Cowie, and Maja Pantic. 2015. AV+EC 2015: The First Affect Recognition Challenge Bridging Across Audio, Video, and Physiological Data. In Proceedings of the 5th International W...
2015
-
[38]
Brian Stasak, Julien Epps, Nicholas Cummins, and Roland Goecke. 2016. An Investigation of Emotional Speech in Depression Classification.. In Interspeech. 485–489
2016
-
[39]
Lei Tong, Qianni Zhang, Abdul Sadka, Ling Li, Huiyu Zhou, et al . 2019. In- verse boosting pruning trees for depression detection on Twitter. arXiv preprint arXiv:1906.00398 (2019)
2019 arXiv
-
[40]
Schuller, Anton Batliner, Dino Seppi, Stefan Steidl, Thurid Vogt, Jo- hannes Wagner, Laurence Devillers, Laurence Vidrascu, Noam Amir, Loïc Kessous, and Vered Aharonson
Björn W. Schuller, Anton Batliner, Dino Seppi, Stefan Steidl, Thurid Vogt, Jo- hannes Wagner, Laurence Devillers, Laurence Vidrascu, Noam Amir, Loïc Kessous, and Vered Aharonson. 2007. The relevance of feature type for the automatic classification of emotional user states: low...
2007
-
[41]
Guangyao Shen, Jia Jia, Liqiang Nie, Fuli Feng, Cunjun Zhang, Tianrui Hu, Tat- Seng Chua, and Wenwu Zhu. 2017. Depression Detection via Harvesting Social Media: A Multimodal Dictionary Learning Solution.. In IJCAI. 3838–3844
2017
-
[42]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. In Proceedings of the 31st International Conference on Neural Informa- tion Processing Systems (NIPS’17)
2017
-
[43]
James R Williamson, Elizabeth Godoy, Miriam Cha, Adrianne Schwarzentru- ber, Pooya Khorrami, Youngjune Gwon, Hsiang-Tsung Kung, Charlie Dagli, and Thomas F Quatieri. 2016. Detecting depression using vocal, facial and seman- tic communication cues. In Proceedings of the 6th Int...
2016
-
[44]
Valstar, Jonathan Gratch, Björn W
Michel F. Valstar, Jonathan Gratch, Björn W. Schuller, Fabien Ringeval, Denis Lalanne, Mercedes Torres, Stefan Scherer, Giota Stratou, Roddy Cowie, and Maja Pantic. 2016. AVEC 2016 - Depression, Mood, and Emotion Recognition Workshop and Challenge. CoRR abs/1605.01600 (2016). ...
2016 arXiv
-
[45]
Michel F Valstar, Enrique Sánchez-Lozano, Jeffrey F Cohn, László A Jeni, Jeffrey M Girard, Zheng Zhang, Lijun Yin, and Maja Pantic. 2017. Fera 2017-addressing head pose in the third facial expression recognition and analysis challenge. In 2017 12th IEEE International Conferenc...
2017
-
[48]
Zheng Zhang, Jeff M Girard, Yue Wu, Xing Zhang, Peng Liu, Umur Ciftci, Shaun Canavan, Michael Reale, Andy Horowitz, Huiyuan Yang, et al. 2016. Multimodal spontaneous emotion corpus for human behavior analysis. In Proceedings of the IEEE Conference on Computer Vision and Patter...
2016
-
[2014]
The distress analysis interview corpus of human and computer interviews.. In LREC. 3123–3128
-
[2019]
AVEC 2019 Workshop and Challenge: State-of-Mind, Depression with AI, and Cross-Cultural Affect Recognition. In Proceedings of the 9th International Workshop on Audio/Visual Emotion Challenge, A VEC’19, co-located with the 27th ACM International Conference on Multimedia, MM 201...
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.