{"id":"f082d473-882e-4b1f-b85d-2e6a21cee3bc","arxiv_id":"2506.18915","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey of machine learning approaches for automatic depression assessment from multimodal human behaviour and brain data, covering 183 papers and ten public datasets.","lead":"This paper reviews 183 machine learning studies that automatically assess depression from brain signals, face, body, voice, language, and their combinations. It maps the field's methods, datasets, and grand challenges, and identifies open problems for researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comprehensiveness claim is undercut by a search protocol that omits PubMed/Scopus and defers PRISMA counts; the 183-paper set may be systematically incomplete for EEG/MRI modalities.","rationale":"The survey is a useful synthesis with clear structure, extensive tables, and a public repository, which gives independent value. The concern is not about the quality of individual summaries but about the central claim of comprehensiveness. The search protocol's database list is a concrete, testable gap: PubMed/Scopus omissions are especially relevant for EEG/MRI modalities. Even if the supplementary contains the search queries, the omission of these databases can bias the corpus. This is an internally verifiable methodological gap rather than a mere disagreement with consensus. A yes/no test exists (re-run search), so the condition can be resolved. The reader's own concern about missing PRISMA/search queries is related, but our emphasis on the database set is more specific. The verdict remains CONDITIONAL; no change to the reader's recommendation is needed because the claim can be verified and the paper still has reference value conditional on addressing the search gaps.","tokens_in":46441,"tokens_out":6225,"duration_ms":76427,"concrete_test":"Retrieve the exact search strings from the Supplementary Material and re-execute the same eligibility criteria (2012–2024, ML-based, English, new ADA method) on PubMed/MEDLINE and Scopus. Count the additional distinct papers meeting the criteria; if the gain is more than ~5% of 183 or concentrates in EEG/MRI, the 'comprehensive' claim is over-stated and the survey should be revised or explicitly re-scoped.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first comprehensive survey of ADA across brain (EEG/MRI) and expressive modalities. That claim requires the Section 3 search to be exhaustive and representative. The reported databases are IEEE Xplore, ACM DL, Elsevier, Springer, ACL Anthology, IET and MIT Press, but not PubMed/MEDLINE, Scopus or Web of Science. EEG and MRI depression-classification work is heavily published in clinical venues indexed by PubMed but not necessarily by those listed databases, so the 183 included papers may be systematically skewed away from the brain-imaging modalities the survey claims to cover. Moreover, the paper defers the exact search strings and PRISMA exclusion counts to the Supplementary Material, and Fig. 3 shows a flow diagram without the stage-wise numbers in the main text. Consequently, a reader cannot currently determine whether the corpus is complete or representative, so the load-bearing assumption of an exhaustive search is insecure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a survey of machine learning-based automatic depression assessment (ADA) methods, covering brain modalities (EEG, MRI), expressive modalities (face, body, audio, language), and multimodal combinations. It synthesizes 183 papers, reviews public datasets and AVEC challenge series, and discusses challenges and opportunities. The central claim is that this is the first comprehensive survey of ADA across both internal brain activities and external expressive behaviours.","tokens_in":46593,"tokens_out":2573,"duration_ms":28629,"significance":"If the survey is indeed comprehensive and accurate, it would be a valuable reference for a rapidly growing field: it aggregates 183 methods, organizes them by modality and ML approach, links depression biomarkers to computational features, and provides a central web repository. The paper also gives due attention to privacy, label quality, and fairness. However, its significance is conditional on the reproducibility of the search protocol and on the accuracy of its challenge statements, since the comprehensiveness claim is load-bearing.","major_comments":[{"comment":"The passage under 'Lack of interpretability and reproducibility' states that 'no existing work has investigated the interpretability, explainability, nor reproducibility of such models.' This is directly contradicted by several papers reviewed in this survey, including Ref. [33] (Gahalawat et al., 'Explainable depression detection via head motion patterns'), Ref. [201] (Zhou et al., 'Visually interpretable representation learning'), Ref. [282] (Kacem et al., 'interpretable representations of motion dynamics'), and Ref. [210] (Moreno et al., 'Expresso-ai: An explainable video-based deep learning model'). The overstatement weakens the credibility of the challenges discussion and should be corrected to a more precise formulation, e.g., that few existing works systematically evaluate interpretability and reproducibility.","section":"Section 6.1"},{"comment":"The central claim of being a 'comprehensive survey' rests on the exhaustiveness and representativeness of the literature search. However, the search databases listed (IEEE Xplore, ACM Digital Library, Elsevier, Springer, ACL Anthology, IET, MIT Press) omit PubMed/MEDLINE, Scopus, and Web of Science, which are the primary indexes for clinical EEG and MRI depression-classification studies. Since the exact search strings, PRISMA flow numbers, and exclusion counts are deferred to the Supplementary Material, and since Fig. 3 shows a flow diagram without stage-wise numbers in the main text, a reader cannot currently reproduce or audit the search. This leaves the 183-paper set potentially systematically skewed away from the very brain-imaging modalities the survey claims to cover. The authors should either justify the database choice, add the missing databases, or report the full search protocol and PRISMA counts in the main text or a clearly reproducible appendix.","section":"Section 3"}],"minor_comments":[{"comment":"The figure caption says 'PRIMA Schema' but should read 'PRISMA Schema.'","section":"Fig. 3"},{"comment":"The phrase 'revealing subtle changes in energy overtime' should be 'revealing subtle changes in energy over time.'","section":"Section 4.5.2"},{"comment":"The word 'behavuours' in 'these behavuours are more prone to data privacy' is a typo and should be 'behaviours.'","section":"Section 4.3.1"},{"comment":"The text mentions 'DAIZ-WOZ depression dataset' twice; the correct acronym for the referenced dataset is DAIC-WOZ (Ref. [64]).","section":"Section 5.2"},{"comment":"In the same paragraph on reproducibility, the paper lists several open-source implementations ([20], [28], [42], [69], [203], [215]), which is useful; consider making this a separate concrete point in the future-directions section rather than embedding it in the challenge statement.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The survey includes several works from the same research groups as the authors (e.g., Exeter, Cambridge, Nottingham). These appear to be peer-reviewed publications and are interleaved with third-party citations, so I do not see a circularity problem. However, the authors could be more transparent about the inclusion of their own works in the survey’s statistics. The main technical concern is that the literature-search protocol is underspecified in the main text; this is fixable and should be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful survey, probably the broadest modality coverage in the ADA survey genre. The EEG/MRI inclusion, plus the summary of depression-linked behavioural evidence in Section 2.2, is the new part, and it is done carefully. I would trust it as an entry point for a new researcher, and I would cite it.\n\nThe paper is not methodologically groundbreaking—it is a synthesis—but the synthesis is the contribution, and the modality breadth is real. The modality-by-modality organization, the summary table, and the dataset/competition review are all serviceable. The authors also make a reasonable case for why prior surveys were narrower.\n\nSoft spots, in order of importance. First, Section 6.1 says 'no existing work has investigated the interpretability, explainability, nor reproducibility of such models,' but the survey itself cites explainable and interpretable ADA papers (for example, references [33] and [210], among others). That is an internal overstatement and should be fixed; it undercuts an otherwise sensible discussion. Second, the comprehensiveness claim in the abstract and conclusion rests on the Section 3 search, and the main text reports only that 1679 papers were retrieved and 183 included. The search strings, exclusion counts, and full PRISMA flow are deferred to the supplementary material, and the database list omits PubMed/Scopus/Web of Science. For EEG/MRI depression work, much of the literature is in clinical venues indexed by PubMed, so I share the stress-test concern that the corpus may be systematically skewed. This is not fatal—the included papers look representative enough for a survey of this scope—but it is a load-bearing assumption for the word 'comprehensive,' and the authors should either add the missing databases or soften the claim. Minor: Fig. 3 is labeled PRIMA, and there are small typos like 'DAIZ-WOZ,' but those are trivial.\n\nThe citation pattern is fine. Some self-citations appear, but the cited works are peer-reviewed and sit alongside many third-party papers; no circularity concern.\n\nWho is it for: graduate students and researchers entering ADA, and anyone needing a modality map. It will not settle any scientific question, but it is a solid reference. I would send it to review, mainly because a corrected version is likely to be a heavily used survey.","headline":"A broad, useful ADA survey whose main overclaim—no existing interpretability work—contradicts its own citations, and whose search protocol needs one more pass before 'comprehensive' sticks.","tokens_in":47111,"tokens_out":1966,"would_cite":true,"duration_ms":23625,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims to be the first comprehensive survey of machine-learning-based automatic depression assessment, spanning EEG and MRI as well as facial, body, audio, and language behaviours.","keywords":["automatic depression assessment","machine learning","deep learning","multimodal behaviour analysis","EEG and MRI brain activity","facial and body behaviour","speech and language cues","ADA datasets and competitions"],"falsifier":"A reader could re-run the stated search across the same seven digital libraries (IEEE Explore, ACM Digital Library, Elsevier, Springer, ACL Anthology, IET, MIT Press) with the stated inclusion window of 2012-2024 and eligibility criteria, and check whether the 183 included papers match the yield; any significant set of qualifying peer-reviewed ADA papers that are missing would falsify the comprehensiveness claim.","tokens_in":46262,"feed_emoji":"🧠","tokens_out":7165,"duration_ms":76343,"temperature":0.7,"pith_summary":"This paper claims to be the first survey to treat automatic depression assessment (ADA) as one field spanning all depression-relevant behaviour modalities: internal brain activity from EEG and MRI, plus external facial, body, audio, and language cues. The authors reviewed 1679 candidate papers and included 183 methods published between 2012 and 2024, organising them by modality, feature type, and predictor, and linking each technical family to the psychological and physiological evidence for why that behaviour carries depression information. If the coverage is right, the survey gives researchers a single map of what has been tried, where the gaps are, and what is actually holding the field back—not model creativity, but small datasets with self-reported labels and unstandardised evaluation.","feed_headline":"183 studies map how AI reads depression from face, voice, brain","feed_subtitle":"A survey of 183 methods and 10 datasets shows depression is readable across brain, face, voice, body, and words—if data improves.","key_machinery":"The organising device is a modality-by-modality taxonomy (EEG, MRI, facial/head, body, audio, language, and multimodal), each reviewed with its typical feature extractors and predictors. The survey is carried by a PRISMA-style systematic search across seven digital libraries that yielded 1,679 papers and 183 included studies, plus an evidence review (Sec. 2.2) that connects each modality to concrete depression biomarkers such as EEG alpha asymmetry, reduced positive facial expressions, monotonous speech, and increased first-person pronouns. This evidence base is what justifies treating the surveyed models as capturing real depression signals rather than dataset artefacts.","core_discovery":"The central claim is that machine-learning-based depression assessment has become a genuinely multi-modal field, and that this is the first review to capture that breadth. The survey asserts that depression is expressed both internally, in brain structure and electrical activity, and externally, in the face, body, voice, and word choice, and that each modality has produced distinct technical pipelines—from hand-crafted features with classical regressors to CNNs, transformers, and graph networks. It further claims that the shared bottleneck across all modalities is data quality: public datasets are small, mostly under 200 participants, and depression labels come largely from self-report questionnaires rather than standardised clinical assessment, which limits both model performance and evaluation objectivity.","pith_inferences":["A testable extension of the survey's evidence review is whether behaviour-grounded hand-crafted features (e.g., spectral EEG features, AU 12/14 patterns, pause and pitch statistics, first-person pronouns) close the accuracy gap with end-to-end deep models on the same datasets.","If label subjectivity is the key data problem, then a shared collection protocol using clinician-administered instruments rather than self-report scales would be a concrete improvement; models trained on such data could be tested for cross-site and cross-cultural generalisation.","The paper's call for a common repository could be operationalised as a public leaderboard with fixed evaluation splits and fairness metrics, which would test whether reported accuracy differences between methods are stable or dataset-driven.","The geographic concentration of publications suggests that the field's comparative results may be fragile; an independent multi-site re-evaluation of top-performing methods would settle whether current rankings reflect generalisable skill or dataset bias."],"forward_implications":["The modality taxonomy and the 183-method table give researchers a direct way to spot underexplored areas, such as body-based ADA and privacy-preserving de-identified inputs.","If the data bottleneck is as central as the survey argues, then better datasets—larger, clinically labelled, demographically diverse—matter more for ADA progress than new network architectures.","The survey's review of competitions shows that de-identified facial primitives and audio features have been workable benchmarks since AVEC 2016, supporting privacy-preserving ADA as a practical default.","The documented lack of interpretability and reproducibility in deep ADA models indicates that future evaluation should routinely report which behaviour cues drive predictions.","The observation that most audio and language methods are fused with visual cues suggests that future ADA systems will be inherently multimodal, making fusion strategy and fairness across demographic groups core design issues."],"supporting_citations":[{"why":"Supplies the PRISMA reporting flow that structures the survey's search and selection method.","marker":"[146]"},{"why":"The AVEC 2013 depression challenge, the first public audio-visual ADA benchmark and the anchor of the field.","marker":"[59]"},{"why":"DAIC-WOZ, the de-identified audio-visual-text corpus used by the AVEC 2016-2019 baselines for depression assessment.","marker":"[64]"},{"why":"Cohn et al. 2009, the facial Action Unit study that grounds visual depression cues in AU patterns.","marker":"[117]"},{"why":"Cummins et al. 2015, the speech-analysis review that provides the audio modality's biomarker evidence.","marker":"[136]"},{"why":"Gupta 2009, the clinical source listing the physical signs of depression that justify behaviour-based ADA.","marker":"[17]"},{"why":"Nilsonne et al. 1988, provides the vocal and linguistic biomarker evidence for pitch and pronoun use.","marker":"[21]"},{"why":"Pennebaker's LIWC tool, the standard lexicon the language-based ADA approaches build on.","marker":"[22]"},{"why":"Lin et al. 2021, introduces the LaB body-gesture dataset used for body-modality ADA.","marker":"[67]"},{"why":"Yoon et al. 2022, the D-vlog naturalistic dataset, the largest public audio-visual ADA corpus.","marker":"[68]"}],"fun_headline_variants":["Depression detection AI spans face, voice, brain in 183-method survey","Survey: depression shows in brain, face, voice, and words—183 methods","AI depression screening: 183 ways to read face, voice, brain","Multi-modal depression AI: 183 methods, but data is the bottleneck","183 methods surveyed: AI reads depression from face, voice, brain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's claim to be comprehensive rests on the assumption that its literature search was exhaustive and that the 183 included papers fairly represent the 2012-2024 ADA literature; the full search queries and exclusion counts are only in the Supplementary Material, so a reader cannot independently reproduce the search.","fun_headline_variants_meta":{"raw":{"variants":["Depression detection AI spans face, voice, brain in 183-method survey","Survey: depression shows in brain, face, voice, and words—183 methods","AI depression screening: 183 ways to read face, voice, brain","Multi-modal depression AI: 183 methods, but data is the bottleneck","183 methods surveyed: AI reads depression from face, voice, brain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000966,"raw_usage":{"total_tokens":4092,"prompt_tokens":906,"completion_tokens":3186,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":3087}},"tokens_in":522,"tokens_out":3186,"duration_ms":27132,"temperature":1.0,"reasoning_tokens":3087,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:24:21.645718+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could re-run the stated search across the same seven digital libraries (IEEE Explore, ACM Digital Library, Elsevier, Springer, ACL Anthology, IET, MIT Press) with the stated inclusion window of 2012-2024 and eligibility criteria, and check whether the 183 included papers match the yield; any significant set of qualifying peer-reviewed ADA papers that are missing would falsify the comprehensiveness claim.","supporting_citations":[],"review_version":1}