Pith. sign in

REVIEW 2 major objections 4 minor 13 references

Non-social-media free-text mental health datasets are mostly English and depression-focused, leaving clear gaps for clinically useful NLP.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 01:44 UTC pith:7YEDVCN5

load-bearing objection Solid first PRISMA inventory of non-social-media free-text mental-health datasets; useful gap map, no load-bearing flaws. the 2 major comments →

arxiv 2607.03540 v1 pith:7YEDVCN5 submitted 2026-07-03 cs.CL

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

classification cs.CL
keywords language resourcesmental health disordersclinical NLPdatasetsdepressionPRISMAannotationnon-social media
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper presents the first systematic review of free-text datasets for mental health disorder detection that come from sources other than social media, such as clinical notes, interviews, discharge summaries, and forums. Using the PRISMA method the authors map 45 datasets across languages, platforms, data types, annotation methods and availability. They show that the resources are dominated by English-language material focused on depression, while other languages and conditions remain scarce. A reader should care because social-media data carry sampling biases, consent problems and limited demographic coverage, so better alternatives are needed for models that can actually support clinical care. The review therefore flags concrete opportunities to build more diverse, transparent and clinically grounded resources.

Core claim

A PRISMA review of non-social-media free-text mental-health datasets finds they are predominantly English and depression-oriented, vary widely in platform, data type, annotation instrument and access regime, and leave major gaps in language coverage, disorder breadth and clinical reliability.

What carries the argument

PRISMA systematic review protocol applied to free-text mental-health datasets collected outside social media, producing a structured inventory of 45 resources by language, disorder, platform, annotation method and availability.

Load-bearing premise

The English keyword search plus backtracking fully captures the relevant literature, so no large body of non-English or differently worded papers was missed.

What would settle it

Identification of a substantial collection of publicly available non-English free-text clinical datasets focused on underrepresented disorders such as schizophrenia or PTSD that existed before the review cutoff and were not recovered by the stated search terms or citation backtracking.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Releasing more public or agreement-governed clinical free-text datasets would expand research opportunities and improve reproducibility.
  • Creating resources in additional languages and for disorders beyond depression would make detection models more comprehensive and generalisable.
  • Standardised reporting of annotation instruments and inter-rater metrics would raise trust and allow reliable comparison across studies.
  • Instruction-tuned language models guided by formal diagnostic criteria could help scale consistent labelling even when clinicians are scarce.
  • Better ethical access to high-quality non-social datasets would facilitate collaborative clinical NLP work while protecting privacy.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The scarcity of non-English clinical free-text data is likely to slow development of culturally adapted detection tools, leaving non-English-speaking populations underserved.
  • Hybrid labelling that pairs self-report scales with clinician ratings may become the practical gold standard once automated prompting methods mature.
  • Higher citation rates for restricted clinical datasets suggest that perceived clinical credibility currently outweighs open access in driving research adoption.
  • Shared tasks built on multi-language non-social free-text collections could test whether models trained outside social media transfer more reliably to real clinical settings.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper presents the first PRISMA-guided systematic review of non-social-media free-text datasets for mental-health disorder detection. After multi-stage screening (keyword search yielding 543 papers, social-media exclusion to 259, novelty/modality/data-type filters to 57, plus backtracking to a final set of 45 datasets published 2004–2025), the authors catalog the resources in Table 1 and analyze distributions by language (62 % English), disorder (depression dominant), platform/data type, demographics, annotation instruments (DSM/ICD, PHQ/BDI/GAD, self-report, manual), availability, and citation impact. They answer three RQs on dataset variety, labeling clinical implications, and adoption factors, then highlight gaps (language/disorder imbalance, opaque labeling) and opportunities (LLM-assisted standardization).

Significance. The work fills a clear gap: prior surveys focus almost exclusively on social-media data, whose sampling, ethical, and demographic biases are well-known. By supplying a transparent PRISMA process, a public GitHub repository, a comprehensive Table 1, and a balanced discussion of clinical versus self-report labeling trade-offs, the paper supplies a practical map for clinical NLP researchers seeking more reliable resources. The descriptive claims (English- and depression-centric landscape, wide methodological heterogeneity) are directly supported by the tabulated evidence and are therefore immediately useful for dataset design, benchmarking, and funding priorities.

major comments (2)
  1. [§2 Methods / Figure 1] §2 (Methods) and Figure 1: The multi-stage screening (543 → 259 → 57 → 45) is described narratively, yet the manuscript never states whether title/abstract and full-text screening were performed independently by two reviewers or reports any inter-rater agreement statistic. Dual independent screening is a core PRISMA recommendation for reducing selection bias; its absence is load-bearing for the claim that the 45-dataset corpus is comprehensive and reproducible.
  2. [§3 / Table 1 / Figure 6] §3 and Table 1: Availability is coded PUB/DUA/RSTR/UNK for every dataset, and Figure 6 normalizes citations by year. However, the large UNK fraction (especially among English clinical datasets) is never quantified or sensitivity-tested. Because RQ3 explicitly links availability type to research adoption, the unexamined UNK category weakens the causal interpretation of the citation distributions.
minor comments (4)
  1. [§3 Figures 4–5] Figure 4 heatmap and Figure 5 stacked bars would benefit from absolute counts printed on each cell/bar; percentages alone make it hard to recover the raw tallies that underlie the “33 depression datasets” claim.
  2. [Table 1] Table 1 column “Annotation Instrument” mixes formal scales (PHQ-9, DSM-V) with informal phrases (“Key-word Search”, “Self-disclosure”). A short legend or secondary coding column would improve scannability.
  3. [Limitations] Limitations correctly notes English-only queries and missing CINAHL coverage; a single sentence quantifying how many of the final 45 were recovered solely via backtracking (rather than the primary keyword search) would further clarify residual risk of under-sampling.
  4. [Throughout] Minor typographic inconsistencies appear (e.g., “Sameto˘glu”, “MinDP” vs. “MinDP”, occasional missing spaces before citations). A final proof-reading pass is warranted.

Circularity Check

0 steps flagged

No circularity: empirical PRISMA survey of published datasets; findings are tallies, not derivations that reduce to inputs.

full rationale

This is a systematic literature review (PRISMA flowchart, keyword search + backtracking, multi-stage filtering to 45 free-text non-social-media mental-health datasets). The strongest claims (English/depression dominance, variation in platforms/annotation/availability, identified gaps) are direct empirical counts and distributions from Table 1 and Figures 2–8, not mathematical derivations, fitted parameters renamed as predictions, or uniqueness theorems. Self-citations (e.g., Bucur et al. 2025a,b on social-media surveys; Raihan et al.) appear only for motivation/comparison and do not load-bear or force the tabulated results. No self-definitional loops, no ansatz smuggled via prior work, no renaming of known results. The acknowledged English-keyword limitation is a coverage caveat, not circularity. Derivation chain is self-contained against the public table of datasets; score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

As a systematic review the paper rests on standard literature-search assumptions and the PRISMA protocol rather than free parameters or invented physical entities. The main load-bearing premises are the completeness of the chosen search strategy and the correctness of the inclusion criteria.

axioms (3)
  • domain assumption PRISMA 2020 guidelines produce a transparent and reproducible systematic review when followed.
    Invoked in §2 Methods as the reporting framework; standard in evidence synthesis.
  • ad hoc to paper Keyword search on ‘depression’, ‘anxiety’, ‘suicidal ideation’ plus detection terms, restricted to non-social-media sources, yields a representative sample of free-text mental-health datasets.
    Core search design in §2; authors themselves note possible missed non-English or atypical-term papers.
  • domain assumption Datasets labeled with DSM/ICD or validated scales (PHQ-9, HAMD, etc.) are clinically more reliable than pure self-disclosure or keyword labels.
    Used throughout §4 to rank annotation quality and discuss clinical implications.

pith-pipeline@v1.1.0-grok45 · 24411 in / 2062 out tokens · 18737 ms · 2026-07-12T01:44:47.618489+00:00 · methodology

0 comments
read the original abstract

Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detection rely on data collected primarily from social media. However, such datasets often have sampling biases and inherent ethical and privacy issues. One avenue to overcome these limitations is non-social media data. We present the first comprehensive review of non-social media, free-text datasets for mental health research. We use the PRISMA methodology to conduct our survey and we review datasets available in multiple languages. We find that non-social media free-text based datasets are predominantly focused on English and on detecting depression. These datasets also vary in demographics, platforms, data types, annotation techniques, and methodologies. This systematic review also reveals key gaps and highlights opportunities to develop more diverse, reliable and clinically-relevant resources.

Figures

Figures reproduced from arXiv: 2607.03540 by Ana-Maria Bucur, Marcos Zampieri, \"Ozlem Uzuner, Sadiya Sayara Chowdhury Puspo, Stevie Chancellor.

Figure 1
Figure 1. Figure 1: PRISMA flowchart illustrating the system [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Distribution of the datasets by year remained minimal and steady up to 2012. In 2014 and 2017, there was a noticeable rise in the num￾ber of proposed datasets, indicating intensified re￾search efforts in this domain during that period. The proposal of such datasets peaked in 2022, which can be attributed to greater attention to men￾tal health concerns after the COVID-19 pandemic (Bucur et al., 2025b). The … view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of datasets into by language [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Heatmap showing the distribution of datasets across language groups and disorders. Platform & Data Type Distribution [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Distribution of citation counts by dataset [PITH_FULL_IMAGE:figures/full_fig_p005_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Availability of datasets across language [PITH_FULL_IMAGE:figures/full_fig_p005_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Frequency of demographic attributes across datasets. Demographic attributes like gender and age domi￾nate, appearing in over 24 datasets, while others like education, race, and ethnicity are moderately included. Fewer datasets report attributes like in￾come, relationship status, or health indicators (e.g., height, weight, blood pressure). Some authors (Hiraga, 2017; Low et al., 2010) use demographic inform… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

13 extracted references · 4 linked inside Pith

  1. [1]

    Ankit Aich, Avery Quynh, Varsha Badal, Amy Pinkham, Philip Harvey, Colin Depp, and Na- talie Parde

    A corpus-based stylistic analysis of online suicide notes retrieved from reddit.Cogent Arts & Humanities, 9(1):2047434. Ankit Aich, Avery Quynh, Varsha Badal, Amy Pinkham, Philip Harvey, Colin Depp, and Na- talie Parde. 2022. Towards intelligent clinically- informed language analyses of people with bipo- lar disorder and schizophrenia. InFindings of EMNLP...

  2. [2]

    Dudley David Blake, Frank W Weathers, Linda M Nagy, Danny G Kaloupek, Fred D Gusman, Den- nis S Charney, and Terence M Keane

    Beck depression inventory–ii.Psychologi- cal assessment. Dudley David Blake, Frank W Weathers, Linda M Nagy, Danny G Kaloupek, Fred D Gusman, Den- nis S Charney, and Terence M Keane. 1995. The development of a clinician-administered ptsd scale.Journal of traumatic stress, 8:75–90. Nick Boettcher. 2021. Studies of depression and anxiety using reddit as a d...

  3. [3]

    https: //www.cdc.gov/mental-health/about-data/ suicidal-thoughts-and-behavior.html

    Suicidal thoughts and behavior. https: //www.cdc.gov/mental-health/about-data/ suicidal-thoughts-and-behavior.html . Ac- cessed 22 February 2026. Stevie Chancellor, Eric PS Baumer, and Munmun De Choudhury. 2019a. Who is the" human" in human-centered machine learning: The case of predicting mental health from social media. Proceedings of the ACM on Human-C...

  4. [4]

    NPJ digital medicine, 3(1):43

    Methods in predictive techniques for men- tal health status on social media: a critical review. NPJ digital medicine, 3(1):43. Moumita Chatterjee, Poulomi Samanta, Piyush Kumar, and Dhrubasish Sarkar. 2022. Suicide ideation detection using multiple feature analysis from twitter data. InIEEE DELCON. Glen Coppersmith, Casey Hilland, Ophir Frieder, and Ryan ...

  5. [5]

    Mario Cruz, Debra L Roter, Robyn F Cruz, Melissa Wieland, Susan Larson, Lisa A Cooper, and Harold Alan Pincus

    Detection of postnatal depression: de- velopment of the 10-item edinburgh postnatal depression scale.The British journal of psychia- try, 150(6):782–786. Mario Cruz, Debra L Roter, Robyn F Cruz, Melissa Wieland, Susan Larson, Lisa A Cooper, and Harold Alan Pincus. 2013. Appointment length, psychiatrists’ communication behaviors, and medication management ...

  6. [6]

    InProceedings of CLPsych

    On the state of social media data for men- tal health research. InProceedings of CLPsych. Qiwei He, Bernard P Veldkamp, Cees AW Glas, and Theo de Vries. 2017. Automated assess- ment of patients’ self-narratives for posttrau- matic stress disorder screening using natural language processing and text mining.Assess- ment, 24(2):157–172. Ronald D Hester. 2017...

  7. [7]

    Institute for Health Metrics and Evaluation

    Two-way messaging therapy for depres- sion and anxiety: longitudinal response trajecto- ries.BMC psychiatry, 20(1):297. Institute for Health Metrics and Evaluation. 2024. Global burden of disease study 2021 (gbd

  8. [8]

    https://vizhub.healthdata

    results. https://vizhub.healthdata. org/gbd-results/. Online database. Seattle, WA. Accessed 13 August 2025. Md Rafiqul Islam, Muhammad Ashad Kabir, Ashir Ahmed, Abu Raihan M Kamal, Hua Wang, and Anwaar Ulhaq. 2018. Depression detection from social network data using machine learning tech- niques.Health information science and systems, 6:1–12. Julia Ive, ...

  9. [9]

    Kurt Kroenke, Robert L Spitzer, and Janet BW Williams

    Identification of maternal depression risk from natural language collected in a mo- bile health app.Procedia computer science, 206:132–140. Kurt Kroenke, Robert L Spitzer, and Janet BW Williams. 2001. The phq-9: validity of a brief depression severity measure.Journal of general internal medicine, 16(9):606–613. Kurt Kroenke, Tara W Strine, Robert L Spitze...

  10. [10]

    In Proceedings of CHI

    Identification of imminent suicide risk among young adults using text messages. In Proceedings of CHI. Jihoon Oh, Taekgyu Lee, Eun Su Chung, Hyon- soo Kim, Kyongchul Cho, Hyunkyu Kim, Jihye Choi, Hyeon-Hee Sim, Jongseo Lee, In Y oung Choi, et al. 2024. Development of depression detection algorithm using text scripts of routine psychiatric interview.Fronti...

  11. [11]

    Biological Psychiatry, 54(5):573–583

    The 16-item quick inventory of depressive symptomatology (qids), clinician rating (qids-c), and self-report (qids-sr): a psychometric evalu- ation in patients with chronic major depression. Biological Psychiatry, 54(5):573–583. Esteban Andrés Ríssola, Mohammad Aliannejadi, and Fabio Crestani. 2020. Beyond modelling: un- derstanding mental disorders in onl...

  12. [12]

    Daun Shin, Kyungdo Kim, Seung-Bo Lee, Chang- woo Lee, Y e Seul Bae, Won Ik Cho, Min Ji Kim, C Hyung Keun Park, Eui Kyu Chie, Nam Soo Kim, et al

    Enhancing depression diagnosis with chain-of-thought prompting.arXiv preprint arXiv:2408.14053. Daun Shin, Kyungdo Kim, Seung-Bo Lee, Chang- woo Lee, Y e Seul Bae, Won Ik Cho, Min Ji Kim, C Hyung Keun Park, Eui Kyu Chie, Nam Soo Kim, et al. 2022a. Detection of depression and suicide risk based on text from clinical interviews using machine learning: possi...

  13. [13]

    Y evhen Tyshchenko

    A narrative review of the assessment of depression in chronic pain.Pain Management Nursing, 23(2):158–167. Y evhen Tyshchenko. 2018. Depression and anxiety detection from blog posts data.Nature Precis. Sci., Inst. Comput. Sci., Univ. Tartu, Tartu, Esto- nia, pages 6–46. Rudolf Uher, Roy H Perlis, Anna Placentino, Mo- jca Zvezdana Dernovšek, Neven Henigsbe...