REVIEW 3 major objections 25 references
People with higher depressive symptoms use ChatGPT more for late-night, recurring mental-health support, but professional redirection does not rise with need.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 03:47 UTC pith:C4A4QE2I
load-bearing objection First large PHQ-linked ChatGPT history study with careful stats and a clear non-clinical framing; the boundary-gap claim leans on unvalidated LLM labels, but the timing/lexical core holds. the 3 major comments →
Depression Symptoms and Relational Patterns in 187k ChatGPT Histories
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Higher-PHQ participants used ChatGPT differently in relational and temporal ways that matter for design: more mental-health and interpersonal content, more late-night and recurring use, more self-focused language and high-disclosure support seeking, while professional redirection did not increase robustly with apparent need, and language-based prediction of the PHQ split remained too weak for screening.
What carries the argument
The PHQ ≥ 10 versus PHQ < 10 split applied to participant-weighted conversation histories (topics, disclosure, support seeking, nocturnal and month-level recurrence, lexical markers, and response-side professional redirection), used to contrast how symptom-severity groups engage ChatGPT as informal support infrastructure.
Load-bearing premise
The analysis rests on unvalidated gpt-4o-mini labels for topics, disclosure, support seeking, and professional redirection being accurate enough for group contrasts even though they cover only targeted subsets and have no human-rater check.
What would settle it
If human-coded disclosure and professional-redirection labels on the same Health/Mental Health subset reverse or erase the PHQ-group differences, or if a language-only model on held-out histories reaches clinically usable screening performance, the paper’s central interpretation fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper links PHQ-8 symptom scores from 766 Prolific participants to 187,093 donated ChatGPT conversation histories, comparing PHQ<10 vs PHQ≥10 groups. Higher-PHQ participants show higher shares of mental-health, interpersonal, loneliness, and negative self-focus conversations; more late-night and recurring month-level use; elevated first-person singular and absolutist language; and more high-disclosure/support-seeking contexts, while professional redirection rates are essentially flat (~16%). Language-only prediction of PHQ≥10 is modest (best AUROC 0.591) and framed as insufficient for screening. The authors position ChatGPT as informal, always-available support infrastructure rather than a clinical tool, with design implications around history-aware boundaries and redirection.
Significance. The contribution is timely and substantial for CSCW/HCI: a large, survey-linked corpus of private LLM histories is rare, and the work carefully separates descriptive use patterns from clinical screening claims. Strengths include participant-weighted contrasts with FDR control, conversation-weighted clustered adjusted models, recency-window sensitivities (14–90 days), health vs non-health subset checks, and an explicit negative result on screening utility. If the main contrasts hold under stronger label validation, the paper would be a durable empirical reference for how depressive-symptom severity co-occurs with relational LLM use and for design debates about professional boundaries in always-on chat systems.
major comments (3)
- Finding 3 and the design-implication claim that professional redirection does not scale with need rest almost entirely on unvalidated gpt-4o-mini labels for disclosure level, support-seeking, and professional_redirect (Methods; Appendix B.2.3–B.2.4; Table 1). Temperature-0 prompting and “exploratory” framing do not substitute for human agreement. Please report inter-rater reliability (or model–human agreement) on a stratified sample of Health/Mental Health turns, and show that PHQ-group differences survive when restricted to high-agreement items. Without this, the boundary-gap result remains the softest load-bearing claim.
- Finding 1’s headline mental-health, interpersonal, loneliness, and negative self-focus shares likewise depend on the same unvalidated conversation-level GPT taxonomy (Appendix B.2.1–B.2.2; Table 1). Deterministic markers (nocturnal share, first-person pronouns, absolutist words) and the AUROC result are independent of those labels and should be presented as the primary evidence backbone. Please either validate topic/construct labels on a human-coded subset or restructure Results so LLM-dependent vs deterministic claims are clearly separated and weighted accordingly.
- §3 and Appendix C.4: PHQ-8 indexes the prior two weeks, while primary contrasts use full exported histories. Recency sensitivities are a strength, but several lexical and advice/information contrasts weaken or lose significance in the 14–30 day windows (Appendix Table 5). The manuscript should state more explicitly which claims remain robust under the survey-anchored 14-day window and avoid treating full-history LLM-label shares as interchangeable with “recent symptom severity” without that qualification.
Circularity Check
No circularity: PHQ groups are external survey measures; chat features and modest out-of-fold AUROC are independent contrasts, not self-defined predictions.
full rationale
This is an observational CSCW/HCI study that splits participants by an external PHQ-8 survey threshold and compares usage, timing, lexical rates, and exploratory LLM-derived conversation labels. The PHQ split is not defined from chat language or annotations, so group contrasts (mental-health share, nocturnal use, first-person pronouns, disclosure, professional redirect) are not true by construction. Language-only PHQ prediction uses regularized logistic regression with repeated stratified out-of-fold CV and reports a modest AUROC (0.591) that the authors explicitly treat as insufficient for screening—not a fitted parameter renamed as a first-principles prediction. LLM labels (gpt-4o-mini) are framed as unvalidated exploratory research annotations for aggregate characterization, not as the definition of the outcome or of the PHQ groups. Self-citations to the authors’ related chatbot-design work appear in related work and acknowledgments but do not supply a uniqueness theorem, ansatz, or load-bearing premise that forces the empirical results. No equation or claim reduces by definition to its own inputs. Circularity score is therefore 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- PHQ-8 moderate threshold =
10
- gpt-4o-mini temperature =
0
- user-text truncation lengths =
1500/2400/1000 chars
axioms (4)
- domain assumption PHQ-8 score ≥10 indexes moderate-or-greater recent depressive symptom severity and is a valid grouping variable for exported ChatGPT histories.
- ad hoc to paper gpt-4o-mini labels for topics, disclosure, support-seeking, professional redirection, and sycophancy items are useful exploratory research annotations for aggregate group contrasts.
- standard math Participant-weighted contrasts with FDR within analysis families adequately control multiple testing and volume bias.
- domain assumption Prolific-recruited ChatGPT users in US/UK/Canada who donate histories are informative for the claimed informal-support-infrastructure pattern, even if not representative of all users.
invented entities (1)
-
ChatGPT as informal support infrastructure
no independent evidence
read the original abstract
Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them. Building on CSCW work on disclosure and peer support, we examine ChatGPT as an emerging informal support infrastructure: private, persistent, responsive, and available outside ordinary hours. We analyze 187,093 ChatGPT conversations from 766 participants who completed the PHQ-8, comparing those below the moderate-symptom threshold (score of 10) with those at or above it. Higher-PHQ participants used ChatGPT more for mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, with pronounced late-night and recurring month-level patterns. Their language contained more first-person singular pronouns and absolutist terms. They more often engaged ChatGPT in high-disclosure contexts, but professional redirection was not higher. Language-based prediction was modest and insufficient for screening (AUROC 0.591). We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure.
Figures
Reference graph
Works this paper leans on
-
[1]
Mohammed Al-Mosaiwi and Tom Johnstone. 2018. In an absolute state: Elevated use of absolutist words is a marker specific to anxiety, depression, and suicidal ideation.Clinical psychological science6, 4 (2018), 529–542
work page 2018
-
[2]
Nazanin Andalibi. 2020. Disclosure, Privacy, and Stigma on Social Media: Examining Non-disclosure of Distressing Experiences.ACM Transactions on Computer-Human Interaction27, 3 (2020). doi:10.1145/3386600
-
[3]
Haimson, Munmun De Choudhury, and Andrea Forte
Nazanin Andalibi, Oliver L. Haimson, Munmun De Choudhury, and Andrea Forte. 2018. Social Support, Reciprocity, and Anonymity in Responses to Sexual Abuse Disclosures on Social Media.ACM Transactions on Computer-Human Interaction25, 5, Article 28 (2018). doi:10.1145/3234942
-
[4]
Nazanin Andalibi and Pinar Öztürk. 2017. Sensitive Self-disclosures, Responses, and Social Support on Instagram: The Case of Depression. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. Association for Computing Machinery. doi:10.1145/2998181.2998243
-
[5]
Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025.How people use chatgpt. Technical Report. National Bureau of Economic Research
work page 2025
-
[6]
Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. 2026. Sycophantic AI decreases prosocial intentions and promotes dependence.Science391, 6792 (2026), eaec8352
work page 2026
-
[7]
Munmun De Choudhury, Michael Gamon, Scott Counts, and Eric Horvitz. 2013. Predicting depression via social media. InProceedings of the international AAAI conference on web and social media, Vol. 7. 128–137
work page 2013
-
[8]
Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng, Cinoo Lee, Rebecca Anselmetti, Robb Willer, Luc Rocher, and Diyi Yang. 2026. Sycophantic AI makes human interaction feel more effortful and less satisfying over time. arXiv:2605.07912 [cs.HC] https://arxiv.org/abs/2605.07912
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[9]
Yucheng Jin, Wanling Cai, Li Chen, Yuwan Dai, and Tonglin Jiang. 2023. Understanding disclosure and support for youth mental health in social music communities.Proceedings of the ACM on human-Computer Interaction7, CSCW1 (2023), 1–32
work page 2023
-
[10]
Kyuha Jung, Gyuho Lee, Yuanhui Huang, and Yunan Chen. 2025. ‘I’ve talked to ChatGPT about my issues last night. ’: Examining Mental Health Conversations with Large Language Models through Reddit Analysis.Proceedings of the ACM on Human-Computer Interaction9, 7 (Oct. 2025), 1–25. doi:10.1145/3757537
-
[11]
Gummadi, Animesh Mukherjee, Ingmar Weber, and Savvas Zannettou
Sai Keerthana Karnam, Abhisek Dash, Krishna P. Gummadi, Animesh Mukherjee, Ingmar Weber, and Savvas Zannettou. 2026. Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI Systems. arXiv:2602.01114 [cs.HC] https://arxiv.org/abs/2602.01114
-
[12]
Kurt Kroenke, Tara W Strine, Robert L Spitzer, Janet BW Williams, Joyce T Berry, and Ali H Mokdad. 2009. The PHQ-8 as a measure of current depression in the general population.Journal of affective disorders114, 1-3 (2009), 163–173
work page 2009
-
[13]
Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024. The opportunities and risks of large language models in mental health.JMIR Mental Health11, 1 (2024), e59479
work page 2024
-
[14]
Tingting Liu, Lyle H Ungar, Brenda Curtis, Garrick Sherman, Kenna Yadeta, Louis Tay, Johannes C Eichstaedt, and Sharath Chandra Guntuku. 2022. Head versus heart: social media reveals differential language of loneliness from depression.Npj Mental Health Research1, 1 (2022), 16
work page 2022
-
[15]
Clifford Nass, Jonathan Steuer, and Ellen R. Tauber. 1994. Computers are social actors.Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(1994). https://api.semanticscholar.org/CorpusID:2739302
work page 1994
-
[16]
Sachin R. Pendse, Faisal M. Lalani, Munmun De Choudhury, Amit Sharma, and Neha Kumar. 2020. “Like Shock Absorbers”: Understanding the Human Infrastructures of Technology-Mediated Mental Health Support. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3313831.3376465
-
[17]
Can I Not Be Suicidal on a Sunday?
Sachin R. Pendse, Amit Sharma, Aditya Vashistha, Munmun De Choudhury, and Neha Kumar. 2021. “Can I Not Be Suicidal on a Sunday?”: Understanding Technology-Mediated Pathways to Mental Health Support. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3411764.3445410
-
[18]
Sunny Rai, Elizabeth C Stade, Salvatore Giorgi, Ashley Francisco, Lyle H Ungar, Brenda Curtis, and Sharath C Guntuku. 2024. Key language markers of depression on social media depend on race.Proceedings of the National Academy of Sciences121, 14 (2024), e2319837121
work page 2024
- [19]
-
[20]
H Andrew Schwartz, Johannes Eichstaedt, Margaret Kern, Gregory Park, Maarten Sap, David Stillwell, Michal Kosinski, and Lyle Ungar. 2014. Towards assessing changes in degree of depression through facebook. InProceedings of the workshop on computational linguistics and clinical psychology: from linguistic signal to clinical reality. 118–125. Manuscript sub...
work page 2014
- [21]
-
[22]
Neil KR Sehgal, Hita Kambhamettu, Sai Preethi Matam, Lyle Ungar, and Sharath Chandra Guntuku. 2025. Exploring Socio-Cultural challenges and opportunities in designing mental health chatbots for adolescents in India. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–7
work page 2025
-
[23]
The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support
Inhwa Song, Sachin R. Pendse, Neha Kumar, and Munmun De Choudhury. 2025. The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support. arXiv:2401.14362 [cs.HC] https://arxiv.org/abs/2401.14362
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[24]
Mina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, and Natalie Parde. 2021. Identifying medical self-disclosure in online communities. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4398–4408
work page 2021
-
[25]
I have never used ChatGPT / I do not have a ChatGPT account
Diyi Yang, Zheng Yao, Joseph Seering, and Robert Kraut. 2019. The channel matters: Self-disclosure, reciprocity and social support in online cancer support groups. InProceedings of the 2019 chi conference on human factors in computing systems. 1–15. A Sample, Recruitment, and Eligibility A.1 Participant Demographics Appendix Table 1. Participant and datas...
work page 2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.