Pith. sign in

REVIEW 3 major objections 25 references

People with higher depressive symptoms use ChatGPT more for late-night, recurring mental-health support, but professional redirection does not rise with need.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 03:47 UTC pith:C4A4QE2I

load-bearing objection First large PHQ-linked ChatGPT history study with careful stats and a clear non-clinical framing; the boundary-gap claim leans on unvalidated LLM labels, but the timing/lexical core holds. the 3 major comments →

arxiv 2607.05685 v1 pith:C4A4QE2I submitted 2026-07-06 cs.HC cs.AI

Depression Symptoms and Relational Patterns in 187k ChatGPT Histories

classification cs.HC cs.AI
keywords ChatGPTdepressive symptomsPHQ-8informal support infrastructuredisclosureprofessional redirectionlate-night uselanguage-based prediction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper links Patient Health Questionnaire-8 scores from 766 participants to 187,093 donated ChatGPT conversation histories and treats ChatGPT as informal support infrastructure rather than a clinical tool. Participants at or above the moderate-symptom threshold (PHQ ≥ 10) brought more mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, used ChatGPT more between 23:00 and 04:59, and showed recurring month-level patterns of those topics. Their language contained more first-person singular pronouns and absolutist words, and they entered higher-disclosure contexts more often, yet professional redirection rates stayed essentially flat across groups. Language-only models predicting the PHQ split reached only modest performance (best AUROC 0.591), which the authors judge insufficient for screening. The claim is that private LLM histories should be read as evidence of how people already use always-on systems for support, not as clinical data for triage.

Core claim

Higher-PHQ participants used ChatGPT differently in relational and temporal ways that matter for design: more mental-health and interpersonal content, more late-night and recurring use, more self-focused language and high-disclosure support seeking, while professional redirection did not increase robustly with apparent need, and language-based prediction of the PHQ split remained too weak for screening.

What carries the argument

The PHQ ≥ 10 versus PHQ < 10 split applied to participant-weighted conversation histories (topics, disclosure, support seeking, nocturnal and month-level recurrence, lexical markers, and response-side professional redirection), used to contrast how symptom-severity groups engage ChatGPT as informal support infrastructure.

Load-bearing premise

The analysis rests on unvalidated gpt-4o-mini labels for topics, disclosure, support seeking, and professional redirection being accurate enough for group contrasts even though they cover only targeted subsets and have no human-rater check.

What would settle it

If human-coded disclosure and professional-redirection labels on the same Health/Mental Health subset reverse or erase the PHQ-group differences, or if a language-only model on held-out histories reaches clinically usable screening performance, the paper’s central interpretation fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. This paper links PHQ-8 symptom scores from 766 Prolific participants to 187,093 donated ChatGPT conversation histories, comparing PHQ<10 vs PHQ≥10 groups. Higher-PHQ participants show higher shares of mental-health, interpersonal, loneliness, and negative self-focus conversations; more late-night and recurring month-level use; elevated first-person singular and absolutist language; and more high-disclosure/support-seeking contexts, while professional redirection rates are essentially flat (~16%). Language-only prediction of PHQ≥10 is modest (best AUROC 0.591) and framed as insufficient for screening. The authors position ChatGPT as informal, always-available support infrastructure rather than a clinical tool, with design implications around history-aware boundaries and redirection.

Significance. The contribution is timely and substantial for CSCW/HCI: a large, survey-linked corpus of private LLM histories is rare, and the work carefully separates descriptive use patterns from clinical screening claims. Strengths include participant-weighted contrasts with FDR control, conversation-weighted clustered adjusted models, recency-window sensitivities (14–90 days), health vs non-health subset checks, and an explicit negative result on screening utility. If the main contrasts hold under stronger label validation, the paper would be a durable empirical reference for how depressive-symptom severity co-occurs with relational LLM use and for design debates about professional boundaries in always-on chat systems.

major comments (3)
  1. Finding 3 and the design-implication claim that professional redirection does not scale with need rest almost entirely on unvalidated gpt-4o-mini labels for disclosure level, support-seeking, and professional_redirect (Methods; Appendix B.2.3–B.2.4; Table 1). Temperature-0 prompting and “exploratory” framing do not substitute for human agreement. Please report inter-rater reliability (or model–human agreement) on a stratified sample of Health/Mental Health turns, and show that PHQ-group differences survive when restricted to high-agreement items. Without this, the boundary-gap result remains the softest load-bearing claim.
  2. Finding 1’s headline mental-health, interpersonal, loneliness, and negative self-focus shares likewise depend on the same unvalidated conversation-level GPT taxonomy (Appendix B.2.1–B.2.2; Table 1). Deterministic markers (nocturnal share, first-person pronouns, absolutist words) and the AUROC result are independent of those labels and should be presented as the primary evidence backbone. Please either validate topic/construct labels on a human-coded subset or restructure Results so LLM-dependent vs deterministic claims are clearly separated and weighted accordingly.
  3. §3 and Appendix C.4: PHQ-8 indexes the prior two weeks, while primary contrasts use full exported histories. Recency sensitivities are a strength, but several lexical and advice/information contrasts weaken or lose significance in the 14–30 day windows (Appendix Table 5). The manuscript should state more explicitly which claims remain robust under the survey-anchored 14-day window and avoid treating full-history LLM-label shares as interchangeable with “recent symptom severity” without that qualification.

Circularity Check

0 steps flagged

No circularity: PHQ groups are external survey measures; chat features and modest out-of-fold AUROC are independent contrasts, not self-defined predictions.

full rationale

This is an observational CSCW/HCI study that splits participants by an external PHQ-8 survey threshold and compares usage, timing, lexical rates, and exploratory LLM-derived conversation labels. The PHQ split is not defined from chat language or annotations, so group contrasts (mental-health share, nocturnal use, first-person pronouns, disclosure, professional redirect) are not true by construction. Language-only PHQ prediction uses regularized logistic regression with repeated stratified out-of-fold CV and reports a modest AUROC (0.591) that the authors explicitly treat as insufficient for screening—not a fitted parameter renamed as a first-principles prediction. LLM labels (gpt-4o-mini) are framed as unvalidated exploratory research annotations for aggregate characterization, not as the definition of the outcome or of the PHQ groups. Self-citations to the authors’ related chatbot-design work appear in related work and acknowledgments but do not supply a uniqueness theorem, ansatz, or load-bearing premise that forces the empirical results. No equation or claim reduces by definition to its own inputs. Circularity score is therefore 0.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 1 invented entities

The central claims rest on standard clinical thresholding, unvalidated but temperature-0 LLM annotations, and a convenience sample of ChatGPT users willing to donate histories. No new physical entities are postulated; free parameters are mainly analysis choices (threshold, temperature, truncations) rather than fitted scientific constants. Domain assumptions about PHQ validity and annotation utility are the main load-bearing premises.

free parameters (3)
  • PHQ-8 moderate threshold = 10
    Binary split at score 10 follows Kroenke et al. convention; results depend on this cut rather than continuous modeling as primary.
  • gpt-4o-mini temperature = 0
    Fixed at 0 for all annotations; deterministic but still model-dependent labels.
  • user-text truncation lengths = 1500/2400/1000 chars
    1,500 / 2,400 / 1,000 character cutoffs for different annotation streams; ad-hoc engineering choices that could affect label coverage.
axioms (4)
  • domain assumption PHQ-8 score ≥10 indexes moderate-or-greater recent depressive symptom severity and is a valid grouping variable for exported ChatGPT histories.
    Invoked throughout Methods and Findings; PHQ covers prior two weeks while histories span months; authors note this but still treat the split as primary.
  • ad hoc to paper gpt-4o-mini labels for topics, disclosure, support-seeking, professional redirection, and sycophancy items are useful exploratory research annotations for aggregate group contrasts.
    Appendix B; authors explicitly call them unvalidated and non-clinical, yet all headline disclosure/redirection findings depend on them.
  • standard math Participant-weighted contrasts with FDR within analysis families adequately control multiple testing and volume bias.
    Methods §Analysis and Appendix C; standard statistical practice.
  • domain assumption Prolific-recruited ChatGPT users in US/UK/Canada who donate histories are informative for the claimed informal-support-infrastructure pattern, even if not representative of all users.
    Limitations section; convenience sample with eligibility screens for prior ChatGPT use and study volume.
invented entities (1)
  • ChatGPT as informal support infrastructure no independent evidence
    purpose: Conceptual framing that organizes the empirical patterns (privacy, persistence, after-hours availability, one-on-one nonhuman support) without treating the system as clinical.
    Introduced in Abstract/Introduction as the interpretive lens; not a physical entity but a design concept whose independent evidence is the observed usage patterns themselves.

pith-pipeline@v1.1.0-grok45 · 20588 in / 3136 out tokens · 32261 ms · 2026-07-11T03:47:36.207055+00:00 · methodology

0 comments
read the original abstract

Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with depressive symptoms use them. Building on CSCW work on disclosure and peer support, we examine ChatGPT as an emerging informal support infrastructure: private, persistent, responsive, and available outside ordinary hours. We analyze 187,093 ChatGPT conversations from 766 participants who completed the PHQ-8, comparing those below the moderate-symptom threshold (score of 10) with those at or above it. Higher-PHQ participants used ChatGPT more for mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, with pronounced late-night and recurring month-level patterns. Their language contained more first-person singular pronouns and absolutist terms. They more often engaged ChatGPT in high-disclosure contexts, but professional redirection was not higher. Language-based prediction was modest and insufficient for screening (AUROC 0.591). We argue these histories should not be treated as clinical screening data but as evidence LLMs are increasingly used as informal support infrastructure.

Figures

Figures reproduced from arXiv: 2607.05685 by Dunigan Folk, Lyle Ungar, Neil K. R. Sehgal, Sharath Chandra Guntuku.

Figure 1
Figure 1. Figure 1: Participant-level distributions for key PHQ markers. Violin shapes show full distribution, boxes show median and interquartile [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Differential language analysis and language-only PHQ prediction. Left panel shows strongest combined user + ChatGPT LIWC [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 25 canonical work pages · 2 internal anchors

  1. [1]

    Mohammed Al-Mosaiwi and Tom Johnstone. 2018. In an absolute state: Elevated use of absolutist words is a marker specific to anxiety, depression, and suicidal ideation.Clinical psychological science6, 4 (2018), 529–542

  2. [2]

    Nazanin Andalibi. 2020. Disclosure, Privacy, and Stigma on Social Media: Examining Non-disclosure of Distressing Experiences.ACM Transactions on Computer-Human Interaction27, 3 (2020). doi:10.1145/3386600

  3. [3]

    Haimson, Munmun De Choudhury, and Andrea Forte

    Nazanin Andalibi, Oliver L. Haimson, Munmun De Choudhury, and Andrea Forte. 2018. Social Support, Reciprocity, and Anonymity in Responses to Sexual Abuse Disclosures on Social Media.ACM Transactions on Computer-Human Interaction25, 5, Article 28 (2018). doi:10.1145/3234942

  4. [4]

    Nazanin Andalibi and Pinar Öztürk. 2017. Sensitive Self-disclosures, Responses, and Social Support on Instagram: The Case of Depression. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing. Association for Computing Machinery. doi:10.1145/2998181.2998243

  5. [5]

    2025.How people use chatgpt

    Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025.How people use chatgpt. Technical Report. National Bureau of Economic Research

  6. [6]

    Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, and Dan Jurafsky. 2026. Sycophantic AI decreases prosocial intentions and promotes dependence.Science391, 6792 (2026), eaec8352

  7. [7]

    Munmun De Choudhury, Michael Gamon, Scott Counts, and Eric Horvitz. 2013. Predicting depression via social media. InProceedings of the international AAAI conference on web and social media, Vol. 7. 128–137

  8. [8]

    Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng, Cinoo Lee, Rebecca Anselmetti, Robb Willer, Luc Rocher, and Diyi Yang. 2026. Sycophantic AI makes human interaction feel more effortful and less satisfying over time. arXiv:2605.07912 [cs.HC] https://arxiv.org/abs/2605.07912

  9. [9]

    Yucheng Jin, Wanling Cai, Li Chen, Yuwan Dai, and Tonglin Jiang. 2023. Understanding disclosure and support for youth mental health in social music communities.Proceedings of the ACM on human-Computer Interaction7, CSCW1 (2023), 1–32

  10. [10]

    Kyuha Jung, Gyuho Lee, Yuanhui Huang, and Yunan Chen. 2025. ‘I’ve talked to ChatGPT about my issues last night. ’: Examining Mental Health Conversations with Large Language Models through Reddit Analysis.Proceedings of the ACM on Human-Computer Interaction9, 7 (Oct. 2025), 1–25. doi:10.1145/3757537

  11. [11]

    Gummadi, Animesh Mukherjee, Ingmar Weber, and Savvas Zannettou

    Sai Keerthana Karnam, Abhisek Dash, Krishna P. Gummadi, Animesh Mukherjee, Ingmar Weber, and Savvas Zannettou. 2026. Bowling with ChatGPT: On the Evolving User Interactions with Conversational AI Systems. arXiv:2602.01114 [cs.HC] https://arxiv.org/abs/2602.01114

  12. [12]

    Kurt Kroenke, Tara W Strine, Robert L Spitzer, Janet BW Williams, Joyce T Berry, and Ali H Mokdad. 2009. The PHQ-8 as a measure of current depression in the general population.Journal of affective disorders114, 1-3 (2009), 163–173

  13. [13]

    Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024. The opportunities and risks of large language models in mental health.JMIR Mental Health11, 1 (2024), e59479

  14. [14]

    Tingting Liu, Lyle H Ungar, Brenda Curtis, Garrick Sherman, Kenna Yadeta, Louis Tay, Johannes C Eichstaedt, and Sharath Chandra Guntuku. 2022. Head versus heart: social media reveals differential language of loneliness from depression.Npj Mental Health Research1, 1 (2022), 16

  15. [15]

    Clifford Nass, Jonathan Steuer, and Ellen R. Tauber. 1994. Computers are social actors.Proceedings of the SIGCHI Conference on Human Factors in Computing Systems(1994). https://api.semanticscholar.org/CorpusID:2739302

  16. [16]

    Like Shock Absorbers

    Sachin R. Pendse, Faisal M. Lalani, Munmun De Choudhury, Amit Sharma, and Neha Kumar. 2020. “Like Shock Absorbers”: Understanding the Human Infrastructures of Technology-Mediated Mental Health Support. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3313831.3376465

  17. [17]

    Can I Not Be Suicidal on a Sunday?

    Sachin R. Pendse, Amit Sharma, Aditya Vashistha, Munmun De Choudhury, and Neha Kumar. 2021. “Can I Not Be Suicidal on a Sunday?”: Understanding Technology-Mediated Pathways to Mental Health Support. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. doi:10.1145/3411764.3445410

  18. [18]

    Sunny Rai, Elizabeth C Stade, Salvatore Giorgi, Ashley Francisco, Lyle H Ungar, Brenda Curtis, and Sharath C Guntuku. 2024. Key language markers of depression on social media depend on race.Proceedings of the National Academy of Sciences121, 14 (2024), e2319837121

  19. [19]

    Jean Rehani, Victoria Oldemburgo de Mello, Dariya Ovsyannikova, Ashton Anderson, and Michael Inzlicht. 2026. The Social Sycophancy Scale: A psychometrically validated measure of sycophancy. arXiv:2603.15448 [cs.HC] https://arxiv.org/abs/2603.15448

  20. [20]

    H Andrew Schwartz, Johannes Eichstaedt, Margaret Kern, Gregory Park, Maarten Sap, David Stillwell, Michal Kosinski, and Lyle Ungar. 2014. Towards assessing changes in degree of depression through facebook. InProceedings of the workshop on computational linguistics and clinical psychology: from linguistic signal to clinical reality. 118–125. Manuscript sub...

  21. [21]

    Neil KR Sehgal, Hita Kambhamettu, Sai Preethi Matam, Lyle Ungar, and Sharath Chandra Guntuku. 2025. Designing Mental-Health Chatbots for Indian Adolescents: Mixed-Methods Evidence, a Boundary-Object Lens, and a Design-Tensions Framework.arXiv preprint arXiv:2511.07729(2025)

  22. [22]

    Neil KR Sehgal, Hita Kambhamettu, Sai Preethi Matam, Lyle Ungar, and Sharath Chandra Guntuku. 2025. Exploring Socio-Cultural challenges and opportunities in designing mental health chatbots for adolescents in India. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–7

  23. [23]

    The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support

    Inhwa Song, Sachin R. Pendse, Neha Kumar, and Munmun De Choudhury. 2025. The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support. arXiv:2401.14362 [cs.HC] https://arxiv.org/abs/2401.14362

  24. [24]

    Mina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, and Natalie Parde. 2021. Identifying medical self-disclosure in online communities. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4398–4408

  25. [25]

    I have never used ChatGPT / I do not have a ChatGPT account

    Diyi Yang, Zheng Yao, Joseph Seering, and Robert Kraut. 2019. The channel matters: Self-disclosure, reciprocity and social support in online cancer support groups. InProceedings of the 2019 chi conference on human factors in computing systems. 1–15. A Sample, Recruitment, and Eligibility A.1 Participant Demographics Appendix Table 1. Participant and datas...