{"id":"f764ca7a-0b0d-4804-8958-ea224d053776","arxiv_id":"2607.05685","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Higher-PHQ users bring more mental-health, relational, late-night, and high-disclosure concerns to ChatGPT without higher professional redirection, and language prediction is too weak for screening (AUROC 0.591).","lead":"People with higher depressive symptoms use ChatGPT more for mental-health, loneliness, and support-seeking talk, especially late at night, with more self-focused language and disclosure. Professional redirection does not rise with need, and language alone cannot screen symptoms, so the paper treats chatbots as informal support infrastructure rather than clinical tools.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Unvalidated gpt-4o-mini labels for disclosure, support-seeking and professional redirect remain the softest link for the boundary-gap claim in Finding 3.","rationale":"The reader correctly isolates the unvalidated LLM annotations as the weakest assumption supporting the strongest non-lexical claims. Deterministic and recency-window checks already shore up the usage, timing and language-marker results, so the paper’s observational core survives; the conditional verdict is therefore appropriate and needs no further downgrade. A human-validation check of the exact labels that drive Finding 3 is the single most decisive next measurement.","tokens_in":16446,"tokens_out":447,"duration_ms":15806,"concrete_test":"Independently double-code a stratified random sample of 200 Health/Mental Health conversations (balanced by PHQ group) for disclosure level (1–5), support-seeking binary, and professional-redirect binary using two human raters blind to PHQ; compute Cohen’s κ against gpt-4o-mini and re-estimate the participant-weighted PHQ contrasts. If κ < 0.60 or the redirect difference reverses/sign flips after human labels, Finding 3 and the boundary-gap claim weaken materially.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s central design implication (higher-PHQ users enter high-disclosure/support-seeking contexts more often, yet professional redirection stays flat at ~16 %) rests almost entirely on unvalidated gpt-4o-mini annotations applied to the Health/Mental Health subset (Methods; Appendix B.2.3–B.2.4). Deterministic markers (nocturnal share, first-person pronouns, absolutist words) and the modest AUROC are independent of those labels, but the relational-support and boundary-gap story is not. Temperature-0 prompting and exploratory framing do not substitute for human agreement; systematic over- or under-labeling of disclosure or redirect language that co-varies with PHQ-associated lexical style would collapse the key contrast without touching the lexical or timing results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This paper links PHQ-8 symptom scores from 766 Prolific participants to 187,093 donated ChatGPT conversation histories, comparing PHQ<10 vs PHQ≥10 groups. Higher-PHQ participants show higher shares of mental-health, interpersonal, loneliness, and negative self-focus conversations; more late-night and recurring month-level use; elevated first-person singular and absolutist language; and more high-disclosure/support-seeking contexts, while professional redirection rates are essentially flat (~16%). Language-only prediction of PHQ≥10 is modest (best AUROC 0.591) and framed as insufficient for screening. The authors position ChatGPT as informal, always-available support infrastructure rather than a clinical tool, with design implications around history-aware boundaries and redirection.","tokens_in":16736,"tokens_out":716,"duration_ms":12441,"significance":"The contribution is timely and substantial for CSCW/HCI: a large, survey-linked corpus of private LLM histories is rare, and the work carefully separates descriptive use patterns from clinical screening claims. Strengths include participant-weighted contrasts with FDR control, conversation-weighted clustered adjusted models, recency-window sensitivities (14–90 days), health vs non-health subset checks, and an explicit negative result on screening utility. If the main contrasts hold under stronger label validation, the paper would be a durable empirical reference for how depressive-symptom severity co-occurs with relational LLM use and for design debates about professional boundaries in always-on chat systems.","major_comments":[{"comment":"Finding 3 and the design-implication claim that professional redirection does not scale with need rest almost entirely on unvalidated gpt-4o-mini labels for disclosure level, support-seeking, and professional_redirect (Methods; Appendix B.2.3–B.2.4; Table 1). Temperature-0 prompting and “exploratory” framing do not substitute for human agreement. Please report inter-rater reliability (or model–human agreement) on a stratified sample of Health/Mental Health turns, and show that PHQ-group differences survive when restricted to high-agreement items. Without this, the boundary-gap result remains the softest load-bearing claim.","section":null},{"comment":"Finding 1’s headline mental-health, interpersonal, loneliness, and negative self-focus shares likewise depend on the same unvalidated conversation-level GPT taxonomy (Appendix B.2.1–B.2.2; Table 1). Deterministic markers (nocturnal share, first-person pronouns, absolutist words) and the AUROC result are independent of those labels and should be presented as the primary evidence backbone. Please either validate topic/construct labels on a human-coded subset or restructure Results so LLM-dependent vs deterministic claims are clearly separated and weighted accordingly.","section":null},{"comment":"§3 and Appendix C.4: PHQ-8 indexes the prior two weeks, while primary contrasts use full exported histories. Recency sensitivities are a strength, but several lexical and advice/information contrasts weaken or lose significance in the 14–30 day windows (Appendix Table 5). The manuscript should state more explicitly which claims remain robust under the survey-anchored 14-day window and avoid treating full-history LLM-label shares as interchangeable with “recent symptom severity” without that qualification.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first large-scale link of survey PHQ-8 to donated longitudinal ChatGPT histories (766 people, 187k conversations). That alone makes it worth reading if you work on HCI, CSCW support infrastructures, or consumer chatbot safety.\n\nWhat is new is the combination: external symptom scores, full exported histories, nocturnal and month-level recurrence, user lexical markers, disclosure/support contexts, professional redirection, and assistant-side language, all under a deliberately non-clinical framing. They do the statistics carefully—participant-weighted Welch tests with FDR by family, conversation-weighted clustered adjusted models, recency windows (14–90 days), and health vs non-health subset checks. Effect sizes are modest (d ~0.2–0.46) but consistent. The prediction result is honest: best AUROC 0.591, not screening-ready. Limitations (convenience sample, PHQ as recent severity not diagnosis, observational design) are stated plainly.\n\nThe soft spot is real but partial. Finding 3’s boundary-gap story (higher disclosure/support-seeking, flat ~16% professional redirect) rests on gpt-4o-mini labels with no human agreement reported. Temperature-0 and “exploratory” framing do not fix that. If those labels systematically mis-track disclosure or redirect language that co-varies with PHQ style, the design implication weakens. Deterministic pieces—nocturnal share (14.2% vs 9.7%), first-person and absolutist rates, mental-health topic share, modest AUROC—do not depend on those labels and still stand. Sycophancy differences largely wash out after adjustment, which they report.\n\nCitation pattern is appropriate CSCW/HCI plus classic depression-language work; no load-bearing circularity. Data are private, so independent re-analysis is limited, but the methods appendix is detailed enough to judge.\n\nThis is for people designing or studying always-on conversational systems and informal support pathways. It deserves a serious referee, not a desk reject. I would engage with it, cite the observational baseline, and treat the redirection gap as provisional until labels are validated.","headline":"First large PHQ-linked ChatGPT history study with careful stats and a clear non-clinical framing; the boundary-gap claim leans on unvalidated LLM labels, but the timing/lexical core holds.","tokens_in":17279,"tokens_out":557,"would_cite":true,"duration_ms":6944,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"People with higher depressive symptoms use ChatGPT more for late-night, recurring mental-health support, but professional redirection does not rise with need.","keywords":["ChatGPT","depressive symptoms","PHQ-8","informal support infrastructure","disclosure","professional redirection","late-night use","language-based prediction"],"falsifier":"If human-coded disclosure and professional-redirection labels on the same Health/Mental Health subset reverse or erase the PHQ-group differences, or if a language-only model on held-out histories reaches clinically usable screening performance, the paper’s central interpretation fails.","tokens_in":17392,"feed_emoji":"💬","tokens_out":641,"duration_ms":7017,"temperature":0.7,"pith_summary":"This paper links Patient Health Questionnaire-8 scores from 766 participants to 187,093 donated ChatGPT conversation histories and treats ChatGPT as informal support infrastructure rather than a clinical tool. Participants at or above the moderate-symptom threshold (PHQ ≥ 10) brought more mental-health, interpersonal, loneliness, self-focused, and support-seeking conversations, used ChatGPT more between 23:00 and 04:59, and showed recurring month-level patterns of those topics. Their language contained more first-person singular pronouns and absolutist words, and they entered higher-disclosure contexts more often, yet professional redirection rates stayed essentially flat across groups. Language-only models predicting the PHQ split reached only modest performance (best AUROC 0.591), which the authors judge insufficient for screening. The claim is that private LLM histories should be read as evidence of how people already use always-on systems for support, not as clinical data for triage.","feed_headline":"Higher-PHQ users lean on ChatGPT for late-night support","feed_subtitle":"Redirection stays flat while disclosure and recurrence rise; language models cannot screen","key_machinery":"The PHQ ≥ 10 versus PHQ < 10 split applied to participant-weighted conversation histories (topics, disclosure, support seeking, nocturnal and month-level recurrence, lexical markers, and response-side professional redirection), used to contrast how symptom-severity groups engage ChatGPT as informal support infrastructure.","core_discovery":"Higher-PHQ participants used ChatGPT differently in relational and temporal ways that matter for design: more mental-health and interpersonal content, more late-night and recurring use, more self-focused language and high-disclosure support seeking, while professional redirection did not increase robustly with apparent need, and language-based prediction of the PHQ split remained too weak for screening.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Higher-PHQ users seek more late-night ChatGPT support","ChatGPT histories show more self-focus and disclosure above PHQ-10","Late-night and recurring ChatGPT use rises with higher PHQ scores","Higher-PHQ chats lean mental-health but redirection stays flat","Language fails to screen PHQ status despite support-seeking patterns"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The analysis rests on unvalidated gpt-4o-mini labels for topics, disclosure, support seeking, and professional redirection being accurate enough for group contrasts even though they cover only targeted subsets and have no human-rater check.","fun_headline_variants_meta":{"raw":{"variants":["Higher-PHQ users seek more late-night ChatGPT support","ChatGPT histories show more self-focus and disclosure above PHQ-10","Late-night and recurring ChatGPT use rises with higher PHQ scores","Higher-PHQ chats lean mental-health but redirection stays flat","Language fails to screen PHQ status despite support-seeking patterns"]},"model":"grok-4.5","effort":"low","cost_usd":0.003472,"raw_usage":{"total_tokens":1104,"prompt_tokens":740,"num_sources_used":0,"completion_tokens":95,"cost_in_usd_ticks":34720000,"prompt_tokens_details":{"text_tokens":740,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":269,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":740,"tokens_out":95,"duration_ms":3440,"temperature":1.0,"reasoning_tokens":269,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T03:47:36.207055+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If human-coded disclosure and professional-redirection labels on the same Health/Mental Health subset reverse or erase the PHQ-group differences, or if a language-only model on held-out histories reaches clinically usable screening performance, the paper’s central interpretation fails.","supporting_citations":[],"review_version":1}