{"id":"f165c356-7555-4cfc-956b-ae610eb2af44","arxiv_id":"2504.18932","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"People with lived depression experience who tried a GPT-4o chatbot prioritized five values: informational support, emotional support, personalization, privacy, and crisis management.","lead":"This study had 17 people with lived experience of depression interact with a GPT-4o chatbot called Zenny in four self-management scenarios, then mapped the values they expressed to potential harms and design recommendations. It gives chatbot builders an empirical, user-side picture of what to protect when designing mental health AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Short, scripted, supervised probe may elicit different values than real long-term chatbot use; this untested transfer premise is load-bearing for the design recommendations.","rationale":"The reader's weakest-assumption analysis identifies exactly the load-bearing premise: the study's one-hour, scripted, supervised interactions with a technology probe stand in for real, long-term, unsupervised chatbot use. Section 3.1 explicitly says the probe was not deployed in a real-world setting, and Section 5.4 acknowledges the absence of long-term testing. This matters because the paper's contribution is a mapping of values to harms and design recommendations intended for actual mental-health chatbots. If the value priorities shift under real use—for example, if crisis management becomes more salient than the probe suggests, or if trust and continuity emerge as new values—the design guidance in Table 4 could be misdirected. The concern is real but not fatal: the paper is appropriately cautious, does not overclaim generalizability, and frames itself as a stepping stone to longitudinal work. The same concern is also visible in the crisis-management value, which was deliberately excluded from the interaction scenarios yet is reported as one of the five values and translated into design recommendations; that part of the mapping rests on hypothetical concerns rather than observed interaction. Thus, the paper's central claim is conditionally supported, and the appropriate verdict remains conditional. A longitudinal within-subjects replication is the concrete check that would settle whether the controlled probe's value-harms mapping transfers to real use.","tokens_in":28301,"tokens_out":6744,"duration_ms":74420,"concrete_test":"Recruit 15–20 participants with lived experience of depression and prior chatbot use; have them use a comparable GPT-4o-based chatbot freely at home for 4–6 weeks with safety monitoring and automated logging; conduct exit interviews and apply the same reflexive thematic analysis to logs and transcripts. Compare the resulting values/harms and their relative salience to the five values and Table 4 mapping. If the five values are not the most salient in naturalistic use, or if new values/harms (e.g., continuity of relationship, escalation to human care, sustained-trust issues) emerge, the controlled one-hour probe does not support the central mapping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that five values from lived experience can ground harm anticipation for real mental-health chatbots—depends on values elicited in the study being the values that govern actual use. Section 3.1 states Zenny was deliberately not deployed in a real-world setting; participants interacted with it only during one-hour interviews under researcher-presented, scripted scenarios (§3.1.1, Table 1). Those prompts directly solicit informational support ('Use Zenny to help you brainstorm... get advice on... talk to Zenny about...'), so the salience of that value is partly an artifact of task design. At the other end, crisis-management value is reported without a single participant quote, and the authors state they 'intentionally excluded crisis-related scenarios' (§4, Crisis management); the corresponding harms row in Table 4 therefore rests on hypothetical concerns, not probe behavior. Section 5.4 concedes the study 'did not directly address long-term interactions or sustained impacts.' Because the design recommendations (§5.2, Table 4) target real deployments (e.g., long-term scaffolding to encourage human connection, crisis governance), the unstated premise that short supervised interactions reveal the same value priorities and harm dynamics as longitudinal private use is load-bearing. The paper is honest about the limits and does not claim generalizability, but the empirical basis for the mapping is exactly this untested transfer.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study in which 17 U.S. adults with clinician-diagnosed depression and prior experience using AI chatbots interacted with Zenny, a GPT-4o-based technology probe, during one-hour remote interviews. Participants worked through two to four scripted self-management scenarios derived from van Grieken et al.'s taxonomy (goal setting, discussing depression with trusted others, leaving the house, and finding a new therapist) and were also invited to ask their own questions. Using reflexive thematic analysis of interview transcripts, chat logs, and notes, the authors derive five values: informational support, emotional support, personalization, privacy, and crisis management. They then map these values to potential harms and design recommendations in Table 4 and claim to contribute a harm mitigation checklist. The central claim is that harms of LLM-based mental health chatbots can be anticipated by grounding them in the values of people with lived experience of depression.","tokens_in":28481,"tokens_out":6411,"duration_ms":65286,"significance":"If the value-harms mapping is accepted as valid for real-world use, the paper is a useful value-sensitive design contribution that grounds harm anticipation in lived experience rather than only expert speculation. The study has genuine strengths: an explicit positionality statement, IRB-approved safety measures, scenario prompts based on patient-centered literature, an open 'ask your own question' channel, and a research team that includes a clinical psychologist. The identified personalization-privacy dilemma, supported by participant quotes about privacy-preserving tactics, is a substantive and actionable tension that deserves attention. However, the small self-selected sample, the short supervised probe context, and the intentional exclusion of crisis scenarios mean the empirical reach is narrower than the title and design recommendations suggest; the value-harms mapping is best read as a hypothesis-generating account rather than a validated basis for deployment decisions.","major_comments":[{"comment":"The values that anchor the central mapping were elicited in a one-hour, researcher-supervised interview in which the scenario prompts explicitly directed participants to request information (Table 1 tasks all ask for brainstorming, advice, or tips). The paper acknowledges in §5.4 that long-term interactions were not addressed, yet Table 4 and §5.2 issue recommendations for deployed systems, including long-term scaffolding to encourage human connection and crisis governance, that require the unstated premise that values from short, scripted, supervised interactions transfer to private, longitudinal use. The crisis-management value is especially thin in evidence: §4 states that crisis scenarios were intentionally excluded, the value is reported without a single participant quote, and the crisis harms row in Table 4 rests on hypothetical concerns rather than probe behavior. Please either reframe the contribution as 'values expressed in hypothetical supervised interactions' or provide a dedicated transferability argument, for example by triangulating the probe-based values with participants' prior real-world chatbot experiences.","section":"§3.1, §3.1.1, §4, §5.4, Table 4"},{"comment":"The Abstract and Section 1 announce 'a harm mitigation checklist,' and Section 2.1 repeats this as a contribution, but no checklist is actually presented anywhere in the manuscript; Table 4 is a mapping of values, harms, and design recommendations, not a checklist with discrete actionable items. This claimed artifact should either be added as a distinct table or section, or the contribution should be corrected to say 'design recommendations' rather than 'harm mitigation checklist.'","section":"Abstract, §1, §2.1, §5.2"},{"comment":"The empirical basis for the five-value taxonomy is a sample of 17 volunteers from a single U.S. registry, all of whom had already used AI chatbots more than once, and the paper does not report how many of the four scenarios each participant completed or which scenarios were skipped. Because several participants could not complete all scenarios and the prompts strongly invoke informational support, the prominence of that value may be partly an artifact of task design and completion rates. The first author led the deductive coding (Section 3.5), and although this is consistent with reflexive thematic analysis, the absence of any reported codebook, coding excerpt, or scenario-completion breakdown makes it difficult to assess how the five values were consolidated. Please include a scenario-completion table and a short illustrative excerpt of the coding structure to strengthen the transferability claim.","section":"§3.2, §3.4, §3.5, Table 2"}],"minor_comments":[{"comment":"There are several typos: 'expereinces' appears in §3.5, 'convinient' appears in §3.5, and 'the the' appears in §4 near 'Considering the the emerging use.'","section":"§3.5, §4"},{"comment":"The Figure 1 caption says the probe is 'built with GPT-4,' but the Abstract, Section 3.1.2, and Section 6 say GPT-4o; these should be made consistent.","section":"Figure 1 caption"},{"comment":"Reference [1] has a malformed URL: 'https://https://platform.openai.com/docs/models/.' Remove the duplicate scheme.","section":"Reference [1]"},{"comment":"Section 3.4 says the scenarios were based on patient perspectives on depression self-management [122], but reference [122] is van Berkel et al. on experience sampling; the intended source appears to be van Grieken et al. [123] (and [124]).","section":"§3.4, references [122] and [123]"},{"comment":"The author name 'Søgaard Neilsen' is misspelled; the cited work by Søgaard Nielsen and Wilson should be spelled 'Søgaard Nielsen.'","section":"§2.3, reference [113]"}],"recommendation":"major_revision","confidential_remarks":"The paper is a sincere and transparent qualitative study, and I did not find evidence of circularity in the value derivation. The two issues I would ask the editor to weigh are the gap between the claimed 'harm mitigation checklist' and the actual manuscript content, and the degree to which the crisis-management value and the broader value-harms mapping are supported by data from which crisis scenarios were deliberately excluded. The authors' heavy reliance on their own prior work in the framing is noticeable but not inappropriate. A revision that adds the missing checklist artifact and a scenario-completion breakdown, and that explicitly limits the scope to values expressed in supervised hypothetical interactions, would strengthen the contribution considerably."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent exploratory qualitative study that maps lived-experience values to potential LLM chatbot harms, and the personalization-privacy dilemma it names is a genuinely useful contribution. The authors do the basics right: clear method, quotes that actually support the themes, safety measures that show care, and a limitations section that does not pretend the sample is representative. I would send it to review.\n\nThe new thing here is the empirical mapping itself—values and harms from people with lived depression experience interacting with a GPT-4o chatbot—plus the explicit dilemma between personalization and privacy, which is a practical safety requirement for designers. The five values are recognizable from prior social support and privacy literature, and the method (probe plus reflexive thematic analysis) is established, so this is incremental but useful.\n\nThe soft spots are real but not fatal. The biggest is the transferability premise: participants used Zenny for an hour in supervised, scripted scenarios, and the scenarios themselves prompt advice-seeking, which inflates the salience of informational support. The paper acknowledges in 5.4 that it didn't test long-term interactions, so the design recommendations for long-term use are reasonable conjectures, not empirical findings. Also, crisis management is the weakest value: crisis scenarios were deliberately excluded, and no participant quote supports it, so that harm row in Table 4 is entirely hypothetical. That's worth flagging in any revision. Two smaller issues: the sample is 17 from a single US registry and required prior chatbot use, and the coding was led by the first author with team review—fine for an exploratory study but worth noting. Also there's a citation mismatch in 3.4 (points to [122] van Berkel instead of [123] van Grieken) that obscures the scenario source; easy to fix.\n\nThe paper doesn't overclaim. It says the goal is transferability, not generalizability, and the recommendations stay close to the data. I'd want the authors to sharpen the distinction between what participants said and what designers can safely infer from one-hour interactions, and to either get crisis data or explicitly frame that row as speculative.\n\nFor peer review: yes, serious desk editors should send this out. It's a solid qualitative contribution to the mental health chatbot literature, and the personalization-privacy dilemma is worth disseminating. My own verdict would be conditional acceptance with revision, not rejection.","headline":"A competent, honest values-harms mapping for mental-health chatbots with one genuinely useful named dilemma, but the short-supervised-interaction basis means the design recommendations are hypotheses, not validated results.","tokens_in":29056,"tokens_out":2643,"would_cite":true,"duration_ms":27700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People with lived experience of depression value five things in mental health chatbots, and harms map to those values.","keywords":["LLM","mental health","depression","chatbot","conversational agent","ethics","harms","self-management"],"falsifier":"A longitudinal field trial in which people use a mental health chatbot in their daily lives for several weeks, with crisis episodes and privacy incidents logged, would settle the claim: if reported harms do not track the five values, or if new values emerge that change the mapping, the central claim is weakened.","tokens_in":28058,"feed_emoji":"💬","tokens_out":8859,"duration_ms":79186,"temperature":0.7,"pith_summary":"The paper tries to establish that the harms of LLM mental health chatbots can be anticipated by asking people with lived experience of depression what they value, rather than relying only on expert opinion. It reports that 17 participants who interacted with a GPT-4o chatbot called Zenny in four self-management scenarios prioritized five values: informational support, emotional support, personalization, privacy, and crisis management. From those values the authors derive a mapping to potential harms and a set of design recommendations, including a harm mitigation checklist. This matters because public enthusiasm for AI chatbots is ahead of evidence about their risks, and a values-based map offers a concrete way to design for safety.","feed_headline":"Five values predict mental-health chatbot harms","feed_subtitle":"Interviews with 17 people with lived depression show misinformation, over-reliance, and privacy risks threaten key values.","key_machinery":"The central object is Zenny, a GPT-4o-based technology probe embedded in a one-hour scenario-based interview. The mechanism is a value-harms mapping built from reflexive thematic analysis of chat logs, interview transcripts, and notes: participants' expressed values are treated as the standard against which potential harms are identified. The analysis also surfaces the personalization-privacy dilemma, the tension where tailored advice requires more sensitive disclosure, as the core dynamic designers must resolve.","core_discovery":"The paper's central claim is that the harms of LLM-based mental health chatbots can be mapped to values held by people with lived experience of depression, and that this mapping yields design guidance. In interviews built around interactions with Zenny, a GPT-4o chatbot used as a technology probe, participants prioritized five values: informational support, emotional support, personalization, privacy, and crisis management. The authors argue that inaccurate or inapplicable advice, over-reliance on chatbot emotional support, the personalization-privacy dilemma, and inadequate crisis handling are best understood as threats to these values, and they offer design recommendations plus a harm mitigation checklist to address them.","pith_inferences":["Editorial inference: If the value-harms mapping generalizes, showing users a visible profile of what the chatbot has inferred about them could be tested as a way to rebuild trust and reduce the privacy harm from invisible inference.","Editorial inference: The personalization-privacy dilemma implies a measurable trade-off between the amount of sensitive information disclosed and the relevance of advice, so interventions could be evaluated by how much they shift that trade-off in the user's favor.","Editorial inference: A longer-term deployment might surface additional values, such as continuity or trust calibration, because one-hour interactions cannot capture how users handle repeated use, memory, and broken advice over weeks.","Editorial inference: The same value-harms mapping approach could be extended to other mental health conditions or to comorbid populations, but only if the values are re-elicited rather than assumed to transfer."],"forward_implications":["Mental health chatbots should explicitly tell users that responses may be inaccurate and encourage cross-checking with clinicians or other sources.","Chatbots should ask follow-up questions about constraints and preferences before giving advice, so suggestions are contextually applicable rather than generic.","Emotional-support features should be paired with prompts and scaffolding that steer users toward human support networks, reducing over-reliance.","Privacy controls should let users see and delete what the chatbot stores and infers, because users already obscure their queries to protect themselves.","Crisis management should be designed in from the start: clear limitations, referral to resources like 988, and governance oversight."],"supporting_citations":[{"why":"Supplies the lived-experience self-management strategies from which the four interview scenarios were built.","marker":"[123]"},{"why":"Validates the helpfulness of those self-management strategies, grounding the scenarios in patient experience.","marker":"[124]"},{"why":"Defines the technology probe method that justifies using Zenny inside interviews.","marker":"[53]"},{"why":"Establishes the value-centric design approach that frames the study's elicitation of stakeholder values.","marker":"[43]"},{"why":"Argues that values are situated in lived experience, motivating recruitment of people with depression.","marker":"[69]"},{"why":"Catalogs potential benefits and harms of LLMs in digital mental health, the prior expert-led baseline the study contrasts with participant enthusiasm.","marker":"[34]"},{"why":"Raises ethical dilemmas for LLM-based chatbots, representing the expert concerns compared against participant views.","marker":"[20]"},{"why":"Supplies the reflexive thematic analysis method used to derive the five values and three themes.","marker":"[19]"},{"why":"Frames value-sensitive algorithm design and motivates engaging stakeholders early in system development.","marker":"[143]"}],"fun_headline_variants":["Chatbot harms mapped to five values from lived experience","Depression chatbot risks tied to five valued features","Lived experience shapes mental-health chatbot design","Five values flag mental-health chatbot dangers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings assume that what people value and fear during short, scripted, one-hour chats with a chatbot in a supervised interview is what they would value and fear in real, long-term use of a mental health chatbot.","fun_headline_variants_meta":{"raw":{"variants":["Chatbot harms mapped to five values from lived experience","Depression chatbot risks tied to five valued features","Lived experience shapes mental-health chatbot design","Five values flag mental-health chatbot dangers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1271,"prompt_tokens":820,"completion_tokens":451,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":394}},"tokens_in":436,"tokens_out":451,"duration_ms":5062,"temperature":1.0,"reasoning_tokens":394,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:05:35.539357+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal field trial in which people use a mental health chatbot in their daily lives for several weeks, with crisis episodes and privacy incidents logged, would settle the claim: if reported harms do not track the five values, or if new values emerge that change the mapping, the central claim is weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the lived-experience self-management strategies from which the four interview scenarios were built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates the helpfulness of those self-management strategies, grounding the scenarios in patient experience."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames value-sensitive algorithm design and motivates engaging stakeholders early in system development."}],"review_version":1}