{"id":"894731d5-19c6-4a20-90bd-98b937819987","arxiv_id":"2607.06371","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"In a vignette study of 354 U.S. participants, deletion-based privacy controls outperformed all other controls in increasing willingness to engage with GenAI chatbots for emotional support, while technically complex controls like local-only processing and model training opt-outs reduced engagement.","lead":"This paper finds that simple deletion controls — not technically sophisticated ones like local processing or training opt-outs — most increase users' willingness to share emotional struggles with AI chatbots. A smart generalist should read it because it maps where user trust breaks down in the rapidly growing intersection of AI and mental health, with direct design and policy implications.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified beyond the reader's correctly identified vignette-behavior gap, which is acknowledged and standard for HCI vignette methodology.","rationale":"The reader's verdict of ACCEPT with HIGH confidence is appropriate. The paper makes a solid empirical contribution with appropriate mixed-methods methodology, transparent limitations, and shared artifacts. The reader correctly identified the most load-bearing concern — the vignette-behavior gap, particularly for the affective urgency finding — and correctly noted that the paper acknowledges this limitation transparently in §3.5. This is a standard tradeoff in HCI vignette research, not a flaw in execution. The central claim (deletion controls outperform technically sophisticated controls) is supported by within-subjects comparisons with appropriate statistical controls, convergent qualitative evidence, and robustness across multiple significance thresholds. The three-gap framework, while inductively derived, is a useful organizing contribution grounded in the data. I found no concern that would warrant adjusting the verdict. The concrete test I propose would provide valuable evidence about the magnitude of the vignette-behavior gap specifically for the affective urgency finding, but its absence does not undermine the paper's contribution as it stands.","tokens_in":26585,"tokens_out":5772,"duration_ms":321135,"concrete_test":"To test whether the affective urgency finding is inflated by the vignette methodology, run a follow-up study with a 2×2 design: (vignette vs. real interaction with a chatbot) × (MFA present vs. absent), where participants in the real-interaction condition are primed with an emotional scenario and must actually complete MFA before chatting. If the MFA willingness penalty (currently β=−0.929) shrinks by more than 50% in the real-interaction condition compared to the vignette condition, the affective urgency gap magnitude is likely inflated by the hypothetical methodology. If the penalty persists or grows, the finding is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the most load-bearing concern: the study's central findings — particularly the affective urgency gap — depend on participants in a calm survey context accurately simulating how they would respond to S&P controls during emotional distress. The paper acknowledges this in §3.5. I considered several alternative concerns and found none that rise to the level of a load-bearing objection:\n\n1. **Comprehension gap as artifact of brief vignettes**: The one-sentence control definitions could inflate comprehension failures for technically complex controls like Local-only Processing. However, the paper notes definitions were 'directly informed by the language used by the reviewed chatbots,' meaning the comprehension gap reflects real-world product language. The paper's own recommendation to 'reframe technical mechanisms as user-facing outcomes' treats this as a genuine finding about current control descriptions, not an artifact.\n\n2. **Self-selected context assignment**: Context was assigned based on participants' highest-rated use case rather than random assignment, meaning the null context effect could be confounded by self-selection. However, this only weakens the secondary claim of context invariance, not the central claim about deletion controls outperforming other controls (which is a within-subjects comparison).\n\n3. **Absolute vs. relative effects**: The CLMM uses Delete Conversation as reference, so we know other controls performed worse relative to deletion but cannot confirm from the model alone that deletion itself increased willingness above the midpoint. However, the qualitative data (227 participants describing deletion as 'essential') provides strong convergent evidence for an absolute positive effect.\n\n4. **Three-gap framework circularity**: The framework is inductively derived, creating a risk of post-hoc fitting. But this is standard for grounded-theory approaches in HCI, and the framework is offered as an organizing contribution,","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper investigates how nine user-facing security and privacy (S&P) controls—derived from a systematic audit of 87 generative AI chatbot applications—influence users' willingness to engage, perceived protection, and perceived efficacy when using GenAI chatbots for emotional support. Using a mixed-methods vignette study (N=354 U.S. participants), the authors fit cumulative link mixed models (CLMMs) with participant random intercepts and supplement quantitative findings with inductive thematic analysis of open-ended responses. The central finding is that deletion-based controls dominate user preferences, significantly outperforming technically sophisticated controls like local-only processing and model training opt-outs, which suffer from comprehension gaps. The authors synthesize their findings into a three-gap framework (comprehension, assurance, affective urgency) and offer design and policy recommendations.","tokens_in":27163,"tokens_out":1187,"duration_ms":153084,"significance":"The paper addresses a timely and important gap at the intersection of usable privacy, conversational AI, and mental health. The saturation-based audit methodology is well-executed and follows HCI norms, and the nine-control taxonomy is derived independently rather than from prior theoretical commitments, avoiding circularity. The CLMMs are appropriate for the ordinal repeated-measures design, and the authors apply both BH FDR and Bonferroni corrections. The mixed-methods integration is a strength: qualitative codes provide mechanistic explanations for quantitative patterns (e.g., MFA's dissociation between protection and willingness). The three-gap framework is actionable and falsifiable. The vignette-behavior gap is acknowledged in §3.5 and is standard for HCI vignette methodology. Open science practices are noted, with survey instruments and analysis code provided in a supplementary repository.","major_comments":[{"comment":"The between-subjects assignment of Context of Disclosure is based on participants' highest-rated use case rather than random assignment. This self-selection mechanism means the null finding for context (Tables 3–5, Table 7: Depression and Interpersonal Tension coefficients near zero, all p > .05) could be confounded. Participants who select 'Anxiety & Stress' may differ systematically from those who select 'Interpersonal Tension' in ways that mask context effects. The paper claims context 'had no significant effect on user perceptions' (§4), but this is a secondary claim that the design cannot strongly support. The authors should explicitly acknowledge this confound when interpreting the null context effect, or soften the claim to note that context invariance is observed only among self-selected context assignments. This does not undermine the central within-subjects finding about deleti","section":"§3.2, Experimental Design"},{"comment":"The policy discussion states: 'Our findings suggest the empirical ground favors the latter orientation' (referring to California's platform-level vetting approach over the federal user-responsibility model). This claim overreaches the evidence. The study measures user perceptions of S&P controls in a hypothetical vignette context; it does not evaluate the effectiveness of regulatory frameworks or platform-level vetting. The finding that users lack comprehension of certain controls supports the general observation that user-responsibility models face challenges, but it does not constitute empirical evidence favoring one regulatory approach over another. The authors should reframe this as suggesting that their findings raise concerns about user-responsibility models, rather than claiming the empirical ground favors a specific regulatory orientation.","section":"§5.4, Policy Implications"}],"minor_comments":[{"comment":"The paper states each participant evaluated 4 of 9 S&P controls. Given the within-subjects design, it would be helpful to report how many participants were exposed to each control and whether any control was systematically under- or over-represented due to the randomization scheme.","section":"§3.2"},{"comment":"The vignette-behavior gap limitation is acknowledged but could be strengthened. The Baruh et al. meta-analysis citation is somewhat general; the authors could note that their vignettes are contextually specific (as Baruh et al. recommend), which is a design strength that partially mitigates the gap.","section":"§3.5"},{"comment":"The rationale for retaining only race/ethnicity as a demographic covariate is briefly mentioned ('covariate selection rationale'). A more explicit justification for why other demographics (e.g., age, gender, mental health care history) were removed would improve transparency.","section":"Table 7 note"},{"comment":"The categorization of controls into preventive, reversibility, and protective is introduced in the Discussion (§5.1–§5.3) but not in the Methods or Results. A brief note in the Methods explaining that this taxonomy emerged from the qualitative analysis would help readers understand its provenance.","section":"Figure 2"},{"comment":"The compensation ($2.50 for ~17 minutes) is noted as consistent with Prolific norms. This is acceptable, but the paper could note the hourly rate (~$8.82/hr) for reader expectations.","section":"Ethical Considerations"},{"comment":"The phrase 'betrayal framing' is used in the section heading but not clearly defined in the body text. A brief operational definition would improve clarity.","section":"§4.1"}],"recommendation":"minor_revision","confidential_remarks":"The reader's report and stress-test note correctly identify the vignette-behavior gap as the primary limitation, which the paper acknowledges. I agree this is standard for HCI vignette methodology and does not constitute a load-bearing objection. The two major comments I raise are about interpretation overreach (policy claims, context null effect) rather than fundamental flaws in the design or analysis. The central claim about deletion controls outperforming other controls is well-supported by both quantitative and qualitative evidence. The paper is a strong contribution to an underexplored area and should be published with minor revisions."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. Both major comments identify legitimate issues that we will address in revision. Below we respond to each point.","responses":[{"response":"The referee is correct. Context of Disclosure was assigned based on each participant's highest-rated use case in a pre-task, not by random assignment. This means participants self-selected into contexts, and systematic differences between groups (e.g., participants who rate 'Anxiety & Stress' highest may differ in clinically relevant ways from those who rate 'Interpersonal Tension' highest) could mask true context effects. Our design cannot rule out this confound. We chose self-selected context assignment to maximize ecological validity—participants evaluated vignettes in the context most relevant to them—but this design decision trades internal validity for contextual relevance, and we should have been more transparent about the tradeoff. In revision, we will (1) add an explicit acknowledgment of the self-selection confound in §3.5 (Limitations), noting that the null context effect is observed only among self-selected context assignments and cannot be interpreted as evidence of true context invariance, and (2) soften the claim in §4 from 'context had no significant effect on user perceptions' to language specifying that 'no significant differences were observed across self-selected disclosure contexts.' We will also note that a fully randomized design would be needed to draw stronger conclusions about context effects. We agree this does not undermine the central within-subjects finding about deletion controls, which is independent of the context assignment mechanism.","revision_made":"yes","referee_comment":"The between-subjects assignment of Context of Disclosure is based on participants' highest-rated use case rather than random assignment. This self-selection mechanism means the null finding for context could be confounded. The authors should explicitly acknowledge this confound or soften the claim."},{"response":"The referee is right that our study does not evaluate regulatory frameworks or platform-level vetting. Our evidence is about user comprehension of and trust in S&P controls, not about the comparative effectiveness of regulatory approaches. The sentence as written conflates 'our findings raise concerns about user-responsibility models' with 'the empirical ground favors a specific regulatory orientation,' which is a stronger claim than our data support. In revision, we will reframe this passage to state that our findings raise concerns about user-responsibility models—specifically, that such models presuppose a deliberative capacity to understand and calibrate S&P controls that our participants often lacked—without claiming that the empirical ground favors any particular regulatory framework. We will also clarify that we offer our three-gap framework as assessment criteria that regulators may find useful, not as empirical validation of any specific regulatory approach.","revision_made":"yes","referee_comment":"The policy discussion states 'Our findings suggest the empirical ground favors the latter orientation' (California's platform-level vetting over the federal user-responsibility model). This overreaches the evidence, as the study measures user perceptions in a hypothetical vignette context, not the effectiveness of regulatory frameworks."}],"tokens_in":26178,"tokens_out":839,"duration_ms":41774,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"This is a well-executed mixed-methods study on a genuinely underexplored question: how do user-facing security and privacy controls shape people's willingness to use GenAI chatbots for emotional support? The headline finding — that simple deletion controls outperform technically sophisticated ones like local-only processing and model training opt-outs — is well-supported and practically useful. The three-gap framework (comprehension, assurance, affective urgency) is a reasonable organizing lens, and the design recommendations are concrete rather than hand-wavy.","headline":"Solid empirical study on how S&P controls shape emotional engagement with GenAI chatbots; deletion controls dominate, three-gap framework is useful, vignette limitation is real but acknowledged.","tokens_in":27308,"tokens_out":866,"would_cite":true,"duration_ms":51910,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Simple deletion controls beat sophisticated privacy tech for AI emotional support","keywords":["generative AI","privacy controls","emotional support","chatbot","human-computer interaction","security usability","vignette study","mental health"],"falsifier":"If a real-world deployment study showed that users in acute emotional distress actually tolerate MFA or other friction-inducing controls at rates similar to their vignette responses, the affective urgency gap would need to be reconsidered. Conversely, if technically sophisticated controls like local-only processing were reframed in plain outcome language and subsequently performed comparably to deletion controls, the comprehension gap — not the controls themselves — would be confirmed as the primarybarrier.","tokens_in":26597,"feed_emoji":"🔒","tokens_out":1397,"duration_ms":74682,"temperature":0.7,"pith_summary":"This paper asks a straightforward question: when people turn to AI chatbots for emotional support, which privacy and security controls actually make them more willing to open up? Through a vignette study of 354 U.S. participants who already use generative AI chatbots for emotional support, the authors tested nine distinct privacy controls — from simple conversation deletion to technically complex options like local-only processing and model training opt-outs — across three emotional contexts (anxiety, depression, interpersonal tension). The central finding is counterintuitive: the simplest controls won decisively. Participants rated deletion-based controls (deleting a conversation or an entire account) as most likely to increase their willingness to engage, their sense of protection, and their confidence in the chatbot's helpfulness. Technically sophisticated controls performed significantly worse — local-only processing and model training opt-outs actually reduced willingness to engage compared to deletion. The paper explains this pattern through three structural gaps. The comprehension gap: users cannot form accurate mental models of how controls like local processing or training opt-outs actually work, leading to confusion and mistrust. The assurance gap: even when users understand a control (like deletion), they doubt the platform will actually honor it — over a third of participants expressed skepticism that deletion truly removes their data. The affective urgency gap: controls that add friction (like multifactor authentication) are recognized as protective but rejected because users in emotional distress need immediate access. Notably, the emotional context (anxiety vs. depression vs. interpersonal tension) had no significant effect on any of these patterns, suggesting the findings are robust across types of distress.","feed_headline":"Delete buttons beat sophisticated privacy tech for AI emotional support","feed_subtitle":"Users seeking emotional support from AI chatbots prefer simple deletion controls over technically advanced privacy features they can'tunder","key_machinery":"Three structural gaps — comprehension, assurance, and affective urgency — form the explanatory framework. Preventive controls (local-only processing, model training opt-out, memory toggle, anonymous chat, non-mandatory login) suffer from the comprehension gap: users lack accurate mental models of how these mechanisms affect their data. Reversibility controls (delete conversation, delete account & data) suffer from the assurance gap: users understand the promise but cannot verify execution. Protective controls (MFA, access/sharing controls) suffer from the affective urgency gap: users recognize protection but experience authentication barriers as prohibitively burdensome during emotionaldist","core_discovery":"The paper's central discovery is that user-facing privacy controls in AI chatbots succeed or fail based on three distinct, identifiable gaps between what a control promises and what users can comprehend, trust, or tolerate under emotional distress. Deletion controls dominate because they are comprehensible, feel reversible, and impose no friction — even though users doubt they truly work. Sophisticated controls fail not because they are technically inferior but because users cannot understand them, cannot verify them, or cannot bear the friction they impose when seeking emotional relief. The paper classifies the nine controls into three categories — preventive (limiting data creation),revers","pith_inferences":["If the comprehension gap generalizes beyond the nine tested controls, then any new privacy mechanism introduced into AI chatbots will face adoption resistance proportional to how difficult it is to explain in plain language — regardless of its technical strength.","The finding that emotional context had no effect on control preferences suggests that privacy control design for AI chatbots can be unified rather than context-specific, simplifying the design space considerably.","The coexistence of desire and doubt (participants who wanted to disclose more but doubted controls would protect them) implies a population of users currently under-disclosing to AI chatbots — representing lost therapeutic value that better assurance mechanisms couldunlock.","The paper's vignette methodology may systematically overstate the affective urgency gap, since calm survey respondents may not accurately simulate the urgency that reduces MFA tolerance during acute emotionaldistress."],"forward_implications":["Designers of AI chatbots used for emotional support should prioritize simple, outcome-framed deletion controls over technically sophisticated privacy mechanisms, since users understand and trust 'delete' far more than 'local-only processing' or 'training opt-out.'","The finding that MFA is perceived as protective but reduces willingness to engage suggests that authentication design for emotional-support contexts needs context-sensitive defaults — frictionless access during acute distress with optional hardening for users who fear physically proximate adversaries.","The assurance gap — where users understand deletion but doubt it works — implies that verifiable transparency mechanisms (showing what data exists before and after deletion) may be more important than adding new controls.","Policy frameworks that place responsibility on users to calibrate their own privacy settings (as in the federal AI policy approach cited) may fail systematically, since users demonstrably cannot understand or verify the controls they are asked tomanage."],"fun_headline_variants":["Simple deletion beats advanced privacy for AI chatbot support","Users trust delete buttons over complex AI privacy controls","For AI emotional support, simple deletion controls win out","AI privacy controls fail users when emotions run high","Friction and confusion limit advanced AI privacy tools"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The study's central design premise is that vignette-based hypothetical responses — where calm participants imagine how they would behave when emotionally distressed — reliably predict real-world behavior in actual emotional crises. If participants cannot accurately simulate the urgency and reduced tolerance for friction that accompanies acute distress, then the magnitude of the affective urgency finding may be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Simple deletion beats advanced privacy for AI chatbot support","Users trust delete buttons over complex AI privacy controls","For AI emotional support, simple deletion controls win out","AI privacy controls fail users when emotions run high","Friction and confusion limit advanced AI privacy tools","Users prefer simple delete options over complex AI privacy","Trust gap limits effectiveness of advanced AI privacy tools"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1015,"prompt_tokens":512,"completion_tokens":503,"prompt_tokens_details":null},"tokens_in":512,"tokens_out":503,"duration_ms":20718,"temperature":1.0,"reasoning_tokens":477,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T08:00:45.340559+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a real-world deployment study showed that users in acute emotional distress actually tolerate MFA or other friction-inducing controls at rates similar to their vignette responses, the affective urgency gap would need to be reconsidered. Conversely, if technically sophisticated controls like local-only processing were reframed in plain outcome language and subsequently performed comparably to deletion controls, the comprehension gap — not the controls themselves — would be confirmed as the primarybarrier.","supporting_citations":[],"review_version":1}