{"id":"48215325-e74b-445a-8f9e-ac285c1a6a10","arxiv_id":"2501.15628","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"People with severe social anxiety symptoms report greater trust in and willingness to use GenAI chatbots, valuing emotional connection, while milder-symptom users emphasize technical reliability.","lead":"This paper surveyed 159 people and interviewed 17 to ask how people with different levels of social anxiety view using generative AI chatbots for support. It finds that people with severe symptoms are more willing to trust and use these chatbots for emotional comfort, while people with milder symptoms care more about technical reliability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central severity–trust claim is confounded: high-severity interview participants are also long-term GenAI users, so emotional trust may reflect usage duration rather than symptom severity.","rationale":"The paper is a genuinely useful, honestly reported mixed-methods study, and the authors explicitly flag the missing high-severity/never-user group in Section 7. That absence is exactly the weak point. The central contribution is the qualitative claim that symptom severity drives different trust priorities—severe users value emotional trust, mild users value technical reliability—but the interview design entangles severity with GenAI usage duration: the high-severity cluster is also the long-term-user cluster, and interviewees were drawn from those clusters. Thus a plausible rival explanation is that familiarity and repeated use, not the severity of social anxiety itself, leads participants to value the chatbot's non-judgmental, emotionally supportive qualities. The survey regression does control for usage duration and frequency as separate main effects when predicting willingness, and it does find a severity effect on willingness, which is a real supporting result. However, it does not measure emotional versus cognitive trust priorities, so it cannot resolve the confound in the qualitative contrast. I also note that calling the trust–willingness association \"strong\" is generous given Spearman's rho = 0.30 and R-squared = 0.13, but that is a secondary presentation issue rather than the deepest threat. Because the claimed severity-based design implications depend on an attribution the current sample cannot identify, the conditional verdict is appropriate; no change to the reader's verdict is needed.","tokens_in":33252,"tokens_out":3610,"duration_ms":37844,"concrete_test":"Recruit a high-severity, short-or-no-usage cell (e.g., at least 5 SPIN-severe participants who have never used GenAI for SA support or have used it for less than 3 months), interview them with the same protocol, and compare their trust priorities with the existing high-severity/long-use and low-severity/long-use quotes. If these never-users emphasize technical reliability, accuracy, or skepticism toward emotional support, then the severity attribution fails and the familiarity explanation survives. A faster check on the current data is to re-code all 17 interviews with severity and usage duration masked and test whether emotional-trust codes are predicted by usage duration as strongly as by severity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3's clustering (Figure 4B) defines the undecided group's high-severity profile as \"long-term GenAI use with high severity,\" and Section 5.1's emotional-trust findings come from that same cluster. The interview sample therefore has no high-severity short-duration or never-use cell, which the authors acknowledge in Section 7. Because usage length and severity do not vary independently, the qualitative contrast between \"severe users value emotional trust\" and \"mild users value technical reliability\" cannot establish that symptom severity is the driver; long-term exposure and familiarity with chatbots are a plausible alternative explanation. The survey regression in Table 5 estimates severity, duration, and frequency as separate main effects and does show a severity effect on willingness, but it does not test trust priorities; it therefore does not rescue the qualitative causal attribution. This is load-bearing because the paper's design implications and the conclusion that trust-building for severe users should emphasize emotional support rest on that attribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Wang et al. report a mixed-methods study (survey n = 159, interviews n = 17) of attitudes toward and trust in generative AI chatbots for social anxiety (SA) support. Quantitative analyses examine the trust–willingness association, an ordinal logistic regression of willingness on SA severity and GenAI usage patterns, and GMM clustering of participants in the 'undecided' willingness group. Qualitative thematic analysis of the interviews contrasts emotional trust among severe-symptom users with cognitive/technical trust among milder-symptom users. The paper derives design implications around emotional versus cognitive trust and discusses ethics and future research directions.","tokens_in":33418,"tokens_out":6413,"duration_ms":58413,"significance":"If fully supported, the paper would make a useful contribution to HCI and mental-health technology by showing that trust in GenAI support is situated and symptom-dependent, with concrete design implications. The study has notable strengths: a mixed-method design, use of the validated SPIN instrument, explicit reporting of statistical procedures and effect sizes, and a candid limitations section. The survey's severity–willingness effect is a meaningful empirical result. However, the headline qualitative claim that symptom severity drives trust priorities is currently confounded with GenAI usage duration, and the abstract's characterization of the trust–willingness relationship as 'strong' overstates the reported moderate effect sizes. As it stands, the findings are suggestive rather than conclusive for the paper's main design recommendation.","major_comments":[{"comment":"The qualitative contrast between severe and mild users' trust priorities is confounded with GenAI usage duration. In Figure 4B, the high-severity cluster is also the long-term-use cluster, and the interview participants in Section 5.1 are recruited from these clusters; Table 2 confirms that no interview participant was a never-user of GenAI. The Limitations (Section 7) acknowledge the absence of a high-severity, never-used group, but this caveat is not carried into the design implications in Section 6.2, which recommend tailoring emotional trust-building to severe-symptom users. The emotional-trust finding could equally be explained by familiarity and usage length rather than by symptom severity. The ordinal logistic regression in Table 5 estimates severity, duration, and frequency as separate main effects on willingness and does not test trust priorities, so it does not rescue the qualitative attribution. Please either substantially soften the causal framing to a description of the observed cluster profiles or provide additional data or analyses that separate severity from usage duration.","section":"§5.1, §4.3, §7"},{"comment":"The paper repeatedly calls the trust–willingness relationship 'strong' in the abstract and in the Section 4 'Takeaways' paragraph, but the reported Spearman rho = 0.30 and R-squared = 0.13 are moderate-to-weak by conventional social-science standards. The body text itself correctly describes the association as 'moderate' in Section 4.1.2. This inconsistency matters because the abstract's central claim overstates the quantitative support. Please harmonize the wording and either drop 'strong' or provide a field-specific benchmark justifying the label.","section":"Abstract, §4.1.2, §4 Takeaways"},{"comment":"The abstract and introduction say that individuals with severe symptoms 'tend to trust and embrace GenAI chatbots more readily,' but the quantitative result in Table 5 concerns willingness to use, not trust. Trust-by-severity is not directly tested in the survey; the qualitative data address trust priorities but are confounded as noted above. Please disambiguate 'trust' and 'willingness' throughout the claims so that the strength of each reported result is accurately represented.","section":"Abstract, §1, §4.2"}],"minor_comments":[{"comment":"The sentence following Table 5 says participants with very severe symptoms showed 'even stronger odds at 1.124,' but the coefficient for the Severe group is 1.233 and is larger than the Very Severe coefficient; please correct this description and avoid implying a strictly monotonic severity effect when Mild and Moderate are non-significant.","section":"§4.2, Table 5"},{"comment":"The cluster figure and its caption contain garbled annotation text (e.g., 'with high severityof SA') and placeholder symbols such as 'group ' with no visible group names; please provide a clean legend and complete sentences in the figure so the three clusters are identifiable.","section":"§4.3, Figure 4"},{"comment":"The GMM cluster count is selected using WCSS and silhouette scores, which are typically associated with k-means; please clarify how these criteria were applied to Gaussian mixture models and whether mclust's model-selection criteria (e.g., BIC) were also considered.","section":"§4.3, Appendix Figure 5"},{"comment":"The trust scale is adapted in part from the authors' own prior work (references [123] and [124]) and is validated with EFA on the same sample used for the main trust–willingness test; this is acceptable for an exploratory study, but please acknowledge explicitly that the factor structure and the association are not independent evidence.","section":"§3.1.2, §4.1.1"},{"comment":"Please correct minor typos: 'Haman at el.' in reference [42] should be 'Haman et al.'; 'its'' in Section 1 should be 'its'; 'might might form' in Section 4 should be 'might form'; and the heading in Section 5.1.3 should read 'GenAI chatbots are slightly better than a bad psychotherapist.'","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within CHI's scope and the quantitative survey portion is usable, but the headline qualitative claim—that symptom severity, rather than usage duration, drives emotional-trust priorities—is not currently supported because severity and GenAI exposure are confounded in the interview sample. The authors should either collect additional cells (high-severity never-users or short-duration users) or substantially re-frame the claims as descriptive of the sampled cluster profiles. The 'strong correlation' wording should also be calibrated to the reported effect sizes. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere is my take on the social-anxiety chatbot trust paper. It is a well-organized mixed-methods study on a timely topic: how people with different levels of social anxiety trust GenAI chatbots. The genuinely new bit is the severity-based split—severe users valuing emotional, non-judgmental support and mild users prioritizing technical reliability. That has not been shown explicitly for GenAI chatbots in SA contexts, and the interview excerpts give it texture. The authors also did the work properly: SPIN stratification, ordinal logistic regression, GMM clustering, and a limitations section that already admits the main gap. Credit where due.\n\nNow the soft spots. The abstract calls the trust–willingness relationship “strong,” but the reported Spearman rho is 0.30 and R-squared is 0.13. That is moderate-to-weak; “strong” is an overstatement. More importantly, the qualitative claim that severe-symptom users prioritize emotional trust is confounded. The high-severity cluster from Section 4.3 is also the long-term-use cluster, and the interviewees were recruited from those clusters. So the emotional-trust pattern could just as easily come from familiarity and frequency of use as from symptom severity. The regression in Table 5 shows a severity effect on willingness, but it does not test trust priorities, so it does not rescue the attribution. The authors acknowledge the missing high-severity never-user cell, which is good, but it means the central design implication—build emotional trust for severe users—rests on weaker evidence than the narrative implies. The EFA on the same sample used for the main tests is a minor circularity, and the scale’s roots in the authors’ prior work are fine, just worth noting.\n\nWho should read this? HCI and digital-mental-health researchers, especially those working on trust in LLM-based support tools. It is a useful empirical data point, not a breakthrough. I would send it to peer review—a serious referee can push for the re-analysis and a rewording of the claims. My verdict: conditional acceptance or major revision, with the confound explicitly addressed.","headline":"Solid mixed-methods study, but the severity–trust split is softer than the abstract suggests and the qualitative attribution is confounded with usage duration.","tokens_in":33940,"tokens_out":2567,"would_cite":false,"duration_ms":23126,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that symptom severity determines whether trust in GenAI chatbots for social anxiety rests on emotional connection or technical reliability, and that both trust modes must be designed for.","keywords":["social anxiety","generative AI","chatbots","trust","mixed methods","emotional trust","cognitive trust"],"falsifier":"Recruit people with very severe SPIN scores who have never used a GenAI chatbot, present all of them with the same first interaction, and ask what grounds their trust; if they do not emphasize non-judgmental emotional connection over technical reliability, the severity explanation fails. Re-analyzing the undecided cluster while controlling for duration of chatbot use would also settle whether the apparent severity effect survives.","tokens_in":33041,"feed_emoji":"💬","tokens_out":9344,"duration_ms":83442,"temperature":0.7,"pith_summary":"The paper asks how people with social anxiety decide to trust generative AI chatbots as mental-health support, and whether that trust depends on symptom severity. Using survey data from 159 respondents and follow-up interviews with 17, it finds that trust and willingness to use move together, and that severity changes what people trust for. People with severe symptoms trust chatbots for emotional reasons—non-judgmental availability, perceived empathy, and a sense of being understood—while people with milder symptoms trust them for technical reasons—accuracy, memory, and user control. The paper argues that designers should build both kinds of trust and that small emotional cues can carry real weight for the most distressed users. If right, this means mental-health chatbots should not be judged only on factual reliability; for the people most likely to adopt them, perceived emotional comprehension is the bridge to trust.","feed_headline":"Severe anxiety shifts chatbot trust from reliability to empathy","feed_subtitle":"159-person survey and 17 interviews show symptom severity decides whether users trust AI for warmth or for accuracy.","key_machinery":"The paper's central instrument is a six-item trust scale measuring competence, honesty, experience, benevolence, reliability, and expectation; factor analysis showed the items load onto one trust factor, and the scale anchors every quantitative comparison. The Social Phobia Inventory (SPIN) supplies the severity grouping, and model-based cluster analysis of the undecided users identifies three profiles—high severity with long use, low severity with long use, and low severity with short use—which guide the interview sampling. Conceptually, the paper organizes trust into two modes: emotional trust, trust rooted in feeling unjudged, understood, and emotionally held, and cognitive trust, trust rooted in factual reliability and predictable competence. The qualitative analysis then uses these two modes to explain why the same technology is trusted for different reasons across severity groups.","core_discovery":"The paper's central discovery is that trust in GenAI chatbots for social anxiety support is not a single thing; it splits by symptom severity. In the survey, trust and willingness to use were significantly linked (rank correlation ρ = 0.30), and respondents with severe or very severe SPIN scores were significantly more willing to adopt chatbots than those with no symptoms. In the interviews, participants with severe symptoms said they trusted chatbots because interactions felt non-judgmental, emotionally attuned, and always available—trust rooted in emotional connection—while participants with milder symptoms said they trusted chatbots only when the model was accurate, remembered context, and left control with the user—trust rooted in technical reliability. The paper also found that even minimal empathy cues, such as a warm tone or a thinking animation, could build emotional trust, and that some users with disappointing therapy experiences considered a chatbot better than an inadequate human psychotherapist. The authors conclude that GenAI chatbot design for social anxiety should deliberately build both cognitive and emotional trust, depending on the user's symptom severity.","pith_inferences":["The paper does not test, but its logic implies that conventional AI evaluation based on factual accuracy may miss the trust mechanism that matters most to highly distressed users; a plausible design corollary is that empathic tone could increase adoption more than a further accuracy gain.","A direct next study would separate severity from familiarity: compare high-severity users with no prior chatbot experience against high-severity long-term users on the same standardized interaction, because the current design cannot tell those two explanations apart.","The finding that minimal cues like a thinking spinner and warm text build emotional trust turns those cues into testable interventions; an A/B experiment varying tone and presence cues would quantify their causal effect on trust.","The cluster pattern hints at a selection loop in which users who feel emotionally held keep using chatbots and report more trust, while distrustful users never accumulate experience; longitudinal data from first session onward would reveal whether trust drives engagement or engagement drives trust."],"forward_implications":["Trust and willingness are coupled: survey respondents who were unwilling to use chatbots had the lowest trust scores, and trust rose as willingness moved from unwilling to undecided to willing, so trust-building should be expected to increase adoption.","Severe social anxiety is linked to significantly higher willingness to adopt GenAI chatbots, with severe and very severe SPIN groups showing the largest odds, making these users the natural early-adopter population.","For severe-symptom users, emotional trust can be built through small and consistent cues—warm wording, non-judgmental tone, constant availability, even minimal simulated empathy—so these design elements are not peripheral.","For mild-symptom users, cognitive trust is the gate: hallucinations, lack of long-term memory, rigid responses, and unclear user control all suppress trust, so reliability and transparency must come first for this group.","A chatbot can occupy a useful niche between no help and inadequate human psychotherapy, providing consistent basic support; the paper frames this as complementing, not replacing, professional care."],"supporting_citations":[{"why":"Supplies the SPIN inventory used to group respondents into none, mild, moderate, severe, and very severe social anxiety categories.","marker":"[19]"},{"why":"Provides the six trust dimensions—competence, honesty, experience, benevolence, reliability, expectation—on which the survey trust scale is built.","marker":"[36]"},{"why":"Supplies the trust-measurement and profiling approach that the six survey items and the trust–willingness analysis adapt.","marker":"[124]"},{"why":"Establishes the prior link between trust and willingness to follow health advice that motivates the paper's trust–willingness analysis.","marker":"[106]"},{"why":"Documents chatbot use among people with social deficits as a non-judgmental, low-pressure practice space, the gap the paper extends to trust dynamics.","marker":"[34]"},{"why":"Supports the claim that severe symptoms heighten the need for safe, non-judgmental interactions, the mechanism behind emotional trust.","marker":"[24]"},{"why":"Supplies the general finding that empathy is critical in mental health support, which the paper applies to even minimal chatbot empathy.","marker":"[63]"}],"fun_headline_variants":["Anxiety severity decides if AI trust is emotional or technical","Severe social anxiety boosts trust in AI chatbots","Empathy over accuracy: how anxiety severity shapes AI trust","Mild anxiety demands reliability; severe anxiety demands empathy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the differing trust priorities reflect symptom severity, but severity and chatbot experience are entangled in the sample, and the paper itself notes there was no high-severity group that had never used GenAI, so prior familiarity could produce the same pattern.","fun_headline_variants_meta":{"raw":{"variants":["Anxiety severity decides if AI trust is emotional or technical","Severe social anxiety boosts trust in AI chatbots","Empathy over accuracy: how anxiety severity shapes AI trust","Mild anxiety demands reliability; severe anxiety demands empathy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1417,"prompt_tokens":909,"completion_tokens":508,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":444}},"tokens_in":525,"tokens_out":508,"duration_ms":5516,"temperature":1.0,"reasoning_tokens":444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:05:16.220295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit people with very severe SPIN scores who have never used a GenAI chatbot, present all of them with the same first interaction, and ask what grounds their trust; if they do not emphasize non-judgmental emotional connection over technical reliability, the severity explanation fails. Re-analyzing the undecided cluster while controlling for duration of chatbot use would also settle whether the apparent severity effect survives.","supporting_citations":[{"cited_title":"Connor, Jonathan R","cited_arxiv_id":null,"evidence_quote":"Supplies the SPIN inventory used to group respondents into none, mild, moderate, severe, and very severe social anxiety categories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the trust-measurement and profiling approach that the six survey items and the trust–willingness analysis adapt."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior link between trust and willingness to follow health advice that motivates the paper's trust–willingness analysis."},{"cited_title":"I Don’t Want to Bother You","cited_arxiv_id":null,"evidence_quote":"Supports the claim that severe symptoms heighten the need for safe, non-judgmental interactions, the mechanism behind emotional trust."}],"review_version":1}