REVIEW 2 major objections 5 minor 9 references
LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics
T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read People's willingness to use health AI chatbots tracks benefits, risks, and personal traits far more than whether the topic is physical or psychological.
desk verdict Solid preregistered factorial survey: benefits/risks and individual traits dominate topic type for chatbot health use; scenario method is the main external-validity limit. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A mixed factorial design (topic type between-subjects × topic sensitivity within-subjects) that presents short scenarios of AI-chatbot interaction for pretested health conditions, then measures perceived benefits, risks, intention, and willingness to self-disclose, with linear mixed models linking those outcomes to both experimental factors and individual covariates.
What would settle it
A field study that logs actual chatbot conversations and subsequent disclosure or continued use for matched physical versus psychological and low- versus high-sensitivity topics, then checks whether topic type still fails to predict behaviour once benefits, risks, and the same individual covariates are controlled.
Extended reading notes
Core claim
In a 2 imes2 mixed experiment with a Dutch representative sample (N = 1,388), perceived benefits positively predicted intention to use an AI chatbot and willingness to self-disclose health information, while perceived risks negatively predicted both. Topic type (physical vs psychological) had negligible univariate effects; the only reliable topic effect was slightly higher usage intention for low-sensitivity than high-sensitivity scenarios. Individual characteristics—prior chatbot use, AI literacy, education, political orientation, institutional trust, and monitoring coping style—also systematically shifted perceptions and intentions. The authors conclude that chatbot use and disclosure for
Load-bearing premise
That asking people to imagine having a given health condition and using a chatbot for it produces ratings that track real-life benefit–risk trade-offs and disclosure decisions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This preregistered online experiment (N=1,388, Dutch LISS panel) uses a 2 (topic type: physical vs psychological, between) × 2 (topic sensitivity: low vs high, within) mixed design with sixteen scenarios to test how topic features and individual characteristics relate to perceived benefits/risks, intention to use AI chatbots for health questions, and willingness to self-disclose. H1/H2 are supported: benefits positively predict intention and disclosure (b ≈ 0.45–0.51), risks negatively (b ≈ −0.14 to −0.18). RQ1 finds only a small sensitivity effect on intention (higher for low-sensitive topics); topic type effects are multivariate but tiny (Pillai = .006) and largely non-significant univariately. RQ2 shows associations with AI experience/literacy, coping style, education, political orientation, and trust. The abstract and §5 conclude that intentions and disclosure are driven primarily by benefit–risk perceptions and personal characteristics rather than topic type.
Significance. The paper supplies timely, large-scale evidence on public benefit–risk trade-offs for LLM chatbots in health, using a representative sample, preregistration, power analysis, measurement-invariance checks, and linear mixed models with alpha correction. Strengths include systematic comparison of physical/psychological and low/high-sensitivity topics (often studied in isolation), open materials/code on OSF, and explicit linkage to UTAUT, privacy calculus, and HBM. If the directional findings hold, they usefully inform designers and policymakers that individual factors and perceived benefits dominate topic framing, while highlighting low overall intention/disclosure. The scenario design limits external validity, but the authors already flag this; the work remains a solid empirical contribution for HCI/health communication.
major comments (2)
- §4.2 and §5.1: The manipulation check is only partially successful—psychological conditions were rated more sensitive than physical ones overall, and residual severity differences remain. The central claim that topic type is unimportant therefore rests on a confounded contrast. Either reframe RQ1 conclusions more cautiously (topic type cannot be cleanly isolated) or report sensitivity-matched subgroup analyses / covariate-adjusted models that partial out residual sensitivity/severity before asserting negligible topic-type effects.
- §3.2 Procedure / §5.1: All outcomes are scenario-based ratings of imagined chatbot use for assigned conditions that many participants have never experienced. This is the load-bearing external-validity assumption for the claim that benefits/risks and individual factors (not topic) drive real intention and disclosure. The limitation is acknowledged, but the manuscript should quantify how many participants had personal experience with each condition and test whether experience moderates the benefit/risk coefficients; without that, the strongest claim over-reaches the design.
minor comments (5)
- Section numbering is inconsistent (§3.1 Pretest and §3.2 Procedure appear after §3.3 Stimuli; later subsections restart at 3.1). Renumber for clarity.
- §3.3.2.1: Monitoring coping style α = .48 is poor; the decision to enter items separately is correct but should be flagged earlier as a measurement limitation.
- Figure 5 / §4.4: Report exact means, SDs, and effect sizes for all four DVs by condition in a table, not only the forest/ANOVA summary, so readers can judge practical significance of the tiny Pillai value.
- §4.5 / Figure 6: Several coefficients (e.g., physiotherapist visits b = −0.64) look large relative to scale; confirm standardisation and units in the figure caption.
- Typos: “Chronbach’s” (multiple places), “UTUAT” vs UTAUT, and occasional missing spaces around × symbols.
Circularity Check
No circularity: empirical associations from a preregistered mixed experiment; benefits/risks are measured predictors, not redefined outcomes.
full rationale
This is a standard social-science experiment (2×2 mixed design, N=1,388) testing preregistered hypotheses H1–H2 and RQs about associations between measured perceived benefits/risks, topic type/sensitivity, individual characteristics, and the outcomes intention-to-use and willingness-to-self-disclose. Benefit and risk items are adapted from external taxonomies (Antes et al. 2021; Weidinger et al. 2022) and used as predictors in linear mixed models; they are not fitted to the outcomes and then re-presented as predictions. Topic effects are tested via MANOVA/ANOVA and found negligible except for a small sensitivity effect on intention. Individual-difference models are exploratory associations, not uniqueness claims. Self-citations are limited to data sources (LISS panel, AlgoSoc wave) and the authors’ own OSF preregistration/materials; none supply a load-bearing theoretical uniqueness theorem or ansatz that forces the central claim. The scenario-based design is an external-validity limitation already acknowledged by the authors, not a circular derivation. No equation, coefficient, or conclusion reduces by construction to its own inputs. Score 0 is therefore appropriate.
Assumptions & free parameters
assumptions (4)
- domain assumption Scenario-based self-reports of intention and willingness approximate real-world chatbot use and disclosure decisions.
- domain assumption Perceived benefits and risks can be measured as averaged multi-item Likert scales adapted from Antes et al. (2021) and Weidinger et al. (2022).
- standard math Linear mixed-effects models with participant random intercepts and fixed effects for benefits/risks/covariates correctly capture the associations of interest.
- ad hoc to paper Pretest-selected health conditions adequately operationalise low vs high sensitivity within physical and psychological categories despite residual sensitivity differences across topic type.
Cite this review
Pith. "Pith review of LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics." pith.science (2026). https://pith.science/paper/NOH2DOBA
@misc{pith2026260709253,
author = {Pith},
title = {Pith review of: LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics},
year = {2026},
howpublished = {\url{https://pith.science/paper/NOH2DOBA}},
note = {Machine review of arXiv:2607.09253}
}
read the original abstract
AI chatbots are increasingly used for answering health-related questions. This study examines the role of topic type discussed with an AI chatbot and individual characteristics on perceived benefits and risks, intention to use an AI chatbot, and willingness to self-disclose health information. We conducted an online experiment with a 2 (topic type: physical versus psychological, between-subjects) x 2 (topic sensitivity: low versus high, within-subjects) mixed design among a Dutch representative sample (N = 1,388). Results showed that perceived benefits were positively associated with intention and willingness to self-disclose, while perceived risks were negatively associated. Moreover, participants reported higher usage intentions for low-sensitive topics compared to high-sensitive topics. Furthermore, perceptions, intention, and willingness to self-disclose varied by individual characteristics. Overall, our findings suggest that intentions to use AI chatbots and self-disclosure of health-related information are primarily related to perceived benefits and risks and to personal characteristics rather than to topic type.
Reference graph
Works this paper leans on
-
[1]
To what extent would you describe the health condition as sensitive to discuss?
1 LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics Gwenn Beets1*, Anniek Jansen1*, Saar Hommes1,2, Ruben D. Vromans1, Leonie Weste...
2024
-
[2]
It is a benefit that the AI chatbot… can answer an unlimited number of questions
(▲). 3.3.1 Dependent variables Measurement invariance was assessed and confirmed for all dependent variables (see OSF Appendix G). Additionally, exploratory factor analyses were conducted for both benefits and risks scales (OSF Appendix H). 3.3.1.1 Perceived benefits ♣ Perceived benefits were measured by asking participants to rate the extent to which the...
2021
-
[3]
These models all returned an error indicating that the number of observations <= number of random effects and was therefore not considered for best fitting model
1 A fourth step of fitting a random slopes model which included random slopes topic type*topic sensitivity was also fitted. These models all returned an error indicating that the number of observations <= number of random effects and was therefore not considered for best fitting model. 10 LLMs for health: Perceived benefits, risks, intention to use AI cha...
2025
-
[4]
It listens better than my therapist
and the privacy calculus (Laufer and Wolfe, 1977), which describe a trade-off between perceived benefits and risks. Notably, effect sizes for benefits outweighed those of risks in shaping behaviours regarding AI chatbot use. A possible explanation may be that certain risks of AI use do not directly affect the user, or they may not have experienced the ris...
-
[5]
Kieskompas (2023) Het politieke landschap. Available at: https://tweedekamer2023.kieskompas.nl/nl/results/compass (accessed 07/04/2026). Kowalski RM and Peipert A (2019) Public- and self-stigma attached to physical versus psychological disabilities. Stigma and Health 4(2): 136-142. Laufer RS and Wolfe M (1977) Privacy as a concept and a social issue: A mu...
-
[6]
Chatting with ChatGPT
Menon D and Shilpa K (2023) “Chatting with ChatGPT”: Analyzing the factors influencing users' intention to use the open ai's ChatGPT using the UTAUT model. Heliyon 9(11): e20962. Miles O, West R and Nadarzynski T (2021) Health chatbots acceptability moderated by perceived stigma and severity: A cross-sectional survey. DIGITAL HEALTH 7: 205520762110630. Ni...
2023
-
[7]
Journal of Medical Internet Research 8(4): e507
Norman CD and Skinner HA (2006) eHEALS: The eHealth literacy scale. Journal of Medical Internet Research 8(4): e507. Nyakhar S and Wang H (2025) Effectiveness of artificial intelligence chatbots on mental health & well-being in college students: A rapid systematic review. Frontiers in Psychiatry
2006
-
[8]
Personality and Social Psychology Review 4(2): 174-185
Omarzu J (2000) A disclosure decision model: Determining how and when individuals will self-disclose. Personality and Social Psychology Review 4(2): 174-185. Platt J, Nong P, Smiddy R, et al. (2024) Public comfort with the use of ChatGPT and expectations for healthcare. Journal of the American Medical Informatics Association 31(9): 1976-1982. R Core Team ...
2000
Show all 9 references
-
[9]
Strathman A, Gleicher F, Boninger DS, et al
Available at: https://www.cbs.nl/nl-nl/onze-diensten/methoden/classificaties/onderwijs-en- beroepen/standaard-onderwijsindeling--soi--/standaard-onderwijsindeling-2021 (accessed 19/05/2026). Strathman A, Gleicher F, Boninger DS, et al. (1994) The consideration of future conseq...
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.