Pith. sign in

REVIEW 2 major objections 5 minor 9 references

LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics

T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read People's willingness to use health AI chatbots tracks benefits, risks, and personal traits far more than whether the topic is physical or psychological.

desk verdict Solid preregistered factorial survey: benefits/risks and individual traits dominate topic type for chatbot health use; scenario method is the main external-validity limit. read the letter →

arxiv 2607.09253 v1 pith:NOH2DOBA submitted 2026-07-10 cs.HC cs.AI

classification cs.HCcs.AI
keywords AIchatbotsLargeLanguageModelstopicsensitivityperceivedbenefitsriskshealthcommunicationself-disclosureintentiontouse
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests whether the kind of health topic people discuss with an AI chatbot—physical versus psychological, low versus high sensitivity—shapes how useful or dangerous the chatbot seems, whether they would use it, and whether they would share personal health details. In a large representative Dutch experiment, the main drivers were not the topic labels themselves. Higher perceived benefits raised intention to use and willingness to disclose; higher perceived risks lowered both. Intention was modestly higher for low-sensitivity topics than high-sensitivity ones, but topic type otherwise did little. Experience with AI chatbots, AI literacy, education, political orientation, trust in institutions, and information-seeking coping styles all moved the outcomes. The practical upshot is that adoption of health chatbots will depend more on how people weigh gains and harms and on who they already are than on which disease category is on the table.

What carries the argument

A mixed factorial design (topic type between-subjects × topic sensitivity within-subjects) that presents short scenarios of AI-chatbot interaction for pretested health conditions, then measures perceived benefits, risks, intention, and willingness to self-disclose, with linear mixed models linking those outcomes to both experimental factors and individual covariates.

What would settle it

A field study that logs actual chatbot conversations and subsequent disclosure or continued use for matched physical versus psychological and low- versus high-sensitivity topics, then checks whether topic type still fails to predict behaviour once benefits, risks, and the same individual covariates are controlled.

Watch

Extended reading notes

Core claim

In a 2 imes2 mixed experiment with a Dutch representative sample (N = 1,388), perceived benefits positively predicted intention to use an AI chatbot and willingness to self-disclose health information, while perceived risks negatively predicted both. Topic type (physical vs psychological) had negligible univariate effects; the only reliable topic effect was slightly higher usage intention for low-sensitivity than high-sensitivity scenarios. Individual characteristics—prior chatbot use, AI literacy, education, political orientation, institutional trust, and monitoring coping style—also systematically shifted perceptions and intentions. The authors conclude that chatbot use and disclosure for

Load-bearing premise

That asking people to imagine having a given health condition and using a chatbot for it produces ratings that track real-life benefit–risk trade-offs and disclosure decisions.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This preregistered online experiment (N=1,388, Dutch LISS panel) uses a 2 (topic type: physical vs psychological, between) × 2 (topic sensitivity: low vs high, within) mixed design with sixteen scenarios to test how topic features and individual characteristics relate to perceived benefits/risks, intention to use AI chatbots for health questions, and willingness to self-disclose. H1/H2 are supported: benefits positively predict intention and disclosure (b ≈ 0.45–0.51), risks negatively (b ≈ −0.14 to −0.18). RQ1 finds only a small sensitivity effect on intention (higher for low-sensitive topics); topic type effects are multivariate but tiny (Pillai = .006) and largely non-significant univariately. RQ2 shows associations with AI experience/literacy, coping style, education, political orientation, and trust. The abstract and §5 conclude that intentions and disclosure are driven primarily by benefit–risk perceptions and personal characteristics rather than topic type.

Significance. The paper supplies timely, large-scale evidence on public benefit–risk trade-offs for LLM chatbots in health, using a representative sample, preregistration, power analysis, measurement-invariance checks, and linear mixed models with alpha correction. Strengths include systematic comparison of physical/psychological and low/high-sensitivity topics (often studied in isolation), open materials/code on OSF, and explicit linkage to UTAUT, privacy calculus, and HBM. If the directional findings hold, they usefully inform designers and policymakers that individual factors and perceived benefits dominate topic framing, while highlighting low overall intention/disclosure. The scenario design limits external validity, but the authors already flag this; the work remains a solid empirical contribution for HCI/health communication.

major comments (2)
  1. §4.2 and §5.1: The manipulation check is only partially successful—psychological conditions were rated more sensitive than physical ones overall, and residual severity differences remain. The central claim that topic type is unimportant therefore rests on a confounded contrast. Either reframe RQ1 conclusions more cautiously (topic type cannot be cleanly isolated) or report sensitivity-matched subgroup analyses / covariate-adjusted models that partial out residual sensitivity/severity before asserting negligible topic-type effects.
  2. §3.2 Procedure / §5.1: All outcomes are scenario-based ratings of imagined chatbot use for assigned conditions that many participants have never experienced. This is the load-bearing external-validity assumption for the claim that benefits/risks and individual factors (not topic) drive real intention and disclosure. The limitation is acknowledged, but the manuscript should quantify how many participants had personal experience with each condition and test whether experience moderates the benefit/risk coefficients; without that, the strongest claim over-reaches the design.
minor comments (5)
  1. Section numbering is inconsistent (§3.1 Pretest and §3.2 Procedure appear after §3.3 Stimuli; later subsections restart at 3.1). Renumber for clarity.
  2. §3.3.2.1: Monitoring coping style α = .48 is poor; the decision to enter items separately is correct but should be flagged earlier as a measurement limitation.
  3. Figure 5 / §4.4: Report exact means, SDs, and effect sizes for all four DVs by condition in a table, not only the forest/ANOVA summary, so readers can judge practical significance of the tiny Pillai value.
  4. §4.5 / Figure 6: Several coefficients (e.g., physiotherapist visits b = −0.64) look large relative to scale; confirm standardisation and units in the figure caption.
  5. Typos: “Chronbach’s” (multiple places), “UTUAT” vs UTAUT, and occasional missing spaces around × symbols.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical associations from a preregistered mixed experiment; benefits/risks are measured predictors, not redefined outcomes.

full rationale

This is a standard social-science experiment (2×2 mixed design, N=1,388) testing preregistered hypotheses H1–H2 and RQs about associations between measured perceived benefits/risks, topic type/sensitivity, individual characteristics, and the outcomes intention-to-use and willingness-to-self-disclose. Benefit and risk items are adapted from external taxonomies (Antes et al. 2021; Weidinger et al. 2022) and used as predictors in linear mixed models; they are not fitted to the outcomes and then re-presented as predictions. Topic effects are tested via MANOVA/ANOVA and found negligible except for a small sensitivity effect on intention. Individual-difference models are exploratory associations, not uniqueness claims. Self-citations are limited to data sources (LISS panel, AlgoSoc wave) and the authors’ own OSF preregistration/materials; none supply a load-bearing theoretical uniqueness theorem or ansatz that forces the central claim. The scenario-based design is an external-validity limitation already acknowledged by the authors, not a circular derivation. No equation, coefficient, or conclusion reduces by construction to its own inputs. Score 0 is therefore appropriate.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard survey-experiment assumptions and adapted multi-item scales rather than free parameters or invented theoretical entities. No numerical constants are fitted to force the result; effect sizes are estimated from the data. Domain assumptions include the validity of scenario imagination and the partial success of the sensitivity manipulation.

assumptions (4)
  • domain assumption Scenario-based self-reports of intention and willingness approximate real-world chatbot use and disclosure decisions.
    Invoked throughout §3 Procedure and outcomes; authors note the limitation in §5.1.
  • domain assumption Perceived benefits and risks can be measured as averaged multi-item Likert scales adapted from Antes et al. (2021) and Weidinger et al. (2022).
    §3.3.1; internal consistencies reported as high.
  • standard math Linear mixed-effects models with participant random intercepts and fixed effects for benefits/risks/covariates correctly capture the associations of interest.
    §3.4 Statistical analyses; model-selection steps described.
  • ad hoc to paper Pretest-selected health conditions adequately operationalise low vs high sensitivity within physical and psychological categories despite residual sensitivity differences across topic type.
    §3.1 Pretest and §4.2 Manipulation check; authors acknowledge partial success.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics." pith.science (2026). https://pith.science/paper/NOH2DOBA

@misc{pith2026260709253,
  author       = {Pith},
  title        = {Pith review of: LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NOH2DOBA}},
  note         = {Machine review of arXiv:2607.09253}
}
read the original abstract

AI chatbots are increasingly used for answering health-related questions. This study examines the role of topic type discussed with an AI chatbot and individual characteristics on perceived benefits and risks, intention to use an AI chatbot, and willingness to self-disclose health information. We conducted an online experiment with a 2 (topic type: physical versus psychological, between-subjects) x 2 (topic sensitivity: low versus high, within-subjects) mixed design among a Dutch representative sample (N = 1,388). Results showed that perceived benefits were positively associated with intention and willingness to self-disclose, while perceived risks were negatively associated. Moreover, participants reported higher usage intentions for low-sensitive topics compared to high-sensitive topics. Furthermore, perceptions, intention, and willingness to self-disclose varied by individual characteristics. Overall, our findings suggest that intentions to use AI chatbots and self-disclosure of health-related information are primarily related to perceived benefits and risks and to personal characteristics rather than to topic type.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

9 extracted references · 1 canonical work pages

  1. [1]

    To what extent would you describe the health condition as sensitive to discuss?

    1 LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics Gwenn Beets1*, Anniek Jansen1*, Saar Hommes1,2, Ruben D. Vromans1, Leonie Weste...

  2. [2]

    It is a benefit that the AI chatbot… can answer an unlimited number of questions

    (▲). 3.3.1 Dependent variables Measurement invariance was assessed and confirmed for all dependent variables (see OSF Appendix G). Additionally, exploratory factor analyses were conducted for both benefits and risks scales (OSF Appendix H). 3.3.1.1 Perceived benefits ♣ Perceived benefits were measured by asking participants to rate the extent to which the...

  3. [3]

    These models all returned an error indicating that the number of observations <= number of random effects and was therefore not considered for best fitting model

    1 A fourth step of fitting a random slopes model which included random slopes topic type*topic sensitivity was also fitted. These models all returned an error indicating that the number of observations <= number of random effects and was therefore not considered for best fitting model. 10 LLMs for health: Perceived benefits, risks, intention to use AI cha...

  4. [4]

    It listens better than my therapist

    and the privacy calculus (Laufer and Wolfe, 1977), which describe a trade-off between perceived benefits and risks. Notably, effect sizes for benefits outweighed those of risks in shaping behaviours regarding AI chatbot use. A possible explanation may be that certain risks of AI use do not directly affect the user, or they may not have experienced the ris...

  5. [5]

    Talk to me, i’m secure

    Kieskompas (2023) Het politieke landschap. Available at: https://tweedekamer2023.kieskompas.nl/nl/results/compass (accessed 07/04/2026). Kowalski RM and Peipert A (2019) Public- and self-stigma attached to physical versus psychological disabilities. Stigma and Health 4(2): 136-142. Laufer RS and Wolfe M (1977) Privacy as a concept and a social issue: A mu...

  6. [6]

    Chatting with ChatGPT

    Menon D and Shilpa K (2023) “Chatting with ChatGPT”: Analyzing the factors influencing users' intention to use the open ai's ChatGPT using the UTAUT model. Heliyon 9(11): e20962. Miles O, West R and Nadarzynski T (2021) Health chatbots acceptability moderated by perceived stigma and severity: A cross-sectional survey. DIGITAL HEALTH 7: 205520762110630. Ni...

  7. [7]

    Journal of Medical Internet Research 8(4): e507

    Norman CD and Skinner HA (2006) eHEALS: The eHealth literacy scale. Journal of Medical Internet Research 8(4): e507. Nyakhar S and Wang H (2025) Effectiveness of artificial intelligence chatbots on mental health & well-being in college students: A rapid systematic review. Frontiers in Psychiatry

  8. [8]

    Personality and Social Psychology Review 4(2): 174-185

    Omarzu J (2000) A disclosure decision model: Determining how and when individuals will self-disclose. Personality and Social Psychology Review 4(2): 174-185. Platt J, Nong P, Smiddy R, et al. (2024) Public comfort with the use of ChatGPT and expectations for healthcare. Journal of the American Medical Informatics Association 31(9): 1976-1982. R Core Team ...

Show all 9 references
  1. [9]

    Strathman A, Gleicher F, Boninger DS, et al

    Available at: https://www.cbs.nl/nl-nl/onze-diensten/methoden/classificaties/onderwijs-en- beroepen/standaard-onderwijsindeling--soi--/standaard-onderwijsindeling-2021 (accessed 19/05/2026). Strathman A, Gleicher F, Boninger DS, et al. (1994) The consideration of future conseq...

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.