{"id":"8dcc5e04-55d0-443d-adb5-d597a7c1d452","arxiv_id":"2607.09253","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Intention to use AI health chatbots and willingness to self-disclose track perceived benefits/risks and individual traits more than physical-vs-psychological topic type, with only a small sensitivity effect on intention.","lead":"A large Dutch survey experiment finds that people’s willingness to use AI health chatbots and share personal data tracks perceived benefits and risks plus personal traits far more than whether the topic is physical or psychological. The result matters for anyone designing or regulating health AI, because it points to user literacy and experience rather than topic stigma as the main levers.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the scenario-validity caveat already flagged by the reader.","rationale":"The reader’s weakest-assumption diagnosis matches the paper’s own limitation statement and is the only soft spot that could affect the strongest claim’s external reach. The statistical support for the claim (large N, preregistration, measurement invariance, consistent benefit/risk associations across models) is solid; residual sensitivity confounds between physical and psychological topics are already noted and do not reverse the “primarily not topic type” conclusion. No further load-bearing attack is warranted. Verdict remains CONDITIONAL pending data release, exactly as the reader concluded.","tokens_in":15731,"tokens_out":379,"duration_ms":4063,"concrete_test":"Once the LISS Archive data are public, re-estimate the linear mixed models for intention and self-disclosure after restricting the sample to participants who reported personal experience with the assigned health condition; if the benefit/risk coefficients remain within 10% of the published values and the topic-type effect stays non-significant, the scenario-validity concern does not materially weaken the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s central claim—that intentions and self-disclosure are driven primarily by perceived benefits/risks and individual characteristics rather than topic type—is supported by the reported coefficients (benefits b≈0.45–0.51, risks b≈−0.14–−0.18) and the negligible topic-type multivariate effect (Pillai=.006). The scenario-based design is the weakest link for external validity, but the authors already acknowledge this limitation in §5.1 and the reader correctly identifies it as the weakest assumption. No additional internal inconsistency, statistical error, or unacknowledged confound rises to a load-bearing concern that would overturn the directional findings.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This preregistered online experiment (N=1,388, Dutch LISS panel) uses a 2 (topic type: physical vs psychological, between) × 2 (topic sensitivity: low vs high, within) mixed design with sixteen scenarios to test how topic features and individual characteristics relate to perceived benefits/risks, intention to use AI chatbots for health questions, and willingness to self-disclose. H1/H2 are supported: benefits positively predict intention and disclosure (b ≈ 0.45–0.51), risks negatively (b ≈ −0.14 to −0.18). RQ1 finds only a small sensitivity effect on intention (higher for low-sensitive topics); topic type effects are multivariate but tiny (Pillai = .006) and largely non-significant univariately. RQ2 shows associations with AI experience/literacy, coping style, education, political orientation, and trust. The abstract and §5 conclude that intentions and disclosure are driven primarily by benefit–risk perceptions and personal characteristics rather than topic type.","tokens_in":15895,"tokens_out":1034,"duration_ms":9467,"significance":"The paper supplies timely, large-scale evidence on public benefit–risk trade-offs for LLM chatbots in health, using a representative sample, preregistration, power analysis, measurement-invariance checks, and linear mixed models with alpha correction. Strengths include systematic comparison of physical/psychological and low/high-sensitivity topics (often studied in isolation), open materials/code on OSF, and explicit linkage to UTAUT, privacy calculus, and HBM. If the directional findings hold, they usefully inform designers and policymakers that individual factors and perceived benefits dominate topic framing, while highlighting low overall intention/disclosure. The scenario design limits external validity, but the authors already flag this; the work remains a solid empirical contribution for HCI/health communication.","major_comments":[{"comment":"§4.2 and §5.1: The manipulation check is only partially successful—psychological conditions were rated more sensitive than physical ones overall, and residual severity differences remain. The central claim that topic type is unimportant therefore rests on a confounded contrast. Either reframe RQ1 conclusions more cautiously (topic type cannot be cleanly isolated) or report sensitivity-matched subgroup analyses / covariate-adjusted models that partial out residual sensitivity/severity before asserting negligible topic-type effects.","section":null},{"comment":"§3.2 Procedure / §5.1: All outcomes are scenario-based ratings of imagined chatbot use for assigned conditions that many participants have never experienced. This is the load-bearing external-validity assumption for the claim that benefits/risks and individual factors (not topic) drive real intention and disclosure. The limitation is acknowledged, but the manuscript should quantify how many participants had personal experience with each condition and test whether experience moderates the benefit/risk coefficients; without that, the strongest claim over-reaches the design.","section":null}],"minor_comments":[{"comment":"Section numbering is inconsistent (§3.1 Pretest and §3.2 Procedure appear after §3.3 Stimuli; later subsections restart at 3.1). Renumber for clarity.","section":null},{"comment":"§3.3.2.1: Monitoring coping style α = .48 is poor; the decision to enter items separately is correct but should be flagged earlier as a measurement limitation.","section":null},{"comment":"Figure 5 / §4.4: Report exact means, SDs, and effect sizes for all four DVs by condition in a table, not only the forest/ANOVA summary, so readers can judge practical significance of the tiny Pillai value.","section":null},{"comment":"§4.5 / Figure 6: Several coefficients (e.g., physiotherapist visits b = −0.64) look large relative to scale; confirm standardisation and units in the figure caption.","section":null},{"comment":"Typos: “Chronbach’s” (multiple places), “UTUAT” vs UTAUT, and occasional missing spaces around × symbols.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Solid, well-powered preregistered work that fits cs.HC / health-communication venues. The two major points are addressable with re-analysis and tighter language rather than new data; I would not require a full redesign. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean, large-N preregistered mixed experiment (N=1,388 Dutch LISS panel) that systematically crosses physical vs psychological topic type with low vs high sensitivity and then models benefits, risks, intention, and willingness to self-disclose plus a broad set of individual covariates. The headline result holds: benefits (b≈0.45–0.51) and risks (b≈−0.14–−0.18) drive intention and disclosure far more than topic type; the multivariate topic-type effect is tiny (Pillai=.006) and the only reliable univariate hit is a small sensitivity effect on intention (higher for low-sensitive topics). Individual factors—prior chatbot experience, AI literacy, monitoring coping style, education, political orientation—matter in expected directions.\n\nWhat is new is the joint factorial design plus the representative sample and the simultaneous individual-difference model. Prior work already had stigma, privacy calculus, and UTAUT pieces; this paper puts them in one design and shows topic type is not the main lever. Methods are careful: power analysis, measurement invariance, exploratory factor analyses, alpha correction, mixed models, OSF prereg and materials. Self-citation is limited to the LISS/AlgoSoc data sources. No circularity or invented constructs.\n\nSoft spots are real but proportionate. The manipulation check only partially succeeded—psychological conditions were rated more sensitive overall, so physical/psychological comparisons are confounded with residual sensitivity. Effect sizes are small. The biggest external-validity limit is the scenario method (imagine you have condition X and use a chatbot); the authors flag this themselves. Full LISS data release is still pending. None of these overturn the directional claims.\n\nThis is useful for health-communication and HCI people who need evidence on whether to prioritise literacy/experience over topic-specific chatbot design. It does not open a new scientific frontier, but it is honest, well-powered, and citable for the benefit–risk and individual-difference findings. I would send it to peer review; a methods-aware referee can handle the residual confounds and scenario caveats. Worth reading if you work on health AI adoption; skip if you only care about model architecture or clinical outcomes.","headline":"Solid preregistered factorial survey: benefits/risks and individual traits dominate topic type for chatbot health use; scenario method is the main external-validity limit.","tokens_in":16493,"tokens_out":548,"would_cite":true,"duration_ms":6476,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"People's willingness to use health AI chatbots tracks benefits, risks, and personal traits far more than whether the topic is physical or psychological.","keywords":["AI chatbots","Large Language Models","topic sensitivity","perceived benefits","perceived risks","health communication","self-disclosure","intention to use"],"falsifier":"A field study that logs actual chatbot conversations and subsequent disclosure or continued use for matched physical versus psychological and low- versus high-sensitivity topics, then checks whether topic type still fails to predict behaviour once benefits, risks, and the same individual covariates are controlled.","tokens_in":16641,"feed_emoji":"💬","tokens_out":673,"duration_ms":5986,"temperature":0.7,"pith_summary":"This paper tests whether the kind of health topic people discuss with an AI chatbot—physical versus psychological, low versus high sensitivity—shapes how useful or dangerous the chatbot seems, whether they would use it, and whether they would share personal health details. In a large representative Dutch experiment, the main drivers were not the topic labels themselves. Higher perceived benefits raised intention to use and willingness to disclose; higher perceived risks lowered both. Intention was modestly higher for low-sensitivity topics than high-sensitivity ones, but topic type otherwise did little. Experience with AI chatbots, AI literacy, education, political orientation, trust in institutions, and information-seeking coping styles all moved the outcomes. The practical upshot is that adoption of health chatbots will depend more on how people weigh gains and harms and on who they already are than on which disease category is on the table.","feed_headline":"Health chatbot use tracks benefits and risks, not topic type","feed_subtitle":"In a Dutch sample of 1,388, personal traits and trade-offs matter more than physical vs psychological labels","key_machinery":"A mixed factorial design (topic type between-subjects × topic sensitivity within-subjects) that presents short scenarios of AI-chatbot interaction for pretested health conditions, then measures perceived benefits, risks, intention, and willingness to self-disclose, with linear mixed models linking those outcomes to both experimental factors and individual covariates.","core_discovery":"In a 2\times2 mixed experiment with a Dutch representative sample (N = 1,388), perceived benefits positively predicted intention to use an AI chatbot and willingness to self-disclose health information, while perceived risks negatively predicted both. Topic type (physical vs psychological) had negligible univariate effects; the only reliable topic effect was slightly higher usage intention for low-sensitivity than high-sensitivity scenarios. Individual characteristics—prior chatbot use, AI literacy, education, political orientation, institutional trust, and monitoring coping style—also systematically shifted perceptions and intentions. The authors conclude that chatbot use and disclosure for","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Benefits and risks, not topic type, drive health chatbot use","Personal traits beat topic labels for AI health chatbots","Self-disclosure to health AIs tracks risk-benefit trade-offs","Usage intention higher for low-sensitive topics only","Dutch study: Benefits, not physical vs psych, shape chatbot intent"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That asking people to imagine having a given health condition and using a chatbot for it produces ratings that track real-life benefit–risk trade-offs and disclosure decisions.","fun_headline_variants_meta":{"raw":{"variants":["Benefits and risks, not topic type, drive health chatbot use","Personal traits beat topic labels for AI health chatbots","Self-disclosure to health AIs tracks risk-benefit trade-offs","Usage intention higher for low-sensitive topics only","Dutch study: Benefits, not physical vs psych, shape chatbot intent"]},"model":"grok-4.5","effort":"low","cost_usd":0.004006,"raw_usage":{"total_tokens":1255,"prompt_tokens":785,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":40060000,"prompt_tokens_details":{"text_tokens":785,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":403,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":785,"tokens_out":67,"duration_ms":4332,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T04:22:09.666446+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A field study that logs actual chatbot conversations and subsequent disclosure or continued use for matched physical versus psychological and low- versus high-sensitivity topics, then checks whether topic type still fails to predict behaviour once benefits, risks, and the same individual covariates are controlled.","supporting_citations":[],"review_version":1}