{"id":"c05fe621-a46a-4cd7-8175-b092003d0880","arxiv_id":"2501.15028","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Chatbot symptom checkers that prompt users to reflect on specific evidence or consider when symptoms do not occur reduce the perceived influence of social media and prevent overestimation of ADHD symptoms.","lead":"Two experiments with about 200 participants tested whether social media posts about adult ADHD distort people's self-diagnosis and whether chatbot-style symptom checkers can correct that distortion. The paper is a design study: it proposes and evaluates chatbot features that push users to reflect on their own evidence instead of relying on recent, resonant social media content.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed mitigation by Evidence Reflection and Counterfactual Thinking strategies rests on unverified chatbot fidelity; Section 6's single adherence assertion without transcript audit cannot rule out generic conversational effects.","rationale":"The reader identified the same load-bearing concern: the paper's assertion that GPT-4 adhered to the intervention instructions is not backed by systematic fidelity evidence. This is the most direct threat to the central claim, because the contribution is specifically about integrating cognitive intervention strategies into CSCs. If the chatbots did not implement those strategies as designed, the observed effects cannot be attributed to Evidence Reflection or Counterfactual Thinking, and the paper's design contribution collapses. I agree with the reader's CONDITIONAL verdict: the study is promising and the direction is plausible, but the missing fidelity check is an addressable weakness. A transcript audit or a scripted re-run of the prompts would settle the concern. I also note a secondary statistical concern: the inattention-score mitigation is supported only by within-group null results rather than a direct group-by-time interaction test, but that is less central to the main attribution claim and the reader's verdict already accounts for the null-result reliance. The proposed concrete test is feasible if logs exist or if the same prompt system is still available, and it directly targets the causal interpretation.","tokens_in":25149,"tokens_out":6474,"duration_ms":67854,"concrete_test":"Audit the raw chat logs from Study 2 (or re-run the identical GPT-4 prompts with scripted user turns) and have two independent coders rate each exchange for strategy-specific behaviors: for ER, requests for concrete symptom evidence and appropriate bypass when sufficient context was given; for CT, counterfactual questions about non-occurrence of symptoms. Report adherence rates per participant; if adherence is not high (e.g., below 80% of required interventions), rerun the primary analyses restricted to adherent sessions. This settles whether the observed mitigation is attributable to the cognitive interventions rather than to generic conversational effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Study 2 is that the two cognitive intervention strategies, Evidence Reflection (ER) and Counterfactual Thinking (CT), mitigated availability bias more than a plain chatbot or a static questionnaire. This attribution requires that the GPT-4-based chatbots actually executed the intended strategies. The only support is one sentence in Section 6: \"As the prompts used in the chatbot development, the AI agent adhered to our instructions.\" No transcripts, no fidelity coding, no count of deviations, and no inter-rater assessment of dialogue behavior are provided. Without such evidence, the observed reductions in social media influence and the absence of significant inattention-score increases in the ER and CT groups could be produced by more generic features of the interaction, such as longer time on task, the obligation to type detailed answers, or the perceived need to appear reflective. The comparison against the plain CSC (which showed no significant reduction in social media influence) narrows but does not eliminate this alternative, because the cognitive interventions also changed the length and structure of the dialogue. A load-bearing assumption of the paper is therefore unsupported: that the ER and CT conditions differed from the CSC condition in the intended cognitive mechanism and not merely in conversational style or effort.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses availability bias in online self-diagnosis. Study 1 (N=104) compares a control condition with exposure to neutral and exaggerated social media posts about adult ADHD, finding that neutral content increases self-reported social media influence and symptom overestimation, while exaggerated content does not. Study 2 (N=100) compares a static questionnaire, a plain chatbot-based symptom checker (CSC), a CSC with an Evidence Reflection strategy, and a CSC with a Counterfactual Thinking strategy, reporting that the two cognitive-intervention CSCs reduce self-reported social media influence and prevent significant increases in inattention scores. The authors interpret these results as evidence that CSCs with cognitive interventions mitigate availability bias, and they discuss design implications for online diagnostic tools and social media platforms.","tokens_in":25351,"tokens_out":5057,"duration_ms":46843,"significance":"If the results hold, the paper makes a practical contribution to the design of chatbot-based symptom checkers and provides useful evidence about how resonant social media content triggers availability bias. The work has clear strengths: explicit hypotheses, power calculations, established instruments (SNAP-IV, ASRS, NVS, SRIS, NASA-TLX), manipulation checks, inter-coder reliability for qualitative coding, and two complementary studies. The finding that neutral, relatable content can be more influential than exaggerated content is interesting and well supported by the qualitative data. However, the central quantitative claims are currently weakened by several internal statistical inconsistencies, and the attribution of Study 2's effects to the specific cognitive interventions rests on an unverified assumption about chatbot fidelity.","major_comments":[{"comment":"The text and Table 1 report conflicting p-values for the pairwise comparisons that support H1.a and H2.a. In §4.2, the Control-vs-Neutral comparison for social media influence is reported as p = 0.22* with Cohen's D = -0.65, while Table 1 gives p = 0.022 for the same comparison; these values lead to opposite conclusions at the 0.05 level. In §4.3.1, the Control-vs-Neutral inattention comparison is reported as p < 0.001 in the text but p = 0.012 in Table 1, and the hyperactivity comparison is reported as p = 0.05* in the text but p = 0.47 in Table 1. Because the support for H1.a and H2.a depends directly on these comparisons, the authors must reconcile the reported statistics and restate which hypotheses are actually supported. The same issue appears in the §4.1 manipulation checks (p = 0.008 vs p = 0.004 for accuracy; p = 0.016 vs p = 0.009 for trustworthiness).","section":"§4.2 and Table 1"},{"comment":"The claim that the Evidence Reflection and Counterfactual Thinking strategies, rather than the general conversational properties of the chatbot, drove the observed reductions in social media influence rests on the single assertion that \"the AI agent adhered to our instructions.\" No transcripts, fidelity coding, deviation counts, or inter-rater assessment of dialogue behavior are provided. Because the ER and CT conditions also changed the length, structure, and required response format of the interaction compared with the plain CSC, the intended cognitive mechanisms are confounded with conversational style and effort. The authors should provide a fidelity audit of a sample of conversations (for example, adherence rates per strategy component and representative transcripts in an appendix), or explicitly weaken the causal attribution to the specific strategies.","section":"§6, first paragraph; §5.1.3 and §5.1.4"},{"comment":"The conclusion that ER and CT \"were effective in addressing the overestimation of inattention scores\" is based on the absence of a statistically significant within-group increase in those conditions, not on a significant difference from the control or plain CSC conditions. A null result in paired tests of this size is weak evidence of effectiveness, especially when the comparable between-condition comparisons are not reported. The authors should either report equivalence bounds or Bayes factors for the pre-post changes, or soften the claim to state that no significant overestimation was detected rather than that the interventions were effective.","section":"§6.2"}],"minor_comments":[{"comment":"The sentence \"To compare the outcomes of the four types of health information\" appears to refer to the three experimental conditions in Study 1; please correct the count.","section":"§3.6"},{"comment":"The heading \"Exaggerated content did not led to overestimation of symptoms\" contains a grammatical error; \"led\" should be \"lead.\"","section":"§4.3.1"},{"comment":"The phrase \"an diagnostic result\" should be \"a diagnostic result.\"","section":"§5.1.1"},{"comment":"The sentence \"This indicates that both treatments with cognitive strategy was effective\" has subject-verb agreement problems; it should read \"both treatments with cognitive strategies were effective.\"","section":"§6.2"},{"comment":"Effect sizes are reported inconsistently as \"Cohen's D\" in some places and \"Cohen's d\" in others; please standardize the notation and use the same form in text, tables, and figures.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and relevant HCI topic and the two-study design is appropriate. The main concerns are fixable in revision: reconciling the reported statistics, adding fidelity evidence for the chatbot interventions, and softening the interpretation of null within-group results. I do not see grounds for rejection if the authors can supply this missing support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper: the authors set out expecting exaggerated health posts to cause the most bias in self-diagnosis, and instead found neutral, resonant posts do. That's a real result, and the qualitative data (people saying the exaggerated posts were too far from their experience) makes it credible. The second study is more of a mixed bag but still worth reading: chatbot-based symptom checkers with evidence-reflection and counterfactual-thinking prompts lowered self-reported social media influence and stopped the significant ADHD inattention score jump that the static questionnaire and plain chatbot produced.\n\nWhat it does well: two clean between-subjects experiments, a sensible choice of adult ADHD as the target condition, validated scales, manipulation checks on the posts, and an honest limitations section. The authors do not hide that H1.b/H2.b failed. That earns credit.\n\nWhere it slips: first, several reported p-values do not match the tables. In Study 1, Section 4.2 says p=0.22* for neutral vs control on social media influence; Table 1 says p=0.022. The hyperactivity comparison says p=0.05* in the text but Table 1 shows C vs N p=0.47. These are not small cosmetic issues; the latter would change the conclusion from 'significant' to 'not significant'. Second, the load-bearing claim of Study 2 is that the two cognitive interventions worked via the intended mechanisms. The only evidence is a single sentence in Section 6 claiming the AI agent adhered to instructions. No transcripts, no fidelity coding, no audit. That cannot rule out the possibility that simply having a longer, more conversational exchange or feeling obliged to type detailed answers did the work. The mental effort data actually points that way: both intervention groups report significantly more mental effort than control. Third, the authors treat the null change in inattention scores for ER and CT as evidence of mitigation. With N=25 per group, a null is weak support; the real positive evidence is the social media influence measure.\n\nFor a reader: the Study 1 finding is the takeaway. Study 2 is a promising design direction but not a definitive proof. The paper deserves peer review if the authors can fix the statistical reporting and provide some fidelity evidence for the chatbot conditions. Otherwise the central mechanism claim stays shaky.\n\nRecommendation: send to review, but flag that the chatbot fidelity issue and the p-value mismatches need to be addressed before acceptance.","headline":"A genuinely interesting counterintuitive finding about neutral content and availability bias, wrapped in a study whose central intervention claim relies on unverified chatbot behavior.","tokens_in":25874,"tokens_out":1961,"would_cite":true,"duration_ms":18821,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that chatbot symptom checkers that ask users to justify their symptoms with concrete evidence, or to think about times the symptom is absent, can counteract the symptom overestimation that follows exposure to relatable…","keywords":["availability bias","online self-diagnosis","chatbot symptom checker","cognitive intervention","evidence reflection","counterfactual thinking","social media health information","adult ADHD"],"falsifier":"Record and code the actual chatbot turns: if an audit shows that the Evidence Reflection bot rarely asked for concrete evidence, or that the Counterfactual Thinking bot rarely posed absence-of-symptom questions, then the observed reductions in social-media influence and the absence of inattention-score inflation cannot be attributed to those cognitive interventions; they could be due to conversational engagement or longer time on task.","tokens_in":24893,"feed_emoji":"🧠","tokens_out":7050,"duration_ms":61930,"temperature":0.7,"pith_summary":"This paper tries to establish that availability bias from social media can be mitigated in online self-diagnosis by chatbot-based symptom checkers that prompt users to reflect on evidence instead of relying on how easily examples come to mind. The authors ran two experiments using adult ADHD as the test case. In the first, people who read factual, relatable posts about ADHD symptoms rated their own symptoms higher than controls, while exaggerated posts did not have this effect because readers distrusted them. In the second, a static questionnaire and a plain chatbot both left the social-media-induced symptom inflation in place, whereas chatbots that asked users to give concrete evidence for symptoms, or to consider when symptoms do not occur, removed the effect and lowered self-reported social media influence. If right, this gives a concrete design recipe: conversational agents that force evidence-based reflection can guard against a cognitive bias that otherwise distorts self-diagnosis.","feed_headline":"Chatbot reflection cuts social-media self-diagnosis bias","feed_subtitle":"Asking users to reflect on evidence or alternatives stops social-media symptom overestimation.","key_machinery":"The load-bearing mechanism is a conversational symptom checker built on GPT-4 that asks follow-up questions instead of presenting a fixed form. Evidence Reflection prompts the user, when they report a symptom, to describe the specific circumstances and concrete details supporting that report, until sufficient context is provided. Counterfactual Thinking prompts the user who says a symptom occurs to also consider how often and under what circumstances it does not occur, forcing them to compare against absence. Both strategies shift the user from System 1 heuristics, where whatever comes to mind dominates, to System 2 analytical reasoning, which is the theorized route through which availability bias is interrupted. The plain chatbot without these strategies served as a control to isolate the effect of conversation itself.","core_discovery":"On the paper's own terms, the central discovery is that availability bias in online self-diagnosis is triggered primarily by content resonance rather than exaggeration, and that the bias can be reversed at the point of self-assessment by cognitive-intervention question strategies. In Study 1, participants who read neutral, relatable social media posts about adult ADHD reported stronger social media influence on their symptom assessment and produced significantly higher inattention and hyperactivity scores than controls; participants who read exaggerated posts did not differ from controls. Qualitative responses showed the mechanism: resonant posts made people recognize their own experiences and recall similar symptoms, causing them to disregard their own evidence. In Study 2, a chatbot with an Evidence Reflection strategy and a chatbot with a Counterfactual Thinking strategy both significantly reduced the self-reported influence of social media relative to a static questionnaire, and eliminated the significant baseline-to-post increase in inattention scores that appeared in the static-questionnaire and plain-chatbot conditions. The authors conclude that CSCs with cognitive intervention strategies mitigate availability bias by guiding users into evidence-based reflective thinking.","pith_inferences":["An extension the authors leave implicit: recommendation algorithms that surface accurate, relatable health stories should be treated as a distinct bias risk, because Study 1 suggests resonance, not exaggeration, is what inflates self-assessed symptoms.","I would predict the same two chatbot strategies will suppress availability bias for other ambiguous, common conditions such as chronic Lyme, migraine, or irritable bowel syndrome, since the mechanism is the ease-of-recall heuristic rather than anything ADHD-specific.","A stricter test of the claimed mechanism would track whether the chatbot's follow-up questions change downstream behavior, such as actual symptom diaries or healthcare visits, rather than relying only on self-report scales, which can be affected by wanting to appear thoughtful to the bot."],"forward_implications":["Static questionnaire symptom checkers are a vulnerable format: in this study they left social-media-induced inattention-score inflation intact, so designers should not assume a well-validated scale alone protects users.","A plain conversational wrapper is not enough; only the two chatbots with active cognitive-intervention questions reduced social media influence and removed the inattention-score jump, pointing to the question design as the active ingredient.","Evidence Reflection and Counterfactual Thinking can be added to existing chatbot symptom checkers without changing the underlying medical questions, making them a cheap bias-mitigation layer for online self-diagnosis.","The bias reduction came with higher self-reported mental effort, so these designs will face a usability trade-off: slower, more demanding reflection versus faster but more biased self-assessment."],"supporting_citations":[{"why":"Defines the availability heuristic as judging likelihood by ease of recall, the bias this paper sets out to mitigate.","marker":"[82]"},{"why":"Provides the heuristic account of why people substitute easy recall for analytic judgment.","marker":"[83]"},{"why":"Basis for the Evidence Reflection strategy: ease of retrieval can be corrected by recalling concrete evidence.","marker":"[71]"},{"why":"Basis for the Evidence Reflection strategy: being accountable for a judgment makes people weigh facts more carefully.","marker":"[78]"},{"why":"Basis for the Counterfactual Thinking strategy: counterfactual reasoning disrupts heuristic thinking.","marker":"[34]"},{"why":"Basis for the Counterfactual Thinking strategy: imagining alternatives broadens the scenarios considered.","marker":"[32]"},{"why":"Supports the claim that AI questioning frameworks actively engage users' reasoning, motivating the chatbot format.","marker":"[17]"},{"why":"Provides the doctor-like probing and emotional-support design on which the plain chatbot condition was built.","marker":"[92]"},{"why":"Supplies the ASRS-v1.1 scale that measures ADHD symptom self-assessment in both studies.","marker":"[16]"},{"why":"Supplies the SNAP-IV items used to measure baseline ADHD levels in Study 1.","marker":"[24]"}],"fun_headline_variants":["Resonant social posts skew self-diagnosis; chatbots correct bias","Chatbot reflection counters social-media symptom overestimation","Evidence-based chatbot prompts reduce self-diagnosis bias","Counterfactual thinking in chatbots mitigates symptom bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim breaks if the GPT-4 chatbots did not actually follow their evidence-reflection and counterfactual-thinking scripts, since the study provides no transcript-level check of what the bots said.","fun_headline_variants_meta":{"raw":{"variants":["Resonant social posts skew self-diagnosis; chatbots correct bias","Chatbot reflection counters social-media symptom overestimation","Evidence-based chatbot prompts reduce self-diagnosis bias","Counterfactual thinking in chatbots mitigates symptom bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000229,"raw_usage":{"total_tokens":1450,"prompt_tokens":890,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":506,"tokens_out":560,"duration_ms":5283,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:41:12.318421+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record and code the actual chatbot turns: if an audit shows that the Evidence Reflection bot rarely asked for concrete evidence, or that the Counterfactual Thinking bot rarely posed absence-of-symptom questions, then the observed reductions in social-media influence and the absence of inattention-score inflation cannot be attributed to those cognitive interventions; they could be due to conversational engagement or longer time on task.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SNAP-IV items used to measure baseline ADHD levels in Study 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the availability heuristic as judging likelihood by ease of recall, the bias this paper sets out to mitigate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the heuristic account of why people substitute easy recall for analytic judgment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Basis for the Evidence Reflection strategy: ease of retrieval can be corrected by recalling concrete evidence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Basis for the Evidence Reflection strategy: being accountable for a judgment makes people weigh facts more carefully."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Basis for the Counterfactual Thinking strategy: counterfactual reasoning disrupts heuristic thinking."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Basis for the Counterfactual Thinking strategy: imagining alternatives broadens the scenarios considered."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the doctor-like probing and emotional-support design on which the plain chatbot condition was built."}],"review_version":1}