{"id":"8ad975b9-85e9-4bf4-a71b-f4b7a9f776b9","arxiv_id":"2607.06913","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"Lightweight LLMs lose 7.2 percentage points of accuracy when false health claims are injected into prompts, but only 1.4 points when medical jargon is replaced with everyday language.","lead":"This paper tests how small language models handle two types of realistic health prompts: misinformation injections and layperson terminology. It finds models are easily swayed by false claims but mostly survive informal language, highlighting deployment risks for health AI.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"MF degradation magnitude is confounded by TF-IDF claim-to-sample matching, which may artificially inflate topical relevance beyond what real user misinformation looks like.","rationale":"The reader correctly identified the TF-IDF matching realism as the weakest assumption. I agree this is the most load-bearing concern because it directly affects the magnitude of the headline finding (7.2 pp). However, I do not think this concern moves the verdict below CONDITIONAL. The paper is presented as an empirical benchmark study, and the finding that models are vulnerable to contextually-relevant misinformation is still valid even if the exact magnitude is an upper bound. The small sample size (n=100 per dataset) and lack of released code/data are additional valid concerns that support a CONDITIONAL rather than ACCEPT verdict, but the TF-IDF confound is the most technically substantive issue. The paper's contribution as an initial exploration of MF vs. LR perturbations in public health holds, provided future work validates the ecological validity of the perturbation methodology. The reader's assessment is well-calibrated.","tokens_in":7811,"tokens_out":548,"duration_ms":96659,"concrete_test":"Re-run the MF evaluation on one dataset (e.g., PubMedQA, n=100) using two matching conditions: (1) the existing TF-IDF retrieval, and (2) a random sampling of unsupported claims from the same source datasets. If the accuracy drop under random matching is significantly smaller than under TF-IDF matching, the 7.2 pp headline figure is inflated by the retrieval method and should be reported as an upper bound rather than a representative degradation rate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MF degrades accuracy by 7.2 pp relies on the realism of the perturbation. Section III.B.1 states that misinformation claims are matched to samples using TF-IDF retrieval. This creates a systematic bias: TF-IDF maximizes lexical overlap, meaning the injected false claims are likely topically closer to the question than a random or naturally-occurring piece of misinformation would be. If a user 'unintentionally carries misinformation into their queries' (as stated in the abstract), the false claim they introduce is unlikely to be as topically optimized as a TF-IDF-retrieved claim. This means the 7.2 pp degradation and 9-38% flip rate may reflect an upper bound on MF vulnerability rather than a representative real-world risk. The paper does not report the semantic relevance distribution of the matched claims or compare against a random-matching baseline, making it impossible to distinguish model vulnerability to misinformation from vulnerability to highly-salient contextual injection.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The manuscript evaluates the robustness of four lightweight LLMs (Llama-3.1-8B, Mistral-7B, Qwen2.5-7B, GPT-4.1-Nano) under two domain-specific prompt perturbations in public health settings: misinformation framing (MF) and layperson rewriting (LR). Using three datasets (PubMedQA, MedQA-USMLE, COVID-19 Vaccine Stance), the authors find that MF causes substantial accuracy degradation (−7.2 pp average) and high flip rates (9–38%), while LR has a comparatively minor effect (−1.4 pp). The paper concludes that these represent distinct deployment risks requiring perturbation-aware evaluation beyond clean benchmarks.","tokens_in":8483,"tokens_out":933,"duration_ms":145252,"significance":"The study addresses a practically important gap: most LLM evaluations in clinical or public health domains assume expert-authored, well-formed inputs, whereas real-world users introduce misinformation and informal language. The finding that explicit disclaimers do not fully mitigate MF vulnerability is actionable for deployment decisions. The comparison of open-source 7–8B models against a commercial lightweight model provides useful guidance for resource-constrained settings. However, the significance of the quantitative claims is tempered by the small sample size (100 examples per dataset) and the TF-IDF-based claim matching, which may not reflect naturally occurring user misinformation. The benchmark design is reproducible (frozen mapping file, deterministic decoding), which is a strength.","major_comments":[{"comment":"Section III.B.1: The TF-IDF retrieval method for matching misinformation claims to samples introduces a potential confound. TF-IDF maximizes lexical overlap, meaning injected claims are likely more topically salient than misinformation a real user would naturally introduce. This could inflate the 7.2 pp MF degradation relative to real-world conditions. The paper does not report the semantic relevance distribution of matched claims or compare against a random-matching baseline, making it difficult to distinguish model vulnerability to misinformation from vulnerability to highly-salient contextual injection. A random-matching control or a relevance analysis would substantially strengthen the central claim.","section":null},{"comment":"Section IV.A and Table II: The sample size of 100 examples per dataset (3,600 total inference records) is small for drawing robust conclusions, especially when broken down by model, dataset, and condition. For instance, Table II reports per-dataset accuracy changes in single-digit percentage points, which correspond to differences of only 1–4 examples. The confidence intervals in Table I are wide (e.g., Mistral-7B MF: 45.3 [39.8–51.0]), and some per-dataset comparisons in Table II likely lack statistical significance. The paper should either acknowledge this limitation more prominently or expand the evaluation to support the granularity of claims made in the discussion.","section":null},{"comment":"Section V.C, Table III, Case 2: The qualitative analysis of hedging as a failure mode is interesting but underdeveloped. The paper states that three models shift from 'yes' to 'maybe' under MF on PubMedQA, but it is unclear how prevalent this pattern is across the full dataset. Without quantifying how often MF induces hedging versus direct answer flips, the qualitative case may not be representative. A frequency count of this failure mode would help assess its importance.","section":null}],"minor_comments":[{"comment":"Section I: 'lightweighted' should be 'lightweight'.","section":null},{"comment":"Section V.A: 'In constrast' should be 'In contrast'.","section":null},{"comment":"Section V.A: 'Each dataset represents a distinct tasks' should be 'distinct tasks'.","section":null},{"comment":"Section VI: 'MF degradation of -$7.2 pp' contains a stray dollar sign.","section":null},{"comment":"Table III: The notation for prediction transitions (e.g., 'B→N×') is compact but could benefit from a clearer legend or expanded explanation in the caption.","section":null},{"comment":"Reference [19]: The access date '2025' is vague; a more specific date would improve reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about TF-IDF matching is valid and is the most important methodological issue. The paper's framing in the abstract ('users unintentionally carry misinformation into their queries') implies ecological validity that the TF-IDF matching may not support. The authors should reframe the MF results as potentially reflecting an upper bound on vulnerability, or provide evidence that the matched claims are representative of real-world user misinformation. The small sample size is a secondary but non-trivial concern for the granularity of the claims."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The headline finding: misinformation framing (MF) degrades LLM accuracy by 7.2 pp on average across four lightweight models, with 9–38% flip rates, even when injected claims carry an explicit disclaimer. Layperson rewriting (LR) causes only 1.4 pp degradation. The asymmetry is the useful result — models handle lexical variation fine but get swayed by false contextual content despite being told it's false. That's a real deployment risk for health AI. The paper does several things right. The perturbation taxonomy is grounded in realistic scenarios (non-clinical users injecting misinformation, using lay terminology). The metrics — accuracy drop and flip rate — are appropriate and the decoupling analysis (flip rate vs. accuracy drop) is a nice touch, showing GPT-4.1-Nano flips more correct predictions while Qwen2.5-7B churns between wrong answers. The per-dataset breakdown revealing that task structure matters (multiple-choice MedQA is more robust than open-ended Vaccine Stance) is a genuine insight. The stress-test concern about TF-IDF matching is legitimate and is the main soft spot. Section III.B.1 says misinformation claims are matched to samples via TF-IDF retrieval, which maximizes lexical overlap. This means injected claims are probably more topically salient than what a real user would accidentally introduce. The 7.2 pp figure likely reflects an upper bound on MF vulnerability, not a representative real-world risk. The paper doesn't report semantic relevance distributions or compare against random matching, so you can't fully separate 'model is vulnerable to misinformation' from 'model is vulnerable to highly-salient contextual injection.' That said, the disclaimer finding still holds regardless of matching method — models can't discount flagged false content even when warned. The sample size (100 per dataset) is small but adequate for a benchmark study. No code or data release is noted, which limits reproducibility. This paper is for practitioners choosing lightweight models for public health deployment and for researchers building robustness benchmarks. It deserves a serious referee who can push on the TF-IDF realism question and request a random-matching baseline. The core empirical contribution is sound enough to warrant review.","headline":"MF degrades accuracy 7.2 pp on average across four lightweight LLMs; LR barely matters. The MF number is likely an upper bound due to TF-IDF claim matching.","tokens_in":8579,"tokens_out":534,"would_cite":false,"duration_ms":56312,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Health AI models swallow misinformation even when warned","keywords":[],"falsifier":"If a model family showed no accuracy drop under misinformation framing despite dropping under layperson rewriting, the paper's central asymmetry claim would be contradicted. Alternatively, if models that received explicit disclaimers about injected claims showed no residual degradation, the claim that warnings are insufficient would fail.","tokens_in":7896,"feed_emoji":"🏥","tokens_out":1709,"duration_ms":81581,"temperature":0.7,"pith_summary":"The paper tests whether four lightweight LLMs maintain accuracy when their inputs are perturbed in ways that reflect real-world public health usage. Two perturbation types are compared: misinformation framing (injecting a false health claim into the prompt) and layperson rewriting (replacing medical terms with everyday language). The central finding is an asymmetry: models are largely resilient to informal vocabulary (1.4 pp average drop) but substantially vulnerable to injected misinformation (7.2 pp average drop, with 9–38% of predictions flipping), even when the false claims are explicitly labeled as unsupported. The paper argues this reveals a deployment risk distinct from the one usually discussed—models do not just fail to understand patients; they get swayed by false beliefs patients carry into their queries.","feed_headline":"Health AI models swallow misinformation even when warned","feed_subtitle":"Models handle informal patient language fine but fold under false claims—even when explicitly told the claims are unsupported","key_machinery":"Two perturbation functions applied to identical test prompts: misinformation framing concatenates a retrieved false health claim (from curated myth-busting datasets) into the prompt, while layperson rewriting applies term-level substitution using a consumer health vocabulary. Accuracy drop and flip rate are measured against a clean baseline across three public health tasks (biomedical QA, clinical reasoning, vaccine stance classification).","core_discovery":"The robustness gap between semantic-level false content injection and lexical-level vocabulary substitution is large and consistent across model families. Lightweight LLMs can bridge professional and lay medical terminology, but they overweight contextual false claims over their own parametric knowledge, even when warned. This means the dominant failure mode in public health AI deployment is not comprehension of informal speech but susceptibility to misinformation carried in by users.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Health AI follows misinformation even when warned","LLMs handle informal speech but fold under misinformation","Misinformation framing hurts health LLMs more than patient jargon","LLMs flip on misinformation even when warned","Misinformation hurts health LLM accuracy more than informal speech"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that its method of matching misinformation claims to test questions via keyword retrieval produces a realistic simulation of how actual non-clinical users introduce false information into their queries. If the injected claims are more topically salient or directly relevant than what real users would naturally include, the 7.2 pp degradation may overstate the real-world risk.","fun_headline_variants_meta":{"raw":{"variants":["Health AI follows misinformation even when warned","LLMs handle informal speech but fold under misinformation","Misinformation framing hurts health LLMs more than patient jargon","LLMs flip on misinformation even when warned","Misinformation hurts health LLM accuracy more than informal speech"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":560,"prompt_tokens":487,"completion_tokens":73,"prompt_tokens_details":null},"tokens_in":487,"tokens_out":73,"duration_ms":19265,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T23:00:31.096695+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If a model family showed no accuracy drop under misinformation framing despite dropping under layperson rewriting, the paper's central asymmetry claim would be contradicted. Alternatively, if models that received explicit disclaimers about injected claims showed no residual degradation, the claim that warnings are insufficient would fail.","supporting_citations":[],"review_version":1}