{"id":"1b3a79da-cb03-4d40-8a46-eadb1ea7322b","arxiv_id":"2505.10472","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"General-purpose LLMs outperformed specialized medical LLMs on linguistic quality and emotional engagement for breast and cervical cancer questions, while medical models were simpler to read but scored worse on safety.","lead":"The paper compares five general-purpose AI chatbots with three medical-specialist chatbots on breast and cervical cancer questions. It finds general models write higher-quality, more empathetic answers while medical models write simpler but less safe and less trustworthy text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed medical-vs-general safety/toxicity duality contradicts the paper's own Table 4: group means on toxicity are essentially equal and medical models have lower mean gender bias, so the central claim is not established.","rationale":"The reader's weakest-assumption concern was construct validity of the metrics (reference-based QA scores, hallucination entropy, PAIR reflection). That is a legitimate concern, but the more immediately load-bearing problem is internal: the paper's own reported numbers do not support the central group-level contrast. The toxicity group means are indistinguishable, gender-bias means point in the opposite direction, and reflection scores differ only in the third decimal. This is not a matter of needing more sophisticated statistics; a simple arithmetic mean from Table 4 already undermines the abstract. I also credit the paper's breadth—multi-dimensional framework, expert qualitative ratings with reported kappa, and explicit limitations—but that does not rescue the headline conclusion. The Meditron-vs-Llama3 contradiction noted by the reader and the unreported significance tests compound the problem, but the aggregation flaw alone is sufficient to reject the paper in its current form. The proposed check is deliberately cheap: recompute group comparisons from the published tables and run a jackknife to see whether the 'medical vs general' effect is driven by single outlier models. If the effect disappears, the central claim must be revised.","tokens_in":19606,"tokens_out":4555,"duration_ms":43727,"concrete_test":"Using the exact values in Tables 3–5 and, if artifacts are provided, the per-response data, recompute the between-group comparison: run Welch's ANOVA on group (general vs medical) for each safety, toxicity, bias, accessibility, and affectiveness metric, plus a leave-one-model-out jackknife (especially dropping Gemma from general and Meditron from medical). Report p-values and Hedges' g. If the between-group effect is non-significant or flips sign under jackknife, the duality claim must be withdrawn or reframed as a per-model finding.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim—that medical LLMs show higher harm/toxicity/bias while general-purpose LLMs are safer—is contradicted by the paper's own summary tables. Averaging Table 4's Toxicity column gives general-purpose mean ≈0.0296 (Alpaca 0.019, Vicuna 0.025, Llama3 0.033, Mistral 0.033, Gemma 0.038) vs medical mean ≈0.0297 (MedAlpaca 0.024, BioMistral 0.028, Meditron 0.037). Gender bias means are ≈1.213 for general models vs ≈0.969 for medical models, the opposite direction of the stated claim. The only medical model with consistently poor safety scores is Meditron, and the only general model with consistently poor safety scores is Gemma; grouping by specialization turns this overlap into a \"duality.\" No between-group significance test is reported despite §3.4.1 naming Welch's ANOVA and Games-Howell; §4 gives no p-values, effect sizes, or group-level contrasts. For affectiveness, Table 5's Reflection Scores are nearly identical across all eight models (-7.70 to -7.91), with MedAlpaca numerically highest, so 'general-purpose LLMs produced outputs of higher affectiveness' is also unsupported. Because the abstract and discussion make a group-level causal claim about medical fine-tuning, the per-model evidence shown is not merely incomplete—it contradicts the headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an evaluation of eight open-source LLMs (five general-purpose: Llama 3, Gemma, Alpaca, Mistral, Vicuna; three medical: MedAlpaca, BioMistral, Meditron) for generating patient-facing breast and cervical cancer communication. It uses a mixed-methods framework covering linguistic quality (BLEURT, BERTScore, ROUGE, a constructed hallucination score, and expert ratings), safety and trustworthiness (Perspective API toxicity, GenBit gender bias, racial-context similarity, and expert harm/trust ratings), and communication accessibility and affectiveness (six readability indices, PAIR reflection scores, and expert empathy/clarity ratings). The central claim, stated in the abstract and discussion, is that general-purpose models yield higher linguistic quality and affectiveness, medical models yield greater accessibility, and medical models exhibit higher harm, toxicity, and bias, suggesting a duality between domain-specific knowledge and safety. The paper also describes Welch's ANOVA, Games-Howell post hoc tests, and Hedges' g for statistical comparisons, though no test statistics are reported.","tokens_in":19881,"tokens_out":7919,"duration_ms":64213,"significance":"If the central claims were robustly supported, the paper would provide a useful comparative benchmark for open-source models in patient-facing cancer communication, with practical implications for fine-tuning strategies and safety evaluation. The evaluation design is mostly external: it relies on established tools (Perspective API, GenBit, sentence-BERT, readability indices) rather than self-referential scoring, and the qualitative ratings include inter-rater reliability. However, the significance is currently limited by internal contradictions between the text and tables, the absence of reported statistical tests, and an underspecified hallucination metric. The paper's headline \"duality\" claim is not established by the presented quantitative data.","major_comments":[{"comment":"The sentence \"Meditron achieved the highest ranking and aggregate scores in BLEURT, BERTScore Recall, and all ROUGE variants, indicating overall superiority in linguistic quality\" directly contradicts both Table 3 and Table 1. In Table 3, Llama 3 has higher BLEURT (0.41 vs 0.32), BERTScore Recall (0.86 vs 0.85), ROUGE-1 (0.51 vs 0.40), ROUGE-2 (0.32 vs 0.23), and ROUGE-L (0.42 vs 0.33) than Meditron. Table 1 ranks Llama 3 first on Bleurt, BERTScore Recall, BERTScore F1, ROUGE-1, and ROUGE-2, while Meditron ranks fourth on Bleurt, third on Recall, and fifth or sixth on the ROUGE variants. This is not a minor wording issue; it inverts the reported evidence and must be corrected, along with the subsequent qualitative discussion that credits medical models with poor linguistic quality (which is consistent) but then attributes the highest linguistic scores to Meditron.","section":"Section 4.1, Table 1, Table 3"},{"comment":"The central claim that \"medical LLMs tend to exhibit higher levels of potential harm, toxicity, and bias\" is not supported by the quantitative data in Table 4. Averaging the Perspective API toxicity scores gives general-purpose models a mean of approximately 0.0296 (Alpaca 0.019, Vicuna 0.025, Llama 3 0.033, Mistral 0.033, Gemma 0.038) and medical models a mean of approximately 0.0297 (MedAlpaca 0.024, BioMistral 0.028, Meditron 0.037) — essentially identical. For gender bias, the general-purpose mean is approximately 1.213 while the medical mean is approximately 0.969, the opposite direction of the claimed effect. The sentence in Section 4.2 that \"BioMistral and Meditron exhibited higher toxicity and bias scores than general LLMs\" is also contradicted by Table 4 for both toxicity and gender bias. The only quantitative support in the claimed direction comes from the racial-context similarity scores in Table 7/Figure 3, which are not validated as a bias measure, and from the qualitative harm ratings in Table 2. The authors should either restrict the safety claim to the specific models and metrics that actually show it, or provide group-level statistical tests (with effect sizes) that justify the generalization.","section":"Abstract, Section 4.4, Table 4"},{"comment":"The methods section states that Welch's ANOVA, Games-Howell post hoc tests, and Hedges' g are used, with significance set at p < 0.05, yet no p-values, confidence intervals, effect sizes, or significance asterisks appear anywhere in Section 4 or in any table. Statements such as \"post hoc analysis identified general LLMs, specifically Llama 3, outperforming medical LLMs\" (Section 4.1) and the toxicity comparisons in Section 4.2 are therefore unverifiable. The authors must report the actual test statistics for at least the headline group comparisons (e.g., linguistic quality, toxicity, gender bias, readability, reflection score), or explicitly state which differences failed to reach significance. Without this, the comparative claims rest on unsupported point estimates. Additionally, the ranking-adjustment procedure described in Section 3.4.1 — \"if the effect size was positive, the first model's rank increased and the second's decreased\" — is nonstandard and could amplify noise into the rankings in Table 1; please clarify how this transformation preserves the original metric values.","section":"Section 3.4.1 vs Section 4"},{"comment":"The hallucination score is a load-bearing metric for the claim that \"medical LLMs hallucinated more frequently than general LLMs\" (Section 4.4), but its construction is not reproducible. Section 3.1.1 says it is based on \"the entropy of named entities and nouns\" and is \"adjusted for varying confidence levels and is normalized,\" but no formula, entity set, entropy estimator, or normalization procedure is given. Moreover, Table 3 reports hallucination scores whose group means are nearly identical: general-purpose models average (0.41 + 0.43 + 0.57 + 0.57 + 0.51)/5 = 0.498, and medical models average (0.54 + 0.44 + 0.52)/3 = 0.500. Thus, even the direction of the claim is contradicted by the reported numbers. Specify the hallucination score precisely, report its distribution, and either present a valid group contrast or amend the claim to reflect the per-model scores.","section":"Section 3.1.1, Table 3, Section 4.4"},{"comment":"The paper repeatedly frames the results as evidence of \"a duality between domain-specific knowledge and safety,\" but no measure of domain-specific knowledge is presented. The evaluation dimensions are linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness; there is no metric of medical accuracy, answer correctness, or domain-knowledge recall. The curated dataset of 4,643 QA instances (Section 3.2) is used only to sample 50 questions per model, and no answer-accuracy evaluation is reported. Without an explicit knowledge metric, the conclusion that medical fine-tuning trades away safety while preserving domain knowledge is not operationalized. Please either add a domain-knowledge metric (e.g., answer accuracy on held-out questions from the dataset) or reframe the discussion to describe the observed trade-off in terms of the constructs actually measured (e.g., readability vs. qualitative harm).","section":"Section 4.4, Section 3.2"}],"minor_comments":[{"comment":"The term \"affectiveness\" is used throughout the abstract, introduction, tables, and discussion; \"affectivity\" or \"affective quality\" would be more standard in the psychology and health-communication literature.","section":"Throughout"},{"comment":"There is a double period in the sentence \"...without assuming homogeneity of variance or equal sample sizes, for this multi-model and multi-metric comparison. .\" The extra period should be removed.","section":"Section 3.4.1"},{"comment":"Table 1 lists \"Racial Bias - - - - - - - -\" with no explanation for why no quantitative racial-bias score is reported in that row, even though Figure 3 and Table 7 report racial-context similarity scores. A table note should clarify that racial bias is measured via similarity, not as a direct score, or the row should be removed.","section":"Table 1"},{"comment":"The example responses for Alpaca and BioMistral contain the literal text \"[Hallucinated Content]\" within the model outputs. If these are annotation placeholders, they should be removed or clearly marked as editorial additions; as printed, they appear to be part of the model outputs and undermine the credibility of the appendix.","section":"Appendix (sample responses)"},{"comment":"The parenthetical scores in Table 1 (e.g., Llama 3 Bleurt Score \"7\", Alpaca Bleurt Score \"-2\") are not defined. The caption says \"rank (score)\" but the score appears to be a rank-based transformation rather than the raw metric value. The transformation should be described in the caption.","section":"Table 1"},{"comment":"The text references both \"the figure 3\" and \"Table 7\" for the same racial-context similarity results; the cross-referencing should be unified and the figure should be explicitly described in the text.","section":"Section 4.2, Figure 3, Table 7"}],"recommendation":"major_revision","confidential_remarks":"The internal contradictions between the text and the tables are substantial enough that the authors should be required to re-check the reported values before resubmission. In particular, the Section 4.1 statement about Meditron and the Section 4.2 statement about BioMistral/Meditron appear to be direct inversions of the tabulated data. The absence of any reported statistical tests, despite a full statistical machinery section, makes the headline claims unverifiable. The \"[Hallucinated Content]\" placeholder in the appendix also suggests possible data-handling errors that the authors should inspect."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: the paper's headline—that medical LLMs trade away safety and empathy—doesn't survive its own tables. The toxicity means for the two groups are essentially identical (0.0296 vs 0.0297), and medical models have lower mean gender bias (0.97 vs 1.21). The real pattern is per-model: Meditron is the most toxic, Gemma the most biased general model, but that's not a group-level duality. I'd be cautious about citing the abstract as a result.\n\nWhat's genuinely useful: the curated dataset (4,643 breast/cervical QA pairs) and the eight-model comparison are new. The framework—linguistic quality, safety, accessibility/affectiveness—with both automatic metrics and masked expert ratings is sensible, and reporting inter-rater kappas is a step up from most LLM evals. The authors also correctly point back to Kursuncu et al. (2025) showing that domain fine-tuning can increase toxicity, so the conceptual seed is already in the literature; the contribution is the cancer communication application.\n\nSoft spots, in order of severity:\n\n1. The central claim is not supported by Table 4. I did the arithmetic; the group averages are essentially equal on toxicity, and medical models are less gender-biased. The abstract's \"duality\" is an artifact of treating two outlier models (Meditron, Gemma) as group trends.\n\n2. Welch's ANOVA and Games-Howell are promised but no p-values, effect sizes, or intervals appear anywhere. The \"significant differences\" in Section 4 are asserted, not shown.\n\n3. Section 4.1 contains a visible contradiction: it credits Meditron with the highest BLEURT/ROUGE scores, while Table 3 shows Llama3 with 0.41 BLEURT and Meditron at 0.32. That looks like a copy-paste error, but it erodes trust.\n\n4. The hallucination score is an uncalibrated entropy-based construction; the PAIR reflection scores range only from -7.70 to -7.91, so the affectiveness ranking rests on noise. Neither is validated as a communication-quality measure.\n\n5. No code, data, or generation settings are released. Replication is impossible as is.\n\nWho it's for: people building or evaluating open-source LLMs for patient-facing health content. The per-model table is a useful starting point, but the analysis needs major revision before the conclusions are actionable. I'd like to see a version that reports the full statistical output, corrects the internal contradictions, and limits claims to the models actually tested.\n\nMy recommendation: send it to peer review—the dataset and framework deserve scrutiny—but expect heavy revision. If the authors fix the stats and stop overinterpreting, this could become a citable benchmark.\n\nBest,\n[You]","headline":"The paper's headline 'duality' is contradicted by its own Table 4; the dataset and framework are worth a look, but the central claim needs major reanalysis.","tokens_in":20430,"tokens_out":3414,"would_cite":false,"duration_ms":30199,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Medical fine-tuning buys readability, not safety, in cancer chatbots.","keywords":["large language models","cancer communication","health literacy","patient-facing AI","toxicity and bias","breast cancer","cervical cancer","mixed-methods evaluation"],"falsifier":"Take the same 400 model responses, mask model identities, and have oncologists and health-literacy experts rate factual safety, emotional support, and comprehension using a clinical rubric that does not depend on the paper's surrogate scores; if a medical model such as MedAlpaca is rated as safe and as empathetic as Llama 3, the claimed duality is not robust.","tokens_in":19414,"feed_emoji":"🩺","tokens_out":10610,"duration_ms":98415,"temperature":0.7,"pith_summary":"This paper asks whether open-source large language models can communicate about breast and cervical cancer safely and clearly enough for patient-facing use. It compares five general-purpose models (Llama 3, Gemma, Alpaca, Mistral, Vicuna) with three medical-domain models (MedAlpaca, BioMistral, Meditron) on 4,643 curated questions, scoring outputs by automated metrics and blinded expert ratings. The central claim is a duality: general-purpose models produced higher linguistic quality, safety, and emotional engagement, while medical models produced more readable text but tended toward more harm, toxicity, bias, and coherence failures. If the claim holds, current medical fine-tuning buys accessibility at the cost of trustworthiness, which matters because patient-facing AI is already being piloted in oncology education and triage. The paper's stated scope is open-source models at or below 8B parameters and existing benchmark datasets, so generalization to larger or proprietary systems is left open.","feed_headline":"Medical fine-tuning buys readability, not safety, in cancer chatbots","feed_subtitle":"General-purpose models were rated more trustworthy and empathetic; medical models were simpler but riskier.","key_machinery":"The load-bearing mechanism is a three-axis mixed-methods evaluation. Linguistic quality is measured by reference-based text metrics (ROUGE, BLEURT, BERTScore), a normalized hallucination score built from named-entity and noun entropy, and expert ratings of accuracy, coherence, jargon, and reasoning. Safety and trustworthiness are measured by an automated toxicity scorer, a gender-bias score, an in-context impersonation test for racial bias, and expert ratings of harm and trust. Accessibility and affectiveness are measured by six readability indices and the PAIR reflection score for empathetic response, plus expert ratings of clarity, empathy, compassion, cue to action, domain relevance, and usability. The framework's work is to let every model be ranked on each axis; an analysis of variance identifies where models differ, post hoc pairwise tests compare models, and an effect-size measure quantifies the differences. This construction is what turns the raw outputs into the claimed duality.","core_discovery":"On the paper's own terms, the discovery is that quality dimensions do not align across model classes. The general-purpose models, especially Llama 3 and Gemma, ranked first on BLEURT, BERTScore, ROUGE, and the hallucination score, and on expert ratings of accuracy, coherence, reasoning, harm avoidance, trust, empathy, compassion, cue to action, and usability. The medical models, especially MedAlpaca and BioMistral, ranked best on readability indices such as Flesch Reading Ease and grade level, meaning their outputs were closer to public-health readability targets. Yet the paper reports that medical fine-tuning was not accompanied by safety: Meditron and BioMistral showed higher toxicity and bias values in the automated metrics, the medical models scored lower on expert-rated harm and trust, and several medical outputs contained obvious hallucinated content. The paper reads this pattern as evidence of a duality between domain-specific knowledge and safety in health communications, and recommends that future medical LLM development add harm mitigation, bias reduction, and affectiveness objectives rather than optimize domain knowledge alone.","pith_inferences":["A testable extension would run the same 400 outputs past patients with limited health literacy; the readability indices are proxies, and the paper does not measure whether the simpler medical-model text is actually better understood.","I would expect the safety gap to shrink with larger, heavily safety-aligned models, because all eight models studied are 7-8B open-source models with comparatively little alignment; this remains speculative.","A practical design suggested by the pattern is a two-stage pipeline: draft with a general-purpose model, then rewrite for readability and filter toxicity before presenting to patients.","The racial-bias measure compares response consistency across prompts, so the reported fairness gap is about consistency of tone and content, not about whether the advice would change treatment outcomes."],"forward_implications":["Healthcare organizations piloting open-source chatbots for cancer education should not treat medical fine-tuning as a safety certification; on this evidence, the medical models were more readable but rated less safe and less trustworthy.","Fine-tuning pipelines for medical LLMs should include explicit objectives for harm reduction, bias mitigation, and supportive tone, not just accuracy on medical benchmarks.","Readability and communicative quality should be measured as separate targets, because the models with the easiest reading levels had the weakest expert-rated accuracy, coherence, and empathy.","Evaluations of domain-specific LLMs should report toxicity and demographic bias alongside medical knowledge metrics, since the models that carried the most domain training also carried the most risk.","The safety gap between model classes provides a concrete baseline for future work: any new medical fine-tune should be required to match general-purpose models on harm, trust, and empathy before deployment."],"supporting_citations":[{"why":"Supplies PubMedQA cases and part of the reference answers used to score linguistic quality and hallucination.","marker":"Singhal et al. (2023)"},{"why":"Supplies MedQA-USMLE cases filtered for breast and cervical cancer questions.","marker":"Jin et al. (2021)"},{"why":"Supplies MedMCQA cases used in the curated question set.","marker":"Pal et al. (2022)"},{"why":"Introduces MedAlpaca, one of the three medical models whose outputs are evaluated.","marker":"Han et al. (2023)"},{"why":"Introduces BioMistral, the medical model that scored best on several readability indices.","marker":"Labrak et al. (2024)"},{"why":"Introduces Meditron, the medical model that showed the highest toxicity and identity-attack scores.","marker":"Z. Chen et al. (2023)"},{"why":"Introduces Llama 3, the general-purpose model that ranked first on most linguistic and safety criteria.","marker":"Dubey et al. (2024)"},{"why":"Introduces Gemma, the general-purpose model that ranked second on linguistic quality and showed low racial bias.","marker":"Team et al. (2024)"},{"why":"Defines the PAIR reflection score used to measure how well responses affirm patient affect.","marker":"Min, Resnicow, Resnicow, & Mihalcea (2022)"},{"why":"Provides the Perspective API that generates the toxicity scores used in the safety dimension.","marker":"Jigsaw (2024)"}],"fun_headline_variants":["Medical LLMs: easier to read, harder to trust","Cancer chatbots: general models safer, medical simpler","Readability vs safety: the cancer chatbot trade-off","For cancer info, general LLMs beat medical on trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the curated benchmark reference answers and the paper's hand-built hallucination and reflection scores are valid proxies for what patients actually need from cancer communication; if they are not, the ranking and the claimed safety-versus-readability duality are artifacts of the metrics.","fun_headline_variants_meta":{"raw":{"variants":["Medical LLMs: easier to read, harder to trust","Cancer chatbots: general models safer, medical simpler","Readability vs safety: the cancer chatbot trade-off","For cancer info, general LLMs beat medical on trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1477,"prompt_tokens":1003,"completion_tokens":474,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":410}},"tokens_in":619,"tokens_out":474,"duration_ms":5050,"temperature":1.0,"reasoning_tokens":410,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:08:34.562273+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 400 model responses, mask model identities, and have oncologists and health-literacy experts rate factual safety, emotional support, and comprehension using a clinical rubric that does not depend on the paper's surrogate scores; if a medical model such as MedAlpaca is rated as safe and as empathetic as Llama 3, the claimed duality is not robust.","supporting_citations":[{"cited_title":", Azizi, S","cited_arxiv_id":null,"evidence_quote":"Supplies PubMedQA cases and part of the reference answers used to score linguistic quality and hallucination."},{"cited_title":", Pan, E","cited_arxiv_id":null,"evidence_quote":"Supplies MedQA-USMLE cases filtered for breast and cervical cancer questions."},{"cited_title":", Umapathi, L.K","cited_arxiv_id":null,"evidence_quote":"Supplies MedMCQA cases used in the curated question set."},{"cited_title":", Resnicow, V.P","cited_arxiv_id":null,"evidence_quote":"Defines the PAIR reflection score used to measure how well responses affirm patient affect."},{"cited_title":"Perspective API","cited_arxiv_id":null,"evidence_quote":"Provides the Perspective API that generates the toxicity scores used in the safety dimension."}],"review_version":1}