REVIEW 5 major objections 6 minor 1 cited by
Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
T0 review · 5 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Medical fine-tuning buys readability, not safety, in cancer chatbots.
desk verdict The paper's headline 'duality' is contradicted by its own Table 4; the dataset and framework are worth a look, but the central claim needs major reanalysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a three-axis mixed-methods evaluation. Linguistic quality is measured by reference-based text metrics (ROUGE, BLEURT, BERTScore), a normalized hallucination score built from named-entity and noun entropy, and expert ratings of accuracy, coherence, jargon, and reasoning. Safety and trustworthiness are measured by an automated toxicity scorer, a gender-bias score, an in-context impersonation test for racial bias, and expert ratings of harm and trust. Accessibility and affectiveness are measured by six readability indices and the PAIR reflection score for empathetic response, plus expert ratings of clarity, empathy, compassion, cue to action, domain relevance, and usability. The framework's work is to let every model be ranked on each axis; an analysis of variance identifies where models differ, post hoc pairwise tests compare models, and an effect-size measure quantifies the differences. This construction is what turns the raw outputs into the claimed duality.
What would settle it
Take the same 400 model responses, mask model identities, and have oncologists and health-literacy experts rate factual safety, emotional support, and comprehension using a clinical rubric that does not depend on the paper's surrogate scores; if a medical model such as MedAlpaca is rated as safe and as empathetic as Llama 3, the claimed duality is not robust.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that quality dimensions do not align across model classes. The general-purpose models, especially Llama 3 and Gemma, ranked first on BLEURT, BERTScore, ROUGE, and the hallucination score, and on expert ratings of accuracy, coherence, reasoning, harm avoidance, trust, empathy, compassion, cue to action, and usability. The medical models, especially MedAlpaca and BioMistral, ranked best on readability indices such as Flesch Reading Ease and grade level, meaning their outputs were closer to public-health readability targets. Yet the paper reports that medical fine-tuning was not accompanied by safety: Meditron and BioMistral showed higher toxicity and bias values in the automated metrics, the medical models scored lower on expert-rated harm and trust, and several medical outputs contained obvious hallucinated content. The paper reads this pattern as evidence of a duality between domain-specific knowledge and safety in health communications, and recommends that future medical LLM development add harm mitigation, bias reduction, and affectiveness objectives rather than optimize domain knowledge alone.
Load-bearing premise
The load-bearing premise is that the curated benchmark reference answers and the paper's hand-built hallucination and reflection scores are valid proxies for what patients actually need from cancer communication; if they are not, the ranking and the claimed safety-versus-readability duality are artifacts of the metrics.
Editorial extensions
If this is right
- Healthcare organizations piloting open-source chatbots for cancer education should not treat medical fine-tuning as a safety certification; on this evidence, the medical models were more readable but rated less safe and less trustworthy.
- Fine-tuning pipelines for medical LLMs should include explicit objectives for harm reduction, bias mitigation, and supportive tone, not just accuracy on medical benchmarks.
- Readability and communicative quality should be measured as separate targets, because the models with the easiest reading levels had the weakest expert-rated accuracy, coherence, and empathy.
- Evaluations of domain-specific LLMs should report toxicity and demographic bias alongside medical knowledge metrics, since the models that carried the most domain training also carried the most risk.
- The safety gap between model classes provides a concrete baseline for future work: any new medical fine-tune should be required to match general-purpose models on harm, trust, and empathy before deployment.
Reading between the lines
- A testable extension would run the same 400 outputs past patients with limited health literacy; the readability indices are proxies, and the paper does not measure whether the simpler medical-model text is actually better understood.
- I would expect the safety gap to shrink with larger, heavily safety-aligned models, because all eight models studied are 7-8B open-source models with comparatively little alignment; this remains speculative.
- A practical design suggested by the pattern is a two-stage pipeline: draft with a general-purpose model, then rewrite for readability and filter toxicity before presenting to patients.
- The racial-bias measure compares response consistency across prompts, so the reported fairness gap is about consistency of tone and content, not about whether the advice would change treatment outcomes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an evaluation of eight open-source LLMs (five general-purpose: Llama 3, Gemma, Alpaca, Mistral, Vicuna; three medical: MedAlpaca, BioMistral, Meditron) for generating patient-facing breast and cervical cancer communication. It uses a mixed-methods framework covering linguistic quality (BLEURT, BERTScore, ROUGE, a constructed hallucination score, and expert ratings), safety and trustworthiness (Perspective API toxicity, GenBit gender bias, racial-context similarity, and expert harm/trust ratings), and communication accessibility and affectiveness (six readability indices, PAIR reflection scores, and expert empathy/clarity ratings). The central claim, stated in the abstract and discussion, is that general-purpose models yield higher linguistic quality and affectiveness, medical models yield greater accessibility, and medical models exhibit higher harm, toxicity, and bias, suggesting a duality between domain-specific knowledge and safety. The paper also describes Welch's ANOVA, Games-Howell post hoc tests, and Hedges' g for statistical comparisons, though no test statistics are reported.
Significance. If the central claims were robustly supported, the paper would provide a useful comparative benchmark for open-source models in patient-facing cancer communication, with practical implications for fine-tuning strategies and safety evaluation. The evaluation design is mostly external: it relies on established tools (Perspective API, GenBit, sentence-BERT, readability indices) rather than self-referential scoring, and the qualitative ratings include inter-rater reliability. However, the significance is currently limited by internal contradictions between the text and tables, the absence of reported statistical tests, and an underspecified hallucination metric. The paper's headline "duality" claim is not established by the presented quantitative data.
major comments (5)
- [Section 4.1, Table 1, Table 3] The sentence "Meditron achieved the highest ranking and aggregate scores in BLEURT, BERTScore Recall, and all ROUGE variants, indicating overall superiority in linguistic quality" directly contradicts both Table 3 and Table 1. In Table 3, Llama 3 has higher BLEURT (0.41 vs 0.32), BERTScore Recall (0.86 vs 0.85), ROUGE-1 (0.51 vs 0.40), ROUGE-2 (0.32 vs 0.23), and ROUGE-L (0.42 vs 0.33) than Meditron. Table 1 ranks Llama 3 first on Bleurt, BERTScore Recall, BERTScore F1, ROUGE-1, and ROUGE-2, while Meditron ranks fourth on Bleurt, third on Recall, and fifth or sixth on the ROUGE variants. This is not a minor wording issue; it inverts the reported evidence and must be corrected, along with the subsequent qualitative discussion that credits medical models with poor linguistic quality (which is consistent) but then attributes the highest linguistic scores to Meditron.
- [Abstract, Section 4.4, Table 4] The central claim that "medical LLMs tend to exhibit higher levels of potential harm, toxicity, and bias" is not supported by the quantitative data in Table 4. Averaging the Perspective API toxicity scores gives general-purpose models a mean of approximately 0.0296 (Alpaca 0.019, Vicuna 0.025, Llama 3 0.033, Mistral 0.033, Gemma 0.038) and medical models a mean of approximately 0.0297 (MedAlpaca 0.024, BioMistral 0.028, Meditron 0.037) — essentially identical. For gender bias, the general-purpose mean is approximately 1.213 while the medical mean is approximately 0.969, the opposite direction of the claimed effect. The sentence in Section 4.2 that "BioMistral and Meditron exhibited higher toxicity and bias scores than general LLMs" is also contradicted by Table 4 for both toxicity and gender bias. The only quantitative support in the claimed direction comes from the racial-context similarity scores in Table 7/Figure 3, which are not validated as a bias measure, and from the qualitative harm ratings in Table 2. The authors should either restrict the safety claim to the specific models and metrics that actually show it, or provide group-level statistical tests (with effect sizes) that justify the generalization.
- [Section 3.4.1 vs Section 4] The methods section states that Welch's ANOVA, Games-Howell post hoc tests, and Hedges' g are used, with significance set at p < 0.05, yet no p-values, confidence intervals, effect sizes, or significance asterisks appear anywhere in Section 4 or in any table. Statements such as "post hoc analysis identified general LLMs, specifically Llama 3, outperforming medical LLMs" (Section 4.1) and the toxicity comparisons in Section 4.2 are therefore unverifiable. The authors must report the actual test statistics for at least the headline group comparisons (e.g., linguistic quality, toxicity, gender bias, readability, reflection score), or explicitly state which differences failed to reach significance. Without this, the comparative claims rest on unsupported point estimates. Additionally, the ranking-adjustment procedure described in Section 3.4.1 — "if the effect size was positive, the first model's rank increased and the second's decreased" — is nonstandard and could amplify noise into the rankings in Table 1; please clarify how this transformation preserves the original metric values.
- [Section 3.1.1, Table 3, Section 4.4] The hallucination score is a load-bearing metric for the claim that "medical LLMs hallucinated more frequently than general LLMs" (Section 4.4), but its construction is not reproducible. Section 3.1.1 says it is based on "the entropy of named entities and nouns" and is "adjusted for varying confidence levels and is normalized," but no formula, entity set, entropy estimator, or normalization procedure is given. Moreover, Table 3 reports hallucination scores whose group means are nearly identical: general-purpose models average (0.41 + 0.43 + 0.57 + 0.57 + 0.51)/5 = 0.498, and medical models average (0.54 + 0.44 + 0.52)/3 = 0.500. Thus, even the direction of the claim is contradicted by the reported numbers. Specify the hallucination score precisely, report its distribution, and either present a valid group contrast or amend the claim to reflect the per-model scores.
- [Section 4.4, Section 3.2] The paper repeatedly frames the results as evidence of "a duality between domain-specific knowledge and safety," but no measure of domain-specific knowledge is presented. The evaluation dimensions are linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness; there is no metric of medical accuracy, answer correctness, or domain-knowledge recall. The curated dataset of 4,643 QA instances (Section 3.2) is used only to sample 50 questions per model, and no answer-accuracy evaluation is reported. Without an explicit knowledge metric, the conclusion that medical fine-tuning trades away safety while preserving domain knowledge is not operationalized. Please either add a domain-knowledge metric (e.g., answer accuracy on held-out questions from the dataset) or reframe the discussion to describe the observed trade-off in terms of the constructs actually measured (e.g., readability vs. qualitative harm).
minor comments (6)
- [Throughout] The term "affectiveness" is used throughout the abstract, introduction, tables, and discussion; "affectivity" or "affective quality" would be more standard in the psychology and health-communication literature.
- [Section 3.4.1] There is a double period in the sentence "...without assuming homogeneity of variance or equal sample sizes, for this multi-model and multi-metric comparison. ." The extra period should be removed.
- [Table 1] Table 1 lists "Racial Bias - - - - - - - -" with no explanation for why no quantitative racial-bias score is reported in that row, even though Figure 3 and Table 7 report racial-context similarity scores. A table note should clarify that racial bias is measured via similarity, not as a direct score, or the row should be removed.
- [Appendix (sample responses)] The example responses for Alpaca and BioMistral contain the literal text "[Hallucinated Content]" within the model outputs. If these are annotation placeholders, they should be removed or clearly marked as editorial additions; as printed, they appear to be part of the model outputs and undermine the credibility of the appendix.
- [Table 1] The parenthetical scores in Table 1 (e.g., Llama 3 Bleurt Score "7", Alpaca Bleurt Score "-2") are not defined. The caption says "rank (score)" but the score appears to be a rank-based transformation rather than the raw metric value. The transformation should be described in the caption.
- [Section 4.2, Figure 3, Table 7] The text references both "the figure 3" and "Table 7" for the same racial-context similarity results; the cross-referencing should be unified and the figure should be explicitly described in the text.
Circularity Check
No significant circularity: the paper is an external benchmark evaluation whose metrics come from independent tools and references, with no fitted parameter or self-referential derivation.
full rationale
The paper is an empirical evaluation rather than a derivation chain, and none of its reported quantities are defined in terms of the outcome being claimed. Linguistic quality scores come from external reference-based metrics (ROUGE, BLEURT, BERTScore) and expert ratings; safety scores come from the Perspective API, GenBit, and sentence-BERT similarity comparisons; accessibility and affectiveness come from standard readability indices and the external PAIR reflection scorer. No parameter is fitted to a subset of the data and then used to predict the same or a closely related quantity. The hallucination score is an internally defined entropy proxy, but it is not calibrated against, nor derived from, the model rankings or group-level conclusions, so it does not make the central claims true by construction. The self-citations that appear (e.g., Erol et al. 2025 for Perspective API attribution and Kursuncu et al. 2025 as related-work motivation for fine-tuning and toxicity) are not load-bearing justifications for the measured results; the actual measurements rely on external tools and benchmarks. Concerns that some group-level claims are unsupported by the paper's own summary tables are evidentiary or statistical issues, not circularity. The evaluation is therefore self-contained against external standards, and no circular step can be identified.
Assumptions & free parameters
free parameters (2)
- Hallucination score construction
- Ranking score transformation
assumptions (4)
- domain assumption Reference answers in the curated medical QA datasets are valid ground truth for assessing patient-facing cancer communication.
- domain assumption The two expert raters' averaged 3-point Likert scores can be treated as interval-level data.
- domain assumption External tools (Perspective API, GenBit, sentenceBERT, PAIR) validly capture toxicity, bias, and affectiveness.
- standard math Welch's ANOVA and Games-Howell post hoc tests are appropriate for the data distributions.
Cite this review
Pith. "Pith review of Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI." pith.science (2026). https://pith.science/paper/ABMIUOGK
@misc{pith2026250510472,
author = {Pith},
title = {Pith review of: Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/ABMIUOGK}},
note = {Machine review of arXiv:2505.10472}
}
read the original abstract
Effective communication about breast and cervical cancers remains a persistent health challenge, with significant gaps in public understanding of cancer prevention, screening, and treatment, potentially leading to delayed diagnoses and inadequate treatments. This study evaluates the capabilities and limitations of Large Language Models (LLMs) in generating accurate, safe, and accessible cancer-related information to support patient understanding. We evaluated five general-purpose and three medical LLMs using a mixed-methods evaluation framework across linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness. Our approach utilized quantitative metrics, qualitative expert ratings, and statistical analysis using Welch's ANOVA, Games-Howell, and Hedges' g. Our results show that general-purpose LLMs produced outputs of higher linguistic quality and affectiveness, while medical LLMs demonstrate greater communication accessibility. However, medical LLMs tend to exhibit higher levels of potential harm, toxicity, and bias, reducing their performance in safety and trustworthiness. Our findings indicate a duality between domain-specific knowledge and safety in health communications. The results highlight the need for intentional model design with targeted improvements, particularly in mitigating harm and bias, and improving safety and affectiveness. This study provides a comprehensive evaluation of LLMs for cancer communication, offering critical insights for improving AI-generated health content and informing future development of accurate, safe, and accessible digital health tools.
Forward citations
Cited by 1 Pith paper
-
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
Fine-tuning GPT-3.5 and Llama 2 on r/Anxiety posts improves readability but raises toxicity and bias while reducing empathy and reflection.
Reference graph
Works this paper leans on
-
[1]
, Mrabet, Y
abacha2019bridging APACrefauthors Abacha, A.B. , Mrabet, Y. , Sharp, M. , Goodwin, T.R. , Shooshan, S.E. Demner-Fushman, D. APACrefauthors \ 2019 . Bridging the gap between consumers’ medication questions and trusted answers Bridging the gap between consumers’ medication questions and trusted answers . MEDINFO 2019: Health and Wellbeing e-Networks for All...
2019
-
[2]
, Khatibi, E
abbasian2024foundation APACrefauthors Abbasian, M. , Khatibi, E. , Azimi, I. , Oniani, D. , Shakeri Hossein Abad, Z. , Thieme, A. others APACrefauthors \ 2024 . Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI Foundation metrics for evaluating effectiveness of healthcare conversations powered by generati...
2024
-
[3]
, Carmel, D
agichtein2015overview APACrefauthors Agichtein, E. , Carmel, D. , Pelleg, D. , Pinter, Y. Harman, D. APACrefauthors \ 2015 . Overview of the TREC 2015 LiveQA Track. Overview of the trec 2015 liveqa track. TREC. Trec
2015
-
[4]
AmericanCancerSociety2023 APACrefauthors American Cancer Society APACrefauthors \ 2023 . Cancer Facts & Figures 2023. Cancer facts & figures 2023. APACrefURL https://acsjournals.onlinelibrary.wiley.com/doi/10.3322/caac.21820 APACrefURL
-
[5]
Key Statistics for Cervical Cancer
acs2024cervical APACrefauthors American Cancer Society APACrefauthors \ 2024 . Key Statistics for Cervical Cancer. Key statistics for cervical cancer. APACrefURL https://www.cancer.org/cancer/types/cervical-cancer/about/key-statistics.html APACrefURL Accessed: 2025-05-11
2024
-
[6]
, Jatoi, I
anderson2010male APACrefauthors Anderson, W.F. , Jatoi, I. , Tse, J. Rosenberg, P.S. APACrefauthors \ 2010 . Male breast cancer: a population-based comparison with female breast cancer Male breast cancer: a population-based comparison with female breast cancer . Journal of Clinical Oncology 28 2 232--239,
2010
-
[7]
Ayers2023ChatGPTvsPhysicians APACrefauthors Ayers, J.W. , Poliak, A. , Dredze, M. , Leas, E.C. , Zhu, Z. , Kelley, J.B. Smith, D.M. APACrefauthors \ 2023 . Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum Comparing physician and artificial intelligence chatbot responses to patient...
arXiv 2023
-
[8]
Simply Put: A guide for creating easy-to-understand materials Simply put: A guide for creating easy-to-understand materials \ [ ]
cdc2009simplyput APACrefauthors Centers for Disease Control and Prevention APACrefauthors \ 2009 . Simply Put: A guide for creating easy-to-understand materials Simply put: A guide for creating easy-to-understand materials \ [ ]. Available at https://www.cdc.gov/healthliteracy/pdf/simply_put.pdf
2009
Show all 87 references
-
[9]
, Avison, K
Chen2025OncologyLLM APACrefauthors Chen, D. , Avison, K. , Alnassar, S. , Huang, R.S. Raman, S. APACrefauthors \ 2025 . Medical accuracy of artificial intelligence chatbots in oncology: a scoping review Medical accuracy of artificial intelligence chatbots in oncology: a scopin...
2025 doi
-
[10]
, Parsa, R
chen2024physician APACrefauthors Chen, D. , Parsa, R. , Hope, A. , Hannon, B. , Mak, E. , Eng, L. Raman, S. APACrefauthors \ 2024 1 . Physician and artificial intelligence chatbot responses to cancer questions from social media Physician and artificial intelligence chatbot res...
2024
-
[11]
, Parsa, R
Chen2024JAMAOnc APACrefauthors Chen, D. , Parsa, R. , Hope, A. , Hannon, B. , Mak, E. , Eng, L. Raman, S. APACrefauthors \ 2024 2 . Physician and Artificial Intelligence Chatbot Responses to Cancer Questions From Social Media Physician and artificial intelligence chatbot respo...
2024
-
[12]
, Cano, A.H
chen2023meditron APACrefauthors Chen, Z. , Cano, A.H. , Romanou, A. , Bonnet, A. , Matoba, K. , Salvi, F. others APACrefauthors \ 2023 . Meditron-70b: Scaling medical pretraining for large language models Meditron-70b: Scaling medical pretraining for large language models . ar...
2023 arXiv
-
[13]
\ Liau, T.L
Coleman1975 APACrefauthors Coleman, M. \ Liau, T.L. APACrefauthors \ 1975 . A Computer Readability Formula Designed for Machine Scoring A computer readability formula designed for machine scoring . Journal of Applied Psychology 60 283--284,
1975
-
[14]
, Magai, C
consedine2004breast APACrefauthors Consedine, N.S. , Magai, C. , Spiller, R. , Neugut, A.I. Conway, F. APACrefauthors \ 2004 . Breast cancer knowledge and beliefs in subpopulations of African American and Caribbean women Breast cancer knowledge and beliefs in subpopulations of...
2004
-
[15]
, Wang, T
deng2024evaluation APACrefauthors Deng, L. , Wang, T. , Zhai, Z. , Tao, W. , Li, J. , Zhao, Y. others APACrefauthors \ 2024 . Evaluation of large language models in breast cancer clinical scenarios: a comparative analysis based on ChatGPT-3.5, ChatGPT-4.0, and Claude2 Evaluati...
2024
-
[16]
, Jauhri, A
dubey2024llama APACrefauthors Dubey, A. , Jauhri, A. , Pandey, A. , Kadian, A. , Al-Dahle, A. , Letman, A. others APACrefauthors \ 2024 . The llama 3 herd of models The llama 3 herd of models . arXiv preprint arXiv:2407.21783 ,
2024 arXiv
-
[17]
, Padhi, T
erol2025playing APACrefauthors Erol, A. , Padhi, T. , Saha, A. , Kursuncu, U. Aktas, M.E. APACrefauthors \ 2025 . Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models Playing devil's advocate: Unmasking toxicity and vulnerabilities i...
2025 arXiv
-
[18]
, Sano, M
federman2009health APACrefauthors Federman, A.D. , Sano, M. , Wolf, M.S. , Siu, A.L. Halm, E.A. APACrefauthors \ 2009 . Health literacy and cognitive performance in older adults Health literacy and cognitive performance in older adults . Journal of the American Geriatrics Soci...
2009
-
[19]
APACrefauthors \ 1948
Flesch1948 APACrefauthors Flesch, R. APACrefauthors \ 1948 . A New Readability Yardstick A new readability yardstick \ ( 32). Journal of Applied Psychology
1948
-
[20]
\ Howell, J.F
games1976pairwise APACrefauthors Games, P.A. \ Howell, J.F. APACrefauthors \ 1976 . Pairwise multiple comparison procedures with unequal n’s and/or variances: a Monte Carlo study Pairwise multiple comparison procedures with unequal n’s and/or variances: a monte carlo study . J...
1976
-
[21]
, Padhi, T
garg2024just APACrefauthors Garg, R. , Padhi, T. , Jain, H. , Kursuncu, U. Kumaraguru, P. APACrefauthors \ 2024 . Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes Just kiddin: Knowledge infusion and distillation for detection of indecent memes ....
2024
-
[22]
, Gallo, R
Goh2024LLMdiagnosticRCT APACrefauthors Goh, E. , Gallo, R. , Hom, J. , Strong, E. , Weng, Y. , Kerman, H. Chen, J.H. APACrefauthors \ 2024 . Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial Large language model influence on diagnostic reasoni...
2024
-
[23]
, Marques, C
grilo2025assessing APACrefauthors Grilo, A. , Marques, C. , Corte-Real, M. , Carolino, E. , Caetano, M. \ . APACrefauthors \ 2025 . Assessing the Quality and Reliability of ChatGPT’s Responses to Radiotherapy-Related Patient Queries: Comparative Study With GPT-3.5 and GPT-4 As...
2025
-
[24]
APACrefauthors \ 1952
Gunning1952 APACrefauthors Gunning, R. APACrefauthors \ 1952 . The Technique of Clear Writing The technique of clear writing . McGraw-Hill
1952
-
[25]
, Adams, L.C
han2023medalpaca APACrefauthors Han, T. , Adams, L.C. , Papaioannou, J M. , Grundmann, P. , Oberhauser, T. , L \"o ser, A. Bressem, K.K. APACrefauthors \ 2023 . MedAlpaca--an open-source collection of medical conversational AI models and training data Medalpaca--an open-source...
2023 arXiv
-
[26]
APACrefauthors \ 1981
hedges1981distribution APACrefauthors Hedges, L.V. APACrefauthors \ 1981 . Distribution theory for Glass's estimator of effect size and related estimators Distribution theory for glass's estimator of effect size and related estimators . journal of Educational Statistics 6 2 107--128,
1981
-
[27]
huang2025long APACrefauthors Huang, A. , Li, D. , Fan, Z. , Chen, J. , Zhang, W. Wu, W. APACrefauthors \ 2025 . Long-term trends in the incidence of male breast cancer and nomogram for predicting survival in male breast cancer patients: a population-based epidemiologic study L...
2025
-
[28]
, Galal, G
huang2022evaluation APACrefauthors Huang, J. , Galal, G. , Etemadi, M. Vaidyanathan, M. APACrefauthors \ 2022 . Evaluation and mitigation of racial bias in clinical machine learning models: scoping review Evaluation and mitigation of racial bias in clinical machine learning mo...
2022
-
[29]
\ Gundersen, O.E
intahchomphoo2020artificial APACrefauthors Intahchomphoo, C. \ Gundersen, O.E. APACrefauthors \ 2020 . Artificial intelligence and race: A systematic review Artificial intelligence and race: A systematic review . Legal Information Management 20 2 74--84,
2020
-
[30]
, Carcioppolo, N
jensen2014cancer APACrefauthors Jensen, J.D. , Carcioppolo, N. , King, A.J. , Scherr, C.L. , Jones, C.L. Niederdeppe, J. APACrefauthors \ 2014 . The cancer information overload (CIO) scale: Establishing predictive and discriminant validity The cancer information overload (cio)...
2014
-
[31]
, Hwang, H
jeong2024olaph APACrefauthors Jeong, M. , Hwang, H. , Yoon, C. , Lee, T. Kang, J. APACrefauthors \ 2024 . OLAPH: Improving Factuality in Biomedical Long-form Question Answering Olaph: Improving factuality in biomedical long-form question answering . arXiv preprint arXiv:2405.12701 ,
2024 arXiv
-
[32]
, Sablayrolles, A
jiang2023mistral APACrefauthors Jiang, A.Q. , Sablayrolles, A. , Mensch, A. , Bamford, C. , Chaplot, D.S. , Casas, D.d.l. others APACrefauthors \ 2023 . Mistral 7B Mistral 7b . arXiv preprint arXiv:2310.06825 ,
2023 arXiv
-
[33]
Perspective API
PerspectiveAPI APACrefauthors Jigsaw APACrefauthors \ 2024 . Perspective API. Perspective api. https://www.perspectiveapi.com/
2024
-
[34]
, Pan, E
jin2021disease APACrefauthors Jin, D. , Pan, E. , Oufattole, N. , Weng, W H. , Fang, H. Szolovits, P. APACrefauthors \ 2021 . What disease does this patient have? a large-scale open domain question answering dataset from medical exams What disease does this patient have? a lar...
2021
-
[35]
\ MacDermid, J.C
jindal2017assessing APACrefauthors Jindal, P. \ MacDermid, J.C. APACrefauthors \ 2017 . Assessing reading levels of health information: uses and limitations of flesch formula Assessing reading levels of health information: uses and limitations of flesch formula . Education for...
2017
-
[36]
, Pollard, T.J
johnson2016mimic APACrefauthors Johnson, A.E. , Pollard, T.J. , Shen, L. , Lehman, L w.H. , Feng, M. , Ghassemi, M. Mark, R.G. APACrefauthors \ 2016 . MIMIC-III, a freely accessible critical care database Mimic-iii, a freely accessible critical care database . Scientific data ...
2016
-
[37]
, Gaur, M
khandelwal2024domain APACrefauthors Khandelwal, V. , Gaur, M. , Kursuncu, U. , Shalin, V.L. Sheth, A.P. APACrefauthors \ 2024 . A domain-agnostic neurosymbolic approach for big social data analysis: Evaluating mental health sentiment on social media during covid-19 A domain-ag...
2024
-
[38]
, Fishburne, R.P
Kincaid1975 APACrefauthors Kincaid, J.P. , Fishburne, R.P. , Rogers, R.L. Chissom, B.S. APACrefauthors \ 1975 . Derivation of New Readability Formulas for Navy Enlisted Personnel Derivation of new readability formulas for navy enlisted personnel \ \ \ Research Branch Report 8-...
1975
-
[39]
, Padhi, T
kursuncu2025fromreddit APACrefauthors Kursuncu, U. , Padhi, T. , Sinha, G.R. , Erol, A. , Mandivarapu, J.K. Larrison, C.R. APACrefauthors \ 2025 . From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data From reddit to ...
2025
-
[40]
, Bazoge, A
labrak2024biomistral APACrefauthors Labrak, Y. , Bazoge, A. , Morin, E. , Gourraud, P A. , Rouvier, M. Dufour, R. APACrefauthors \ 2024 . Biomistral: A collection of open-source pretrained large language models for medical domains Biomistral: A collection of open-source pretra...
2024 arXiv
-
[41]
\ Koch, G.G
landis1977application APACrefauthors Landis, J.R. \ Koch, G.G. APACrefauthors \ 1977 . An application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers An application of hierarchical kappa-type statistics in the assessment o...
1977
-
[42]
, Gao, Y.F
leong2024efficient APACrefauthors Leong, H.Y. , Gao, Y.F. , Shuai, J. , Zhang, Y. Pamuksuz, U. APACrefauthors \ 2024 . Efficient fine-tuning of large language models for automated medical documentation Efficient fine-tuning of large language models for automated medical docume...
2024
-
[43]
, Karver, T.S
levy2024evaluating APACrefauthors Levy, S. , Karver, T.S. , Adler, W.D. , Kaufman, M.R. Dredze, M. APACrefauthors \ 2024 . Evaluating Biases in Context-Dependent Health Questions Evaluating biases in context-dependent health questions . arXiv preprint arXiv:2403.04858 ,
2024 arXiv
-
[44]
, Gao, Q
li2023kappa APACrefauthors Li, M. , Gao, Q. Yu, T. APACrefauthors \ 2023 . Kappa statistic considerations in evaluating inter-rater reliability between two raters: which, when and context matters Kappa statistic considerations in evaluating inter-rater reliability between two ...
2023
-
[45]
li2023chatdoctor APACrefauthors Li, Y. , Li, Z. , Zhang, K. , Dan, R. , Jiang, S. Zhang, Y. APACrefauthors \ 2023 . Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge Chatdoctor: A medical chat model fine-tuned ...
2023
-
[46]
, Bommasani, R
liang2022holistic APACrefauthors Liang, P. , Bommasani, R. , Lee, T. , Tsipras, D. , Soylu, D. , Yasunaga, M. others APACrefauthors \ 2022 . Holistic evaluation of language models Holistic evaluation of language models . arXiv preprint arXiv:2211.09110 ,
2022 arXiv
-
[47]
APACrefauthors \ 2004
lin2004rouge APACrefauthors Lin, C Y. APACrefauthors \ 2004 . Rouge: A package for automatic evaluation of summaries Rouge: A package for automatic evaluation of summaries . Text summarization branches out Text summarization branches out \ ( \ 74--81)
2004
-
[48]
, Ronn, N
manes2024k APACrefauthors Manes, I. , Ronn, N. , Cohen, D. , Ber, R.I. , Horowitz-Kugler, Z. Stanovsky, G. APACrefauthors \ 2024 . K-qa: A real-world medical q&a benchmark K-qa: A real-world medical q&a benchmark . arXiv preprint arXiv:2401.14493 ,
2024 arXiv
-
[49]
APACrefauthors \ 1969
McLaughlin1969 APACrefauthors McLaughlin, G.H. APACrefauthors \ 1969 . SMOG Grading—A New Readability Formula Smog grading—a new readability formula . Journal of Reading 12 639--646,
1969
-
[50]
, Lin, Z
meng2021knowledge APACrefauthors Meng, H. , Lin, Z. , Yang, F. , Xu, Y. Cui, L. APACrefauthors \ 2021 . Knowledge distillation in medical data mining: a survey Knowledge distillation in medical data mining: a survey . 5th International Conference on Crowd Science and Engineeri...
2021
-
[51]
, Resnicow, V.P
Min2022 APACrefauthors Min, J. , Resnicow, V.P. , Resnicow, K. Mihalcea, R. APACrefauthors \ 2022 . PAIR: Prompt-aware margIn ranking for counselor reflection scoring in motivational interviewing Pair: Prompt-aware margin ranking for counselor reflection scoring in motivationa...
2022
-
[52]
, Andrzejak, S.E
moore2023exploring APACrefauthors Moore, J.X. , Andrzejak, S.E. , Jones, S. Han, Y. APACrefauthors \ 2023 . Exploring the intersectionality of race/ethnicity with rurality on breast cancer outcomes: SEER analysis, 2000--2016 Exploring the intersectionality of race/ethnicity wi...
2023
-
[53]
, Holaday, B
moore2017cues APACrefauthors Moore de Peralta, A. , Holaday, B. Hadoto, I.M. APACrefauthors \ 2017 . Cues to cervical cancer screening among US Hispanic women Cues to cervical cancer screening among us hispanic women . Hispanic Health Care International 15 1 5--12,
2017
-
[54]
, Shirzad, A
morava2024acute APACrefauthors Morava, A. , Shirzad, A. , Van Riesen, J. , Elshawish, N. , Ahn, J. Prapavessis, H. APACrefauthors \ 2024 . Acute stress negatively impacts on-task behavior and lecture comprehension Acute stress negatively impacts on-task behavior and lecture co...
2024
-
[55]
, Powers, B
obermeyer2019dissecting APACrefauthors Obermeyer, Z. , Powers, B. , Vogeli, C. Mullainathan, S. APACrefauthors \ 2019 . Dissecting racial bias in an algorithm used to manage the health of populations Dissecting racial bias in an algorithm used to manage the health of populatio...
2019
-
[56]
, Banerjee, H.N
olusola2019human APACrefauthors Olusola, P. , Banerjee, H.N. , Philley, J.V. Dasgupta, S. APACrefauthors \ 2019 . Human papilloma virus-associated cervical cancer and health disparities Human papilloma virus-associated cervical cancer and health disparities . Cells 8 6 622,
2019
-
[57]
, Umapathi, L.K
pal2022medmcqa APACrefauthors Pal, A. , Umapathi, L.K. Sankarasubbu, M. APACrefauthors \ 2022 . Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question...
2022
-
[58]
, Fayyaz, H
poulain2024bias APACrefauthors Poulain, R. , Fayyaz, H. Beheshti, R. APACrefauthors \ 2024 . Bias patterns in the application of LLMs for clinical decision support: A comprehensive study Bias patterns in the application of llms for clinical decision support: A comprehensive st...
2024 arXiv
-
[59]
APACrefauthors \ 2019
reimers2019sentence APACrefauthors Reimers, N. APACrefauthors \ 2019 . Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks Sentence-bert: Sentence embeddings using siamese bert-networks . arXiv preprint arXiv:1908.10084 ,
2019 arXiv
-
[60]
, Alaniz, S
salewski2024context APACrefauthors Salewski, L. , Alaniz, S. , Rio-Torto, I. , Schulz, E. Akata, Z. APACrefauthors \ 2024 . In-context impersonation reveals Large Language Models' strengths and biases In-context impersonation reveals large language models' strengths and biases...
2024
-
[61]
\ Mayer, J.D
salovey1990emotional APACrefauthors Salovey, P. \ Mayer, J.D. APACrefauthors \ 1990 . Emotional intelligence Emotional intelligence . Imagination, cognition and personality 9 3 185--211,
1990
-
[62]
, Das, D
Sellam2020 APACrefauthors Sellam, T. , Das, D. Parikh, A.P. APACrefauthors \ 2020 . BLEURT: Learning Robust Metrics for Text Generation Bleurt: Learning robust metrics for text generation . Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics...
2020
-
[63]
, Maher, R
sengupta2021genbit APACrefauthors Sengupta, K. , Maher, R. , Groves, D. Olieman, C. APACrefauthors \ 2021 . GenBiT: measure and mitigate gender bias in language datasets Genbit: measure and mitigate gender bias in language datasets . Microsoft Journal of Applied Research 16 63--71,
2021
-
[64]
\ Smith, E.A
Senter1967 APACrefauthors Senter, R.J. \ Smith, E.A. APACrefauthors \ 1967 . Automated Readability Index Automated readability index \ \ \ AMRL-TR-66-220 . Aerospace Medical Research Laboratories, Wright-Patterson Air Force Base
1967
-
[65]
, Adimi, S
shool2025systematic APACrefauthors Shool, S. , Adimi, S. , Saboori Amleshi, R. , Bitaraf, E. , Golpira, R. Tara, M. APACrefauthors \ 2025 . A systematic review of large language model (LLM) evaluations in clinical medicine A systematic review of large language model (llm) eval...
2025
-
[66]
, Miller, K.D
Siegel2023 APACrefauthors Siegel, R. , Miller, K.D. , Wagle, H.F. \ . APACrefauthors \ 2023 . Cancer statistics, 2023 Cancer statistics, 2023 . CA: A Cancer Journal for Clinicians 73 1 17--48, APACrefURL https://acsjournals.onlinelibrary.wiley.com/doi/10.3322/caac.21763 APACrefURL
2023 doi
-
[67]
, Azizi, S
singhal2023large APACrefauthors Singhal, K. , Azizi, S. , Tu, T. , Mahdavi, S.S. , Wei, J. , Chung, H.W. others APACrefauthors \ 2023 . Large language models encode clinical knowledge Large language models encode clinical knowledge . Nature 620 7972 172--180,
2023
-
[68]
singhal2025toward APACrefauthors Singhal, K. , Tu, T. , Gottweis, J. , Sayres, R. , Wulczyn, E. , Amin, M. others APACrefauthors \ 2025 . Toward expert-level medical question answering with large language models Toward expert-level medical question answering with large languag...
2025
-
[69]
Singhal2025MedPaLM2 APACrefauthors Singhal, K. , Tu, T. , Gottweis, J. , Sayres, R. , Wulczyn, E. , Lee, D. Valiant, A. APACrefauthors \ 2025 . Toward expert-level medical question answering with large language models Toward expert-level medical question answering with large l...
2025 doi
-
[70]
, Larrison, C.R
sinha2023comparing APACrefauthors Sinha, G.R. , Larrison, C.R. , Brooks, I. Kursuncu, U. APACrefauthors \ 2023 . Comparing naturalistic mental health expressions on student loan debts using reddit and twitter Comparing naturalistic mental health expressions on student loan deb...
2023
-
[71]
, Barash, Y
sorin2023large APACrefauthors Sorin, V. , Barash, Y. , Konen, E. Klang, E. APACrefauthors \ 2023 . Large language models for oncological applications Large language models for oncological applications . Journal of Cancer Research and Clinical Oncology 149 11 9505--9508,
2023
-
[72]
, Kim, J.J
spencer2023racial APACrefauthors Spencer, J.C. , Kim, J.J. , Tiro, J.A. , Feldman, S.J. , Kobrin, S.C. , Skinner, C.S. others APACrefauthors \ 2023 . Racial and ethnic disparities in cervical cancer screening from three US healthcare settings Racial and ethnic disparities in c...
2023
-
[73]
, Lazer, D
swire2020public APACrefauthors Swire-Thompson, B. , Lazer, D. \ . APACrefauthors \ 2020 . Public health and online misinformation: challenges and recommendations Public health and online misinformation: challenges and recommendations . Annu Rev Public Health 41 1 433--451,
2020
-
[74]
, Gulrajani, I
taori2023alpaca APACrefauthors Taori, R. , Gulrajani, I. , Zhang, T. , Dubois, Y. , Li, X. , Guestrin, C. Hashimoto, T.B. APACrefauthors \ 2023 . Alpaca: A strong, replicable instruction-following model Alpaca: A strong, replicable instruction-following model . Stanford Center...
2023
-
[75]
, Mesnard, T
team2024gemma APACrefauthors Team, G. , Mesnard, T. , Hardin, C. , Dadashi, R. , Bhupatiraju, S. , Pathak, S. others APACrefauthors \ 2024 . Gemma: Open models based on gemini research and technology Gemma: Open models based on gemini research and technology . arXiv preprint a...
2024 arXiv
-
[76]
, Hassan, R
Thirunavukarasu2023GP_AKT APACrefauthors Thirunavukarasu, A.J. , Hassan, R. , Mahmood, S. , Sanghera, R. , Barzangi, K. , El Mukashfi, M. Shah, S. APACrefauthors \ 2023 . Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observatio...
2023
-
[77]
APACrefauthors \ 1962
tomkins1962affect APACrefauthors Tomkins, S. APACrefauthors \ 1962 . Affect imagery consciousness: Volume I: The positive affects Affect imagery consciousness: Volume i: The positive affects . Springer publishing company
1962
-
[78]
, Tamimi, R.M
warner2012time APACrefauthors Warner, E.T. , Tamimi, R.M. , Hughes, M.E. , Ottesen, R.A. , Wong, Y N. , Edge, S.B. others APACrefauthors \ 2012 . Time to diagnosis and breast cancer stage by race/ethnicity Time to diagnosis and breast cancer stage by race/ethnicity . Breast ca...
2012
-
[79]
APACrefauthors \ 1951
welch1951comparison APACrefauthors Welch, B.L. APACrefauthors \ 1951 . On the comparison of several mean values: an alternative approach On the comparison of several mean values: an alternative approach . Biometrika 38 3/4 330--336,
1951
-
[80]
\ Leung, S O
wu2017can APACrefauthors Wu, H. \ Leung, S O. APACrefauthors \ 2017 . Can Likert scales be treated as interval scales?—A simulation study Can likert scales be treated as interval scales?—a simulation study . Journal of social service research 43 4 527--532,
2017
-
[81]
, Cui, H
xu2023knowledge APACrefauthors Xu, R. , Cui, H. , Yu, Y. , Kan, X. , Shi, W. , Zhuang, Y. Yang, C. APACrefauthors \ 2023 . Knowledge-infused prompting: Assessing and advancing clinical text data generation with large language models Knowledge-infused prompting: Assessing and a...
2023 arXiv
-
[82]
, Kishore, V
Zhang2020 APACrefauthors Zhang, T. , Kishore, V. , Wu, F. , Weinberger, K.Q. Artzi, Y. APACrefauthors \ 2020 . BERTScore: Evaluating Text Generation with BERT Bertscore: Evaluating text generation with bert . International Conference on Learning Representations. International ...
2020
-
[83]
, Qiu, L
zhang2023enhancing APACrefauthors Zhang, T. , Qiu, L. , Guo, Q. , Deng, C. , Zhang, Y. , Zhang, Z. Fu, L. APACrefauthors \ 2023 . Enhancing uncertainty-based hallucination detection with stronger focus Enhancing uncertainty-based hallucination detection with stronger focus . a...
2023 arXiv
-
[84]
, Chen, T
Zhu2025CancerMyth APACrefauthors Zhu, W.B. , Chen, T. , Lin, C Y. , Law, J. Smith, R. APACrefauthors \ 2025 . Cancer-Myth : Evaluating Large Language Models on Patient Questions with False Presuppositions Cancer-Myth : Evaluating large language models on patient questions with...
2025
-
[85]
, Le, T.D
zitu2025large APACrefauthors Zitu, M.M. , Le, T.D. , Duong, T. , Haddadan, S. , Garcia, M. , Amorrortu, R. Thieu, T. APACrefauthors \ 2025 . Large language models in cancer: potentials, risks, and safeguards Large language models in cancer: potentials, risks, and safeguards \ ...
2025
-
[86]
write newline
" write newline " cite write " FUNCTION editor.postfix editor num.names #1 > "( )" "( )" if FUNCTION editor.trans.postfix editor num.names #1 > "( )" "( )" if FUNCTION trans.postfix translator num.names #1 > "( )" "( )" if FUNCTION authors.editors.reflist.apa5 'field := 'dot :...
-
[87]
write newline
" write newline "" before.all 'output.state := FUNCTION string.to.integer 't := t text.length 'k := #1 'char.num := t char.num #1 substring 's := s is.num s "." = or char.num k = not and char.num #1 + 'char.num := while char.num #1 - 'char.num := t #1 char.num substring FUNCTI...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.