REVIEW 3 major objections 5 minor 40 references
Communication Styles and Reader Preferences of LLM and Human Experts in Explaining Health Information
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Lay readers prefer LLM-generated fact-checking articles over human ones for clarity, completeness, and persuasiveness, even though the AI text scores lower on standard health-communication quality measures.
desk verdict A well-measured blind preference study whose central 'despite lower quality' claim is undercut by the hallucinated-citation confound the authors themselves acknowledge. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a three-part operationalization of health communication style—information (automated certainty scores, psycholinguistic cognitive-process counts, and readability indices), sender (a hierarchical persuasive-strategy model scoring credibility, evidence, and impact), and receiver (social-value and moral-foundation alignment scores)—combined with a blind paired-preference protocol in which 99 readers each rate five human/LLM article pairs without knowing which text is AI-generated. The divergence between the automated style scores and the reader ratings is the discovery: the metrics say LLMs are weaker on persuasion, certainty, and values, while readers say LLMs are clearer, more complete, and more persuasive.
What would settle it
Repeat the same blind paired-preference protocol using full-length human fact-checking articles rather than only the 250–350-word subset, and add a condition in which participants are told which article is AI-generated; if the LLM preference disappears or reverses under either change, the claimed disconnect between style scores and reader preference is an artifact of length matching or of perceived rather than actual quality.
Extended reading notes
Core claim
This paper reports a disconnect between automated measures of health communication style and lay readers' actual preferences. Across 495 blinded ratings, more than 60% preferred LLM-generated fact-checking articles on language clarity, inclusion of necessary information, and persuasiveness, while the same LLM texts scored significantly lower than human writing on persuasive strategies, certainty expressions, alignment with social values, and moral foundations. Readers attributed their preference to focused presentation, an objective and neutral tone, a professional appearance, and accessible language. The paper argues that LLMs' structured way of presenting information may be more effective at engaging readers despite scoring lower on traditional quality dimensions in fact-checking and health communication.
Load-bearing premise
The automated measures of certainty, persuasion, and value alignment really capture the dimensions of health communication that matter to readers, and the 250–350-word human articles used in the preference test fairly represent how human fact-checkers write.
Editorial extensions
If this is right
- Fact-checking organizations could use LLMs to restructure verified information into a focused, neutral, accessible format while keeping human experts in charge of evidence and values.
- Automated style benchmarks alone are not enough to judge health communication; reader perception needs to be part of the evaluation because the two can disagree sharply.
- Because LLM articles can create an impression of completeness and professionalism while citing sources that may be incorrect or nonexistent, deploying them without human oversight risks increasing trust in text that is less rigorous than it appears.
- The mismatch between higher readability-complexity scores for LLM articles and readers' perception of clearer language suggests that readability formulas miss the structural clarity readers actually experience.
Reading between the lines
- The preference test matched human articles to LLM length by selecting only human articles of 250–350 words, while the human dataset averages 765 words; the LLM advantage may therefore partly reflect brevity and structure rather than a general superiority, and a test with full-length human articles would separate these.
- The 'neutral' and 'professional' perceptions readers reported could weaken if AI authorship were disclosed before rating; a simple extension is to repeat the blind protocol with a disclosure condition.
- The findings imply a testable prediction: among texts about the same claim, readers will systematically prefer the most structured, shortest, and least emotionally charged version, even when it contains fewer evidence cues—this could be checked with controlled rewrites of a single article.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript compares communication styles of LLM-generated fact-checking articles versus human fact-checking articles on health misinformation, and then measures reader preferences in a blinded study. Using a corpus of 1,498 human articles and LLM responses generated with zero-shot and few-shot chain-of-thought prompting, the authors measure linguistic certainty, cognitive processing, readability, persuasive strategies, and value/moral alignment using published computational tools. They find that LLM articles score significantly lower than human articles on several expert-oriented dimensions, including persuasive impact, evidence use, certainty, and moral/value alignment. Yet in a randomized blinded evaluation with 99 participants and 495 pairwise ratings, readers preferred LLM articles on clarity, completeness, and persuasiveness in over 60% of responses. The authors interpret this as evidence that LLMs' structured presentation may be more engaging despite lower scores on traditional quality measures, while also acknowledging that LLMs may cite non-existent or incorrect sources, which could contribute to perceived professionalism.
Significance. If the core comparison is robust, the paper addresses an important practical question: whether expert-derived communication benchmarks align with what lay readers actually value in AI-generated health fact-checking. The study has notable strengths: a large corpus of 1,498 human articles, multiple LLMs with zero- and few-shot prompting, use of published and externally validated measures (e.g., Pei-Jurgens certainty, the Chen-Yang persuasive-strategy model, Mformer, Schwartz values), and a blinded, randomized, attention-checked human evaluation with quantitative ratings and qualitative open-ended reasons. The finding that readers prefer LLM articles despite lower automated style scores is potentially consequential for health communication, fact-checking workflows, and LLM interface design. However, the central interpretive claim depends on ruling out confounds such as fabricated citations and article-length mismatch, and on the validity of the automated style measures as proxies for quality; the manuscript does not yet fully establish this.
major comments (3)
- [Discussion; Results, Human Evaluation] The paper's central "despite" conclusion is undercut by the acknowledged presence of fabricated citations. In the Discussion, the authors write that "LLMs' tendency to cite reputable sources, even when citations are incorrect or non-existent, contributes to a perception of professionalism and evidence support." The human-preference study reports no citation-verification step, so the 62% preference for LLM articles on "necessary information" and "persuasiveness" may reflect readers being convinced by false source cues rather than by the LLM's structural presentation. This is a direct confound for the headline claim that LLMs' structured approach "may be more effective at engaging readers despite scoring lower on traditional measures." The authors need either to add a control condition that removes or verifies citations, or to reframe the conclusion as being about current, hallucination-prone LLM output rather than about the value of LLM presentation style per se.
- [Methods, Human Evaluation] The length filter is applied to human articles only. The text states that the preference subset was selected "where the human fact-checking articles were within 250-350 words to ensure no identifiable difference between LLM-generated articles and human articles," but no analogous filter is described for LLM-generated articles. Given that the full LLM dataset has average length 335 words and range [9-2135], the preference pairs may compare length-matched human articles to unfiltered LLM articles, confounding the outcome with article length and overall structure. The authors should report the length distribution of the LLM articles actually used in the preference pairs and demonstrate that they are comparable, or apply the same 250-350-word criterion to the LLM articles.
- [Results, Human Evaluation (Table 4)] The regression analysis treats the 495 responses as independent observations, but each of the 99 participants contributed five article-pair ratings, creating repeated measurements. The multivariate ordinal logistic regressions in Table 4 do not appear to include participant-level random effects or cluster-robust standard errors. This can underestimate standard errors and overstate the significance of demographic and content predictors. The authors should fit a mixed-effects ordinal model with participant as a random intercept, or use cluster-robust standard errors by participant, to support the regression-based claims about which factors predict preference.
minor comments (5)
- [Results, Language clarity] The text contains an internal contradiction: it says "education levels and content readability were not significantly related to language clarity preferences," then immediately states "participants with higher education levels were more likely to find LLMs clearer, and that LLM articles with lower readability scores (i.e., easier to read) were rated as clearer." Please reconcile this statement with Table 4 and with the reported nonsignificant odds ratios for education and ARI.
- [Table 5] The "meandiff" column in Table 5 does not state its direction (whether it is Human minus LLM or LLM minus Human). Without this definition, the supplemented values are not interpretable, and the Flesch-Kincaid row for GPT-4 (negative meandiff) appears inconsistent with the ARI row (positive meandiff) and with the text claiming that LLM articles have higher readability scores, meaning they are more complex. Please clarify the sign convention and check the reported directions.
- [Methods, LLM Fact-checking Articles] The few-shot generation produced 1,492 articles versus 1,498 for zero-shot, but the reason for the six missing few-shot articles is not explained. Please clarify whether this was due to API failures, output-length limits, or another exclusion criterion.
- [Results, Correctness] The correctness section reports only recall for the veracity classification task. Precision, F1, and the classification prompt's true/false balance should be reported to give a complete picture of the LLMs' fact-checking accuracy.
- [Methods, Human Evaluation] The description of the rating protocol should specify whether the five article pairs were sampled with or without replacement, and whether a given claim could appear in more than one pair rated by the same participant. This affects the independence assumptions of the Wilcoxon and regression analyses.
Circularity Check
No circularity: preference outcomes are directly observed from blinded participants and style dimensions come from independent published measures.
full rationale
The paper's derivation chain is empirical rather than definitional. The automated style comparisons (certainty, cognitive processing, readability, persuasive strategies, social values, moral foundations) are computed with externally published tools (Pei & Jurgens, LIWC, Chen & Yang's persuasive-strategy model, Schwartz values, Mformer) and are not fitted to the preference outcome. The preference result is directly measured: 99 blind participants provided 495 pairwise ratings ('Without informing the involvement of AI generations and with the order of human-AI article pairs randomized, participants answered three rating questions on a seven-point scale'), and the paper reports raw counts such as '308 out of 495 responses (62.23%) preferred LLM-generated articles.' The central 'despite' framing contrasts these two independent measurements; it is an interpretive synthesis, not a quantity derived from the same inputs. The one author-overlapping citation (Ref. 21, Zhou et al., used to construct the prompting guideline) is a method provenance choice, not a load-bearing premise: the reader-preference conclusion does not reduce to that guideline, and no uniqueness claim or ansatz is imported from it. The paper's own admission that 'LLMs' tendency to cite reputable sources, even when citations are incorrect or non-existent, contributes to a perception of professionalism and evidence support' is a post-hoc explanation of the observed preference and raises a validity/confound concern about hallucinated citations, but it does not make the preference outcome equivalent to an input. No fitted parameter is renamed as a prediction, and no equation reduces to its own input. Hence no circular step is exhibited.
Assumptions & free parameters
assumptions (2)
- domain assumption The automated NLP measures (certainty model, LIWC, Mformer, Schwartz value codings) are treated as valid operationalizations of the health communication constructs they claim to measure.
- domain assumption The 99 Prolific participants are representative of lay U.S. readers for drawing general preference conclusions.
Cite this review
Pith. "Pith review of Communication Styles and Reader Preferences of LLM and Human Experts in Explaining Health Information." pith.science (2026). https://pith.science/paper/SXCYPXUB
@misc{pith2026250508143,
author = {Pith},
title = {Pith review of: Communication Styles and Reader Preferences of LLM and Human Experts in Explaining Health Information},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXCYPXUB}},
note = {Machine review of arXiv:2505.08143}
}
read the original abstract
With the wide adoption of large language models (LLMs) in information assistance, it is essential to examine their alignment with human communication styles and values. We situate this study within the context of fact-checking health information, given the critical challenge of rectifying conceptions and building trust. Recent studies have explored the potential of LLM for health communication, but style differences between LLMs and human experts and associated reader perceptions remain under-explored. In this light, our study evaluates the communication styles of LLMs, focusing on how their explanations differ from those of humans in three core components of health communication: information, sender, and receiver. We compiled a dataset of 1498 health misinformation explanations from authoritative fact-checking organizations and generated LLM responses to inaccurate health information. Drawing from health communication theory, we evaluate communication styles across three key dimensions of information linguistic features, sender persuasive strategies, and receiver value alignments. We further assessed human perceptions through a blinded evaluation with 99 participants. Our findings reveal that LLM-generated articles showed significantly lower scores in persuasive strategies, certainty expressions, and alignment with social values and moral foundations. However, human evaluation demonstrated a strong preference for LLM content, with over 60% responses favoring LLM articles for clarity, completeness, and persuasiveness. Our results suggest that LLMs' structured approach to presenting information may be more effective at engaging readers despite scoring lower on traditional measures of quality in fact-checking and health communication.
Figures
Reference graph
Works this paper leans on
-
[1]
medical Internet research25, e47621 (2023)
Kuroiwa, T.et al.The potential of chatgpt as a self-diagnostic tool in common orthopedic diseases: exploratory study.J. medical Internet research25, e47621 (2023)
work page 2023
-
[2]
The role of using chatgpt ai in writing medical scientific articles.J
Benichou, L. The role of using chatgpt ai in writing medical scientific articles.J. Stomatol. Oral Maxillofac. Surg.124, 101456, DOI: https://doi.org/10.1016/j.jormas.2023.101456 (2023)
arXiv 2023
-
[3]
Wang, Z., Yang, Z., Azimi, I. & Rahmani, A. M. Differential private federated transfer learning for mental health monitoring in everyday settings: A case study on stress detection. InProceedings of the 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)(2024)
work page 2024
- [4]
-
[5]
Guo, Z.et al.Large language models for mental health applications: Systematic review.JMIR Ment Heal.11, e57400, DOI: 10.2196/57400 (2024)
doi:10.2196/57400 2024
-
[6]
Rashkin, H., Choi, E., Jang, J. Y ., V olkova, S. & Choi, Y . Truth of varying shades: Analyzing language in fake news and political fact-checking. In Palmer, M., Hwa, R. & Riedel, S. (eds.)Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2931–2937, DOI: 10.18653/v1/D17-1317 (Association for Computational Linguistics...
-
[7]
Kreuter, M. W. & McClure, S. M. The role of culture in health communication.Annu. Rev. Public Heal.25, 439–455 (2004). 8.Schiavo, R.Health communication: From theory to practice(John Wiley & Sons, 2013)
work page 2004
-
[9]
Mheidly, N. & Fares, J. Leveraging media and health communication strategies to overcome the covid-19 infodemic.J. public health policy41, 410–420 (2020)
work page 2020
Show all 40 references
-
[10]
R., Zhang, Y
Bautista, J. R., Zhang, Y . & Gwizdka, J. Healthcare professionals’ acts of correcting health misinformation on social media. Int. J. Med. Informatics148, 104375 (2021)
2021
-
[11]
Rev Public Heal.41, 433–451 (2020)
Swire-Thompson, B., Lazer, D.et al.Public health and online misinformation: challenges and recommendations.Annu. Rev Public Heal.41, 433–451 (2020)
2020
-
[12]
& Vlachos, A
Guo, Z., Schlichtkrull, M. & Vlachos, A. A survey on automated fact-checking.Transactions Assoc. for Comput. Linguist. 10, 178–206, DOI: 10.1162/tacl_a_00454 (2022). https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl_a_00454/1987018/ tacl_a_00454.pdf
2022 doi
-
[13]
& Toutanova, K
Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C. & Solorio, T. (eds.)Proceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...
2023 doi
-
[15]
& Gao, W
Zhang, X. & Gao, W. Towards llm-based fact verification on news claims with a hierarchical step-by-step prompting method (2023). 2310.00305
2023 arXiv
-
[16]
S., Chinn, S
Hart, P. S., Chinn, S. & Soroka, S. Politicization and polarization in covid-19 news coverage.Sci. communication42, 679–697 (2020)
2020
-
[17]
Tasnim, S., Hossain, M. M. & Mazumder, H. Impact of rumors and misinformation on covid-19 in social media.J. preventive medicine public health53, 171–174 (2020)
2020
-
[18]
A., Kalathur Gopal, S
Saenz, J. A., Kalathur Gopal, S. R. & Shukla, D. Covid-19 fake news infodemic research dataset (covid19-fnir dataset), DOI: 10.21227/b5bt-5244 (2021)
2021 doi
-
[19]
Shahi, G. K. & Nandini, D.FakeCovid- A Multilingual Cross-domain Fact Check News Dataset for COVID-19(ICWSM, 2020)
2020
-
[20]
neural information processing systems35, 24824–24837 (2022)
Wei, J.et al.Chain-of-thought prompting elicits reasoning in large language models.Adv. neural information processing systems35, 24824–24837 (2022)
2022
-
[21]
Zhou, J., Zhang, Y ., Luo, Q., Parker, A. G. & De Choudhury, M. Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, 1–20, DOI: 10.1145...
2023
-
[25]
& Rosenstiel, T.Blur: How to know what’s true in the age of information overload(Bloomsbury Publishing USA, 2011)
Kovach, B. & Rosenstiel, T.Blur: How to know what’s true in the age of information overload(Bloomsbury Publishing USA, 2011)
2011
-
[26]
X.et al.A structured response to misinformation: Defining and annotating credibility indicators in news articles
Zhang, A. X.et al.A structured response to misinformation: Defining and annotating credibility indicators in news articles. InCompanion Proceedings of the The Web Conference 2018, 603–612 (2018)
2018
-
[27]
& Lee, M
Parnami, A. & Lee, M. Learning from few examples: A summary of approaches to few-shot learning.arXiv preprint arXiv:2203.04291(2022)
2022 arXiv
-
[28]
& De Choudhury, M
Mittal, S., Jung, H., ElSherief, M., Mitra, T. & De Choudhury, M. Online myths on opioid use disorder: A comparison of reddit and large language model.Proc. ICWSM(2025)
2025
-
[29]
& Karkar, R
Saha, K., Jain, Y ., Liu, C., Kaliappan, S. & Karkar, R. Ai vs. humans for online support: Comparing the language of responses from llms and online communities of alzheimer’s disease.ACM Transactions on Comput. for Healthc.(2025). 30.Berry, D.Risk, communication and health psy...
2025
-
[31]
& Jurgens, D
Pei, J. & Jurgens, D. Measuring sentence-level and aspect-level (un)certainty in science communications. In Moens, M.-F., Huang, X., Specia, L. & Yih, S. W.-t. (eds.)Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 9959–10011, DOI: 10.186...
2001 doi
-
[33]
& Wilson, C
Jiang, S. & Wilson, C. Linguistic signals under misinformation and fact-checking: Evidence from user comments on social media.Proc. ACM on Human-Computer Interact.2, 1–23 (2018)
2018
-
[34]
Su, Q., Wan, M., Liu, X., Huang, C.-R.et al.Motivations, methods and metrics of misinformation detection: an nlp perspective.Nat. Lang. Process. Res.1, 1–13 (2020)
2020
-
[35]
& Ghorbani, A
Zhang, X. & Ghorbani, A. A. An overview of online fake news: Characterization, detection, and discussion.Inf. Process. & Manag.57, 102025 (2020)
2020
-
[36]
Smith, E. A. & Senter, R.Automated readability index, vol. 66 (Aerospace Medical Research Laboratories, Aerospace Medical Division, Air . . . , 1967). 37.Flesch, R. Flesch-kincaid readability test.Retrieved Oct.26, 2007 (2007)
2007
-
[38]
& Mao, J
Chen, S., Xiao, L. & Mao, J. Persuasion strategies of misinformation-containing posts in the social media.Inf. Process. & Manag.58, 102665 (2021)
2021
-
[39]
& Yang, D
Chen, J. & Yang, D. Weakly-supervised hierarchical models for predicting persuasive strategies in good-faith textual requests (2021). 2101.06351
2021 arXiv
-
[40]
Basic human values: Theory, measurement, and applications.Revue Francaise de Sociol.47, 929– 968+977+981 (2006)
Schwartz, S. Basic human values: Theory, measurement, and applications.Revue Francaise de Sociol.47, 929– 968+977+981 (2006)
2006
-
[41]
Van Der Meer, M., V ossen, P., Jonker, C. M. & Murukannaiah, P. K. Do differences in values influence disagreements in online discussions?arXiv preprint arXiv:2310.15757(2023). 42.Nguyen, T. D.et al.Measuring moral dimensions in social media with mformer (2024). 2311.10219
2023 arXiv
-
[43]
InAdvances in experimental social psychology, vol
Graham, J.et al.Moral foundations theory: The pragmatic validity of moral pluralism. InAdvances in experimental social psychology, vol. 47, 55–130 (Elsevier, 2013)
2013
-
[44]
InProceedings of the 12th ACM conference on web Science, 1–10 (2020)
Im, J.et al.Still out there: Modeling and identifying russian troll accounts on twitter. InProceedings of the 12th ACM conference on web Science, 1–10 (2020)
2020
-
[45]
Kuehn, K. M. & Salter, L. A. Assessing digital threats to democracy, and workable solutions: a review of the recent literature.Int. J. Commun.14, 22 (2020)
2020
-
[46]
Thomas, D. R. A general inductive approach for analyzing qualitative evaluation data.Am. journal evaluation27, 237–246 (2006)
2006
-
[47]
The persuasiveness of source credibility: A critical review of five decades’ evidence.J
Pornpitakpan, C. The persuasiveness of source credibility: A critical review of five decades’ evidence.J. applied social psychology34, 243–281 (2004). 13/15
2004
-
[48]
medical Internet research5, e893 (2003)
Dutta-Bergman, M.et al.Trusted online sources of health information: differences in demographics, health beliefs, and health-information orientation.J. medical Internet research5, e893 (2003)
2003
-
[49]
Thon, F. M. & Jucks, R. Believing in expertise: How authors’ credentials and language use influence the credibility of online health information.Heal. communication32, 828–836 (2017). 14/15 Supplement Dimensions LLM meandiffp-adj Persuasive Strategies Impact GPT4 -0.2292 *** L...
2017
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.