{"id":"e72d29d3-aca6-431f-9d60-b7dd80dbb92a","arxiv_id":"2505.08143","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM-generated health fact-checking articles score lower on expert-style communication metrics but are preferred by lay readers, suggesting structured presentation may outweigh traditional quality cues.","lead":"Researchers compared 1,498 human-written fact-checking articles about COVID-19 misinformation with AI-generated versions and asked 99 readers which they preferred. They found AI articles scored lower on traditional quality measures like persuasion and value alignment, but readers preferred them for clarity, completeness, and persuasiveness.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The preference result is confounded by hallucinated citations, an issue the paper itself acknowledges; the 'despite lower quality scores' claim needs a citation-free control condition.","rationale":"The reader's weakest assumption concerned proxy validity and subset representativeness; those are legitimate but secondary. The hallucinated-citation issue is more load-bearing because it determines whether the 'despite' framing is valid at all. The manuscript itself states that citations are sometimes incorrect or non-existent and that these citations contribute to perceived professionalism and evidence support. That admission means the observed preference for LLM 'completeness' and 'persuasiveness' could be an artifact of fake source cues, not a genuine advantage of the LLM's structured presentation style. The paper does not report any audit of citation accuracy for the specific articles rated by participants, so the central conclusion is not separated from this confound. A citation audit and a citation-free replication would settle the issue: if preference survives removal or correction of references, the central claim is supported; if not, the claim should be narrowed to 'fabricated citations can increase perceived persuasiveness,' which is a different and more cautionary result. Because the paper's existing conditional verdict already flags this area but does not require the decisive test, I recommend keeping the verdict conditional while making the citation audit and citation-free control an explicit condition of acceptance.","tokens_in":12892,"tokens_out":6174,"duration_ms":66487,"concrete_test":"Extract all references from the 495 LLM articles shown to participants and verify each citation (title, venue, DOI) against a source database; report the proportion that are incorrect or non-existent. Then run a follow-up preference study with the same paired design and three arms: original LLM articles, LLM articles with references removed, and LLM articles with references replaced by verified, real sources. If LLM preference for completeness and persuasiveness is not significantly above chance in the citation-free or citation-corrected arms, the observed preference is attributable to fabricated source cues rather than to the LLM's structured presentation style, and the central claim should be revised accordingly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing problem is that the human-preference result cannot be attributed to communication style while LLM articles contain fabricated references. In the Discussion, the authors write that 'LLMs' tendency to cite reputable sources, even when citations are incorrect or non-existent, contributes to a perception of professionalism and evidence support.' This admission directly undercuts the central 'despite' framing: readers rated LLM articles higher on 'inclusion of necessary information' and 'persuasiveness' in a blinded comparison, and the qualitative reasons include 'references from reputable sources.' Since no citation-verification step is reported, the 495 preference ratings may reflect participants being convinced by fake source cues rather than by the LLM's 'structured approach to presenting information.' The automated analyses also show LLM articles score lower on the persuasive-strategy Evidence dimension, so the apparent paradox is exactly what one would expect if hallucinated citations are supplying the missing evidence. Without controlling for citation accuracy, the paper's claim that LLM structure 'may be more effective at engaging readers despite scoring lower on traditional measures' is not established for truthful LLM output. The concern is not that hallucination is unmentioned; it is that the paper uses it as a post-hoc explanation but still frames the preference as evidence for the value of LLM presentation style.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares communication styles of LLM-generated fact-checking articles versus human fact-checking articles on health misinformation, and then measures reader preferences in a blinded study. Using a corpus of 1,498 human articles and LLM responses generated with zero-shot and few-shot chain-of-thought prompting, the authors measure linguistic certainty, cognitive processing, readability, persuasive strategies, and value/moral alignment using published computational tools. They find that LLM articles score significantly lower than human articles on several expert-oriented dimensions, including persuasive impact, evidence use, certainty, and moral/value alignment. Yet in a randomized blinded evaluation with 99 participants and 495 pairwise ratings, readers preferred LLM articles on clarity, completeness, and persuasiveness in over 60% of responses. The authors interpret this as evidence that LLMs' structured presentation may be more engaging despite lower scores on traditional quality measures, while also acknowledging that LLMs may cite non-existent or incorrect sources, which could contribute to perceived professionalism.","tokens_in":13112,"tokens_out":5014,"duration_ms":49611,"significance":"If the core comparison is robust, the paper addresses an important practical question: whether expert-derived communication benchmarks align with what lay readers actually value in AI-generated health fact-checking. The study has notable strengths: a large corpus of 1,498 human articles, multiple LLMs with zero- and few-shot prompting, use of published and externally validated measures (e.g., Pei-Jurgens certainty, the Chen-Yang persuasive-strategy model, Mformer, Schwartz values), and a blinded, randomized, attention-checked human evaluation with quantitative ratings and qualitative open-ended reasons. The finding that readers prefer LLM articles despite lower automated style scores is potentially consequential for health communication, fact-checking workflows, and LLM interface design. However, the central interpretive claim depends on ruling out confounds such as fabricated citations and article-length mismatch, and on the validity of the automated style measures as proxies for quality; the manuscript does not yet fully establish this.","major_comments":[{"comment":"The paper's central \"despite\" conclusion is undercut by the acknowledged presence of fabricated citations. In the Discussion, the authors write that \"LLMs' tendency to cite reputable sources, even when citations are incorrect or non-existent, contributes to a perception of professionalism and evidence support.\" The human-preference study reports no citation-verification step, so the 62% preference for LLM articles on \"necessary information\" and \"persuasiveness\" may reflect readers being convinced by false source cues rather than by the LLM's structural presentation. This is a direct confound for the headline claim that LLMs' structured approach \"may be more effective at engaging readers despite scoring lower on traditional measures.\" The authors need either to add a control condition that removes or verifies citations, or to reframe the conclusion as being about current, hallucination-prone LLM output rather than about the value of LLM presentation style per se.","section":"Discussion; Results, Human Evaluation"},{"comment":"The length filter is applied to human articles only. The text states that the preference subset was selected \"where the human fact-checking articles were within 250-350 words to ensure no identifiable difference between LLM-generated articles and human articles,\" but no analogous filter is described for LLM-generated articles. Given that the full LLM dataset has average length 335 words and range [9-2135], the preference pairs may compare length-matched human articles to unfiltered LLM articles, confounding the outcome with article length and overall structure. The authors should report the length distribution of the LLM articles actually used in the preference pairs and demonstrate that they are comparable, or apply the same 250-350-word criterion to the LLM articles.","section":"Methods, Human Evaluation"},{"comment":"The regression analysis treats the 495 responses as independent observations, but each of the 99 participants contributed five article-pair ratings, creating repeated measurements. The multivariate ordinal logistic regressions in Table 4 do not appear to include participant-level random effects or cluster-robust standard errors. This can underestimate standard errors and overstate the significance of demographic and content predictors. The authors should fit a mixed-effects ordinal model with participant as a random intercept, or use cluster-robust standard errors by participant, to support the regression-based claims about which factors predict preference.","section":"Results, Human Evaluation (Table 4)"}],"minor_comments":[{"comment":"The text contains an internal contradiction: it says \"education levels and content readability were not significantly related to language clarity preferences,\" then immediately states \"participants with higher education levels were more likely to find LLMs clearer, and that LLM articles with lower readability scores (i.e., easier to read) were rated as clearer.\" Please reconcile this statement with Table 4 and with the reported nonsignificant odds ratios for education and ARI.","section":"Results, Language clarity"},{"comment":"The \"meandiff\" column in Table 5 does not state its direction (whether it is Human minus LLM or LLM minus Human). Without this definition, the supplemented values are not interpretable, and the Flesch-Kincaid row for GPT-4 (negative meandiff) appears inconsistent with the ARI row (positive meandiff) and with the text claiming that LLM articles have higher readability scores, meaning they are more complex. Please clarify the sign convention and check the reported directions.","section":"Table 5"},{"comment":"The few-shot generation produced 1,492 articles versus 1,498 for zero-shot, but the reason for the six missing few-shot articles is not explained. Please clarify whether this was due to API failures, output-length limits, or another exclusion criterion.","section":"Methods, LLM Fact-checking Articles"},{"comment":"The correctness section reports only recall for the veracity classification task. Precision, F1, and the classification prompt's true/false balance should be reported to give a complete picture of the LLMs' fact-checking accuracy.","section":"Results, Correctness"},{"comment":"The description of the rating protocol should specify whether the five article pairs were sampled with or without replacement, and whether a given claim could appear in more than one pair rated by the same participant. This affects the independence assumptions of the Wilcoxon and regression analyses.","section":"Methods, Human Evaluation"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and societally relevant topic for a human-computer interaction or health-communication audience. The main concern is not the novelty but the causal interpretation of the preference result, which needs to be made robust against the citation-hallucination and length confounds. I believe these are addressable within a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this paper gives a large, careful measurement of how LLM and human fact-checking articles differ on standard communication-style dimensions, and then shows, in a blinded test, that lay readers prefer the LLM articles on clarity, completeness, and persuasiveness. That preference is real, as measured. The problem is the authors' framing of it as evidence that LLM presentation style can compensate for lower quality scores. They can't make that leap, because the LLM articles contain fabricated citations, and the participants explicitly said those references made them seem professional and well-supported. In other words, the preference may be for the illusion of evidence, not for the structure.\n\nWhat's genuinely new: the combination of automated style measures (Pei-Jurgens certainty, persuasive-strategy model, Mformer, Schwartz values) with a randomized, blinded human evaluation on the same set of factual claims. That's a solid design, and the style differences themselves (less certainty, less evidence language, more complex readability, lower authority/fairness) are worth having. The authors also do an honest qualitative analysis of why readers preferred LLM text, and they don't hide the hallucination issue: they bring it up directly in the Discussion. Credit where due: the data collection is substantial (1,498 human articles, multiple LLMs, 99 participants, 495 paired ratings), and the core stats are largely appropriate.\n\nSoft spots, in order of weight. First, the citation confound. The paper never checks whether the LLM-generated references actually exist. Given that GPT-4 and other models are known to hallucinate, and the authors themselves say 'citations are incorrect or non-existent,' the preference result cannot be attributed to communication style. A reader who sees 'references from reputable sources' is not responding to structure alone; they're responding to an authenticity cue. A citation-free control condition, or at least a citation-verification step, is needed before the 'despite lower quality scores' claim holds.\n\nSecond, the length filtering is asymmetrical. They selected human articles of 250-350 words to match LLM output, but they don't say they filtered the LLM articles the same way. LLM articles range from 9 to 2,135 words, and the human subset is a narrow band. That could introduce structural differences unrelated to style. Moderate concern.\n\nThird, the regression reporting doesn't account for repeated measures: each participant rated five pairs, so the 495 responses are not independent. They use ordinal logistic regression without obvious clustering or random effects. Fixable, and it might shift some of the demographic p-values.\n\nIs it worth engaging? Yes, but the conclusion needs reframing: readers preferred LLM articles in this setting, yet the reasons include potentially misleading citation cues, so we cannot yet say whether truthful LLM output would keep that edge. This is a good starting point for a follow-up that controls for citation accuracy. I'd send it to review, expecting the revision to add that control or substantially soften the claim.\n\nRecommendation: serious referee, conditional accept after major revision.","headline":"A well-measured blind preference study whose central 'despite lower quality' claim is undercut by the hallucinated-citation confound the authors themselves acknowledge.","tokens_in":13616,"tokens_out":3049,"would_cite":false,"duration_ms":30707,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Lay readers prefer LLM-generated fact-checking articles over human ones for clarity, completeness, and persuasiveness, even though the AI text scores lower on standard health-communication quality measures.","keywords":["health misinformation","fact-checking","large language models","health communication","reader preferences","communication style","value alignment","persuasive strategies"],"falsifier":"Repeat the same blind paired-preference protocol using full-length human fact-checking articles rather than only the 250–350-word subset, and add a condition in which participants are told which article is AI-generated; if the LLM preference disappears or reverses under either change, the claimed disconnect between style scores and reader preference is an artifact of length matching or of perceived rather than actual quality.","tokens_in":12720,"feed_emoji":"🩺","tokens_out":8711,"duration_ms":79406,"temperature":0.7,"pith_summary":"The paper asks whether large language models explain health misinformation the way human experts do, and whether readers care. It builds a dataset of 1,498 human fact-checking articles about COVID-19 claims, generates LLM counterparts with chain-of-thought prompting, and measures communication style along three components: information, sender, and receiver. The LLM articles score significantly lower on persuasive strategies, certainty, and value and moral alignment, and their readability metrics suggest more complex language. Yet in a blinded study with 99 readers, over 60% of 495 ratings preferred the LLM articles for clarity, completeness, and persuasiveness. The paper concludes that structured, neutral presentation can engage readers more effectively than traditional markers of high-quality health communication.","feed_headline":"Readers prefer LLM fact-checks despite lower style scores","feed_subtitle":"Blind test: 62% preferred AI fact-checks for clarity, completeness, persuasiveness, despite lower expert quality scores.","key_machinery":"The carrying mechanism is a three-part operationalization of health communication style—information (automated certainty scores, psycholinguistic cognitive-process counts, and readability indices), sender (a hierarchical persuasive-strategy model scoring credibility, evidence, and impact), and receiver (social-value and moral-foundation alignment scores)—combined with a blind paired-preference protocol in which 99 readers each rate five human/LLM article pairs without knowing which text is AI-generated. The divergence between the automated style scores and the reader ratings is the discovery: the metrics say LLMs are weaker on persuasion, certainty, and values, while readers say LLMs are clearer, more complete, and more persuasive.","core_discovery":"This paper reports a disconnect between automated measures of health communication style and lay readers' actual preferences. Across 495 blinded ratings, more than 60% preferred LLM-generated fact-checking articles on language clarity, inclusion of necessary information, and persuasiveness, while the same LLM texts scored significantly lower than human writing on persuasive strategies, certainty expressions, alignment with social values, and moral foundations. Readers attributed their preference to focused presentation, an objective and neutral tone, a professional appearance, and accessible language. The paper argues that LLMs' structured way of presenting information may be more effective at engaging readers despite scoring lower on traditional quality dimensions in fact-checking and health communication.","pith_inferences":["The preference test matched human articles to LLM length by selecting only human articles of 250–350 words, while the human dataset averages 765 words; the LLM advantage may therefore partly reflect brevity and structure rather than a general superiority, and a test with full-length human articles would separate these.","The 'neutral' and 'professional' perceptions readers reported could weaken if AI authorship were disclosed before rating; a simple extension is to repeat the blind protocol with a disclosure condition.","The findings imply a testable prediction: among texts about the same claim, readers will systematically prefer the most structured, shortest, and least emotionally charged version, even when it contains fewer evidence cues—this could be checked with controlled rewrites of a single article."],"forward_implications":["Fact-checking organizations could use LLMs to restructure verified information into a focused, neutral, accessible format while keeping human experts in charge of evidence and values.","Automated style benchmarks alone are not enough to judge health communication; reader perception needs to be part of the evaluation because the two can disagree sharply.","Because LLM articles can create an impression of completeness and professionalism while citing sources that may be incorrect or nonexistent, deploying them without human oversight risks increasing trust in text that is less rigorous than it appears.","The mismatch between higher readability-complexity scores for LLM articles and readers' perception of clearer language suggests that readability formulas miss the structural clarity readers actually experience."],"supporting_citations":[{"why":"Supplies the COVID-19 fake-news fact-checking dataset from which human articles were drawn.","marker":"[18]"},{"why":"Supplies the FakeCovid cross-domain fact-check dataset, the second source of human articles.","marker":"[19]"},{"why":"Provides the chain-of-thought prompting method used to generate LLM fact-checking articles.","marker":"[20]"},{"why":"Supplies the journalism information-assessment guideline used to construct the fact-checking prompts.","marker":"[21]"},{"why":"Provides the sentence- and aspect-level certainty measure used to score information style.","marker":"[31]"},{"why":"Provides the psycholinguistic lexicon used to measure cognitive-processing expressions.","marker":"[32]"},{"why":"Provides the hierarchical model used to score persuasive strategies such as credibility, evidence, and impact.","marker":"[39]"},{"why":"Provides the social-values framework used to measure receiver value alignment.","marker":"[40]"},{"why":"Provides the moral-foundation scoring model used to measure authority and fairness.","marker":"[42]"}],"fun_headline_variants":["Readers prefer LLM fact-checks despite lower style scores","LLM fact-checks win readers, lose on style metrics","Blind study: 60% prefer LLM explanations for health fact-checks","Health fact-checks: readers favor LLMs, not human style","LLM clarity beats human style in health fact-checking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The automated measures of certainty, persuasion, and value alignment really capture the dimensions of health communication that matter to readers, and the 250–350-word human articles used in the preference test fairly represent how human fact-checkers write.","fun_headline_variants_meta":{"raw":{"variants":["Readers prefer LLM fact-checks despite lower style scores","LLM fact-checks win readers, lose on style metrics","Blind study: 60% prefer LLM explanations for health fact-checks","Health fact-checks: readers favor LLMs, not human style","LLM clarity beats human style in health fact-checking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2786,"prompt_tokens":947,"completion_tokens":1839,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":1750}},"tokens_in":563,"tokens_out":1839,"duration_ms":12773,"temperature":1.0,"reasoning_tokens":1750,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:01:46.385791+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same blind paired-preference protocol using full-length human fact-checking articles rather than only the 250–350-word subset, and add a condition in which participants are told which article is AI-generated; if the LLM preference disappears or reverses under either change, the claimed disconnect between style scores and reader preference is an artifact of length matching or of perceived rather than actual quality.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the FakeCovid cross-domain fact-check dataset, the second source of human articles."},{"cited_title":"Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual Requests","cited_arxiv_id":"2101.06351","evidence_quote":"Provides the hierarchical model used to score persuasive strategies such as credibility, evidence, and impact."},{"cited_title":"Basic human values: Theory, measurement, and applications.Revue Francaise de Sociol.47, 929– 968+977+981 (2006)","cited_arxiv_id":null,"evidence_quote":"Provides the social-values framework used to measure receiver value alignment."}],"review_version":1}