{"id":"d2613929-6f7a-4c08-aaf1-327e5f7b12f0","arxiv_id":"2507.13919","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Conversational AI is most persuasive when post-trained for persuasion and prompted to flood the dialogue with claims, and those persuasion gains come with measurably lower factual accuracy.","lead":"In three large online experiments with about 77,000 UK adults, the authors measured how well 19 chatbots could shift political opinions during live conversations. The results suggest that persuasion comes mainly from post-training and information-dense prompting, not from model size or personalization, and that these same methods reduce factual accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fact-checker validation covers only one model; between-condition accuracy differences may reflect judge bias rather than true accuracy, threatening the persuasion–accuracy tradeoff claim.","rationale":"The paper's central contribution is two-part: (i) post-training and prompting are larger levers than scale or personalization, and (ii) these gains come at a cost to factual accuracy. Part (i) is supported by randomized comparisons (prompt, RM) and a careful scaling meta-regression; its main threats (confounded scale in developer models, selection in maximal-persuasion estimate) do not overturn the direction. Part (ii) is entirely dependent on the LLM fact-checker, and the validation is insufficient for between-condition inference. The validation sample is 198 messages from one model; the paper uses the judge to compare dozens of models and eight prompts. Correlation alone cannot exclude differential bias. This is the load-bearing weakness because the abstract's 'strikingly' claim and the policy-relevant 'truthfulness tradeoff' would be false if the judge is biased. Alternative concerns (e.g., maximal-persuasion selection bias, wide CI on information-density correlation) affect secondary magnitudes, not the main conclusion. Thus the reader's identified assumption is correct and is the primary reason to keep the verdict conditional pending additional validation.","tokens_in":19107,"tokens_out":4892,"duration_ms":56262,"concrete_test":"Sample 50 messages from each of the key comparison cells in Figure 4 (e.g., GPT-4o 3/25 and 8/24 with information vs. other prompts; Llama-8B base vs. RM; GPT-4.5; Grok-3). Have professional fact-checkers rate all claims. Compute the mean human-minus-judge bias within each cell. If bias varies across cells by more than 5pp, apply cell-level bias corrections and re-estimate the accuracy differences; if the persuasion–accuracy tradeoff in Figure 4B/C disappears, the headline claim is not supported. If bias is constant or the tradeoff survives, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'where they increased AI persuasiveness they also systematically decreased factual accuracy' rests on accuracy ratings from a single LLM judge (gpt-4o-search-preview). The validation (Methods 1.8) compares this judge to professional human fact-checkers on 198 messages drawn from a single model (GPT-4o 8/24, Study 1 Chat 2), and reports only a correlation (r=0.84). A high correlation does not establish measurement invariance: the judge could be systematically more lenient or harsher for particular models, prompts, or message styles while still correlating with humans. The paper's accuracy comparisons are differences between conditions: information prompt vs. other prompts, GPT-4o 3/25 vs. 8/24, RM vs. base. If the judge rates claim-dense, information-prompted messages more harshly (or rates outputs from newer GPT-4o versions more leniently), the observed 'tradeoff' could emerge even if true accuracy is unchanged. The validation sample is too small and too narrow to rule out such differential bias. Since the abstract's 'strikingly' clause depends on this tradeoff, the accuracy subclaim is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses three large-scale randomized survey experiments (N = 76,977) in the UK to measure the persuasiveness of 19 conversational LLMs across 707 political issues, manipulating model scale (via effective compute), post-training (supervised fine-tuning and reward modeling), prompting strategy (eight rhetorical prompts, including an information-focused prompt), and personalization. Persuasion is measured as the difference in post-conversation attitude agreement relative to a no-conversation control, and factual accuracy is measured by an LLM-based fact-checking pipeline that rated 466,769 extracted claims on a 0–100 veracity scale. The authors report that scale has a positive but modest association with persuasiveness among uniformly chat-tuned models; reward-model post-training and information-dense prompting produce the largest persuasion gains; personalization has small effects; and the factors that increase persuasion also tend to increase information density and decrease the rated accuracy of claims. They extrapolate that a 'maximal-persuasion' AI condition would shift attitudes by 15.9pp overall and 26.5pp among initial disagreers, with nearly 30% of its claims rated inaccurate.","tokens_in":19247,"tokens_out":9114,"duration_ms":106967,"significance":"If the results hold, they would substantially revise current debates about AI persuasion: the primary levers would be developer post-training and prompting choices that encourage information-dense argumentation, rather than model scale or user personalization, and these levers would come at a measurable cost to factual accuracy. The evidentiary base is unusually strong for this literature: three pre-registered experiments with large samples, a range of open- and closed-source models, robustness checks including attrition imputation and a one-month durability follow-up, and publicly available code and materials. The main weakness is that the accuracy and information-density measurements rely on an LLM-based pipeline validated on only 198 messages from one model; because the headline tradeoff claim depends on cross-condition comparisons, this validation gap is the key threat to the paper's central conclusion.","major_comments":[{"comment":"The paper's two central measurement claims—the information-density mechanism and the persuasion–accuracy tradeoff—both depend on an LLM-based measurement pipeline (GPT-4o for claim extraction, gpt-4o-search-preview for fact-checking) whose validation is limited to 198 messages from a single model (GPT-4o 8/24, Study 1 Chat 2). The reported correlations (r = 0.87 for counts, r = 0.84 for accuracy) demonstrate overall agreement with human raters but do not establish measurement invariance across the models, prompts, and post-training conditions whose differences drive the paper's conclusions. A judge or extractor that is differentially lenient or strict for particular conditions (e.g., information-prompted, claim-dense messages or newer GPT-4o versions) could produce the observed pattern—higher information density, lower rated accuracy, and higher persuasion—even if the models' true characteristics were unchanged. Please extend validation to a stratified sample spanning each model family, prompt type, and post-training condition, and report within-stratum human–machine agreement; if this is infeasible, the abstract and Discussion should explicitly hedge the accuracy and mechanism claims.","section":"Methods §1.8, Results 'How do models persuade?' and 'How accurate is the information provided by the models?'"},{"comment":"The accuracy analyses use the per-conversation proportion of claims rated >50/100 as the dependent variable. Because the information prompt and reward modeling also increase the number of claims per conversation, a judge whose ratings are noisier or systematically lower for additional or more marginal claims would make the accuracy tradeoff appear larger than it is, even if per-claim truthfulness were unchanged. The paper should test the robustness of the Figure 4B–C contrasts to controlling for claim count (e.g., condition-level regressions of accuracy on persuasion levers that include total claims) or report the accuracy comparison restricted to the first k claims of each conversation. This is important because the headline claim that persuasion and accuracy are in systematic tension is stated without such a control.","section":"Results 'How accurate is the information provided by the models?', Figure 4"}],"minor_comments":[{"comment":"The fact-checking pipeline was implemented between April 1st and May 18th, 2025; if the gpt-4o-search-preview model version changed during that window, temporal drift could confound the across-model accuracy comparisons. Please report the exact model version(s) used and test for time trends in ratings.","section":"Methods §1.8 Fact-checking"},{"comment":"The reference to 'The Brms Book' as an early draft ([6]) is not a stable citation for a journal submission; please replace it with a published reference or the package documentation.","section":"References"},{"comment":"Figure 4A removes some model labels 'for clarity'; given that the accuracy findings for GPT-4.5 and GPT-3.5 are surprising, a full version with all labels (or a corresponding table in the main text) would aid the reader.","section":"Figure 4A"},{"comment":"The maximal-persuasion estimate (15.9pp overall; 26.5pp among initial disagreers) is reported as an observed mean of the 500 conversations predicted to be most persuasive by a cross-fit random forest, but no uncertainty interval or correction for selection is provided; please report a bootstrap CI or state that this is an exploratory upper-bound estimate.","section":"Results 'maximal-persuasion' analysis"},{"comment":"The persistence analysis (36–42% of the effect after one month) is only conducted in Study 1 with one model; consider reporting whether this differs across conditions or noting the limitation.","section":"Results 'durability' analysis"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to attract wide attention and the experimental program is impressive, but the central 'systematic accuracy tradeoff' claim currently rests on a validation sample too narrow to support cross-model and cross-prompt inferences. If the authors can add stratified human validation or materially soften the claim, the paper would be suitable for publication; the current version is not yet there."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: the paper's central claim is likely right, but the accuracy subclaim rests on a weaker measurement foundation than the persuasion results. It deserves a serious referee and will be heavily cited.\n\nWhat's actually new: the combination of scale (17 base models spanning four orders of magnitude), post-training (SFT, RM), eight prompts, personalization, and fact-checking of 466k claims in one design, with pre-registration, attrition imputation, and a durability follow-up. The persuasion reward model is a genuinely new method and the RM-vs-scale comparison (Llama-8B + RM matching GPT-4o) is the kind of result people will chew on. The information-density mechanism is well supported by the prompt-level analysis, even if the meta-analytic slope has a wide CI.\n\nThe soft spots are in proportion. The stress-test note is correct: the fact-checking validation used 198 messages from GPT-4o 8/24 only, and a high human-machine correlation on that sample does not establish that the judge rates claims from other models, prompts, and post-training conditions with equal validity. Since the headline \"where they increased persuasion they decreased accuracy\" compares conditions, differential judge leniency could partly produce the pattern. This is a measurement-invariance problem, not a fatal one: the real model differences could still be there, but the current validation does not rule out the artifact. The maximal-persuasion estimate (selecting top 500 predicted conversations then reporting observed effects) is also optimistically biased; the authors flag it as a cross-fit ML estimate, but the reporting of observed effect makes it look more like a direct measurement than it is. Those are the two real issues.\n\nI'd push back mildly on the reader's circularity concern: the RM was trained on prior conversations and evaluated on new participants, which is legitimately external. The fact-checker being the same model family is a concern only through the measurement-invariance lens above, not as a circularity-by-construction.\n\nThe citation pattern looks fine; self-citations are to their own prior single-message work and are appropriate.\n\nBottom line: send it to review. Ask for per-model or at least per-condition fact-checker validation, or a clear statement of the limitation; and an unbiased estimator or explicit selection-adjusted interpretation of the maximal-persuasion result. The persuasion-levers conclusion will survive those revisions; the accuracy-tradeoff claim might shrink a bit but probably not vanish.","headline":"A large, carefully built empirical study whose persuasion-levers conclusion holds up, while the accuracy-tradeoff headline needs better fact-checker validation before it can be taken at face value.","tokens_in":19870,"tokens_out":2610,"would_cite":true,"duration_ms":27954,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In three large experiments, the paper finds that AI persuasion is driven mainly by post-training and prompting that increase how much information a model deploys, and that these levers systematically reduce factual accuracy.","keywords":["AI persuasion","political attitudes","large language models","information density","reward modeling","factual accuracy","conversational AI","persuasion post-training"],"falsifier":"Ask professional fact-checkers, blind to condition, to rate a stratified sample of claims from the information-prompted and non-information conditions of GPT-4o (3/25), GPT-4.5, and the reward-model-tuned Llama models, and compare their ratings with the gpt-4o-search-preview ratings; if the human ratings do not reproduce the reported accuracy drops, such as 62% versus 78% for GPT-4o (3/25) or the 2.22pp accuracy drop from reward modeling, then the persuasion-accuracy trade-off is at least partly an artifact of the judge.","tokens_in":18840,"feed_emoji":"🗳️","tokens_out":9821,"duration_ms":107628,"temperature":0.7,"pith_summary":"Across three experiments with 76,977 U.K. participants, the authors had 19 large language models argue for a stance on one of 707 political issues, then measured attitude shifts and fact-checked 466,769 claims the models produced. They find that model scale shows a reliable positive association with persuasion when post-training is held constant, but that post-training and prompting are larger levers: reward-model post-training boosted persuasion by as much as 51%, and an information prompt that tells the model to deploy facts and evidence was 27% more persuasive than a basic prompt. The mechanism identified is information density: factors that raised the number of fact-checkable claims per conversation also raised persuasion, while the effect of personalization was consistently small, below 1pp. The paper also reports a consistent accuracy cost: the same factors that increased persuasion lowered the proportion of claims rated factually accurate, with the most persuasive conditions producing 29.7% inaccurate claims versus a 16.0% average. The authors conclude that the persuasive power of current and near-future AI will come mainly from post-training and prompting choices, not from scaling or user data.","feed_headline":"77,000-person study: AI persuades by claiming more, not by scaling","feed_subtitle":"Persuasion gains come from training tweaks and information-dense prompts—and accuracy falls.","key_machinery":"The central mechanism is information density, defined as the number of fact-checkable claims a model makes per conversation, extracted from 91,000 conversations by GPT-4o and validated against professional human fact-checkers with a correlation of r = 0.87 for claim counts. This measure carries the argument: across randomized conditions, information density explains roughly 44% of the variability in persuasive effects, and every major persuasion-boosting intervention examined, information prompting, reward-model post-training, and newer frontier post-training, also increases information density. The second key mechanism is the persuasion reward model, a GPT-4o fine-tuned on 56,283 conversations to predict belief change at each turn and used to select the best of 12 to 20 candidate replies; this is the post-training procedure that turns a small open model into a frontier-competitive persuader while increasing inaccurate claims. Claim accuracy is measured by a search-enabled GPT-4o judge validated against professional human fact-checkers at r = 0.84.","core_discovery":"On the paper's own terms, the central discovery is that conversational AI persuades by mobilizing large volumes of information, and that the methods which most increase this information flow also reduce factual accuracy. When post-training is held constant, larger models are reliably more persuasive, with an order-of-magnitude compute increase buying roughly 1.6 to 1.8 percentage points of persuasion, but this scaling return is eclipsed by post-training: a newer deployment of GPT-4o was 3.50pp more persuasive than an older deployment of the same model, which exceeds the predicted gain from a 10-fold or even 100-fold compute increase. Applying a reward model trained to pick persuasive replies raised a small open-source model, Llama3.1-8B, from about 6pp to about 9pp, matching or beating GPT-4o (8/24). Prompt-level analysis shows that information density, defined as the number of fact-checkable claims per conversation, predicts persuasiveness with a meta-analytic correlation of r = 0.76 and explains 44% of the variability in persuasive effects across randomized conditions. Accuracy analyses show that the same levers reduce factual accuracy: the information prompt dropped GPT-4o (3/25) accuracy from 78% to 62%, and reward modeling on chat-tuned models added 2.32pp of persuasion while reducing the proportion of accurate claims by 2.22pp.","pith_inferences":["Editorial extension: if the persuasion-accuracy trade-off is a stable property of optimizing for persuasion, then treating persuasiveness as an unregularized training target will keep pulling models toward inaccuracy; a testable fix is adding explicit accuracy penalties to the reward model and checking whether the persuasion gains survive.","Editorial extension: because the AI fact-checker was validated on only 198 messages from one model, a replication that validates the judge across GPT-4.5, Grok-3, and reward-model-tuned Llama outputs would show whether the reported accuracy decline is real or partly a measurement artifact.","Editorial extension: the paper leaves open whether the accuracy drop comes from generating more claims or from reward selection favoring confident-sounding falsehoods; an experiment holding information density constant while varying accuracy incentives would separate the two.","Editorial extension: the large controlled-condition effects may overstate real-world influence, since people outside a paid survey may not sustain long political conversations; voluntary-exposure field studies would test this bottleneck."],"forward_implications":["Fine-tuning a small open-source model with a persuasion reward model can make it as persuasive as a frontier model, so high-persuasion AI is accessible to actors who cannot afford frontier-scale compute.","Frontier persuasion gains are more likely to come from developer post-training updates than from scale; the paper estimates the gain from one GPT-4o post-training update exceeded predicted gains from a 10x or 100x compute increase.","The single most effective prompting strategy among the eight tested is telling the model to provide information, and it works by increasing the number of fact-checkable claims per conversation.","Optimizing AI for persuasion is associated with a systematic loss of factual accuracy, including in the most persuasive conditions where nearly 30% of claims were rated inaccurate.","Conversational AI is 41% to 52% more persuasive than a static AI message, so interactive deployment is the relevant risk surface for near-future persuasion."],"supporting_citations":[{"why":"Provides the scaling-law framing and the effective-compute measure used to test whether larger models are more persuasive.","marker":"[34]"},{"why":"Supplies the effective compute estimates for proprietary frontier models whose true scale is not public.","marker":"[18]"},{"why":"Provides the instruction-following and reward-model training approach on which the paper's persuasion post-training (SFT and RM) is based.","marker":"[44]"},{"why":"Supplies the Ultrachat conversations used for the uniform chat-tuning of the open-source base models.","marker":"[14]"},{"why":"Provides the theory and evidence that information-based argumentation persuades, the mechanism the paper identifies as the main driver.","marker":"[11]"},{"why":"Supports the claim that interactive conversation is a uniquely persuasive format, which the paper validates against static messages.","marker":"[1]"},{"why":"Gives a prior demonstration that AI dialogue can durably shift beliefs, used as a comparison for the persuasion magnitudes observed.","marker":"[12]"},{"why":"Provides a recent smaller-scale test of AI political persuasion, which the paper contrasts to justify studying conversational rather than static influence.","marker":"[2]"}],"fun_headline_variants":["AI persuades with info density, not model size","Persuasiveness grows with claims, accuracy doesn't","Post-training beats scaling for AI persuasion","Fact-dense prompts boost AI influence, hurt truth","77k study: AI persuasion from info, not compute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the AI fact-checker scores claims from every model, prompt, and post-training condition with the same accuracy standard; it was checked against human fact-checkers on only 198 messages from one deployment of GPT-4o, so a judge biased toward certain styles could make the accuracy drop look larger or smaller than it is.","fun_headline_variants_meta":{"raw":{"variants":["AI persuades with info density, not model size","Persuasiveness grows with claims, accuracy doesn't","Post-training beats scaling for AI persuasion","Fact-dense prompts boost AI influence, hurt truth","77k study: AI persuasion from info, not compute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":1978,"prompt_tokens":976,"completion_tokens":1002,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":942}},"tokens_in":592,"tokens_out":1002,"duration_ms":12983,"temperature":1.0,"reasoning_tokens":942,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:14:33.156915+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask professional fact-checkers, blind to condition, to rate a stratified sample of claims from the information-prompted and non-information conditions of GPT-4o (3/25), GPT-4.5, and the reward-model-tuned Llama models, and compare their ratings with the gpt-4o-search-preview ratings; if the human ratings do not reproduce the reported accuracy drops, such as 62% versus 78% for GPT-4o (3/25) or the 2.22pp accuracy drop from reward modeling, then the persuasion-accuracy trade-off is at least partly an artifact of the judge.","supporting_citations":[{"cited_title":"Scaling laws for neural language models.arXiv, 1 2020","cited_arxiv_id":null,"evidence_quote":"Provides the scaling-law framing and the effective-compute measure used to test whether larger models are more persuasive."},{"cited_title":"Key trends and figures in machine learning, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the effective compute estimates for proprietary frontier models whose true scale is not public."},{"cited_title":"Persuasion in Parallel","cited_arxiv_id":null,"evidence_quote":"Provides the theory and evidence that information-based argumentation persuades, the mechanism the paper identifies as the main driver."},{"cited_title":"Scaling up interactive argumentation by providing counterarguments with a chatbot.Nature Human Behaviour, 6(4):579–592, February 2022","cited_arxiv_id":null,"evidence_quote":"Supports the claim that interactive conversation is a uniquely persuasive format, which the paper validates against static messages."},{"cited_title":"Costello, Gordon Pennycook, and David Rand","cited_arxiv_id":null,"evidence_quote":"Gives a prior demonstration that AI dialogue can durably shift beliefs, used as a comparison for the persuasion magnitudes observed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a recent smaller-scale test of AI political persuasion, which the paper contrasts to justify studying conversational rather than static influence."}],"review_version":1}