REVIEW 3 major objections 5 minor 3 cited by
GPT-4o is as effective at increasing belief in conspiracy theories as at reducing them, and only a truth-prompt restores a truth advantage.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 11:49 UTC pith:WC5RZM2H
load-bearing objection A solid, preregistered demonstration of bunk/debunk symmetry in immediate self-reported belief, worth publishing after the authors fix an abstract that promises more than the full text delivers. the 3 major comments →
Large language models can effectively convince people to believe conspiracies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that GPT-4o, a frontier large language model, is as effective at increasing belief in conspiracy theories as it is at reducing them when it is not explicitly constrained to be truthful. In Study 1, using a jailbreak-tuned variant, a pro-conspiracy conversation increased focal belief by 13.7 points and a debunking conversation decreased it by 12.1 points, with the difference nonsignificant (p = .22). Study 2 replicated this with standard GPT-4o (+11.9 vs -12.9, p = .47), showing that default safety guardrails did not prevent the model from promoting conspiracies. A truth-constrained prompt in Study 3 sharply reduced the bunking effect to about 4.8 points while leaving deb
What carries the argument
The experimental design is the core mechanism: participants select a conspiracy they genuinely feel uncertain about (operationalized as a baseline rating between 25 and 75 on a 0-100 slider), then are randomly assigned to a back-and-forth text conversation with GPT-4o prompted either to argue for ('bunking') or against ('debunking') that specific theory. The primary outcome is direction-aligned pre-to-post change on the same belief slider. Two supporting mechanisms make the interpretation possible: an automated 'attempt to persuade' evaluator verifies that the model actually complied with its assigned direction, and a claim-level fact-checking pipeline rates every factual statement for verac
Load-bearing premise
The paper treats a single self-report slider taken immediately after the chat as evidence of persuasion, so if those shifts reflect demand or short-lived acquiescence to a one-sided AI rather than durable internalized belief change, the central symmetry may not describe real-world persuasion.
What would settle it
Run the same paradigm with a neutral-chat control group (same AI, no stance) and a delayed belief measure weeks later: the central claim fails if bunking versus debunking differences are not significantly larger than the no-stance control or do not persist to the follow-up.
If this is right
- With default guardrails, GPT-4o is roughly as effective at increasing as at decreasing belief in a target conspiracy; the persuasive power does not inherently favor truth.
- Guardrails alone did not block conspiracy promotion: jailbroken and standard GPT-4o produced similar pro-conspiracy belief changes.
- Bunking-induced belief increases are reversible: a corrective conversation lowered belief below the participant's original baseline.
- A simple truth-constrained prompt reduced bunking effectiveness by roughly half to two-thirds while leaving debunking effectiveness unchanged, showing a practical path to favoring accurate beliefs.
- Bunking was experienced more positively than debunking—rated more informative, collaborative, and trust-building—which may make AI-spread misinformation harder to detect.
Where Pith is reading between the lines
- Because belief is measured once, immediately after the chat, the symmetry may describe short-run acquiescence to a one-sided AI rather than durable persuasion; a delayed follow-up would test this.
- If the symmetry generalizes beyond conspiracy theories, any domain where an LLM can be prompted to argue both sides—political, medical, scientific—faces the same dual-use risk, and the same truth-prompt fix deserves testing there.
- The persistence of bunking under a truth constraint suggests paltering—selecting and framing true facts to mislead—may be the harder failure mode to detect, since individual claims check out as accurate.
- The equivocal-window sampling means the results apply to uncertain people, not committed believers; the same interventions may behave differently in echo-chamber settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports three preregistered, between-subjects experiments (total N = 2,724 after exclusions) in which American participants who were uncertain about a conspiracy theory had a text conversation with GPT-4o instructed either to argue for ("bunking") or against ("debunking") the theory. Study 1 used a jailbroken GPT-4o variant; Study 2 used standard GPT-4o; Study 3 used standard GPT-4o with an explicit truth-constraint prompt. The primary outcome is pre-to-post change in a 0–100 self-reported belief slider. The authors report large, roughly symmetric belief shifts in Studies 1 and 2 (+13.7 vs −12.1, p = .22; +11.9 vs −12.9, p = .47), a corrective debrief that reverses bunking-induced increases, a truth prompt that sharply reduces bunking while preserving debunking, and higher subjective ratings of the bunking AI compared with the debunking AI. The paper concludes that, absent explicit truth-oriented design, LLM persuasive power is symmetric between true and false claims.
Significance. If the symmetry result holds, it is a substantive empirical contribution to the debate about whether LLM persuasion inherently favors true claims. The study is exemplary in transparency and execution: preregistered, randomized, large samples, public data and code, browseable conversation transcripts, baseline-adjusted models with HC3 robust standard errors, APE compliance checks, counterfactual-compliance sensitivity analyses, and claim-level automated fact-checking. The replication of symmetry with standard GPT-4o, and the selective reduction of bunking under a truth prompt, strengthen the causal narrative. The main caveat is that the outcome is an immediate self-report measure and no equivalence test or neutral-chat control is provided; the applied implications therefore depend on assumptions about demand characteristics and durability of belief change.
major comments (3)
- [§2.1, §2.3] The central claim that bunking and debunking are "as effective" is inferred from non-significant p-values for the difference (z = 1.26, p = .22 in Study 1; p = .47 in Study 2). A non-significant difference is not evidence of equivalence. No confidence interval for the bunking-debunking difference, and no pre-specified equivalence margin or TOST test, is reported. Given N ≈ 1,000 per study, the difference may be imprecisely estimated; the data could be consistent with a meaningful truth advantage or disadvantage. Please report the CI for the difference and/or an equivalence test with a pre-specified margin, and hedge the symmetry claim accordingly.
- [§4.2.5, §2.1] The primary outcome is a single 0–100 self-report slider administered immediately after a one-sided AI conversation. There is no neutral-chat or no-conversation control condition, so the absolute "persuasion" effect cannot be separated from a general tendency to agree with an interactive AI or from experimenter demand. The symmetric shifts in the two active conditions are exactly the pattern expected under acquiescence. The debrief disclosure occurs after the primary post-conversation measure, so it is not the direct problem, but participants know the AI is arguing a position. The paper's prior paradigm (ref [1]) included a 2-month follow-up; here no delayed measurement is reported. The title and abstract ("convince people to believe") therefore overstate what was established. At minimum, restrict the claims to immediate self-reported belief change and explicitly discuss demand/acquiesce
- [Abstract] The abstract block preceding the main text describes four experiments, N = 3,996, GPT 5.2, and social-media sharing results, but the full text reports three experiments, N = 2,724, GPT-4o, and no social-media measure. This is a serious inconsistency that must be resolved before publication; it is unclear which findings are current. All mentions of study count, sample size, and model names should be made consistent throughout the manuscript.
minor comments (5)
- [§4.6.1] The automated fact-checking pipeline description is duplicated verbatim over two consecutive paragraphs; remove the duplicate.
- [§4.2.6] Typo in the debrief-chat prompt: "rebbutting" should be "rebutting."
- [Title/Abstract] The results are based on a single model family (GPT-4o). The title's "large language models" overgeneralizes; consider specifying the model family or acknowledge the scope in the title/abstract.
- [§2.1, Figure 2B] The distributional asymmetry is important: debunking produced very large shifts (≥40 points) twice as often as bunking (16% vs 8%). This nuance is mentioned in text but absent from the abstract, which states symmetry unconditionally. A brief qualifier would improve precision.
- [§4.4] DBSCAN parameters (eps = 3.6, minPts = 25) are reasonable but appear to be chosen by the authors; please state whether these were fixed a priori or selected post hoc, and describe the sensitivity to these choices.
Circularity Check
No significant circularity: the central claims rest on fresh randomized experimental measurements, not on fitted inputs or self-citation chains.
full rationale
The paper's central estimates—bunking increasing focal conspiracy belief by ~12–14 points and debunking decreasing it by ~12–13 points—are directly observed pre-to-post changes from randomized between-subjects experiments. There is no model fitting to the outcome followed by a 'prediction' of the same outcome, no normalization that forces the symmetry conclusion, and no uniqueness theorem imported from the authors' prior work to justify the design. The dialogue paradigm, APE compliance evaluator, and Perplexity fact-checking pipeline originate in prior or affiliated work, but they are used as measurement/manipulation checks rather than as inputs that mathematically determine the belief-change results; the APE evaluator is reported to have been validated against human judgment (84% agreement), and the fact-checking approach is likewise an external evaluator. Self-citations to prior dialogue-based debunking studies provide context and replication, but the new bunking effects are measured afresh. The paper even acknowledges the need for further research on whether debunking has an advantage for lasting belief updates, which further indicates the claims are not presented as forced by prior results. The main limitations—single immediate self-report slider, no neutral-chat control, debrief disclosure timing—are construct-validity concerns, not circularity in the derivation sense. No load-bearing step reduces by construction to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (3)
- Equivocal-window bounds on baseline belief =
25 < pre-belief < 75
- DBSCAN topic-clustering parameters =
eps = 3.6, minPts = 25
- Low-veracity claim cutoff =
40/100
axioms (5)
- domain assumption The 0-100 slider response immediately after the conversation measures genuine belief, not acquiescence or demand.
- domain assumption Participants' self-chosen 'uncertain' conspiracies, screened by the LLM equivocality classifier and the 25-75 baseline window, are a valid proxy for epistemically low-credence claims.
- domain assumption The APE LLM evaluator correctly identifies persuasion attempts (84% human agreement).
- domain assumption Perplexity Sonar's automated 0-100 veracity scores validly measure factual accuracy of claims.
- domain assumption Belief change measured minutes after the chat would persist and is not merely momentary.
read the original abstract
Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against ("debunking") or for ("bunking") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.
Figures
Forward citations
Cited by 3 Pith papers
-
A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
PERSUASIONTRACE introduces a Bayesian-network simulated target for multi-turn persuasion that matches human belief dynamics (81 vs 80) better than LLM baselines (64) and enables process-level evaluation.
-
On the Effectiveness of Fact Checking Information from Politically Congruent and Incongruent Large Language Models
LLM fact-checkers shift trust in political headlines across partisan lines, with perceived chatbot politics mattering only for politically distant true headlines.
-
Spontaneous Persuasion: An Audit of Model Persuasiveness in Everyday Conversations
LLMs engage in spontaneous persuasion in virtually all multi-turn conversations by favoring information-based strategies like logic and evidence, in contrast to human responses that rely more on social influence and n...
Reference graph
Works this paper leans on
-
[1]
H., Pennycook, G
Costello, T. H., Pennycook, G. & Rand, D. G. Durably reducing conspiracy beliefs through dialogues with AI.Science385, eadq1814 (2024)
2024
-
[2]
Lin, H. et al. Persuading Voters using Human-AI Dialogues.Nature(2025)
2025
-
[3]
J., Smith, A
Hornsey, M. J., Smith, A. E., Pearson, S., Bretter, C. & Nylund, J. L. Using conversational AI to reduce science skepticism.Curr. Opin. Psychol.67, 102216 (2026)
2026
-
[4]
Murphy, B. et al. Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility. Preprint athttps://doi.org/10.48550/arXiv.2507.11630(2025)
-
[5]
OpenAI et al. GPT-4 Technical Report. Preprint athttps://doi.org/10.48550/arXiv. 2303.08774(2024)
-
[6]
Jones, C. R. & Bergen, B. K. Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models. Preprint athttps://doi.org/10. 48550/arXiv.2412.17128(2024)
-
[7]
H., Spinoza-Martín, D., Rand, D
Boissin, E., Costello, T. H., Spinoza-Martín, D., Rand, D. G. & Pennycook, G. Dialogues with large language models reduce conspiracy beliefs even when the AI is perceived as human. PNAS Nexus4, (2025)
2025
-
[8]
Bretter, C. et al. Mapping, understanding and reducing belief in misinformation about electric vehicles.Nat. Energy10, 869–879 (2025). 22
2025
-
[9]
Czarnek, G. et al. Addressing climate change skepticism and inaction using human-AI dia- logues. Preprint athttps://doi.org/10.31234/osf.io/mqcwj_v1(2025)
-
[10]
Hou, Z. et al. A vaccine chatbot intervention for parents to improve HPV vaccination uptake among middle school girls: a cluster randomized trial.Nat. Med.31, 1855–1862 (2025)
2025
-
[11]
Hackenburg, K. et al. The Levers of Political Persuasion with Conversational AI. Preprint at https://doi.org/10.48550/arXiv.2507.13919(2025)
-
[12]
Costello, T. H., Pennycook, G. & Rand, D. Just the facts: How dialogues with AI reduce conspiracy beliefs. Preprint athttps://doi.org/10.31234/osf.io/h7n8u_v1(2025)
-
[13]
& Evans, J
Farrell, H., Gopnik, A., Shalizi, C. & Evans, J. Large AI models are cultural and social technologies.Science387, 1153–1156 (2025)
2025
-
[14]
Hornsey, M. J. et al. The promise and limitations of using GenAI to reduce climate scepticism. Nat. Clim. Change15, 1183–1189 (2025)
2025
-
[15]
Kowal, M. et al. It’s the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics. Preprint athttps://doi.org/10.48550/arXiv.2506.02873 (2025)
-
[16]
Q.47, 1–15(1983)
DAVISON,W.P.TheThird-PersonEffectinCommunication.Public Opin. Q.47, 1–15(1983)
1983
-
[17]
The Argumentative Theory: Predictions and Empirical Evidence.Trends Cogn
Mercier, H. The Argumentative Theory: Predictions and Empirical Evidence.Trends Cogn. Sci.20, 689–700 (2016)
2016
-
[18]
Carpenter, C. J. A Meta-Analysis of the Elm’s Argument Quality×Processing Type Predic- tions.Hum. Commun. Res.41, 501–534 (2015)
2015
-
[19]
(Prince- ton University Press, 2020)
Mercier, H.Not Born Yesterday: The Science of Who We Trust and What We Believe. (Prince- ton University Press, 2020). doi:10.1515/9780691198842
-
[20]
Simon, F. M. & Altay, S. Don’t Panic (Yet): Assessing the Evidence and Discourse Around Generative AI and Elections
-
[21]
How people are using ChatGPT.https://openai.com/index/ how-people-are-using-chatgpt/(2025)
2025
-
[22]
Summerfield, C. et al. The impact of advanced AI systems on democracy.Nat. Hum. Behav. 1–11 (2025) doi:10.1038/s41562-025-02309-z
-
[23]
Goldstein, J. A. et al. Generative Language Models and Automated Influence Operations: EmergingThreatsandPotentialMitigations.Preprintathttps://doi.org/10.48550/arXiv. 2301.04246(2023). 23
-
[24]
Dentith, M. R. X. Conspiracy theories on the basis of the evidence.Synthese196, 2243–2261 (2019)
2019
-
[25]
M., Costello, T
Bowes, S. M., Costello, T. H., Ma, W. & Lilienfeld, S. O. Looking under the tinfoil hat: Clarifying the personological and psychopathological correlates of conspiracy beliefs.J. Pers. 89, 422–436 (2021)
2021
-
[26]
Brotherton, R., French, C. C. & Pickering, A. D. Measuring Belief in Conspiracy Theories: The Generic Conspiracist Beliefs Scale.Front. Psychol.4, 279 (2013)
2013
-
[27]
& Aral, S
Vosoughi, S., Roy, D. & Aral, S. The spread of true and false news online.Science359, 1146–1151 (2018)
2018
-
[28]
& Baldi, P
Itti, L. & Baldi, P. Bayesian surprise attracts human attention.Vision Res.49, 1295–1306 (2009)
2009
-
[29]
Rogers, T., Zeckhauser, R., Gino, F., Norton, M. I. & Schweitzer, M. E. Artful paltering: The risks and rewards of using truthful statements to mislead others.J. Pers. Soc. Psychol.112, 456–473 (2017)
2017
-
[30]
S., Goldstein, S., O’Gara, A., Chen, M
Park, P. S., Goldstein, S., O’Gara, A., Chen, M. & Hendrycks, D. AI deception: A survey of examples, risks, and potential solutions.Patterns5, (2024)
2024
-
[31]
Schoen, B. et al. Stress Testing Deliberative Alignment for Anti-Scheming Training. Preprint athttps://doi.org/10.48550/arXiv.2509.15541(2025)
-
[32]
JFK assassination,
Lin, W. Agnostic notes on regression adjustments to experimental data: Reexamining Freed- man’s critique.Ann. Appl. Stat.7, 295–318 (2013). 24 Supplementary Information Supplementary Figures Supplementary Figure S1.Exclusion and attrition pipeline for Study 1. 25 Supplementary Figure S2.Exclusion and attrition pipeline for Study 2. 26 Supplementary Figure...
2013
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.