REVIEW 3 major objections 5 minor 51 references
Personalized Large Language Models Can Increase the Belief Accuracy of Social Networks
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A personalized, accuracy-guarded LLM inside a social network shifted members' beliefs toward the truth and steered their follow choices toward accurate peers, in a pre-registered experiment of 1,265 people around the 2024 US election.
desk verdict Novel and worth taking seriously, but the batch-clustering analysis and the missing non-personalized arm need to be fixed before the causal claims are clean. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the personalized bot pipeline: a traditional machine-learning model converts each user's demographics and Big Five personality ratings into a predicted preferred news source and rhetorical style (ethos, pathos, or logos); a retrieval system over roughly 70,000 full-length articles from ten ideologically balanced outlets (three right, three center, four left) gathers evidence for or against the statement; and GPT-4o-mini summarizes the evidence and reframes it in the predicted style, with prompts that forbid hate speech. The measure that carries the network claim is the "follow signal": each participant's average factual-accuracy score (0 = most accurate to 4 = least accurate) of the human peers they chose to follow each round, with the single most-accurate entry removed so that the bot's mere presence cannot inflate the score. The protocol — initial belief rating with rationale, exposure to peer responses (plus the bot in treatment), then a three-of-six follow choice, repeated over three rounds in batches of 10 to 28 participants — is what lets the paper observe belief revision and network rewiring in the same session.
What would settle it
Re-run the main comparisons with standard errors clustered by experimental session: if the treatment-control differences in belief shift or follow signal lose significance, the central claims fail. Separately, if a bot that confidently states falsehoods attracts followers and pulls follow-networks toward its own inaccurate positions as strongly as the accurate bot did, then the mechanism is similarity-seeking toward whatever the bot says, not a pull toward the truth.
Extended reading notes
Core claim
The paper reports three pre-registered findings from a three-round networked experiment. First, hypothesis 1: compared with controls who saw three peer responses, participants who also saw a personalized bot response shifted their belief ratings toward the objectively correct answer (average signed shift of approximately $-0.35$ in the treatment versus approximately $-0.1$ in the control, $p < 0.0001$; negative means toward the truth), a difference that held across statement factuality, topic salience, economic relevance, media skepticism, demographics, and pre- versus post-election timing. Second, hypothesis 2a: 70.2% of treatment participants followed the bot at least once, and 54.4% of those kept following it across all three rounds. Third, hypothesis 2b: even after excluding the single most accurate option, treatment participants' "follow signal" — the mean factual-accuracy score of the people they chose to follow, on a 0-to-4 scale — was closer to the truth (1.55 versus 2.02, $p < 0.0001$), and their networks became more assortative, meaning they did not choose followers at random. The authors conclude that a truthful, personalized agent can function both as a direct informant and as a catalyst that reconfigures the surrounding network toward accuracy.
Load-bearing premise
The analysis treats each participant as an independent observation, but participants ran in synchronous batches of 10 to 28 people who saw each other's answers and influenced one another's follow choices; if outcomes are correlated within batches, the t-tests and regressions, computed without cluster-robust standard errors, overstate significance.
Editorial extensions
If this is right
- Deploying accurate, personalized LLM agents inside real online communities could shift not just what individuals believe but also whom they choose to follow, nudging local information environments toward verified claims.
- The belief-shift effect held across statement factuality, topic salience, economic relevance, self-reported media skepticism, demographic groups, and pre- versus post-election timing, indicating the corrective effect is not limited to a single issue or audience.
- Belief shifts occurred without back-and-forth dialogue and despite partial awareness of the bot, so brief and partly transparent AI messages may be enough to move beliefs in real-world settings.
- The authors' risk analysis: the same pipeline with the accuracy guardrails replaced by a malicious agenda could amplify misinformation and reshape networks around false beliefs, marking where oversight is needed.
Reading between the lines
- Because the follow-signal difference survived dropping the single most-accurate entry, the bot appears to have changed the criterion people used for choosing social ties, not merely added one accurate node; a direct test would remove the bot in a fourth round and check whether follow choices stay accurate.
- Treatment participants' updated, truthward answers were shown to other people in later rounds, so the bot's influence may spill over to participants who never saw it; the paper does not measure this, though its design could.
- Control participants also preferentially followed a "most truthful" human when one was identifiable, so a probe that replaces the bot with an equally accurate, well-written human argument would reveal whether AI identity adds anything beyond making an accurate opinion salient.
- Editorial notes on the manuscript text: the means in main-text Table 1 appear swapped relative to the prose and the supplementary tables, and one citation marker in the hypotheses section is empty ('()'); the directional claims rest on the prose and supplementary numbers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-registered online experiment (N=1265) conducted in synchronous batches during the 2024 U.S. presidential election period. Participants rated the veracity of political statements, then revised their ratings after seeing peer responses, and in the treatment condition, a personalized LLM-generated response tailored to their profile. In the third stage, participants selected peers to follow, with selections influencing later rounds. The authors report that treatment-group participants shifted their beliefs toward factual accuracy significantly more than controls, followed the bot at high rates, and constructed more accurate follow networks. The paper concludes that personalized LLMs can improve individual and network-level belief accuracy.
Significance. If the findings hold, this is a valuable contribution to the literature on LLM persuasion and network effects: it is one of the first studies to embed an LLM inside a dynamic social network and measure both belief updating and network construction. The study has notable strengths, including a large sample, pre-registration, and extensive robustness checks across demographics, statement types, and time periods. However, the statistical analysis ignores the batch structure of the data, and the main results section contains a table that contradicts the text. The design also confounds personalization with the mere presence of an accurate LLM agent. These issues preclude using the paper's current evidence to support the headline claims, though they are potentially addressable.
major comments (3)
- [Materials and Methods (SM), Results, Tables 1–2] The analysis treats each of the N=1265 participants as an independent observation, but the experiment was run in synchronous batches of 10–28 participants (median 16) with three rounds, and participants saw each other's responses and influenced which peers appeared in later rounds through their follow choices. This creates within-batch correlation in both the belief-shift and follow-signal outcomes. The reported t-tests and confidence intervals (e.g., Tables 1 and 2; Figure 2A-B and 3B) do not account for this clustering with cluster-robust standard errors, a multilevel model, or a batch-level analysis. If outcomes are correlated within batches, the effective sample size is the number of batches rather than 1265, and the reported p-values (often <0.0001) overstate the strength of the evidence. Because the paper's central claims rest on these tests, the authors should re-estimate the effects with appropriate clustering and report cluster-robust confidence intervals and p-values.
- [Table 1, Results] Table 1 as printed labels the Treatment mean as -0.102 and the Control mean as -0.354, but the text states the opposite ('an average shift of -0.35 in the treatment and -0.1 in the control') and Figure 2 shows the treatment distribution shifted more negative (toward truth). As printed, the table contradicts the central claim that the treatment moved beliefs toward accuracy. This is not a minor typo because the table is the primary quantitative evidence for Hypothesis 1; the authors must correct the labels or the values and verify the corrected table against the underlying data and SM Tables (e.g., Table S13).
- [Abstract, Results (H1), SM Section 2.2] The title and abstract attribute the effect to 'personalized' LLMs, and Hypothesis 1 is framed around personalized LLM messages. However, the experiment only contrasts a personalized-bot condition with a no-bot control; it does not include a condition with a non-personalized bot or with a bot delivering the same accurate content without tailoring. Consequently, the design cannot separate the effect of personalization from the effect of having an additional accurate information source in the network. The authors should either add such a condition in future work or, for the present paper, reframe the claims to state that an accurate, personalized LLM agent, relative to no agent, improves belief accuracy. Without this change, the personalization-specific conclusion is not supported.
minor comments (5)
- [Figure 2B caption] The caption states 'The intervals are non-overlapping, implying significance'; non-overlap of 95% confidence intervals is a conservative heuristic, not a significance test, and overlapping intervals do not imply non-significance. Report the actual t-statistics and p-values.
- [Table S12] In Table S12, the Control 'Post-Election, All Conditions' mean is listed as -0.91, which appears to be a typo (likely -0.091); please correct and verify all SM tables for similar errors.
- [SM Sections 3.3 and 3.4] The text refers to 'Strata' where it means the statistical software 'Stata'.
- [Main text, Discussion] The discussion should acknowledge the SM Section 3.9 limitation that the exact causal actor (bot vs. peer) cannot be isolated, as the authors themselves note in SM, to help readers correctly interpret the follow-signal results.
- [Materials and Methods] The main text should state the batch-based, interactive design (10–28 participants per session, three rounds) in the Methods rather than only in the SM, since this design feature is central to the statistical analysis.
Circularity Check
No significant circularity: the belief-shift and follow-signal results are measured participant outcomes, not constructions from the bot parameters.
full rationale
This paper is an empirical intervention study rather than a derivation chain, so the circularity patterns do not apply. The belief-shift outcome is computed from participants' own initial and updated ratings, with the sign set by the externally labeled veracity of each statement; the personalized-LLM message is generated from user profiles plus factual sources, but those personalization predictors are fitted in a separate pre-processing stage and do not enter the outcome equations. The follow-signal outcome is measured from participants' actual follow choices, and the metric deliberately excludes the bot or the most-truthful peer (SM Section 1, Equation S1), so the treatment effect is not guaranteed by construction. The one self-citation (ref 35, used for demographic response categories) is not load-bearing for any headline result. The batch-level clustering concern about synchronous participant batches is a statistical robustness issue, not a definitional or self-referential reduction.
Assumptions & free parameters
assumptions (4)
- domain assumption A five-point belief rating is an interval scale and can be averaged across different statements and rounds into a single belief-shift score.
- domain assumption Participant outcomes within a batch are independent.
- domain assumption Bot-generated messages are factually accurate.
- domain assumption Dropping the most truthful followed person from the follow signal is a fair comparison.
Cite this review
Pith. "Pith review of Personalized Large Language Models Can Increase the Belief Accuracy of Social Networks." pith.science (2026). https://pith.science/paper/RBHHJONR
@misc{pith2026250606153,
author = {Pith},
title = {Pith review of: Personalized Large Language Models Can Increase the Belief Accuracy of Social Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/RBHHJONR}},
note = {Machine review of arXiv:2506.06153}
}
read the original abstract
Large language models (LLMs) are increasingly involved in shaping public understanding on contested issues. This has led to substantial discussion about the potential of LLMs to reinforce or correct misperceptions. While existing literature documents the impact of LLMs on individuals' beliefs, limited work explores how LLMs affect social networks. We address this gap with a pre-registered experiment (N = 1265) around the 2024 US presidential election, where we empirically explore the impact of personalized LLMs on belief accuracy in the context of social networks. The LLMs are constructed to be personalized, offering messages tailored to individuals' profiles, and to have guardrails for accurate information retrieval. We find that the presence of a personalized LLM leads individuals to update their beliefs towards the truth. More importantly, individuals with a personalized LLM in their social network not only choose to follow it, indicating they would like to obtain information from it in subsequent interactions, but also construct subsequent social networks to include other individuals with beliefs similar to the LLM -- in this case, more accurate beliefs. Therefore, our results show that LLMs have the capacity to influence individual beliefs and the social networks in which people exist, and highlight the potential of LLMs to act as corrective agents in online environments. Our findings can inform future strategies for responsible AI-mediated communication.
Reference graph
Works this paper leans on
-
[1]
Nyhan, Facts and myths about misperceptions
B. Nyhan, Facts and myths about misperceptions. Journal of Economic Perspectives 34 (3), 220–236 (2020)
work page 2020
-
[2]
H. Matatov, M. Naaman, O. Amir, Stop the [Image] steal: The role and dynamics of visual content in the 2020 US Election Misinformation Campaign. Proceedings of the ACM on Human-Computer Interaction 6 (CSCW2), 1–24 (2022)
work page 2022
-
[3]
Oehmichen, et al., Not all lies are equal
A. Oehmichen, et al., Not all lies are equal. A study into the engineering of political misinfor- mation in the 2016 US Presidential Election. IEEE access 7, 126305–126314 (2019)
work page 2019
- [4]
-
[5]
D. M. Lazer, et al., The science of fake news. Science 359 (6380), 1094–1096 (2018)
work page 2018
-
[6]
J. A. Tucker, et al., Social media, political polarization, and political disinformation: A review of the scientific literature. Political polarization, and political disinformation: a review of the scientific literature (March 19, 2018)(2018)
work page 2018
-
[7]
Eggertson, Lancet retracts 12-year-old article linking autism to MMR vaccines
L. Eggertson, Lancet retracts 12-year-old article linking autism to MMR vaccines. Canadian Medical Association Journal 182 (4), E199–E200 (2010), doi:10.1503/cmaj.109-3179, http: //www.cmaj.ca/cgi/doi/10.1503/cmaj.109-3179
-
[8]
I. Skafle, A. Nordahl-Hansen, D. S. Quintana, R. Wynn, E. Gabarron, Misinformation About COVID-19 Vaccines on Social Media: Rapid Review. Journal of Medical Internet Research 24 (8), e37367 (2022), doi:10.2196/37367, https://www.jmir.org/2022/8/e37367
doi:10.2196/37367 2022
Show all 51 references
-
[9]
Ognyanova, D
K. Ognyanova, D. Lazer, R. E. Robertson, C. Wilson, Misinformation in action: Fake news exposure is linked to lower trust in media, higher trust in government when your side is in power. Harvard Kennedy School Misinformation Review (2020), doi:10.37016/mr-2020-024, https://mis...
2020 doi
-
[10]
Huang, L
Y. Huang, L. Sun, FakeGPT: Fake News Generation, Explanation and Detection of Large Language Models. arXiv preprint arXiv:2310.05046 (2024)
2024 arXiv
-
[11]
C. Chen, K. Shu, Can LLM-Generated Misinformation Be Detected?, in The Twelfth Interna- tional Conference on Learning Representations (ICLR) (2024)
2024
-
[12]
Pan, et al., Fact-checking complex claims with program-guided reasoning
L. Pan, et al., Fact-checking complex claims with program-guided reasoning. arXiv preprint arXiv:2305.12744 (2023)
2023 arXiv
-
[13]
Wang, et al., Mmidr: Teaching large language model to interpret multimodal misinformation via knowledge distillation
L. Wang, et al., Mmidr: Teaching large language model to interpret multimodal misinformation via knowledge distillation. arXiv preprint arXiv:2403.14171 (2024)
2024 arXiv
-
[14]
J. Lucas, et al., Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elu- sive Disinformation, in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, K. Bali, Eds. (Association for Computational Lin...
2023 doi
-
[15]
T. H. Costello, G. Pennycook, D. G. Rand, Durably reducing conspiracy beliefs through dialogues with AI. Science 385 (6714), eadq1814 (2024)
2024
-
[16]
S. M. Breum, D. V. Egdal, V. G. Mortensen, A. G. Møller, L. M. Aiello, The persuasive power of large language models, in Proceedings of the International AAAI Conference on Web and Social Media, vol. 18 (2024), pp. 152–163
2024
-
[17]
Karinshak, S
E. Karinshak, S. X. Liu, J. S. Park, J. T. Hancock, Working with AI to persuade: Examining a large language model’s ability to generate pro-vaccination messages.Proceedings of the ACM on Human-Computer Interaction 7 (CSCW1), 1–29 (2023)
2023
-
[18]
Matz, et al., The potential of generative AI for personalized persuasion at scale
S. Matz, et al., The potential of generative AI for personalized persuasion at scale. Scientific Reports 14 (1), 4692 (2024)
2024
-
[19]
Simchon, M
A. Simchon, M. Edwards, S. Lewandowsky, The persuasive effects of political microtargeting in the age of generative artificial intelligence. PNAS nexus 3 (2), pgae035 (2024). 14
2024
-
[20]
When” not “If
J. D. Teeny, S. C. Matz, We need to understand “When” not “If” generative AI can enhance per- sonalized persuasion.Proceedings of the National Academy of Sciences121 (43), e2418005121 (2024)
2024
-
[21]
Pennycook, D
G. Pennycook, D. G. Rand, The Psychology of Fake News. Trends in Cognitive Sci- ences 25 (5), 388–402 (2021), doi:https://doi.org/10.1016/j.tics.2021.02.007, https://www. sciencedirect.com/science/article/pii/S1364661321000516
2021 doi
-
[22]
Bolsen, J
T. Bolsen, J. N. Druckman, F. L. Cook, The Influence of Partisan Motivated Reasoning on Public Opinion. Political Behavior 36 (2), 235–262 (2014), doi:10.1007/s11109-013-9238-0, http://link.springer.com/10.1007/s11109-013-9238-0
2014 doi
-
[23]
Ahmed, H
S. Ahmed, H. W. Tan, Personality and perspicacity: Role of personality traits and cognitive ability in political misinformation discernment and sharing behavior.Personality and Individual Differences 196, 111747 (2022), doi:10.1016/j.paid.2022.111747, https://linkinghub. elsev...
2022
-
[24]
S. Chen, L. Xiao, J. Mao, Persuasion strategies of misinformation-containing posts in the social media. Information Processing & Management 58 (5), 102665 (2021)
2021
-
[25]
Fazio, et al., Combating misinformation: A megastudy of nine interventions designed to reduce the sharing of and belief in false and misleading headlines (2024)
L. Fazio, et al., Combating misinformation: A megastudy of nine interventions designed to reduce the sharing of and belief in false and misleading headlines (2024)
2024
-
[26]
Porter, T
E. Porter, T. J. Wood, Factual corrections: Concerns and current evidence. Current Opinion in Psychology 55, 101715 (2024)
2024
-
[27]
Pennycook, et al., Shifting attention to accuracy can reduce misinformation online
G. Pennycook, et al., Shifting attention to accuracy can reduce misinformation online. Nature 592 (7855), 590–595 (2021)
2021
-
[28]
Sinclair, The social citizen: Peer networks and political behavior (University of Chicago Press) (2012)
B. Sinclair, The social citizen: Peer networks and political behavior (University of Chicago Press) (2012)
2012
-
[29]
R. M. Bond, R. K. Garrett, Engagement with fact-checked posts on Reddit. PNAS nexus 2 (3), pgad018 (2023). 15
2023
-
[30]
Sharma, H
M. Sharma, H. C. Siu, R. Paleja, J. D. Pe ˜na, Why Would You Suggest That? Human Trust in Language Model Responses. arXiv preprint arXiv:2406.02018 (2024)
2024 arXiv
-
[31]
M. R. DeVerna, H. Y. Yan, K.-C. Yang, F. Menczer, Fact-checking information from large language models can decrease headline discernment. Proceedings of the National Academy of Sciences 121 (50), e2322823121 (2024)
2024
-
[32]
Klar, Partisanship in a social setting
S. Klar, Partisanship in a social setting. American journal of political science 58 (3), 687–704 (2014)
2014
-
[33]
J. N. Druckman, M. S. Levendusky, A. McLain, No need to watch: How the effects of partisan media can spread via interpersonal discussions. American Journal of Political Science 62 (1), 99–112 (2018)
2018
-
[34]
J. N. Druckman, Experimental thinking (Cambridge University Press) (2022)
2022
-
[35]
A. M. Proma, et al., Exploring the role of randomization on belief rigidity in online social networks. arXiv preprint arXiv:2407.01820 (2024)
2024 arXiv
-
[36]
Rammstedt, O
B. Rammstedt, O. P. John, Measuring personality in one minute or less: A 10-item short version of the Big Five Inventory in English and German. Journal of research in Personality 41 (1), 203–212 (2007)
2007
-
[37]
Boos, et al
D. Boos, et al. , A comparison of tests of equality of variances (1999), https://www. sciencedirect.com/science/article/pii/0167947395000542
1999
-
[38]
E. D. Sun, et al., Spatial transcriptomic clocks reveal cell proximity effects in brain ageing (2024), https://www.nature.com/articles/s41586-024-08334-8
2024
-
[39]
Jurkowitz, A
M. Jurkowitz, A. Mitchell, E. Shearer, M. Walker, Democrats report much higher levels of trust in a number of news sources than Republicans. Pew Research Center(2020)
2020
-
[40]
Schaffner, S
B. Schaffner, S. Ansolabehere, S. Luks, Cooperative election study common content, 2020. Harvard Dataverse 1 (10.7910) (2021). 16
2021
-
[41]
Harvard Library, Research Guides: News Media Across the Political Spectrum: Starting Point:
-
[42]
”The Chart” — guides.library.harvard.edu, https://guides.library.harvard.edu/ newsleans/thechart (2025), [Accessed 02-13-2025]
2025
-
[43]
Ground News — ground.news, https://ground.news/, [Accessed 08-11-2024]
2024
-
[44]
Brenan, Economy most important issue to 2024 presidential vote (2024), https://news
M. Brenan, Economy most important issue to 2024 presidential vote (2024), https://news. gallup.com/poll/651719/economy-important-issue-2024-presidential-vote. aspx
2024
-
[45]
J. M. Jones, Immigration surges to top of most important problem list (2025),https://news. gallup.com/poll/611135/immigration-surges-top-important-problem-list. aspx
2025
-
[46]
S. A. Banducci, J. A. Karp, How elections change the way citizens view the political system: Campaigns, media effects and electoral outcomes in comparative perspective. British Journal of Political Science 33 (3), 443–467 (2003)
2003
-
[47]
C. G. Wilson, A. T. Nusbaum, P. Whitney, J. M. Hinson, Age-differences in cognitive flexibility when overcoming a preexisting bias through feedback. Journal of clinical and experimental neuropsychology 40 (6), 586–594 (2018)
2018
-
[48]
L. P. Argyle, et al. , Testing theories of political persuasion using AI. Proceedings of the National Academy of Sciences 122 (18), e2412815122 (2025), doi:10.1073/pnas.2412815122, https://pnas.org/doi/10.1073/pnas.2412815122
2025 doi
-
[49]
A. Luttrell, Dual process models of persuasion(Oxford University Press) (2018), doi:10.1093/ acrefore/9780190236557.013.319, http://psychology.oxfordre.com/view/10.1093/ acrefore/9780190236557.001.0001/acrefore-9780190236557-e-319
2018
-
[50]
R. E. Petty, J. A. Krosnick, Attitude strength: Antecedents and consequences (Psychology Press) (2014)
2014
-
[51]
response entry
J. N. Druckman, K. R. Nelson, Framing and deliberation: How citizens’ conversations limit elite influence. American journal of political science 47 (4), 729–745 (2003). S1 Supplementary Materials for Personalized Large Language Models Can Increase the Belief Accuracy of Social...
2003
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.