{"id":"066e030c-b197-4a5e-8b41-85a293b9ce5f","arxiv_id":"2506.06153","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A pre-registered experiment finds that adding a personalized, factually grounded LLM to an online discussion moves individuals' beliefs toward the truth and makes them build more accurate social networks.","lead":"In a 1,265-person experiment around the 2024 US election, people who saw a personalized AI chatbot message alongside peer opinions updated their beliefs toward the factual answer more than people who saw only peers. The same people were also more likely to follow the bot and to follow other humans whose beliefs were closer to the truth.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Batch-level clustering is not accounted for in the statistical tests, so the reported p-values and confidence intervals may overstate the significance of the treatment effects.","rationale":"The reader's weakest assumption correctly identifies the batch-level clustering as the most load-bearing statistical threat. I agree because the central claims are about effects measured inside interacting networks, and the paper's own Materials and Methods describe exactly the kind of interdependence that makes independent-unit inference inappropriate. A quick back-of-the-envelope check reinforces the concern: with ~79 batches and a moderate intraclass correlation, the effective sample size drops substantially, yet the main text reports p < 0.0001 with no adjustment. This is not a mere robustness detail—it determines whether the headline effects are statistically credible. The Table 1 label swap is a serious manuscript error and must be corrected, but the SM consistently supports the intended direction, so it is more likely a typo than a substantive contradiction. The absence of a non-personalized LLM arm is a design limitation that narrows the interpretation of 'personalized' but does not invalidate the comparison as run. The clustering issue, by contrast, directly undermines the inferential basis for every main result, so it is the single most load-bearing concern. A concrete batch-level reanalysis would settle it, and until that is done the paper should remain conditional rather than fully acceptable.","tokens_in":25719,"tokens_out":9001,"duration_ms":98178,"concrete_test":"Re-run the main belief-shift and follow-signal comparisons with batch as the unit of analysis: compute each batch's mean outcome, compare treatment vs. control batches with a two-sample t-test or Mann-Whitney, and also fit linear mixed models with a random intercept for batch and cluster-robust standard errors. If the treatment effects remain significant at p < 0.05 and the adjusted confidence intervals exclude zero, the clustering concern is resolved; if not, the reported findings are overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Materials and Methods state that participants were run in synchronous batches of 10–28 (median 16) and that every experiment had three rounds; the design is explicitly network-mediated, with participants seeing each other's responses and influencing each other's follow choices. Yet all main analyses—the belief-shift t-test, the follow-signal t-test, and the confidence intervals in Tables 1 and 2—treat the N=1265 participants as independent observations. If condition was assigned at the batch level, or even if within-batch interactions induce correlated outcomes, the effective sample size is the number of batches (roughly 50–120), not 1265. The reported standard errors are therefore too small and the p-values too strong. This is load-bearing because the central claims—that a personalized LLM shifts beliefs toward accuracy and alters subsequent network composition—rest entirely on these two significance tests. The problem is compounded for the follow-signal outcome, where one participant's follow choices directly determine which peers appear in another participant's later rounds, creating mechanical within-batch dependence. Without cluster-robust standard errors, a batch-level analysis, or a multilevel model with batch as a random effect, the headline conclusions are not statistically supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pre-registered online experiment (N=1265) conducted in synchronous batches during the 2024 U.S. presidential election period. Participants rated the veracity of political statements, then revised their ratings after seeing peer responses, and in the treatment condition, a personalized LLM-generated response tailored to their profile. In the third stage, participants selected peers to follow, with selections influencing later rounds. The authors report that treatment-group participants shifted their beliefs toward factual accuracy significantly more than controls, followed the bot at high rates, and constructed more accurate follow networks. The paper concludes that personalized LLMs can improve individual and network-level belief accuracy.","tokens_in":25853,"tokens_out":6420,"duration_ms":59791,"significance":"If the findings hold, this is a valuable contribution to the literature on LLM persuasion and network effects: it is one of the first studies to embed an LLM inside a dynamic social network and measure both belief updating and network construction. The study has notable strengths, including a large sample, pre-registration, and extensive robustness checks across demographics, statement types, and time periods. However, the statistical analysis ignores the batch structure of the data, and the main results section contains a table that contradicts the text. The design also confounds personalization with the mere presence of an accurate LLM agent. These issues preclude using the paper's current evidence to support the headline claims, though they are potentially addressable.","major_comments":[{"comment":"The analysis treats each of the N=1265 participants as an independent observation, but the experiment was run in synchronous batches of 10–28 participants (median 16) with three rounds, and participants saw each other's responses and influenced which peers appeared in later rounds through their follow choices. This creates within-batch correlation in both the belief-shift and follow-signal outcomes. The reported t-tests and confidence intervals (e.g., Tables 1 and 2; Figure 2A-B and 3B) do not account for this clustering with cluster-robust standard errors, a multilevel model, or a batch-level analysis. If outcomes are correlated within batches, the effective sample size is the number of batches rather than 1265, and the reported p-values (often <0.0001) overstate the strength of the evidence. Because the paper's central claims rest on these tests, the authors should re-estimate the effects with appropriate clustering and report cluster-robust confidence intervals and p-values.","section":"Materials and Methods (SM), Results, Tables 1–2"},{"comment":"Table 1 as printed labels the Treatment mean as -0.102 and the Control mean as -0.354, but the text states the opposite ('an average shift of -0.35 in the treatment and -0.1 in the control') and Figure 2 shows the treatment distribution shifted more negative (toward truth). As printed, the table contradicts the central claim that the treatment moved beliefs toward accuracy. This is not a minor typo because the table is the primary quantitative evidence for Hypothesis 1; the authors must correct the labels or the values and verify the corrected table against the underlying data and SM Tables (e.g., Table S13).","section":"Table 1, Results"},{"comment":"The title and abstract attribute the effect to 'personalized' LLMs, and Hypothesis 1 is framed around personalized LLM messages. However, the experiment only contrasts a personalized-bot condition with a no-bot control; it does not include a condition with a non-personalized bot or with a bot delivering the same accurate content without tailoring. Consequently, the design cannot separate the effect of personalization from the effect of having an additional accurate information source in the network. The authors should either add such a condition in future work or, for the present paper, reframe the claims to state that an accurate, personalized LLM agent, relative to no agent, improves belief accuracy. Without this change, the personalization-specific conclusion is not supported.","section":"Abstract, Results (H1), SM Section 2.2"}],"minor_comments":[{"comment":"The caption states 'The intervals are non-overlapping, implying significance'; non-overlap of 95% confidence intervals is a conservative heuristic, not a significance test, and overlapping intervals do not imply non-significance. Report the actual t-statistics and p-values.","section":"Figure 2B caption"},{"comment":"In Table S12, the Control 'Post-Election, All Conditions' mean is listed as -0.91, which appears to be a typo (likely -0.091); please correct and verify all SM tables for similar errors.","section":"Table S12"},{"comment":"The text refers to 'Strata' where it means the statistical software 'Stata'.","section":"SM Sections 3.3 and 3.4"},{"comment":"The discussion should acknowledge the SM Section 3.9 limitation that the exact causal actor (bot vs. peer) cannot be isolated, as the authors themselves note in SM, to help readers correctly interpret the follow-signal results.","section":"Main text, Discussion"},{"comment":"The main text should state the batch-based, interactive design (10–28 participants per session, three rounds) in the Methods rather than only in the SM, since this design feature is central to the statistical analysis.","section":"Materials and Methods"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and, after a careful reanalysis with batch-level clustering and correction of the table, could be a strong contribution. Please ask the authors to also justify the personalization attribution or soften it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core result here is new and important: a personalized LLM embedded in a live social network moves both individual beliefs toward factual accuracy and, even with the bot excluded, shifts the set of people participants choose to follow toward more accurate peers. That follow-signal outcome is a genuinely new measurement, and the experimental design is clever, with real-time networked interaction and repeated rounds. The robustness checks are extensive, and the SM confirms the direction of the main effects, so the reader's verdict of \"conditional\" is about right.\n\nThe paper's soft spots are real but mostly fixable. First, Table 1 in the main text has the treatment and control means swapped as printed: the text and SM both say treatment shift is -0.35 and control -0.1, but the table prints the opposite. This is clearly a labeling error, and the SM's Table S12/S13 resolves it, but it has to be corrected before publication because as printed the table contradicts the abstract. Second, and more substantively, the stress-test note is right: participants were run in synchronous batches of 10-28, condition appears to be assigned at the batch level, and participants influenced each other's follow choices. The reported t-tests treat all 1265 participants as independent. That inflates significance. Given the effect size (a 0.25-point shift on a 4-point scale) and roughly 80 batches, the findings would likely survive cluster-robust standard errors or a multilevel model, but the current p-values are overconfident and the confidence intervals in Tables 1 and 2 are too narrow. This needs to be reanalyzed, not just acknowledged. Third, there is no non-personalized LLM arm, so the claim that personalization is the active ingredient is not identified; what is identified is that a factual, guardrailed LLM message in a network beats no such message. The authors should either temper the claim or add the condition. Fourth, no data, code, or preregistration link is shipped, which is a reproducibility gap for a pre-registered study.\n\nDespite these issues, this is not a desk-reject. It is a strong, interesting study that deserves serious peer review, with the expectation of substantial revision. The batch-level reanalysis is the load-bearing fix; the rest is presentation and framing. I would bring it to a reading group and would cite the follow-signal result once the clustering correction is in. Send it to review.","headline":"Novel and worth taking seriously, but the batch-clustering analysis and the missing non-personalized arm need to be fixed before the causal claims are clean.","tokens_in":26465,"tokens_out":2504,"would_cite":true,"duration_ms":28244,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91D30","62P25"],"pacs":[],"model":"deepseek-v4-flash","headline":"A personalized, accuracy-guarded LLM inside a social network shifted members' beliefs toward the truth and steered their follow choices toward accurate peers, in a pre-registered experiment of 1,265 people around the 2024 US election.","keywords":["personalized large language models","belief accuracy","social network formation","political misinformation","persuasive AI","pre-registered experiment","misinformation correction","network following behavior"],"falsifier":"Re-run the main comparisons with standard errors clustered by experimental session: if the treatment-control differences in belief shift or follow signal lose significance, the central claims fail. Separately, if a bot that confidently states falsehoods attracts followers and pulls follow-networks toward its own inaccurate positions as strongly as the accurate bot did, then the mechanism is similarity-seeking toward whatever the bot says, not a pull toward the truth.","tokens_in":25473,"feed_emoji":"🧭","tokens_out":12034,"duration_ms":106119,"temperature":0.7,"pith_summary":"People form and revise beliefs inside social networks, yet almost all previous tests of whether large language models change minds have examined one-on-one interactions. This paper argues that the more consequential effect appears when a personalized, accuracy-guarded LLM is embedded in the network itself: in a pre-registered experiment with 1,265 participants run around the 2024 US presidential election, people who saw the bot's tailored, fact-checked response moved their beliefs toward the truth about three times as far as controls. The deeper claim is social: most participants chose to follow the bot, and the human peers they followed scored measurably closer to factual accuracy than those followed in the control condition, so the agent's influence reached the structure of the network, not just individual ratings. If right, an accurate AI agent can act as a corrective force in online environments; the paper itself notes that the same mechanism, with the accuracy guardrails stripped out, could just as plausibly spread false beliefs.","feed_headline":"Personalized AI bots can steer social networks toward the truth","feed_subtitle":"In a 1,265-person experiment, an accurate chatbot improved beliefs and who people chose to follow.","key_machinery":"The load-bearing object is the personalized bot pipeline: a traditional machine-learning model converts each user's demographics and Big Five personality ratings into a predicted preferred news source and rhetorical style (ethos, pathos, or logos); a retrieval system over roughly 70,000 full-length articles from ten ideologically balanced outlets (three right, three center, four left) gathers evidence for or against the statement; and GPT-4o-mini summarizes the evidence and reframes it in the predicted style, with prompts that forbid hate speech. The measure that carries the network claim is the \"follow signal\": each participant's average factual-accuracy score (0 = most accurate to 4 = least accurate) of the human peers they chose to follow each round, with the single most-accurate entry removed so that the bot's mere presence cannot inflate the score. The protocol — initial belief rating with rationale, exposure to peer responses (plus the bot in treatment), then a three-of-six follow choice, repeated over three rounds in batches of 10 to 28 participants — is what lets the paper observe belief revision and network rewiring in the same session.","core_discovery":"The paper reports three pre-registered findings from a three-round networked experiment. First, hypothesis 1: compared with controls who saw three peer responses, participants who also saw a personalized bot response shifted their belief ratings toward the objectively correct answer (average signed shift of approximately $-0.35$ in the treatment versus approximately $-0.1$ in the control, $p < 0.0001$; negative means toward the truth), a difference that held across statement factuality, topic salience, economic relevance, media skepticism, demographics, and pre- versus post-election timing. Second, hypothesis 2a: 70.2% of treatment participants followed the bot at least once, and 54.4% of those kept following it across all three rounds. Third, hypothesis 2b: even after excluding the single most accurate option, treatment participants' \"follow signal\" — the mean factual-accuracy score of the people they chose to follow, on a 0-to-4 scale — was closer to the truth (1.55 versus 2.02, $p < 0.0001$), and their networks became more assortative, meaning they did not choose followers at random. The authors conclude that a truthful, personalized agent can function both as a direct informant and as a catalyst that reconfigures the surrounding network toward accuracy.","pith_inferences":["Because the follow-signal difference survived dropping the single most-accurate entry, the bot appears to have changed the criterion people used for choosing social ties, not merely added one accurate node; a direct test would remove the bot in a fourth round and check whether follow choices stay accurate.","Treatment participants' updated, truthward answers were shown to other people in later rounds, so the bot's influence may spill over to participants who never saw it; the paper does not measure this, though its design could.","Control participants also preferentially followed a \"most truthful\" human when one was identifiable, so a probe that replaces the bot with an equally accurate, well-written human argument would reveal whether AI identity adds anything beyond making an accurate opinion salient.","Editorial notes on the manuscript text: the means in main-text Table 1 appear swapped relative to the prose and the supplementary tables, and one citation marker in the hypotheses section is empty ('()'); the directional claims rest on the prose and supplementary numbers."],"forward_implications":["Deploying accurate, personalized LLM agents inside real online communities could shift not just what individuals believe but also whom they choose to follow, nudging local information environments toward verified claims.","The belief-shift effect held across statement factuality, topic salience, economic relevance, self-reported media skepticism, demographic groups, and pre- versus post-election timing, indicating the corrective effect is not limited to a single issue or audience.","Belief shifts occurred without back-and-forth dialogue and despite partial awareness of the bot, so brief and partly transparent AI messages may be enough to move beliefs in real-world settings.","The authors' risk analysis: the same pipeline with the accuracy guardrails replaced by a malicious agenda could amplify misinformation and reshape networks around false beliefs, marking where oversight is needed."],"supporting_citations":[{"why":"Establishes that personalized generative-AI messages persuade more than non-personalized or human messages, the premise for tailoring the bot to each user.","marker":"[18]"},{"why":"Shows AI dialogue can durably reduce conspiratorial beliefs; this is the one-on-one baseline the study extends into networked settings.","marker":"[15]"},{"why":"Documents the persuasive power of large language models, motivating the possibility that a bot can move beliefs at all.","marker":"[16]"},{"why":"Provides the truth-seeking account of why people would follow an accurate bot and accurate peers, underpinning hypotheses 2a and 2b.","marker":"[27]"},{"why":"Supports building the bot around accurate factual corrections rather than neutral persuasive messaging.","marker":"[26]"},{"why":"Supplies the prior political-science design with a small number of actors in a social network that the experimental platform follows.","marker":"[32]"},{"why":"Provides the networked opinion-formation rationale for studying belief change in social settings rather than in isolation.","marker":"[33]"},{"why":"Supplies the 10-item Big Five personality instrument used to personalize the bot's messages.","marker":"[36]"},{"why":"Underlies the partisan-motivated-reasoning account used to interpret why a mix of cross-partisan source cues boosted persuasion.","marker":"[22]"}],"fun_headline_variants":["AI bots nudge social networks toward accuracy","Personalized LLMs shift beliefs and follow choices toward truth","Truthful AI chatbots reshape social networks for accuracy","How personalized AI improves belief accuracy in networks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis treats each participant as an independent observation, but participants ran in synchronous batches of 10 to 28 people who saw each other's answers and influenced one another's follow choices; if outcomes are correlated within batches, the t-tests and regressions, computed without cluster-robust standard errors, overstate significance.","fun_headline_variants_meta":{"raw":{"variants":["AI bots nudge social networks toward accuracy","Personalized LLMs shift beliefs and follow choices toward truth","Truthful AI chatbots reshape social networks for accuracy","How personalized AI improves belief accuracy in networks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1295,"prompt_tokens":1050,"completion_tokens":245,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":186}},"tokens_in":666,"tokens_out":245,"duration_ms":2845,"temperature":1.0,"reasoning_tokens":186,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:59:35.657257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the main comparisons with standard errors clustered by experimental session: if the treatment-control differences in belief shift or follow signal lose significance, the central claims fail. Separately, if a bot that confidently states falsehoods attracts followers and pulls follow-networks toward its own inaccurate positions as strongly as the accurate bot did, then the mechanism is similarity-seeking toward whatever the bot says, not a pull toward the truth.","supporting_citations":[{"cited_title":"Matz, et al., The potential of generative AI for personalized persuasion at scale","cited_arxiv_id":null,"evidence_quote":"Establishes that personalized generative-AI messages persuade more than non-personalized or human messages, the premise for tailoring the bot to each user."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows AI dialogue can durably reduce conspiratorial beliefs; this is the one-on-one baseline the study extends into networked settings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the persuasive power of large language models, motivating the possibility that a bot can move beliefs at all."},{"cited_title":"Pennycook, et al., Shifting attention to accuracy can reduce misinformation online","cited_arxiv_id":null,"evidence_quote":"Provides the truth-seeking account of why people would follow an accurate bot and accurate peers, underpinning hypotheses 2a and 2b."},{"cited_title":"Porter, T","cited_arxiv_id":null,"evidence_quote":"Supports building the bot around accurate factual corrections rather than neutral persuasive messaging."},{"cited_title":"Klar, Partisanship in a social setting","cited_arxiv_id":null,"evidence_quote":"Supplies the prior political-science design with a small number of actors in a social network that the experimental platform follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the networked opinion-formation rationale for studying belief change in social settings rather than in isolation."},{"cited_title":"Rammstedt, O","cited_arxiv_id":null,"evidence_quote":"Supplies the 10-item Big Five personality instrument used to personalize the bot's messages."},{"cited_title":"Bolsen, J","cited_arxiv_id":null,"evidence_quote":"Underlies the partisan-motivated-reasoning account used to interpret why a mix of cross-partisan source cues boosted persuasion."}],"review_version":1}