{"id":"5f99db99-2a1c-4be0-8751-68fd1dbb9d92","arxiv_id":"2506.21620","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4 impersonating Reddit users in 2016 election threads produces comments that lean toward consensus and are semantically separable from real human comments.","lead":"This paper tests whether GPT-4 can pass as a human Reddit user in 2016 election threads, and finds it generates consensus more easily than dissent. The study also shows that real and AI-written comments form separate clusters in a semantic embedding space, pointing to a potential bot-detection signal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The consensus-over-dissent result depends on unvalidated GPT-4 labels whose in-context examples may bias anti-candidate classification; a human-annotated validation set is needed before the claim can be accepted.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test agrees that the main threat is the unvalidated GPT-4 classifier. I see this as the single most load-bearing assumption because every quantitative statement about 'consensus over dissent' is computed from those labels; the embedding/SVC result (Section S8) is independent and reasonably supported, but it does not bear on the consensus claim. The specific risk is sharper than a general 'LLM-as-judge' worry: the in-context examples are extreme, and GPT-4's safety training may make it reluctant to label non-explicit text as anti-candidate, inflating neutral. Because generator and evaluator are the same model, this would create exactly the observed pattern. The paper includes no human validation, no agreement statistics, and no variance estimates for the 5 runs. The proposed concrete test—human annotation of a stratified sample and recalibration—is feasible with modest resources and would settle whether the effect is real. If the human labels reproduce Figure 2, the finding is credible; if not, the paper should be revised or the claim weakened. Thus the verdict stays CONDITIONAL pending this check.","tokens_in":17415,"tokens_out":4321,"duration_ms":49296,"concrete_test":"Sample 200 comments from each real and generated condition (real, Scenario 1, 2, 3, null) per subreddit. Have at least two independent human annotators, blind to condition, assign the same three-category party alignment labels. Compute Cohen's kappa between annotators and per-class accuracy of GPT-4 labels against adjudicated human labels. Then re-estimate Figure 2 using human labels on the full dataset (or a GPT-4 classifier recalibrated on the human-annotated subset). The central claim survives only if, under human labels, the anti-candidate share in Scenario 3 remains significantly below the pro-candidate share in Scenario 2 for both subreddits (e.g., non-overlapping confidence intervals or a pre-registered hypothesis test).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that GPT-4 creates consensus more easily than dissent—rests entirely on the party-alignment labels assigned by GPT-4 in Section III-A. These labels are never checked against human judgments. The classification prompt is anchored with only extreme, explicitly abusive examples ('Hillary Clinton is a whore!' -> -1; 'Donald Trump is a piece of shit!' -> -1). This anchor can teach the model to reserve the 'anti-candidate' category for profanity-laden attacks and to class more measured or implicit criticism as neutral. Because the same model both generates and evaluates the comments, the low anti-candidate shares in Scenario 3 (32% for Clinton, 22% for Trump) and the high neutral shares may reflect a labeling bias, not a generation bias. No inter-annotator agreement, human benchmark, or error analysis is provided, and the reported percentages are averages over 5 runs without variance or significance tests. If the evaluator systematically undercounts dissent, the main finding dissolves; the paper's public-release claims (arXiv, no code/data) prevent external verification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates GPT-4's ability to generate Reddit comments in the context of the 2016 US presidential election, using three scenarios: impersonation of real users (Scenario 1), supportive bots (Scenario 2), and dissenting bots (Scenario 3), plus a null model with no user history. The authors classify generated and real comments for party alignment, sentiment, and violence using GPT-4 itself, and compare their semantic embeddings via t-SNE and a linear SVC. The central claims are that GPT-4 produces realistic comments but tends to create consensus more easily than dissent, and that real and artificial comments are separable in embedding space despite being indistinguishable by human inspection.","tokens_in":17549,"tokens_out":6569,"duration_ms":64173,"significance":"If the claims hold, the paper offers a useful empirical benchmark for LLM behavior in politically loaded online settings, with implications for bot detection and discourse manipulation. The use of real Reddit data, a null model, and multiple prompting conditions is a strength, and the supplementary SVC analysis provides a quantitative complement to the t-SNE visualization. However, the main finding is only as strong as the validity of GPT-4's self-ratings, which are not validated against human labels; a human-annotation study and statistical error bars would make the contribution far more credible.","major_comments":[{"comment":"The party-alignment and sentiment labels used throughout the paper are produced by GPT-4 with a single in-context example that is explicitly abusive ('Hillary Clinton is a whore!' -> -1, -1, 1). This anchoring is likely to teach the model to reserve the 'anti-candidate' category for profanity-laced attacks and to classify measured or implicit criticism as neutral, which would inflate the neutral share and deflate the anti-candidate share. The paper reports no human validation, inter-annotator agreement, or error analysis for these labels. Since the 'consensus over dissent' finding rests entirely on these labels, the authors must provide human-annotated validation (e.g., a few hundred comments labeled by 2-3 annotators) and report agreement metrics and a confusion matrix for the GPT-4 classifier.","section":"Section III-A (Party Alignment classification prompt)"},{"comment":"All percentages in Figures 2 and 3 are averages over 5 runs with no error bars, confidence intervals, or significance tests. The text draws strong conclusions from these point estimates: 'generated dissenting comments almost vanish' (Scenario 1: 0% and 1% anti-candidate) and 'the anti-candidate counterpart is steadily below 30%' (Scenario 3). With only five runs, these differences could be within run-to-run variability; the authors should report the per-run spread (e.g., standard deviation or bootstrap intervals) and, ideally, a permutation test comparing the generated anti-candidate shares against the real-comment baseline.","section":"Figures 2 and 3, Section III-A"},{"comment":"The claim that real and generated comments are 'indistinguishable by manual inspection' is not supported by a systematic human study. The text states only that the authors themselves found it difficult to tell them apart; no protocol, sample size, or accuracy measure is given. Additionally, the 'clear separation' in the t-SNE plots (Figure 4) is a visual judgement. The quantitative SVC experiment (Supplementary S8) is more convincing, but it is not cited in the main text and does not appear in the main figures. The authors should integrate a quantitative evaluation of separability (e.g., the SVC accuracy with confidence intervals) into the main text and either run or explicitly retract the anonymous human-inspection claim.","section":"Section III-B"}],"minor_comments":[{"comment":"The paper does not clearly state how many target comments were simulated per user; it reports that 100 and 387 users contributed 5,220 and 5,488 comments in 2016, but the number of simulated outputs per run is never specified. This makes it difficult to assess the effective sample sizes behind the percentages in Figures 2 and 3.","section":"Section II-B"},{"comment":"The sentence 'The model generates a little more dissent in Trump’s Subreddit (1.8% against 0.6% for Clinton), yet this difference may be explained by looking at the classification of the posts' presents a post-hoc explanation without testing; a simple bootstrap comparison would be more appropriate.","section":"Section III-A.1"},{"comment":"The t-SNE figures (Figure 4) do not report the perplexity or other hyperparameters, and the axis scales are not labeled with counts; for reproducibility, these details should be given.","section":"Section III-B"},{"comment":"The statement 'this behavior holds for both Trump and Clinton’s Subreddits' is too strong given the visible differences between the two subreddits in Scenario 3 (32% vs 22% anti-candidate); the conclusion should be qualified.","section":"Conclusion"},{"comment":"No code or data are released (arXiv submission, no links), which prevents external reproduction of the t-SNE and SVC analyses; a reproducibility statement would be desirable.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of cs.CL and would benefit from a human-annotation validation of the classifier; given the reliance on GPT-4 as both generator and evaluator, I recommend asking for this before publication. The lack of code/data release is also a concern for a quantitative claim; many journals now expect a reproducibility statement. The authors' claim that the model's behavior shows no political bias is based on the same unvalidated labels, so it should be tempered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a real attempt to test whether GPT-4 can impersonate Reddit users in a divisive political setting, using a real 2016 corpus. The three-scenario design is sensible, and the embedding-based separation between real and generated comments is the most convincing part. The SVC numbers (around 79% accuracy, with real comments almost perfectly separated) are reported with standard deviations over 10 runs, and the robustness check with Sentence-BERT is a good sign. If you need a paper that demonstrates LLM text has a detectable semantic footprint in social media, this one gives you reproducible evidence.\n\nThe soft spot is the central claim about consensus versus dissent. The percentages in Figure 2 come from GPT-4 classifying its own outputs. There is no human validation, no inter-annotator agreement, no error analysis, and the in-context example is an explicitly abusive sentence ('Hillary Clinton is a whore!'). That can teach the model to reserve 'anti-candidate' for profanity and call measured criticism neutral. So the low anti-candidate shares in Scenario 3 could be a labeling bias rather than a generation bias. The reader's stress-test note is right: this is the load-bearing assumption, and it is not tested. Also, the manual claim that real and generated comments are indistinguishable by inspection is just an assertion; no systematic human study. And the main percentages lack error bars, even though they averaged over 5 runs.\n\nThe paper would be more honest if it presented the consensus-over-dissent result as suggestive. The authors do list limitations, but not the classifier-validation problem. That omission matters because the conclusion leans heavily on it.\n\nWho is this for? People working on LLM agent simulation, bot detection, and AI-driven manipulation. The embedding result and the scenario framework are worth citing. The classifier issue is fixable; a human-annotated validation set of a few hundred comments would settle it.\n\nRecommendation: this deserves serious peer review. It is not a desk reject; the design is original and the embedding result is solid. But a referee should ask for the validation set, error bars, and code/data release before acceptance. My own verdict would be conditional, not accept.","headline":"Solid, useful simulation study on Reddit that overclaims its main finding because the labels come from the same model that wrote the comments.","tokens_in":18127,"tokens_out":1551,"would_cite":true,"duration_ms":17139,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4 impersonates Reddit users convincingly but generates consensus more readily than dissent.","keywords":["large language models","GPT-4","Reddit","political alignment","consensus bias","embedding space","bot detection","2016 US election"],"falsifier":"Have a panel of human annotators label the same real and generated comments for candidate alignment, and compare those labels with GPT-4's; if human labels raise the anti-candidate share of Scenario 3 toward the pro-candidate share of Scenario 2, or if the null model stops looking neutral, the paper's consensus-over-dissent claim collapses. A second check: train the five-class embedding classifier on one subreddit and test on the other; if accuracy drops to chance, the reported separability may be dataset-specific rather than a general bot trace.","tokens_in":17164,"feed_emoji":"🤖","tokens_out":7578,"duration_ms":80241,"temperature":0.7,"pith_summary":"This paper asks whether a large language model can pass as a human participant in politically charged online discussions, and what traces it leaves if it can. Using Reddit threads from the 2016 US presidential election, the authors prompt GPT-4 to write comments as a real user, as a pro-candidate supporter, or as an anti-candidate opponent, and compare the output to the original human comments. The central result is an asymmetry: GPT-4 reproduces the style of human discourse, but it generates consensus much more readily than dissent, so even an explicitly anti-candidate identity yields mostly neutral or pro-candidate text. The paper also finds that real and synthetic comments separate into distinct clusters in semantic embedding space, offering a quantitative route toward catching such bots, even though the generated texts are hard to tell apart by eye.","feed_headline":"GPT-4 fakes Reddit users, yet rarely produces dissent","feed_subtitle":"Told to write against a candidate, it mostly stays neutral or agreeable, while its output leaves a detectable embedding trace.","key_machinery":"The experimental machinery combines three impersonation scenarios that differ only in the user history fed to the prompt (real past comments, a fictitious pro-candidate history, or a fictitious anti-candidate history), plus a null model with no history. The generated texts are labeled by GPT-4 itself with three three-valued tags (party alignment, sentiment, violence) using in-context learning, and are embedded with a text-embedding model, projected with a dimensionality-reduction technique, and finally separated with a linear support vector machine in a five-class task. The load-bearing object is the prompt-to-history pairing: it isolates the effect of the supplied political identity on the model's output.","core_discovery":"GPT-4, when asked to impersonate a user in a Reddit thread, produces comments that look human but are measurably different from what people actually write. The most important difference is that the model avoids dissent: with a fictitious history that strongly opposes the subreddit's candidate, the generated comment is still anti-candidate only about 32% of the time on the Clinton side and 22% on the Trump side, while a supportive history produces pro-candidate comments over half the time. With a real user's history, generated comments become more pro-candidate than the user's actual comments, even for users whose history is anti-candidate. At the same time, synthetic comments occupy their own region of the embedding space: a linear classifier trained on text embeddings separates real comments from four types of generated ones with roughly 79% average accuracy, and a two-dimensional projection shows the clusters clearly, though human inspection cannot reliably distinguish the texts.","pith_inferences":["The consensus bias the authors observe may reflect the model's alignment training rather than a property of language itself; a useful next test would be to run the same prompts on models with different reinforcement-learning-from-human-feedback policies and compare Scenario 3 dissent rates.","The embedding-space separation could be a moving target: as future models are trained to imitate human style more closely, the linear classifier accuracy may degrade, so the method should be re-benchmarked periodically.","If the paper's result generalizes, platform moderators could treat unusually high dissent rates in a community as a signal of human activity, while unusually consensus-bound responding might flag automation.","The five-class classifier was trained and tested on the same subreddits; a fairer estimate of detection value would require a held-out election cycle or a different platform."],"forward_implications":["A partisan bot built on GPT-4 could seed supportive comments in a friendly community more reliably than it could sow dissent in an enemy community.","Because generated comments cluster separately from real ones in embedding space, a linear classifier offers a practical, though not perfect, screen for LLM-written political posts.","The near-total absence of violent generated text suggests built-in content constraints carry over even when the model is explicitly role-playing an aggressive user.","The finding that longer prompts push generated comments toward the subreddit's leaning implies that context-rich threads make consensus easier to elicit."],"supporting_citations":[{"why":"Supplies the curated comment dataset for the 2016 Reddit political threads used as the real-world testbed.","marker":"[27]"},{"why":"Provides the large-scale Reddit archive from which posts and user histories are extracted.","marker":"[28]"},{"why":"Documents the specific generative model whose impersonation behavior is the subject of the study.","marker":"[4]"},{"why":"Establishes the prior result that humans struggle to distinguish synthetic from real conversations, which motivates the realism test.","marker":"[15]"},{"why":"Provides the dimensionality-reduction technique used to visualize the separation between real and generated comments.","marker":"[30]"},{"why":"Supplies the alternative embedding model used in the supplementary robustness check that the real/generated clusters persist with a different encoder.","marker":"[35]"}],"fun_headline_variants":["GPT-4 fakes Reddit users, but dodges dissent","LLM posts as Reddit users, but rarely dissents","Mimicked Reddit users: human-like but detectable","AI impersonates Reddit commenters, avoids disagreement","GPT-4's Reddit personas: realistic, detectable, rare dissent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire consensus-versus-dissent result rests on GPT-4's own three-number labels for party alignment, sentiment, and violence, which are used without validation against human annotators; if the evaluator systematically calls its own generated comments neutral or pro-candidate, the asymmetry would be an artifact of the measurer.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 fakes Reddit users, but dodges dissent","LLM posts as Reddit users, but rarely dissents","Mimicked Reddit users: human-like but detectable","AI impersonates Reddit commenters, avoids disagreement","GPT-4's Reddit personas: realistic, detectable, rare dissent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001423,"raw_usage":{"total_tokens":5749,"prompt_tokens":956,"completion_tokens":4793,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":4717}},"tokens_in":572,"tokens_out":4793,"duration_ms":33583,"temperature":1.0,"reasoning_tokens":4717,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:15:29.142757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human annotators label the same real and generated comments for candidate alignment, and compare those labels with GPT-4's; if human labels raise the anti-candidate share of Scenario 3 toward the pro-candidate share of Scenario 2, or if the null model stops looking neutral, the paper's consensus-over-dissent claim collapses. A second check: train the five-class embedding classifier on one subreddit and test on the other; if accuracy drops to chance, the reported separability may be dataset-specific rather than a general bot trace.","supporting_citations":[{"cited_title":"Pierrehumbert","cited_arxiv_id":null,"evidence_quote":"Supplies the curated comment dataset for the 2016 Reddit political threads used as the real-world testbed."},{"cited_title":"Social simulacra: Creating populated prototypes for social computing systems","cited_arxiv_id":null,"evidence_quote":"Establishes the prior result that humans struggle to distinguish synthetic from real conversations, which motivates the realism test."},{"cited_title":"Visualizing data using t-SNE","cited_arxiv_id":null,"evidence_quote":"Provides the dimensionality-reduction technique used to visualize the separation between real and generated comments."},{"cited_title":"in reply to:","cited_arxiv_id":null,"evidence_quote":"Supplies the alternative embedding model used in the supplementary robustness check that the real/generated clusters persist with a different encoder."}],"review_version":1}