{"id":"a884b16c-614c-44dd-a615-577bf23957c1","arxiv_id":"2412.02712","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using LLM-based stance classification, the paper finds that Republican candidates tweeted more anti-Democrat content than Democrats tweeted anti-Republican, and that replies to candidate tweets skewed Republican-aligned.","lead":"This paper uses three AI language models to sort 2024 U.S. election candidate tweets and replies into pro/anti party categories, finding that Republican candidates posted more attacks on Democrats than the reverse. It also reports that replies skewed Republican-aligned and that major events shifted tweet stances, though the comparisons have timing and design limitations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RQ1's headline comparison is not time-matched: Republican tweets are only from Aug 12–Nov 1 while Democratic tweets span May–Nov, so the higher anti-Democrat rate could be an artifact of the later, more heated campaign period.","rationale":"The reader's weakest assumption is the same one I would put at the center: the RQ1 comparison pools samples from non-overlapping periods, and the paper explicitly reports the cause of the imbalance without providing a common-window analysis. This is a genuine confound, not a disagreement with consensus: if the time-matched test is non-significant, the abstract's first empirical claim would not follow. I do not see a stronger objection. The LLM validation (91% accuracy, kappa >= 0.86) is reasonable evidence that the classification pipeline is not the main issue; the event-study windows in RQ4 are also overlapping, but that concern is secondary to RQ1. Thus the reader's CONDITIONAL verdict remains appropriate, pending a time-matched recomputation.","tokens_in":6605,"tokens_out":6583,"duration_ms":62620,"concrete_test":"Re-run RQ1 on the common time window Aug 12–Nov 1, 2024: keep all 128 Republican candidate tweets and the Democratic candidate tweets from the same dates, then recompute the 2x2 chi-square test comparing anti-opposite-party proportions. If chi-squared no longer reaches p < 0.05 (or the direction flips), the headline claim is an artifact of the unequal observation windows; report the matched sample sizes and exact p-value.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central RQ1 finding depends on comparing the proportion of anti-opposite-party tweets across the two parties, but the samples are drawn from different time windows. In the Validation paragraph of Section 2 the authors note that Trump returned to Twitter on August 12, 2024, leaving 128 Republican candidate tweets versus 1,107 Democratic candidate tweets; Section 3 (RQ1) nevertheless pools all tweets from May 1 to November 1. The Republican sample therefore covers only the post-Biden-withdrawal period, after Harris entered and after the assassination attempt, while the Democratic sample includes five months of earlier, less polarized messaging. Because opposition-focused messaging generally intensifies close to an election, the observed 40.6% vs 26.4% difference and chi-squared = 11.55 (p < 0.001) may simply reflect the phase of the campaign rather than a stable party difference. The Limitations section only claims that collection-frequency effects apply uniformly to both parties; it never addresses unequal observation windows. No time-matched subsample or date-controlled model is reported, so the headline claim is not yet supported as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 1,235 tweets from U.S. presidential candidates (Joe Biden, Kamala Harris, Tim Walz, Donald Trump, JD Vance) and 63,322 replies to those tweets, together with 32,832 event-period tweets, in the lead-up to the 2024 U.S. election. Using a stance-classification pipeline that combines three LLMs (GPT-4o, Gemini-Pro, Claude-Opus) with human validation, the authors classify tweets into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. They report four research questions: RQ1 finds Republican candidates tweet a significantly higher proportion of anti-opposite-party content than Democratic candidates; RQ2 finds replies skew Republican regardless of the candidate; RQ3 examines engagement patterns of reply stances; RQ4 uses regression discontinuity in time around three major political events to estimate shifts in ideological stance. The paper concludes that Republican candidate messaging was more opposition-focused and that major events shifted public discourse toward support or criticism of Trump or the Republican party.","tokens_in":6815,"tokens_out":4041,"duration_ms":33699,"significance":"If the findings hold, the paper contributes a timely empirical description of asymmetric political messaging on Twitter during a U.S. election cycle, using a publicly available dataset and a transparent classification pipeline. The study has concrete strengths: the annotation pipeline is validated against human coders with reported accuracy above 90% and inter-rater agreement above 0.79; the approach uses consensus among three LLMs with human adjudication; and the authors provide an open GitHub repository with annotated data and reproduction code. These elements support the reproducibility of the measurement layer. However, the central RQ1 claim depends crucially on the comparability of the two parties' tweet samples, which is currently undermined by a temporal mismatch, and the RQ4 regression-discontinuity analysis has overlapping treatment windows and uncorrected multiple testing. The paper's significance would be materially strengthened by a time-matched robustness check for RQ1 and a reworked RQ4 that addresses identification and inference concerns.","major_comments":[{"comment":"The central RQ1 comparison pools 1,107 Democratic candidate tweets spanning May 1–November 1 with 128 Republican candidate tweets that begin only on August 12, after Donald Trump's return to Twitter (stated in §2, Validation). Because campaign messaging becomes more opposition-focused as an election approaches, the higher anti-Democrat proportion among Republican tweets (40.6% vs. 26.4%; chi-squared = 11.55, p < 0.001) may be an artifact of the later, more heated observation window rather than a stable party difference. The Limitations section addresses only collection-frequency uniformity and does not confront unequal observation windows. Please provide a time-matched subsample (e.g., Republican tweets vs. Democratic tweets from August 12–November 1) or a date-controlled regression to support the headline claim.","section":"§2 (Validation) and §3 (RQ1)"},{"comment":"The RDiT analyses use two-week windows that overlap substantially across events: the debate window (June 27 ± 14 days = June 13–July 11) overlaps the Supreme Court ruling window (July 1 ± 14 days = June 17–July 15) and the assassination window (July 13 ± 14 days = June 29–July 27). Treatment effects estimated in one window may therefore absorb effects of adjacent events, so the reported causal attribution to a single event is not identified. In addition, the section reports point estimates and p-values without confidence intervals, and the many significance tests across five stances and three events are not corrected for multiple comparisons. Please report confidence intervals, non-overlapping or explicitly robust windows, and an account of multiple-testing correction.","section":"§3 (RQ4)"}],"minor_comments":[{"comment":"In the sentence referring to the work of Ye et al., 'socket puppet driven experiment' should read 'sock puppet driven experiment'.","section":"Introduction"},{"comment":"The name 'Krippendorf' is misspelled; the standard spelling is 'Krippendorff'.","section":"Table 1 and §2"},{"comment":"The test statistic is described as 'two-sided independent t-test; z = 15.19' but a t-test yields a t-statistic; please clarify the test and the statistic actually used.","section":"§3 (RQ2)"},{"comment":"The text describing Figure 1 appears to swap the labels: the paragraph references 'Fig. 1C, for replies to Republican candidates' and 'Figure 1D, while the majority of replies received by Democratic candidates', whereas the caption assigns (C) to Democrat candidates and (D) to Republican candidates.","section":"Figure 1"},{"comment":"Use 'vice versa' instead of 'vice-versa' in the RQ1 paragraph and in the abstract.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The temporal mismatch in RQ1 is the main obstacle; the authors should be asked to supply the matched-window analysis before publication. The RDiT overlapping windows and lack of confidence intervals also require substantive revision. If these points are addressed convincingly, the paper would be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a competent, transparent LLM-based stance-detection study of candidate and reply tweets from the 2024 U.S. election cycle, and it ships annotated data and reproduction code. But the headline RQ1 claim — Republican candidates tweet more anti-Democrat criticism than the reverse — is not supported as stated. The comparison pools non-comparable time windows, and that is a load-bearing flaw.\n\nWhat it does well: the five-way label scheme (pro-party vs. anti-opponent) is a reasonable adaptation of standard stance detection, and the validation is solid — three independent coders, over 90% accuracy across tasks, and inter-rater and inter-LLM agreement statistics are reported. The authors are transparent about the data imbalance, and the public code and data are a real plus.\n\nThe soft spots: Trump was off Twitter until August 12, so all 128 Republican candidate tweets come from the final stretch of the campaign, while the 1,107 Democratic tweets span May through November. Anti-opposition messaging tends to intensify late in a campaign, so the observed 40.6% vs. 26.4% difference and the chi-squared test do not separate party from phase. The limitations section only mentions collection-frequency effects, which is not the issue. The same caveat applies to the RQ2/RQ3 comparisons between replies to Democratic and Republican candidates, since those replies inherit the same temporal imbalance. No time-matched subsample or date-controlled model is provided.\n\nRQ4 is also shaky: the two-week windows around the three events overlap (the debate window contains part of the SCOTUS-ruling window), which undermines the sharp-cutoff RDiT assumption, and the paper reports only p-values, no confidence intervals. Multiple testing is a secondary concern.\n\nOverall, the empirical description is plausible, but the central comparative claims need rework before they can be believed. A time-matched analysis with non-overlapping event windows would be a straightforward fix and would determine whether the party asymmetry survives.\n\nFor a referee: yes, send this to review. It is a descriptive empirical paper with a clear pipeline and a fixable confound; a good reviewer can help the authors make it credible. I would not cite the headline claim as it stands.","headline":"Solid LLM annotation work with a transparent pipeline, but the headline party-asymmetry claim is confounded by unequal observation windows and needs a time-matched reanalysis.","tokens_in":7321,"tokens_out":2876,"would_cite":false,"duration_ms":25891,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Republican candidates on Twitter posted a significantly higher share of tweets criticizing the Democratic Party than Democratic candidates did of Republicans, while replies showed a different, more Republican-leaning pattern.","keywords":["Twitter","2024 U.S. election","stance detection","text classification","large language models","political polarization","ideological framing","regression discontinuity in time"],"falsifier":"Take the 128 Republican candidate tweets (all posted after August 12) and compare them to a time-matched random sample of Democratic candidate tweets from August 12 to November 1. If the proportion of Anti-Democrat tweets among Republicans is no longer significantly higher than the proportion of Anti-Republican tweets among Democrats, the paper's central claim fails.","tokens_in":6445,"feed_emoji":"🗳️","tokens_out":4591,"duration_ms":38605,"temperature":0.7,"pith_summary":"The paper tries to establish that in the months before the 2024 U.S. presidential election, Republican and Democratic candidates on Twitter positioned their messages differently: Republican candidates devoted a significantly larger share of their tweets to criticizing the Democratic Party than Democratic candidates devoted to criticizing Republicans, while both sides devoted similar shares to supporting their own party. It also argues that replies to candidate tweets do not follow the same pattern, instead skewing toward Republican-aligned stances regardless of which party's candidate was replied to. If correct, the results would show an asymmetry in candidate messaging on a major public platform and a decoupling between elite messaging and constituent replies, with implications for how polarization is understood and addressed.","feed_headline":"Republican tweets attack Democrats far more than the reverse","feed_subtitle":"LLM-based analysis of 1,235 candidate tweets and 63,322 replies in the six months before the election.","key_machinery":"The central object is a five-way stance classification scheme that separates support for a party from opposition to the opposing party, labeling each tweet as Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, or Neutral. The scheme is applied through an LLM consensus pipeline (GPT-4o and Gemini-Pro first, Claude-Opus as tiebreaker, and human adjudication as final), validated against human annotations at over 90% accuracy. It carries the argument because the central asymmetry in candidate tweets is defined in terms of the Anti-Democrat versus Anti-Republican categories, and the same categories structure the reply and event analyses.","core_discovery":"The central claim is that Republican candidates on Twitter were more opposition-focused than Democratic candidates during the study period. Among 1,107 Democratic candidate tweets, 26.4% were classified as Anti-Republican, while among 128 Republican candidate tweets, 40.6% were Anti-Democrat, and a chi-squared test rejects the hypothesis that these proportions are equal ($\\chi^2 = 11.55$, $p<0.001$). The paper also claims that replies to candidate tweets were more Republican-aligned (Pro-Republican or Anti-Democrat) than Democrat-aligned overall, regardless of which candidate was replied to, and that major political events produced shifts in the stance mix of public tweets, typically increasing both Pro-Republican and Anti-Republican replies while decreasing Pro-Democrat and Anti-Democrat replies.","pith_inferences":["If the candidate asymmetry is real, it may reflect a deliberate Republican campaign strategy of oppositional messaging; a direct test would compare candidate tweets to those of other Republican and Democratic officeholders over the same period, which the paper does not do.","The dominance of Republican-aligned replies to both parties' candidates could stem from asymmetrical platform activity by partisan users rather than from differences in candidate messaging; the dataset cannot distinguish these explanations, but polling other platforms could.","The event-driven increase in both Pro-Republican and Anti-Republican replies suggests that major Trump-related events mobilize both his supporters and his opponents simultaneously, a dynamic that could be tested on later events such as the election result itself.","The paper's time-mismatched samples make the headline asymmetry fragile; a re-analysis with a time-matched subsample would either confirm the finding or reveal it as an artifact of the campaign calendar."],"forward_implications":["Candidate messaging on Twitter is asymmetric in opposition framing: Republican candidates lean more heavily on attacks on the Democratic Party than Democratic candidates do on attacks on Republicans.","Replies to candidates do not mirror candidate framing; Republican-aligned replies dominate replies to both parties' candidates, suggesting constituent activity is skewed toward the Republican side.","For Democratic candidates, the most common reply type (Anti-Democrat) is not the most engaged-with; Pro-Democrat replies receive more engagement, pointing to a disconnect between reply volume and engagement.","Each of the three major political events studied produced a significant rise in both Anti-Republican and Pro-Republican tweets and a fall in Pro-Democrat and Anti-Democrat tweets, indicating event-driven discourse centered on Trump and the Republican Party.","The asymmetry in candidate framing coexists with an asymmetric reply pattern, meaning the two levels of discourse are driven by different dynamics rather than a single polarization mechanism."],"supporting_citations":[{"why":"Supplies the full Twitter dataset of candidate tweets and replies from which the analysis is drawn.","marker":"[2]"},{"why":"Shows that zero-shot LLM prompting can reliably identify stance, the basis for the classification pipeline.","marker":"[13]"},{"why":"Demonstrates that LLMs can outperform crowd workers at text annotation, supporting the model-as-annotator approach.","marker":"[6]"},{"why":"Provides evidence that LLMs can scale ideology classification of U.S. political figures from public statements.","marker":"[11]"},{"why":"Shows LLMs outperform expert coders and supervised classifiers on political social-media annotation, justifying the validation strategy.","marker":"[10]"}],"fun_headline_variants":["GOP candidates tweet more anti-Democrat than Dems anti-GOP","Republican candidates' tweets are more anti-Democrat, study finds","On Twitter, GOP attacks Democrats more than Dems attack GOP","Candidate tweets: Republicans aim more attacks at Democrats","Study: Republican candidates' tweets hit Democrats harder"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison of candidate tweet framing assumes the two parties' tweet samples are comparable, but Republican tweets in the dataset begin only in mid-August (after Trump returned to Twitter) while Democratic tweets span May to November, so the higher rate of anti-Democrat messaging among Republicans could simply reflect the more aggressive messaging of the late campaign period.","fun_headline_variants_meta":{"raw":{"variants":["GOP candidates tweet more anti-Democrat than Dems anti-GOP","Republican candidates' tweets are more anti-Democrat, study finds","On Twitter, GOP attacks Democrats more than Dems attack GOP","Candidate tweets: Republicans aim more attacks at Democrats","Study: Republican candidates' tweets hit Democrats harder"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000351,"raw_usage":{"total_tokens":1900,"prompt_tokens":920,"completion_tokens":980,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":896}},"tokens_in":536,"tokens_out":980,"duration_ms":8822,"temperature":1.0,"reasoning_tokens":896,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:41:47.849368+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 128 Republican candidate tweets (all posted after August 12) and compare them to a time-matched random sample of Democratic candidate tweets from August 12 to November 1. If the proportion of Anti-Democrat tweets among Republicans is no longer significantly higher than the proportion of Anti-Republican tweets among Democrats, the paper's central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that zero-shot LLM prompting can reliably identify stance, the basis for the classification pipeline."},{"cited_title":"Chatgpt outperforms crowd workers for text-annotation tasks","cited_arxiv_id":null,"evidence_quote":"Demonstrates that LLMs can outperform crowd workers at text annotation, supporting the model-as-annotator approach."},{"cited_title":"Y., Nagler, J., Tucker, J","cited_arxiv_id":null,"evidence_quote":"Provides evidence that LLMs can scale ideology classification of U.S. political figures from public statements."},{"cited_title":"Large language models outperform expert coders and supervised classifiers at annotating political social media messages","cited_arxiv_id":null,"evidence_quote":"Shows LLMs outperform expert coders and supervised classifiers on political social-media annotation, justifying the validation strategy."}],"review_version":1}