Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election

T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Republican candidates on Twitter posted a significantly higher share of tweets criticizing the Democratic Party than Democratic candidates did of Republicans, while replies showed a different, more Republican-leaning pattern.

desk verdict Solid LLM annotation work with a transparent pipeline, but the headline party-asymmetry claim is confounded by unequal observation windows and needs a time-matched reanalysis. read the letter →

arxiv 2412.02712 v1 pith:KPJFOR5P submitted 2024-11-28 cs.SI cs.CY

classification cs.SIcs.CY
keywords Twitter2024U.S.electionstancedetectiontextclassificationlargelanguagemodelspoliticalpolarizationideologicalframingregressiondiscontinuityintime
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that in the months before the 2024 U.S. presidential election, Republican and Democratic candidates on Twitter positioned their messages differently: Republican candidates devoted a significantly larger share of their tweets to criticizing the Democratic Party than Democratic candidates devoted to criticizing Republicans, while both sides devoted similar shares to supporting their own party. It also argues that replies to candidate tweets do not follow the same pattern, instead skewing toward Republican-aligned stances regardless of which party's candidate was replied to. If correct, the results would show an asymmetry in candidate messaging on a major public platform and a decoupling between elite messaging and constituent replies, with implications for how polarization is understood and addressed.

What carries the argument

The central object is a five-way stance classification scheme that separates support for a party from opposition to the opposing party, labeling each tweet as Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, or Neutral. The scheme is applied through an LLM consensus pipeline (GPT-4o and Gemini-Pro first, Claude-Opus as tiebreaker, and human adjudication as final), validated against human annotations at over 90% accuracy. It carries the argument because the central asymmetry in candidate tweets is defined in terms of the Anti-Democrat versus Anti-Republican categories, and the same categories structure the reply and event analyses.

What would settle it

Take the 128 Republican candidate tweets (all posted after August 12) and compare them to a time-matched random sample of Democratic candidate tweets from August 12 to November 1. If the proportion of Anti-Democrat tweets among Republicans is no longer significantly higher than the proportion of Anti-Republican tweets among Democrats, the paper's central claim fails.

Watch

Extended reading notes

Core claim

The central claim is that Republican candidates on Twitter were more opposition-focused than Democratic candidates during the study period. Among 1,107 Democratic candidate tweets, 26.4% were classified as Anti-Republican, while among 128 Republican candidate tweets, 40.6% were Anti-Democrat, and a chi-squared test rejects the hypothesis that these proportions are equal ($\chi^2 = 11.55$, $p<0.001$). The paper also claims that replies to candidate tweets were more Republican-aligned (Pro-Republican or Anti-Democrat) than Democrat-aligned overall, regardless of which candidate was replied to, and that major political events produced shifts in the stance mix of public tweets, typically increasing both Pro-Republican and Anti-Republican replies while decreasing Pro-Democrat and Anti-Democrat replies.

Load-bearing premise

The comparison of candidate tweet framing assumes the two parties' tweet samples are comparable, but Republican tweets in the dataset begin only in mid-August (after Trump returned to Twitter) while Democratic tweets span May to November, so the higher rate of anti-Democrat messaging among Republicans could simply reflect the more aggressive messaging of the late campaign period.

Editorial extensions

If this is right

  • Candidate messaging on Twitter is asymmetric in opposition framing: Republican candidates lean more heavily on attacks on the Democratic Party than Democratic candidates do on attacks on Republicans.
  • Replies to candidates do not mirror candidate framing; Republican-aligned replies dominate replies to both parties' candidates, suggesting constituent activity is skewed toward the Republican side.
  • For Democratic candidates, the most common reply type (Anti-Democrat) is not the most engaged-with; Pro-Democrat replies receive more engagement, pointing to a disconnect between reply volume and engagement.
  • Each of the three major political events studied produced a significant rise in both Anti-Republican and Pro-Republican tweets and a fall in Pro-Democrat and Anti-Democrat tweets, indicating event-driven discourse centered on Trump and the Republican Party.
  • The asymmetry in candidate framing coexists with an asymmetric reply pattern, meaning the two levels of discourse are driven by different dynamics rather than a single polarization mechanism.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the candidate asymmetry is real, it may reflect a deliberate Republican campaign strategy of oppositional messaging; a direct test would compare candidate tweets to those of other Republican and Democratic officeholders over the same period, which the paper does not do.
  • The dominance of Republican-aligned replies to both parties' candidates could stem from asymmetrical platform activity by partisan users rather than from differences in candidate messaging; the dataset cannot distinguish these explanations, but polling other platforms could.
  • The event-driven increase in both Pro-Republican and Anti-Republican replies suggests that major Trump-related events mobilize both his supporters and his opponents simultaneously, a dynamic that could be tested on later events such as the election result itself.
  • The paper's time-mismatched samples make the headline asymmetry fragile; a re-analysis with a time-matched subsample would either confirm the finding or reveal it as an artifact of the campaign calendar.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper analyzes 1,235 tweets from U.S. presidential candidates (Joe Biden, Kamala Harris, Tim Walz, Donald Trump, JD Vance) and 63,322 replies to those tweets, together with 32,832 event-period tweets, in the lead-up to the 2024 U.S. election. Using a stance-classification pipeline that combines three LLMs (GPT-4o, Gemini-Pro, Claude-Opus) with human validation, the authors classify tweets into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. They report four research questions: RQ1 finds Republican candidates tweet a significantly higher proportion of anti-opposite-party content than Democratic candidates; RQ2 finds replies skew Republican regardless of the candidate; RQ3 examines engagement patterns of reply stances; RQ4 uses regression discontinuity in time around three major political events to estimate shifts in ideological stance. The paper concludes that Republican candidate messaging was more opposition-focused and that major events shifted public discourse toward support or criticism of Trump or the Republican party.

Significance. If the findings hold, the paper contributes a timely empirical description of asymmetric political messaging on Twitter during a U.S. election cycle, using a publicly available dataset and a transparent classification pipeline. The study has concrete strengths: the annotation pipeline is validated against human coders with reported accuracy above 90% and inter-rater agreement above 0.79; the approach uses consensus among three LLMs with human adjudication; and the authors provide an open GitHub repository with annotated data and reproduction code. These elements support the reproducibility of the measurement layer. However, the central RQ1 claim depends crucially on the comparability of the two parties' tweet samples, which is currently undermined by a temporal mismatch, and the RQ4 regression-discontinuity analysis has overlapping treatment windows and uncorrected multiple testing. The paper's significance would be materially strengthened by a time-matched robustness check for RQ1 and a reworked RQ4 that addresses identification and inference concerns.

major comments (2)
  1. [§2 (Validation) and §3 (RQ1)] The central RQ1 comparison pools 1,107 Democratic candidate tweets spanning May 1–November 1 with 128 Republican candidate tweets that begin only on August 12, after Donald Trump's return to Twitter (stated in §2, Validation). Because campaign messaging becomes more opposition-focused as an election approaches, the higher anti-Democrat proportion among Republican tweets (40.6% vs. 26.4%; chi-squared = 11.55, p < 0.001) may be an artifact of the later, more heated observation window rather than a stable party difference. The Limitations section addresses only collection-frequency uniformity and does not confront unequal observation windows. Please provide a time-matched subsample (e.g., Republican tweets vs. Democratic tweets from August 12–November 1) or a date-controlled regression to support the headline claim.
  2. [§3 (RQ4)] The RDiT analyses use two-week windows that overlap substantially across events: the debate window (June 27 ± 14 days = June 13–July 11) overlaps the Supreme Court ruling window (July 1 ± 14 days = June 17–July 15) and the assassination window (July 13 ± 14 days = June 29–July 27). Treatment effects estimated in one window may therefore absorb effects of adjacent events, so the reported causal attribution to a single event is not identified. In addition, the section reports point estimates and p-values without confidence intervals, and the many significance tests across five stances and three events are not corrected for multiple comparisons. Please report confidence intervals, non-overlapping or explicitly robust windows, and an account of multiple-testing correction.
minor comments (5)
  1. [Introduction] In the sentence referring to the work of Ye et al., 'socket puppet driven experiment' should read 'sock puppet driven experiment'.
  2. [Table 1 and §2] The name 'Krippendorf' is misspelled; the standard spelling is 'Krippendorff'.
  3. [§3 (RQ2)] The test statistic is described as 'two-sided independent t-test; z = 15.19' but a t-test yields a t-statistic; please clarify the test and the statistic actually used.
  4. [Figure 1] The text describing Figure 1 appears to swap the labels: the paragraph references 'Fig. 1C, for replies to Republican candidates' and 'Figure 1D, while the majority of replies received by Democratic candidates', whereas the caption assigns (C) to Democrat candidates and (D) to Republican candidates.
  5. [Throughout] Use 'vice versa' instead of 'vice-versa' in the RQ1 paragraph and in the abstract.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical measurements from human-validated LLM classifications, with no fitted parameters or self-citation chain used to derive the results.

full rationale

The paper's central finding—that Republican candidates authored significantly more anti-Democrat tweets than vice versa (chi-squared = 11.55, p < 0.001)—is derived directly from counts of tweets classified into predefined stance categories. The classification pipeline uses three LLMs with human adjudication and is validated against human annotations, but this validation is an accuracy check, not an input to the statistical inference. No parameter is fitted to the outcome being predicted, and no equation defines the conclusion in terms of its own inputs. The stance categories are fixed before analysis, and the comparison is a standard frequency test on the resulting labels. The only self-citation (reference [7]) appears in the introduction as background on prior political-content analysis and is not load-bearing for any result. The paper's limitation that Republican candidate tweets are fewer and collected after Trump's return to Twitter in August is a validity concern about time-window comparability, not a circularity: even if the comparison were confounded by timing, the claim is not true by construction. No circular step can be exhibited from the text, so the appropriate score is 0.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The analysis rests on the validity of LLM-based stance classification and on the comparability of the samples being compared. The main free parameters are the RDiT bandwidth and daily sample size, both chosen by the authors. The stance categories are a domain assumption that forces single-label classification. The RDiT identification assumption is especially fragile because the post-event windows overlap.

free parameters (2)
  • RDiT bandwidth = 14 days (also 3, 7, 10 days)
    The choice of window width for the regression discontinuity in time analysis is researcher-selected, not data-driven, and the central event-effect claims are reported at the 14-day bandwidth.
  • Daily tweet sample size = 1000 tweets per day
    The event analysis samples 1000 tweets per day; the choice is arbitrary and may affect the stability of proportion estimates.
assumptions (5)
  • domain assumption The five stance categories (Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, Neutral) are mutually exclusive and exhaustive for political tweets.
    Used in the classification prompt; a tweet that simultaneously supports one party and attacks the other must be forced into a single category, potentially misrepresenting the tweet's stance (Section 2).
  • domain assumption The LLM consensus classification, validated on a few hundred examples, generalizes to the full dataset.
    Validation was on 250+250 replies and 128+128 candidate tweets; the accuracy of ~90% is assumed to hold for the remaining 63k replies and event tweets, but no category-level accuracy or per-stratum validation is reported (Section 2).
  • domain assumption The dataset from Balasubramanian et al. is representative of political tweets on Twitter during the study period.
    The paper relies on this dataset for all candidate tweets, replies, and event tweets; sampling biases in the original collection would propagate (Section 2).
  • domain assumption The RDiT design identifies the causal effect of each event under the assumption that no other event coincides within the window.
    The two-week post-debate window includes the July 1 SCOTUS ruling and July 13 assassination, violating this assumption for the debate effect (Section 3, RQ4).
  • standard math Tweets are independent observations for chi-squared and t-tests.
    Tweets by the same candidate or replies to the same tweet are likely correlated, but the analyses do not cluster standard errors by account or parent tweet (Section 3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election." pith.science (2026). https://pith.science/paper/KPJFOR5P

@misc{pith2026241202712,
  author       = {Pith},
  title        = {Pith review of: Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KPJFOR5P}},
  note         = {Machine review of arXiv:2412.02712}
}
read the original abstract

Social media platforms play a pivotal role in shaping public opinion and amplifying political discourse, particularly during elections. However, the same dynamics that foster democratic engagement can also exacerbate polarization. To better understand these challenges, here, we investigate the ideological positioning of tweets related to the 2024 U.S. Presidential Election. To this end, we analyze 1,235 tweets from key political figures and 63,322 replies, and classify ideological stances into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. Using a classification pipeline involving three large language models (LLMs)-GPT-4o, Gemini-Pro, and Claude-Opus-and validated by human annotators, we explore how ideological alignment varies between candidates and constituents. We find that Republican candidates author significantly more tweets in criticism of the Democratic party and its candidates than vice versa, but this relationship does not hold for replies to candidate tweets. Furthermore, we highlight shifts in public discourse observed during key political events. By shedding light on the ideological dynamics of online political interactions, these results provide insights for policymakers and platforms seeking to address polarization and foster healthier political dialogue.

Figures

Figures reproduced from arXiv: 2412.02712 by the authors.

Figure 1
Figure 1. The proportion of candidate tweets (A) and candidate tweet replies (B) classified as Pro-Democrat, Anti-Republican, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. RDiT analysis of the ideological stance of tweets in the two weeks surrounding major political events. (A) First [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TikTok's recommendations skewed towards Republican content during the 2024 U.S. presidential race

    cs.SI 2025-01 conditional novelty 7.0 of 10

    TikTok's recommendation algorithm served Republican-seeded test accounts more co-partisan content than Democratic-seeded accounts during the 2024 U.S. presidential race.

  2. Political-LLM: Large Language Models in Political Science

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A survey and taxonomy of LLM applications in political science, with a case study suggesting that larger LLMs reproduce ANES 2016 voting patterns more accurately than smaller ones.

Reference graph

Works this paper leans on

14 extracted references · 11 canonical work pages · cited by 2 Pith papers

  1. [1]

    Your stance is exposed! analysing possible factors for stance detection on social media

    Aldayel, A., and Magdy, W. Your stance is exposed! analysing possible factors for stance detection on social media. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019), 1–20

  2. [2]

    A public dataset tracking social media discourse about the 2024 us presidential election on twitter/x

    Balasubramanian, A., Zou, V., Narayana, H., You, C., Luceri, L., and Ferrara, E. A public dataset tracking social media discourse about the 2024 us presidential election on twitter/x. arXiv preprint arXiv:2411.00376 (2024)

  3. [3]

    Facebooking it to the polls: A study in online social networking and political behavior

    Bode, L. Facebooking it to the polls: A study in online social networking and political behavior. Journal of Information Technology & Politics 9 , 4 (2012)

  4. [4]

    The social structure of political echo chambers: Variation in ideological homophily in online networks.Political psychology (2017)

    Boutyline, A., and Willer, R. The social structure of political echo chambers: Variation in ideological homophily in online networks.Political psychology (2017)

  5. [5]

    The echo chamber effect on social media

    Cinelli, M., De Francisci Morales, G., Galeazzi, A., Quattrociocchi, W., and Starnini, M. The echo chamber effect on social media. Proceedings of the National Academy of Sciences 118 , 9 (2021), e2023301118

  6. [6]

    Chatgpt outperforms crowd workers for text-annotation tasks

    Gilardi, F., Alizadeh, M., and Kubli, M. Chatgpt outperforms crowd workers for text-annotation tasks. PNAS 120, 30 (2023), e2305016120

  7. [7]

    Youtube’s recom- mendation algorithm is left-leaning in the united states

    Ibrahim, H., AlDahoul, N., Lee, S., Rahwan, T., and Zaki, Y. Youtube’s recom- mendation algorithm is left-leaning in the united states. PNAS nexus (2023)

  8. [8]

    Linegar, M., Kocielnik, R., and Alvarez, R. M. Large language models and political science. Frontiers in Political Science 5 (2023), 1257092

Show all 14 references
  1. [9]

    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020)

    Stefanov, P., Darwish, K., Atanasov, A., and Nakov, P.Predicting the topical stance and political leaning of media using tweets. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020)

  2. [10]

    Large language models outperform expert coders and supervised classifiers at annotating political social media messages

    Törnberg, P. Large language models outperform expert coders and supervised classifiers at annotating political social media messages. Social Science Computer Review (2024), 08944393241286471

  3. [11]

    Y., Nagler, J., Tucker, J

    Wu, P. Y., Nagler, J., Tucker, J. A., and Messing, S. Large language models can be used to scale the ideologies of politicians in a zero-shot learning setting, april

  4. [12]

    Auditing political exposure bias: Algorithmic amplification on twitter/x approaching the 2024 us presidential election

    Ye, J., Luceri, L., and Ferrara, E. Auditing political exposure bias: Algorithmic amplification on twitter/x approaching the 2024 us presidential election. arXiv preprint arXiv:2411.01852 (2024)

  5. [13]

    Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., and Y ang, D.Can large lan- guage models transform computational social science? Computational Linguistics 50, 1 (2024), 237–291. 5

  6. [2023]

    arXiv preprint arXiv:2303.12057 (2023)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.