REVIEW 2 major objections 5 minor 2 cited by
Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Republican candidates on Twitter posted a significantly higher share of tweets criticizing the Democratic Party than Democratic candidates did of Republicans, while replies showed a different, more Republican-leaning pattern.
desk verdict Solid LLM annotation work with a transparent pipeline, but the headline party-asymmetry claim is confounded by unequal observation windows and needs a time-matched reanalysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a five-way stance classification scheme that separates support for a party from opposition to the opposing party, labeling each tweet as Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, or Neutral. The scheme is applied through an LLM consensus pipeline (GPT-4o and Gemini-Pro first, Claude-Opus as tiebreaker, and human adjudication as final), validated against human annotations at over 90% accuracy. It carries the argument because the central asymmetry in candidate tweets is defined in terms of the Anti-Democrat versus Anti-Republican categories, and the same categories structure the reply and event analyses.
What would settle it
Take the 128 Republican candidate tweets (all posted after August 12) and compare them to a time-matched random sample of Democratic candidate tweets from August 12 to November 1. If the proportion of Anti-Democrat tweets among Republicans is no longer significantly higher than the proportion of Anti-Republican tweets among Democrats, the paper's central claim fails.
Extended reading notes
Core claim
The central claim is that Republican candidates on Twitter were more opposition-focused than Democratic candidates during the study period. Among 1,107 Democratic candidate tweets, 26.4% were classified as Anti-Republican, while among 128 Republican candidate tweets, 40.6% were Anti-Democrat, and a chi-squared test rejects the hypothesis that these proportions are equal ($\chi^2 = 11.55$, $p<0.001$). The paper also claims that replies to candidate tweets were more Republican-aligned (Pro-Republican or Anti-Democrat) than Democrat-aligned overall, regardless of which candidate was replied to, and that major political events produced shifts in the stance mix of public tweets, typically increasing both Pro-Republican and Anti-Republican replies while decreasing Pro-Democrat and Anti-Democrat replies.
Load-bearing premise
The comparison of candidate tweet framing assumes the two parties' tweet samples are comparable, but Republican tweets in the dataset begin only in mid-August (after Trump returned to Twitter) while Democratic tweets span May to November, so the higher rate of anti-Democrat messaging among Republicans could simply reflect the more aggressive messaging of the late campaign period.
Editorial extensions
If this is right
- Candidate messaging on Twitter is asymmetric in opposition framing: Republican candidates lean more heavily on attacks on the Democratic Party than Democratic candidates do on attacks on Republicans.
- Replies to candidates do not mirror candidate framing; Republican-aligned replies dominate replies to both parties' candidates, suggesting constituent activity is skewed toward the Republican side.
- For Democratic candidates, the most common reply type (Anti-Democrat) is not the most engaged-with; Pro-Democrat replies receive more engagement, pointing to a disconnect between reply volume and engagement.
- Each of the three major political events studied produced a significant rise in both Anti-Republican and Pro-Republican tweets and a fall in Pro-Democrat and Anti-Democrat tweets, indicating event-driven discourse centered on Trump and the Republican Party.
- The asymmetry in candidate framing coexists with an asymmetric reply pattern, meaning the two levels of discourse are driven by different dynamics rather than a single polarization mechanism.
Reading between the lines
- If the candidate asymmetry is real, it may reflect a deliberate Republican campaign strategy of oppositional messaging; a direct test would compare candidate tweets to those of other Republican and Democratic officeholders over the same period, which the paper does not do.
- The dominance of Republican-aligned replies to both parties' candidates could stem from asymmetrical platform activity by partisan users rather than from differences in candidate messaging; the dataset cannot distinguish these explanations, but polling other platforms could.
- The event-driven increase in both Pro-Republican and Anti-Republican replies suggests that major Trump-related events mobilize both his supporters and his opponents simultaneously, a dynamic that could be tested on later events such as the election result itself.
- The paper's time-mismatched samples make the headline asymmetry fragile; a re-analysis with a time-matched subsample would either confirm the finding or reveal it as an artifact of the campaign calendar.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes 1,235 tweets from U.S. presidential candidates (Joe Biden, Kamala Harris, Tim Walz, Donald Trump, JD Vance) and 63,322 replies to those tweets, together with 32,832 event-period tweets, in the lead-up to the 2024 U.S. election. Using a stance-classification pipeline that combines three LLMs (GPT-4o, Gemini-Pro, Claude-Opus) with human validation, the authors classify tweets into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. They report four research questions: RQ1 finds Republican candidates tweet a significantly higher proportion of anti-opposite-party content than Democratic candidates; RQ2 finds replies skew Republican regardless of the candidate; RQ3 examines engagement patterns of reply stances; RQ4 uses regression discontinuity in time around three major political events to estimate shifts in ideological stance. The paper concludes that Republican candidate messaging was more opposition-focused and that major events shifted public discourse toward support or criticism of Trump or the Republican party.
Significance. If the findings hold, the paper contributes a timely empirical description of asymmetric political messaging on Twitter during a U.S. election cycle, using a publicly available dataset and a transparent classification pipeline. The study has concrete strengths: the annotation pipeline is validated against human coders with reported accuracy above 90% and inter-rater agreement above 0.79; the approach uses consensus among three LLMs with human adjudication; and the authors provide an open GitHub repository with annotated data and reproduction code. These elements support the reproducibility of the measurement layer. However, the central RQ1 claim depends crucially on the comparability of the two parties' tweet samples, which is currently undermined by a temporal mismatch, and the RQ4 regression-discontinuity analysis has overlapping treatment windows and uncorrected multiple testing. The paper's significance would be materially strengthened by a time-matched robustness check for RQ1 and a reworked RQ4 that addresses identification and inference concerns.
major comments (2)
- [§2 (Validation) and §3 (RQ1)] The central RQ1 comparison pools 1,107 Democratic candidate tweets spanning May 1–November 1 with 128 Republican candidate tweets that begin only on August 12, after Donald Trump's return to Twitter (stated in §2, Validation). Because campaign messaging becomes more opposition-focused as an election approaches, the higher anti-Democrat proportion among Republican tweets (40.6% vs. 26.4%; chi-squared = 11.55, p < 0.001) may be an artifact of the later, more heated observation window rather than a stable party difference. The Limitations section addresses only collection-frequency uniformity and does not confront unequal observation windows. Please provide a time-matched subsample (e.g., Republican tweets vs. Democratic tweets from August 12–November 1) or a date-controlled regression to support the headline claim.
- [§3 (RQ4)] The RDiT analyses use two-week windows that overlap substantially across events: the debate window (June 27 ± 14 days = June 13–July 11) overlaps the Supreme Court ruling window (July 1 ± 14 days = June 17–July 15) and the assassination window (July 13 ± 14 days = June 29–July 27). Treatment effects estimated in one window may therefore absorb effects of adjacent events, so the reported causal attribution to a single event is not identified. In addition, the section reports point estimates and p-values without confidence intervals, and the many significance tests across five stances and three events are not corrected for multiple comparisons. Please report confidence intervals, non-overlapping or explicitly robust windows, and an account of multiple-testing correction.
minor comments (5)
- [Introduction] In the sentence referring to the work of Ye et al., 'socket puppet driven experiment' should read 'sock puppet driven experiment'.
- [Table 1 and §2] The name 'Krippendorf' is misspelled; the standard spelling is 'Krippendorff'.
- [§3 (RQ2)] The test statistic is described as 'two-sided independent t-test; z = 15.19' but a t-test yields a t-statistic; please clarify the test and the statistic actually used.
- [Figure 1] The text describing Figure 1 appears to swap the labels: the paragraph references 'Fig. 1C, for replies to Republican candidates' and 'Figure 1D, while the majority of replies received by Democratic candidates', whereas the caption assigns (C) to Democrat candidates and (D) to Republican candidates.
- [Throughout] Use 'vice versa' instead of 'vice-versa' in the RQ1 paragraph and in the abstract.
Circularity Check
No significant circularity: the paper's claims are empirical measurements from human-validated LLM classifications, with no fitted parameters or self-citation chain used to derive the results.
full rationale
The paper's central finding—that Republican candidates authored significantly more anti-Democrat tweets than vice versa (chi-squared = 11.55, p < 0.001)—is derived directly from counts of tweets classified into predefined stance categories. The classification pipeline uses three LLMs with human adjudication and is validated against human annotations, but this validation is an accuracy check, not an input to the statistical inference. No parameter is fitted to the outcome being predicted, and no equation defines the conclusion in terms of its own inputs. The stance categories are fixed before analysis, and the comparison is a standard frequency test on the resulting labels. The only self-citation (reference [7]) appears in the introduction as background on prior political-content analysis and is not load-bearing for any result. The paper's limitation that Republican candidate tweets are fewer and collected after Trump's return to Twitter in August is a validity concern about time-window comparability, not a circularity: even if the comparison were confounded by timing, the claim is not true by construction. No circular step can be exhibited from the text, so the appropriate score is 0.
Assumptions & free parameters
free parameters (2)
- RDiT bandwidth =
14 days (also 3, 7, 10 days)
- Daily tweet sample size =
1000 tweets per day
assumptions (5)
- domain assumption The five stance categories (Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, Neutral) are mutually exclusive and exhaustive for political tweets.
- domain assumption The LLM consensus classification, validated on a few hundred examples, generalizes to the full dataset.
- domain assumption The dataset from Balasubramanian et al. is representative of political tweets on Twitter during the study period.
- domain assumption The RDiT design identifies the causal effect of each event under the assumption that no other event coincides within the window.
- standard math Tweets are independent observations for chi-squared and t-tests.
Cite this review
Pith. "Pith review of Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election." pith.science (2026). https://pith.science/paper/KPJFOR5P
@misc{pith2026241202712,
author = {Pith},
title = {Pith review of: Analyzing political stances on Twitter in the lead-up to the 2024 U.S. election},
year = {2026},
howpublished = {\url{https://pith.science/paper/KPJFOR5P}},
note = {Machine review of arXiv:2412.02712}
}
read the original abstract
Social media platforms play a pivotal role in shaping public opinion and amplifying political discourse, particularly during elections. However, the same dynamics that foster democratic engagement can also exacerbate polarization. To better understand these challenges, here, we investigate the ideological positioning of tweets related to the 2024 U.S. Presidential Election. To this end, we analyze 1,235 tweets from key political figures and 63,322 replies, and classify ideological stances into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. Using a classification pipeline involving three large language models (LLMs)-GPT-4o, Gemini-Pro, and Claude-Opus-and validated by human annotators, we explore how ideological alignment varies between candidates and constituents. We find that Republican candidates author significantly more tweets in criticism of the Democratic party and its candidates than vice versa, but this relationship does not hold for replies to candidate tweets. Furthermore, we highlight shifts in public discourse observed during key political events. By shedding light on the ideological dynamics of online political interactions, these results provide insights for policymakers and platforms seeking to address polarization and foster healthier political dialogue.
Figures
Forward citations
Cited by 2 Pith papers
-
TikTok's recommendations skewed towards Republican content during the 2024 U.S. presidential race
TikTok's recommendation algorithm served Republican-seeded test accounts more co-partisan content than Democratic-seeded accounts during the 2024 U.S. presidential race.
-
Political-LLM: Large Language Models in Political Science
A survey and taxonomy of LLM applications in political science, with a case study suggesting that larger LLMs reproduce ANES 2016 voting patterns more accurately than smaller ones.
Reference graph
Works this paper leans on
-
[1]
Your stance is exposed! analysing possible factors for stance detection on social media
Aldayel, A., and Magdy, W. Your stance is exposed! analysing possible factors for stance detection on social media. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019), 1–20
work page 2019
-
[2]
Balasubramanian, A., Zou, V., Narayana, H., You, C., Luceri, L., and Ferrara, E. A public dataset tracking social media discourse about the 2024 us presidential election on twitter/x. arXiv preprint arXiv:2411.00376 (2024)
arXiv 2024
-
[3]
Facebooking it to the polls: A study in online social networking and political behavior
Bode, L. Facebooking it to the polls: A study in online social networking and political behavior. Journal of Information Technology & Politics 9 , 4 (2012)
work page 2012
-
[4]
Boutyline, A., and Willer, R. The social structure of political echo chambers: Variation in ideological homophily in online networks.Political psychology (2017)
work page 2017
-
[5]
The echo chamber effect on social media
Cinelli, M., De Francisci Morales, G., Galeazzi, A., Quattrociocchi, W., and Starnini, M. The echo chamber effect on social media. Proceedings of the National Academy of Sciences 118 , 9 (2021), e2023301118
work page 2021
-
[6]
Chatgpt outperforms crowd workers for text-annotation tasks
Gilardi, F., Alizadeh, M., and Kubli, M. Chatgpt outperforms crowd workers for text-annotation tasks. PNAS 120, 30 (2023), e2305016120
work page 2023
-
[7]
Youtube’s recom- mendation algorithm is left-leaning in the united states
Ibrahim, H., AlDahoul, N., Lee, S., Rahwan, T., and Zaki, Y. Youtube’s recom- mendation algorithm is left-leaning in the united states. PNAS nexus (2023)
work page 2023
-
[8]
Linegar, M., Kocielnik, R., and Alvarez, R. M. Large language models and political science. Frontiers in Political Science 5 (2023), 1257092
work page 2023
Show all 14 references
-
[9]
In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020)
Stefanov, P., Darwish, K., Atanasov, A., and Nakov, P.Predicting the topical stance and political leaning of media using tweets. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (2020)
2020
-
[10]
Large language models outperform expert coders and supervised classifiers at annotating political social media messages
Törnberg, P. Large language models outperform expert coders and supervised classifiers at annotating political social media messages. Social Science Computer Review (2024), 08944393241286471
2024
-
[11]
Y., Nagler, J., Tucker, J
Wu, P. Y., Nagler, J., Tucker, J. A., and Messing, S. Large language models can be used to scale the ideologies of politicians in a zero-shot learning setting, april
-
[12]
Auditing political exposure bias: Algorithmic amplification on twitter/x approaching the 2024 us presidential election
Ye, J., Luceri, L., and Ferrara, E. Auditing political exposure bias: Algorithmic amplification on twitter/x approaching the 2024 us presidential election. arXiv preprint arXiv:2411.01852 (2024)
2024 arXiv
-
[13]
Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., and Y ang, D.Can large lan- guage models transform computational social science? Computational Linguistics 50, 1 (2024), 237–291. 5
2024
-
[2023]
arXiv preprint arXiv:2303.12057 (2023)
2023 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.