REVIEW 3 major objections 4 minor 7 references
Analysis of Propaganda in Tweets From Politically Biased Sources
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Journalists at extremely biased outlets tweet more propaganda language than those at mildly biased outlets, per a new 1,874-tweet dataset.
desk verdict The JMBX dataset and LLM benchmark are real contributions; the central extreme-vs-mild finding needs a re-analysis without the official-account center category before it can be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the JMBX dataset: a balanced, expert-annotated collection of 1,874 tweets from five journalists at each of ten outlets, with outlet bias (left, lean left, center, lean right, right) assigned by distant supervision from a multi-partisan media bias rating service. Tweets are labeled using the 18 propaganda techniques defined in a widely used shared-task annotation scheme, and the central comparison is the proportion of propaganda-labeled tweets across bias categories. The same technique set also drives the LLM prompts in the later experiments, so the propaganda taxonomy is the common thread connecting the dataset claim, the detection experiments, and the cost analysis.
What would settle it
Collect tweets from individual journalists at centrist outlets using the same selection criteria, annotate them with the same 18-technique rubric, and compare the propaganda fraction with the left and right journalist groups; if the fraction is similar to the extreme groups, the paper's central claim fails.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is a relationship between institutional bias and individual language: in the JMBX dataset, the fraction of tweets labeled as propaganda is higher for journalists affiliated with outlets rated 'left' or 'right' than for those rated 'lean left', 'lean right', or 'center'. The paper further finds that the distribution of propaganda techniques is dominated by loaded language (48%), followed by exaggeration or minimization (21%) and name calling (9%), and that only centrist-outlet tweets show neutral sentiment in the sample. As a secondary discovery, the authors find that LLMs outperform a fine-tuned BERT model on propaganda detection, with chain-of-thought prompting yielding further gains on the tweet dataset but not consistently on the news-article dataset.
Load-bearing premise
The extreme-versus-mild comparison assumes that tweets from the official accounts of two centrist outlets are comparable to tweets from individual journalists at the other outlets, even though no journalist-level tweets were collected for the center category.
Editorial extensions
If this is right
- Outlet bias ratings could be used as a cheap proxy to flag journalists whose social media output is likely to contain propaganda, supporting media-monitoring tools.
- LLM-based propaganda detection, especially with chain-of-thought prompting, can serve as a scalable alternative to fine-tuned models on tweet-length text.
- The documented monetary and carbon costs per classification (about 0.44 kg CO2 for the full 200-tweet experiment) can inform decisions about deploying LLM detectors at scale.
- The annotated JMBX dataset provides a public benchmark for propaganda detection specifically in social media microblogs, complementing existing news-article corpora.
- Because loaded language dominates, detection systems and future interventions could focus on that single technique to capture most propaganda in tweets.
Reading between the lines
- The paper's center-baseline limitation implies a testable extension: if journalist-level tweets from centrist outlets were collected, the extreme-versus-mild gap might shrink, since the current comparison contrasts individual journalists with official institutional accounts.
- The finding that neutral sentiment appears only in centrist tweets could be an artifact of institutional accounts' editorial voice rather than of political centrism; analyzing journalist tweets from center outlets would separate these factors.
- The propaganda-rate ranking by bias could be confounded by outlet-specific editorial policies or beat assignments; re-annotating tweets by the same journalists on non-political topics would test whether the effect is about bias or topic.
- The LLM cost analysis implies that at scale, a full newsroom or platform monitoring campaign using the best zero-shot model could incur non-trivial emissions, so the reported per-task numbers offer a unit cost for further extrapolation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces JMBX, a dataset of 1,874 tweets from journalists affiliated with ten news outlets spanning the AllSides political-bias spectrum, annotated for 18 propaganda techniques with substantial inter-annotator agreement (Cohen's κ = 0.79). The authors address three research questions: (RQ1) whether journalists from extreme-biased outlets use more propaganda-like language than those from mildly biased outlets; (RQ2) how eight OpenAI and Google LLMs compare to a fine-tuned BERT model for propaganda detection in news articles and tweets, including chain-of-thought prompting; and (RQ3) the environmental cost of the LLM evaluations. The central claim is that journalists affiliated with extremely biased outlets are more likely to use propaganda-like language than those affiliated with mildly biased outlets; the paper also reports that LLMs generally outperform BERT, with CoT further improving performance on the JMBX dataset, and estimates a total carbon emission of about 0.44 kg for the LLM experiments.
Significance. The dataset and annotation effort are a useful contribution: JMBX is a new publicly available Twitter corpus with fine-grained propaganda labels, and the systematic comparison of eight LLMs against BERT on both news and social media text is a valuable benchmark, especially given the inclusion of Google Gemini models, which the authors note had not previously been tested for propaganda detection. The carbon-footprint estimate is a commendable addition that contextualizes the cost of LLM-based analysis. However, the primary scientific claim in the abstract and conclusion is currently not adequately supported because the RQ1 analysis compares journalist-authored tweets against official institutional accounts, and no statistical tests accompany the percentage differences. The strengths of the dataset and LLM evaluation are real, but the headline finding needs a substantially more careful analysis before it can be accepted.
major comments (3)
- [3.1, 3.3, Figure 3, Abstract, Conclusion] The central claim comparing 'extreme' versus 'mild' bias categories is confounded by author type. Section 3.1 states that for the center outlets (Forbes and Reuters), 'instead of individual journalists, the official Twitter handle ... are used to collect tweets and are labeled as political center.' All other categories consist of individual journalists' personal accounts. Figure 3 then shows the center category having the lowest propaganda proportion, and the abstract and conclusion assert that extreme-biased journalists use more propaganda than mild-biased ones. This comparison measures institutional-versus-personal authorship as much as political bias, because official accounts predominantly broadcast wire-style headlines and service updates, whereas individual journalists' accounts contain opinion and commentary. The Limitations section acknowledges that this 'complicates direct comparison,' but the headline claim still depends on the center baseline. The authors must either re-run RQ1 excluding the center category (comparing only extreme versus lean journalist accounts) or collect journalist-level tweets for center outlets; the conclusion should be revised accordingly.
- [3.3, Figure 3] No statistical test, confidence interval, or effect-size measure is reported for the key percentage differences in Figure 3. The claim that extreme categories have 'a higher proportion of propaganda' than mild categories is presented as a visual comparison of percentages, but with only ten outlets and 1,874 tweets, sampling variability could be substantial. The authors should report per-category counts and proportions, apply a chi-square test or logistic regression with outlet-level clustering, and provide confidence intervals for the propaganda rates. Without this, the central RQ1 result cannot be distinguished from noise.
- [3.2] The dataset construction is described as 'a balanced set of 2000 tweets, each with an equal number of instances featuring positive and negative sentiment scores,' and sentiment is known to correlate strongly with propaganda techniques such as loaded language and name-calling. If the final 1,874 tweets retain this sentiment-balancing scheme, the propaganda distribution across bias categories may be an artifact of the sampling design rather than a reflection of the underlying population of journalists' tweets. The authors should report the sentiment distribution per bias category in the final dataset (Figure 4 provides some information but not the per-category counts) and test whether the RQ1 result survives adjustment for sentiment. This is load-bearing because the selection procedure directly manipulates a variable that plausibly mediates the propaganda label.
minor comments (4)
- [3.1] The outlet name 'Breibart News' appears to be a typo for 'Breitbart News'; please correct it.
- [4.2, Tables 2 and 3] The model label '40613' in Tables 2 and 3 should be written as 'GPT-4-0613' (or 'GPT-4 0613') for clarity and to match the text's naming convention.
- [4.2] The sentence 'Results presented in Table 1 show that adding definition of propaganda increases the performance on all LLMs' appears to refer to Tables 2 and 3 (the LLM results), not Table 1 (the BERT results). Please correct the cross-reference.
- [References] Several references are inconsistently formatted, for example 'Barr´on-Cedeno et al. 2019' with an accented character inside a name field, and 'Baron, David P (2006)' missing the journal volume/page formatting used elsewhere. A consistent citation style would improve readability.
Circularity Check
No significant circularity: the central propaganda-rate claim is an empirical measurement, and the LLM comparison is externally benchmarked; the center-category confound is a validity issue, not a circularity.
full rationale
The central claim (Abstract, Section 3.3, Conclusion) that journalists affiliated with extremely biased outlets use more propaganda-like language than those with mild leans is an empirical finding derived from counts in the annotated JMBX dataset (Figure 3). The annotation pipeline is external: it uses the SemEval/18-technique guidelines and reports a Cohen's kappa of 0.79, so the labels are not obtained from the model or from the paper's own conclusion. The LLM propaganda-detection evaluation is benchmarked against the external PTC corpus, not against values fitted from JMBX, so the detection results do not reduce to their inputs by construction. The only self-cited work (Shokri et al. 2024) appears in a related-work sentence about using LLMs to identify subjective language and supplies no parameter, threshold, fitted value, or target result used in the headline finding; it is therefore not load-bearing. The acknowledged limitation that center-category tweets come from official Reuters and Forbes accounts rather than individual journalists is a real threat to the extreme-vs-mild comparison, but it is a data-validity and construct-validity concern, not a circularity: it does not make any prediction equivalent to an input by definition, and no fitted quantity is renamed as a prediction. No uniqueness theorem or ansatz is imported from prior work. Accordingly, the paper shows no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Journalists' political alignment is represented by the bias rating of their affiliated outlet, even when they tweet in a personal capacity.
- domain assumption AllSides bias ratings are an accurate ground-truth measure of outlet bias.
- domain assumption The 18 propaganda techniques of Martino et al. 2020 are the correct ontology, and the commercial annotator applied them consistently.
Cite this review
Pith. "Pith review of Analysis of Propaganda in Tweets From Politically Biased Sources." pith.science (2026). https://pith.science/paper/3LNVGAC2
@misc{pith2026250708169,
author = {Pith},
title = {Pith review of: Analysis of Propaganda in Tweets From Politically Biased Sources},
year = {2026},
howpublished = {\url{https://pith.science/paper/3LNVGAC2}},
note = {Machine review of arXiv:2507.08169}
}
read the original abstract
News outlets are well known to have political associations, and many national outlets cultivate political biases to cater to different audiences. Journalists working for these news outlets have a big impact on the stories they cover. In this work, we present a methodology to analyze the role of journalists, affiliated with popular news outlets, in propagating their bias using some form of propaganda-like language. We introduce JMBX(Journalist Media Bias on X), a systematically collected and annotated dataset of 1874 tweets from Twitter (now known as X). These tweets are authored by popular journalists from 10 news outlets whose political biases range from extreme left to extreme right. We extract several insights from the data and conclude that journalists who are affiliated with outlets with extreme biases are more likely to use propaganda-like language in their writings compared to those who are affiliated with outlets with mild political leans. We compare eight different Large Language Models (LLM) by OpenAI and Google. We find that LLMs generally performs better when detecting propaganda in social media and news article compared to BERT-based model which is fine-tuned for propaganda detection. While the performance improvements of using large language models (LLMs) are significant, they come at a notable monetary and environmental cost. This study provides an analysis of both the financial costs, based on token usage, and the environmental impact, utilizing tools that estimate carbon emissions associated with LLM operations.
Reference graph
Works this paper leans on
-
[1]
Azure, Microsoft (2024).Microsoft Azure PUE, WUE.https://azure.microsoft.com/en- us/blog/how- microsoft- measures- datacenter- water- and- energy- use- to-improve-azure-cloud-sustainability/[Accessed: (09/3/24)]. Baron, David P (2006). “Persistent media bias”. In:Journal of Public Economics90.1-2, pp. 1–36. Barr´on-Cedeno, Alberto et al. (2019). “Proppy: ...
arXiv 2024
-
[4]
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
Citeseer, pp. 1–49. Lin, Luyang et al. (2024). “Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception”. In:arXiv preprint arXiv:2403.14896. Liu, Ye et al. (2024). “Detect, Investigate, Judge and Determine: A Novel LLM-based Framework for Few-shot Fake News Detection”. In:arXiv preprint arXiv:2407.08952. Liu, Yinhan ...
arXiv 2024
-
[9]
Subjectivity Detection in English News using Large Language Models
Shokri, Mohammad et al. (Aug. 2024). “Subjectivity Detection in English News using Large Language Models”. In:Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sen- timent, & Social Media Analysis. Association for Computational Linguistics. Song, X Carol et al. (2022). “Anvil-System Architecture and Experiences from Deployment ...
arXiv 2022
-
[34]
ThatiAR: Subjectivity Detection in Arabic News Sentences
09, pp. 13693–13696. Suwaileh, Reem et al. (2024). “ThatiAR: Subjectivity Detection in Arabic News Sentences”. In:arXiv preprint arXiv:2406.05559. Vijayaraghavan, Prashanth and Soroush V osoughi (2022). “TWEETSPIN: Fine-grained propaganda de- tection in social media using multi-view representations”. In:Proceedings of the 2022 Conference 14 of the North A...
work page Pith review arXiv 2024
-
[454]
How open are journalists on Twitter? Trends towards the end-user journalism
Noguera-Vivo, Jos´e Manuel (2013). “How open are journalists on Twitter? Trends towards the end-user journalism”. In. Paul, Richard and Linda Elder (2006). “How to detect media bias & propaganda”. In:Dillon Beach, CA: Foundation for Critical Thinking. Radford, Alec et al. (2019). “Language models are unsupervised multitask learners”. In:OpenAI blog 1.8, p
work page 2013
-
[2020]
Energy and policy considerations for modern deep learning research
Accessed: 1/26/25.URL:https: / / www . statista . com / statistics / 1023881 / organized - social - media - manipulation-campaigns-worldwide/. Strubell, Emma, Ananya Ganesh, and Andrew McCallum (2020). “Energy and policy considerations for modern deep learning research”. In:Proceedings of the AAAI conference on artificial intelligence. V ol
work page 2020
-
[5646]
Prta: A system to support the analysis of propaganda tech- niques in the news
Da San Martino, Giovanni et al. (2020). “Prta: A system to support the analysis of propaganda tech- niques in the news”. In:Proceedings of the 58th Annual Meeting of the Association for Computa- tional Linguistics: System Demonstrations, pp. 287–293. Environmental Protection Agency (2024).EPA calculator.https://www.epa.gov/energy/ greenhouse-gas-equivalen...
arXiv 2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.