Pith. sign in

REVIEW 3 major objections 4 minor 24 references

Toxicity in State Sponsored Information Operations

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that toxic language in state-sponsored information operations is rare (1.53% of posts) but functions as an engagement amplifier, averaging 20.63 engagements per toxic post versus 3.46 for non-toxic posts, with Russian…

desk verdict A genuinely useful descriptive measurement of toxicity in influence operations, but the headline Russia engagement claim lacks the statistical support to back it. read the letter →

arxiv 2507.10936 v1 pith:E7QO7TYL submitted 2025-07-15 cs.SI

classification cs.SI
keywords ToxicContentInformationOperationsInfluenceOnlineSocialMediaAnalysisTwitter/XEngagementPerspectiveAPI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that state-linked influence operators deploy toxic language as a selective rhetorical tactic rather than a universal one, and that toxicity measurably amplifies engagement inside these campaigns. Analyzing 56 million posts from 42,405 accounts tied to 18 countries, it finds only 1.53% of posts are toxic, yet about a third of operators post at least one toxic item. Across the dataset, toxic posts earn roughly six times the engagement of non-toxic posts, and the paper's headline result is that Russian-origin toxic posts draw substantially more engagement than toxic posts from any other country in the data. If correct, the work matters because content moderation and detection systems would need to treat toxicity not just as a content violation but as a strategic amplifier that some state campaigns deliberately use.

What carries the argument

The central machinery is Google's Perspective API, a toxicity-scoring service applied to every preprocessed post; it returns probability scores for six attributes (toxicity, severe toxicity, identity attack, insult, profanity, and threat), and a post is marked toxic if any score exceeds 0.5. The other load-bearing object is X/Twitter's publicly released archive of state-sponsored information operations, which defines the population of 42,405 operators and 56 million posts. The engagement-amplifier interpretation rests on prior mechanisms, negativity bias and emotional arousal, which the paper cites to explain why toxic posts attract disproportionate sharing and interaction.

What would settle it

A native-speaker re-annotation of a multilingual random sample of posts (Russian, Spanish, Chinese, Arabic, Persian, and Bangla), followed by recomputed country-level engagement gaps with controls for language, topic, and post length, would settle the claim: if Russia's toxic-engagement advantage disappears after those controls, the paper's central result fails.

Watch

Extended reading notes

Core claim

The central discovery is that toxicity in state-sponsored information operations is rare at the post level but behaviorally concentrated and engagement-rich: 859,285 of 56,359,247 posts (1.53%) are toxic, but 14,200 of 42,405 operators (33.47%) posted at least one toxic post; toxic posts average 20.63 total engagements versus 3.46 for non-toxic posts (t = 94.11, p < 0.001). The paper further shows that toxic operators as a class out-engage non-toxic operators (2.72 versus 1.67 average engagements per post), and that Russia stands apart: 5.28% of Russian posts are toxic, 80.27% of Russian operators are toxic posters, and Russian toxic posts average 86.08 engagements versus 27.88 for non-toxic Russian posts. This pattern is presented as evidence that toxicity is deployed strategically by some countries, notably Russia, Venezuela, and Cuba, while other large operations such as China's and Saudi Arabia's keep toxicity low despite high volume.

Load-bearing premise

The whole cross-country comparison depends on Perspective API toxicity scores being valid and comparable across all languages in the data, yet the manual validation covers only 500 English posts and the paper itself notes the API is biased against some non-English languages.

Editorial extensions

If this is right

  • Platforms should treat toxicity in state-linked accounts as a signal of coordinated influence rather than isolated abuse, because a small toxic minority can dominate engagement.
  • Detection systems should weigh geopolitical context, since Russia, Cuba, and Venezuela show high concentrations of toxic operators while China and Saudi Arabia run large operations with low toxicity.
  • Moderation policies that suppress toxic content would directly cut the main engagement advantage that these campaigns appear to exploit.
  • Cross-national toxicity comparisons need language-aware scoring, because the paper itself notes that the API's language bias can distort non-English rates.
  • The sixfold engagement gap implies that toxic posts are disproportionately visible, so studying only aggregate engagement may understate the reach of harmful content.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because toxicity is concentrated in specific countries, a natural extension would be to test whether spikes in toxic posting align with known geopolitical events or election cycles, which this paper does not examine.
  • Comparing toxic-post engagement inside state-run operations with organic toxic posts would clarify whether the amplifier effect is specific to coordinated campaigns or simply a property of toxicity in general.
  • The 1.53% base rate combined with the sixfold engagement gap suggests an asymmetry worth quantifying: a small toxic budget buys a disproportionate share of impressions, which could be tested as a per-operator rate-versus-reach trade-off.
  • Releasing per-post toxicity scores and multilingual rescoring would let others verify whether Russia's engagement advantage survives language and topic controls.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper analyzes the X/Twitter state-sponsored information operations archive (56 million posts, 42,405 operators, 18 geopolitical entities), labeling posts as toxic or non-toxic with Google's Perspective API across six attributes and a 0.5 score threshold. It reports that 1.53% of all posts are toxic, that about one-third of operators posted toxic content at least once, and that toxic posts receive higher average engagement than non-toxic posts (20.63 vs. 3.46 likes/reposts/replies combined). The authors further report country-level variation, highlighting Russia as having the highest toxic-post proportion (5.28%) and the largest toxic engagement gap (86.08 vs. 27.88). The abstract's headline claim is that toxic content from Russian influence operations receives significantly higher engagement than toxic content from any other country. The paper includes manual validation of 500 English posts, with Cohen's kappa around 0.79-0.85 against the API, and releases code on GitHub.

Significance. If the central claims were established, the paper would be a useful large-scale descriptive contribution: it applies a standard toxicity measurement to the complete public IO archive, reports engagement differences, and makes code available. The manual annotation agreement, the use of the full X/Twitter disclosure dataset, and the reproducibility-oriented release are concrete strengths. However, the headline cross-country claim about Russia is not supported by the analysis as reported. The pooled toxic-versus-non-toxic contrast is confounded with country baselines, and the cross-country comparison rests on toxicity labels whose language invariance is both questioned in Section 6 and validated only on English. The paper's current value is therefore mostly descriptive and pooled; the Russia-specific engagement claim requires substantially stronger analysis before it can be accepted.

major comments (3)
  1. [Section 5 (RQ2) and Abstract] The abstract's central claim that toxic content from Russian influence operations receives significantly higher user engagement than influence operations from any other country is not backed by the reported statistics. Section 5 gives a pooled t-test across all posts (t=94.11) and then lists Russia's toxic mean 86.08 versus its non-toxic mean 27.88, but it provides no pairwise comparison of Russia against any other country, no country-by-toxicity interaction test, and no regression with country fixed effects. This matters because Russia's non-toxic mean (27.88) already exceeds the overall toxic mean (20.63), so the observed Russian 'advantage' may reflect baseline engagement differences (audience, language, campaign targeting, account characteristics) rather than an amplifying effect of toxicity. The authors should add a model with country fixed effects and a country-by-toxicity interaction, or at minimum pairwise tests restricted to toxic posts with appropriate multiple-comparison correction.
  2. [Sections 4 and 6] The cross-country comparison requires that Perspective API toxicity scores are valid and comparable across the languages in the dataset, and that assumption is load-bearing for the Russia claim. Section 6 acknowledges that Persian, Bangla, and other unsupported languages were excluded and that the API has documented systematic language bias (reference [14]), while Section 4 reports manual validation only on 500 English posts. Because country-level toxicity rates and engagement comparisons mix languages with different API behavior, the observed ordering of countries could be an artifact of language bias rather than of strategic toxicity use. The authors should provide a multilingual validation subsample or restrict the cross-country engagement claim to languages for which measurement invariance has been demonstrated, with a sensitivity analysis.
  3. [Section 5 (engagement t-tests)] The two-sample t-tests reported for engagement (t=94.11 and t=2.62) treat individual posts or operators as independent observations, but posts are nested within operators and within campaigns, and with roughly 56 million posts near any nonzero mean difference will produce a small p-value. The reported significance levels therefore overstate the practical weight of the pooled differences. The authors should report cluster-robust standard errors, a mixed-effects model, or effect sizes (e.g., Cohen's d or explained variance) so that the engagement contrast can be evaluated on a meaningful scale.
minor comments (4)
  1. [Table 1 caption] The caption says 'On average, Russian operators spread the most toxic content (5.28%),' but 'on average' is ambiguous here since the table reports a percentage of posts; moreover Russia (5.28%) and Mexico (5.24%) are nearly tied, and no statistical test is given for that ordering.
  2. [Section 2 and Abstract] The paper calls itself 'the first comprehensive analysis' and 'the first large-scale, cross-national analysis,' but reference [18] is the authors' own prior study of sentiment, emotion, and hate speech in state-sponsored influence operations; the authors should clarify the specific incremental contribution over that prior work.
  3. [Section 5 (engagement metric)] The engagement metric is defined only implicitly as the sum of likes, reposts, and replies; the text should state explicitly whether quote tweets, views, or other interaction types were included, since this affects comparability with other engagement studies.
  4. [Section 5 (country-level exceptions)] The paragraph noting that toxicity is not universally effective (Honduras, Saudi Arabia, Catalonia) gives no numerical values; adding a small table or explicitly citing Figure 3 values would make the claim verifiable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the engagement and toxicity results are direct empirical measurements, not derived from fitted parameters or self-cited premises.

full rationale

The paper's central claims are empirical measurements: posts are labeled toxic by Google's Perspective API using a fixed 0.5 threshold (Section 4), and engagements are taken directly from X/Twitter's archived influence-operation dataset (Section 3). The reported statistics, such as the pooled toxic mean of 20.63 engagements versus 3.46 for non-toxic posts and the t-test (t=94.11), are direct comparisons of measured quantities; no parameter is fitted to a subset and then renamed as a prediction, and no equation defines one reported quantity in terms of another. The abstract's Russia-specific claim is likewise a descriptive comparison of means (Russia toxic mean 86.08 versus non-toxic mean 27.88), not a model output. Whether that comparison is statistically validated or confounded by country-level baseline engagement is a correctness and robustness concern, not a circularity concern. The paper cites the authors' own prior work [18] and [16], but these citations appear only as background or interpretive context, not as load-bearing inputs to the toxicity labels, engagement counts, or any computed test statistic. Hence there is no circular derivation chain to flag.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities or fitted theoretical parameters. The only adjustable choice is the toxicity threshold; the main assumptions concern the completeness of the platform dataset, the cross-lingual validity of the Perspective API, and the statistical model used for engagement comparisons.

free parameters (1)
  • toxicity threshold = 0.5 (applied to any of six Perspective attributes)
    A post is labeled toxic if any of six Perspective scores exceeds 0.5. This cutoff is chosen without sensitivity analysis, but all counts and engagement comparisons depend on it.
assumptions (4)
  • domain assumption The X/Twitter archive of state-sponsored information operations is a complete and representative enumeration of such accounts and posts for 2018-2021.
    The analysis relies entirely on X's public disclosure; the paper does not independently validate coverage or completeness. Invoked implicitly in Section 3.
  • domain assumption Perspective API toxicity scores are valid, comparable measures of toxicity across all languages in the dataset, despite known language biases and the English-only validation.
    Section 4 uses API scores for all posts; Section 6 admits language limitations and bias (citing ref [14]) but does not correct for them.
  • domain assumption Engagement metrics (likes, reposts, replies) are a valid measure of amplification and audience response for comparing toxic vs non-toxic posts.
    RQ2 operationalizes user engagement exclusively as these three counts, Section 5.
  • domain assumption Posts are independent units for the reported t-tests.
    Two-sample t-tests in Section 5 treat posts as independent despite clear nesting within accounts and countries; this assumption is violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toxicity in State Sponsored Information Operations." pith.science (2026). https://pith.science/paper/E7QO7TYL

@misc{pith2026250710936,
  author       = {Pith},
  title        = {Pith review of: Toxicity in State Sponsored Information Operations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E7QO7TYL}},
  note         = {Machine review of arXiv:2507.10936}
}
read the original abstract

State-sponsored information operations (IOs) increasingly influence global discourse on social media platforms, yet their emotional and rhetorical strategies remain inadequately characterized in scientific literature. This study presents the first comprehensive analysis of toxic language deployment within such campaigns, examining 56 million posts from over 42 thousand accounts linked to 18 distinct geopolitical entities on X/Twitter. Using Google's Perspective API, we systematically detect and quantify six categories of toxic content and analyze their distribution across national origins, linguistic structures, and engagement metrics, providing essential information regarding the underlying patterns of such operations. Our findings reveal that while toxic content constitutes only 1.53% of all posts, they are associated with disproportionately high engagement and appear to be strategically deployed in specific geopolitical contexts. Notably, toxic content originating from Russian influence operations receives significantly higher user engagement compared to influence operations from any other country in our dataset. Our code is available at https://github.com/shafin191/Toxic_IO.

Figures

Figures reproduced from arXiv: 2507.10936 by the authors.

Figure 1
Figure 1. Architecture for Toxicity Analysis of State [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Comparison of post length between toxic and non-toxic posts across countries. In most countries, toxic posts tend to [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Comparison of post engagement between toxic and non-toxic posts across countries. On average, toxic posts received [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages

  1. [18]

    Ashfaq Ali Shafin and Khandaker Mamun Ahmed. 2025. The Language of In- fluence: Sentiment, Emotion, and Hate Speech in State Sponsored Influence Operations. arXiv:2505.07212 [cs.SI] https://arxiv.org/abs/2505.07212

  2. [14]

    Gianluca Nogara, Francesco Pierri, Stefano Cresci, Luca Luceri, Petter Törnberg, and Silvia Giordano. 2025. Toxic Bias: Perspective API Misreads German as More Toxic. Proceedings of the International AAAI Conference on Web and Social Media 19, 1 (Jun. 2025), 1346–1357. doi:10.1609/icwsm.v19i1.35876

  3. [1]

    Perspective API

    2017. Perspective API. https://www.perspectiveapi.com. Accessed: 2025-01-06

  4. [2]

    Roy F Baumeister, Ellen Bratslavsky, Catrin Finkenauer, and Kathleen D Vohs

  5. [3]

    George Beknazar-Yuzbashev, Rafael Jiménez-Durán, Jesse McCrosky, and Ma- teusz Stalinski. 2022. Toxic Content and User Engagement on Social Me- dia: Evidence from a Field Experiment. A vailable at SSRN (November 2022). doi:10.2139/ssrn.4307346

  6. [4]

    Jonah Berger and Katherine L. Milkman. 2012. What Makes Online Content Viral? Journal of Marketing Research 49, 2 (2012), 192–205. doi:10.1509/jmr.10.0353

  7. [5]

    2018.The Tactics & Tropes of the Internet Research Agency

    Renee DiResta, Kris Shaffer, Becky Ruppel, David Sullivan, Robert Matney, Ryan Fox, Jonathan Albright, and Ben Johnson. 2018.The Tactics & Tropes of the Internet Research Agency. Technical Report. https://digitalcommons.unl.edu/senatedocs/ 2/

  8. [6]

    Dominique Geissler, Dominik Bär, Nicolas Pröllochs, and Stefan Feuerriegel. 2023. Russian propaganda on social media during the 2022 invasion of Ukraine. EPJ Data Science 12, 1 (2023), 35

Show all 24 references
  1. [7]

    Linvill and Patrick L

    Darren L. Linvill and Patrick L. Warren. 2020. Troll Factories: Man- ufacturing Specialized Disinformation on Twitter. Political Communica- tion 37, 4 (2020), 447–467. doi:1 0 . 1 0 8 0 / 1 0 5 8 4 6 0 9 . 2 0 2 0 . 1 7 1 8 2 5 7 arXiv:https://doi.org/10.1080/10584609.2020.1718257

  2. [8]

    Luca Luceri, Silvia Giordano, and Emilio Ferrara. 2020. Detecting Troll Behavior via Inverse Reinforcement Learning: A Case Study of Russian Trolls in the 2016 US Election. Proceedings of the International AAAI Conference on Web and Social Media 14, 1 (May 2020), 417–427. doi:...

  3. [9]

    Luca Luceri, Valeria Pantè, Keith Burghardt, and Emilio Ferrara. 2024. Unmasking the Web of Deceit: Uncovering Coordinated Activity to Expose Information Operations on Twitter. InProceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Compu...

  4. [10]

    Abdurahman Maarouf, Nicolas Pröllochs, and Stefan Feuerriegel. 2024. The Virality of Hate Speech on Social Media. Proc. ACM Hum.-Comput. Interact. 8, CSCW1, Article 186 (April 2024), 22 pages. doi:10.1145/3641025

  5. [11]

    Binny Mathew, Ritam Dutt, Pawan Goyal, and Animesh Mukherjee. 2019. Spread of Hate Speech in Online Social Media. InProceedings of the 10th ACM Conference on Web Science (Boston, Massachusetts, USA) (WebSci ’19). Association for Com- puting Machinery, New York, NY, USA, 173–18...

  6. [12]

    Marco Minici, Luca Luceri, Francesco Fabbri, and Emilio Ferrara. 2025. IOHunter: Graph Foundation Model to Uncover Online Information Operations.Proceedings of the AAAI Conference on Artificial Intelligence 39, 27 (Apr. 2025), 28258–28266. doi:10.1609/aaai.v39i27.35046

  7. [13]

    Mainack Mondal, Leandro Araújo Silva, and Fabrício Benevenuto. 2017. A Measurement Study of Hate Speech in Social Media. In Proceedings of the 28th ACM Conference on Hypertext and Social Media (Prague, Czech Republic) (HT ’17) . Association for Computing Machinery, New York, N...

  8. [15]

    Diogo Pacheco, Alessandro Flammini, and Filippo Menczer. 2020. Unveiling Coor- dinated Groups Behind White Helmets Disinformation. InCompanion Proceedings of the Web Conference 2020 (Taipei, Taiwan) (WWW ’20). Association for Com- puting Machinery, New York, NY, USA, 611–616. ...

  9. [16]

    Ruben Recabarren, Bogdan Carbunar, Nestor Hernandez, and Ashfaq Ali Shafin

  10. [17]

    Yoel Roth. 2019. Information operations on Twitter: principles, process, and disclosure. https://blog.x.com/en_us/topics/company/2019/information-ops-on- twitter. Accessed: April - 29 - 2025

  11. [19]

    Ali Siddiquee

    Md. Ali Siddiquee. 2020. The portrayal of the Rohingya genocide and refugee crisis in the age of post-truth politics. Asian Journal of Com- parative Politics 5, 2 (2020), 89–103. doi:10.1177/2057891119864454 arXiv:https://doi.org/10.1177/2057891119864454

  12. [20]

    Luis Vargas, Patrick Emami, and Patrick Traynor. 2020. On the Detection of Disinformation Campaign Activity with Network Analysis. In Proceedings of the 2020 ACM SIGSAC Conference on Cloud Computing Security Workshop (Virtual Event, USA) (CCSW’20). Association for Computing Ma...

  13. [21]

    Padinjaredath Suresh Vishnuprasad, Gianluca Nogara, Felipe Cardoso, Stefano Cresci, Silvia Giordano, and Luca Luceri. 2024. Tracking Fringe and Coordinated Activity on Twitter Leading Up to the US Capitol Attack. Proceedings of the International AAAI Conference on Web and Soci...

  14. [1570]

    doi:10.1609/icwsm.v18i1.31409

  15. [2001]

    Review of General Psychology 5, 4 (2001), 323–370

    Bad is stronger than good. Review of General Psychology 5, 4 (2001), 323–370. doi:10.1037/1089-2680.5.4.323

  16. [2023]

    In 32nd USENIX Security Symposium (USENIX Security 23)

    Strategies and vulnerabilities of participants in venezuelan influence operations. In 32nd USENIX Security Symposium (USENIX Security 23) . 6683– 6700

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.