{"id":"72f32bdb-2f70-4f3f-8434-d5eddc5cd604","arxiv_id":"2507.10936","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Analyzing 56 million posts from 42,405 state-linked accounts, toxic posts (1.53% of the total) drew about 6 times more engagement than non-toxic posts, with Russian operations showing the largest effect.","lead":"The paper measures toxic language in 56 million posts from state-sponsored influence operations on X/Twitter, finding that toxic posts are rare but attract more engagement than non-toxic posts. It claims Russian operations receive the most engagement on their toxic content, and offers the dataset and code for verification.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Russia engagement claim lacks a controlled cross-country comparison; country baseline engagement may drive it.","rationale":"I read the paper as a descriptive measurement study. The dataset is a valuable public artifact; the manual validation (500 English posts, κ=0.81) is a real check. However, the headline claim in the abstract goes beyond description: it asserts a Russia-specific engagement premium for toxic posts. The reported evidence is a set of unadjusted means. Russia's non-toxic mean (27.88) being higher than the pooled toxic mean (20.63) makes the confounding concern concrete and falsifiable from the paper's own numbers. The reader's language-bias concern is valid and acknowledged (Section 6, ref [14]), but it affects toxicity prevalence rates more than the engagement comparison's internal validity. Language bias would mainly shift which posts are labeled toxic; the engagement premium claim is more directly threatened by the absence of any baseline or comparison control. Hence I partially agree with the reader. The required fix is feasible with the public dataset, so conditional acceptance remains appropriate.","tokens_in":7892,"tokens_out":7119,"duration_ms":83509,"concrete_test":"Using the released dataset and code, compute per-country mean engagement for toxic and non-toxic posts. Fit a negative binomial or OLS model of engagement on toxicity indicator + country fixed effects + toxicity×country interaction, with standard errors clustered by account. Test (a) whether Russia's toxicity interaction coefficient is significantly larger than every other country's interaction coefficient (pairwise Wald tests with multiple-comparison correction), and (b) whether Russia toxic mean remains significantly higher than each other country's toxic mean in a bootstrap comparison restricted to a language-matched subset (e.g., English-only posts). If (a) or (b) fails, the abstract's Russia-specific claim should be weakened to a within-country association.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim—that toxic content from Russian influence operations receives significantly higher user engagement than any other country—is not supported by the reported analysis. Section 5 reports only pooled toxic vs non-toxic means (20.63 vs 3.46) and a t-test across all posts (t=94.11). For the Russia-specific claim, it gives Russia's toxic mean 86.08 and non-toxic mean 27.88, but no pairwise comparison against other countries and no test statistic. Critically, Russia's non-toxic mean (27.88) exceeds the overall toxic mean (20.63), so Russian IO posts attract high engagement regardless of toxicity. The observed Russia 'advantage' may therefore reflect country-level baseline engagement (audience, language, campaign targeting, account age or follower count) rather than an amplifying effect of toxicity. Without a model that includes country fixed effects, or at minimum a country × toxicity interaction test, the claim is confounded. This concern is independent of the Perspective API language-bias issue: even with perfect toxicity labels, the cross-country comparison as reported cannot establish that Russia's toxic content out-engages other countries' toxic content after controlling for baseline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes the X/Twitter state-sponsored information operations archive (56 million posts, 42,405 operators, 18 geopolitical entities), labeling posts as toxic or non-toxic with Google's Perspective API across six attributes and a 0.5 score threshold. It reports that 1.53% of all posts are toxic, that about one-third of operators posted toxic content at least once, and that toxic posts receive higher average engagement than non-toxic posts (20.63 vs. 3.46 likes/reposts/replies combined). The authors further report country-level variation, highlighting Russia as having the highest toxic-post proportion (5.28%) and the largest toxic engagement gap (86.08 vs. 27.88). The abstract's headline claim is that toxic content from Russian influence operations receives significantly higher engagement than toxic content from any other country. The paper includes manual validation of 500 English posts, with Cohen's kappa around 0.79-0.85 against the API, and releases code on GitHub.","tokens_in":8079,"tokens_out":3145,"duration_ms":40004,"significance":"If the central claims were established, the paper would be a useful large-scale descriptive contribution: it applies a standard toxicity measurement to the complete public IO archive, reports engagement differences, and makes code available. The manual annotation agreement, the use of the full X/Twitter disclosure dataset, and the reproducibility-oriented release are concrete strengths. However, the headline cross-country claim about Russia is not supported by the analysis as reported. The pooled toxic-versus-non-toxic contrast is confounded with country baselines, and the cross-country comparison rests on toxicity labels whose language invariance is both questioned in Section 6 and validated only on English. The paper's current value is therefore mostly descriptive and pooled; the Russia-specific engagement claim requires substantially stronger analysis before it can be accepted.","major_comments":[{"comment":"The abstract's central claim that toxic content from Russian influence operations receives significantly higher user engagement than influence operations from any other country is not backed by the reported statistics. Section 5 gives a pooled t-test across all posts (t=94.11) and then lists Russia's toxic mean 86.08 versus its non-toxic mean 27.88, but it provides no pairwise comparison of Russia against any other country, no country-by-toxicity interaction test, and no regression with country fixed effects. This matters because Russia's non-toxic mean (27.88) already exceeds the overall toxic mean (20.63), so the observed Russian 'advantage' may reflect baseline engagement differences (audience, language, campaign targeting, account characteristics) rather than an amplifying effect of toxicity. The authors should add a model with country fixed effects and a country-by-toxicity interaction, or at minimum pairwise tests restricted to toxic posts with appropriate multiple-comparison correction.","section":"Section 5 (RQ2) and Abstract"},{"comment":"The cross-country comparison requires that Perspective API toxicity scores are valid and comparable across the languages in the dataset, and that assumption is load-bearing for the Russia claim. Section 6 acknowledges that Persian, Bangla, and other unsupported languages were excluded and that the API has documented systematic language bias (reference [14]), while Section 4 reports manual validation only on 500 English posts. Because country-level toxicity rates and engagement comparisons mix languages with different API behavior, the observed ordering of countries could be an artifact of language bias rather than of strategic toxicity use. The authors should provide a multilingual validation subsample or restrict the cross-country engagement claim to languages for which measurement invariance has been demonstrated, with a sensitivity analysis.","section":"Sections 4 and 6"},{"comment":"The two-sample t-tests reported for engagement (t=94.11 and t=2.62) treat individual posts or operators as independent observations, but posts are nested within operators and within campaigns, and with roughly 56 million posts near any nonzero mean difference will produce a small p-value. The reported significance levels therefore overstate the practical weight of the pooled differences. The authors should report cluster-robust standard errors, a mixed-effects model, or effect sizes (e.g., Cohen's d or explained variance) so that the engagement contrast can be evaluated on a meaningful scale.","section":"Section 5 (engagement t-tests)"}],"minor_comments":[{"comment":"The caption says 'On average, Russian operators spread the most toxic content (5.28%),' but 'on average' is ambiguous here since the table reports a percentage of posts; moreover Russia (5.28%) and Mexico (5.24%) are nearly tied, and no statistical test is given for that ordering.","section":"Table 1 caption"},{"comment":"The paper calls itself 'the first comprehensive analysis' and 'the first large-scale, cross-national analysis,' but reference [18] is the authors' own prior study of sentiment, emotion, and hate speech in state-sponsored influence operations; the authors should clarify the specific incremental contribution over that prior work.","section":"Section 2 and Abstract"},{"comment":"The engagement metric is defined only implicitly as the sum of likes, reposts, and replies; the text should state explicitly whether quote tweets, views, or other interaction types were included, since this affects comparability with other engagement studies.","section":"Section 5 (engagement metric)"},{"comment":"The paragraph noting that toxicity is not universally effective (Honduras, Saudi Arabia, Catalonia) gives no numerical values; adding a small table or explicitly citing Figure 3 values would make the claim verifiable.","section":"Section 5 (country-level exceptions)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's real contribution is descriptive and reproducible, but the headline Russia claim is currently unsupported by the reported analysis and the measurement-comparability issue is central rather than peripheral. I would consider a revised version acceptable if the authors add a proper country-by-toxicity analysis with fixed effects or pairwise tests and address the multilingual validity of the toxicity labels; otherwise the abstract should be weakened to what the current evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real first pass at measuring toxic language across 18 state-sponsored influence operations, with code released and an honest limitations section, but the banner claim about Russian engagement is not supported by the reported analysis. The pooled toxic-vs-non-toxic engagement gap (20.63 vs 3.46, t=94.11) is robust as far as it goes, and the per-country descriptive rates are a useful artifact. The manual validation of 500 English posts (kappa 0.79–0.85) is careful, and reporting six toxicity attributes separately is a plus.\n\nThe soft spots are the cross-country comparisons. First, the Perspective API language bias is acknowledged in Section 6 but not mitigated; excluding Persian, Bangla, and other unsupported languages means the country-level toxicity rates rest on non-representative language samples. Second, and more serious, the Russia claim has no pairwise test. Russia's toxic mean is 86.08 versus 27.88 for non-toxic, but Russia's non-toxic mean already exceeds the overall toxic mean of 20.63. Without a country fixed-effect model or at least a country×toxicity interaction, the engagement premium could be baseline country engagement, audience, or campaign targeting rather than an effect of toxicity. The paper even notes countries where non-toxic posts outperform toxic ones, so toxicity is not uniformly effective; that makes the Russia claim more, not less, in need of a proper statistical contrast. The stress-test note is right on this.\n\nMinor point: the abstract says \"significantly higher\" without reporting a test for the specific country comparison; the only t-test reported is pooled across all posts. That is fixable but a real gap. The overlap with the authors' prior work [18] on the same dataset is handled honestly as related work, and the incremental contribution—multi-attribute toxicity profiling and engagement linkage—is fair.\n\nBottom line: a useful descriptive study with an honest limitations section, but the central comparative claim outruns the evidence. Who is this for? Researchers studying IO rhetoric and platform moderation; they will find the per-country descriptive rates and the code useful. It deserves peer review, but I would expect heavy revision: add pairwise significance tests or a regression with country fixed effects, do language-specific validation on a sample, and either soften the Russia claim or support it properly. I would not desk-reject.","headline":"A genuinely useful descriptive measurement of toxicity in influence operations, but the headline Russia engagement claim lacks the statistical support to back it.","tokens_in":8592,"tokens_out":1901,"would_cite":false,"duration_ms":24302,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that toxic language in state-sponsored information operations is rare (1.53% of posts) but functions as an engagement amplifier, averaging 20.63 engagements per toxic post versus 3.46 for non-toxic posts, with Russian…","keywords":["Toxic Content","Information Operations","Influence Operations","Online Social Media","Social Media Analysis","Twitter/X","Engagement","Perspective API"],"falsifier":"A native-speaker re-annotation of a multilingual random sample of posts (Russian, Spanish, Chinese, Arabic, Persian, and Bangla), followed by recomputed country-level engagement gaps with controls for language, topic, and post length, would settle the claim: if Russia's toxic-engagement advantage disappears after those controls, the paper's central result fails.","tokens_in":7691,"feed_emoji":"☠️","tokens_out":4866,"duration_ms":53161,"temperature":0.7,"pith_summary":"The paper tries to establish that state-linked influence operators deploy toxic language as a selective rhetorical tactic rather than a universal one, and that toxicity measurably amplifies engagement inside these campaigns. Analyzing 56 million posts from 42,405 accounts tied to 18 countries, it finds only 1.53% of posts are toxic, yet about a third of operators post at least one toxic item. Across the dataset, toxic posts earn roughly six times the engagement of non-toxic posts, and the paper's headline result is that Russian-origin toxic posts draw substantially more engagement than toxic posts from any other country in the data. If correct, the work matters because content moderation and detection systems would need to treat toxicity not just as a content violation but as a strategic amplifier that some state campaigns deliberately use.","feed_headline":"Toxic posts in state influence ops get 6x the engagement","feed_subtitle":"Study of 56 million X posts: 1.53 percent are toxic, but they dominate likes and reposts—and Russia leads.","key_machinery":"The central machinery is Google's Perspective API, a toxicity-scoring service applied to every preprocessed post; it returns probability scores for six attributes (toxicity, severe toxicity, identity attack, insult, profanity, and threat), and a post is marked toxic if any score exceeds 0.5. The other load-bearing object is X/Twitter's publicly released archive of state-sponsored information operations, which defines the population of 42,405 operators and 56 million posts. The engagement-amplifier interpretation rests on prior mechanisms, negativity bias and emotional arousal, which the paper cites to explain why toxic posts attract disproportionate sharing and interaction.","core_discovery":"The central discovery is that toxicity in state-sponsored information operations is rare at the post level but behaviorally concentrated and engagement-rich: 859,285 of 56,359,247 posts (1.53%) are toxic, but 14,200 of 42,405 operators (33.47%) posted at least one toxic post; toxic posts average 20.63 total engagements versus 3.46 for non-toxic posts (t = 94.11, p < 0.001). The paper further shows that toxic operators as a class out-engage non-toxic operators (2.72 versus 1.67 average engagements per post), and that Russia stands apart: 5.28% of Russian posts are toxic, 80.27% of Russian operators are toxic posters, and Russian toxic posts average 86.08 engagements versus 27.88 for non-toxic Russian posts. This pattern is presented as evidence that toxicity is deployed strategically by some countries, notably Russia, Venezuela, and Cuba, while other large operations such as China's and Saudi Arabia's keep toxicity low despite high volume.","pith_inferences":["Because toxicity is concentrated in specific countries, a natural extension would be to test whether spikes in toxic posting align with known geopolitical events or election cycles, which this paper does not examine.","Comparing toxic-post engagement inside state-run operations with organic toxic posts would clarify whether the amplifier effect is specific to coordinated campaigns or simply a property of toxicity in general.","The 1.53% base rate combined with the sixfold engagement gap suggests an asymmetry worth quantifying: a small toxic budget buys a disproportionate share of impressions, which could be tested as a per-operator rate-versus-reach trade-off.","Releasing per-post toxicity scores and multilingual rescoring would let others verify whether Russia's engagement advantage survives language and topic controls."],"forward_implications":["Platforms should treat toxicity in state-linked accounts as a signal of coordinated influence rather than isolated abuse, because a small toxic minority can dominate engagement.","Detection systems should weigh geopolitical context, since Russia, Cuba, and Venezuela show high concentrations of toxic operators while China and Saudi Arabia run large operations with low toxicity.","Moderation policies that suppress toxic content would directly cut the main engagement advantage that these campaigns appear to exploit.","Cross-national toxicity comparisons need language-aware scoring, because the paper itself notes that the API's language bias can distort non-English rates.","The sixfold engagement gap implies that toxic posts are disproportionately visible, so studying only aggregate engagement may understate the reach of harmful content."],"supporting_citations":[{"why":"Supplies the X/Twitter state-sponsored operations archive that defines the population of 42,405 operators and 56 million posts.","marker":"[17]"},{"why":"Provides the Perspective API that generates the six toxicity scores used as the paper's dependent variable.","marker":"[1]"},{"why":"Documents the API's systematic cross-language bias, which the paper's limitations section relies on.","marker":"[14]"},{"why":"Offers field-experiment evidence that toxic content drives user engagement, supporting the amplifier interpretation.","marker":"[3]"},{"why":"Shows that toxic posts from verified accounts tend to spread farther, providing the baseline the paper extends to state operators.","marker":"[10]"},{"why":"Documents Russian Internet Research Agency tactics, used to explain Russia's high-toxicity strategy.","marker":"[5]"},{"why":"Characterizes specialized Russian troll disinformation, supporting the strategic-toxicity reading.","marker":"[7]"}],"fun_headline_variants":["Toxic posts in state ops get 6x engagement, Russia leads","Rare toxic posts from state influence ops dominate engagement","Russia's toxic posts in influence ops drive 6x engagement","1.53% of state influence posts are toxic, but get 6x engagement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole cross-country comparison depends on Perspective API toxicity scores being valid and comparable across all languages in the data, yet the manual validation covers only 500 English posts and the paper itself notes the API is biased against some non-English languages.","fun_headline_variants_meta":{"raw":{"variants":["Toxic posts in state ops get 6x engagement, Russia leads","Rare toxic posts from state influence ops dominate engagement","Russia's toxic posts in influence ops drive 6x engagement","1.53% of state influence posts are toxic, but get 6x engagement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1722,"prompt_tokens":932,"completion_tokens":790,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":714}},"tokens_in":548,"tokens_out":790,"duration_ms":8831,"temperature":1.0,"reasoning_tokens":714,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:20:10.558257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A native-speaker re-annotation of a multilingual random sample of posts (Russian, Spanish, Chinese, Arabic, Persian, and Bangla), followed by recomputed country-level engagement gaps with controls for language, topic, and post length, would settle the claim: if Russia's toxic-engagement advantage disappears after those controls, the paper's central result fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the X/Twitter state-sponsored operations archive that defines the population of 42,405 operators and 56 million posts."},{"cited_title":"Perspective API","cited_arxiv_id":null,"evidence_quote":"Provides the Perspective API that generates the six toxicity scores used as the paper's dependent variable."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the API's systematic cross-language bias, which the paper's limitations section relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers field-experiment evidence that toxic content drives user engagement, supporting the amplifier interpretation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that toxic posts from verified accounts tend to spread farther, providing the baseline the paper extends to state operators."},{"cited_title":"2018.The Tactics & Tropes of the Internet Research Agency","cited_arxiv_id":null,"evidence_quote":"Documents Russian Internet Research Agency tactics, used to explain Russia's high-toxicity strategy."}],"review_version":1}