Pith. sign in

REVIEW 3 major objections 4 minor 57 references

Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Routinely reported 'hate' and 'toxicity' rates for state-backed influence operations over-count hate by about twofold, because the detectors behind them actually measure a broader category of hostile or divisive out-group targeting.

desk verdict The core measurement argument holds up—the 'hate' gate overstates hate ~2x on these operations—but the precise composition split rests on an LLM-derived target field the paper admits it never validated at scale. read the letter →

arxiv 2607.14491 v1 pith:MFBCGN7C submitted 2026-07-16 cs.SI cs.CL

classification cs.SIcs.CL
keywords hatespeechdetectioninfluenceoperationsmeasurementvaliditycontentmoderationsocialmediamanipulationLLMannotationmanufactureddivisivenesspartisanhostility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the widely reported 'hate' and 'toxicity' rates for state-backed influence campaigns are inflated by a measurement error: the detectors used to produce them are validated to catch a broader phenomenon—hostile or divisive attacks on an out-group—not hate speech proper. On 25 million posts from seven government-attributed campaigns, it separates the flagged content into three types and shows that under its stated criteria only about 19% of the flagged posts are both identity-directed and dehumanizing or inciting. Reporting every flagged post as 'hate' thus overstates hate roughly twofold. The mix also differs systematically by operation, with six of the seven campaigns sorting into three construct regimes, and the paper introduces 'manufactured divisiveness' as the shared product. A sympathetic reader would care because the correction changes how we measure platform harm and how we compare state actors.

What carries the argument

The argument is carried by a two-stage instrument. Stage one is a broad-construct gate: a language model scores each (tweet, target) pair for hostile or divisive out-group targeting, and an item is positive only if two prompts (one permissive, one strict) both flag it; the gate is validated against human gold at Cohen's kappa = 0.82. Stage two is an auditable rule over a frozen 11-dimension characterization taxonomy produced by the model: the rule deterministically maps the target-group and narrative fields to one of three constructs (identity hate, partisan divisiveness, state/geopolitical invective) and flags a dehumanizing/inciting hard core. The division of labor is key: the broad gate i

What would settle it

Re-code a stratified sample (e.g., 300 posts) of the 5,457 gate-positive posts with human labels for whether the target is an identity group, a partisan actor, or a state; if the human-validated identity share falls far below the model-derived share, or if human-coded hate-speech prevalence matches the broad flag, the central overstatement claim would be refuted.

Watch

Extended reading notes

Core claim

The central discovery is that the construct a detector measures matters: a gate validated at high agreement as a detector of hostile or divisive out-group targeting is not a hate-speech detector. Applied to 5,457 gate-positive posts across seven operations, an auditable typing rule assigns 50.1% to identity-based attacks on people, 30.4% to partisan attacks, and 19.5% to invective against states and foreign policy; only 18.7% meets the narrower standard of identity-directed dehumanizing or inciting content. Consequently a single 'hate rate' compresses categorically different content and overstates hate by a factor of about two (up to five against the strictest core). The composition is not u

Load-bearing premise

The decomposition rests on the model-derived field that identifies whom each post targets; the paper does not validate that field at scale against human labels, and 69% of its typing decisions depend on it, so any systematic error in target assignment would shift the split and the overstatement factor.

Editorial extensions

If this is right

  • Reported hate rates for these seven operations overstate hate speech by about 2x; measured against the narrowest defensible core the gap widens to about 5x.
  • The three construct regimes show that operations scoring similarly under a broad gate produce categorically different hostile content, so cross-actor comparisons based on a single rate are misleading.
  • The divisive/hate boundary has no annotator-stable reading: three experts agree only moderately and no model exceeds kappa 0.601 against the expert majority, so any single gold label inherits an idiosyncratic reading.
  • The leading Russia-attributed operation remains the only one with a non-trivial prevalence floor at the dehumanizing-identity-hate core, so its elevation persists—and becomes more distinct—under the narrowed construct.
  • Account-level concentration is construct-specific: one operation's high concentration under the broad gate disappears when attention is restricted to identity hate, implying that broad-construct concentration should not be read as identity-hate concentration.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same over-attribution mechanism should apply to any study reporting a toxicity or hate rate from a detector validated against a broad hostility construct; the correction factor is set by how much of the hostile tail in that corpus is non-identity invective and can be estimated by re-typing a sample.
  • If the target-group attribute carries bias (the paper's unvalidated field), the 50/30/20 split is a lower bound on identity hate; the overstatement conclusion still holds within an envelope of 1.5–2.2x, so the core finding is unlikely to reverse.
  • Platforms and researchers could adopt the two-stage design and report composition (identity vs. partisan vs. geopolitical) alongside any rate, which would change how influence-operation harm is compared across actors.
  • Because even experts disagree on the boundary, automated enforcement keyed to the identity boundary should be treated as triage, not as a fixed threshold; the paper frames this as a reporting recommendation, which a reader can extend to policy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper analyzes 25.08M tweets from seven state-attributed influence operations and argues that widely reported 'hate' or 'toxicity' rates for such content rest on a measurement error: the detectors validate a broad construct of hostile/divisive out-group targeting rather than hate speech proper. The authors first validate a two-prompt LLM gate (κ=0.82 on a 100-item gold), then apply an auditable rule to the 5,457 gate-positive items, typing them as identity-based hate (50.1%), partisan divisiveness (30.4%), or geopolitical invective (19.5%); only 18.7% are both identity-directed and dehumanizing/inciting. They report that the broad flag therefore overstates hate by roughly 2–5×, that six of seven operations fall into three construct regimes (identity hate, geopolitical invective, partisan divisiveness), and that the divisive/hate boundary is itself unstable across expert annotators (κ=0.37–0.50) and models (best κ=0.601 against the expert majority). The paper frames the contribution as typing the construct before counting it.

Significance. If the quantitative composition is robust, this is a valuable measurement critique for computational social science and platform governance: it provides a transparent, auditable rule, separates the broad hostility construct from narrower hate, and shows that a single scalar 'hate rate' flattens heterogeneous operations. Strengths include the separate validation of gate and typing rule, the boundary-enriched human gold with three experts and nineteen models, the explicit boundary sweep (identity share 45.1–66.9%), the conservative two-prompt consensus, the concurrent validation on the IRA role-labeled corpus, and the clearly scoped, falsifiable transfer claim. The main reservation is that the decisive target-group attribute is LLM-derived and not validated at scale, so the exact split and overstatement margins are provisional despite the paper's careful internal validation.

major comments (3)
  1. [§5.1; Table 2] The central numbers—50.1% identity, 30.4% partisan, 19.5% geopolitical, 18.7% core, and the 2× overstatement—all rest on the LLM-derived target-group attribute. The paper concedes in §7 that this attribute 'is itself LLM-derived and is not separately validated at scale,' and §4.3 reports that 69.0% of assignments are target-decisive. The §5.1 boundary sweep varies the decision rule while holding target labels fixed, so it does not bound characterization-model bias on that field. A systematic error in labeling, for example calling partisan attacks identity-based or state-targeted invective group-based, would move the 50.1/30.4/19.5 split and the 2× margin. Please validate the target field on a stratified sample of the 5,457 positives (including the 850 machine-translated items) against human gold, or provide a sensitivity analysis that perturbs target labels and recomputes the composition
  2. [§5.1; Table 2] The headline composition percentages are point estimates with no uncertainty despite moderate typing reliability (κ=0.52) and widely varying positive-set sizes across operations (e.g., BD-op has tens of items). Report bootstrap confidence intervals for the overall and per-operation splits, or at least clearly flag rows whose composition is statistically unstable. The 45.1–66.9% envelope in §5.1 is a sensitivity range for the rule's boundary choices, not a confidence interval for the underlying target attribute, and should not be read as covering characterization-model error.
  3. [Abstract, §1, §6] The abstract and introduction assert that existing reported 'hate'/'toxicity' rates 'rest on a measurement error,' but the 2–5× overstatement is computed for the authors' own two-prompt Qwen gate, not for the detectors used in the cited prior work [46,45]. The transfer to prior detectors is assumed rather than directly tested. The internal comparison (5,457 vs. 2,733; 5,457 vs. 1,023) is sound for this gate, but the manuscript should either test the actual prior detectors on a shared sample or soften the framing to a conditional claim about how broad-hostility detectors behave.
minor comments (4)
  1. [Abstract; Table 1] Typo: 'shared productmanufactured divisiveness' needs a space; Table 1 caption has 'T able'. Also 'dehumanizing hardcore' is inconsistently hyphenated as 'hard-core' elsewhere.
  2. [§5.3] Table 3 is dense but effective; consider adding a note that greedy decoding was used for all models and that API model sampling variability was not assessed, since the 'no model exceeds κ=0.601' claim is single-run.
  3. [§5.4] The concurrent validation on the Clemson IRA corpus is a useful check, but its candidate selection uses the same cue-and-target filter, so the role-level percentages are not corpus base rates; the text acknowledges this, but a one-sentence reminder in the figure caption would help readers.
  4. [References] Reference [9] is a companion preprint by the same author; consider clarifying the relationship in one sentence to avoid any appearance of redundancy with this manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline composition is a rule-based count validated against human/expert labels, not a fitted prediction or self-citation-derived result.

full rationale

The derivation chain is self-contained: a gate for the broad construct is validated against human gold (Cohen's kappa = 0.82, precision 0.96 on a 100-item sample); the typing rule is validated against an independent expert (kappa = 0.52, accuracy 68%) and falls within the human-human agreement band (kappa = 0.44 on a contested-boundary subset). The 5,457-item positive census, the 50.1/30.4/19.5 composition, the 18.7% dehumanizing-identity core, and the 2-5x overstatement margins are all computed by applying the stated deterministic rule to the frozen target/narrative taxonomy; no parameter was optimized to reproduce those percentages, and the margins are arithmetic ratios of the resulting counts. The rule's single refinement (kappa 0.42 to 0.52) was an expert-agreement improvement, not a fit to the headline composition. The concurrent validation on the Clemson IRA corpus uses independently known role labels with roles withheld, providing an external check. Self-citations appear but are contextual or corroborative, not load-bearing: for example, [2] on 2016 manipulation traces, [8]/[9] as companion work, [18] on annotator inconsistency, [24] on political-leaning inference, and [36]/[48] as backdrop on bots and negative content. No uniqueness theorem or prior ansatz is imported to force the construct split. The Section 7 admission that 'the decisive target-group attribute is itself LLM-derived and is not separately validated at scale' is an external-validity risk, not a circular reduction by construction; the paper explicitly frames the split as 'the rule's partition under these stated definitions.' Thus the central derivation does not reduce to its own inputs.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The central claim depends on hand-designed rule boundaries and on the LLM-derived target taxonomy (free parameters and domain assumption), while the statistical tools are standard. The invented entity is a label, not a mechanism.

free parameters (3)
  • Typing rule boundary choices = identity targets list, ambiguous-category arbitration, text-level override
    Hand-chosen categorical boundaries in the auditable rule: which target categories count as protected/ascriptive, how ethnic-national and immigration targets are arbitrated by narrative, and the override re-typing items that name terrorist orgs/foreign states to geopolitical. These were refined to raise expert agreement from κ=0.42 to 0.52 before validation on a second expert (Section 4.3).
  • Two-prompt consensus gate = CLEAN ∧ STRICT2 both flag
    Design choice requiring both a permissive and stricter prompt to flag; conservative by construction. Hand-chosen, not fitted, though the prompts themselves are crafted (Section 4.1).
  • Sensitivity envelope endpoints = identity share 45.1%–66.9%; overstatement 1.5–2.2x
    The two extremes of the boundary sweep are chosen to bracket the most identity-favorable and most identity-conservative readings; used for robustness, not point estimates (Section 5.1).
assumptions (4)
  • domain assumption Archive attribution labels are accepted
    The seven campaigns are treated as government-attributed based on the Twitter Information Operations archive labels; the paper avoids asserting its own attribution but relies on the archive's. (Section 3, Table 1).
  • domain assumption The gate represents the broad construct used by prior detectors
    The argument that prior reported 'hate' rates over-attribute assumes the detectors in [45,46] are validated to catch the same hostile/divisive construct the paper's gate measures. This is not directly tested. (Section 1).
  • domain assumption Target-group attribute is the reliable axis for typing
    The rule partitions by 'whom the content targets,' presupposing that target assignment is more objective than an is-it-hate verdict, despite the LLM derivation of that attribute and the contested boundary. (Section 4.3, 7).
  • standard math Permutation/statistical conventions
    Use of Cohen's kappa, Fleiss kappa, Cramér's V with item-level permutation tests, Jeffreys prior intervals, Mann-Whitney U with BH-FDR correction. Standard and appropriate. (Sections 4, 5).
invented entities (1)
  • Manufactured divisiveness
    purpose: Shared product of the seven influence operations, encompassing identity hate, partisan divisiveness, and geopolitical invective; narrows to identity hate and to a dehumanizing/inciting core.
    A framing term, not a measured quantity with its own falsifiable handle. The paper does offer a transferable prediction about over-attribution margins, but the entity itself is a descriptive construct. (Section 6).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations." pith.science (2026). https://pith.science/paper/MFBCGN7C

@misc{pith2026260714491,
  author       = {Pith},
  title        = {Pith review of: Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MFBCGN7C}},
  note         = {Machine review of arXiv:2607.14491}
}
abstract

State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed at an out-group, and so over-attribute hate to content better described as partisan or geopolitical invective. Across 25.08M tweets from seven government-attributed campaigns in the Twitter Information Operations archive (8,275 accounts), we separate hate from the other forms of divisiveness. We first validate a two-prompt LLM-based detector, matching human labels at Cohen's $\kappa=0.82$, to identify the broader hostility; we then develop an auditable rule, agreeing with an expert at $\kappa=0.52$, to further classify this content (5,457 posts) into three sub-categories. About 50.1% are identity-based attacks on people, whereas 30.4% are partisan attacks and 19.5% invective against states and their foreign policy. Reporting all of it as hate therefore overstates hate roughly twofold; only 18.7% is both identity-based and dehumanizing or inciting. Six of seven campaigns sort into three regimes that a single ``hate'' rate flattens, namely identity hate (RU-op and IRA, both Russia-attributed), geopolitical invective (both Iran operations), and partisan divisiveness (both Venezuela operations). We call the shared product $manufactured divisiveness$. The line to separate these constructs itself remains unsettled: on the hardest cases three independent human experts agree only moderately (pairwise $\kappa=0.37$--$0.50$), and the best of nineteen LLM models tops out at $\kappa=0.601$ against the experts' majority. Our findings can help redefine the study of hate in the context of influence campaigns and broader online discourse.

Figures

Figures reproduced from arXiv: 2607.14491 by the authors.

Figure 1
Figure 1. Construct composition of positive content, by op￾eration (row-normalized; n positive items per operation in parentheses). Bars decompose each operation’s hostile content into identity-based hate, partisan divisiveness, and geopolit￾ical invective. The typing is a rule applied to the frozen target/narrative taxonomy over items the human-validated broad gate admitted, not an item-level human classification of hate. Th… view at source ↗
Figure 2
Figure 2. Inter-annotator agreement on the broad construct over the boundary-enriched 102-item gold: pairwise Cohen’s κ among the nineteen models and the three human experts, ordered by hierarchical clustering on 1 − κ (average linkage), so the dendrogram groups annotators that label the borderline region alike. Red boxes mark the subblocks isolated by a single depth cut of the dendrogram (dashed line). The three experts (bol… view at source ↗
Figure 3
Figure 3. Concurrent validation on the Clemson IRA corpus, where operator roles are independently known [27, 12, 1]; roles were withheld from the model. (a) The broad gate confirms hostility on the four troll roles (filled circles) but stays near zero on the News Feed and Commercial service roles (open squares); bars are Wilson 95% intervals, n is candidate items per role. (b) Among identity-hate items, the right-flank role t… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Prevalence floors (Jeffreys 95% lower bound) by construct subset (broad gate, identity hate, and the dehu￾manizing hardcore), grouped bars per operation on a log axis (hatching distinguishes subsets in grayscale). Floors fall as the construct narrows; RU-op has the hig…
Figure 5
Figure 5. Figure 5: Cross-operation Cram´er’s V per dimension, by con￾struct subset (n in legend; rows sorted by the broad-construct value). Each dimension shows three points (broad construct, circle; identity subset, square; dehumanizing hardcore, trian￾gle), with bootstrap 95% confidenc…
Figure 6
Figure 6. Figure 6: Account concentration (Gini of hostile items per posting account), as a dumbbell from the broad construct to the identity-hate subset. RU-op and IR-op-A stay highly concentrated at both layers; for VE-op-A concentration falls (0.77 → 0.13), so its broad-construct conce…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 2 linked inside Pith

  1. [1]

    Acting the part: Examining information operations within #BlackLivesMatter discourse.Proceedings of the ACM on Human-Computer Interaction (CSCW), 2:1–27, 2018

    Ahmer Arif, Leo Graiden Stewart, and Kate Starbird. Acting the part: Examining information operations within #BlackLivesMatter discourse.Proceedings of the ACM on Human-Computer Interaction (CSCW), 2:1–27, 2018

  2. [2]

    Analyzing the digital traces of political manipulation: The 2016 Russian interference Twitter campaign

    Adam Badawy, Emilio Ferrara, and Kristina Lerman. Analyzing the digital traces of political manipulation: The 2016 Russian interference Twitter campaign. In Proceedings of the IEEE/ACM International Confer- ence on Advances in Social Networks Analysis and Mining (ASONAM), pages 258–265, 2018

  3. [3]

    A unified taxonomy of harmful content

    Michele Banko, Brendon MacKeen, and Laurie Ray. A unified taxonomy of harmful content. InProceedings of the Fourth Workshop on Online Abuse and Harms (ACL), pages 125–137, 2020

  4. [4]

    Brady, Julian A

    William J. Brady, Julian A. Wills, John T. Jost, Joshua A. Tucker, and Jay J. Van Bavel. Emotion shapes the diffusion of moralized content in social networks.Proceedings of the National Academy of Sciences, 114(28):7313–7318, 2017

  5. [5]

    Dealing with disagreements: Looking beyond the majority vote in subjective an- notations.Transactions of the Association for Com- putational Linguistics (TACL), 10:92–110, 2022

    Aida Mostafazadeh Davani, Mark D ´ ıaz, and Vinod- kumar Prabhakaran. Dealing with disagreements: Looking beyond the majority vote in subjective an- notations.Transactions of the Association for Com- putational Linguistics (TACL), 10:92–110, 2022

  6. [6]

    Automated hate speech detection and the problem of offensive language

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. Automated hate speech detection and the problem of offensive language. InProceedings of the International AAAI Conference on Web and Social Media (ICWSM), pages 512–515, 2017

  7. [7]

    Latent hatred: A benchmark for understanding implicit hate speech

    Mai ElSherief, Caleb Ziems, David Muchlinski, Vaish- navi Anupindi, Jordyn Seybolt, Munmun De Choud- hury, and Diyi Yang. Latent hatred: A benchmark for understanding implicit hate speech. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 345–363, 2021

  8. [8]

    Cultural targets, structural frames, binding morals: A cross-lingual audit of online hate in multicultural singapore, 2026

    Emilio Ferrara. Cultural targets, structural frames, binding morals: A cross-lingual audit of online hate in multicultural singapore, 2026. arXiv preprint arXiv:2606.21996

Show all 57 references
  1. [9]

    Five myths about influence oper- ations: What 25 million tweets across seven state campaigns reveal

    Emilio Ferrara. Five myths about influence oper- ations: What 25 million tweets across seven state campaigns reveal. Preprints, 2026

  2. [10]

    Characterizing social media manipulation in the 2020 U.S

    Emilio Ferrara, Herbert Chang, Emily Chen, Goran Muric, and Jaimin Patel. Characterizing social media manipulation in the 2020 U.S. presidential election. First Monday, 25(11), 2020

  3. [11]

    Finkel, Christopher A

    Eli J. Finkel, Christopher A. Bail, Mina Cikara, Pe- ter H. Ditto, Shanto Iyengar, Samara Klar, Lilliana Mason, Mary C. McGrath, Brendan Nyhan, David G. Rand, Linda J. Skitka, Joshua A. Tucker, Jay J. Van Bavel, Cynthia S. Wang, and James N. Druck- man. Political sectarianism ...

  4. [12]

    Russian troll tweets

    FiveThirtyEight. Russian troll tweets. https://github.com/fivethirtyeight/ russian-troll-tweets, 2018. Internet Re- search Agency tweets collected and categorized by Linvill and Warren, Clemson University

  5. [13]

    A survey on auto- matic detection of hate speech in text.ACM Com- puting Surveys, 51(4):1–30, 2018

    Paula Fortuna and S´ ergio Nunes. A survey on auto- matic detection of hate speech in text.ACM Com- puting Surveys, 51(4):1–30, 2018

  6. [14]

    Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets

    Paula Fortuna, Juan Soler, and Leo Wanner. Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets. InProceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 6786–6794, 2020

  7. [15]

    Large scale crowd- sourcing and characterization of Twitter abusive be- havior

    Antigoni-Maria Founta, Constantinos Djouvas, De- spoina Chatzakou, Ilias Leontiadis, Jeremy Black- burn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. Large scale crowd- sourcing and characterization of Twitter abusive be- havior. InProceeding...

  8. [16]

    Black trolls matter: Racial and ideological asymme- tries in social media disinformation.Social Science Computer Review, 40(3):560–578, 2022

    Deen Freelon, Michael Bossetta, Chris Wells, Josephine Lukito, Yiping Xia, and Kirsten Adams. Black trolls matter: Racial and ideological asymme- tries in social media disinformation.Social Science Computer Review, 40(3):560–578, 2022

  9. [17]

    Frimer, Reihane Boghrati, Jonathan Haidt, Jesse Graham, and Morteza Dehghani

    Jeremy A. Frimer, Reihane Boghrati, Jonathan Haidt, Jesse Graham, and Morteza Dehghani. Moral foun- dations dictionary for linguistic analyses 2.0. Unpub- lished manuscript, distributed via OSF, 2019

  10. [18]

    Position: RLHF may not reflect genuine preferences

    Bijean Ghafouri, Eun Cheol Choi, Priyanka Dey, and Emilio Ferrara. Position: RLHF may not reflect genuine preferences. InProceedings of the 43rd Inter- national Conference on Machine Learning (ICML). PMLR, 2026

  11. [19]

    ChatGPT outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

    Fabrizio Gilardi, Meysam Alizadeh, and Ma¨ el Kubli. ChatGPT outperforms crowd workers for text-annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

  12. [20]

    Jesse Graham, Jonathan Haidt, and Brian A. Nosek. Liberals and conservatives rely on different sets of moral foundations.Journal of Personality and Social Psychology, 96(5):1029–1046, 2009. 11

  13. [21]

    Still out there: Modeling and identifying Russian troll accounts on Twitter

    Jane Im, Eshwar Chandrasekharan, Jackson Sargent, Paige Lighthammer, Taylor Denby, Ankit Bhargava, Libby Hemphill, David Jurgens, and Eric Gilbert. Still out there: Modeling and identifying Russian troll accounts on Twitter. InProceedings of the 12th ACM Conference on Web Scie...

  14. [22]

    Westwood

    Shanto Iyengar, Yphtach Lelkes, Matthew Leven- dusky, Neil Malhotra, and Sean J. Westwood. The origins and consequences of affective polarization in the united states.Annual Review of Political Science, 22:129–146, 2019

  15. [23]

    Jacobs and Hanna Wallach

    Abigail Z. Jacobs and Hanna Wallach. Measurement and fairness. InProceedings of the 2021 ACM Confer- ence on Fairness, Accountability, and Transparency (F AccT), pages 375–385, 2021

  16. [24]

    Retweet- BERT: Political leaning detection using language features and information diffusion on social networks

    Julie Jiang, Xiang Ren, and Emilio Ferrara. Retweet- BERT: Political leaning detection using language features and information diffusion on social networks. InProceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 17, pages 459–469, 2023

  17. [25]

    Reliability in content analy- sis: Some common misconceptions and recommenda- tions.Human Communication Research, 30(3):411– 433, 2004

    Klaus Krippendorff. Reliability in content analy- sis: Some common misconceptions and recommenda- tions.Human Communication Research, 30(3):411– 433, 2004

  18. [26]

    Watch your language: Investigating content moderation with large language models

    Deepak Kumar, Yousef Anees AbuHashem, and Za- kir Durumeric. Watch your language: Investigating content moderation with large language models. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 18, pages 865–878, 2024

  19. [27]

    Linvill and Patrick L

    Darren L. Linvill and Patrick L. Warren. Troll fac- tories: Manufacturing specialized disinformation on Twitter.Political Communication, 37(4):447–467, 2020

  20. [28]

    Creating a general Russian sentiment lexicon

    Natalia Loukachevitch and Anatolii Levchik. Creating a general Russian sentiment lexicon. InProceedings of the Tenth International Conference on Language Resources and Evaluation (LREC), pages 1171–1176, 2016

  21. [29]

    University of Chicago Press, Chicago, IL, 2018

    Lilliana Mason.Uncivil Agreement: How Politics Became Our Identity. University of Chicago Press, Chicago, IL, 2018

  22. [30]

    From dogwhistles to bullhorns: Unveil- ing coded rhetoric with language models

    Julia Mendelsohn, Ronan Le Bras, Yejin Choi, and Maarten Sap. From dogwhistles to bullhorns: Unveil- ing coded rhetoric with language models. InProceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), pages 15162– 15180, 2023

  23. [31]

    Mohammad and Peter D

    Saif M. Mohammad and Peter D. Turney. Crowd- sourcing a word–emotion association lexicon.Com- putational Intelligence, 29(3):436–465, 2013

  24. [32]

    Carley.Bots, Bias, and Influence: The Hidden Architects of Social Media

    Lynnette Hui Xian Ng and Kathleen M. Carley.Bots, Bias, and Influence: The Hidden Architects of Social Media. Cambridge Scholars Publishing, 2026

  25. [33]

    Keeping humans in the loop: Human-centered automated annotation with generative AI

    Nick Pangakis and Sam Wolken. Keeping humans in the loop: Human-centered automated annotation with generative AI. InProceedings of the Interna- tional AAAI Conference on Web and Social Media (ICWSM), volume 19, pages 1471–1492, 2025

  26. [34]

    Pennebaker, Ryan L

    James W. Pennebaker, Ryan L. Boyd, Kayla Jordan, and Kate Blackburn. The development and psycho- metric properties of LIWC2015. Technical report, University of Texas at Austin, 2015

  27. [35]

    The “problem” of human label vari- ation: On ground truth in data, modeling and eval- uation

    Barbara Plank. The “problem” of human label vari- ation: On ground truth in data, modeling and eval- uation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 10671–10682, 2022

  28. [36]

    Measuring bot and human behavioral dynamics.Frontiers in Physics, 8:125, 2020

    Iacopo Pozzana and Emilio Ferrara. Measuring bot and human behavioral dynamics.Frontiers in Physics, 8:125, 2020

  29. [37]

    Van Bavel, and Sander van der Linden

    Steve Rathje, Jay J. Van Bavel, and Sander van der Linden. Out-group animosity drives engagement on social media.Proceedings of the National Academy of Sciences, 118(26):e2024292118, 2021

  30. [38]

    Beyond incivility: Understanding patterns of uncivil and intolerant discourse in online political talk.Communication Research, 49(3):399– 425, 2022

    Patr ´ ıcia Rossini. Beyond incivility: Understanding patterns of uncivil and intolerant discourse in online political talk.Communication Research, 49(3):399– 425, 2022

  31. [39]

    XSTest: A test suite for identifying exaggerated safety behaviours in large language models

    Paul R¨ ottger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. XSTest: A test suite for identifying exaggerated safety behaviours in large language models. InProceedings of the 2024 Conference of the North American Chap- ter of the Associ...

  32. [40]

    Unrav- eling the web of disinformation: Exploring the larger context of state-sponsored influence campaigns on Twitter

    Mohammad Hammas Saeed, Shiza Ali, Pujan Paudel, Jeremy Blackburn, and Gianluca Stringhini. Unrav- eling the web of disinformation: Exploring the larger context of state-sponsored influence campaigns on Twitter. InProceedings of the 27th International Symposium on Research in A...

  33. [41]

    short is the road that leads from fear to hate

    Punyajoy Saha, Binny Mathew, Kiran Garimella, and Animesh Mukherjee. “short is the road that leads from fear to hate”: Fear speech in Indian WhatsApp groups. InProceedings of the Web Conference 2021 (WWW), pages 1110–1121, 2021. 12

  34. [42]

    Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. The risk of racial bias in hate speech detection. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1668–1678, 2019

  35. [43]

    Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. An- notators with attitudes: How annotator beliefs and identities bias toxic language detection. InProceed- ings of the 2022 Conference of the North American Chapter of the Association fo...

  36. [44]

    Nwala, Lake Yin, Luca Luceri, Alessandro Flammini, and Filippo Menczer

    Ozgur Can Seckin, Manita Pote, Alexander C. Nwala, Lake Yin, Luca Luceri, Alessandro Flammini, and Filippo Menczer. Labeled datasets for research on information operations. InProceedings of the Inter- national AAAI Conference on Web and Social Media (ICWSM), volume 19, pages 2...

  37. [45]

    The language of influence: Sentiment, emotion, and hate speech in state sponsored influence operations

    Ashfaq Ali Shafin and Khandaker Mamun Ahmed. The language of influence: Sentiment, emotion, and hate speech in state sponsored influence operations. InProceedings of the 18th International Conference on PErvasive Technologies Related to Assistive Envi- ronments (PETRA), 2025. ...

  38. [46]

    Toxicity in state sponsored information operations

    Ashfaq Ali Shafin and Khandaker Mamun Ahmed. Toxicity in state sponsored information operations. In Proceedings of the 36th ACM Conference on Hypertext and Social Media (HT ’25), 2025

  39. [47]

    Disinformation’s spread: Bots, trolls and all of us.Nature, 571(7766):449, 2019

    Kate Starbird. Disinformation’s spread: Bots, trolls and all of us.Nature, 571(7766):449, 2019

  40. [48]

    Bots increase exposure to negative and inflammatory content in online social systems

    Massimo Stella, Emilio Ferrara, and Manlio De Domenico. Bots increase exposure to negative and inflammatory content in online social systems. Proceedings of the National Academy of Sciences, 115(49):12435–12440, 2018

  41. [49]

    Temporal dynamics of co- ordinated online behavior: Stability, archetypes, and influence.Proceedings of the National Academy of Sciences, 121(20):e2307038121, 2024

    Serena Tardelli, Leonardo Nizzoli, Maurizio Tesconi, Mauro Conti, Preslav Nakov, Giovanni Da San Mar- tino, and Stefano Cresci. Temporal dynamics of co- ordinated online behavior: Stability, archetypes, and influence.Proceedings of the National Academy of Sciences, 121(20):e23...

  42. [50]

    Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio

    Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. Learning from disagreement: A survey.Journal of Artificial Intelligence Research, 72:1385–1470, 2021

  43. [51]

    Directions in abusive language training data: A systematic review of garbage in, garbage out.PLOS ONE, 15(12):e0243300, 2020

    Bertie Vidgen and Leon Derczynski. Directions in abusive language training data: A systematic review of garbage in, garbage out.PLOS ONE, 15(12):e0243300, 2020

  44. [52]

    Evidence of inter-state coordination amongst state-backed information operations.Scien- tific Reports, 13:7716, 2023

    Xinyu Wang, Jiayi Li, Eesha Srivatsavaya, and Sarah Rajtmajer. Evidence of inter-state coordination amongst state-backed information operations.Scien- tific Reports, 13:7716, 2023

  45. [53]

    Understanding abuse: A typology of abusive language detection subtasks

    Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. Understanding abuse: A typology of abusive language detection subtasks. InProceedings of the First Workshop on Abusive Language Online (ACL), pages 78–84, 2017

  46. [54]

    Hateful symbols or hateful people? predictive features for hate speech detection on Twitter

    Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on Twitter. InProceedings of the NAACL Student Research Workshop, pages 88–93, 2016

  47. [55]

    Disinformation warfare: Understanding state-sponsored trolls on Twitter and their influence on the web

    Savvas Zannettou, Tristan Caulfield, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringh- ini, and Jeremy Blackburn. Disinformation warfare: Understanding state-sponsored trolls on Twitter and their influence on the web. InCompanion Proceed- ings of the World Wide Web...

  48. [56]

    Who let the trolls out? towards under- standing state-sponsored trolls

    Savvas Zannettou, Tristan Caulfield, William Setzer, Michael Sirivianos, Gianluca Stringhini, and Jeremy Blackburn. Who let the trolls out? towards under- standing state-sponsored trolls. InProceedings of the 10th ACM Conference on Web Science (WebSci), pages 353–362, 2019

  49. [57]

    Can large lan- guage models transform computational social science? Computational Linguistics, 50(1):237–291, 2024

    Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large lan- guage models transform computational social science? Computational Linguistics, 50(1):237–291, 2024. 13

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.