Pith. sign in

REVIEW 3 major objections 6 minor 58 references

Hostility Detection in UK Politics: A Dataset on Online Abuse Targeting MPs

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A new dataset maps two years of hostile tweets aimed at UK MPs and tests which models can spot them.

desk verdict A solid, much-needed dataset for UK political hostility with careful annotation, but the headline descriptive claims about party and identity differences are not supported by the sampling design. read the letter →

arxiv 2412.04046 v1 pith:LKKJZLVD submitted 2024-12-05 cs.CL

classification cs.CL
keywords UKpoliticshostilitydetectionMPabuseTwitterdatasetidentity-basedhateannotationlargelanguagemodelsspeech
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper constructs a publicly available dataset of 3,320 tweets directed at UK MPs over two years, each manually annotated for hostility and, when hostile, for the identity targeted: race, gender, religion, or none. The authors argue this fills a gap, because existing UK political hostility datasets either lack manual hostility labels, cover only a short period, or focus on a single identity type. They show that hostility tracks contemporaneous issues such as Brexit, illegal immigration, and the cost-of-living crisis, that Conservative MPs in their sample receive more race-based hostility, and that MPs from racial and religious minorities receive more identity-based hostility. They then benchmark five models, finding that a domain-adapted RoBERTa reaches the best macro F1 (73.03) on binary hostility detection, while GPT-3.5 with in-prompt definitions leads identity classification in a hierarchical setup (55.98 macro F1).

What carries the argument

The central object is the annotation taxonomy and dataset construction pipeline. An umbrella 'hostile' definition merges hate, abuse, and toxicity; a hierarchical task structure first asks hostile or not, then religion, gender, race, or none; and each label carries a 1–5 confidence score. Three gold-label sets are derived: majority vote (Set 1), confidence-filtered (Set 2), and intersectionality-preserving (Set 3). The sampling pipeline selects 18 MPs balanced by party and identity, takes their five highest-posting-activity days, and oversamples tweets flagged by an abusive-language classifier, producing 3,320 tweets. This machinery is what lets the paper claim the dataset is identity-aware, temporally diverse, and quality-controlled.

What would settle it

Re-run the analysis on a random sample of tweets from all 568 MPs with active accounts across the full two-year period, instead of the 18 selected MPs and their five highest-activity days. If the Conservative Party no longer shows a higher rate of race-based hostility than Labour, the paper's headline descriptive claim fails. A second check: re-annotate the same tweets with a fresh pool of annotators and recompute Fleiss' kappa; if identity-label agreement falls below substantial (0.4), the identity labels do not support the comparative findings.

Watch

Extended reading notes

Core claim

The central claim is that the two-year, identity-labelled dataset is a step-change resource for UK political hostility detection, enabling both descriptive analysis and model training that prior resources could not support. In the paper's own framing, it 'bridges the gap' left by datasets that lack hostility labels, cover short windows, or address only Islamophobia. The authors demonstrate the resource's value with three gold-label sets, a confidence-aware annotation scheme, and a comparison of fine-tuned transformers (BERT, RoBERTa, RoBERTa-Hate) and zero-shot LLMs (LLaMA-3-8B, GPT-3.5) on binary hostility identification and multi-class identity classification.

Load-bearing premise

The paper assumes that 18 hand-selected MPs and their five highest-posting-activity days represent hostility towards UK MPs as a whole; if those MPs or days are unrepresentative, the aggregate party and identity comparisons do not generalize.

Editorial extensions

If this is right

  • The released dataset gives researchers a UK-specific, identity-labeled resource for training and evaluating hostility detectors.
  • Confidence-filtered labels (Set 2) improve model performance, so future annotation efforts should record and use per-annotation confidence.
  • Because hostility tracks contemporaneous issues, classifiers trained on this two-year period will likely need periodic updating as new issues emerge.
  • Adding label definitions to LLM prompts yields large gains in identity classification, suggesting prompt design is a key lever for zero-shot abuse detection.
  • Race-based hostility is the most common identity category in the sample, with illegal immigration discussions carrying twice the hostility of non-hostile tweets on that topic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 18-MP sample, skewed to non-white and female MPs, means the paper's descriptive statistics describe the sample, not the population; a population-weighted sample could confirm or overturn the Conservative-party finding.
  • The taxonomy's umbrella 'hostile' definition could transfer to other countries, but the paper's own Brexit and immigration examples suggest models will need country-specific retraining.
  • Annotator confusion between race and religion labels for Muslim and Jewish targets implies downstream models may inherit that confusion; expert post-correction was needed, so automated systems should expect ambiguity at that boundary.
  • The dataset's intersectional labels are few (43), so claims about intersectional hostility are suggestive, not statistically powerful.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces a manually annotated dataset of 3,320 English tweets directed at UK MPs, collected between November 2020 and December 2022. Tweets are labeled for hostility (binary) and, when hostile, for targeted identity characteristics (race, gender, religion, none), with per-annotator confidence scores; three gold-label sets are derived. The authors describe their sampling and annotation protocol, present linguistic (BOW, LIWC) and topic (BERTopic) analyses, and benchmark several pretrained language models and large language models on binary hostility detection and multi-class identity classification in flat and hierarchical settings. The dataset is released publicly on Zenodo.

Significance. If the dataset is used as a benchmark for UK-specific political hostility detection, it fills a genuine gap: existing UK political abuse datasets either lack hostility labels or cover only Islamophobia. The annotation pipeline is thorough (48 annotators, three per tweet, training and testing, hostility Fleiss kappa 0.68-0.79), and the public release with confidence scores and multiple gold sets is a useful resource. The model evaluation provides reproducible baselines. The descriptive claims about party- and identity-based differences are not, however, supported by the sampling design as currently presented.

major comments (3)
  1. [Data Sampling; Dataset (Figures 2-3)] The sampling design fixes, for each of the 18 hand-selected MPs, exactly 17 classifier-positive and 20 classifier-negative tweets per MP on each of the 5 highest-posting-activity days. Consequently, the raw counts in Figures 2 and 3 and the statements in the Dataset section ('MPs belonging to the Conservative Party receive more race-based hostility'; 'non-white and non-Christian MPs face significantly higher levels') are not estimates of hostility rates in the population of UK MPs: the quota on hostile tweets makes the number of hostile tweets per MP-day an artifact of sampling, and the hand-picked MPs and peak-activity days are not a probability sample. The authors should either analyze the full collection or a proper random sample with appropriate denominators, or explicitly restrict these claims to the annotated sample and support them with appropriate statistical tests.
  2. [Data Sampling] The initial screening uses the Gorrell et al. (2020) abusive-language classifier to select the 17 'hostile' tweets per MP-day. The manual annotations therefore apply to a set that is conditional on this classifier's decisions. If the classifier's error rates vary by party or identity group, the annotated dataset and every descriptive statistic derived from it inherit that selection bias. The manuscript does not report any validation of this screening classifier on the 2020-2022 period or any analysis of its error distribution across parties or identities. Please quantify the screening precision and recall on a sample and discuss the implications, or weaken the descriptive claims accordingly.
  3. [Data Sampling] The claim that sampling the '5 different highest posting activity days for each MP' ensures 'a long temporal span' is not substantiated. Peak-activity days are likely to cluster around discrete political events such as elections, scandals, or policy announcements, and the authors do not report the actual date distribution of the sampled tweets over November 2020 to December 2022. Without this information, the conclusion that the two-year span provides a broader range of topics and improves generalizability is not supported. Please report the temporal distribution of the sample and, if necessary, adjust the sampling design or the conclusions.
minor comments (6)
  1. [Data Annotation] Please specify how many annotations were corrected by experts in the race/religion confusion cases, and whether the correction was applied before or after computing agreement statistics.
  2. [Table 5] The value 'moral 1.51' lacks a leading zero and is inconsistent with the other correlations; it should presumably be 0.151.
  3. [Data Characterisation] The example tweets are numbered 'Tweet 7' twice; renumber the examples to avoid ambiguity.
  4. [Experimental Set-up] The exact prompts used for LLaMA and GPT are not given; include the full prompt templates in the appendix for reproducibility.
  5. [References] The BERT citation is given as 'Kenton and Toutanova 2019'; the standard citation is Devlin et al. (2019).
  6. [Dataset Availability] The availability section mentions both an anonymous review URL and a Zenodo record; the final version should state one canonical URL for the dataset.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the dataset labels and model evaluations are independently grounded; only minor non-load-bearing self-citations appear.

full rationale

The central derivation chain is self-contained. The paper constructs a dataset whose gold labels come from 48 human annotators, with majority-vote and confidence-based aggregation, and the Gorrell et al. (2020) classifier is used only to sample candidate tweets, not to define the annotated labels. The model evaluations are 5-fold cross-validated against these manually derived gold labels, so no reported F1 or accuracy score is forced by construction. The descriptive comparisons by party and identity are subject to a real sampling-design limitation (18 hand-selected MPs, 5 highest-activity days per MP, fixed 17-hostile and 20-non-hostile quotas per MP-day), but that is a validity and generalizability concern, not a circularity one: the counts in Figures 2 and 3 are not equivalent to the screening classifier's output because annotators can and do overturn screening labels. Several self-citations appear (Gorrell et al. 2020; Jin et al. 2023; Bakir, Farrell, and Bontcheva 2024; Wilby et al. 2023), but they serve as background, a screening tool, a data-collection precedent, or an annotation-platform reference; none of them defines the outcome labels or the evaluation results. No specific reduction of a claimed result to an input by definition or by fitted parameter is exhibited, so the paper does not rise to circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the data collection and annotation process, which is transparently described and the data is released. The key assumptions are domain assumptions about construct validity, the pre-filter's accuracy, and sample representativeness. No invented entities or fitted model parameters are needed.

free parameters (2)
  • Confidence threshold for Set 2 gold labels = 3
    Hand-chosen threshold: annotations with confidence below 3 are removed when constructing Set 2 gold labels; this improves Fleiss' kappa from 0.68 to 0.79 for hostility but is an investigator choice, not a fitted value.
  • Sampling ratio of hostile to non-hostile tweets per MP per day = 17 hostile, 20 non-hostile
    Chosen to create a balanced annotation sample; affects the label distribution and downstream model training but is not fitted to any target.
assumptions (4)
  • domain assumption The umbrella definition of 'hostile' (combining hate, abuse, toxicity, and offence) is a coherent construct that annotators apply consistently.
    Annotation guidelines in Table 3 define the categories; annotators were trained and tested, and Fleiss' kappa is 0.68 to 0.79, but construct validity across identity types is assumed.
  • domain assumption The Gorrell et al. (2020) abusive language classifier accurately identifies likely hostile tweets for sampling.
    Data Sampling section uses this classifier to select 17 hostile and 20 non-hostile tweets per day; its error rate on this two-year corpus is not reported, so the sample may inherit classifier bias.
  • domain assumption The 18 selected MPs and the 5 highest-posting-activity days per MP yield a sample representative enough for the paper's descriptive claims about hostility by party and identity.
    Data Sampling section; Figures 2-3 interpret these aggregates as patterns, but the sample oversamples non-white and female MPs and event-heavy days.
  • domain assumption Self-declared public information about MPs' race, gender, and religion is accurate.
    Footnote 2 in Data Sampling states: 'The MPs' identity characteristics are based on self-declared public information.'

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hostility Detection in UK Politics: A Dataset on Online Abuse Targeting MPs." pith.science (2026). https://pith.science/paper/LKKJZLVD

@misc{pith2026241204046,
  author       = {Pith},
  title        = {Pith review of: Hostility Detection in UK Politics: A Dataset on Online Abuse Targeting MPs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LKKJZLVD}},
  note         = {Machine review of arXiv:2412.04046}
}
read the original abstract

Numerous politicians use social media platforms, particularly X, to engage with their constituents. This interaction allows constituents to pose questions and offer feedback but also exposes politicians to a barrage of hostile responses, especially given the anonymity afforded by social media. They are typically targeted in relation to their governmental role, but the comments also tend to attack their personal identity. This can discredit politicians and reduce public trust in the government. It can also incite anger and disrespect, leading to offline harm and violence. While numerous models exist for detecting hostility in general, they lack the specificity required for political contexts. Furthermore, addressing hostility towards politicians demands tailored approaches due to the distinct language and issues inherent to each country (e.g., Brexit for the UK). To bridge this gap, we construct a dataset of 3,320 English tweets spanning a two-year period manually annotated for hostility towards UK MPs. Our dataset also captures the targeted identity characteristics (race, gender, religion, none) in hostile tweets. We perform linguistic and topical analyses to delve into the unique content of the UK political data. Finally, we evaluate the performance of pre-trained language models and large language models on binary hostility detection and multi-class targeted identity type classification tasks. Our study offers valuable data and insights for future research on the prevalence and nature of politics-related hostility specific to the UK.

Figures

Figures reproduced from arXiv: 2412.04046 by the authors.

Figure 1
Figure 1. Annotation platform user interface. ensured high-quality annotations. The entire annotation pro￾cess was conducted using the collaborative web-based anno￾tation tool Teamware 2 (Wilby et al. 2023). 1. Training sessions: Training sessions were conducted in which annotators received in-person presentations ex￾plaining label definitions with detailed examples. Anno￾tators were also guided on setting up their annotator … view at source ↗
Figure 2
Figure 2. Comparing political party-based differences in the [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 5
Figure 5. Top 100 BOW bigrams associated with hostile and non-hostile tweets. The larger the text size, the higher the Pearson correlation coefficient r, and vice versa. posted to vent anger or dissatisfaction at politicians, rang￾ing from questioning their abilities and distrusting their poli￾cies to insulting their personal traits. There are also emojis like “face with symbols on mouth,” and “face vomiting,” which represent… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Top 100 BOW unigrams associated with hostile and non-hostile tweets. The larger the text size, the higher the Pearson correlation coefficient r, and vice versa. other. Data Characterisation Linguistic Analysis To investigate the difference between both the use of lan￾g…
Figure 6
Figure 6. Figure 6: Proportion of topic-related tweets belonging to [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 46 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Agarwal, P.; Hawkins, O.; Amaxopoulou, M.; Dempsey, N.; Sastry, N.; and Wood, E. 2021. Hate speech in political discourse: A case study of UK MPs on Twitter. In Proceedings of the 32nd ACM conference on hypertext and social media, 5--16

  4. [4]

    Agarwal, P.; Sastry, N.; and Wood, E. 2019. Tweeting mps: Digital engagement between citizens and members of parliament in the uk. In Proceedings of the International AAAI Conference on Web and Social Media, volume 13, 26--37

  5. [5]

    Fight, die, and if required kill

    Amarasingam, A.; Umar, S.; and Desai, S. 2022. “Fight, die, and if required kill”: Hindu nationalism, misinformation, and Islamophobia in India. Religions, 13(5): 380

  6. [6]

    Artstein, R.; and Poesio, M. 2008. Inter-Coder Agreement for Computational Linguistics . Computational linguistics, 34(4): 555--596

  7. [7]

    E.; Farrell, T.; and Bontcheva, K

    Bakir, M. E.; Farrell, T.; and Bontcheva, K. 2024. Abuse in the time of COVID-19: the effects of Brexit, gender and partisanship. Online Information Review

  8. [8]

    Basile, V.; Bosco, C.; Fersini, E.; Nozza, D.; Patti, V.; Pardo, F. M. R.; Rosso, P.; and Sanguinetti, M. 2019. Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter. In Proceedings of the 13th international workshop on semantic evaluation, 54--63

Show all 58 references
  1. [9]

    O'Reilly Media, Inc

    Bird, S.; Klein, E.; and Loper, E. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O'Reilly Media, Inc."

  2. [10]

    L.; Ashokkumar, A.; Seraj, S.; and Pennebaker, J

    Boyd, R. L.; Ashokkumar, A.; Seraj, S.; and Pennebaker, J. W. 2022. The development and psychometric properties of LIWC-22. Austin, TX: University of Texas at Austin, 1--47

  3. [11]

    Carson, A.; Mikolajczak, G.; Ruppanner, L.; and Foley, E. 2024. From online trolls to ‘Slut Shaming’: Understanding the role of incivility and gender abuse in local government. Local Government Studies, 50(2): 427--450

  4. [12]

    Collignon, S.; and R \"u dig, W. 2021. Increasing the cost of female representation? The gendered effects of harassment, abuse and intimidation towards Parliamentary candidates in the UK. Journal of elections, public opinion and parties, 31(4): 429--449

  5. [13]

    Enock, F.; Johansson, P.; Bright, J.; and Margetts, H. Z. 2023. Tracking experiences of online harms and attitudes towards online safety interventions: Findings from a large-scale, nationally representative survey of the british public. Nationally Representative Survey of the ...

  6. [14]

    Esposito, E.; and Breeze, R. 2022. Gender and politics in a digitalised world: Investigating online hostility against UK female MPs. Discourse & Society, 33(3): 303--323

  7. [15]

    How dare you call her a pig, I know several pigs who would be upset if they knew

    Esposito, E.; and Zollo, S. A. 2021. “How dare you call her a pig, I know several pigs who would be upset if they knew” A multimodal critical discursive approach to online misogyny against UK MPs on YouTube. Journal of language aggression and conflict, 9(1): 47--75

  8. [16]

    Farrell, T.; Bakir, M.; and Bontcheva, K. 2021. Mp twitter engagement and abuse post-first covid-19 lockdown in the uk: White paper. arXiv preprint arXiv:2103.02917

  9. [17]

    Fleiss, J. L. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5): 378

  10. [18]

    FORCE11 . 2020. The FAIR Data principles. https://force11.org/info/the-fair-data-principles/

  11. [19]

    Fortuna, P.; Soler, J.; and Wanner, L. 2020. Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets. In Proceedings of the Twelfth Language Resources and Evaluation Conference, 6786--6794

  12. [20]

    Fuchs, T.; and Sch \"a fer, F. 2021. Normalizing misogyny: hate speech and verbal abuse of female politicians on Japanese Twitter. In Japan forum, volume 33, 553--579. Taylor & Francis

  13. [21]

    Goodman, S.; and Locke, A. 2024. Supporting and challenging hate in an online discussion of a controversial refugee policy. Discourse Studies, 14614456231225448

  14. [22]

    E.; Greenwood, M

    Gorrell, G.; Bakir, M. E.; Greenwood, M. A.; Roberts, I.; and Bontcheva, K. 2019. Race and Religion in Online Abuse towards UK Politicians: Working Paper. arXiv preprint ArXiv:1910.00920 [Cs]

  15. [23]

    E.; Roberts, I.; Greenwood, M

    Gorrell, G.; Bakir, M. E.; Roberts, I.; Greenwood, M. A.; and Bontcheva, K. 2020. Which politicians receive abuse? Four factors illuminated in the UK general election 2019. EPJ Data Science, 9(1): 18

  16. [24]

    Gorrell, G.; Greenwood, M.; Roberts, I.; Maynard, D.; and Bontcheva, K. 2018. Twits, twats and twaddle: Trends in online abuse towards uk politicians. In Proceedings of the International AAAI Conference on Web and Social Media, volume 12

  17. [25]

    Grimminger, L.; and Klinger, R. 2021. Hate Towards the Political Opponent: A T witter Corpus Study of the 2020 US Elections on the Basis of Offensive Speech and Stance Detection. In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and...

  18. [26]

    Grootendorst, M. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794

  19. [27]

    Gross, J.; Baltz, S.; Suttmann-Lea, M.; Merivaki, L.; and Stewart III, C. 2023. Online Hostility Towards Local Election Officials Surged in 2020. Available at SSRN 4351996

  20. [28]

    Guellil, I.; Adeel, A.; Azouaou, F.; Chennoufi, S.; Maafi, H.; and Hamitouche, T. 2020. Detecting hate speech against politicians in Arabic community on social media. International Journal of Web Information Systems, 16(3): 295--313

  21. [29]

    H kansson, S. 2024. Explaining citizen hostility against women political leaders: A survey experiment in the United States and Sweden. Politics & Gender, 20(1): 1--28

  22. [30]

    Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volu...

  23. [31]

    Hua, Y.; Naaman, M.; and Ristenpart, T. 2020. Characterizing twitter users who engage in adversarial interactions against political candidates. In Proceedings of the 2020 CHI conference on human factors in computing systems, 1--13

  24. [32]

    A.; Siddiqui, M

    Jafri, F. A.; Siddiqui, M. A.; Thapa, S.; Rauniyar, K.; Naseem, U.; and Razzak, I. 2023. Uncovering Political Hate Speech During Indian Election Campaign: A New Low-Resource Dataset and Baselines. arXiv e-prints

  25. [33]

    S.; and Oussalah, M

    Jahan, M. S.; and Oussalah, M. 2023. A systematic review of Hate Speech automatic detection using Natural Language Processing. Neurocomputing, 126232

  26. [34]

    Jin, M.; Mu, Y.; Maynard, D.; and Bontcheva, K. 2023. Examining temporal bias in abusive language detection. arXiv preprint arXiv:2309.14146

  27. [35]

    Kenton, J. D. M.-W. C.; and Toutanova, L. K. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, volume 1, 2

  28. [36]

    Kuperberg, R. 2018. Intersectional violence against women in politics. Politics & Gender, 14(4): 685--690

  29. [37]

    Kuperberg, R. 2021. Incongruous and illegitimate: Antisemitic and Islamophobic semiotic violence against women in politics in the United Kingdom. Journal of Language Aggression and Conflict, 9(1): 100--126

  30. [38]

    C.; Farrell, T.; Third, A.; and Fernandez, M

    Kwarteng, J.; Perfumi, S. C.; Farrell, T.; Third, A.; and Fernandez, M. 2022. Misogynoir: challenges in detecting intersectional hate. Social Network Analysis and Mining, 12(1): 166

  31. [39]

    Lavalley, R.; and Johnson, K. R. 2022. Occupation, injustice, and anti-Black racism in the United States of America. Journal of Occupational Science, 29(4): 487--499

  32. [40]

    Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692

  33. [41]

    MacAvaney, S.; Yao, H.-R.; Yang, E.; Russell, K.; Goharian, N.; and Frieder, O. 2019. Hate speech detection: Challenges and solutions. PloS one, 14(8): e0221152

  34. [42]

    Mansur, Z.; Omar, N.; and Tiun, S. 2023. Twitter hate speech detection: a systematic review of methods, taxonomy analysis, challenges, and opportunities. IEEE Access, 11: 16226--16249

  35. [43]

    M.; Biemann, C.; Goyal, P.; and Mukherjee, A

    Mathew, B.; Saha, P.; Yimam, S. M.; Biemann, C.; Goyal, P.; and Mukherjee, A. 2021. Hatexplain: A benchmark dataset for explainable hate speech detection. In Proceedings of the AAAI conference on artificial intelligence, volume 35, 14867--14875

  36. [44]

    Mollas, I.; Chrysopoulou, Z.; Karlos, S.; and Tsoumakas, G. 2022. ETHOS: a multi-label hate speech detection dataset. Complex & Intelligent Systems, 8(6): 4663--4678

  37. [45]

    Pavlopoulos, J.; Sorensen, J.; Dixon, L.; Thain, N.; and Androutsopoulos, I. 2020. Toxicity Detection: Does Context Really Matter? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4296--4305

  38. [46]

    S.; Wich, M.; Kiening, T.; and Groh, G

    Rieger, D.; K \"u mpel, A. S.; Wich, M.; Kiening, T.; and Groh, G. 2021. Assessing the extent and types of hate speech in fringe communities: A case study of alt-right communities on 8chan, 4chan, and Reddit. Social Media+ Society, 7(4): 20563051211052906

  39. [47]

    C.; Carvalho, J

    Rosa, H.; Pereira, N.; Ribeiro, R.; Ferreira, P. C.; Carvalho, J. P.; Oliveira, S.; Coheur, L.; Paulino, P.; Sim \ a o, A. V.; and Trancoso, I. 2019. Automatic cyberbullying detection: A systematic review. Computers in Human Behavior, 93: 333--345

  40. [48]

    R \"o ttger, P.; Vidgen, B.; Hovy, D.; and Pierrehumbert, J. 2022. Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...

  41. [49]

    Scott, J. 2019. Women MPs say abuse forcing them from politics. BBC News

  42. [50]

    Solovev, K.; and Pr \"o llochs, N. 2022. Hate speech in the political discourse on social media: Disparities across parties, gender, and ethnicity. In Proceedings of the ACM Web Conference 2022, 3656--3661

  43. [51]

    everyday

    Southern, R.; and Harmer, E. 2021. Twitter, incivility and “everyday” gendered othering: An analysis of tweets sent to UK members of parliament. Social science computer review, 39(2): 259--275

  44. [52]

    Vidgen, B.; and Yasseri, T. 2020. Detecting weak and strong Islamophobic hate speech on social media. Journal of Information Technology & Politics, 17(1): 66--78

  45. [53]

    Walther, J. B. 2022. Social media and online hate. Current Opinion in Psychology, 45: 101298

  46. [54]

    Wang, C.-C.; Day, M.-Y.; and Wu, C.-L. 2022. Political hate speech detection and lexicon building: A study in taiwan. IEEE Access, 10: 44337--44346

  47. [55]

    Ward, S.; and McLoughlin, L. 2020. Turds, traitors and tossers: the abuse of UK MPs via Twitter. The Journal of Legislative Studies, 26(1): 47--73

  48. [56]

    Waseem, Z.; Davidson, T.; Warmsley, D.; and Weber, I. 2017. Understanding Abuse: A Typology of Abusive Language Detection Subtasks. In Proceedings of the First Workshop on Abusive Language Online, 78--84

  49. [57]

    Wilby, D.; Karmakharm, T.; Roberts, I.; Song, X.; and Bontcheva, K. 2023. GATE Teamware 2: An open-source tool for collaborative document classification annotation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: ...

  50. [58]

    Zampieri, M.; Malmasi, S.; Nakov, P.; Rosenthal, S.; Farra, N.; and Kumar, R. 2019. SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval). In Proceedings of the 13th International Workshop on Semantic Evaluation, 75--86

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.