Pith. sign in

REVIEW 4 major objections 5 minor 19 references

Web(er) of Hate: A Survey on How Hate Speech Is Typed

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read No single definition of hate speech exists, 135-dataset survey argues.

desk verdict Worth refereeing, not yet citable: the Weberian framing is genuinely useful, but the survey's headline counts are too shaky to anchor the argument. read the letter →

arxiv 2506.16190 v1 pith:WIWUVCYY submitted 2025-06-19 cs.CL

classification cs.CL
keywords hatespeechdatasetsdatasetcurationidealtypesMaxWeberreflexiveapproachannotationsurveynaturallanguageprocessing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey examines 135 published hate speech datasets and argues that the diversity in how hate speech is defined, labelled, and collected is not a defect to be repaired but an unavoidable consequence of curators' frames of reference. Drawing on Max Weber's ideal types of social action, the paper claims that no fully accurate and comprehensive decomposition of hate speech can exist. It therefore proposes a reflexive approach to dataset creation, in which researchers document their value judgments, annotator composition, and design trade-offs rather than pursuing definitional completeness. The result matters because hate speech detection models inherit whatever assumptions their training data encode, and cross-dataset generalisation failures may be symptoms of these differences rather than solvable technical problems.

What carries the argument

The central mechanism is Max Weber's ideal types of social action, an analytical heuristic that sorts complex social phenomena into four non-exclusive motivational types: goal-rational (strategic calculation), value-rational (belief-driven), affectual (emotion-driven), and traditional (habit and custom). The paper uses this lens in two ways: it interprets hate speech content as expressing these motivations, and it treats each dataset itself as an ideal-typical construction shaped by its curator's frame of reference. The ideal-type framework does the argument's work by turning the observed heterogeneity of datasets from a problem into an expected and legitimate outcome, grounding the call for reflexive documentation.

What would settle it

A complete census of hate speech datasets built by snowball citation tracing, multilingual scholarly databases, and industry sources that showed English-language datasets declining since 2023 and definitions converging would refute the survey's empirical generalisations; alternatively, exhibiting a single annotation scheme whose labels and guidelines reproduce every other dataset's labels at ceiling accuracy would refute the claim that no fully accurate comprehensive decomposition exists.

Watch

Extended reading notes

Core claim

The paper's central claim is that every hate speech dataset is an ideal-typical construction: an observer's purposeful simplification that foregrounds some aspects of hate speech and suppresses others. Because these constructions depend on the curator's frame of reference, the paper argues, no operationalisation can fully encapsulate hate speech. It supports this by reviewing 135 datasets and showing persistent heterogeneity in definitions (23 report none, 71 borrow prior ones, 41 state their own), goals, languages, collection methods, annotation tasks, annotator demographics, and quality assurance. It also finds that English dominance has not declined, in contrast to an earlier review, and that most datasets do not report annotator demographics or quality assurance. The prescriptive conclusion is that researchers should treat datasets as purpose-relative ideal types and explicitly document their perspectives and assumptions, and that annotation aggregation by majority vote is inappropriate when the goal is to capture the diversity of human judgments.

Load-bearing premise

The empirical patterns reported for the 135 datasets depend on a sample gathered from the community-maintained catalogue plus the first three pages of Google Scholar results for two venues and a general query; if that sample over-represents English-language and publicly available corpora, the observed stability of English dominance and the counts of unstated definitions could be artefacts of the search rather than properties of the field.

Editorial extensions

If this is right

  • If the claim is right, researchers should stop trying to produce a single all-purpose hate speech definition or benchmark and instead evaluate datasets and models against the specific purpose for which they were built.
  • Dataset papers would be expected to include a reflexivity statement documenting curatorial stance, definitional choices, annotator demographics, and known omissions, much as they now report annotation statistics.
  • Majority-vote aggregation should be avoided or supplemented when a dataset aims to capture descriptive diversity of opinion, since it erases exactly the disagreement that is informative.
  • Cross-dataset and cross-domain generalisation failures should be interpreted as evidence of differing ideal-typical constructions, not simply as model shortcomings.
  • Annotator pool design should be matched to the guideline paradigm: small hand-picked pools for prescriptive guidelines, diverse crowdsourced pools for descriptive ones, with demographic reporting to verify actual diversity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not say so, but its reflexive documentation proposal could be operationalised as a standardised 'dataset stance' statement, analogous to model cards, covering definitional basis, target categories, annotator demographics, and disagreement policy.
  • The framework should transfer to other subjective annotation tasks such as misinformation, sentiment, or toxicity, where similar diversity across curators is likely and would benefit from the same documentation discipline.
  • The stability of English dominance is reported against a search restricted to English-language queries and specific venues; a broader multilingual census might find different trends, which would test the strength of that particular empirical observation.
  • A testable extension would be to measure whether datasets that explicitly document their ideal-typical assumptions actually produce more predictable cross-dataset transfer than those that do not.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper surveys 135 hate speech datasets retrieved from the community-maintained hatespeechdata.com catalogue and supplementary Google Scholar searches, and it codes them along dimensions including definitions, goals, languages, data collection, annotation practices, quality assurance, and ethics. Drawing on Max Weber's notion of ideal types, the authors argue that hate speech operationalisation is inherently subjective and that 'a fully accurate and comprehensive decomposition of hate speech might not exist'; instead, they recommend a reflexive approach in which researchers explicitly document their perspectives and assumptions. The paper also offers recommendations about annotator composition, annotation aggregation, and reporting of demographics, and it applies Weber's four types of social action to interpret hate speech content.

Significance. If the argument is accepted, the paper provides a principled reframing of dataset diversity in hate speech research: rather than treating heterogeneity as a defect or pursuing definitional completeness, researchers should view each dataset as an ideal-typical construction tied to the curator's frame of reference. This is a potentially useful contribution to the methodology literature on abusive language datasets. The paper's strengths include its broad coverage relative to earlier curation-focused surveys, its explicit selection criteria, its detailed appendix tables, and its honest limitations section. The Weberian lens yields concrete, falsifiable recommendations (e.g., report annotator demographics, avoid majority-vote aggregation under descriptive paradigms) that could guide future dataset documentation. However, the empirical counts that motivate the recommendations contain numerous internal inconsistencies, and the survey's sampling strategy is acknowledged to be selection-sensitive; both issues need to be addressed before the empirical basis can be considered reliable.

major comments (4)
  1. [§5.5.1, Table 6] The task-formulation counts in Table 6 are internally inconsistent: the listed row counts sum to 133 rather than 135 datasets, and Trajano et al. (2024) appears in both the 'Hierarchical' and 'not reported' rows while Shekhar et al. (2022) appears in both 'Multi-label' and 'not reported'. Because the paper relies on these counts to characterise the diversity of task formulations, the table should be corrected and the coding criteria should state whether the categories are mutually exclusive or overlapping.
  2. [§5.5.5, Table 10] The quality-assurance counts disagree between the text and the appendix: §5.5.5 reports n=69 for datasets that 'do not report or are unclear about their QA procedures', while Table 10 lists 70 datasets in that row. In addition, the row 'Metrics-based selection (crowdsourcing)' reports n=10 but lists only 7 datasets, and the 'Tests' row reports n=12 but lists 11. These discrepancies affect the claim that 'around half' of datasets lack reported QA, so the counts should be reconciled.
  3. [§3] The statement that the survey is 'the most comprehensive to date' is difficult to reconcile with the paper's own citation of Yu et al. (2024), which reviews 492 datasets. The claim should be qualified to specify the sense in which the survey is more comprehensive (e.g., the number of curation-process dimensions covered) or removed.
  4. [§4, §5.3] The retrieval strategy—Google Scholar restricted to the first three pages, the literal term 'dataset', and two venues, with no snowballing—may introduce selection bias, and the paper itself demonstrates in §5.3 that its English-dominance finding is sensitive to search scope relative to Tonneau et al. (2024). Since the headline percentages in §5.1 (17% no definition) and §5.5.3 (58% no annotator demographics) motivate the reflexive-approach recommendation, the paper should either add a robustness check (e.g., compare with a snowballed or expanded search) or explicitly temper the empirical claims to acknowledge that they are sample-dependent.
minor comments (5)
  1. [Table 3] The language label 'Barzilian Portuguese' is a typo and should read 'Brazilian Portuguese'.
  2. [§5.4] The sentence 'All but one of the very large datasets (n = 7), which contain entries numbering in the millions, do not not use any filtering' contains a double negative; it should read 'do not use any filtering'.
  3. [§5.6] The phrase 'most of the more recent papers are least partly addressing ethical issues' should be 'are at least partly addressing'.
  4. [§5.2, Table 2] The goal category 'enabling comparison studies' is reported as n=3, but Table 2 lists only two datasets (Waseem 2016; Basile et al. 2019); please correct the count or the list.
  5. [Table 3] The Italian row reports a count of 8 but lists only 6 datasets; please correct the count or the list.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical survey counts are observational, the Weberian frame is an openly adopted interpretive lens, and no central claim reduces to its inputs by construction.

full rationale

This paper is a survey and interpretive argument, not a derivation chain with fitted parameters or predictive claims. The headline counts (e.g., 23/135 with no definition, 78/135 with no annotator demographics) are direct observations of the retrieved dataset descriptions, and the normative recommendation to document perspectives and assumptions is argued from the openly adopted Weberian framework rather than extracted from those counts by construction. There are no author self-citations serving as load-bearing evidence, no invoked uniqueness theorems from prior work by the same authors, and no ansatz smuggled in via citation. The Limitations section explicitly acknowledges that the discussion of ideal types and annotation paradigms is interpretative and that alternative frameworks could yield different viewpoints, which further undercuts any suggestion that the conclusion is presented as a forced empirical result. The potential internal inconsistencies in the appendix coding (e.g., Trajano et al. listed under both 'hierarchical' and 'not reported'; task-type counts summing to 133 over 135 datasets) are reliability concerns about the survey's manual coding, not evidence that the argument is circular. Likewise, the discrepancy with Tonneau et al. (2024) on English-language dominance is attributed by the authors themselves to differing search scopes and methods, which is a candid acknowledgment of methodology dependence rather than a circular step. Overall, the reasoning is self-contained: the empirical observations stand independently of the Weberian interpretation, and the interpretation is not disguised as a prediction derived from those observations.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters and no new entities, such as particles or theoretical constructs. Its central claim rests on three domain assumptions: the applicability of the Weberian analogy, the representativeness of the surveyed sample, and the value of reflexive documentation.

assumptions (3)
  • domain assumption Weber's ideal types are a valid analytical lens for understanding hate speech dataset curation.
    The entire framework imports a sociological concept from Weber (1904, 1930, 1978) into NLP dataset design. If the analogy is inapt, the central recommendation loses its theoretical grounding. Invoked throughout Sections 2 and 6.
  • domain assumption The 135 retrieved datasets are representative of the hate speech dataset landscape.
    All empirical trends, such as the prevalence of definitions, annotator demographics, and English dominance, are drawn from this convenience sample, which is constrained by the search protocol in Section 4. A biased sample would undermine the observed patterns.
  • domain assumption Documenting value judgments will improve transparency and methodological rigour.
    The reflexive approach assumes that explicit documentation of curator perspectives is both feasible and beneficial. This is a normative stance, not an empirically established outcome, as acknowledged in the Limitations section.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Web(er) of Hate: A Survey on How Hate Speech Is Typed." pith.science (2026). https://pith.science/paper/WIWUVCYY

@misc{pith2026250616190,
  author       = {Pith},
  title        = {Pith review of: Web(er) of Hate: A Survey on How Hate Speech Is Typed},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WIWUVCYY}},
  note         = {Machine review of arXiv:2506.16190}
}
read the original abstract

The curation of hate speech datasets involves complex design decisions that balance competing priorities. This paper critically examines these methodological choices in a diverse range of datasets, highlighting common themes and practices, and their implications for dataset reliability. Drawing on Max Weber's notion of ideal types, we argue for a reflexive approach in dataset creation, urging researchers to acknowledge their own value judgments during dataset construction, fostering transparency and methodological rigour.

Figures

Figures reproduced from arXiv: 2506.16190 by the authors.

Figure 1
Figure 1. The number of datasets published in each year, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. A prototypical hierarchical categorisation of hate speech taxonomy. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

19 extracted references · 16 canonical work pages

  1. [2]

    In 2018 IEEE/ACM International Confer- ence on Advances in Social Networks Analysis and Mining (ASONAM), pages 69–76

    Are they our brothers? Analysis and detec- tion of religious hate speech in the Arabic twitter- sphere. In 2018 IEEE/ACM International Confer- ence on Advances in Social Networks Analysis and Mining (ASONAM), pages 69–76. Abdullah Albanyan, Ahmed Hassan, and Eduardo Blanco. 2023. Not all counterhate tweets elicit the same replies: A fine-grained analysis....

  2. [6]

    Overview of the task on automatic misogyny identification at IberEval 2018. In Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval 2018), 34th Conference of the Spanish Society for Natural Language Processing (SEPLN 2018), page 214–228. International World Wide Web Conferences Steering Committee. ...

  3. [8]

    In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 138–149, Online

    Towards a comprehensive taxonomy and large- scale annotated corpus for online slur usage. In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 138–149, Online. Association for Computational Linguistics. Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, and Alice Oh. 2024. Exploring cross-cultural differenc...

  4. [13]

    In Proceedings of the Twelfth Lan- guage Resources and Evaluation Conference, pages 3498–3508, Marseille, France

    Offensive language and hate speech detec- tion for Danish. In Proceedings of the Twelfth Lan- guage Resources and Evaluation Conference, pages 3498–3508, Marseille, France. European Language Resources Association. Aakash Singh, Deepawali Sharma, and Vivek Kumar Singh. 2024. MIMIC: Misogyny identification in multimodal internet content in Hindi-English cod...

  5. [14]

    In Proceedings of the Thirteenth Language Resources and Evaluation Conference , pages 2215–2225, Marseille, France

    Large-scale hate speech detection with cross- domain transfer. In Proceedings of the Thirteenth Language Resources and Evaluation Conference , pages 2215–2225, Marseille, France. European Lan- guage Resources Association. Douglas Trajano, Rafael H. Bordini, and Renata Vieira

  6. [15]

    Language Re- sources and Evaluation, 58(4):1263–1289

    OLID-BR: offensive language identification dataset for Brazilian Portuguese. Language Re- sources and Evaluation, 58(4):1263–1289. Francielle Vargas, Samuel Guimarães, Shamsud- deen Hassan Muhammad, Diego Alves, Ibrahim Said Ahmad, Idris Abdulmumin, Diallo Mohamed, Thiago Pardo, and Fabrício Benevenuto. 2024. HausaHate: An expert annotated corpus for Haus...

  7. [16]

    In The 7th Workshop on Online Abuse and Harms (WOAH), pages 202–214, Toronto, Canada

    HOMO-MEX: A Mexican Spanish anno- tated corpus for LGBT+phobia detection on Twitter. In The 7th Workshop on Online Abuse and Harms (WOAH), pages 202–214, Toronto, Canada. Associa- tion for Computational Linguistics. Bertie Vidgen and Leon Derczynski. 2021. Directions in abusive language training data, a systematic re- view: Garbage in, garbage out. PLOS O...

  8. [17]

    In Social Informatics: 12th International Conference, SocInfo 2020, Pisa, Italy, October 6–9, 2020, Proceedings , page 427–439, Berlin, Heidelberg

    ALONE: A dataset for toxic behavior among adolescents on Twitter. In Social Informatics: 12th International Conference, SocInfo 2020, Pisa, Italy, October 6–9, 2020, Proceedings , page 427–439, Berlin, Heidelberg. Springer-Verlag. Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017. Ex Machina: Personal attacks seen at scale. In Pro- ceedings of the 26th ...

Show all 19 references
  1. [18]

    Predicting the type and target of offensive posts in social media. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1415–1420, Minneapolis,...

  2. [19]

    mixed languages

    Reducing unintended identity bias in Russian 16 hate speech detection. In Proceedings of the Fourth Workshop on Online Abuse and Harms, pages 65–69, Online. Association for Computational Linguistics. A Appendix 17 A.1 Breakdowns of Reviewed Datasets Datasets Subcategories Basi...

  3. [2016]

    Call me sexist, but

    Measuring the reliability of hate speech anno- tations: The case of the European refugee crisis. Ramsha Saeed, Hammad Afzal, Sadaf Abdul Rauf, and Naima Iltaf. 2023. Detection of offensive language and its severity for low resource language. ACM Trans. Asian Low-Resour. Lang. ...

  4. [2017]

    In Proceedings of the First Workshop on Abu- sive Language Online, pages 52–56, Vancouver, BC, Canada

    Abusive language detection on Arabic social media. In Proceedings of the First Workshop on Abu- sive Language Online, pages 52–56, Vancouver, BC, Canada. Association for Computational Linguistics. Hala Mulki and Bilal Ghanem. 2021. Let-mi: An Ara- bic Levantine Twitter dataset...

  5. [2018]

    Procedia Computer Science, 142:174–181

    Dataset construction for the detection of anti- social behaviour in online communication in Arabic. Procedia Computer Science, 142:174–181. Arabic Computational Linguistics. Nuha Albadi, Maram Kurdi, and Shivakant Mishra

  6. [2019]

    In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA

    SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA. Association for Compu- tational Linguistics. Mohit Bhardwaj...

  7. [2020]

    Accademia University Press

    AMI @ EVALITA2020: Automatic Misogyny Identification, page 21–28. Accademia University Press. Elisabetta Fersini, Paolo Rosso, and Maria Anzovino

  8. [2021]

    HateCheck: Functional tests for hate speech detection models. In Proceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online....

  9. [2022]

    Money Rules the World, but Who Rules the Money?

    Detecting abusive Albanian. Preprint, arXiv:2107.13592. Anaïs Ollagnier, Elena Cabrio, Serena Villata, and Catherine Blaya. 2022. CyberAgressionAdo-v1: a dataset of annotated online aggressions in French col- lected through a role-playing game. In Proceedings of the Thirteenth...

  10. [2023]

    In Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023), pages 1–13, Dubrovnik, Croatia

    Analyzing zero-shot transfer scenarios across Spanish variants for hate speech detection. In Tenth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial 2023), pages 1–13, Dubrovnik, Croatia. Association for Computational Linguistics. Amanda Cercas Curry, Gavi...

  11. [2024]

    Improving adversarial data collection by sup- porting annotators: Lessons from GAHD, a German hate speech dataset. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.