Pith. sign in

REVIEW 2 major objections 6 minor 13 references

Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech

T0 review · 2 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Counterspeech NLP has drifted from the communities it aims to serve, a review of 74 studies argues.

desk verdict A solid, reproducible survey plus honest NGO data; the 'growing disconnect' claim needs a year-wise breakdown or a softer verb. read the letter →

arxiv 2508.04638 v1 pith:22ZNK36I submitted 2025-08-06 cs.CL

classification cs.CL
keywords counterspeechhatespeechparticipatorydesignstakeholderinvolvementonlinegender-basedviolencesystematicreviewdatasetreusenaturallanguageprocessing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that counterspeech research in NLP has drifted from its original stakeholder-centred mission. Reviewing 74 studies, it finds that close to half reuse existing datasets, that two-thirds of those reuse the CONAN family of nichesourced counterspeech datasets created between 2019 and 2021, and that most human involvement comes from the researchers themselves or paid crowdworkers rather than from communities affected by hate speech. Through focus groups with five UK NGOs that fight online gender-based violence, it shows that practitioners reason about when hate speech was posted, how far it has spread, who the perpetrator is relative to the target, and which strategy fits which role—information current counterspeech datasets do not carry. If the paper is right, benchmark-driven counterspeech research is optimising for the wrong inputs: it ignores the contextual metadata that real-world responders treat as decisive. The paper's remedy is a concrete set of data-collection recommendations intended to re-centre stakeholder expertise.

What carries the argument

The argument runs on two coupled instruments. First, a systematic review protocol that codes each of 74 papers for where its hate-speech and counterspeech data come from (crawling, crowdsourcing, nichesourcing, automated generation, existing datasets), what human input occurred at which stage, and who supplied it (NGO workers, crowdworkers, academics, or no one). That coding exposes the field's dependence on the CONAN dataset family and the absence of affected-community participation. Second, structured interactive focus groups with five UK oGBV charities, run with feminist co-creation and participatory action design practices, produce a list of stakeholder-informed practices: treating metad

What would settle it

Re-analyse the 74 reviewed papers by publication year and compute the share with stakeholder involvement (the '✓' rows in Tables 1 and 5). The paper's 'growing disconnect' claim predicts this share falls over time; if later papers are no less likely to involve NGOs or affected communities than the 2019–2021 CONAN work, the claim is refuted.

Watch

Extended reading notes

Core claim

A systematic review of 74 NLP publications on counterspeech finds that the field has moved away from the participatory, NGO-centred approach that produced its earliest resources. Close to half of the surveyed studies draw hate speech or counterspeech from existing datasets, and roughly two-thirds of those reuse the CONAN family (CONAN, Multi-Target CONAN, DIALOCONAN, MTKGCONAN); stakeholder involvement beyond the research team is rare, with many resources relying on the authors themselves or other academics for annotation and evaluation. In parallel, structured focus groups with five UK NGOs working on online gender-based violence (oGBV) surface practices the datasets ignore: the date, reach

Load-bearing premise

The recommendations depend on the assumption that five UK charities focused on online gender-based violence speak for the full range of communities that counterspeech is meant to serve—including religious, racial, and LGBTQ+ groups—which the paper's limitations section explicitly acknowledges may not hold.

Editorial extensions

If this is right

  • Counterspeech generation systems should be conditioned on hate-speech metadata—creation date, views, shares, and perpetrator reach—not just on the text of the hate speech.
  • Datasets should record the sub-category of abuse (e.g. harassment vs. dogpiling) and the role being addressed (perpetrator, target, bystander), since strategies differ by role.
  • Evaluation of counterspeech should involve stakeholders or bystander perspectives rather than relying only on automated metrics and academic annotation.
  • Reusing CONAN-family datasets without adding stakeholder input limits progress: models may hit a ceiling and inherit outdated examples of hate speech.
  • NLP counterspeech should treat NGOs as partners in dataset creation and resource compilation, not merely as annotators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the 'growing disconnect' claim is available in the paper's own data: if the proportion of papers with NGO or affected-community involvement is plotted against publication year, the paper's conclusion predicts a downward trend; computing that trend would confirm or refute it.
  • The metadata priorities reported here—perpetrator reach, familiarity, and abuse sub-category—are plausibly shared by counterspeech aimed at religious, racial, and LGBTQ+ hate, but that generalisation is untested; replicating the focus groups with those communities would show whether the recommended features are universal or oGBV-specific.
  • The stakeholder resistance to gendered bot personas suggests that anthropomorphism is not a neutral design choice: deploying female-presenting counterspeech bots could undermine the very messages they deliver, a risk that current generation and evaluation pipelines do not model.
  • Adopting the paper's recommendations would push the field from fully automated pipelines toward human-in-the-loop systems, which may improve safety and relevance but would also raise questions about scalability and who bears the labour of stakeholder consultation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper reports a PRISMA-style systematic review of 74 NLP publications on counterspeech, asking to what extent affected stakeholders are represented in dataset creation, model development, and evaluation. It complements the review with a participatory case study involving five UK NGOs that work on online Gender-Based Violence, conducted through structured focus groups. The authors quantify dataset reuse (about half of resources reuse existing datasets; of those, ~66% use CONAN-family datasets), report that unambiguous stakeholder involvement is rare (10 'Yes' out of 74, with 9 'Possibly'), and derive recommendations around contextual metadata, fine-grained hate-speech subcategories, and attention to roles (perpetrator/target/bystander). The paper's central claim is that there is a 'growing disconnect' between NLP counterspeech research and community needs, and it calls for re-centring stakeholder expertise.

Significance. If the review's conclusions hold—at least in the weaker form of a persistent and widespread disconnect—this is a valuable contribution to counterspeech research and participatory NLP. The full resource table and the explicit coding of data sources and stakeholder involvement are a useful reference for the community. The case study is methodologically strong: IRB approval, informed consent, compensation, and concrete NGO-derived insights (e.g., algospeak, perpetrator reach, gendered bot concerns) that are rarely discussed in the NLP counterspeech literature. The paper also makes a fair point that reuse of legacy CONAN-family datasets can create a ceiling effect and contamination risk, though this part is speculative and the authors themselves flag it as such.

major comments (2)
  1. [Abstract; §3.1; §5; Figure 4] The headline claim of a 'growing disconnect' is a longitudinal claim, but the review is a cross-sectional snapshot (searches in March 2025) and no temporal breakdown of stakeholder involvement is reported. Figure 4 only gives raw counts of publications per year; it does not show the proportion of resources with stakeholder involvement or the rate of dataset reuse over time. The aggregate statistics (e.g., ~66% of resources using existing datasets use CONAN variants; 10 'Yes' vs. 9 'Possibly' in Figure 3) establish a persistent or widespread disconnect, but not growth. Since the abstract, introduction, and conclusions ('there has been something of a downturn' in §5) rest on this temporal claim, please either remove the 'growing'/'downturn' language or add the missing temporal analysis, e.g., stakeholder-involvement rate by two-year bins, with appropriate caution about partial-year 2025 da
  2. [§4.2; Limitations] The concrete recommendations in §4.2 are valuable, but they are presented as general data features for counterspeech research (e.g., sub-category of hate, roles, illegal language) when the evidence base is five UK NGOs focused on oGBV. The Limitations section acknowledges that religious and LGBTQ+ communities, among others, may not be represented, but the main text and abstract do not carry this caveat. Please hedge the recommendations as oGBV-derived and propose them as hypotheses to be validated with other stakeholder groups, rather than as the implied missing requirements for all counterspeech NLP.
minor comments (6)
  1. [Tables 5 and 6] Tables 5 and 6 appear to duplicate each other. Consolidate them into one canonical appendix table and refer to it consistently.
  2. [Figure 3] The stacked-bar figure is hard to parse. The x-axis labels are ambiguous (e.g., 'academics authors', 'crowdworkersno human input') and the caption does not explain the categories. Add a clear legend and axis description.
  3. [Figure 4] Since searches ended in March 2025, the 2025 bar is a partial-year count. Add a note or marker so readers do not compare it directly to full years.
  4. [§3.1, Figure 2] Report the underlying counts alongside the percentages (N=88 sources) and state explicitly which denominator is used, since the text and figure captions currently leave room for ambiguity.
  5. [§3.1] The term 'nichesourcing' is used without a definition. Define it at first use in the Background or Methods.
  6. [Tables 5/6] Small typographical issues: 'Rodrguez et al. (2023)' is missing an accent, and 'Bonaldi et al. (2024b).' has a stray period in Table 5.

Circularity Check

0 steps flagged · score 1.0 of 10

No substantive circularity; the review's central claim is an empirical literature assessment, not a derivation from its own inputs.

full rationale

The paper's main claims are empirical: a systematic review of 74 counterspeech papers and a participatory focus-group case study with five NGOs. The headline 'growing disconnect' is a temporal inference, and the paper does not provide a year-wise cross-tabulation of stakeholder involvement, so that particular claim is under-supported. However, under-support is a validity concern, not circularity. No equation or fitted parameter is reused as a prediction; no unique-solution theorem is imported; no ansatz is smuggled in through citation. The self-citations (e.g., Abercrombie et al. 2023b for shortcomings of oGBV datasets, Bonaldi et al. 2024a as a prior survey) are used as background and motivation, not as the load-bearing evidence for the paper's central finding. The central statistics—66% of reused datasets are CONAN variants, roughly 26 of 74 resources use academics/authors as annotators, and only a small number have clear stakeholder involvement—are coded directly from the surveyed papers in Tables 1, 5, and 6 and Figure 3. The focus-group findings are primary qualitative data. The limitations section explicitly acknowledges the narrow NGO sample and the lack of experimental validation, further indicating that the authors do not present the case study as a formal derivation. No step reduces to its own input by definition, and the paper is not circular in the sense targeted by this analysis.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No mathematical free parameters or invented entities appear. The paper's argument rests on the representativeness and reliability of its review corpus and its small NGO sample, both of which are stated as limitations.

assumptions (3)
  • domain assumption DBLP search with the specified keyword list and inclusion/exclusion criteria identifies the relevant population of NLP counterspeech research.
    The entire review corpus (74 papers) is built on this assumption; the Limitation section acknowledges restriction to peer-reviewed NLP/computational social science publications, potentially missing workshop papers or non-DBLP venues.
  • domain assumption The authors' qualitative labeling of 'stakeholder involvement', 'expert', and 'possibly bystander' across the reviewed papers is trustworthy without inter-annotator agreement.
    Section 3.1 describes the labeling as done by two authors and cross-checked by a third, but no reliability statistic or detailed codebook is provided. All subsequent percentages (e.g., 66% CONAN reuse) depend on this coding.
  • domain assumption Five UK-based NGOs focused on online gender-based violence provide insights generalizable to counterspeech for other hate speech domains.
    The case study in Section 4 draws all recommendations from these five NGOs; the Limitations section explicitly notes they may not capture perspectives of religious or LGBTQ+ groups.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech." pith.science (2026). https://pith.science/paper/22ZNK36I

@misc{pith2026250804638,
  author       = {Pith},
  title        = {Pith review of: Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/22ZNK36I}},
  note         = {Machine review of arXiv:2508.04638}
}
read the original abstract

Counterspeech, i.e. the practice of responding to online hate speech, has gained traction in NLP as a promising intervention. While early work emphasised collaboration with non-governmental organisation stakeholders, recent research trends have shifted toward automated pipelines that reuse a small set of legacy datasets, often without input from affected communities. This paper presents a systematic review of 74 NLP studies on counterspeech, analysing the extent to which stakeholder participation influences dataset creation, model development, and evaluation. To complement this analysis, we conducted a participatory case study with five NGOs specialising in online Gender-Based Violence (oGBV), identifying stakeholder-informed practices for counterspeech generation. Our findings reveal a growing disconnect between current NLP research and the needs of communities most impacted by toxic online content. We conclude with concrete recommendations for re-centring stakeholder expertise in counterspeech research.

Figures

Figures reproduced from arXiv: 2508.04638 by the authors.

Figure 1
Figure 1. Search and selection protocol. counterspeech tools (Mun et al., 2024a). In this re￾view, we uncover the extent to which stakeholders participate in NLP counterspeech research design and resource creation. Online Gender-Based Violence or oGBV is a framework used by international organisations such as the UN and WHO, and covers harmful effects on all genders, particularly women.3 Misogynistic abuse affects around 50% … view at source ↗
Figure 2
Figure 2. Counterspeech sources and datasets. The percentage reflects the proportion of total sources (N = 88), [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Stakeholder identity by participation. ‘Possi [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: shows the resources we surveyed by pub￾lication year, with a notable recent spike. B Full table of resources for counterspeech [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

13 extracted references · 11 canonical work pages

  1. [5]

    In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5792–5809, Toronto, Canada

    Counterspeeches up my sleeve! intent dis- tribution learning and persistent fusion for intent- conditioned counterspeech generation. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5792–5809, Toronto, Canada. Association for Computational Linguistics. Sadaf Md. Halim, Saquib Irtiz...

  2. [6]

    CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs

    Reinforcement learning-based counter- misinformation response generation: a case study of COVID-19 vaccine misinformation. In Proceedings of the ACM Web Conference 2023, pages 2698–2709. Amey Hengle, Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. CSEval: Towards auto- mated, multi-dimensional, and reference-free coun- terspeech evaluation us...

  3. [7]

    In Proceedings of the 2022 Conference on Empiri- cal Methods in Natural Language Processing, pages 10818–10833, Abu Dhabi, United Arab Emirates

    KOLD: Korean offensive language dataset. In Proceedings of the 2022 Conference on Empiri- cal Methods in Natural Language Processing, pages 10818–10833, Abu Dhabi, United Arab Emirates. As- sociation for Computational Linguistics. Aiqi Jiang, Xiaohan Yang, Yang Liu, and Arkaitz Zu- biaga. 2022. SWSR: A chinese dataset and lexicon for online sexism detecti...

  4. [8]

    Korean Online Hate Speech Dataset for Multilabel Classification: How Can Social Science Improve Dataset on Hate Speech?

    Korean online hate speech dataset for mul- tilabel classification: How can social science im- prove dataset on hate speech? arXiv preprint arXiv:2204.03262. Hannah Kirk, Wenjie Yin, Bertie Vidgen, and Paul Röttger. 2023. SemEval-2023 task 10: Explainable detection of online sexism. In Proceedings of the 17th International Workshop on Semantic Evaluation (...

  5. [9]

    Trauma, Violence, & Abuse, 24(3):1727– 1742

    A systematic review exploring variables re- lated to bystander intervention in sexual violence contexts. Trauma, Violence, & Abuse, 24(3):1727– 1742. Binny Mathew, Navish Kumar, Pawan Goyal, and Ani- mesh Mukherjee. 2020. Interaction dynamics be- tween hate and counter users on Twitter.Proceedings of the 7th ACM IKDD CoDS and 25th COMAD. Binny Mathew, Nav...

  6. [10]

    In NeurIPS 2023 Computational Sustainability: Promises and Pitfalls from Theory to Deployment

    AI for whom? Shedding critical light on AI for social good. In NeurIPS 2023 Computational Sustainability: Promises and Pitfalls from Theory to Deployment. David L. Morgan. 1996. Focus groups. Annual Review of Sociology, 22(V olume 22, 1996):129–152. Michael J Muller and Sarah Kuhn. 1993. Participatory design. Communications of the ACM, 36(6):24–28. Jimin ...

  7. [13]

    Proceedings of the 2021 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining

    Racism is a virus: anti-asian hate and counter- speech in social media during the COVID-19 crisis. Proceedings of the 2021 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining. I. Zubiaga, A. Soroa, and R. Agerri. 2024a. A LLM- based ranking method for the evaluation of automatic counter-narrative generation. pages 9572–958...

  8. [2017]

    Marcus Tomalin and Stefanie Ullmann, editors

    #DistractinglySexy: How Social Media was used as a Counter Narrative on Gender in STEM. Marcus Tomalin and Stefanie Ullmann, editors. 2023. Counterspeech. Multidisciplinary Perspectives on Countering Dangerous Speech. Taylor & Francis. Vittoria Tonini, Simona Frenda, M. Stranisci, and Vi- viana Patti. 2024. How do we counter hate speech in Italy? María Es...

Show all 13 references
  1. [2019]

    In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA

    SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA. Association for Compu- tational Linguistics. Susan Benesch,...

  2. [2020]

    In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 102–112, Online

    Countering hate on social media: Large scale classification of hate and counter speech. In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 102–112, Online. Association for Computational Linguistics. Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Dur- mu...

  3. [2022]

    arXiv preprint arXiv:2203.03584

    Counter hate speech in social media: A survey. arXiv preprint arXiv:2203.03584. Ghadi Alyahya and Abeer Aldayel. 2024. Hatred stems from ignorance! Distillation of the persuasion modes in countering conversational hate speech. ArXiv preprint, abs/2403.15449. I. Arpinar, Ugur K...

  4. [2023]

    EPJ Data Science, 12:1

    Correction: Impact and dynamics of hate and counter speech online. EPJ Data Science, 12:1. Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent Hébert-Dufresne, and Mirta Galesic

  5. [2024]

    ArXiv preprint, abs/2412.15453

    Northeastern Uni at multilingual counter- speech generation: Enhancing counter speech gener- ation with LLM alignment through direct preference optimization. ArXiv preprint, abs/2412.15453. Haiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao, Minghao Hu, and Bin Zhou. 2024a. F2RL...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.