REVIEW 2 major objections 6 minor 13 references
Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech
T0 review · 2 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Counterspeech NLP has drifted from the communities it aims to serve, a review of 74 studies argues.
desk verdict A solid, reproducible survey plus honest NGO data; the 'growing disconnect' claim needs a year-wise breakdown or a softer verb. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on two coupled instruments. First, a systematic review protocol that codes each of 74 papers for where its hate-speech and counterspeech data come from (crawling, crowdsourcing, nichesourcing, automated generation, existing datasets), what human input occurred at which stage, and who supplied it (NGO workers, crowdworkers, academics, or no one). That coding exposes the field's dependence on the CONAN dataset family and the absence of affected-community participation. Second, structured interactive focus groups with five UK oGBV charities, run with feminist co-creation and participatory action design practices, produce a list of stakeholder-informed practices: treating metad
What would settle it
Re-analyse the 74 reviewed papers by publication year and compute the share with stakeholder involvement (the '✓' rows in Tables 1 and 5). The paper's 'growing disconnect' claim predicts this share falls over time; if later papers are no less likely to involve NGOs or affected communities than the 2019–2021 CONAN work, the claim is refuted.
Extended reading notes
Core claim
A systematic review of 74 NLP publications on counterspeech finds that the field has moved away from the participatory, NGO-centred approach that produced its earliest resources. Close to half of the surveyed studies draw hate speech or counterspeech from existing datasets, and roughly two-thirds of those reuse the CONAN family (CONAN, Multi-Target CONAN, DIALOCONAN, MTKGCONAN); stakeholder involvement beyond the research team is rare, with many resources relying on the authors themselves or other academics for annotation and evaluation. In parallel, structured focus groups with five UK NGOs working on online gender-based violence (oGBV) surface practices the datasets ignore: the date, reach
Load-bearing premise
The recommendations depend on the assumption that five UK charities focused on online gender-based violence speak for the full range of communities that counterspeech is meant to serve—including religious, racial, and LGBTQ+ groups—which the paper's limitations section explicitly acknowledges may not hold.
Editorial extensions
If this is right
- Counterspeech generation systems should be conditioned on hate-speech metadata—creation date, views, shares, and perpetrator reach—not just on the text of the hate speech.
- Datasets should record the sub-category of abuse (e.g. harassment vs. dogpiling) and the role being addressed (perpetrator, target, bystander), since strategies differ by role.
- Evaluation of counterspeech should involve stakeholders or bystander perspectives rather than relying only on automated metrics and academic annotation.
- Reusing CONAN-family datasets without adding stakeholder input limits progress: models may hit a ceiling and inherit outdated examples of hate speech.
- NLP counterspeech should treat NGOs as partners in dataset creation and resource compilation, not merely as annotators.
Reading between the lines
- A direct test of the 'growing disconnect' claim is available in the paper's own data: if the proportion of papers with NGO or affected-community involvement is plotted against publication year, the paper's conclusion predicts a downward trend; computing that trend would confirm or refute it.
- The metadata priorities reported here—perpetrator reach, familiarity, and abuse sub-category—are plausibly shared by counterspeech aimed at religious, racial, and LGBTQ+ hate, but that generalisation is untested; replicating the focus groups with those communities would show whether the recommended features are universal or oGBV-specific.
- The stakeholder resistance to gendered bot personas suggests that anthropomorphism is not a neutral design choice: deploying female-presenting counterspeech bots could undermine the very messages they deliver, a risk that current generation and evaluation pipelines do not model.
- Adopting the paper's recommendations would push the field from fully automated pipelines toward human-in-the-loop systems, which may improve safety and relevance but would also raise questions about scalability and who bears the labour of stakeholder consultation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a PRISMA-style systematic review of 74 NLP publications on counterspeech, asking to what extent affected stakeholders are represented in dataset creation, model development, and evaluation. It complements the review with a participatory case study involving five UK NGOs that work on online Gender-Based Violence, conducted through structured focus groups. The authors quantify dataset reuse (about half of resources reuse existing datasets; of those, ~66% use CONAN-family datasets), report that unambiguous stakeholder involvement is rare (10 'Yes' out of 74, with 9 'Possibly'), and derive recommendations around contextual metadata, fine-grained hate-speech subcategories, and attention to roles (perpetrator/target/bystander). The paper's central claim is that there is a 'growing disconnect' between NLP counterspeech research and community needs, and it calls for re-centring stakeholder expertise.
Significance. If the review's conclusions hold—at least in the weaker form of a persistent and widespread disconnect—this is a valuable contribution to counterspeech research and participatory NLP. The full resource table and the explicit coding of data sources and stakeholder involvement are a useful reference for the community. The case study is methodologically strong: IRB approval, informed consent, compensation, and concrete NGO-derived insights (e.g., algospeak, perpetrator reach, gendered bot concerns) that are rarely discussed in the NLP counterspeech literature. The paper also makes a fair point that reuse of legacy CONAN-family datasets can create a ceiling effect and contamination risk, though this part is speculative and the authors themselves flag it as such.
major comments (2)
- [Abstract; §3.1; §5; Figure 4] The headline claim of a 'growing disconnect' is a longitudinal claim, but the review is a cross-sectional snapshot (searches in March 2025) and no temporal breakdown of stakeholder involvement is reported. Figure 4 only gives raw counts of publications per year; it does not show the proportion of resources with stakeholder involvement or the rate of dataset reuse over time. The aggregate statistics (e.g., ~66% of resources using existing datasets use CONAN variants; 10 'Yes' vs. 9 'Possibly' in Figure 3) establish a persistent or widespread disconnect, but not growth. Since the abstract, introduction, and conclusions ('there has been something of a downturn' in §5) rest on this temporal claim, please either remove the 'growing'/'downturn' language or add the missing temporal analysis, e.g., stakeholder-involvement rate by two-year bins, with appropriate caution about partial-year 2025 da
- [§4.2; Limitations] The concrete recommendations in §4.2 are valuable, but they are presented as general data features for counterspeech research (e.g., sub-category of hate, roles, illegal language) when the evidence base is five UK NGOs focused on oGBV. The Limitations section acknowledges that religious and LGBTQ+ communities, among others, may not be represented, but the main text and abstract do not carry this caveat. Please hedge the recommendations as oGBV-derived and propose them as hypotheses to be validated with other stakeholder groups, rather than as the implied missing requirements for all counterspeech NLP.
minor comments (6)
- [Tables 5 and 6] Tables 5 and 6 appear to duplicate each other. Consolidate them into one canonical appendix table and refer to it consistently.
- [Figure 3] The stacked-bar figure is hard to parse. The x-axis labels are ambiguous (e.g., 'academics authors', 'crowdworkersno human input') and the caption does not explain the categories. Add a clear legend and axis description.
- [Figure 4] Since searches ended in March 2025, the 2025 bar is a partial-year count. Add a note or marker so readers do not compare it directly to full years.
- [§3.1, Figure 2] Report the underlying counts alongside the percentages (N=88 sources) and state explicitly which denominator is used, since the text and figure captions currently leave room for ambiguity.
- [§3.1] The term 'nichesourcing' is used without a definition. Define it at first use in the Background or Methods.
- [Tables 5/6] Small typographical issues: 'Rodrguez et al. (2023)' is missing an accent, and 'Bonaldi et al. (2024b).' has a stray period in Table 5.
Circularity Check
No substantive circularity; the review's central claim is an empirical literature assessment, not a derivation from its own inputs.
full rationale
The paper's main claims are empirical: a systematic review of 74 counterspeech papers and a participatory focus-group case study with five NGOs. The headline 'growing disconnect' is a temporal inference, and the paper does not provide a year-wise cross-tabulation of stakeholder involvement, so that particular claim is under-supported. However, under-support is a validity concern, not circularity. No equation or fitted parameter is reused as a prediction; no unique-solution theorem is imported; no ansatz is smuggled in through citation. The self-citations (e.g., Abercrombie et al. 2023b for shortcomings of oGBV datasets, Bonaldi et al. 2024a as a prior survey) are used as background and motivation, not as the load-bearing evidence for the paper's central finding. The central statistics—66% of reused datasets are CONAN variants, roughly 26 of 74 resources use academics/authors as annotators, and only a small number have clear stakeholder involvement—are coded directly from the surveyed papers in Tables 1, 5, and 6 and Figure 3. The focus-group findings are primary qualitative data. The limitations section explicitly acknowledges the narrow NGO sample and the lack of experimental validation, further indicating that the authors do not present the case study as a formal derivation. No step reduces to its own input by definition, and the paper is not circular in the sense targeted by this analysis.
Assumptions & free parameters
assumptions (3)
- domain assumption DBLP search with the specified keyword list and inclusion/exclusion criteria identifies the relevant population of NLP counterspeech research.
- domain assumption The authors' qualitative labeling of 'stakeholder involvement', 'expert', and 'possibly bystander' across the reviewed papers is trustworthy without inter-annotator agreement.
- domain assumption Five UK-based NGOs focused on online gender-based violence provide insights generalizable to counterspeech for other hate speech domains.
Cite this review
Pith. "Pith review of Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech." pith.science (2026). https://pith.science/paper/22ZNK36I
@misc{pith2026250804638,
author = {Pith},
title = {Pith review of: Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech},
year = {2026},
howpublished = {\url{https://pith.science/paper/22ZNK36I}},
note = {Machine review of arXiv:2508.04638}
}
read the original abstract
Counterspeech, i.e. the practice of responding to online hate speech, has gained traction in NLP as a promising intervention. While early work emphasised collaboration with non-governmental organisation stakeholders, recent research trends have shifted toward automated pipelines that reuse a small set of legacy datasets, often without input from affected communities. This paper presents a systematic review of 74 NLP studies on counterspeech, analysing the extent to which stakeholder participation influences dataset creation, model development, and evaluation. To complement this analysis, we conducted a participatory case study with five NGOs specialising in online Gender-Based Violence (oGBV), identifying stakeholder-informed practices for counterspeech generation. Our findings reveal a growing disconnect between current NLP research and the needs of communities most impacted by toxic online content. We conclude with concrete recommendations for re-centring stakeholder expertise in counterspeech research.
Figures
Reference graph
Works this paper leans on
-
[5]
Counterspeeches up my sleeve! intent dis- tribution learning and persistent fusion for intent- conditioned counterspeech generation. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5792–5809, Toronto, Canada. Association for Computational Linguistics. Sadaf Md. Halim, Saquib Irtiz...
work page 2023
-
[6]
Reinforcement learning-based counter- misinformation response generation: a case study of COVID-19 vaccine misinformation. In Proceedings of the ACM Web Conference 2023, pages 2698–2709. Amey Hengle, Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. CSEval: Towards auto- mated, multi-dimensional, and reference-free coun- terspeech evaluation us...
work page Pith review arXiv 2023
-
[7]
KOLD: Korean offensive language dataset. In Proceedings of the 2022 Conference on Empiri- cal Methods in Natural Language Processing, pages 10818–10833, Abu Dhabi, United Arab Emirates. As- sociation for Computational Linguistics. Aiqi Jiang, Xiaohan Yang, Yang Liu, and Arkaitz Zu- biaga. 2022. SWSR: A chinese dataset and lexicon for online sexism detecti...
work page 2022
-
[8]
Korean online hate speech dataset for mul- tilabel classification: How can social science im- prove dataset on hate speech? arXiv preprint arXiv:2204.03262. Hannah Kirk, Wenjie Yin, Bertie Vidgen, and Paul Röttger. 2023. SemEval-2023 task 10: Explainable detection of online sexism. In Proceedings of the 17th International Workshop on Semantic Evaluation (...
work page Pith review arXiv 2023
-
[9]
Trauma, Violence, & Abuse, 24(3):1727– 1742
A systematic review exploring variables re- lated to bystander intervention in sexual violence contexts. Trauma, Violence, & Abuse, 24(3):1727– 1742. Binny Mathew, Navish Kumar, Pawan Goyal, and Ani- mesh Mukherjee. 2020. Interaction dynamics be- tween hate and counter users on Twitter.Proceedings of the 7th ACM IKDD CoDS and 25th COMAD. Binny Mathew, Nav...
arXiv 2020
-
[10]
In NeurIPS 2023 Computational Sustainability: Promises and Pitfalls from Theory to Deployment
AI for whom? Shedding critical light on AI for social good. In NeurIPS 2023 Computational Sustainability: Promises and Pitfalls from Theory to Deployment. David L. Morgan. 1996. Focus groups. Annual Review of Sociology, 22(V olume 22, 1996):129–152. Michael J Muller and Sarah Kuhn. 1993. Participatory design. Communications of the ACM, 36(6):24–28. Jimin ...
arXiv 2023
-
[13]
Racism is a virus: anti-asian hate and counter- speech in social media during the COVID-19 crisis. Proceedings of the 2021 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining. I. Zubiaga, A. Soroa, and R. Agerri. 2024a. A LLM- based ranking method for the evaluation of automatic counter-narrative generation. pages 9572–958...
work page 2021
-
[2017]
Marcus Tomalin and Stefanie Ullmann, editors
#DistractinglySexy: How Social Media was used as a Counter Narrative on Gender in STEM. Marcus Tomalin and Stefanie Ullmann, editors. 2023. Counterspeech. Multidisciplinary Perspectives on Countering Dangerous Speech. Taylor & Francis. Vittoria Tonini, Simona Frenda, M. Stranisci, and Vi- viana Patti. 2024. How do we counter hate speech in Italy? María Es...
work page 2023
Show all 13 references
-
[2019]
In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA
SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 54–63, Min- neapolis, Minnesota, USA. Association for Compu- tational Linguistics. Susan Benesch,...
2019 arXiv
-
[2020]
In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 102–112, Online
Countering hate on social media: Large scale classification of hate and counter speech. In Pro- ceedings of the Fourth Workshop on Online Abuse and Harms, pages 102–112, Online. Association for Computational Linguistics. Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Dur- mu...
2024
-
[2022]
arXiv preprint arXiv:2203.03584
Counter hate speech in social media: A survey. arXiv preprint arXiv:2203.03584. Ghadi Alyahya and Abeer Aldayel. 2024. Hatred stems from ignorance! Distillation of the persuasion modes in countering conversational hate speech. ArXiv preprint, abs/2403.15449. I. Arpinar, Ugur K...
2024 arXiv
-
[2023]
EPJ Data Science, 12:1
Correction: Impact and dynamics of hate and counter speech online. EPJ Data Science, 12:1. Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent Hébert-Dufresne, and Mirta Galesic
-
[2024]
ArXiv preprint, abs/2412.15453
Northeastern Uni at multilingual counter- speech generation: Enhancing counter speech gener- ation with LLM alignment through direct preference optimization. ArXiv preprint, abs/2412.15453. Haiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao, Minghao Hu, and Bin Zhou. 2024a. F2RL...
1994 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.