Pith. sign in

REVIEW 3 major objections 5 minor 73 references

Echoes of Discord: Forecasting Hater Reactions to Counterspeech

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Predicting hater reactions in one step beats a two-stage pipeline

desk verdict The ReEco dataset is a real contribution, but the headline 3-way-vs-two-stage comparison rests on a baseline that is internally inconsistent with the paper's own component tables. read the letter →

arxiv 2501.16235 v2 pith:RB6S2EWI submitted 2025-01-27 cs.CL

classification cs.CL
keywords hatespeechcounterspeechhaterreentrypredictionconversationoutcomeforecastingRedditdatasetmulti-tasklearninglinguisticanalysisthree-wayclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the immediate reaction of a hate-speech author to a counterspeech can be forecast from the words of the hate comment and the counterspeech, and that forecasting all three possible reactions at once is more accurate than first predicting whether the hater comes back and then whether that reentry is hateful. To test this, the authors built ReEco, a dataset of 5,723 Reddit conversations with hate speech, a counterspeech reply, and the hater's labeled follow-up reaction (no reentry, hateful reentry, or non-hateful reentry). They also ran linguistic analyses that link particular counterspeech styles, such as aggressive or negative wording, to hateful reentry, and respectful or forgiving wording to non-hateful reentry. If the claims hold, platforms and counterspeech writers can screen response drafts for likely hater reactions, and model builders get a concrete comparison showing that a single three-way classifier beats the decomposed two-stage pipeline on this task.

What carries the argument

The central object is the ReEco corpus: 5,723 triple-turn Reddit threads in which a hate-speech post, the counterspeech reply to it, and the hater's follow-up comment are preserved, with the follow-up labeled as no reentry, hateful reentry, or non-hateful reentry. The load-bearing mechanism is the three-way classifier, a BERT-based multi-task model that takes the hate speech and counterspeech as a single paired input and predicts the three-way outcome directly. The comparison that carries the argument is the two-stage reaction predictor, which chains a reentry yes/no classifier and a reentry-type classifier; the paper argues the three-way design avoids compounding errors between these stages. The dataset's labels come from three-way-agreement classifiers for hate speech and counterspeech, so the automatic labeling pipeline is part of the machinery the conclusions rest on.

What would settle it

Re-annotate a large random sample of ReEco with human judgments for hate speech, counterspeech, and reentry type, retrain both the three-way classifier and the two-stage predictor on those human labels, and compare; if the two-stage predictor then matches or beats the three-way classifier, the claimed advantage is an artifact of automatic-label noise or error propagation rather than a genuine property of the forecasting task.

Watch

Extended reading notes

Core claim

The paper's central finding is that a BERT-based multi-task model that sees the hate speech and the counterspeech concatenated as a pair and directly outputs one of three labels—no reentry, hateful reentry, non-hateful reentry—predicts haters' reactions with a weighted F1 of 0.77, while the two-stage predictor that first decides whether the hater reenters and then decides whether the reentry is hateful achieves only 0.53. The same pair input also gives the best results for both subtasks when they are trained separately, and models that see only the hate speech or only the counterspeech are consistently weaker. The paper further reports that fine-tuned Llama 3 models, and zero-shot LLM prompting in particular, underperform BERT-based models on these predictions, and that linguistic markers such as aggression, exclamation, and negative emotion in counterspeech are associated with hateful reentry, while respect, power, worship, and forgiveness words are associated with non-hateful reentry.

Load-bearing premise

The load-bearing premise is that the automatic classifiers used to identify hate speech and counterspeech, and to label reentry comments as hateful, are accurate enough that the dataset's labels—validated on only 200 examples—reflect what the models learn and what the linguistic comparisons show.

Editorial extensions

If this is right

  • A single three-way classifier should be preferred over a two-stage pipeline for forecasting hater reactions to counterspeech on this type of data, because it reaches weighted F1 0.77 versus 0.53.
  • Including both the hate speech and the counterspeech as input, rather than either text alone, improves prediction across almost every model tested.
  • Counterspeech drafts that contain aggression or exclamation wording are associated with hateful reentry, while wording that signals respect, power, worship, or forgiveness is associated with non-hateful reentry; these signals can be used as linguistic guidelines for writing counterspeech.
  • Large language models, even fine-tuned, do not outperform smaller BERT-based models at predicting hater reactions in this setup, so model choice matters for this forecasting task.
  • The ReEco dataset, with real user-generated counterspeech and labeled hater outcomes, provides a benchmark for training or evaluating counterspeech generation systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the advantage of joint three-way prediction over a chained pipeline is a design lesson likely to transfer to other conversation-outcome forecasting tasks, such as predicting derailment or thread-ending posts, where intermediate decisions are currently chained.
  • Editorial extension: the linguistic markers could be turned into a lightweight scoring rule for counterspeech drafts, and a direct test would be to rewrite counterspeech to increase respect and forgiveness markers and reduce aggression markers while measuring hater reentry rates on a held-out Reddit sample.
  • Editorial extension: because only 200 samples were human-validated, the dataset's automatic labels may contain systematic noise; a fully human-annotated version of ReEco would show whether the reported linguistic differences and model comparisons survive cleaner labels.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces ReEco, a new dataset of Reddit triple-turn conversations in which an initial hate speech (HS) post receives counterspeech and the hater's reaction is categorized as no reentry, hateful reentry, or non-hateful reentry. The authors report linguistic analyses of counterspeech associated with each outcome and compare several models—BERT, BERT-MTL, and Llama 3 (zero-shot and fine-tuned)—under two prediction strategies: a two-stage pipeline (first predict reentry, then predict reentry type) and a direct 3-way classifier. The central claim is that the 3-way BERT-MTL model using HS/counterspeech pairs (weighted F1=0.77) outperforms the two-stage predictor (weighted F1=0.53). The paper also includes an error analysis of the best model.

Significance. The task itself is under-explored: modeling the hater's immediate reaction to counterspeech in real conversations is a valuable complement to prior work that focuses on bystanders or on synthetic counterspeech. The ReEco dataset is a concrete contribution and is released publicly. The linguistic findings (e.g., that respectful or forgiving counterspeech correlates with non-hateful reentry) are plausible and potentially actionable. However, the empirical support for the headline claim is currently weakened by internal inconsistencies in the reported evaluation tables and by the fact that the reaction-type labels are produced entirely by automatic classifiers without a dedicated human validation of those labels. If these issues are corrected and the main comparison is re-established, the paper would be a useful addition to the counterspeech and conversation-forecasting literature.

major comments (3)
  1. [Tables 5, 6, and 7] The per-class precision, recall, and weighted-average values are internally inconsistent with the test-set class priors. In Table 5, the majority baseline (reentry P=0.69, R=1.00) implies a test set with roughly 69% reentry. However, the BERT-MTL Pair row (reentry P=0.86, R=0.79; no-reentry P=0.81, R=0.87) is compatible only with a reentry prior near 50%; using the confusion-matrix identities FP = TP*(1/P - 1) and FN = (1-R)*N_actual yields p≈0.50. Table 6 shows the same pattern: the baseline implies 70% non-hateful among reentries, while the BERT-MTL Pair row implies balanced classes. In Table 7, the BERT-MTL Pair row's per-class P/R values imply class priors of roughly 6% hateful, 13% non-hateful, and 81% no-reentry, which contradicts the dataset distribution (20/47/33) and the baseline row's implied 48% non-hateful. The reported weighted F1 values thus cannot all be computed from the same test set. The authors need to recompute all metrics on a single fixed test split and report the actual label distribution alongside the results.
  2. [Table 7, Two-stage row] The 'Two-stage (BERT-MTL) Pair' row is not credible as a cascade of the best component models. The text states that the two-stage predictor combines the best models, i.e., BERT-MTL Pair for reentry (Table 5) and BERT-MTL Pair for reentry type (Table 6). In such a strict cascade, every true no-reentry instance that the first stage classifies as no reentry remains no reentry; the first stage's no-reentry recall is 0.87, so the cascade's no-reentry recall cannot be lower than 0.87. The reported value is 0.26. This suggests either a different first-stage model was used, a threshold was applied, or there is an error in the reported numbers. The authors must describe the exact two-stage implementation and re-derive its predictions from the component models.
  3. [Section 3 and Table 11] The hater reaction labels (hateful vs non-hateful reentry) are generated entirely by applying the authors' HS classifiers to the reentry comments, with no human validation of this specific label assignment. The 200-sample validation described in Section 3 (and Table 11) covers HS and counterspeech identification, not the reaction labels. Since both the model comparisons and the linguistic analyses (Tables 3, 4, and 7) depend on these labels, the authors should either validate the reaction-type labels on a human-annotated sample or provide a sensitivity analysis showing that the conclusions are robust to realistic levels of label noise in the reaction-type annotations.
minor comments (5)
  1. [Figure 1] The caption contains a typo: 'countersppech' should be 'counterspeech'.
  2. [Section 6] The phrase 'which complex the model understanding' should be 'which complicates the model's understanding'.
  3. [Appendix E, Table 13] The subreddit r/PurplePillDebate is listed twice under the Discussion category; please remove the duplicate.
  4. [Section 5.2, Reentry Type Prediction] The text refers to 'BERT-MLT models'; this should be 'BERT-MTL models' to match the rest of the paper.
  5. [Section 3, Data Collection] The definition of 'no reentry' should explicitly state that it means the hater does not appear in any subsequent reply within the collected thread; the current wording 'shows up in the follow-up conversation' is ambiguous regarding how far the follow-up window extends.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central 3-way versus two-stage comparison is a supervised model comparison on the authors' disclosed auto-labeled corpus; the same-author citations are not load-bearing.

full rationale

The paper's core claim is empirical: a 3-way classifier trained directly on three outcome labels outperforms a two-stage pipeline that first predicts reentry and then reentry type. This is a standard supervised model comparison, not a derivation that reduces to its inputs. The target labels are machine-generated (e.g., 'Based on the prediction of HS classifiers, the reentry comments are further labeled as hateful or non-hateful'), but this is a disclosed annotation pipeline with a 200-sample human validation subset; it creates a validity caveat about label quality, not a circular step. The same-author citations (Yu et al., 2022, used as one counterspeech training source and for the HS definition; Yu et al., 2023, used for annotator examples in error analysis) are supporting references to prior published work and do not by themselves force the experimental outcome. The comparison between the BERT-MTL 3-way model (weighted F1=0.77) and the two-stage predictor (weighted F1=0.53) is not an identity or a fitted-parameter rename; it is an evaluated performance difference on held-out data. Even if the two-stage baseline is implausibly weak or the tables contain inconsistencies, that is a correctness concern, not circularity. No step in the paper's derivation chain is equivalent by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is an empirical supervised-learning study; no hand-fitted constants or new theoretical entities are introduced. The main assumptions are about the validity of automatically generated labels and the representativeness of the filtered Reddit conversations.

assumptions (4)
  • domain assumption The three hate speech datasets (Vidgen et al., 2019; Qian et al., 2019; Davidson et al., 2017) provide valid training signal for hate speech detection.
    The paper fine-tunes RoBERTa on these datasets and relies on the resulting classifiers to label hate speech and hateful reentry in ReEco (Section 3).
  • domain assumption The SEANCE tool and spaCy accurately measure the linguistic features used in the corpus analysis.
    Section 4 uses SEANCE for sentiment and social cognition factors and spaCy for entity recognition; the conclusions about counterspeech language rely on these tools.
  • standard math The statistical tests used (Wilcoxon rank-sum, McNemar, Bonferroni correction) are appropriate for the data and comparisons.
    Sections 4 and 5 apply these tests to support linguistic differences and model comparisons.
  • domain assumption The exclusion of conversation pairs without follow-up replies does not systematically bias the no-reentry class.
    In Section 3 the authors exclude pairs at the end of the dialogue tree to avoid topic exhaustion and external interruptions, but this changes the definition of no reentry.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Echoes of Discord: Forecasting Hater Reactions to Counterspeech." pith.science (2026). https://pith.science/paper/RB6S2EWI

@misc{pith2026250116235,
  author       = {Pith},
  title        = {Pith review of: Echoes of Discord: Forecasting Hater Reactions to Counterspeech},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RB6S2EWI}},
  note         = {Machine review of arXiv:2501.16235}
}
read the original abstract

Hate speech (HS) erodes the inclusiveness of online users and propagates negativity and division. Counterspeech has been recognized as a way to mitigate the harmful consequences. While some research has investigated the impact of user-generated counterspeech on social media platforms, few have examined and modeled haters' reactions toward counterspeech, despite the immediate alteration of haters' attitudes being an important aspect of counterspeech. This study fills the gap by analyzing the impact of counterspeech from the hater's perspective, focusing on whether the counterspeech leads the hater to reenter the conversation and if the reentry is hateful. We compile the Reddit Echoes of Hate dataset (ReEco), which consists of triple-turn conversations featuring haters' reactions, to assess the impact of counterspeech. To predict haters' behaviors, we employ two strategies: a two-stage reaction predictor and a three-way classifier. The linguistic analysis sheds insights on the language of counterspeech to hate eliciting different haters' reactions. Experimental results demonstrate that the 3-way classification model outperforms the two-stage reaction predictor, which first predicts reentry and then determines the reentry type. We conclude the study with an assessment showing the most common errors identified by the best-performing model.

Figures

Figures reproduced from arXiv: 2501.16235 by the authors.

Figure 1
Figure 1. Hater’s non-hateful reentry as a conversation [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

73 extracted references · 57 canonical work pages

  1. [1]

    Dhafar Hamed Abd, Ayad R Abbas, and Ahmed T Sadiq. 2021. Analyzing sentiment system to specify polarity by lexicon-based. Bulletin of Electrical Engineering and Informatics, 10(1):283--289

  2. [2]

    Sindhu Abro, Sarang Shaikh, Zahid Hussain Khand, Ali Zafar, Sajid Khan, and Ghulam Mujtaba. 2020. Automatic hate speech detection using machine learning: A comparative study. International Journal of Advanced Computer Science and Applications, 11(8)

  3. [3]

    Doris E Acheme, Chris Anderson, and Claude Miller. 2024. The effects of language features and accents on the arousal of psychological reactance and communication outcomes. Communication Research, page 00936502241229883

  4. [4]

    Abdullah Albanyan, Ahmed Hassan, and Eduardo Blanco. 2023. Finding authentic counterhate arguments: A case study with public figures. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13862--13876

  5. [5]

    Alain Auger and Jean Roy. 2008. Expression of uncertainty in linguistic data. In 2008 11th International Conference on Information Fusion, pages 1--8. IEEE

  6. [6]

    Lars Backstrom, Jon Kleinberg, Lillian Lee, and Cristian Danescu-Niculescu-Mfizil. 2013. Characterizing and curating conversation threads: expansion, focus, volume, re-entry. In Proceedings of the sixth ACM international conference on Web search and data mining, pages 13--22

  7. [7]

    Fabienne Baider. 2023. Accountability issues, online covert hate speech, and the efficacy of counter-speech. Politics and Governance, 11(2):249--260

  8. [8]

    Jiajun Bao, Junjie Wu, Yiming Zhang, Eshwar Chandrasekharan, and David Jurgens. 2021. Conversations gone alright: Quantifying and predicting prosocial outcomes in online conversations. In Proceedings of the Web Conference 2021, pages 1134--1145

Show all 73 references
  1. [9]

    Helena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, and Marco Guerini. 2022. Human-machine collaboration approaches to build a dialogue dataset for hate speech countering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8031--8049

  2. [10]

    Anna M Borghi et al. 2019. Linguistic relativity and abstract words. Paradigmi, 37(3):429--448

  3. [11]

    Bianca Cepollaro, Maxime Lepoutre, and Robert Mark Simpson. 2023. Counterspeech. Philosophy Compass, 18(1):e12890

  4. [12]

    Jonathan P Chang and Cristian Danescu-Niculescu-Mizil. 2019. Trouble on the horizon: Forecasting the derailment of online conversations as they develop. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Co...

  5. [13]

    Naganna Chetty and Sreejith Alathur. 2018. Hate speech review in the context of online social networks. Aggression and violent behavior, 40:108--118

  6. [14]

    enthusiasm

    Jihyang Choi and Jiyoung Lee. 2021. “enthusiasm” toward the other side matters: Emotion and willingness to express disagreement in social media political conversation. The Social Science Journal, pages 1--17

  7. [15]

    Yi-Ling Chung, Gavin Abercrombie, Florence Enock, Jonathan Bright, and Verena Rieser. 2023. Understanding counterspeech for online harm mitigation. arXiv preprint arXiv:2307.04761

  8. [16]

    Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiro g lu, and Marco Guerini. 2019. Conan-counter narratives through nichesourcing: a multilingual dataset of responses to fight online hate speech. In Proceedings of the 57th Annual Meeting of the Association for Computational ...

  9. [17]

    Yi-Ling Chung, Serra Sinem Tekiroglu, and Marco Guerini. 2020. Italian counter narrative generation to fight online hate speech. In Proceedings of the Seventh Italian Conference on Computational Linguistics (CLIC-it 2020), volume 2769

  10. [18]

    Yi-Ling Chung, Serra Sinem Tekiro g lu, and Marco Guerini. 2021. Towards knowledge-grounded counter narrative generation for hate speech. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 899--914

  11. [19]

    Leigh Clark, Nadia Pantidi, Orla Cooney, Philip Doyle, Diego Garaialde, Justin Edwards, Brendan Spillane, Emer Gilmartin, Christine Murad, Cosmin Munteanu, et al. 2019. What makes a good conversation? challenges in designing truly conversational agents. In Proceedings of the 2...

  12. [20]

    Scott A Crossley, Kristopher Kyle, and Danielle S McNamara. 2017. Sentiment analysis and social cognition engine (seance): An automatic tool for sentiment, social cognition, and social-order analysis. Behavior research methods, 49:803--821

  13. [21]

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the international AAAI conference on web and social media, volume 11, pages 512--515

  14. [22]

    Jim Dillard and Cindy Harmon-Jones. 2002. A cognitive dissonance theory perspective on persuasion. The persuasion handbook: Developments in theory and practice, page 99

  15. [23]

    Rehab Duwairi, Amena Hayajneh, and Muhannad Quwaider. 2021. A deep learning framework for automatic detection of hate speech embedded in arabic tweets. Arabian Journal for Science and Engineering, 46:4001--4014

  16. [24]

    Simona Frenda, Alessandra Teresa Cignarella, Valerio Basile, Cristina Bosco, Viviana Patti, and Paolo Rosso. 2022. The unbearable hurtfulness of sarcasm. Expert Systems with Applications, 193:116398

  17. [25]

    Asma Ghandeharioun, Daniel McDuff, Mary Czerwinski, and Kael Rowan. 2019. Towards understanding emotional intelligence for behavior change chatbots. In 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), pages 8--14. IEEE

  18. [26]

    Seth Green, Megan Stiles, Katherine Harton, Samantha Garofalo, and Donald E Brown. 2017. Computational analysis of religious and ideological linguistic behavior. In 2017 Systems and Information Engineering Design Symposium (SIEDS), pages 359--364. IEEE

  19. [27]

    Edita Grolman, Hodaya Binyamini, Asaf Shabtai, Yuval Elovici, Ikuya Morikawa, and Toshiya Shimizu. 2022. Hateversarial: Adversarial attack against hate speech detection algorithms on twitter. In Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personaliz...

  20. [28]

    Ella Guest, Bertie Vidgen, Alexandros Mittos, Nishanth Sastry, Gareth Tyson, and Helen Margetts. 2021. An expert annotated dataset for the detection of online misogyny. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguisti...

  21. [29]

    Jeffrey T Hancock, Kailyn Gee, Kevin Ciaccio, and Jennifer Mae-Hwah Lin. 2008. I'm sad you're sad: emotional contagion in cmc. In Proceedings of the 2008 ACM conference on Computer supported cooperative work, pages 295--298

  22. [30]

    Dominik Hangartner, Gloria Gennaro, Sary Alasiri, Nicholas Bahrich, Alexandra Bornhoft, Joseph Boucher, Buket Buse Demirci, Laurenz Derksen, Aldo Hall, Matthias Jochum, et al. 2021. Empathy-based counterspeech can reduce racist hate speech in a social media field experiment. P...

  23. [31]

    Bing He, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, and Srijan Kumar. 2021. Racism is a virus: Anti-asian hate and counterspeech in social media during the covid-19 crisis. In Proceedings of the 2021 IEEE/ACM International Conference on Advances in Social Networ...

  24. [32]

    M Honnibal and I Montani. 2017. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. neural machine translation. In Proceedings of the Association for Computational Linguistics (ACL), pages 688--697

  25. [33]

    Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations

  26. [34]

    Yunhao Jiao, Cheng Li, Fei Wu, and Qiaozhu Mei. 2018. Find the conversation killers: A predictive study of thread-ending posts. In Proceedings of the 2018 World Wide Web Conference, pages 1145--1154

  27. [35]

    Nathan Lambert, Kristofer Pister, and Roberto Calandra. 2022. Investigating compounding prediction errors in learned dynamics models. arXiv preprint arXiv:2203.09637

  28. [36]

    Larissa Leonhard, Christina Rue , Magdalena Obermaier, and Carsten Reinemann. 2018. Perceiving threat and feeling responsible. how severity of hate speech, number of bystanders, and prior reactions of others affect bystanders’ intention to counterargue against hate speech on f...

  29. [37]

    Ping Liu, Joshua Guberman, Libby Hemphill, and Aron Culotta. 2018. Forecasting the presence and intensity of hostility on instagram using linguistic and social features. In Proceedings of the International AAAI Conference on Web and Social Media, volume 12

  30. [38]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. http://arxiv.org/abs/1907.11692 Roberta: A robustly optimized bert pretraining approach

  31. [39]

    Nurul Lubis, Sakriani Sakti, Koichiro Yoshino, and Satoshi Nakamura. 2019. Positive emotion elicitation in chat-based dialogue systems. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(4):866--877

  32. [40]

    George E Marcus, Michael MacKuen, and W Russell Neuman. 2011. Parsimony and complexity: Developing and testing theories of affective intelligence. Political Psychology, 32(2):323--336

  33. [41]

    Binny Mathew, Punyajoy Saha, Hardik Tharad, Subham Rajgaria, Prajwal Singhania, Suman Kalyan Maity, Pawan Goyal, and Animesh Mukherjee. 2019. Thou shalt not hate: Countering online hate speech. In Proceedings of the international AAAI conference on web and social media, volume...

  34. [42]

    Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153--157

  35. [43]

    Saif M Mohammad. 2021. Sentiment analysis: Automatically detecting valence, emotions, and other affectual states from text. In Emotion measurement, pages 323--379. Elsevier

  36. [44]

    Prerna Nadathur and Sven Lauer. 2020. Causal necessity, causal sufficiency, and the implications of causative verbs. Glossa: a journal of general linguistics, 5(1)

  37. [45]

    Rob Procter, Helena Webb, Marina Jirotka, Pete Burnap, William Housley, Adam Edwards, and Matt Williams. 2019. A study of cyber hate on twitter with implications for social media governance strategies. arXiv preprint arXiv:1908.11732

  38. [46]

    Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019. A benchmark dataset for learning to intervene in online hate speech. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Co...

  39. [47]

    Louis Reynolds and Henry Tuck. 2016. The counter-narrative monitoring & evaluation handbook. Institute for Strategic Dialogue

  40. [48]

    Julian Risch and Ralf Krestel. 2020. Top comment or flop comment? predicting and explaining user engagement in online news discussions. In Proceedings of the International AAAI Conference on Web and Social Media, volume 14, pages 579--589

  41. [49]

    Georgios Rizos, Symeon Papadopoulos, and Yiannis Kompatsiaris. 2016. Predicting news popularity by mining online discussions. In Proceedings of the 25th international conference companion on world wide web, pages 737--742

  42. [50]

    Paul R \"o ttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet B Pierrehumbert. 2020. Hatecheck: Functional tests for hate speech detection models. arXiv preprint arXiv:2012.15606

  43. [51]

    Punyajoy Saha, Kanishk Singh, Adarsh Kumar, Binny Mathew, and Animesh Mukherjee. 2022. Countergedi: A controllable approach to generate polite, detoxified and emotional counterspeech. arXiv preprint arXiv:2205.04304

  44. [52]

    Carla Schieb and Mike Preuss. 2016. Governing hate speech by means of counterspeech on facebook. In 66th ica annual conference, at fukuoka, japan, pages 1--23

  45. [53]

    Anna Schmidt and Michael Wiegand. 2017. A survey on hate speech detection using natural language processing. In Proceedings of the fifth international workshop on natural language processing for social media, pages 1--10

  46. [54]

    Sarah Shugars and Nicholas Beauchamp. 2019. Why keep arguing? predicting engagement in political conversations online. Sage Open, 9(1):2158244019828850

  47. [55]

    Wolfgang Stroebe. 2008. Strategies of attitude and behaviour change. Introduction to social psychology. Oxford, UK: Blackwell

  48. [56]

    Rinji Suzuki and Akiyo Nadamoto. 2020. Extracting rhetorical question from twitter. In Proceedings of the 22nd International Conference on Information Integration and Web-based Applications & Services, pages 290--299

  49. [57]

    Serra Sinem Tekiroglu, Helena Bonaldi, Margherita Fanton, and Marco Guerini. 2022. Using pre-trained language models for producing counter narratives against hate speech: a comparative study. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3099--3114

  50. [58]

    Luk Van Mensel, Mieke Vandenbroucke, Robert Blackwood, Ofelia Garc \' a, Nelson Flores, and Massimiliano Spotti. 2016. The oxford handbook of language and society

  51. [59]

    Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019. Challenges and frontiers in abusive content detection. In Proceedings of the third workshop on abusive language online. Association for Computational Linguistics

  52. [60]

    Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, and Rebekah Tromble. 2021. Introducing cad: the contextual abuse dataset. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolo...

  53. [61]

    Anthony J Viera, Joanne M Garrett, et al. 2005. Understanding interobserver agreement: the kappa statistic. Fam med, 37(5):360--363

  54. [62]

    Sebastian Wachs, Ludwig Bilz, Alexander Wettstein, Michelle F Wright, Norman Krause, Cindy Ballaschk, and Julia Kansok-Dusche. 2022. The online hate speech cycle of violence: Moderating effects of moral disengagement and empathy in the victim-to-perpetrator relationship. Cyber...

  55. [63]

    Lingzhi Wang, Xingshan Zeng, Huang Hu, Kam-Fai Wong, and Daxin Jiang. 2021. Re-entry prediction for online conversations via self-supervised learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2127--2137

  56. [64]

    Eric W Weisstein. 2004. Bonferroni correction. https://mathworld. wolfram. com/

  57. [65]

    Galen Weld, Amy X Zhang, and Tim Althoff. 2022. What makes online communities ‘better’? measuring values, consensus, and conflict across thousands of subreddits. In Proceedings of the International AAAI Conference on Web and Social Media, volume 16, pages 1121--1132

  58. [66]

    Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2022. Hate speech and counter speech detection: Conversational context does matter. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p...

  59. [67]

    Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2023. Hate cannot drive out hate: Forecasting conversation incivility following replies to hate speech. arXiv preprint arXiv:2312.04804

  60. [68]

    Xingshan Zeng, Jing Li, Lu Wang, and Kam-Fai Wong. 2019. Joint effects of context and user history for predicting online conversation re-entries. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2809--2818

  61. [69]

    Justine Zhang, Jonathan P Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Nithum Thain, and Dario Taraborelli. 2018. Conversations gone awry: Detecting early signs of conversational failure. arXiv preprint arXiv:1805.05345

  62. [70]

    Yu Zhang and Qiang Yang. 2018. An overview of multi-task learning. National Science Review, 5(1):30--43

  63. [71]

    Wanzheng Zhu and Suma Bhat. 2021. Generate, prune, select: A pipeline for counterspeech generation against online hate speech. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 134--149

  64. [72]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  65. [73]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.