REVIEW 3 major objections 4 minor 82 references
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Add the user's history, not fine-tuning, to win with AI counterspeech: a base LLM prompted with conversation context and the toxic user's past comments outperforms generic replies on perceived adequacy and persuasiveness.
desk verdict Systematic, pre-registered comparison of 36 counterspeech configurations; the headline result about lightweight context is plausible but rests on near-zero inter-rater reliability, so the claim needs a mixed-effects reanalysis before it is solid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is 'lightweight contextual prompting' of the base model: the input prompt is augmented with (i) the two messages preceding the toxic comment in the thread and (ii) ten comments from the toxic user's history, while the model weights remain untouched. This is contrasted with fine-tuning on counterspeech datasets (MultiCONAN, RHSI) or community comments (political subreddits). The argument is carried by comparing 36 configurations across algorithmic indicators and a pre-registered crowdsourced human evaluation, with the [Ba Pr Hi] versus [Ba] contrast as the pivotal result.
What would settle it
A field experiment on a real platform: deploy [Ba Pr Hi] and [Ba] counterspeech to users who posted toxic comments (matched design), and measure subsequent comment toxicity, deletion, or engagement. If [Ba Pr Hi] does not outperform [Ba] on actual behavioral outcomes, the central claim—that contextualized counterspeech is more persuasive—would be falsified for real-world persuasion. Alternatively, a re-analysis of the crowdsourced data using a model with rater-level random effects could show whether the configuration effect survives when accounting for rater idiosyncrasies.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a base instruction-tuned LLaMA2-13B model prompted with up to two preceding conversation messages and ten of the toxic user's previous Reddit comments (configuration [Ba Pr Hi]) significantly outperforms the generic baseline [Ba] on perceived adequacy and perceived persuasiveness toward the toxic author, with consistent gains on relevance, truthfulness, and bystander persuasion. The paper further shows that this improvement is specific to supplying context at inference time: configurations that add the same information via fine-tuning, e.g., [Mu Re Pr Hi], fail dramatically, with 90.6% of outputs lacking any moderation framing and 8.6%
Load-bearing premise
The study assumes that average Likert ratings from crowdworkers, despite near-zero inter-rater agreement (Krippendorff's alpha around 0.002–0.008), capture a meaningful shared signal of counterspeech adequacy and persuasiveness; if those ratings are mostly idiosyncratic noise, the configuration-level comparisons lose their footing.
Editorial extensions
If this is right
- If true, moderation systems can improve counterspeech simply by feeding conversation context and user history into a generic instruction-tuned model, avoiding expensive and risky fine-tuning.
- Fine-tuning on counterspeech or community data should be treated as a potential failure source; the choice of corpus determines which aspect of the moderation function is lost.
- The negative correlation between automatic metrics and human judgments means algorithmic screening alone can mislead: a configuration can rank among the best under metrics yet among the worst for humans.
- The register of a moderation act (imperative mood, normative appeals) strongly predicts human adequacy ratings, suggesting future systems should explicitly enforce this register.
- Effective counterspeech may need to vary by toxicity subtype and emotional tone: anger hurts persuasiveness, while trust and anticipation help.
Reading between the lines
- The very low inter-rater reliability (Krippendorff's alpha around 0.004) raises the possibility that the reported configuration-level differences reflect shared rating patterns rather than a stable property of the messages; a re-analysis with alternative aggregation or rater-level modeling could either reinforce or dissolve the headline effect.
- The paper measures perceived persuasiveness from third-party raters, not actual behavior change. A direct field experiment sending real counterspeech to toxic users and measuring subsequent toxicity or re-engagement would be the natural test of whether perceived persuasiveness translates to real-world persuasion.
- The finding that older raters systematically give lower scores suggests that the demographic composition of crowdsourcing samples could tilt conclusions; future studies should preregister demographic weighting or stratified sampling.
- The success of [Ba Pr Hi] may not generalize to other base models or domains; the appendix's robustness check with another LLM shows model-dependent differences, so the mechanism may be specific to instruction-tuned models rather than universal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes and evaluates configurations for contextualized/personalized counterspeech generation using LLaMA2-13B, varying fine-tuning corpora (MultiCONAN, RHSI, Reddit), community adaptation, conversation prefix, user-history, and user-summary factors (36 factorial but not fully crossed). Algorithmic indicators and a pre-registered crowdsourcing study (N=2,444 non-contextual, N=2,353 contextual) compare seven selected configurations against the generic baseline [Ba]. The headline finding is that [Ba Pr Hi] (base model with conversation context and user's comment history) is rated significantly better than [Ba] in adequacy and perceived persuasiveness toward the toxic author, while many fine-tuned and/or more heavily contextualized configurations degrade perceived quality. The paper also reports that automatic metrics (ROUGE/BLEU/BERTScore) rank configurations essentially opposite to human ratings, and presents a failure-mode analysis showing fine-tuning produces unframed/toxic/register-shifted outputs.
Significance. The study is valuable for its systematic factorial design, pre-registered protocol, large participant pool, open models, and the split between algorithmic and human evaluation. If the human-rating result survives the reliability concerns, the finding that lightweight inference-time contextualization (without fine-tuning) improves perceived counterspeech quality—and that fine-tuning on counter-narrative corpora can destroy the moderation register—is an actionable, non-obvious result for the moderation community. The negative correlation between automatic indicators and human judgments is a useful cautionary result, and the failure-mode analysis (unframed, toxic, degenerate) is a concrete step toward functional evaluation.
major comments (3)
- [§6.2, Appendix Table 5, §4.2.4] The central claim rests entirely on crowdworkers' Likert ratings. The manuscript's own reliability statistics show near-zero inter-rater reliability (Krippendorff's α=0.002–0.008; mean pairwise Spearman=0.002–0.005; pairwise agreement≈0.29). Although low α can coexist with stable aggregate means, it places the burden on the analysis to show that configuration-level differences reflect a shared perception rather than rater-specific scale use or noise. The Friedman and Wilcoxon tests in §4.2.4 operate on raw scores and do not include random effects for participants or items; with N≈2,400 and 7 configurations even small idiosyncratic tendencies can reach significance. I request a re-analysis with cumulative-link mixed models containing random intercepts for participants and for toxic-message instances (or threads), reporting the [Ba Pr Hi] vs [Ba] contrasts for adequacy and the two persuasi
- [Fig. 3, §6.2.1] The significance legend uses '*: p<0.1' as a significance level, while the text describes [Ba Pr Hi] as 'significantly better' on adequacy and persuasiveness-to-author. The paper states that a Bonferroni correction was applied but does not report the number of comparisons or the adjusted threshold, nor exact p-values. If these findings are significant only at the uncorrected or p<0.1 level, the abstract's claim is overstated. Please provide a full contrast table with test statistic, raw p, adjusted p, and effect size with confidence interval for each dimension and each configuration, and use a uniform α=0.05 threshold or clearly label p<0.1 results as marginal.
- [§5, §6.2.4, Table 3, title/abstract] The evidence base is narrow: 128 toxic comments, 85 from r/politics, predominantly obscene/insulting subtypes; the outcome is third-party perceived persuasiveness, not actual behavior change. The authors acknowledge these limitations in the Discussion and Limitations section, but the title and abstract state 'contextualized counterspeech can be more persuasive than generic counterspeech' without those qualifiers. Please temper the abstract and title to 'perceived persuasiveness in U.S. political Reddit threads' or otherwise bound the claim, since the current phrasing over-generalizes beyond the demonstrated scope.
minor comments (4)
- [§6.3] The text says 'computed over all 6,912 generations,' but §5 reports 36 configurations × 128 toxic messages = 4,608 generated responses. Please reconcile or clarify where 6,912 comes from (e.g., additional LLM robustness runs).
- [Appendix Table 5] The caption states that statistics are reported separately for the non-contextual condition, the contextual condition, and both conditions pooled together, but the table as shown has only one row per question with no condition split. Please include the condition breakdown or correct the caption.
- [Fig. 3 and Fig. 4] Using '*' for p<0.1 alongside '**' for p<0.05 is unusual and easily misread. If retained, explicitly state in the caption that '*' is nominal and whether any multiple-comparison correction is applied.
- [§8, §9] Typographical issues: 'pre-registreted' in §8, and 'the the PNRR' in the Acknowledgments. Also, the super-ranking optimization in §4.2.2 is described only verbally; a formal statement of the minimized objective would aid reproducibility.
Circularity Check
No significant circularity: the headline claim rests on pre-registered human judgments, and the paper explicitly shows algorithmic indicators rank configurations opposite to humans.
full rationale
The central claim—that [Ba Pr Hi] improves perceived adequacy and persuasiveness over [Ba]—is an empirical result from a pre-registered crowdsourcing experiment (Section 4.2.3, Section 6.2.1), not a quantity derived from the model or from fitted indicators. The paper's own comparison in Section 6.2.3 shows algorithmic rankings are negatively correlated with human rankings (Kendall τ from -0.05 to -0.71), which is the opposite of fitting the outcome into the metric. The configuration-selection pipeline uses algorithmic indicators only to choose representative configurations and, as the paper acknowledges in Section 7, this may exclude human-preferred outputs; it does not define the evaluated outcome. The low inter-rater reliability reported in Appendix Table 5 (Krippendorff's α ≈ 0.002–0.008) and the comment that non-parametric tests do not jointly model participant- and stimulus-level variability (Section 7) are measurement-validity limitations, not circularity: ratings are external to the generation process. Self-citations such as [17] introduce prior strategies and background, but they are not load-bearing for the persuasion claim, which is measured de novo. No equation or fitted parameter reduces to the target result by construction, so no circular step can be exhibited.
Assumptions & free parameters
free parameters (7)
- Toxicity threshold (Perspective API) =
0.5
- Minimum user activity for personalization =
20 comments
- Comment-history window (Hi) =
10 comments
- Summary window (Su) =
20 comments
- Centroid-selection sample size =
20 messages per configuration
- Reddit fine-tuning dataset size =
≈7,500 (§4.1.1) vs ≈5K (Appendix A)
- Power-analysis target =
≈2,500 participants per condition; achieved N ≈ 2,444 / 2,353
assumptions (8)
- domain assumption Crowdworkers' five-point Likert ratings operationalize relevance, adequacy, truthfulness, artificiality, and persuasiveness.
- domain assumption Perspective API scores with threshold ≥ 0.5 identify toxic comments and toxic outputs.
- domain assumption ROUGE/BLEU/BERTScore overlap operationalize relevance, diversity, adaptation, and lexical personalization.
- ad hoc to paper The twenty messages closest to each configuration's centroid in indicator space are representative of that configuration.
- standard math Nonparametric significance tests (Friedman, Wilcoxon, Mann-Whitney U) with Bonferroni correction are valid for ordinal Likert responses.
- domain assumption A single generation per toxic message is enough to characterize each configuration.
- domain assumption LLaMA2-13B instruction-tuned model follows the prompts and uses supplied context as intended.
- domain assumption User summaries generated by LLaMA2 do not infer sensitive or protected attributes.
Cite this review
Pith. "Pith review of Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech." pith.science (2026). https://pith.science/paper/TUNKF6MC
@misc{pith2026260726236,
author = {Pith},
title = {Pith review of: Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech},
year = {2026},
howpublished = {\url{https://pith.science/paper/TUNKF6MC}},
note = {Machine review of arXiv:2607.26236}
}
read the original abstract
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Ana Aleksandric, Sayak Saha Roy, Hanani Pankaj, Gabriela Mustata Wilson, and Shirin Nilizadeh. 2024. Users’ behavioral and emotional response to toxicity in Twitter conversations. InAAAI ICWSM
2024
-
[2]
Anirban Saha Anik, Xiaoying Song, Elliott Wang, Bryan Wang, Bengisu Yarimbas, and Lingzi Hong. 2025. Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation. InCOLM
2025
-
[3]
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift Reddit dataset. InAAAI ICWSM
2020
-
[4]
Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. InEACL
2024
-
[5]
Helena Bonaldi, Yi-Ling Chung, Gavin Abercrombie, and Marco Guerini. 2024. NLP for counterspeech against hate: A survey and how-to guide. InNAACL
2024
-
[6]
Angana Borah, Rada Mihalcea, and Verónica Pérez-Rosas. 2026. Persuasion at play: Understanding misinformation dynamics in demographic-aware human-LLM interactions. InEACL
2026
-
[7]
Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur, and Dilek Hakkani-Tur. 2026. Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. InACM CAIS
2026
-
[8]
Simon Martin Breum, Daniel Vædele Egdal, Victor Gram Mortensen, Anders Giovanni Møller, and Luca Maria Aiello
Show all 82 references
-
[9]
Dominique Brunato, Andrea Cimino, Felice Dell’Orletta, Giulia Venturi, and Simonetta Montemagni. 2020. Profiling-UD: A tool for linguistic profiling of texts. InLREC
2020
-
[10]
Dominik Bär, Abdurahman Maarouf, and Stefan Feuerriegel. 2024. Generative AI may backfire for counterspeech. arXiv:2411.14986(2024)
2024 arXiv
-
[11]
Aldo Cerulli, Lorenzo Cima, Benedetta Tessa, Serena Tardelli, and Stefano Cresci. 2026. The Big Ban Theory: A pre-and post-intervention dataset of online content moderation actions. InAAAI ICWSM
2026
-
[12]
Aldo Cerulli, Benedetta Tessa, Giuseppe La Selva, Oronzo Mazzeo, Lorenzo Cima, Lucia Monacis, and Stefano Cresci
-
[13]
Eshwar Chandrasekharan, Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2022. Quarantined! Examining the effects of a community-wide moderation intervention on Reddit.ACM TOCHI29, 4 (2022)
2022
-
[14]
Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, and Jonathan May. 2024. Can language model moderators improve the health of online discourse? NAACL(2024)
2024
-
[15]
Yi-Ling Chung, Gavin Abercrombie, Florence Enock, Jonathan Bright, and Verena Rieser. 2024. Understanding Counterspeech for Online Harm Mitigation.Northern European Journal of Language Technology10, 1 (2024)
2024
-
[16]
Yi-Ling Chung, Serra Sinem Tekiroğlu, and Marco Guerini. 2021. Towards knowledge-grounded counter narrative generation for hate speech. InACL-IJCNLP
2021
-
[17]
Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avvenuti, Felice Dell’Orletta, and Stefano Cresci. 2025. Contextualized counterspeech: Strategies for adaptation, personalization, and evaluation. InACM WWW
2025
-
[18]
Lorenzo Cima, Benedetta Tessa, Amaury Trujillo, Stefano Cresci, and Marco Avvenuti. 2025. Investigating the heterogeneous effects of a massive content moderation intervention via Difference-in-Differences.Online Social Networks and Media48 (2025), 100320
2025
-
[19]
Jordi Guillem Condom Tibau, Angelina Voggenreiter, Jürgen Pfeffer, et al. 2025. Prevalence, Substance and Responses to Hate Speech Against LGBTQ Communities on TikTok. InAAAI ICWSM
2025
-
[20]
Thomas H Costello, Gordon Pennycook, and David G Rand. 2024. Durably reducing conspiracy beliefs through dialogues with AI.Science385 (2024)
2024
-
[21]
Stefano Cresci, Roberto Di Pietro, Marinella Petrocchi, Angelo Spognardi, and Maurizio Tesconi. 2014. A criticism to society (as seen by Twitter analytics). InIEEE ICDCS Workshops
2014
-
[22]
Stefano Cresci, Amaury Trujillo, and Tiziano Fagni. 2022. Personalized interventions for online moderation. InACM Hypertext
2022
-
[23]
Francisco Cribari-Neto and Achim Zeileis. 2010. Beta regression in R.Journal of statistical software34 (2010), 1–24
2010
-
[24]
Mekselina Doğanç and Ilia Markov. 2023. From generic to personalized: Investigating strategies for generating targeted counter narratives against hate speech. InACL CS4OA
2023
-
[25]
Cynthia Dwork, Chris Hays, Jon Kleinberg, and Manish Raghavan. 2024. Content moderation and the formation of online communities: A theoretical framework. InACM WWW
2024
-
[26]
Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroğlu, and Marco Guerini. 2021. Human-in-the-Loop for data collection: A multi-target counter narrative dataset to fight online hate speech. InACL-IJCNLP
2021
-
[27]
Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra, Yudai Yamazaki, Yasutaka Nishimura, Sina J Semnani, Kazushi Ikeda, Weiyan Shi, and Monica S Lam. 2024. Zero-shot persuasive chatbots with LLM-generated strategies and information , Vol. 1, No. 1, Article . Publication date: Jul...
2024
-
[28]
John D Gallacher, Marc W Heerdink, and Miles Hewstone. 2021. Online engagement between opposing political protest groups via social media is linked to physical violence of offline encounters.Social Media + Society7, 1 (2021)
2021
-
[29]
Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent Hébert-Dufresne, and Mirta Galesic. 2022. Impact and dynamics of hate and counter speech online.EPJ Data Science11, 1 (2022)
2022
-
[30]
Gloria Gennaro, Laurenz Derksen, Aya Abdelrahman, Emma Broggini, Mariya Alexandra Green, Victoria Andrea Haerter, Elia Heer, Isabel Heidler, Fiona Kauer, Han-Nuri Kim, et al. 2025. Counterspeech encouraging users to adopt the perspective of minority groups reduces hate speech ...
2025
-
[31]
Tarleton Gillespie. 2020. Content moderation, AI, and the question of scale.Big Data & Society7, 2 (2020)
2020
-
[32]
Tommaso Giorgi, Lorenzo Cima, Tiziano Fagni, Marco Avvenuti, and Stefano Cresci. 2025. Human and LLM biases in hate speech annotations: A socio-demographic analysis of annotators and targets.AAAI ICWSM(2025)
2025
-
[33]
Natasha Goel, Thomas Bergeron, Blake Lee-Whiting, Thomas Galipeau, Danielle Bohonos, Sarah Lachance, Sonja Savolainen, Clareta Treger, and Eric Merkley. 2024. Artificial influence? Comparing AI and human persuasion in reducing belief certainty. (2024). https://doi.org/10.31219...
2024 doi
-
[34]
Pierpaolo Goffredo, Valerio Basile, Bianca Cepollaro, Viviana Patti, et al. 2022. Counter-TWIT: An Italian corpus for online counterspeech in ecological contexts. InACL WOAH
2022
-
[35]
Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2024. How persuasive is AI-generated propaganda?PNAS Nexus3, 2 (2024)
2024
-
[36]
Jarod Govers, Eduardo Velloso, Vassilis Kostakos, and Jorge Goncalves. 2024. AI-Driven Mediation Strategies for Audience Depolarisation in Online Debates. InACM CHI
2024
-
[37]
Kobi Hackenburg and Helen Margetts. 2024. Evaluating the persuasive influence of political microtargeting with large language models.PNAS121, 24 (2024)
2024
-
[38]
Kobi Hackenburg, Ben M Tappin, Paul Röttger, Scott A Hale, Jonathan Bright, and Helen Margetts. 2025. Scaling language model size yields diminishing returns for single-message political persuasion.PNAS122, 10 (2025)
2025
-
[39]
Sadaf MD Halim, Saquib Irtiza, Yibo Hu, Latifur Khan, and Bhavani Thuraisingham. 2023. WokeGPT: Improving counterspeech generation against online hate speech by intelligently augmenting datasets using a novel metric. In IEEE IJCNN
2023
-
[40]
Sabit Hassan and Malihe Alikhani. 2023. DisCGen: A framework for discourse-informed counterspeech generation. In IJCNLP-AACL
2023
-
[41]
Bing He, Mustaque Ahamad, and Srijan Kumar. 2023. Reinforcement learning-based counter-misinformation response generation: A case study of COVID-19 vaccine misinformation. InACM WWW
2023
-
[42]
Amey Hengle, Aswini Kumar Padhi, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs. InNAACL
2025
-
[43]
Daniel Hickey, Daniel MT Fessler, Matheus Schmitz, Paul Smaldino, Kristina Lerman, Goran Murić, and Keith Burghardt
-
[44]
Lingzi Hong, Pengcheng Luo, Eduardo Blanco, and Xiaoying Song. 2024. Outcome-constrained large language models for countering hate speech.EMNLP(2024)
2024
-
[45]
Manoel Horta Ribeiro, Shagun Jhaver, Savvas Zannettou, Jeremy Blackburn, Gianluca Stringhini, Emiliano De Cristofaro, and Robert West. 2021. Do platform migrations compromise content moderation? Evidence from r/The_Donald and r/Incels. InACM CSCW
2021
-
[46]
InAAAI ICWSM
Assessing How Hate, Counterspeech, and Toxicity Affect Hate Group Newcomers. InAAAI ICWSM
-
[47]
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. PersonaLLM: Investigating the ability of large language models to express personality traits.NAACL(2024)
2024
-
[48]
Shuyu Jiang, Wenyi Tang, Xingshu Chen, Rui Tang, Haizhou Wang, and Wenxian Wang. 2025. ReZG: Retrieval- augmented zero-shot counter narrative generation for hate speech.Neurocomputing620 (2025)
2025
-
[49]
Evey Jiaxin Huang, Abhraneel Sarma, Sohyeon Hwang, Eshwar Chandrasekharan, and Stevie Chancellor. 2024. Opportunities, tensions, and challenges in computational approaches to addressing online harassment. InACM DIS
2024
-
[50]
Shirish Karande, V Santhosh, and Yash Bhatia. 2024. Persuasion games with large language models. InACL ICON
2024
-
[51]
Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning. InACL
2025
-
[52]
Cameron Jones and Benjamin Bergen. 2026. Lies, damned lies, and language statistics: a comprehensive review of risks from manipulation, persuasion, and deception with large language models.Artificial Intelligence Review59, 4 (2026), 116
2026
-
[53]
Rohan Leekha, Olga Simek, and Charlie Dagli. 2024. War of Words: Harnessing the Potential of Large Language Models and Retrieval Augmented Generation to Classify, Counter and Diffuse Hate Speech. InAAAI FLAIRS. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextual...
2024
-
[54]
Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A new generation of perspective API: Efficient multilingual character-level transformers. InACM KDD
2022
-
[55]
Nihal Kumarswamy, Mohit Singhal, and Shirin Nilizadeh. 2025. Causal Insights into Parler’s Content Moderation Shift: Effects on Toxicity and Factuality. InACM WWW
2025
-
[56]
Mikel K Ngueajio, Flor Miriam Plaza-del Arco, Yi-Ling Chung, Danda B Rawat, and Amanda Cercas Curry. 2025. Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate. InACL WOAH
2025
-
[57]
Amalie Brogaard Pauli, Isabelle Augenstein, and Ira Assent. 2024. Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language. InNAACL
2024
-
[58]
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey.ACM Computing Surveys56, 9 (2024)
2024
-
[59]
Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019. A benchmark dataset for learning to intervene in online hate speech. InEMNLP-IJCNLP
2019
-
[60]
Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, and Roberto Di Pietro. 2026. A Geometric Analysis of Small-sized Language Model Hallucinations. InForty-third International Conference on Machine Learning
2026
-
[61]
Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden Persuaders: LLMs’ Political Leaning and Their Influence on Voters. InACL EMNLP
2024
-
[62]
Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2025. On the conversational persuasiveness of GPT-4.Nature Human Behaviour9, 8 (2025), 1645–1653
2025
-
[63]
Gautam Kishore Shahi, Benedetta Tessa, Amaury Trujillo, and Stefano Cresci. 2025. A Year of the DSA Transparency Database: What it (Does Not) Reveal About Platform Moderation During the 2024 European Parliament Election. In ICWSM Workshops
2025
-
[64]
Punyajoy Saha, Kanishk Singh, Adarsh Kumar, Binny Mathew, and Animesh Mukherjee. 2022. CounterGeDi: A controllable approach to generate polite, detoxified and emotional counterspeech. InIJCAI
2022
-
[65]
Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support. InACM CHI
2021
-
[66]
Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. 2024. Investigating moderation challenges to combating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on Reddit. InUSENIX
2024
-
[67]
Xiaoying Song, Sujana Mamidisetty, Eduardo Blanco, and Lingzi Hong. 2025. Assessing the human likeness of AI-generated counterspeech. InACL COLING
2025
-
[68]
Serra Sinem Tekiroğlu, Yi-Ling Chung, and Marco Guerini. 2020. Generating counter narratives against online hate speech: Data and strategies. InACL
2020
-
[69]
Benedetta Tessa, Lorenzo Cima, Amaury Trujillo, Marco Avvenuti, and Stefano Cresci. 2025. Beyond trial-and-error: Predicting user abandonment after a moderation intervention.Engineering Applications of Artificial Intelligence162 (2025), 112375
2025
-
[70]
Serra Sinem Tekiroglu, Helena Bonaldi, Margherita Fanton, and Marco Guerini. 2022. Using pre-trained language models for producing counter narratives against hate speech: A comparative study. InACL
2022
-
[71]
Amaury Trujillo and Stefano Cresci. 2023. One of many: Assessing user-level effects of moderation interventions on r/The_Donald. InACM WebSci
2023
-
[72]
Amaury Trujillo, Tiziano Fagni, and Stefano Cresci. 2025. The DSA Transparency Database: Auditing self-reported moderation actions by social media. InACM CSCW
2025
-
[73]
Amaury Trujillo and Stefano Cresci. 2022. Make Reddit Great Again: Assessing community effects of moderation interventions on r/The_Donald. InACM CSCW
2022
-
[74]
Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2024. Hate cannot drive out hate: Forecasting conversation incivility following replies to hate speech. InAAAI ICWSM
2024
-
[75]
Yi Zheng, Björn Ross, and Walid Magdy. 2026. Validating Automatic Evaluation of Controllable Counterspeech Generation: Rankings Matter More Than Scores. InEACL
2026
-
[76]
Haiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao, Minghao Hu, and Bin Zhou. 2024. F2rl: Factuality and faithfulness reinforcement learning framework for claim-guided evidence-supported counterspeech generation. InACL EMNLP
2024
-
[77]
Irune Zubiaga, Aitor Soroa, and Rodrigo Agerri. 2024. A LLM-based ranking method for the evaluation of automatic counter-narrative generation. InACL EMNLP
2024
-
[78]
Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, and Matus Mesarcik
-
[79]
Nawaal Zubair, Shazia Hashmat, and Ume Aimen. 2025. Artificial Intelligence and the Generational Divide: A Study on Trust and Acceptance.Annual Methodological Archive Research Review3, 6 (2025), 19–44
2025
-
[2024]
InAAAI ICWSM
The persuasive power of large language models. InAAAI ICWSM
-
[2025]
You are a moderator of a subreddit and you come across a comment that exhibits hate speech. Generate a brief response that aims to counteract this comment in a respectful manner
Evaluation of LLM vulnerabilities to being misused for personalized disinformation generation. InACL. , Vol. 1, No. 1, Article . Publication date: July 2026. 26 Cima et al. Appendix A Dataset For counterspeech generation, we retrieved 128 toxic comments from 49 Reddit threads....
2026
-
[2026]
Dark personality traits and online toxicity: Linking self-reports to reddit activity.Computers in Human Behavior (2026), 109085
2026
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.