Pith. sign in

REVIEW 3 cited by

CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17581 v2 pith:EXNY7MV2 submitted 2025-01-29 cs.CL cs.AIcs.CYcs.SI

classification cs.CLcs.AIcs.CYcs.SI
keywords counterspeechevaluationautomatedauto-calibratedhumanmetricsaggressivenessauto-cseval
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Counterspeech has emerged as a popular and effective strategy for combating online hate speech, sparking growing research interest in automating its generation using language models. However, the field still lacks standardised evaluation protocols and reliable automated evaluation metrics that align with human judgement. Current automatic evaluation methods, primarily based on similarity metrics, do not effectively capture the complex and independent attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence. This has led to an increased dependency on labor-intensive human evaluations to assess automated counter-speech generation methods. To address these challenges, we introduce CSEval, a novel dataset and framework for evaluating counterspeech quality across four dimensions: contextual-relevance, aggressiveness, argument-coherence, and suitableness. Furthermore, we propose Auto-Calibrated COT for Counterspeech Evaluation (Auto-CSEval), a prompt-based method with auto-calibrated chain-of-thoughts (CoT) for scoring counterspeech using large language models. Our experiments show that Auto-CSEval outperforms traditional metrics like ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can NLP Tackle Hate Speech in the Real World? Stakeholder-Informed Feedback and Survey on Counterspeech

    cs.CL 2025-08 conditional novelty 6.0 of 10

    NLP counterspeech research increasingly relies on recycled datasets and excludes the affected communities, according to a systematic review and NGO case study.

  2. Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    HiPPrO generates counterspeech conditioned on both a strategy and an emotion via hierarchical prefix learning plus preference optimization, and reports gains on a new emotion-labeled corpus.

  3. Contextualized Counterspeech: Strategies for Adaptation, Personalization, and Evaluation

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Contextualized counterspeech from LLaMA2-13B using conversation and user history is rated more adequate and persuasive than generic counterspeech, but algorithmic metrics rank configurations inconsistently with humans.

Pith tools