{"id":"b1cbd006-0c8f-4347-8507-046589a09d4c","arxiv_id":"2501.00697","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The first Chinese-language counterspeech dataset of paired hate speech and counterspeech instances, created via an LLM-as-a-Judge pipeline with only partial human verification.","lead":"This paper introduces PANDA, a dataset of paired Chinese hate speech and counterspeech built with an LLM-based scoring pipeline and human review. It also reports that an LLM judge in Chinese tends to favor AI-generated responses over human-written ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset validity hinges on unverified HS labels: only 41.3% of 785 annotated entries confirmed as hate speech, and 2,189 of 2,974 pairs lack human annotation.","rationale":"Reviewing in good faith, the paper's central deliverable is a first-of-its-kind Chinese paired hate-speech/counterspeech dataset. The paper is commendably transparent about pipeline limitations: Figure 3 reports the human-label distribution, and Appendix A acknowledges mislabeling in source data. However, the central claim depends on the released pairs actually being HS-CS pairs. That condition is least secure for the 2,189 of 2,974 pairs that were never human-annotated. The filtering step (Section 3.3) selected entries by Llama-3.1 score and length, but the human check on the first 785 entries showed only 41.3% confirmed as HS. Since the unannotated portion was selected by the same filter, it likely contains a similar or worse mix, meaning a majority of the released 'HS' entries may not be hate speech. The paper does not clarify whether the released dataset contains all 2,974 pairs or only the 785 annotated ones; this ambiguity is itself a correctness risk for the contribution. Additionally, no inter-annotator agreement is reported, leaving label reliability unknown even for the verified subset. These are not merely 'outside consensus' issues but internal consistency issues between the claimed paired structure and the measured label distribution. The concrete test of re-annotating a sample of the unannotated entries, plus reporting agreement, would determine whether the dataset can support the 'first Chinese counterspeech dataset' claim as released, or whether it must be scoped to a smaller verified subset. If the latter, the paper's contribution is still non-trivial but the headline claim needs adjustment. This does not change the reader's conditional verdict, but it sharpens the condition.","tokens_in":11513,"tokens_out":3447,"duration_ms":30681,"concrete_test":"Download the released GitHub dataset and determine how many of the 2,974 HS entries were among the 785 human-annotated instances. Then draw a random sample of 100 entries from the remaining 2,189 unannotated HS entries, have two independent annotators label each as HS, CS, or neither using the paper's functional definition, and compute the confirmed-HS proportion and Cohen's kappa. If the confirmed-HS rate is near 41.3% (or below 50%) or agreement is low, the released corpus must be relabeled or scoped to the verified subset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 selects 2,974 'hate speech' entries using Llama-3.1 hate scores and a length threshold. Section 3.6 reports that human annotators confirmed only 41.3% of the first 785 selected entries as hate speech, with 31.0% labeled counterspeech and 27.7% neither. Because only these 785 entries were human-annotated, the remaining 2,189 pairs in the released dataset have no human verification of the HS side. If the unannotated entries follow the same distribution, roughly 1,200 of the 2,974 pairs are not hate speech at all, which directly undermines the central claim of a paired hate-speech/counterspeech resource. The paper acknowledges this mislabeling but still presents the full 2,974-pair corpus as the contribution; no inter-annotator agreement statistic is reported, so the reliability of even the 785 verified labels is unknown. The most load-bearing concern is therefore that the dataset's core pairing is unverified for most entries, making the 'first Chinese counterspeech dataset' claim contingent on the unannotated portion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PANDA, a dataset of paired Chinese hate speech and counterspeech (CS) entries, reportedly the first such resource for an East Asian language. The construction pipeline combines filtering from existing Chinese hate speech corpora using Llama-3.1 hate scores and string length thresholds, counterspeech generation via simulated annealing over multiple open-source LLMs, and a JudgeLM-based round-robin ranking step. Human annotators then scored the hate speech label of each entry and selected/edited the best CS response. Only 785 of the proposed 2,974 pairs received human annotation; of these, 41.3% were confirmed as hate speech, 31.0% as counterspeech, and 27.7% as neither. The paper also reports a statistical analysis suggesting that JudgeLM systematically downgrades human-preferred responses, interpreting this as a limitation of LLM-as-a-judge evaluation for Chinese counterspeech.","tokens_in":11822,"tokens_out":4193,"duration_ms":39054,"significance":"If the dataset were reliable, it would fill a genuine gap in non-English counterspeech resources and provide a valuable testbed for cross-lingual hate speech intervention research. The paper is also open about the limitations of existing Chinese hate speech labels and contributes a detailed description of the LLM-in-the-loop pipeline, which is useful for replication. Its explicit documentation of mislabeling rates and an LLM-based evaluation bias is a candid and useful lesson for the community. However, the central claim of a paired hate-speech/counterspeech corpus is currently not supported for the majority of the released pairs: only a small fraction has human verification, and even that verification reveals a low hate speech confirmation rate. The significance therefore hinges on whether the authors can either verify the remaining pairs or reframe the resource honestly. The JudgeLM bias analysis is interesting but methodologically confounded by the selection procedure, as noted below.","major_comments":[{"comment":"The central contribution—a paired hate-speech/counterspeech dataset—is undermined by the unverified hate speech labels. Section 3.3 selects 2,974 entries as hate speech using Llama-3.1 hate scores and a length threshold, but Figure 3 (reported in Section 3.6) shows that of the first 785 entries annotated by humans, only 41.3% were confirmed as hate speech, with 31.0% judged counterspeech and 27.7% neither. Since the remaining 2,189 pairs received no human verification, the released corpus likely contains a large fraction of non-hate-speech entries, contradicting the abstract's claim of a paired hate-speech/counterspeech resource. The paper must either verify all 2,974 pairs, release only the verified subset, or clearly reframe the resource as a noisy candidate set with confidence scores; the current abstract and contributions overstate what has been produced.","section":"§3.3, §3.6, Fig. 3"},{"comment":"No inter-annotator agreement statistic is reported, despite four annotators independently scoring the same entries. Cohen's kappa or Krippendorff's alpha is standard for this type of annotation and is essential for assessing whether the 41.3% hate speech confirmation rate is a reliable estimate. Without it, even the human-verified subset cannot be used with confidence for training or evaluation, and the reported distribution in Figure 3 may be dominated by annotation idiosyncrasies.","section":"§3.5–§3.6"},{"comment":"The conclusion that JudgeLM is biased against human-preferred counterspeech is confounded by the selection procedure. The AI responses shown to annotators were themselves selected by JudgeLM through simulated annealing, so the comparison is not between independent human and AI candidates but between judge-selected candidates and human-edited versions of those same candidates. The one-sample t-test demonstrates a ranking difference but cannot support the attribution to 'goal misalignment' or 'bias' without controlling for this selection effect, for example by comparing against responses generated by a different judge or against random AI responses.","section":"§3.6, Table 3"}],"minor_comments":[{"comment":"Section 4.1 states that human annotators were able to create only 785 out of the proposed 2,974 pairs, but the released dataset is described elsewhere as containing 2,974 pairs. Please clarify what is actually released: 785 verified pairs, 2,974 unverified pairs, or a mixture with explicit flags.","section":"§4.1"},{"comment":"The corpus is abbreviated as 'CHSD' in the text but 'CSHD' in Table 1; please make the abbreviation consistent. Also, the 'Political' dataset (Wang et al., 2022) is listed in the table even though it is not open-source and not used in preprocessing; explain why it is included in the survey.","section":"Table 1"},{"comment":"The reference for JudgeLM (Zubiaga et al., 2024) appears to have a corrupted author list ('Irune Zubiaga, Aitor Soroa, and Rodrigo Agerri' seems to be missing at least one author); please verify and correct this citation.","section":"References"},{"comment":"The abstract contains a grammatical error: 'we introduce the a corpus' should be 'we introduce a corpus'. Please proofread the manuscript for similar typos.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses an important resource gap, but the central dataset claim is not currently supported by the evidence. The authors may be able to remedy this by re-annotating or clearly delimiting the verified subset, but as written the paper overstates the validity of the full corpus. I would encourage the editor to ask for a revised version that either provides verification for the full set or explicitly redefines the contribution to match what was actually produced."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I've read the PANDA paper. Bottom line: it's a genuine first effort at a Chinese HS-CS paired dataset, and it reports an honest negative result about LLM-as-a-Judge bias in Chinese. But the central deliverable is shakier than the abstract suggests. Only 785 of the 2,974 pairs were human-annotated, and of those, only 41.3% of the selected \"hate speech\" entries were confirmed as hate speech by the annotators. 31% were themselves counterspeech, and 27.7% neither. That means the unannotated ~2,189 pairs likely follow a similar distribution, so a large chunk of the dataset is not actually HS-CS pairs as claimed. This is a load-bearing flaw, not a minor caveat.\n\nWhat's genuinely new: this is, as far as I know, the first published Chinese counterspeech dataset with paired HS and CS entries. The authors also surface a real problem—JudgeLM consistently ranks human-preferred, logically-engaged responses below AI-generated paraphrases that restate the hate speech. That's a useful empirical finding for anyone building LLM-based evaluation in non-English settings. The paper is also refreshingly transparent about the pipeline's failure modes: the mislabeling statistics, the annotator demographics, and the cost in human hours.\n\nSoft spots beyond the verification issue. The \"novel metrics\" claim in the contributions list is inflated; the metrics are standard BLEU/ROUGE/BERTScore/Novelty/Genlen. The threshold selection for filtering (score ≥51, length ≥53) is described via a heat-map but there's no clear justification or sensitivity analysis. No inter-annotator agreement is reported, so we don't know how reliable even the 785 labels are. Code and full data aren't available at the link in the paper, which hurts reproducibility. The hamming distance for deduplication is acknowledged as crude.\n\nThe authors do flag most of these limitations themselves in Section A, which earns them credit. But the paper still presents the full 2,974-pair corpus as the resource, and the abstract says \"first Chinese counterspeech dataset\" without conditioning on verification. If the dataset is released with clear flags about which pairs are verified, it could be a useful starting point. As is, it's a promising proof-of-concept with a strong negative result, but not yet a clean resource.\n\nFor peer review: I'd send it back for major revisions. The core idea and the negative finding deserve referee time, but the dataset claim needs to be rescaled and the verification statistics need to be much more prominent. A serious referee could help the authors turn this into something solid.","headline":"A genuinely first Chinese counterspeech dataset attempt with an honest negative result about LLM-as-a-Judge bias, but the core pairing is only one-quarter human-verified and the verification rate is low, making the contribution a promising scaffold rather than a finished resource.","tokens_in":12278,"tokens_out":2311,"would_cite":false,"duration_ms":20048,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the first Mandarin Chinese counterspeech corpus, pairing each hate speech instance with generated and human-edited responses.","keywords":["counterspeech","Chinese hate speech","Mandarin","paired dataset","LLM-as-a-judge","simulated annealing","hate speech annotation","evaluation bias"],"falsifier":"Randomly sample the final 2,974 pairs, have a diverse group of native Chinese annotators independently label each hate speech entry as hate speech, counterspeech, or neither, and compare with the pipeline's selection; if the confirmation rate does not rise above the 41.3% found in the first 785 instances, the filtering step has not isolated hate speech. Separately, on a held-out set of pairs, count how often JudgeLM ranks the human-preferred response first; the paper's t-tests predict it will rarely do so.","tokens_in":11291,"feed_emoji":"💬","tokens_out":5686,"duration_ms":48764,"temperature":0.7,"pith_summary":"The paper sets out to close a gap: counterspeech resources for Chinese are virtually nonexistent, and existing hate speech datasets are mostly English or Eurocentric. It introduces PANDA, a paired Mandarin hate-speech–counterspeech corpus of nearly 3,000 designed pairs, built from three open-source Chinese hate speech datasets. The construction pipeline uses an LLM-as-a-Judge to score and rank counterspeech candidates, a simulated annealing search over LLM-generated responses, and a round-robin tournament, followed by manual verification and editing. The paper also reports two findings that matter beyond the dataset itself: human annotators confirmed only 41.3% of the algorithmically selected hate speech entries as actual hate speech, and the LLM judge systematically ranked AI-generated paraphrases above human-edited counterspeech.","feed_headline":"First Mandarin counterspeech dataset pairs hate speech with replies","feed_subtitle":"Built from open Chinese data with LLM judges and human edits, it exposes judge bias toward AI-style replies.","key_machinery":"The load-bearing mechanism is the scoring-and-search pipeline. Hate speech candidates are filtered by a Llama-3.1-70B hate score threshold (at least 51) and a string length threshold (at least 53 characters), then counterspeech candidates are generated by six LLMs and refined through simulated annealing, where acceptance probability follows a Boltzmann distribution over JudgeLM scores. JudgeLM is an LLM-based ranking method that scores relevance, fluency, and effectiveness; a round-robin tournament ranks the surviving candidates, and four per hate speech instance are passed to human annotators for selection and editing. The paper's evidence that JudgeLM ranks human-edited responses below AI paraphrases is what makes the pipeline's output depend on human oversight.","core_discovery":"The paper claims that it is possible to build the first Chinese counterspeech dataset by combining open-source hate speech corpora with an LLM-based scoring pipeline, and that doing so surfaces structural problems in both Chinese hate speech resources and LLM-based evaluation. Its central discovery is the paired PANDA corpus: hate speech instances selected from COLD, SWSR, and CHSD, each paired with four LLM-generated counterspeech candidates refined by human annotators. The accompanying analysis finds that existing Chinese hate speech labels are unreliable — a substantial share of labeled hate speech is actually counterspeech or neutral content — and that JudgeLM's rankings are misaligned with human preferences, favoring responses that restate or rephrase the original hate speech over responses that directly rebut its argument.","pith_inferences":["A natural extension would be to use the 785 human-verified pairs as a held-out test set for calibrating LLM judges, measuring how often a judge's top pick matches the human-preferred response.","The 41.3% confirmation rate suggests that the amount of usable Chinese hate speech data may be far smaller than the raw corpus counts imply, so estimates of Chinese hate speech prevalence based on existing labels could be inflated.","If paraphrase-favoring bias generalizes, automated counterspeech systems evaluated only by LLM judges may drift toward echoing hate speech rather than rebutting it, which has consequences for how such systems are deployed."],"forward_implications":["The PANDA corpus gives Chinese-language hate speech and counterspeech research a paired resource where none existed, enabling direct study of counterspeech strategies in Mandarin.","If the reported judge bias is systematic, LLM-based evaluation of counterspeech in Chinese will need human-aligned scoring or multiple judges before it can be trusted.","The pipeline can be transferred to other East Asian languages with scarce counterspeech resources, such as Korean or Japanese, though cultural adaptation of annotation is required.","The low confirmation rate of hate speech labels in the source datasets argues for re-examining and re-annotating existing Chinese hate speech corpora before they are reused."],"supporting_citations":[{"why":"Provides the COLD dataset, one of the three open-source Chinese hate speech sources used to build the corpus.","marker":"Deng et al. (2022)"},{"why":"Provides the SWSR sexism-labeled Chinese dataset, included to increase representation of sexist hate speech.","marker":"Jiang et al. (2022)"},{"why":"Provides the CHSD dataset of Chinese hate speech used in preprocessing.","marker":"Rao et al. (2023)"},{"why":"Supplies CONAN, the paired hate-speech/counterspeech format the paper adopts for Chinese.","marker":"Chung et al. (2019b)"},{"why":"Defines JudgeLM, the LLM-based ranking method central to scoring and selection of counterspeech candidates.","marker":"Zubiaga et al. (2024)"},{"why":"Shows that LLMs can generate counterspeech zero-shot, the basis for the generation step.","marker":"Saha et al. (2024)"},{"why":"Provides the eight counterspeech strategies that guide annotation and evaluation.","marker":"Chung et al. (2023)"}],"fun_headline_variants":["First Mandarin counterspeech dataset built with LLM judge pipeline","PANDA corpus reveals LLM judges favor rephrasing over rebuttals","Chinese hate speech labels unreliable in new counterspeech dataset","New PANDA dataset pairs Chinese hate speech with human-edited replies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that a Llama-3.1 hate score of at least 51 and a string length of at least 53 select genuine hate speech, but human annotators confirmed only 41.3% of the selected entries, so if that filtering assumption fails the dataset's paired structure is built on mislabeled inputs.","fun_headline_variants_meta":{"raw":{"variants":["First Mandarin counterspeech dataset built with LLM judge pipeline","PANDA corpus reveals LLM judges favor rephrasing over rebuttals","Chinese hate speech labels unreliable in new counterspeech dataset","New PANDA dataset pairs Chinese hate speech with human-edited replies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000151,"raw_usage":{"total_tokens":1180,"prompt_tokens":904,"completion_tokens":276,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":203}},"tokens_in":520,"tokens_out":276,"duration_ms":3520,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:43:30.125450+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample the final 2,974 pairs, have a diverse group of native Chinese annotators independently label each hate speech entry as hate speech, counterspeech, or neither, and compare with the pipeline's selection; if the confirmation rate does not rise above the 41.3% found in the first 785 instances, the filtering step has not isolated hate speech. Separately, on a held-out set of pairs, count how often JudgeLM ranks the human-preferred response first; the paper's t-tests predict it will rarely do so.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CHSD dataset of Chinese hate speech used in preprocessing."},{"cited_title":"A LLM-Based Ranking Method for the Evaluation of Automatic Counter-Narrative Generation","cited_arxiv_id":"2406.15227","evidence_quote":"Defines JudgeLM, the LLM-based ranking method central to scoring and selection of counterspeech candidates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that LLMs can generate counterspeech zero-shot, the basis for the generation step."}],"review_version":1}