{"id":"6f1aa224-81fe-496f-9986-a00ca92cd6e5","arxiv_id":"2501.15451","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"STATE ToxiCN provides the first span-level Chinese hate speech dataset with 9,533 target-argument-hateful-group quadruples and a 830-term annotated hateful slang lexicon, and baseline results show fine-tuned models outperform LLM APIs.","lead":"STATE ToxiCN is a new Chinese hate speech resource that marks the exact spans of targets and arguments in 8,029 posts, plus a lexicon of 830 explained hateful slang terms. The authors show that fine-tuned open models clearly outperform general-purpose LLM APIs on this fine-grained span extraction benchmark.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unquantified Section 3.2 filtering may make STATE ToxiCN a benchmark for explicitly targeted posts only, narrowing the 'first span-level Chinese hate speech dataset' claim.","rationale":"The paper's stated goal is to provide a reusable span-level Chinese hate speech resource, so the claim that matters is not the raw model scores but that STATE ToxiCN is a valid span-level sample. That validity rests on Section 3.2's filtering: if excluding posts without a clear Target-Argument structure silently drops most implicit, context-dependent hate, the resource is really an 'explicitly-targeted' subset. Because TOXICN contains implicit toxic language by design, the comparison is available. The paper reports no per-rule exclusion counts or retained/excluded comparisons, so the scope restriction is unquantified. Even if this bias exists, the dataset is still useful, but the headline 'first span-level Chinese hate speech dataset' would need qualification as first for explicitly targeted posts. The test-set learning-rate selection in Appendix F and the unspecified construction of the 502-post hateful-slang subset in Section 5.3.1 are separate problems: they undermine the numerical baselines and the RQ3 comparison, but not the resource itself. Given the public data/code release, diverse annotator pool, and reported IAA, the right disposition remains conditional: require the authors to quantify the filtering effect and fix the evaluation protocol.","tokens_in":15488,"tokens_out":8115,"duration_ms":78177,"concrete_test":"Obtain the public TOXICN corpus and reconstruct the Section 3.2 filtering pipeline, recording the number of posts removed by each rule (length, non-text, no clear Target-Argument, etc.). Compare retained and excluded posts on TOXICN's explicit/implicit toxicity labels, on target specificity, and on category distribution. If excluded posts are substantially more likely to be implicit or to lack a specific target, the benchmark is narrower than the headline claim; if the two groups are comparable on these axes, the filtering concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a resource claim, so the load-bearing condition is that the 8,029 posts in STATE ToxiCN are a usable sample of Chinese hate speech rather than an artificial subset chosen for annotatability. Section 3.2 states that posts 'lacking a clear Target-Argument structure' were excluded, but no counts, per-rule statistics, or comparison against the source TOXICN corpus are reported. The original TOXICN dataset explicitly includes both explicit and implicit toxic language, so it is plausible that the excluded posts are disproportionately implicit, vague, or context-dependent. If that is true, the benchmark and every span-level result in Tables 6 and 7 are conditional on an explicitly targeted sub-population, and the 'first span-level Chinese hate speech dataset' headline overstates the scope. The paper's own Section 5.3.1 acknowledges that removing posts can remove 'implicit hate expressions,' which confirms that the mechanism is real. The fix is not harder annotation; it is quantifying the selection effect and stating the scope.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces STATE ToxiCN, a span-level Chinese hate speech dataset of 8,029 posts derived from the post-level TOXICN dataset, annotated with 9,533 Target-Argument-Hateful-Group quadruples. It also presents a lexicon of 830 Chinese hateful slang terms with group labels and explanatory definitions. The authors evaluate twelve open-source, safety-domain, and closed-source LLMs on target/argument span extraction, hatefulness classification, and group classification under hard and soft matching, and conduct case studies of LLM understanding of hateful slang. The central claims are that STATE ToxiCN is the first span-level Chinese hate speech dataset and that the lexicon is the first annotated Chinese hateful slang lexicon with interpretive annotations.","tokens_in":15673,"tokens_out":5519,"duration_ms":46249,"significance":"If the resource is representative, it fills a genuine gap: Chinese hate speech research lacks fine-grained span-level annotations, and the slang lexicon addresses a real evasion phenomenon. The paper ships public code and data, reports inter-annotator agreement on all label types, and provides detailed annotation guidelines and quality-control procedures. The evaluation across twelve models, including safety-specific models, gives a useful baseline. The significance of the resource claims, however, depends on how the filtering from TOXICN affects representativeness and on whether the reported model comparisons are unbiased, so I treat those as load-bearing.","major_comments":[{"comment":"The exclusion of posts 'lacking a clear Target-Argument structure' is not quantified. The paper reports no counts, per-rule statistics, or comparison of the retained posts with the source TOXICN corpus. Given that TOXICN includes explicit and implicit toxic language, and Section 5.3.1 states that the removed posts include 'implicit hate expressions,' the filtering may systematically drop implicit or vaguely targeted posts. If so, STATE ToxiCN is a benchmark for explicitly targeted posts only, and the headline 'first span-level Chinese hate speech dataset' overstates its scope. Please report the number of posts removed at each filtering step, annotate a sample of excluded posts to characterize their target-argument structure, and discuss the resulting coverage of implicit hate speech.","section":"§3.2 (Data Source and Filtering)"},{"comment":"Learning rate selection is performed on the test set: the authors train models for each learning rate, select the result with the highest F1 on the test set, and then calculate the final performance by weighted averaging. This makes the reported numbers in Tables 6 and 7 optimistic and not unbiased estimates of generalization. The benchmark comparison cannot support conclusions about model ranking without a proper validation split or nested cross-validation. Please introduce a development set for hyperparameter selection, report the chosen learning rates, and clarify the 'weighted averaging' procedure.","section":"Appendix F (Detailed Information of the Fine-tuning)"},{"comment":"The procedure for identifying the 502-post hateful-slang subset is not specified. It is unclear whether posts were selected by lexicon matching, manual annotation, or some other rule, and what coverage of the lexicon this subset represents. Since RQ3 compares model performance on this subset with the full test set, a confounded selection method could drive the observed differences. Please specify the selection algorithm, its validation, and report how many of the 830 lexicon terms appear in the subset.","section":"§5.3.1 (Impact of Hateful Slang on Chinese Hate Speech Detection)"}],"minor_comments":[{"comment":"The text states 6,034 hateful quadruples (63.60%), while Table 4 reports 6,063 hateful quadruples (63.60%); the counts are inconsistent and the percentage does not match the stated total. The text also cites 854 multi-group instances while Table 4 lists 866. Please reconcile these numbers.","section":"§3.4 and Table 4"},{"comment":"The caption reads 'Annotators Demographics' but the table lists training hyperparameters; the captions for Tables 8 and 9 appear to be swapped.","section":"Table 8"},{"comment":"The word 'Slience' should be 'Silence' in the English gloss for '默'.","section":"Table 1"},{"comment":"The difference values are printed as plain numbers immediately after the F1 scores (e.g., '56.83 2.32'), but the text refers to superscripts; using explicit '±' or bracketed differences would improve readability.","section":"Table 7"},{"comment":"The phrase 'To reduce hyperparameter sensitivity' is misleading because the described test-set selection actually increases sensitivity to the test set; please rephrase once the methodology is fixed.","section":"Appendix F"}],"recommendation":"major_revision","confidential_remarks":"The paper builds directly on TOXICN by overlapping authors; this is a source dependency rather than a circularity problem, but the requested quantification of the filtering step is important for assessing the scope of the new resource claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper delivers a genuinely new resource: the first span-level Chinese hate speech dataset with Target-Argument-Hateful-Group quadruples, plus an 830-term annotated Chinese hateful slang lexicon with explanations. That claim holds up against the Chinese datasets they cite, which are all post-level. The annotation process is careful: multi-stage, with Fleiss' kappa in the 0.61–0.75 range, which is acceptable for span boundaries in Chinese. The evaluation is thorough: 12 models, fine-tuned and API, hard and soft matching, with error bars in the main tables. The finding that fine-tuned models only reach ~25% hard-match on quadruples, and APIs stay in single digits, is a useful reality check.\n\nThe soft spots are real but fixable. First, Section 3.2 excludes posts that lack a clear Target-Argument structure but gives no counts or comparison against the source TOXICN corpus. Since TOXICN explicitly includes implicit toxic language, the filtering likely skews STATE toward explicitly targeted posts. The paper never states this as a scope limitation, which makes the \"first span-level Chinese hate speech dataset\" headline a bit broader than the actual data. The authors' own discussion in 5.3.1 shows that a removal step can drop implicit hate expressions, so the same risk applies here. Quantifying the selection effect would fix it. Second, the learning rate is selected on test-set F1 (Appendix F). That is test-set tuning, so the absolute scores are probably slightly optimistic, even if relative model comparisons are less affected. Third, the 502-post hateful-slang subset in Section 5.3.1 is defined without saying how the posts were identified as containing slang. That should be specified for reproducibility.\n\nNone of these are fatal. The dataset and lexicon are substantial additions, and the paper is honest about many limitations (annotator subjectivity, Chinese boundary ambiguity, lexicon coverage). I would like to see a revision with filtering statistics, validation-set hyperparameter selection, and a documented slang-subset procedure. With those changes, the benchmark numbers would be on solid ground.\n\nWho this is for: anyone working on Chinese content moderation, span-level extraction, or hate speech resources. It deserves a serious referee, not a desk reject.\n\nRecommendation: send to review with a request for the filtering analysis and hyperparameter clarification.","headline":"Genuinely useful first span-level Chinese hate speech dataset and slang lexicon, but the unquantified filtering step needs to be addressed before the benchmark scope claim is accepted.","tokens_in":16189,"tokens_out":5396,"would_cite":true,"duration_ms":45558,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces STATE ToxiCN, the first span-level Chinese hate speech dataset, with 8,029 posts annotated as Target-Argument-Hateful-Group quadruples, alongside the first annotated Chinese hateful slang lexicon of 830 terms.","keywords":["Chinese hate speech detection","span-level annotation","target-argument extraction","hateful slang lexicon","large language model evaluation","benchmark dataset","toxicity detection"],"falsifier":"Take a random sample of the TOXICN posts that were excluded from STATE ToxiCN and annotate them with the same quadruple scheme; if a substantial fraction turn out to express identifiable hate toward a clear target, the filtering has biased the benchmark. A simpler version: compare the proportion of implicit-hate labels in the excluded set versus the included set, and look for a large imbalance.","tokens_in":15307,"feed_emoji":"🏷️","tokens_out":13055,"duration_ms":102078,"temperature":0.7,"pith_summary":"Chinese hate speech detection has lagged behind English, and existing Chinese datasets label whole posts rather than the spans that actually carry hate. This paper tries to close that gap by constructing STATE ToxiCN, which it presents as the first span-level Chinese hate speech dataset: 8,029 posts annotated with 9,533 Target-Argument-Hateful-Group quadruples that record who is attacked, what is said about them, whether the statement is hateful, and which group is targeted. The paper also builds the first annotated Chinese hateful slang lexicon, 830 terms gathered from real online forums with explanations of their meaning, usage, and targeted groups. Using these resources, it evaluates twelve large language models and finds that fine-tuned open models clearly outperform API-only models, while both struggle with exact span boundaries and with slang that hides hate behind homophones, merged characters, and historical allusions. The contribution is a fine-grained resource that lets researchers measure not just whether a Chinese post is hateful but how the hate is directed and disguised.","feed_headline":"New benchmark splits Chinese hate speech into target, slur, group","feed_subtitle":"An 8,029-post corpus with 9,533 labeled quadruples and an 830-term slang lexicon shows where large models fail.","key_machinery":"The load-bearing mechanism is the Target-Argument-Hateful-Group quadruple, a labeled tuple (Target, Argument, Hateful, Group) extracted from a single post. Target is the attacked span, Argument is the claim or characterization made about it, Hateful marks whether the pair constitutes hate speech, and Group names the targeted category (sexism, racism, region, LGBTQ, others; multiple groups are allowed). This schema converts post-level classification into a structured extraction task, forcing models to say exactly which spans carry hate and toward whom. It is paired with hard- and soft-matching metrics, where hard matching demands exact boundary agreement and soft matching credits overlapping predictions, and with the 830-term annotated Chinese hateful slang lexicon that supplies the cultural background needed to interpret disguised hate.","core_discovery":"The paper's central claim is that span-level Chinese hate speech detection is a tractable and necessary next step, and that the missing piece has been fine-grained annotation resources. Its discovery is the dataset itself: by annotating Target-Argument-Hateful-Group quadruples, the authors show that the target and argument of Chinese hate speech can be extracted at span level even though Chinese lacks word delimiters and permits flexible word order. In their evaluation, no model solves the task: the best fine-tuned models reach hard-match F1 scores below 30% on full quadruples, and API-only LLMs fall below 12%, identifying precise span identification as the main bottleneck. For hateful slang, LLM APIs display better background knowledge than fine-tuned models but still miss culturally specific terms such as merged-character insults. The paper reads these results as evidence that the benchmark exposes actionable gaps for future Chinese hate speech detection.","pith_inferences":["A testable extension the paper leaves open: annotate the TOXICN posts that were filtered out for lacking a clear Target-Argument structure and compare their implicit-hate rate with the included posts; if the excluded posts are systematically harder, the benchmark overstates how well models handle indirect Chinese hate speech.","The paper's own limitation section concedes that flexible grammar and ambiguous boundaries make precise span annotations imperfect; since hard matching depends entirely on boundary precision, some of the reported model shortfall may reflect annotation boundary noise rather than pure model failure.","The limitation section also warns that the 830-term lexicon will miss rapidly evolving internet slang, so the difficulty of the slang portion will drift and will need periodic lexicon refresh to stay representative.","Because the dataset includes non-hate Target-Argument pairs, it supports contrastive learning between hateful and non-hateful statements about the same target, an avenue the paper does not exploit."],"forward_implications":["Fine-tuned open models such as LLaMA3-8B and Qwen2.5-7B can identify target and argument spans with soft-match F1 scores near 70% or higher, but full quadruple hard-match scores stay below 30%, so the benchmark resets expectations for joint extraction.","API-only LLMs without task-specific fine-tuning lag far behind, with quadruple hard-match F1 below 12%, showing that few-shot prompting alone is not enough for span-level Chinese hate extraction.","Hateful slang degrades fine-tuned models' target and argument extraction while slightly improving hatefulness classification, indicating that the main slang problem is span and group identification, not hate detection.","LLM APIs outperform fine-tuned models at explaining culturally grounded slang such as '冉闵', which points to background-knowledge infusion as a concrete path for improving smaller models.","The hard- and soft-matching evaluation protocol provides a reusable way to compare future models despite the ambiguity of Chinese span boundaries."],"supporting_citations":[{"why":"Supplies the TOXICN posts and post-level labels that STATE ToxiCN filters and re-annotates into span-level quadruples.","marker":"Lu et al. (2023)"},{"why":"Defines the Target-Argument-Harmful triple task whose schema STATE ToxiCN extends by adding a Group label.","marker":"Zampieri et al. (2023)"},{"why":"Introduced toxic span detection, the span-level task lineage this paper brings to Chinese for the first time.","marker":"Pavlopoulos et al. (2021)"},{"why":"Provides the COLD Chinese post-level offensive-language benchmark used to show that prior Chinese resources stop at post level.","marker":"Deng et al. (2022)"},{"why":"Supplies a prior Chinese sexism dataset and lexicon that lack span-level annotation and interpretable slang explanations.","marker":"Jiang et al. (2022)"},{"why":"Provides the CDial-Bias Chinese dialogue dataset, evidence that existing Chinese resources remain post- or utterance-level.","marker":"Zhou et al. (2022)"},{"why":"Documents cloaking perturbations such as homophones and character merging that motivate the hateful slang lexicon.","marker":"Xiao et al. (2024)"},{"why":"Supplies the soft-matching algorithm the paper uses to score predictions when exact Chinese span boundaries are ambiguous.","marker":"Han et al. (2023)"},{"why":"Supplies the inter-annotator agreement measure used to report annotation reliability across target, argument, hateful, and group labels.","marker":"Fleiss (1971)"}],"fun_headline_variants":["First Chinese hate-speech dataset with span-level quadruples","Chinese hate speech parsed into target, argument, slur, group","Span-level Chinese hate speech dataset exposes model gaps","LLMs miss Chinese hate speech spans in new benchmark"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's representativeness rests on the assumption that removing posts without a clear Target-Argument structure removes only unusable noise, not a meaningful share of real-world Chinese hate speech; if that assumption fails, the dataset oversamples explicit, well-formed hate and under-represents the implicit cases that are hardest to detect.","fun_headline_variants_meta":{"raw":{"variants":["First Chinese hate-speech dataset with span-level quadruples","Chinese hate speech parsed into target, argument, slur, group","Span-level Chinese hate speech dataset exposes model gaps","LLMs miss Chinese hate speech spans in new benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000866,"raw_usage":{"total_tokens":3728,"prompt_tokens":895,"completion_tokens":2833,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":2768}},"tokens_in":511,"tokens_out":2833,"duration_ms":19739,"temperature":1.0,"reasoning_tokens":2768,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:16:11.434596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the TOXICN posts that were excluded from STATE ToxiCN and annotate them with the same quadruple scheme; if a substantial fraction turn out to express identifiable hate toward a clear target, the filtering has biased the benchmark. A simpler version: compare the proportion of implicit-hate labels in the excluded set versus the included set, and look for a large imbalance.","supporting_citations":[],"review_version":1}