{"id":"c1c2a994-9576-4e0e-b9ed-0e8103621658","arxiv_id":"2505.15773","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"ToxicTone provides a 52k-clip Mandarin spoken-toxicity dataset with form and source labels, and a multimodal detector that outperforms off-the-shelf text baselines.","lead":"The paper presents ToxicTone, a 52,062-clip, 93-hour Mandarin audio dataset annotated for toxic speech and the tone behind it. The authors also show that combining speech, emotion, and text models detects toxicity better than off-the-shelf text-only tools, though the data filtering may bias toward explicitly rude language.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §3.2 text-toxicity prefilter conditions the corpus on lexically visible toxicity, so the dataset cannot support the central claim of revealing hidden prosodic toxicity; the speech-cue modeling conclusion also lacks a matched text baseline.","rationale":"The dataset's existence and scale are credible: a release link is provided, the collection and annotation pipeline is described in detail, 11 native annotators spent about 900 hours, and the comparison in Table 2 supports the 'largest public Mandarin spoken toxicity dataset' claim. The central weakness is not fabrication but sampling design. The §3.2 text-toxicity gate selects on lexical toxicity, so the resulting corpus cannot by itself demonstrate the paper's headline contribution of capturing hidden prosodic toxicity. The modeling conclusion is similarly underdetermined: the best configuration outperforms a frozen text encoder (ST), but this is not the same as outperforming a strong text-only model trained on the same transcripts. These are fixable through additional annotation of low-score clips and a matched text baseline, which is exactly what a conditional acceptance should request. The reader's weakest assumption identifies the same prefilter bias, and the reader's CONDITIONAL verdict remains appropriate; no adjustment is needed.","tokens_in":8617,"tokens_out":6167,"duration_ms":56064,"concrete_test":"Annotate a stratified sample of ~1,000 clips drawn from the original 770k pool below the 0.75 threshold (e.g., 500 clips scoring 0.3–0.75 and 500 scoring <0.3), using the same 11-annotator protocol, and measure the human toxic rate and the distribution of toxicity sources. If a non-negligible fraction (say, >10%) of low-score clips are toxic by tone or source, the §3.2 filter has excluded the very hidden-toxicity population the paper claims to capture; then re-run Table 3's X+ST+E versus a BERT baseline fine-tuned on the same transcripts to see whether the speech-cue advantage persists. If nearly all low-score clips are non-toxic, the filter is not the fatal bias and the current conclusions stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2's prefilter is the load-bearing step for the paper's central novelty. All 52,062 clips are retained because an ASR transcript scored above 0.75 on the Alibaba-pai text toxicity classifier (plus 600 porn-rule hits). This conditions the corpus on text-visible toxicity: clips whose harm is carried by prosody alone—sarcastic praise, dismissive politeness, threatening calm—are almost certainly discarded. Yet the abstract and §6 claim the dataset reveals hidden toxic expressions and that speech cues are essential. Those claims cannot be sustained by a sample selected to contain toxic words. The source-label distribution is consistent with this bias: Specific Words dominates (~8,000 clips) while Threatening has ~560, and sarcastic/satirical sources are comparatively rare. The F1 comparison in Table 3 is also affected: ST is a frozen SONAR text encoder, not a strong in-domain text baseline; no BERT fine-tuned on the same transcripts is reported, so the X+ST+E improvement may reflect encoder capacity rather than speech-specific information. The 32% human-toxic rate among supposedly prefiltered clips (16,727/52,062) further shows the text score is a weak, noisy gate. The dataset-scale claim likely survives, but the hidden-toxicity and speech-cue conclusions are not supported without a low-score holdout and a matched text baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript introduces ToxicTone, a Mandarin audio dataset of 52,062 two-to-ten-second clips (about 93 hours), collected from web-crawled audio, diarized and transcribed with public models, then prefiltered by a text-based Chinese toxicity classifier at a 0.75 threshold plus 600 rule-based pornographic samples, and annotated by 11 native speakers for both the form and the source of toxicity. The paper reports binary toxicity detection and multi-label source classification using SONAR, XLS-R, and Emotion2Vec features, and claims that the best ensemble (X+ST+E, F1 64.16%) outperforms text-only baselines, underscoring the essential role of speech-specific cues and the ability to reveal hidden toxic expressions.","tokens_in":8897,"tokens_out":4437,"duration_ms":37906,"significance":"If its central claims were supported, ToxicTone would be a valuable resource: it is the largest public Mandarin audio toxicity dataset by utterance count and duration, with a two-axis annotation scheme (form and source) and a public GitHub release. The annotation effort (11 annotators, approximately 900 hours) and the use of real-world topical categories are concrete strengths. The paper also ships a reproducible pipeline built on public models, which is a practical asset. However, the evidence currently supports the dataset's scale and internal statistics, not the stronger claims about non-lexical prosodic toxicity, because the collection pipeline conditions on text-visible toxicity and the modeling comparisons lack a matched in-domain text baseline.","major_comments":[{"comment":"The text-toxicity prefilter is the load-bearing step for the paper's central novelty, and the manuscript provides no evidence that the retained corpus contains hidden, non-lexical toxicity. All 52,062 clips (plus 600 porn-rule clips) were selected because the Alibaba-pai classifier scored their ASR transcripts above 0.75; clips with benign text but toxic prosody—sarcastic praise, dismissive politeness, calmly delivered threats—are discarded by construction. The source-label distribution in Section 3.3 (Specific Words nearly 8,000, Threatening about 560, Sarcastic/Satirical relatively rare) is consistent with lexical bias. The claims in the Abstract and Section 6 that ToxicTone 'uncovers toxic content that may be hidden behind seemingly polite words' therefore require a held-out low-score sample: the authors should annotate a random sample of clips with scores at or below 0.75 and report the prefilter's recall against human labels, along with threshold sensitivity. Without this, the dataset can be described as a large audio corpus of lexically prompted toxic speech, but not as evidence about prosodic-only toxicity.","section":"3.2 (Preprocessing)"},{"comment":"Table 3 does not include a matched text-only model trained on ToxicTone transcripts, so the claimed superiority of X+ST+E does not establish that speech cues are essential. ST is a frozen SONAR text encoder, COLDETECTOR is fine-tuned on the unrelated COLD dataset, and ETOX is a lexicon-based system. The correct control is a text encoder (e.g., bert-base-chinese or RoBERTa) fine-tuned on the ASR transcripts of the same train/dev/test splits with the same labels and decision threshold, reported with the same metrics. If a fine-tuned text model matches or exceeds the reported F1 of 64.16%, the multimodal advantage disappears; if it does not, the result gains support. This experiment is required to back the abstract's 'essential role of speech-specific cues.'","section":"4.2 and Table 3"},{"comment":"No inter-annotator agreement is reported for the human labels. Given that four annotators per sample can still produce two-to-two ties requiring a fifth review, and that only 32% of the prefiltered clips (16,727/52,062, Table 1) are labeled toxic despite the text prefilter, label reliability is a load-bearing property of the dataset. The authors should report agreement statistics (e.g., Fleiss' kappa or Krippendorff's alpha) overall and per form/source label, and they should follow through on the Section 5 promise to release annotator-level annotations.","section":"3.3 (Human Annotation)"}],"minor_comments":[{"comment":"The paragraph justifying the 0.75 threshold is a single sentence; the authors should report precision/recall of the prefilter on a validation sample and the effect of the threshold on the final dataset composition.","section":"3.2 (Preprocessing)"},{"comment":"The paper does not report ASR word error rate or diarization error for the models cited as [16] and [17]; these errors propagate into the text prefilter and into the ST text encoder, so at least a qualitative estimate would help readers interpret the pipeline's noise.","section":"3.2 (Preprocessing)"},{"comment":"The ETOX and COLDETECTOR baselines are not described in enough detail to reproduce their audio-vs-text input configuration; for example, it is unclear how audio is fed to COLDETECTOR and which language resources ETOX uses for Mandarin.","section":"4.3 and Table 3"},{"comment":"For the one-vs-all source classifiers, the manuscript should state whether the same test clip can receive multiple source labels and how the decision threshold is applied, since Figure 3 reports both F1 and accuracy for each source category.","section":"4.4 and Figure 3"},{"comment":"There are several formatting and naming inconsistencies: 'COLDETECTOR' appears as 'COLD ETECTOR' in one line, and 'Etox'/'ETOX' are used interchangeably; please normalize the baseline names.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the dataset resource has clear value for the spoken-language toxicity community, but the current framing overstates what the collection pipeline can support. I would view the manuscript as potentially acceptable after the authors add a low-score holdout evaluation, a matched fine-tuned text baseline, and inter-annotator agreement statistics. I would also ask the authors to address consent and privacy for the web-crawled audio clips, since the manuscript is silent on this despite releasing raw audio."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is real and worth having. ToxicTone fills a clear gap: a public Mandarin spoken-toxicity corpus at this scale did not exist, and the two-axis annotation scheme (form and source) is a useful addition. The 52k clips, 93 hours, diverse topical categories, and 11 native annotators with four judgments per clip all point to a substantial, careful collection effort. That alone justifies attention.\n\nThe soft spot is real, and it is load-bearing. Section 3.2 keeps only clips whose ASR transcript scores above 0.75 on a text toxicity classifier, plus 600 porn-rule hits. That conditions the corpus on lexically visible toxicity. Sarcastic praise, dismissive politeness, or threatening calm without toxic words mostly get discarded before human annotation. So the abstract's claim that the dataset reveals hidden toxic expressions carried by prosody is not supported by the sampling procedure. The fact that annotators judged only 32% of these prefiltered clips as toxic also tells us the text gate is noisy, and we learn nothing about the 718k discarded clips. A low-score holdout, even a few thousand clips, would directly test whether prosodic-only toxicity exists in the discarded population. Without it, the dataset is best described as a corpus of text-toxic candidates, not a corpus of prosodically hidden toxicity.\n\nThe modeling comparison has a second, related weakness. ST is a frozen SONAR text encoder, and COLDETECTOR is not trained on ToxicTone transcripts. That is not a matched text-only baseline. A fine-tuned Chinese BERT or COLDETECTOR on the same transcripts would let the authors isolate whether the X+ST+E gain comes from speech-specific cues or just from a stronger encoder. That experiment is straightforward and should be done.\n\nMinor issues: no inter-annotator agreement, no error bars, and the dataset release is promised but not fully documented. All fixable.\n\nNone of this kills the paper. The dataset is still a contribution, and the taxonomy plus annotator-level labels could make it a standard resource. The right move is to send it to peer review with a request to address the selection-bias question, add a matched text baseline, and report agreement. If those experiments weaken the prosody claim, the authors should soften the conclusion; the dataset remains valuable either way.","headline":"ToxicTone is a genuinely useful Mandarin spoken-toxicity resource, but the paper's central claim that it captures hidden prosodic toxicity is undercut by its own text-based prefilter.","tokens_in":9487,"tokens_out":2501,"would_cite":true,"duration_ms":25436,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new 52,062-clip Mandarin speech dataset labels both the form and the source of toxicity, and a model that hears tone as well as words detects it best.","keywords":["Mandarin speech toxicity","spoken toxicity detection","audio dataset annotation","toxic tone sources","multimodal ensemble","prosody","sarcasm detection","hate speech detection"],"falsifier":"Take a random sample of the clips that scored below 0.75 on the text filter and have native Mandarin speakers annotate them with the same two-layer scheme; if a substantial fraction are toxic, particularly via angry, dismissive, or sarcastic tone, the dataset's claim to capture tone-hidden toxicity fails.","tokens_in":8419,"feed_emoji":"🎙️","tokens_out":6263,"duration_ms":50733,"temperature":0.7,"pith_summary":"This paper introduces ToxicTone, a public Mandarin audio dataset of 52,062 short clips (93 hours) in which each clip is annotated both for the form of toxicity (profanity, hate speech, pornographic language, bullying, sarcasm, other) and for the source of toxicity (specific words, angry or violent tone, dismissive or impatient tone, sarcastic or satirical tone, threatening tone). The authors argue that text-only toxicity detection misses spoken Mandarin toxicity because harm often lives in prosody, such as intonation, emphasis, and rhythm, rather than in the words themselves. They demonstrate that a model combining text, speech, and emotion embeddings detects toxicity better than text-only baselines, with the best configuration reaching 64.16% F1. If the dataset and result hold, they provide the first large-scale benchmark for studying how tone carries toxicity in Mandarin.","feed_headline":"93 hours of Mandarin speech labeled for toxic tone","feed_subtitle":"The largest public Mandarin speech-toxicity corpus tags angry and sarcastic delivery—and audio models that hear it win.","key_machinery":"The load-bearing object is the annotation scheme: a two-axis label space (form of toxicity and source of toxicity) applied to real-world web-crawled audio, split into 2-to-10-second clips. The pipeline first runs speaker diarization, ASR transcription via K2D, and a text-based BERT toxicity filter (score above 0.75) that reduces 770k candidate clips to 52k, augmented with 600 rule-matched pornographic-language samples. Human annotation by 11 native Mandarin speakers resolves ties with a fifth annotator. The detection model is an ensemble of three pre-trained encoders—XLS-R 1B (acoustic), SONAR text (linguistic), and Emotion2Vec+ Large (emotional)—feeding a three-layer linear classifier; the same features are also used for one-vs-all source classification.","core_discovery":"ToxicTone is, to the authors' knowledge, the largest public spoken toxicity dataset for Mandarin, and it is built to separate what is said from how it is said. Each of the 52,062 two-to-ten-second clips carries two annotation layers: the form of toxicity and its source, where the source layer captures tone-based toxicity such as dismissiveness, sarcasm, anger, and threat that can co-occur with innocuous words. The paper's central empirical finding is that the multimodal ensemble $X+S_T+E$—concatenating XLS-R speech features, SONAR text embeddings of ASR transcripts, and Emotion2Vec+ emotion embeddings—outperforms all text-only and single-encoder baselines on binary toxicity detection (F1 64.16% vs. 50.54% for the best text baseline), with the gains concentrated on the tone-driven source categories. The authors read this as evidence that speech-specific prosodic cues are not optional extras but necessary signal for spoken toxicity detection.","pith_inferences":["The text-based filter likely sets an upper bound on how much purely prosodic toxicity the dataset can contain; auditing rejected clips would quantify that ceiling.","Because source labels are fine-grained, they could be used as auxiliary supervision or as targets for a hierarchical model, potentially improving the weakest categories (sarcasm, threat).","The form and source annotation scheme is language-neutral and could transfer to other dialects or languages where indirect toxicity is common.","Re-running the same pipeline with an audio-based or multimodal prefilter on the original 770k candidate clips could grow the dataset substantially and test whether tone-driven toxicity was being screened out."],"forward_implications":["ToxicTone provides a public 52,062-clip benchmark with both binary toxicity labels and fine-grained form and source labels for Mandarin speech.","Combining acoustic (XLS-R), linguistic (SONAR text), and emotional (Emotion2Vec+) features outperforms text-only and single-modality models on binary toxicity detection.","The best configuration reaches 64.16% F1, while the strongest text baseline reaches 50.54%, showing that speech cues carry signal text misses.","Source classification is strongest for word-based toxicity and weakest for threat, reflecting label imbalance and the subtlety of tone-based categories."],"supporting_citations":[{"why":"COLD benchmark for Chinese offensive text; supplies the COLDETECTOR baseline and the text-only comparison point.","marker":"[3]"},{"why":"ToxiCN's hierarchical text labels motivate the paper's finer-grained form and source annotation.","marker":"[6]"},{"why":"DeToxy, an English spoken-utterance toxicity dataset, is the prior art on audio toxicity the paper extends from.","marker":"[9]"},{"why":"MuTox provides the multilingual audio toxicity detector used as a baseline and the comparison dataset in Table 2.","marker":"[10]"},{"why":"Speaker diarization pipeline used to segment multi-speaker web audio into clips.","marker":"[16]"},{"why":"K2D ASR model that transcribes clips for filtering and for the text encoder input.","marker":"[17]"},{"why":"SONAR supplies the sentence-level text and speech embeddings used as $S_T$ and $S_A$.","marker":"[23]"},{"why":"XLS-R supplies the acoustic speech representation $X$ used in the best ensemble.","marker":"[24]"},{"why":"Emotion2Vec+ supplies the emotional embedding $E$ that, combined with $X$ and $S_T$, gives the top F1.","marker":"[26]"}],"fun_headline_variants":["Largest Mandarin speech-toxicity dataset tags tone, not just words","Hear the toxicity: audio beats text for Mandarin hate speech","93 hours of Mandarin audio tagged for both form and tone of toxicity","Audio reveals hidden sarcasm and anger: ToxicTone dataset covers 52K clips","Multimodal beats text-only for Mandarin toxicity detection on new dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that a text-based classifier scoring above 0.75 is a sufficient gate for what counts as toxic speech, so utterances whose toxicity lives mainly in tone rather than words may be discarded before any human sees them.","fun_headline_variants_meta":{"raw":{"variants":["Largest Mandarin speech-toxicity dataset tags tone, not just words","Hear the toxicity: audio beats text for Mandarin hate speech","93 hours of Mandarin audio tagged for both form and tone of toxicity","Audio reveals hidden sarcasm and anger: ToxicTone dataset covers 52K clips","Multimodal beats text-only for Mandarin toxicity detection on new dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2916,"prompt_tokens":902,"completion_tokens":2014,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1919}},"tokens_in":518,"tokens_out":2014,"duration_ms":11504,"temperature":1.0,"reasoning_tokens":1919,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:11:07.588758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the clips that scored below 0.75 on the text filter and have native Mandarin speakers annotate them with the same two-layer scheme; if a substantial fraction are toxic, particularly via angry, dismissive, or sarcastic tone, the dataset's claim to capture tone-hidden toxicity fails.","supporting_citations":[{"cited_title":"Definition of toxicity We define toxicity via two aspects: theformof toxicity and the sourceof toxicity","cited_arxiv_id":null,"evidence_quote":"COLD benchmark for Chinese offensive text; supplies the COLDETECTOR baseline and the text-only comparison point."},{"cited_title":"Unlike prior text-based datasets, our dataset incorporates prosodic cues and detailed toxicity labels, enabling a more nuanced understanding of harmful speech","cited_arxiv_id":null,"evidence_quote":"ToxiCN's hierarchical text labels motivate the paper's finer-grained form and source annotation."},{"cited_title":"On- line networks of racial hate: A systematic review of 10 years of research on cyber-racism,","cited_arxiv_id":null,"evidence_quote":"DeToxy, an English spoken-utterance toxicity dataset, is the prior art on audio toxicity the paper extends from."},{"cited_title":"COLD: A benchmark for Chinese offensive language detection,","cited_arxiv_id":null,"evidence_quote":"MuTox provides the multilingual audio toxicity detector used as a baseline and the comparison dataset in Table 2."},{"cited_title":"DeToxy: A Large-Scale Multimodal Dataset for Toxicity Clas- sification in Spoken Utterances,","cited_arxiv_id":null,"evidence_quote":"Speaker diarization pipeline used to segment multi-speaker web audio into clips."},{"cited_title":"toxic” and “non-toxic","cited_arxiv_id":null,"evidence_quote":"K2D ASR model that transcribes clips for filtering and for the text encoder input."},{"cited_title":"Building a Taiwanese Mandarin Spoken Language Model: A First Attempt,","cited_arxiv_id":null,"evidence_quote":"SONAR supplies the sentence-level text and speech embeddings used as $S_T$ and $S_A$."},{"cited_title":"Leave no knowledge behind during knowledge distillation: Towards practical and effective knowledge distillation for code-switching asr using realistic data,","cited_arxiv_id":null,"evidence_quote":"XLS-R supplies the acoustic speech representation $X$ used in the best ensemble."},{"cited_title":"Antiso- cial behavior in online discussion communities,","cited_arxiv_id":null,"evidence_quote":"Emotion2Vec+ supplies the emotional embedding $E$ that, combined with $X$ and $S_T$, gives the top F1."}],"review_version":1}