Pith. sign in

REVIEW 4 cited by

ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12223 v1 pith:GSGZUFYO submitted 2024-06-18 cs.CL cs.CY

classification cs.CLcs.CY
keywords offensivelanguageperturbationscontentdetectionchinesecloakingdetecting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment. This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations. We introduce \textsf{ToxiCloakCN}, an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations. Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied. We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness. Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lost in Pronunciation: Detecting Chinese Offensive Language Disguised by Phonetic Cloaking Replacement

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new 500-post benchmark of naturally occurring phonetic cloaking shows LLMs detect such Chinese offensive language with F1 at most 0.672, and Pinyin-augmented prompting partially repairs the gap.

  2. Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon

    cs.CL 2025-05 conditional novelty 6.0 of 10

    C2TU combines a Chinese pronunciation graph, a toxic lexicon, and language-model probability checking to find and correct homophone-cloaked toxic words without any training.

  3. Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new Chinese detoxification dataset and 17-model benchmark show that LLMs can remove toxic words but often distort emotional tone, especially for emoji, homophone, and dialogue-based toxicity.

  4. Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuned BERT-like models outperform zero-shot and internal-state LLM methods on four of six challenging text classification datasets.

Pith tools