Pith. sign in

REVIEW 2 cited by

SCCD: A Session-based Dataset for Chinese Cyberbullying Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15042 v1 pith:AZXHPMHQ submitted 2025-01-25 cs.CL

classification cs.CL
keywords chinesecyberbullyingdetectiondatasetsccdcommentdatasetsfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rampant spread of cyberbullying content poses a growing threat to societal well-being. However, research on cyberbullying detection in Chinese remains underdeveloped, primarily due to the lack of comprehensive and reliable datasets. Notably, no existing Chinese dataset is specifically tailored for cyberbullying detection. Moreover, while comments play a crucial role within sessions, current session-based datasets often lack detailed, fine-grained annotations at the comment level. To address these limitations, we present a novel Chinese cyber-bullying dataset, termed SCCD, which consists of 677 session-level samples sourced from a major social media platform Weibo. Moreover, each comment within the sessions is annotated with fine-grained labels rather than conventional binary class labels. Empirically, we evaluate the performance of various baseline methods on SCCD, highlighting the challenges for effective Chinese cyberbullying detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ChineseHarm-Bench is a six-category, 6,000-sample Chinese harmful content detection benchmark with a human-annotated knowledge rule base, and a knowledge-augmented fine-tuning baseline that brings small models to near...

  2. Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new Chinese detoxification dataset and 17-model benchmark show that LLMs can remove toxic words but often distort emotional tone, especially for emoji, homophone, and dialogue-based toxicity.

Pith tools