ChineseHarm-Bench is a six-category, 6,000-sample Chinese harmful content detection benchmark with a human-annotated knowledge rule base, and a knowledge-augmented fine-tuning baseline that brings small models to near state-of-the-art performance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
ChineseHarm-Bench is a six-category, 6,000-sample Chinese harmful content detection benchmark with a human-annotated knowledge rule base, and a knowledge-augmented fine-tuning baseline that brings small models to near state-of-the-art performance.