CANDY, a Chinese misinformation fact-checking benchmark, shows LLMs reach only ~76% accuracy on contamination-free claims and frequently fabricate supporting evidence, while serving better as human assistants than autonomous judges.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
CANDY, a Chinese misinformation fact-checking benchmark, shows LLMs reach only ~76% accuracy on contamination-free claims and frequently fabricate supporting evidence, while serving better as human assistants than autonomous judges.