GB/T-Bench injects 7,306 artificial errors into 488 Chinese standards to test LLM review, and a multi-agent reviewer lifts the top score from 0.328 to 0.509, but conflicting taxonomy definitions undermine the evaluation.
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents
GB/T-Bench injects 7,306 artificial errors into 488 Chinese standards to test LLM review, and a multi-agent reviewer lifts the top score from 0.328 to 0.509, but conflicting taxonomy definitions undermine the evaluation.