A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.
Provide four possible answers, including the correct one, and make sure the incorrect choices are similar or related to the right answer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.