A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.
The multiple-choice question should provide four answer options, with non-correct options being similar or related to the correct answer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.