A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.
For example, if the original question asks about the result of a specific action, the reversed question could ask about the conditions re- quired for that result to occur
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity
A new 4,369-question bilingual benchmark measures LLM cybersecurity performance across 42 categories and three cognitive levels.