A 315-task benchmark shows top LLMs surpass human experts on cryptography knowledge questions, approach them on proofs, but lag 25 to 30 points behind on capture-the-flag exploitation.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AICrypto: Evaluating Cryptography Capabilities of Large Language Models
A 315-task benchmark shows top LLMs surpass human experts on cryptography knowledge questions, approach them on proofs, but lag 25 to 30 points behind on capture-the-flag exploitation.