SRE-Bench, a contamination-controlled reverse-engineering benchmark with 262 realistic binary instances, shows the strongest tested AI agent fully solves only 31.5% of instances.
Livecodebench: Holistic and contamination free evalua- tion of large language models for code
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
SRE-Bench, a contamination-controlled reverse-engineering benchmark with 262 realistic binary instances, shows the strongest tested AI agent fully solves only 31.5% of instances.