DiagChain, a five-stage diagnostic benchmark for evidence-grounded attack chain reconstruction, finds the strongest of six LLM agents completes only 39.6% of reference steps and that failures shift from evidence use to evidence ordering as model scale grows.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
DiagChain, a five-stage diagnostic benchmark for evidence-grounded attack chain reconstruction, finds the strongest of six LLM agents completes only 39.6% of reference steps and that failures shift from evidence use to evidence ordering as model scale grows.