{"id":"985f6d91-fa43-4700-9ad8-b50d4e93acdb","arxiv_id":"2608.03591","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DiagChain, a five-stage diagnostic benchmark for evidence-grounded attack chain reconstruction, finds the strongest of six LLM agents completes only 39.6% of reference steps and that failures shift from evidence use to evidence ordering as model scale grows.","lead":"A new benchmark scores LLM agents on five separate stages of cyberattack chain reconstruction instead of one final accuracy number. Tested on it, the best of six AI models still gets only 39.6% of attack steps right, with small models losing evidence they found and large models misordering it.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gold reference packages rest on a single-author audit with no inter-annotator agreement; if step boundaries or reference order are subjective, every stage-wise score and the bottleneck-shift claim inherit that bias.","rationale":"The reader's weakest-assumption diagnosis is correct: the gold reference packages are the linchpin of every quantitative claim in the paper. The paper itself discloses the single-author audit and the absence of inter-annotator agreement, and it is the right place to focus a conditional verdict. I considered alternative concerns—single-run RQ1 results without confidence intervals, the RQ4 operating point chosen from a GLM-only sweep, and the correlation of source family with chain length—but these are secondary: even with error bars or a different operating point, the core finding of low absolute performance and the qualitative bottleneck shift would likely survive. Gold-label correctness, by contrast, could invalidate the entire measurement if step boundaries and ordering are systematically non-objective. The paper has substantial independent support: SHA-256 manifests, a zero-finding leakage audit, a C0/C1 evidence-dependence control, a 60/60 manual validation of evaluator outputs, and a frozen, reproducible protocol. These justify keeping the verdict at CONDITIONAL rather than moving to REJECT; the independent-annotation test would determine whether the condition can be lifted.","tokens_in":40514,"tokens_out":5129,"duration_ms":53795,"concrete_test":"Commission two independent annotators with security/forensics experience, blinded to the gold packages, to reconstruct reference chains from the same model-visible evidence cards for a stratified sample of 12 MAIN-69 cases (4 AutoLabel, 4 ExCyTIn, 4 OTRF), using the paper's stated step-unit definitions. Compute per-case B3 grouping F1 and pairwise order accuracy between each annotator's chain and the released gold, plus Cohen's kappa on step-boundary and support-assignment decisions. Pre-register a threshold (e.g., median agreement >= 0.7). If agreement is below threshold, the gold standard cannot support the precise 39.6% figure or the bottleneck-shift claim; if agreement is high, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing premise is gold-standard correctness. Appendix B ('Reference Construction and Verification') states that one author audited all 69 reference packages, 849 steps, and 780 edges, and Table 8 reports zero corrections; the paper explicitly says 'This was a full author audit rather than independent double annotation, so inter-annotator agreement is not reported.' Every DiagChain metric (Ret., Grp., Ord., Grd., Gap) and the E1–E4/OK funnel in Appendix D are computed against these packages, and the headline '39.6% of the 849 reference steps' and the size-ordered E2-to-E4 bottleneck shift are proportions over reference steps. If reference step granularity, supporting-evidence assignment, or reference order are systematically wrong or reflect one author's subjective choices, then the E2 versus E4 rates—and the central claim that larger models fail at ordering rather than evidence incorporation—could be an artifact of the gold segmentation rather than a fact about model behavior. The risk is heightened because 39 of 69 cases derive from AutoLabel, a dataset produced by the same research group, so source labels and reference construction share assumptions. The paper is admirably transparent about this limitation, but transparency does not remove the empirical dependence of the central claims on unvalidated ground truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DiagChain, a diagnostic benchmark and agent workflow for evidence-grounded attack chain reconstruction. It contributes MAIN-69, a suite of 69 scenarios built from AutoLabel, ExCyTIn-Bench, and OTRF APT29 sources under clean/noisy/raw evidence profiles; ECRAG, a retrieval-augmented generation kernel coupled with an evolving structured working chain; and five stage-wise metrics (Ret., Grp., Ord., Grd., Gap) plus a progressive E1–E4/OK failure funnel. Six LLM configurations are evaluated. The headline result is that even the strongest configuration completes only 39.6% of the 849 reference steps, with smaller models predominantly failing at E2 (observed but unused evidence) and larger models at E4 (ordering). Ablations (RQ3) and budget sweeps (RQ4) support the claimed shift, rather than elimination, of failures as evidence access improves. The paper also reports a leakage audit with zero findings, an artifact-consistency audit, a 12-case blind manual validation of evaluator outputs (60/60 agreement), and a C0/C1 evidence-dependence control.","tokens_in":40693,"tokens_out":9133,"duration_ms":84568,"significance":"If the central results hold, DiagChain would be a valuable contribution to cybersecurity LLM evaluation. The stage-wise diagnostic framing moves beyond end-to-end accuracy and yields an actionable, if sobering, characterization of where current agents fail. The paper is unusually careful in several respects: it ships machine-checkable reproducibility artifacts (SHA-256 manifests, frozen run records, and verification scripts), reports a leakage audit with zero findings, reports an artifact-consistency audit with no mismatches, provides 60/60 reviewer agreement on a purposive 12-case evaluator-output sample, and includes a C0/C1 control that demonstrates dependence on supplied evidence. These strengths are genuinely commendable. However, the benchmark's central quantitative claims—the 39.6% headline, the E1–E4 proportions, and the size-ordered E2-to-E4 bottleneck shift—are all computed against gold reference packages whose validity rests on a single-author audit. That dependence tempers the significance until the reference standard is independently validated.","major_comments":[{"comment":"The gold reference packages are the sole ground truth for every DiagChain metric (Ret., Grp., Ord., Grd., Gap), for the RQ2 E1–E4/OK funnel (Table 19), and for the headline '39.6% of the 849 reference steps.' Appendix B states that the validation was 'a full author audit rather than independent double annotation, so inter-annotator agreement is not reported,' and Table 8 reports zero corrections across 69 packages, 849 steps, and 780 edges. Because 39 of the 69 cases derive from AutoLabel, a dataset produced by the same research group, the independence of the gold standard is limited. If the reference step segmentation, support-to-evidence assignment, or reference order contain systematic bias, every stage-wise score and the central E2-versus-E4 bottleneck shift could be an artifact of the gold labels rather than a fact about model behavior. The paper's transparency is commendable, but transparency does not remove the empirical dependence. A concrete remedy within scope would be a second, independent annotation on a purposive sample spanning all three sources and chain lengths, with reported inter-annotator agreement on step boundaries, support IDs, and ordering; additionally, a sensitivity analysis that merges or splits adjacent reference steps would show whether the model-level E2/E4 ordering is stable under plausible reference-granularity perturbations.","section":"Appendix B, 'Reference Construction and Verification', Table 8"},{"comment":"The common operating point k=32, T=15 was selected as a quality–cost operating point from the RQ4 sweeps on R24, which is a subset of the same benchmark (MAIN-69) on which RQ1/RQ2 are then reported. This means the RQ1/RQ2 results are evaluated at a setting chosen by looking at the test distribution itself, and no alternative operating point is reported for the full six-model panel. The authors correctly label the setting 'an empirical default rather than a universal optimum,' but the 39.6% headline and the cross-model E2/E4 comparison could shift under a different k/T choice. To establish that the bottleneck shift is robust rather than an artifact of the selected operating point, the authors should re-run the RQ2 funnel for all six models at a second, independently justified operating point (for example k=16 with T=10, or k=48 with T=20) and confirm that the relative E2/E4 ordering across model sizes persists.","section":"Experiments, 'Setup' and 'Budget Sensitivity (RQ4)'"},{"comment":"All model-comparison conclusions rest on a single run per model–case condition, with temperature 0 for the Ollama, DeepSeek, and GLM backends but provider-default sampling for GPT-5.5. The paper reports no bootstrap confidence intervals or repeated-run variance for the 39.6% headline, the E1–E4 proportions in Table 19, or the cross-model differences in Figure 3. Because the central claim is a cross-model difference in failure stages, and because steps within a case are likely correlated, the absence of uncertainty quantification makes it hard to assess how much of the large E2-versus-E4 gaps (e.g., Qwen-3-32b E2=39.2% vs. E4=13.4%; GPT-5.5 E2=5.8% vs. E4=30.6%) reflects a stable population-level bottleneck rather than run-to-run variation. The authors should provide case-bootstrap confidence intervals for the funnel proportions (or repeated runs on a representative subset of model–case pairs) to support the generalization from this single-run panel.","section":"Experiments, RQ1/RQ2 and Table 19"}],"minor_comments":[{"comment":"The text contains 'howerrorsariseandpropagateacrossintermediatereasoning stages' with missing spaces; this formatting/typo should be corrected.","section":"Abstract"},{"comment":"The B3 clustering F1 definition would benefit from an explicit statement of how evidence items that are not in the reference support set (for example, background cards in the evidence universe) are treated in the per-item B3 precision and recall, since the phrase 'the same evidence universe' is currently ambiguous.","section":"Section 'Diagnostic Evaluator', Grp. definition"},{"comment":"Because Table 8 reports 'Corrected 0' across all categories, consider adding a note clarifying that the audit was a confirmation pass against source records rather than a re-annotation from scratch; this would make the sense of 'no correction was required' more precise.","section":"Appendix B, Table 8"},{"comment":"The faint/outlined markers for individual case runs are difficult to distinguish in grayscale; increasing marker contrast or using distinct symbols would improve readability.","section":"Figure 2(c)"},{"comment":"The C1–C0 difference intervals are described as 'case-bootstrap,' but with n=15 cases the interpretation depends on whether resampling is over cases or steps; the authors should state explicitly that cases are the resampling unit.","section":"Appendix F, Table 18"},{"comment":"The main text presents Ord. as a neutral 'ordering accuracy,' while Appendix B clarifies that reference edges and ordering are temporal, not fully verified causal claims; the main text should state this temporal, non-causal interpretation up front to avoid overclaiming.","section":"Main text, ordering metric"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and unusually careful about reproducibility, but the gold-standard validity is the central risk. The reported effect sizes for the E2/E4 bottleneck shift are large, so the finding may well be robust; however, the absence of any inter-annotator agreement on the reference packages, combined with the overlap between the AutoLabel source and the same research group, makes the current evidence insufficient for acceptance. A second-annotator study on a representative sample, plus a sensitivity check at an alternative operating point, would address the two principal load-bearing concerns. If those are provided, I would support acceptance; without them, the headline quantitative claims remain conditional on unvalidated ground truth."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a real benchmark paper, not a benchmark stunt: MAIN-69, the five stage-wise metrics, and the finding that failure bottlenecks shift from evidence incorporation to ordering as models get bigger are all genuinely useful. Second, the whole edifice rests on gold reference packages that one author audited alone; that is the soft spot to probe.\n\nWhat is actually new: the cross-source suite (AutoLabel, ExCyTIn, OTRF APT29) with three controlled noise profiles; metrics that separate retrieval coverage, grouping, ordering, grounding, and attribution; and an ECRAG retrieval workflow that couples evidence retrieval with an evolving structured chain. The empirical result is striking: even the best configuration completes only 39.6% of reference steps, and the E2-to-E4 bottleneck shift with scale is the kind of actionable finding the security-agent community needs. The paper is also unusually careful for this area: SHA-256 manifests, a leakage audit with zero findings, artifact-consistency checks, a 60/60 manual validation of evaluator outputs, and a C0/C1 control showing models depend on supplied evidence rather than recalling public incidents. Those are real marks of discipline.\n\nThe soft spots are mostly disclosed, but they are real. The gold labels were checked by one author with no inter-annotator agreement, and 39 of 69 cases derive from the same group's AutoLabel. If step boundaries or reference order are subjective, every stage-wise rate inherits that bias, and the E2-versus-E4 story could shift. I don't think it collapses: the pattern is large and consistent across six models, and the observed-vs-cited recall analysis in Figure 2(c) supports the same conclusion without using the funnel. But the exact percentages should be treated as provisional. Elsewhere: RQ1 is single-run with no error bars; RQ3 and RQ4 use purposive subsets; the headline operating point was picked from the R24 sweep; and RQ4 varies budgets for one model only. None of these makes the paper wrong; they all argue for conditional, not unconditional, acceptance.\n\nWho this is for: anyone building or evaluating LLM agents for security investigation. It deserves a serious referee. My recommendation: send it out, and ask the authors for independent double annotation on a stratified subset, variance estimates or multi-run results on RQ1, and a clearer statement of the step-unit rules across sources.","headline":"A worthwhile diagnostic benchmark with a real bottleneck-shift finding; the single-author gold audit is the main thing to probe before trusting the exact numbers.","tokens_in":41343,"tokens_out":3447,"would_cite":true,"duration_ms":28987,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even the strongest LLM agent completes only 39.6% of 849 reference steps of evidence-grounded attack chains without an earlier failure.","keywords":["attack chain reconstruction","LLM agents","diagnostic benchmark","evidence grounding","retrieval-augmented generation","cyber threat investigation","attribution gap","stage-wise evaluation"],"falsifier":"Independently re-annotate a random sample of MAIN-69's reference packages with a second annotator blind to the authors' labels, and measure agreement on step boundaries, supporting evidence IDs, and edge ordering; if disagreement changes the E1–E4 first-failure assignment of more than a few percent of steps, the reported 39.6% figure and the bottleneck shift are not robust.","tokens_in":40234,"feed_emoji":"🛡️","tokens_out":4983,"duration_ms":40814,"temperature":0.7,"pith_summary":"DiagChain is a benchmark that tests whether LLM agents can reconstruct an ordered cyberattack chain from retrieved logs, alerts, and host events. Its central claim is that today's agents, even with retrieval support, complete only 39.6% of the 849 reference steps without an earlier failure. The benchmark's diagnostic value is a stage-wise failure funnel that separates evidence discovery, grouping, ordering, grounding, and attribution, and it shows that smaller models lose evidence they already observed, while larger models fail at ordering it. These findings argue that end-to-end accuracy hides where reconstruction breaks and that better retrieval or larger budgets only relocate failures.","feed_headline":"Even the strongest LLM agent completes 39.6% of attack-chain steps","feed_subtitle":"Stage-wise benchmark shows small models lose observed evidence; large models misorder it.","key_machinery":"The key machinery is a stage-wise evaluator built on a progressive failure funnel that assigns each coverable reference step to its first failure stage: E1 evidence not observed, E2 observed but unused, E3 partial attribution, E4 misordering, or OK. It is driven by five metrics—retrieval step coverage, $B^3$ grouping F1, pairwise ordering accuracy, evidence grounding F1, and attribution gap rate—and an agent workflow, Evidence-Centric Retrieval-Augmented Generation (ECRAG), which couples TF-IDF retrieval with entity and temporal expansion around seed evidence while maintaining an evolving structured working chain. This combination lets the authors localize where reconstruction breaks instead of reporting a single aggregate score.","core_discovery":"The central discovery is a size-ordered bottleneck shift in evidence-grounded attack chain reconstruction: smaller models (Qwen-3-32b, Llama4-17b-Scout) are dominated by steps where supporting evidence was observed but never cited, whereas larger or reasoning-enhanced models (DeepSeek-V4-Pro, GLM-5.2 variants, GPT-5.5) proceed further and fail primarily at ordering the evidence they have acquired. Across 69 scenarios with three noise profiles, the strongest configuration (GPT-5.5) reconstructs only 39.6% of reference steps without an earlier failure. The paper also shows that raw evidence impairs discovery, long chains expose assembly limits, and larger retrieval budgets or turn ceilings expand evidence exposure without consistently improving final chain quality.","pith_inferences":["The same E1–E4 funnel could diagnose other evidence-to-structure tasks, such as paper-grounded scientific reasoning or document-grounded timeline reconstruction, where discovery and ordering are separable.","If gold-label error is small, the bottleneck shift suggests different interventions: retrieval re-ranking or memory compression for small models, and chain-planning or ordering constraints for large models.","The near-miss trace (one local permutation causing failure with perfect retrieval, grouping, and grounding) implies that even strong agents could improve with ordering-specific reflection, a testable extension.","The closed-book control suggests models do not silently rely on memorized public incidents, but an adversarial probe with a re-named variant of a public APT scenario would directly test contamination."],"forward_implications":["Stage-wise evaluation shows that end-to-end accuracy masks where reconstruction fails; benchmarks should report evidence discovery and ordering separately.","Improving retrieval scope or interaction budgets alone will not fix chain assembly; failures relocate downstream as evidence access improves.","Smaller models need better mechanisms for retaining and citing observed evidence, while larger models need better global ordering and planning.","Raw evidence disproportionately impairs evidence discovery, while long chains expose grouping and ordering limits.","The benchmark and ECRAG scaffold provide an auditable testbed for evidence-grounded security agents."],"supporting_citations":[{"why":"Provides the retrieval-augmented generation paradigm that ECRAG adapts to evidence cards.","marker":"(Lewis et al. 2020)"},{"why":"SLEUTH motivates DiagChain's evidence-linked chain model from audit-event provenance.","marker":"(Hossain et al. 2017)"},{"why":"ExCyTIn-Bench incidents and alert-graph sources feed 24 of the MAIN-69 scenarios.","marker":"(Wu et al. 2026)"},{"why":"AutoLabel scenarios and their rule/AI-assisted labels supply 39 cases and step candidates.","marker":"(Peng et al. 2025)"},{"why":"OTRF APT29 host-event datasets supply six long-chain Windows cases.","marker":"(Rodriguez 2020)"},{"why":"Supplies SynthChain's attack-chain formulation adopted for reference steps and edges.","marker":"(Tan et al. 2026a)"},{"why":"Supplies the B^3 clustering F1 used for the grouping metric.","marker":"(Bagga and Baldwin 1998)"},{"why":"Supplies the TF-IDF ranking that seeds ECRAG's retrieval and expansion.","marker":"(Salton and Buckley 1988)"}],"fun_headline_variants":["LLM agents complete only 39.6% of attack-chain steps","Small models drop evidence, large models misorder it","DiagChain pinpoints why LLM agents fail at attack chains","Attack-chain benchmark: best LLM scores 39.6%","Evidence handling shifts from citing to ordering in LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gold reference packages—849 steps and 780 edges—were individually checked by a single author rather than independently double-annotated, so errors in the source labels or in that audit propagate into every model score and every failure-stage label.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents complete only 39.6% of attack-chain steps","Small models drop evidence, large models misorder it","DiagChain pinpoints why LLM agents fail at attack chains","Attack-chain benchmark: best LLM scores 39.6%","Evidence handling shifts from citing to ordering in LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1276,"prompt_tokens":935,"completion_tokens":341,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":551,"tokens_out":341,"duration_ms":4081,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:48:31.827642+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-annotate a random sample of MAIN-69's reference packages with a second annotator blind to the authors' labels, and measure agreement on step boundaries, supporting evidence IDs, and edge ordering; if disagreement changes the E1–E4 first-failure assignment of more than a few percent of steps, the reported 39.6% figure and the bottleneck shift are not robust.","supporting_citations":[],"review_version":1}