{"id":"1774b159-186a-4e1a-89bd-5a04ea8b1297","arxiv_id":"2608.09567","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A pre-registered audit of 102 cases plus 18 papers found artifact-embedded oracles unreliable: 20 of 30 signals also triggered on patched builds, so runnable artifacts rarely confirm the CVE.","lead":"This paper audited 104 papers on LLM/agent-driven vulnerability validation and executed artifacts from 18 papers plus all 102 cases of one benchmark in isolated Docker containers. It found that most proof-of-concept triggers still fire on patched builds, so a runnable artifact is not proof that the CVE was actually reproduced.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"G3b patched-counterfactual verdicts are invalid for apt-level EOL cases: Section 5.4 admits 'apt-level cases whose EOL sources carry no fix' are among the 20 dirty, so the 'patched' build may still be vulnerable, inflating the oracle-unreliability rates and the 60%/45% matrix.","rationale":"The reader's weakest assumption—that G3b patched-counterfactual verdicts may be invalid for EOL apt-level cases—is directly confirmed by the paper's own text in Section 5.4. This is not a speculative worry but an explicit inclusion of unpatched builds in the 'dirty' set. The confusion matrix and the 20/30 statistic therefore cannot be taken at face value, and the paper's quantitative headline on oracle unreliability is overstated. However, the central qualitative claim does not rest solely on the rates: the paper documents concrete cases (CVE-2020-1967, CVE-2021-31162, CVE-2022-22816, CVE-2023-0217) where a signal fires on both vulnerable and genuinely patched builds, and the protocol itself (G1–G3b ladder, pre-registration, evidence bundles) is a genuine contribution with independent value. The correct remedy is a revision that reclassifies the EOL apt-level cases as 'unbuildable' or 'not audited' and recomputes the matrix, not a rejection of the paper's core message. Since the reader already reached CONDITIONAL, my assessment does not move the verdict; it sharpens the specific condition that must be met: validate the patched builds before trusting the matrix.","tokens_in":11938,"tokens_out":5570,"duration_ms":47055,"concrete_test":"Re-audit all 20 dirty G3b verdicts by verifying the patched build actually contains the fix before scoring. For apt-level EOL cases, rebuild using a maintained mirror or a backported patch (or a later base image) and confirm the vulnerable version is gone (e.g., check `dpkg -l` for the fixed version, or `grep` the vulnerable symbol). For image-tag and pin replacements, verify the runtime version inside the container (`openssl version`, `python --version`, `cargo --version`) actually changed to the patched release. Recompute TP/FP/FN/TN with corrected labels: if any previously dirty case becomes clean, the 20/30 rate and the 60%/45% matrix change; report the corrected numbers and their Wilson CIs. A smaller concrete variant: list which of the 30 verdicts used EOL apt sources and recalculate the matrix excluding them.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative core of the paper—the oracle confusion matrix in §5.6 (sensitivity 60%, specificity 45%; 20/30 G3b dirty)—depends on the G3b patched-counterfactual verdicts being valid ground truth. Section 5.4 states that the 20 dirty verdicts 'include ... apt-level cases whose EOL sources carry no fix.' For those cases, the patched build was constructed by apt upgrade injection, but because the base image is EOL, no fix is available; the build is therefore not actually patched. Counting a signal that persists on an unpatched build as 'dirty' mislabels the oracle as producing a false positive when it may be correctly detecting the still-present vulnerability. The same section also includes 2 Go cases 'whose base images already contain the fix yet the signal persists,' which are legitimate dirty verdicts, but the EOL apt-level cases are not. If even a few of the 20 dirty verdicts belong to the not-actually-patched category, the headline rates overstate oracle unreliability. The paper's own limitation section (8.2) acknowledges the EOL problem only as a construction failure, yet the affected cases are retained in the denominator. The central qualitative conclusion (trigger on vulnerable build is not sufficient; specific counterexamples like CVE-2020-1967) is independently supported and would survive, but the quantitative matrix as reported is not a valid measurement of oracle reliability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pre-registered reproducibility audit of LLM/agent-driven vulnerability validation artifacts. It builds a 104-paper consensus corpus, finds 59 reachable artifacts, executes 18 paper-level artifacts at R0/R1, executes all 102 cases of the anchor benchmark (arXiv:2509.24037), and audits 30 signal-producing cases with patched-counterfactual (G3b) builds and 19 with matched negative controls (G3a). The three headline findings are: 58/102 anchor cases have a script-internal CVE identifier that diverges from the declared directory CVE; 10/18 (55.6%) artifacts complete their workflow at R0, rising to 11/18 after R1; and artifact-embedded oracles appear unreliable, with 20/30 patched-counterfactual verdicts dirty and an oracle confusion matrix of sensitivity 60% and specificity 45%. The paper concludes that a trigger on the vulnerable build is not evidence of CVE-specific reproduction unless the patched counterfactual is clean.","tokens_in":12276,"tokens_out":5116,"duration_ms":44905,"significance":"If the quantitative measurement were fully valid, this would be a valuable contribution: it operationalizes semantic confirmation for security reproducibility, provides a reusable G1-G3 evidence ladder, and documents concrete failure modes of artifact-embedded oracles. The qualitative ladder claim is convincingly supported by the CVE-2020-1967 case in Section 5.4, where a clean segmentation fault occurs on both the vulnerable and the patched OpenSSL builds, showing that signal production does not imply CVE-specific reproduction. The frozen protocol, evidence bundles, per-case execution logs, and the explicit execution ledger in Table 2 are real strengths that make the audit auditable. However, the headline oracle-reliability rates currently rest on a patched-counterfactual ground truth that is not valid for a subset of the included cases, so the quantitative findings need revision before the paper's central rates can be accepted as stated.","major_comments":[{"comment":"The G3b patched-counterfactual verdicts are not valid ground truth for the EOL apt-level cases, yet these cases are retained in the 20/30 dirty rate and in Table 5. Section 5.4 explicitly lists among the 20 dirty verdicts 'apt-level cases whose EOL sources carry no fix.' For those cases, the patched build was constructed by apt upgrade injection, but since the base image is EOL, no fix is actually available; the resulting 'patched' build may still be vulnerable. A signal that persists on such a build is therefore not evidence that the oracle is non-specific. Section 8.2 itself states that 'EOL base images carry no apt-level fixes, so patched-counterfactual construction failed,' yet those same cases are counted as dirty rather than as unbuildable or unassessable. The authors should either exclude these cases from the dirty set, classify them as unbuildable, or demonstrate for each retained EOL apt-level case that a genuine patched build was available and applied; the 20/30 rate, the specificity estimate of 45%, and the confusion matrix must then be recomputed.","section":null},{"comment":"Table 5's confusion matrix pools non-random calibration cases with full-corpus cases, yet reports Wilson 95% confidence intervals as though the 30 verdicts form a well-defined random sample. Section 5.4 states that the calibration G3b subset 'is not a random sample, so the 4/5 rate below is not an estimate of a population proportion.' The G3b-executed set comprises 29 of the 34 full-corpus G1 signals plus four additional calibration-only cases (Section 5.4), so the n=30 matrix is a convenience set rather than a probability sample. The confidence intervals for sensitivity and specificity are therefore not sampling-based estimates for a defined population. The authors should present the full-corpus and calibration-only verdicts separately, or clearly label the matrix as a descriptive summary of the audited cases and omit inferential intervals.","section":null},{"comment":"The abstract and Section 5.6 state that artifact-embedded oracles 'prove unreliable' and headline the 60%/45% matrix, but Section 8.2 correctly cautions that these are single-corpus, partial-coverage results that 'should not be read as a population estimate until the multi-paper sample replicates it.' This internal tension is load-bearing because the G3b coverage is incomplete (5 of 34 G1 signals lack a patched counterfactual), the E1 confirmation set contains only two cases, and the G3b ground-truth issues described above directly affect the matrix. The authors should align the abstract and conclusion with the stated exploratory scope, or provide a formal justification for why the audited 30 cases support a systematic-reliability claim despite the non-random sampling and the EOL construction failures.","section":null}],"minor_comments":[{"comment":"The bullet 'F ailure taxonomy' contains a spacing typo and should read 'Failure taxonomy.'","section":null},{"comment":"The inconsistent spacing in 'PASS/F AIL' and 'PASS/F AIL' should be normalized to 'PASS/FAIL' throughout the manuscript and tables.","section":null},{"comment":"The paper describes the protocol as pre-registered but provides only self-hosted paths such as 'protocol/search protocol v1.json' and no external registry, DOI, or third-party timestamp. Please provide a public, immutable deposit or registry entry so readers can independently verify the pre-registration claim.","section":null},{"comment":"The exploratory claim gap on the calibration cases (claimed success 0.80 vs. E1 rate 0.20, computed on five G3b-verdict cases) should be explicitly labeled as a pilot illustration rather than an estimate, especially since the preceding paragraph states that the four-fifths calibration dirty rate is not a population proportion.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's central qualitative claim lands: a trigger signal on a vulnerable build is not evidence of CVE-specific reproduction without a clean patched counterfactual. The CVE-2020-1967 example alone justifies that sentence—the PoC calls SSL_check_chain(NULL), which crashes any OpenSSL version, and the artifact's own comment says 'simulate vulnerability.' That is a clean, reproducible counterexample to trigger-string oracles. The four-level availability/runnable/signal/confirmed ladder, the R0/R1 repair ladder, and the G1–G3 evidence levels are genuinely new relative to prior reproducibility work, which stops at build-and-run. The 58/102 script-internal CVE divergence is also a fresh, concrete failure mode. The paper ships frozen protocols, claim sheets, and execution logs; this is how reproducibility audits should be built.\n\nNow the soft spots, in proportion. The quantitative headline—20/30 dirty patched counterfactuals, oracle sensitivity 60% / specificity 45%—is weaker than it looks. Section 5.4 explicitly states that the dirty verdicts include apt-level cases whose EOL base images carry no fix. For those cases, the 'patched' build was constructed by apt upgrade injection, but no fix exists, so the build is still vulnerable. Counting a persistent signal on a still-vulnerable build as 'dirty' mislabels the oracle as falsely positive when it may be correctly detecting the real vulnerability. The paper's own limitation section acknowledges the EOL problem only as a construction failure, yet the affected cases stay in the denominator. That is a measurement error in the instrument, not a mere caveat. Second, the n=30 matrix pools 4 non-random calibration cases with the full-corpus G1 cases, then calls itself the 'full-corpus matrix.' That is a denominator integrity problem. Third, the pre-registration is self-hosted and the evidence-bundle URLs are not independently verifiable; minor, but worth noting.\n\nNone of this breaks the qualitative takeaway. The paper is honest about exploratory status, and the specific counterexamples stand on their own. The fix is straightforward: reclassify or exclude the not-actually-patched EOL cases, report the matrix separately for calibration and full-corpus, and say which of the 20 dirty verdicts survive that reclassification. The central argument about patched-counterfactual checks does not depend on the exact 20/30 rate.\n\nThis paper deserves a serious referee. It gives the security reproducibility community a reusable protocol and a concrete, correct lesson. Send it to review with a request to clean up the G3b instrument before publication.\n\nReading group: yes. I'd cite the ladder and the CVE-divergence finding, not the confusion matrix as-is.","headline":"The qualitative core is right and useful, but the headline oracle-reliability numbers overstate the case because some 'patched' builds are not actually patched and calibration cases are mixed into the full-corpus matrix.","tokens_in":762,"tokens_out":2346,"would_cite":true,"duration_ms":37689,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a trigger signal on a vulnerable build is not evidence of CVE-specific reproduction unless a clean patched counterfactual is present, and that embedded oracles are unreliable (sensitivity 60%, specificity 45%).","keywords":["reproducibility audit","vulnerability validation artifacts","LLM agents","patched counterfactual","oracle reliability","CVE reproduction","artifact evaluation","security benchmarks"],"falsifier":"Run the 30 G3b cases again with patched builds whose fix status is independently confirmed (for example, compile from the official fixed release tags and verify the vulnerable code path is gone via diff), then recompute the dirty rate and the oracle confusion matrix; if the dirty rate drops well below 20/30, the oracle-unreliability conclusion weakens.","tokens_in":11718,"feed_emoji":"🧪","tokens_out":5531,"duration_ms":42003,"temperature":0.7,"pith_summary":"The paper tries to establish that a security artifact that publicly exists, runs, and prints an alarming signal is still a long way from having reproduced the claimed CVE. It argues that artifact-embedded oracles—scripts that declare \"VULNERABILITY TRIGGERED\"—are systematically unreliable, because they rarely check the one thing that would make a signal CVE-specific: a clean patched counterfactual. On a 102-case benchmark plus an 18-paper sample, 20 of 30 signal-producing cases still fired on the patched build, 7 of 19 matched negative controls fired on benign input, and the oracle confusion matrix showed 60% sensitivity and 45% specificity. A sympathetic reader should care because claimed per-case success rates in this literature are widely used to compare agents and to certify vulnerability-reproduction skill, and this audit says those claims can overstate confirmation by a wide margin.","feed_headline":"Trigger signals fail patched-counterfactual tests 20 of 30 times","feed_subtitle":"An audit of 102 CVE cases finds artifact oracles at 60% sensitivity and 45% specificity—trigger strings alone overstate confirmation.","key_machinery":"The machinery is a four-layer semantic-confirmation ladder: candidates at the \"available\" level, \"runnable\" at R0 or after environment-only R1 repair, \"signal-producing\" when a run yields a candidate marker (G1), and \"semantically confirmed\" when the signal matches a CVE-specific post-condition (G2), is absent on benign input (G3a negative control), and is absent on the patched build (G3b patched counterfactual). The patched-counterfactual oracle is the load-bearing instrument: it turns a yes/no \"did it print something scary\" check into a test of whether the signal is actually CVE-specific, and it is what generates the paper's confusion matrix.","core_discovery":"The central discovery is that availability, runnability, signal production, and semantic confirmation form a strict ladder, and the gap between layers is large. On the anchor corpus, the repository was reachable, the workflow completed at R0 for only 10 of 18 papers (11 of 18 after environment-only R1 repair), 34 of 87 executed runs (39.1%) produced a candidate signal, 23 of those 34 satisfied their pre-registered post-condition, and only 10 of 30 audited cases survived the patched-counterfactual check. A trigger on the vulnerable build is therefore not evidence of CVE-specific reproduction unless the patched counterfactual is clean. The paper also found that 58 of 102 anchor cases contain a script-internal CVE identifier that differs from the directory's declared CVE, meaning the target being tested can differ from the target claimed.","pith_inferences":["Editorial inference: the same patched-counterfactual discipline likely applies to traditional (non-LLM) PoC exploits; a cheaper spot-check of existing exploit databases against patched builds would test this directly.","Editorial inference: the 56.9% CVE-ID divergence rate, if it generalizes beyond the anchor corpus, means any benchmark evaluated by matching directory labels to trigger strings inherits label noise that could silently inflate agent scores.","Editorial inference: oracle false positives concentrated in manually forced crashes (NULL dereferences, designed-in panics) suggest a concrete fix—require PoCs to exercise the vulnerable code path rather than any crash.","Editorial inference: the study's calibration E1 rate bounds (1/20 to 5/20) could be sharpened by completing G3a/G3b audits on the remaining cases; the pre-registered bounds policy already gives readers worst/best cases rather than single imputed numbers."],"forward_implications":["Artifact-embedded trigger checks should be treated as screening tools, not confirmation, until a patched-counterfactual check is included.","Published per-case success rates for LLM/agent vulnerability validation are upper bounds; the true semantically-confirmed rate on the anchor corpus is far lower, with only two strict E1 confirmations out of the signal-producing corpus.","Benchmark maintainers need to reconcile script-internal CVE identifiers with declared targets before using case outcomes for agent evaluation.","Reproducibility audits of security papers should add a semantic-confirmation layer rather than stopping at buildability.","The pre-registered protocol (R0/R1 ladder, G1–G3 evidence levels, patched-counterfactual oracle) offers a reusable template for future audits."],"supporting_citations":[{"why":"Anchor benchmark (arXiv:2509.24037) whose 102 case directories form the execution corpus for all oracle and counterfactual audits.","marker":"[6]"},{"why":"NVD record for CVE-2020-1967 used to define the post-condition and to demonstrate the NULL-crash false positive on the patched OpenSSL build.","marker":"[8]"},{"why":"NVD record for CVE-2021-31162 used to show a designed-in panic being mistaken for the vulnerability signal in an artifact oracle.","marker":"[9]"},{"why":"NVD record for CVE-2021-44228 used to show a genuine JNDI RCE signal missed because the oracle required an absent marker file.","marker":"[10]"},{"why":"NVD record for CVE-2025-30223, the one calibration case that passed G3b and reached strict E1 confirmation.","marker":"[14]"},{"why":"Previous tier-1 security reproducibility study that measured buildability but stopped short of semantic confirmation, defining the gap this audit fills.","marker":"[16]"},{"why":"PVBench result that patched-validated patches can fail rigorous testing, cited as the patch-side analogue of the oracle-reliability gap.","marker":"[21]"},{"why":"CVE-Bench supplies the cross-paper feasibility check and is part of the case-level sample pool for the planned multi-paper confirmatory phase.","marker":"[23]"}],"fun_headline_variants":["CVE reproducibility audit: only 10/18 workflows finish at R0","Artifact oracles: 60% sensitivity, 45% specificity in CVE validation","Script-internal CVE mismatches: 58/102 anchor cases diverge from directory","Trigger alone fails as CVE evidence: 66% patched builds still signal","Independent audit: 4-layer gap from runnable to verified CVE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The patched-counterfactual verdicts are treated as ground truth for whether a signal is CVE-specific, but some patched builds are constructed by swapping image tags or pinned versions and may still contain the vulnerability, so a dirty verdict can mean \"not actually patched\" rather than \"oracle unreliable.\"","fun_headline_variants_meta":{"raw":{"variants":["CVE reproducibility audit: only 10/18 workflows finish at R0","Artifact oracles: 60% sensitivity, 45% specificity in CVE validation","Script-internal CVE mismatches: 58/102 anchor cases diverge from directory","Trigger alone fails as CVE evidence: 66% patched builds still signal","Independent audit: 4-layer gap from runnable to verified CVE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000916,"raw_usage":{"total_tokens":4014,"prompt_tokens":1109,"completion_tokens":2905,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":725,"completion_tokens_details":{"reasoning_tokens":2797}},"tokens_in":725,"tokens_out":2905,"duration_ms":17957,"temperature":1.0,"reasoning_tokens":2797,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:54:50.709922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the 30 G3b cases again with patched builds whose fix status is independently confirmed (for example, compile from the official fixed release tags and verify the vulnerable code path is gone via diff), then recompute the dirty rate and the oracle confusion matrix; if the dirty rate drops well below 20/30, the oracle-unreliability conclusion weakens.","supporting_citations":[{"cited_title":"Cve-2020-1967","cited_arxiv_id":null,"evidence_quote":"NVD record for CVE-2020-1967 used to define the post-condition and to demonstrate the NULL-crash false positive on the patched OpenSSL build."},{"cited_title":"Cve-2021-31162","cited_arxiv_id":null,"evidence_quote":"NVD record for CVE-2021-31162 used to show a designed-in panic being mistaken for the vulnerability signal in an artifact oracle."},{"cited_title":"Cve-2021-44228","cited_arxiv_id":null,"evidence_quote":"NVD record for CVE-2021-44228 used to show a genuine JNDI RCE signal missed because the oracle required an absent marker file."},{"cited_title":"Cve-2025-30223","cited_arxiv_id":null,"evidence_quote":"NVD record for CVE-2025-30223, the one calibration case that passed G3b and reached strict E1 confirmation."}],"review_version":1}