{"id":"dae6071c-56b9-41c2-ac1b-68e15bd69c4a","arxiv_id":"2608.03110","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"PECR is a transparent nine-factor vulnerability prioritization protocol for SD-WAN, with synthetic stress tests showing it ranks differently than CVSS/EPSS/KEV queues, but no real outcome validation.","lead":"A new decision method, PECR, ranks software vulnerabilities on SD-WAN networks by combining nine factors such as severity, current exploit evidence, reachability, and data quality. The paper tests the ranking only on synthetic records, so it demonstrates the method's behavior and limits, not its real-world effectiveness.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated graph factors: F6/F7 are sampled from role/zone profiles (Sec. 8.1), not computed by the Sec. 4.2 reachability procedure, so the 'full factor set refines order' result is conditional on an untested shortcut.","rationale":"The paper is explicit and honest about its limits, and the strongest claim is appropriately modest. The single most load-bearing condition is that the two network-derived factors actually represent what PECR would compute. Since they are sampled from hand-declared profiles rather than produced by the Section 4.2 algorithm, the synthetic validation does not exercise the part of the specification that distinguishes PECR from prior context-aware methods. If the graph algorithm yields similar distributions, the concern is resolved; if not, the reported τ values and top-10 overlap are artifacts of the generator's profiling choice rather than of PECR itself. I agree with the reader's weakest_assumption, and the reader's CONDITIONAL verdict remains appropriate. The missing archival identifier is a reproducibility barrier, but the deeper scientific issue is the unexecuted reachability algorithm. Therefore no verdict change is needed; the condition for acceptance should explicitly include releasing the artifact and validating the graph-derived factors against the specified procedure.","tokens_in":11946,"tokens_out":5088,"duration_ms":50475,"concrete_test":"Implement the Section 4.2 BFS reachability procedure on a directed graph consistent with the artifact's 62 assets and declared role/zone edge policies; compute F6 and F7 for the 100 primary records; rerun Table 6 and Table 8 with those values, keeping every other factor, weight, and tie-break unchanged. If the context-lite Kendall τ or top-10 overlap moves by more than ±0.05 or ±1 record, the sampled F6/F7 shortcut is load-bearing and the 'refines order' claim is unsubstantiated for the graph-aware factors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 8.1 states that 'F6 and F7 are sampled from declared role and zone profiles; it does not validate the graph-construction algorithm itself.' Yet F6 (path reachability) and F7 (blast radius) are the only factors implementing the paper's network-aware contribution (Section 4.2). The central claim, that the full factor set refines the context-lite order, is supported by the difference between context-lite (τ=0.755) and PECR in Table 6. That difference is generated by the relative ordering effects of F5–F7 and F9. If a real execution of Section 4.2 produced F6/F7 values that are nearly constant, dominated by F4, or differently correlated with the latent threat z, the refinement margin and the differentiation from simpler queues could shrink or grow arbitrarily. The paper itself flags this in Section 10 ('it still samples reachability and blast radius instead of executing the graph procedure'), but the flag does not reduce the dependence: the central behavioral claims about ordering are not yet grounded in the algorithm that defines PECR. Without a check that computed F6/F7 yield similar Table 6 and Table 8 results, the specification remains executable only for the non-graph factors.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper specifies PECR, a decision-support method for SD-WAN vulnerability prioritization that combines nine normalized factors (severity, exploit evidence, EPSS probability, exposure, privilege, path reachability, blast radius, consequence, and control concentration), reports confidence and interval bounds separately, and maintains a distinct cryptographic-migration (C-BOM) queue. The evaluation is entirely synthetic: 100 role- and zone-consistent records over 62 assets from a disclosed generator, 30 seeds of replication, five baseline queues, 30,000 joint Dirichlet weight draws for sensitivity, and one 10% missing-cell masking test. The headline results are that a simple context-lite five-factor subset captures much of PECR's top-10 behavior (Kendall tau = 0.755), the full factor set refines order further, and the ranking is reasonably stable under weight perturbation (median tau 0.836-0.924 across broad, moderate, and narrow regimes). The paper explicitly disclaims predictive accuracy and production effectiveness. However, F6 and F7 are sampled from role/zone profiles rather than computed by the Section 4.2 graph procedure, and the promised archival artifact is not yet available.","tokens_in":12199,"tokens_out":5390,"duration_ms":52749,"significance":"The contribution is a detailed, auditable decision specification plus an honest synthetic stress-test methodology. Strengths include explicit equations, deterministic queue semantics, separation of score/confidence/bounds/migration, broad joint-weight sensitivity analysis, transparency about limitations, and a clearly stated absence of predictive-performance claims. If accompanied by an accessible artifact and a check that the graph-construction procedure yields similar ordering behavior, the paper would be a useful reference point for telemetry-informed prioritization. As it stands, the headline numerical results are plausible but the advertised 'executable specification' is not fully realized, because the two graph-derived factors are not exercised by the validation and the artifact cannot currently be checked.","major_comments":[{"comment":"The reproducibility claim is not currently verifiable: Section 12 states that 'A permanent archival identifier will be added in a subsequent version' and provides no artifact URL, DOI, or checksum. Because the paper's central contribution is an executable, auditable specification, this is a load-bearing issue rather than a cosmetic one. Please provide a stable archival link to the supplementary artifact (including the generator, seeds, CSV, and analysis scripts) at submission time, or explicitly rebrand the paper as a specification-only manuscript whose quantitative results cannot yet be independently reproduced.","section":"Section 12"},{"comment":"The validation never executes the directed-reachability procedure of Section 4.2. As stated in Section 8.1, F6 and F7 are sampled from declared role and zone profiles, and Section 10 acknowledges this. Since F6 and F7 are the only factors that implement the paper's network-aware contribution, the differences between PECR and context-lite in Tables 6 and 8 are conditional on an unvalidated shortcut. If the actual procedure produced F6/F7 values that are nearly constant, dominated by F4, or differently correlated with the latent threat z, the reported refinement margin and queue-differentiation numbers could change materially. The paper should either run the Section 4.2 procedure on the synthetic topology and show whether the Table 6 and Table 8 results survive, or restrict the behavioral claims to the non-graph factors and relabel F6/F7 as externally supplied inputs.","section":"Section 8.1 and Section 4.2 / Tables 6 and 8"},{"comment":"The missing-evidence test appears to be a single realization of a 10% mask (90 cells), yielding a single set of counts (63 records lose a factor; 47 intervals cross a band boundary). Repeating the mask many times would show whether these counts are stable or a chance draw; a single realization is not a stress test in the same sense as the 30,000 weight draws. Please report a distribution over repeated masks (for example, median and 90% interval for the number of evidence-limited records), and note clearly that only uniform missingness is covered.","section":"Section 8.4"}],"minor_comments":[{"comment":"Context-lite is a reweighted five-factor subset of PECR's own factors, so labeling it a 'stronger five-factor context-lite comparator' in the abstract and Table 6 is potentially misleading; it is an ablation, not an independent baseline. The paper's prose in Section 8.2 is transparent about this, but the abstract and table caption should use 'ablation' or 'feature-subset comparator' to avoid the impression of independent validation.","section":"Section 8.2 and Table 6"},{"comment":"The phrase 'median interval width' appears to be a median over records for one mask, not a median over mask realizations. Clarify this in the text so readers do not confuse it with a distribution over missingness patterns.","section":"Section 8.4"},{"comment":"The AHP consistency ratio is reported (0.0023), but the reciprocal matrix itself is not shown in the text; since independent implementation requires the exact matrix, consider including it in an appendix or explicitly referencing the artifact if the artifact were available.","section":"Section 5.1"},{"comment":"The acknowledged overlap between F7 and F9 in centralized environments could be quantified; a correlation or redundancy analysis for these two factors on the primary cohort would help readers understand the incremental information contributed by F7.","section":"Section 10"},{"comment":"The 30-seed replication tests seed variation within one generator family only; the text acknowledges this, but the phrase 'repeatability across seeds' should not be read as robustness to alternative data-generating models.","section":"Section 8.5"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This is a thoughtfully bounded paper that does not overclaim. The main risks are that the claimed reproducibility contribution cannot be checked as submitted (no artifact) and that the graph-derived factors are not exercised by the validation. These are fixable within the manuscript's scope, so I would not reject. A major revision that archives the artifact and adds a graph-procedure check would put the paper in good shape; a distribution over masking draws would also strengthen the missing-evidence analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is the short version: PECR is a carefully specified decision-support method for SD-WAN vulnerability prioritization, and the paper is unusually honest about what its synthetic validation does and doesn't show. The main thing to know is that the central claims are modest and mostly hold up as stated; the one real soft spot is that the two network-graph factors (F6, F7) are sampled from role/zone profiles, not computed by the Section 4.2 reachability algorithm, so the \"full factor set refines the order\" result is conditional on an untested shortcut. The paper flags this itself, so it is not a hidden flaw, but it does mean the headline differentiation from simpler queues is not yet grounded in the actual graph procedure.\n\nWhat is new: the combination of nine telemetry-informed factors with separate confidence and interval reporting, plus a separate cryptographic-migration queue, into one executable protocol. The AHP weights are transparent, the sensitivity analysis over 30,000 Dirichlet draws is broad, and the replication across 30 seeds checks seed dependence. The paper also correctly notes that nominal weights differ from realized score mass, which is a useful operational point. The comparators are declared in advance, and the context-lite result (tau = 0.755) is honestly described as partly structural: it is a reweighted subset of PECR's own factors, so high agreement is expected. The paper explicitly disavows predictive accuracy and production effectiveness.\n\nSoft spots, in proportion. The biggest is F6/F7: without executing the graph algorithm, the refinement margin over context-lite could shrink or grow under real telemetry. The stress-test note is right about that, and it is the main reason the verdict should be conditional, not accept. Also, the artifact is not currently accessible—no archival DOI—so \"reproducible\" is still a promise. The masking test uses a single draw and one missingness mechanism, and the generator is hand-authored with a latent threat structure that drives both EPSS and KEV, so all results are conditional on that family. The unpopulated Critical band shows threshold calibration is an open issue, as the paper concedes.\n\nOverall, this is a serious, clear-headed specification paper with an honest validation agenda. The math is explicit, the limitations are stated in the text and in the conclusion, and the comparison protocol is fair. It deserves a serious referee. If the authors release the artifact and run a real graph-based check of F6/F7, the contribution becomes much stronger. I would send it to review, with a request that the F6/F7 shortcut be addressed or explicitly hived off as a separate validation step.","headline":"PECR is a transparent, well-scoped prioritization specification whose main claims are conditional on a sampled graph-factors shortcut; worth sending to review if the artifact and an F6/F7 check are required.","tokens_in":12732,"tokens_out":2058,"would_cite":false,"duration_ms":18714,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PECR specifies a nine-factor, telemetry-informed scoring method for SD-WAN vulnerability prioritization that separates score, confidence, and cryptographic migration; on synthetic records it substantially reorders remediation versus…","keywords":["vulnerability prioritization","SD-WAN security","telemetry-informed scoring","evidence quality","uncertainty intervals","global sensitivity analysis","cryptographic migration","synthetic validation"],"falsifier":"Run the PECR protocol on a real, multi-environment SD-WAN dataset with recorded exploitation outcomes and compare top-k precision and time-to-mitigate for later-exploited findings against the five-factor context-lite comparator; if context-lite matches or beats PECR's nine-factor queue on validated outcomes while PECR's extra factors merely reorder the queue, the central refinement claim fails. Alternatively, regenerate the synthetic cohort with a different latent-threat structure and check whether the full queue still reorders the top 10 and keeps median Kendall τ with context-lite above about 0.7 under the same weight stresses.","tokens_in":11710,"feed_emoji":"🛡️","tokens_out":8076,"duration_ms":74258,"temperature":0.7,"pith_summary":"PECR is a decision-support specification for ordering vulnerability remediation in SD-WAN environments. It combines nine normalized factors — base severity, exploit evidence, exploit probability, untrusted accessibility, attainable privilege, path reachability, blast radius, operational consequence, and control concentration — into a single score, while reporting confidence and plausible score intervals separately and placing cryptographic migration in an independent queue. The paper's central finding is calibrated and modest: on a 100-record synthetic cohort, PECR's operational order differs sharply from severity-, probability-, and KEV-only queues (Kendall τ ≈ 0.196–0.301), but a simpler five-factor context-lite comparator is far closer (τ = 0.755), indicating that most of the differentiation comes from adding basic context rather than the full factor set. The queue order remains reasonably stable under 30,000 joint Dirichlet weight draws (median τ ≈ 0.836–0.924), and masking 10% of factor cells labels 47 records as evidence-limited. The paper claims executable specification, queue differentiation, and stress-test robustness — not predictive accuracy or production effectiveness.","feed_headline":"Richer score reorders SD-WAN patches beyond CVSS and EPSS","feed_subtitle":"On synthetic SD-WAN records, nine-factor PECR stays stable under 30,000 weight draws and shifts the top 10 versus severity-only queues.","key_machinery":"The load-bearing object is the weighted aggregation identity $R = 100\\sum_{i=1}^{9} w_i f_i$ over nine normalized factors, with an AHP-elicited weight vector and a companion confidence formula $C = \\sum_{i=1}^{9} w_i q_i$, plus interval-valued factor bounds that yield $R^-$ and $R^+$. The mechanism around it does the work: a bounded directed reachability computation produces $F_6 = 1/(1+h)$ (path distance from an untrusted origin) and $F_7$ (two-hop blast radius), and the interval rule routes any record whose $R^-$ and $R^+$ cross a band boundary into an evidence-limited verification workflow. Cryptographic dependencies are scored by $g(d) = 0.25c_1 + 0.25c_2 + 0.20c_3 + 0.10c_4 + 0.10c_5 + 0.10c_6$, aggregated as $G(a) = \\max_d g(d)$, and sent to an independent migration queue. The final-band, descending-score, stable-identifier sort key makes the operational order deterministic and replayable.","core_discovery":"The paper establishes that a fully specified, vendor-neutral scoring protocol can turn SD-WAN telemetry into a deterministic work queue in which each record carries a score $R$, a confidence $C$, and lower/upper bounds, and in which records whose bounds cross a criticality band are routed to verification. The score is the weighted sum of nine normalized factors, with weights derived from an analytic hierarchy process matrix; the operational sort key is final band, then descending $R$, then record identifier, with an exception that forces any confirmed-exploited and Internet-accessible finding to at least High. A separate cryptographic-readiness queue, scored by $g(d)$ and maximized over an asset's cryptographic dependencies, keeps post-quantum migration planning distinct from vulnerability remediation. In the synthetic evaluation, this full protocol reorders the queue relative to CVSS-only, EPSS-only, KEV-first, and CVSS×EPSS comparators, but the context-lite comparator shows that basic context already accounts for most of the top-10 behavior, with PECR's extra factors refining within- and near-band order. The author explicitly frames the result as evidence of incremental differentiation and synthetic feasibility, not a demonstration of avoided loss or accuracy.","pith_inferences":["If the synthetic differentiation transfers, the first deployment win is likely to come from assembling a five-factor context-lite queue, because it captures most of the top-10 difference with far fewer inputs and less topology dependency.","A natural extension the paper does not run: replace the sampled $F_6$ and $F_7$ with actual graph-construction outputs on a small real SD-WAN topology and re-measure the Kendall τ versus context-lite; this would test whether the reachability machinery adds order information beyond its sampled distribution.","The evidence-limited interval rule implies a cost-of-honesty problem: adversarial or contradictory inputs can inflate the number of records routed to verification, so a production deployment would need per-source rate limits and a queue floor for low-confidence, high-upper-bound records; the paper identifies but does not solve this.","The unpopulated Critical band suggests the default band thresholds are uncalibrated for this generator, so adoption would require a governed calibration step on historical outcomes, not just configuration; this is the author's own limitation, but the corollary is that an outcome-feedback loop is needed for operational use."],"forward_implications":["An operator can implement PECR from the specification alone: the sort key, factors, interval rules, and thresholds are defined precisely enough for independent, reproducible implementation.","Because score and confidence are reported separately and intervals are preserved, missing or contradictory evidence changes the workflow — a record is sent to verification rather than silently scored at a midpoint.","The E1 exception floor guarantees that any finding with confirmed exploitation and Internet accessibility is at least High in the queue regardless of its computed score.","Cryptographic migration planning is decoupled from vulnerability remediation: an asset with a classical-crypto dependency that fails the horizon check appears in the migration queue via $G(a)$ without changing its vulnerability score.","The synthetic results imply that organizations adopting such a method should expect the biggest queue changes from basic context (exploit evidence, probability, accessibility, consequence), with privilege, path, blast-radius, and control-concentration factors mostly refining near-band order."],"supporting_citations":[{"why":"Supplies the intrinsic severity scores used to normalize F1 and to build the CVSS-only comparator queue.","marker":"[4]"},{"why":"Defines EPSS as a 30-day exploitation probability, the source for F3 and the EPSS-only comparator.","marker":"[5]"},{"why":"Defines the confirmed-exploitation signal used in F2 and by the KEV-first comparator.","marker":"[7]"},{"why":"Provides the stakeholder-specific context-aware decision-tree approach that PECR positions against.","marker":"[8]"},{"why":"Documents score disagreement across vulnerability scoring systems, motivating the need for a decision-specific, evidence-preserving protocol.","marker":"[11]"},{"why":"Supplies the logic-based attack-graph approach that motivates PECR's bounded directed reachability for F6 and F7.","marker":"[18]"},{"why":"Provides the analytic hierarchy process used to derive the starting weight vector for the nine factors.","marker":"[37]"},{"why":"Provides the compositional-data simplex distribution used for the 30,000-draw joint weight sensitivity test.","marker":"[39]"},{"why":"Supplies the CycloneDX carrier format that the separate cryptographic-migration queue builds on.","marker":"[29]"}],"fun_headline_variants":["Nine-factor PECR reorders SD-WAN patch queues versus CVSS and EPSS","Synthetic stress test: PECR ranking stable across 30k weight draws","PECR scores SD-WAN vulnerabilities with intervals, not just severity","PECR patch queue reorders CVSS, EPSS on synthetic SD-WAN data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The validation rests on the synthetic generator: real SD-WAN telemetry could correlate severity, exploitability, reachability, and consequence differently than the hand-declared role and zone profiles, and F6 and F7 are sampled rather than produced by executing the graph-construction procedure, so the reported queue differentiation and stability numbers may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Nine-factor PECR reorders SD-WAN patch queues versus CVSS and EPSS","Synthetic stress test: PECR ranking stable across 30k weight draws","PECR scores SD-WAN vulnerabilities with intervals, not just severity","PECR patch queue reorders CVSS, EPSS on synthetic SD-WAN data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00107,"raw_usage":{"total_tokens":4542,"prompt_tokens":1067,"completion_tokens":3475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":3398}},"tokens_in":683,"tokens_out":3475,"duration_ms":22369,"temperature":1.0,"reasoning_tokens":3398,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:52:36.504331+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the PECR protocol on a real, multi-environment SD-WAN dataset with recorded exploitation outcomes and compare top-k precision and time-to-mitigate for later-exploited findings against the five-factor context-lite comparator; if context-lite matches or beats PECR's nine-factor queue on validated outcomes while PECR's extra factors merely reorder the queue, the central refinement claim fails. Alternatively, regenerate the synthetic cohort with a different latent-threat structure and check whether the full queue still reorders the top 10 and keeps median Kendall τ with context-lite above about 0.7 under the same weight stresses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the intrinsic severity scores used to normalize F1 and to build the CVSS-only comparator queue."},{"cited_title":"Exploit Prediction Scoring System: The EPSS Model","cited_arxiv_id":null,"evidence_quote":"Defines EPSS as a 30-day exploitation probability, the source for F3 and the EPSS-only comparator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the confirmed-exploitation signal used in F2 and by the KEV-first comparator."},{"cited_title":"Spring, Eric Hatleback, Allen Householder, Art Manion, and Deana Shick","cited_arxiv_id":null,"evidence_quote":"Provides the stakeholder-specific context-aware decision-tree approach that PECR positions against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the logic-based attack-graph approach that motivates PECR's bounded directed reachability for F6 and F7."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CycloneDX carrier format that the separate cryptographic-migration queue builds on."}],"review_version":1}