{"id":"aa01241b-2c3c-4be2-b2be-f994490ae627","arxiv_id":"2508.13644","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"The abstract claims a first-of-kind, outcome-linked comparison of four vulnerability scoring systems showing major ranking disagreements, but the submitted full text is an unrelated paper, leaving the study unevaluable.","lead":"This paper compares four vulnerability scoring systems, CVSS, SSVC, EPSS, and the Exploitability Index, on 600 real-world vulnerabilities to see which one best captures actual exploitation risk. The abstract claims the systems rank the same bugs very differently, but the manuscript body provided is an unrelated machine-learning paper, so the study itself cannot be verified.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The provided full text is an unrelated paper (Text2Weight), so the claimed vulnerability-scoring study—methods, dataset, and results—is entirely absent; the abstract's empirical claims are unverifiable.","rationale":"The reader's verdict of UNVERDICTED is correct and remains unchanged. The reader's stated weakest assumption—accuracy and independence of exploitation-outcome labels—is a valid concern at the abstract level, but it is not the most load-bearing concern. The more fundamental issue is that the manuscript contains no methods, dataset, or analysis for the described vulnerability study; it is an entirely different paper. Without the study's methodology, one cannot even determine what the labels are, how they were constructed, or whether they overlap with EPSS training data. The label concern is therefore downstream of the document mismatch. I partially agree with the reader because they did identify the absent methods as a reason for UNVERDICTED in their rationale, but their formal weakest assumption was about label quality. There is no need to adjust the verdict: the submission remains unverifiable. I also note the reviewing rule that the full text is in-scope evidence; here it clearly contradicts the abstract's subject matter, which is a concrete internal inconsistency rather than a mere lack of detail. No further speculation about author intent is warranted—the concern is about the argument's support, not the authors.","tokens_in":17637,"tokens_out":2230,"duration_ms":22740,"concrete_test":"Retrieve the actual PDF for arXiv:2508.13644 from arXiv and compare its full text with the abstract. Specifically, search the PDF for any section describing: (a) the 600 vulnerabilities from four months of Microsoft Patch Tuesday disclosures, (b) CVSS/SSVC/EPSS/Exploitability Index scoring, or (c) exploitation-outcome labels. If no such section exists—as in the provided full text, which is the Text2Weight paper—the described empirical study is absent and the abstract's findings cannot be verified. Also check the arXiv metadata (title and author list) to confirm whether this is a submission error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical comparison of CVSS, SSVC, EPSS, and the Exploitability Index on 600 Patch-Tuesday vulnerabilities, with findings about ranking disparities and real-world exploitation risk. For that claim to be assessable, the manuscript must contain the methodology, dataset construction, outcome-label definitions, and analysis. It does not: the full text under arXiv:2508.13644 is 'Text2Weight: Bridging Natural Language and Neural Network Weight Spaces' by Tian et al. (arXiv:2508.13633), on an unrelated topic. No section of the submitted text describes vulnerability scoring systems, the 600-vulnerability dataset, or exploitation-outcome labels. The abstract is the only in-scope content for the claimed study, and it offers no methods or evidence. Consequently, every downstream concern—including the accuracy and independence of exploitation labels relative to EPSS training data—cannot even be examined, because the supporting content is missing. This is an internal inconsistency in the submission, not a disagreement with consensus. The abstract's claims may be true, but the document as submitted provides no way to check them, so the central claim is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of arXiv:2508.13644 claims a large-scale, outcome-linked empirical comparison of four vulnerability scoring systems (CVSS, SSVC, EPSS, and the Exploitability Index) on 600 Patch Tuesday vulnerabilities, reporting significant ranking disparities and assessing the systems' ability to capture real-world exploitation risk. However, the full text supplied with the submission is an entirely unrelated manuscript, 'Text2Weight: Bridging Natural Language and Neural Network Weight Spaces' (arXiv:2508.13633), which describes a diffusion transformer for generating neural network weights from text. No section of the provided document describes the vulnerability dataset, the scoring systems, the outcome-label definitions, or any analysis of the claimed study. The abstract is therefore the only in-scope content, and its central claims are unsupported by any methods, data, or results in the submission.","tokens_in":17754,"tokens_out":2203,"duration_ms":24463,"significance":"If the claimed vulnerability-scoring study existed and were methodologically sound, it would be a valuable contribution to vulnerability management, providing independent evidence about the agreement and predictive validity of widely used scoring systems. It could also inform organizations that rely on these scores for risk-based prioritization. However, as submitted, the paper contains none of the supporting content: no dataset, no methodology, no outcome-label definitions, no tables, no statistical analysis, and no reproducibility artifacts. There are no machine-checked proofs, parameter-free derivations, or falsifiable predictions that could be assessed. Because the central empirical claim is entirely unverifiable from the manuscript, the significance of the work cannot be evaluated in its current form.","major_comments":[{"comment":"The body of the submission is not the study described in the abstract. The full text is 'Text2Weight,' an unrelated machine-learning paper, and it contains no methods, dataset construction, outcome-label definitions, scoring-system implementations, tables, or figures relevant to vulnerability scoring. The claimed comparison of CVSS, SSVC, EPSS, and the Exploitability Index on 600 Patch Tuesday vulnerabilities is therefore unsupported. This is an internal inconsistency in the submission, not a matter of disagreement with consensus; the central claim cannot be checked at all.","section":"Full text (entire body)"},{"comment":"The abstract states that the study is 'outcome-linked' and evaluates the systems' 'ability to capture the real-world exploitation risk.' This premise is load-bearing because EPSS is itself a model trained on exploitation-in-the-wild data. If the outcome labels are drawn from the same public exploitation feeds used to train EPSS, the comparison would be partly self-referential; if the labels are incomplete or noisy, the performance of all systems would be underestimated. Since no methods or data are provided, the independence and accuracy of the outcome labels cannot be assessed. This concern cannot be resolved from the submitted manuscript.","section":"Abstract"},{"comment":"The abstract claims that this is 'the first large-scale, outcome-linked empirical comparison' of the four scoring systems. This novelty claim cannot be verified without a detailed description of the dataset, the vulnerability selection process, the outcome-label sources, and a comparison with prior work. None of these are present in the submission. Without this information, the claim of 'first' is unsubstantiated.","section":"Abstract"}],"minor_comments":[{"comment":"The arXiv identifier and title in the abstract do not match the content of the full text, which is marked as arXiv:2508.13633. This appears to be a compilation or submission error that should be corrected.","section":"Metadata"},{"comment":"The abstract refers to 'the Exploitability Index' without defining which index is meant or specifying its version. In a proper methods section this would need to be clarified; here, the absence of any definition compounds the lack of content.","section":"Abstract"},{"comment":"Figures, equations, and references in the body all pertain to neural network weight generation, not vulnerability scoring. The document is internally inconsistent as a submission and cannot be reviewed as a coherent manuscript.","section":"Full text (body)"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error: the full text supplied corresponds to a different paper (Text2Weight, arXiv:2508.13633) rather than the vulnerability-scoring study claimed in the abstract. The editor may wish to verify the arXiv listing and whether the intended manuscript exists elsewhere. As it stands, the claimed study is entirely absent from the submission, so no amount of revision to the current document could make the central claim assessable; a fresh submission with the correct content would be needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know upfront: the submission is not internally coherent. The metadata and abstract describe an empirical study of CVSS, SSVC, EPSS, and the Exploitability Index on 600 Patch Tuesday vulnerabilities. The full text under that arXiv ID is a different paper entirely—“Text2Weight,” about generating neural network weights from text, by different authors. There is no section in the body about vulnerability scoring, no dataset, no tables of ranking disparities, no outcome-label definitions. The claimed study exists only as an abstract.\n\nI agree with the reader’s take: this is unverdictable. The abstract’s claim of being the “first large-scale, outcome-linked comparison” of those four systems is plausible on its face—pairwise CVSS/EPSS and CVSS/SSVC comparisons exist, but a four-way outcome-linked comparison would be a useful within-subfield measurement. The abstract is readable and the research question is legitimate. If the actual study existed, it could matter for vulnerability management triage.\n\nBut it doesn’t exist in the submitted document. The reader’s weakest-assumption point about EPSS circularity—if the exploitation-outcome labels come from the same public feeds used to train EPSS, the comparison becomes partly self-referential—is a real hazard, but it’s a secondary concern. You can’t even examine that hazard because the methods are missing. The same goes for label accuracy and independence. The document provides no way to check any of it.\n\nWhat’s good here? Very little, other than the abstract’s framing. There’s no formal verification, no shipped data or code for the claimed study. The unrelated Text2Weight paper does have some experimental details, but it’s not what’s being submitted for review under this title, so it doesn’t count as evidence for the abstract’s claims.\n\nMy recommendation: this should be desk rejected, not sent to peer review. A referee has nothing to assess. The abstract alone doesn’t warrant referee time, and the internal contradiction means the paper cannot be evaluated on its own terms. The authors may have uploaded the wrong file; if so, they should resubmit with the actual manuscript. But as submitted, this is a no.","headline":"The abstract promises a useful four-way comparison of vulnerability scoring systems, but the manuscript body is an unrelated neural-network paper, so there is nothing to evaluate.","tokens_in":18381,"tokens_out":2086,"would_cite":false,"duration_ms":20222,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that four public vulnerability scoring systems rank the same real-world vulnerabilities substantially differently, and that some of them fail to track actual exploitation risk, based on a dataset of 600 Patch Tuesday vulne","keywords":["vulnerability scoring","CVSS","SSVC","EPSS","Exploitability Index","Patch Tuesday","vulnerability prioritization","real-world exploitation risk"],"falsifier":"Re-run the comparison on the same 600 vulnerabilities using independently verified exploitation labels (for example, from vendor incident reports rather than public exploit feeds). If the four scoring systems then agree closely in their rankings, the claim of significant disparity collapses; alternatively, if EPSS's predictive performance on independently verified labels is much worse than on the public-feed labels, the paper's assessment of EPSS is undercut.","tokens_in":17396,"feed_emoji":"🛡️","tokens_out":3454,"duration_ms":33891,"temperature":0.7,"pith_summary":"The paper aims to provide the first large-scale, outcome-linked empirical comparison of four publicly available vulnerability scoring systems: CVSS, SSVC, EPSS, and the Exploitability Index. Using 600 real-world vulnerabilities from four months of Microsoft Patch Tuesday disclosures, it examines how the systems rank the same vulnerabilities, how they categorize them across triage tiers, and how well they capture real-world exploitation risk. The reported finding is significant disparity across systems, with implications for security teams that rely on these metrics for data-driven prioritization. The abstract stands alone here because the full text supplied with this record is a different paper; the summary therefore rests on the abstract and the stated discovery.","feed_headline":"Vulnerability scoring systems sharply disagree on the same bugs","feed_subtitle":"A 600-vulnerability, outcome-linked study finds some scores miss real-world exploitation risk.","key_machinery":"The comparison is carried by an outcome-linked dataset: 600 real-world vulnerabilities from four months of Microsoft Patch Tuesday disclosures, each connected to whether it was actually exploited. The four scoring systems—the Common Vulnerability Scoring System (CVSS), the Stakeholder-Specific Vulnerability Categorization (SSVC), the Exploit Prediction Scoring System (EPSS), and the Exploitability Index—are the objects whose rankings and risk assessments are compared against these exploitation outcomes. The outcome link is what allows the paper to test whether the scores capture real-world exploitation risk, not just internal severity.","core_discovery":"The central claim is that the four scoring systems diverge markedly when ranking the same vulnerabilities, and that at least some of them do not adequately reflect whether a vulnerability was actually exploited in the wild. The study builds an outcome-linked dataset of 600 vulnerabilities drawn from four months of Microsoft Patch Tuesday disclosures, linking each vulnerability to its real-world exploitation status. The authors report significant disparities in how the systems rank the same vulnerabilities, and they assess the systems' ability to capture real-world exploitation risk. If the paper is right, organizations using a single scoring system to triage vulnerabilities may be making sys","pith_inferences":["The full text supplied with this record is a different paper (on generating neural-network weights from text descriptions), not the vulnerability-scoring study described in the abstract; this extraction is therefore based on the abstract and reader-pass notes, and the methods section behind the comparison is unavailable for verification.","If EPSS was trained on the same public exploitation feeds used to label the 600 vulnerabilities as exploited or not, then the claimed test of EPSS is partly self-referential and would overstate its agreement with the outcome labels.","Real-world exploitation data is incomplete and noisy: many exploited vulnerabilities are never publicly documented, so the outcome labels may be incomplete, which would make every scoring system look worse than it is.","A testable extension is to run the same comparison with independently verified exploitation labels (e.g., from vendor incident reports) to see whether the measured disparities persist."],"forward_implications":["Security teams should not treat any single severity score as a reliable proxy for real-world exploitation risk.","Organizations that triage vulnerabilities using one scoring system may be ordering their remediation priorities differently than they would under another system.","The measured disparities motivate more transparent and consistent definitions of exploitability, risk, and severity.","The findings suggest that relying on a single metric could lead to systematically misallocated patching effort, and that a consensus or blended approach may be worth exploring."],"supporting_citations":[],"fun_headline_variants":["Scoring systems clash on vulnerability risk, study finds","Vulnerability scores often fail to predict actual exploits","600 bugs reveal scoring systems rank risk differently","Bug scoring systems miss real-world exploitation risk","Patch Tuesday study: scoring systems give mixed signals"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The comparison rests on the premise that each of the 600 vulnerabilities' true exploitation status is known and correctly recorded, and that those labels are not drawn from the same public exploitation feeds that EPSS was trained on.","fun_headline_variants_meta":{"raw":{"variants":["Scoring systems clash on vulnerability risk, study finds","Vulnerability scores often fail to predict actual exploits","600 bugs reveal scoring systems rank risk differently","Bug scoring systems miss real-world exploitation risk","Patch Tuesday study: scoring systems give mixed signals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1323,"prompt_tokens":691,"completion_tokens":632,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":561}},"tokens_in":435,"tokens_out":632,"duration_ms":7340,"temperature":1.0,"reasoning_tokens":561,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:57:42.595041+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison on the same 600 vulnerabilities using independently verified exploitation labels (for example, from vendor incident reports rather than public exploit feeds). If the four scoring systems then agree closely in their rankings, the claim of significant disparity collapses; alternatively, if EPSS's predictive performance on independently verified labels is much worse than on the public-feed labels, the paper's assessment of EPSS is undercut.","supporting_citations":[],"review_version":1}