{"id":"6e1e2424-9636-41aa-8c0d-8df2f8fd58c2","arxiv_id":"2607.14587","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 109 Qubes Security Bulletins (2011–2025), ~80% are attributable to upstream Xen/hypervisor, CPU, or integration components rather than Qubes-core logic; the advisory rate shifted up in 2015Q1 and has been statistically flat since 2018.","lead":"This paper counts and labels 14 years of Qubes OS security bulletins to see where the security burden actually sits. It finds most bulletins trace to Xen, CPU, or other upstream components, that disclosure activity jumped in 2015 then plateaued after 2018, and that standard S-curve models do not beat simple baselines for forecasting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Attribution codebook validation is post hoc; an independent blind full-corpus audit is required before accepting the 79.8% upstream share.","rationale":"I agree with the reader's weakest-assumption analysis: the attribution codebook is the single most load-bearing premise, and the post hoc 30-QSB validation does not constitute an independent accuracy check. The headline 79.8% upstream share is a simple proportion, so a small number of reclassifications could shift it materially. The paper's own mixed-title example (QSB-034) shows the exact failure mode. Other potential objections—such as the forced 2018 breakpoint, low-power Diebold–Mariano tests, or 2025 right-censoring—are acknowledged and supported by sensitivity checks, so they are less consequential. The proposed independent full-corpus audit is feasible because the artifact contains per-QSB assignments and raw trigger matches; it would settle whether the upstream-dominance claim is a property of the record or an artifact of tie-breaking. The reader's CONDITIONAL verdict is appropriate; I would keep it unchanged pending that check.","tokens_in":11745,"tokens_out":5190,"duration_ms":60703,"concrete_test":"Have two independent annotators, blinded to the codebook, the paper's labels, and each other, manually classify all 109 QSBs from full bulletin texts using a pre-registered coding manual with the same four categories but with no precedence rule and no title-only tokens. Compute inter-rater agreement (Cohen's kappa) and the upstream share under consensus labels. If kappa < 0.80 or the consensus upstream share falls outside the reported 95% Wilson CI (71.3%–86.3%), the codebook's tie-breaking is not a stable basis for the headline and the claim should be re-expressed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing quantitative claim is the 79.8% (87/109) upstream share. It rests entirely on a deterministic, title-driven codebook whose fixed precedence order (CPU/microarchitecture > Xen/hypervisor > Qubes-core > Upstream integration) is the only thing breaking mixed-title ties. The reported 30-QSB validation is post hoc: the subset was drawn from the same corpus used to author the rules, and the initial manual pass, while blind to the outputs, was performed by the same researcher who designed the codebook. High agreement on this subset shows internal consistency, not out-of-sample accuracy. The precedence rule itself is not neutral: any title containing both an XSA/Xen token and a qrexec/GUI/policy token is counted upstream, and moving even five of the 109 bulletins from upstream to Qubes-core changes the headline from 79.8% to 75.2% (nine would bring it to 71.6%). The paper's only identified primary-label disagreement, QSB-034, is exactly such a mixed GUI/Xen bulletin. Additionally, the 'Upstream integration' category includes tokens like Salt and RPM, which may capture Qubes packaging/integration work rather than purely upstream code. The XSA-tracker statistic (113/464) is less vulnerable because it uses Qubes' official relevance labels, but the headline specifically depends on the QSB attribution. Without an independent full-corpus audit, the upstream-dominance finding could be an artifact of tie-breaking rather than a property of the public record.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a protocol-driven longitudinal analysis of the public Qubes OS advisory record: all 109 Qubes Security Bulletins (QSBs, 2011–2025), the official Qubes-maintained Xen Security Advisory (XSA) tracker, and a secondary annualized vulnerability-event series. It makes four main claims: (1) under the paper's primary component-attribution codebook, 79.8% of QSBs are attributable to upstream components (Xen, CPU/microarchitecture, or other upstream integration) rather than Qubes-core logic; (2) the quarterly advisory series exhibits a dominant change-point around 2015Q1, with a statistically flat annual rate after 2018; (3) the attribution and stable-regime conclusions survive multiple sensitivity checks (weighted/incidence views, overdispersion diagnostics, endpoint exclusion); and (4) S-shaped vulnerability discovery models fit descriptively but do not significantly outperform simple rolling baselines in short-horizon forecasts. The paper positions itself as measuring the public record, not latent vulnerability incidence, and repeatedly cautions against over-interpretation.","tokens_in":12203,"tokens_out":7741,"duration_ms":77837,"significance":"If the headline claims hold, the paper makes a useful empirical contribution: it provides a fully reproducible QSB ledger and XSA-tracker snapshot, along with a transparent multi-scheme attribution protocol, and it reports a genuinely useful negative result on VDM forecasting. The overdispersion and censoring sensitivity checks are methodologically sound, and the change-point analysis converges on a consistent break across four methods. The XSA-tracker statistic (113/464, using Qubes' official relevance labels) is an independent, less codebook-dependent corroboration of upstream dependence. The paper is carefully scoped and its stated limitations are generally conservative; this is the kind of self-aware measurement study that the security community needs. The main risk is that the 79.8% headline and the 'stable post-2018' wording are stronger than the evidence from the post hoc, single-coder attribution validation and the low-power slope tests, respectively.","major_comments":[{"comment":"The validation of the deterministic attribution codebook is post hoc and single-coder: the 30-QSB audit subset is drawn from the same corpus used to author the token rules, and the manual pass was performed by the same researcher who designed the codebook. High agreement (96.7% primary-label accuracy, macro-F1 0.969) documents internal consistency, not out-of-sample accuracy. This matters because the primary-label scheme with fixed precedence (CPU > Xen > Qubes-core > Upstream integration) is the basis for the headline 79.8% upstream share. The paper acknowledges this limitation in Sec. 6 but does not provide a mitigating analysis. Please add an independent full-corpus audit by a second coder blind to the codebook, or a formal development/test split with the codebook frozen before validation. At minimum, report a sensitivity analysis that varies the precedence order (e.g., Qubes-core fir","section":"Section 3.2, Table 2"},{"comment":"The 'Upstream integration' category conflates genuinely upstream code with Qubes-specific integration and packaging work (Salt, RPM, Linux netback, kernel driver, template packaging, domU integration issue). Since the paper's central interpretive claim is that the public-record burden is concentrated in 'upstream trust anchors' (Abstract, Sec. 5), this category can inflate the upstream share if a QSB concerns Qubes-maintained packaging/integration logic rather than an upstream component. Report the upstream share with 'Upstream integration' excluded (i.e., Xen/hypervisor + CPU/microarchitecture only) under each of the four attribution views, and discuss how the conclusion changes. The XSA-tracker statistic (113/464) is a useful independent check but covers only Xen, not the full 'upstream' grouping used in the headline.","section":"Section 4.2, Tables 1 and 4"},{"comment":"The 'stable post-2018 regime' conclusion is supported by failing to reject the null that the net post-2018 slope β1+β2 is zero (p=0.208 with the partial 2025 endpoint; p=0.989 without it). In a series of only 14–15 annual observations this test has low power, so non-significance is weak evidence of a zero slope. To support the wording 'statistically flat,' report a 95% confidence interval for β1+β2 (or an equivalence test with a pre-specified bound). The same concern applies to the secondary-series trend test (z=-0.92, p=0.356). The change-point analysis around 2015Q1 is more convincing, but the post-2018 'stability' claim needs stronger statistical support.","section":"Section 4.1, Eq. (1)"}],"minor_comments":[{"comment":"'Distribution-light' appears to be a typo for 'distribution-free' in the description of the Mann–Kendall test.","section":"Section 3.3"},{"comment":"The full attribution codebook is only in the external artifact. Since the codebook is a central piece of the methodology, consider including the complete token list and precedence rules as an appendix rather than representative triggers only.","section":"Data Availability / Appendix A"},{"comment":"For the incidence view, make explicit that percentages sum to more than 100% because one bulletin can contribute to multiple categories; the footnote says 'incidence mass,' but a parenthetical in the table header would improve clarity.","section":"Table 1"},{"comment":"'Collapses essentially to zero' is stronger than the reported estimates justify. The NB2 overdispersion parameter is near zero, but the one-sided profile-likelihood upper bounds (α < 0.15 annually, α < 0.21 quarterly) are not negligible; consider phrasing such as 'small' or 'consistent with negligible overdispersion.'","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope for a security-measurement venue and the artifact is a real strength. The headline attribution figure is the main risk: the post hoc, single-coder validation is not sufficient for the precision claimed, though the weighted/incidence views and the XSA-tracker statistic provide partial robustness. The change-point and VDM-forecast analyses are generally sound, and the negative forecasting result is valuable. I would be willing to review a revised version that adds an independent audit or a formal sensitivity analysis of the attribution, and that reports confidence intervals for the post-2018 slope. No concerns about the authors' conduct; the self-identified limitations are candid."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know the main thing about this paper: it is a genuinely careful, reproducible measurement study of the Qubes OS public advisory record, and its central claims hold up. The headline numbers — 79.8% upstream attribution and a flat post-2018 disclosure rate — are not overclaimed, and the paper is unusually honest about what the public record can and cannot tell you.\n\nWhat is actually new: a complete longitudinal dataset of all 109 QSBs, an XSA-tracker snapshot, and a protocol that combines deterministic attribution, change-point analysis, overdispersion checks, and baseline-aware VDM forecast evaluation. The negative VDM result — S-shaped curves fit well but do not beat a rolling three-year mean in short-horizon forecasts — is a useful revalidation of the Massacci–Nguyen caution on a new case. That is a real contribution, especially since the data bundle is being released.\n\nWhat the paper does well: it scopes itself to the advisory record, says so repeatedly, and then does the technical work properly. The 2015Q1 break is recovered by four different change-point methods. The Poisson assumption is actually checked, with negative-binomial sensitivity and profile-likelihood bounds. The right-censored 2025 endpoint is handled by exclusion, and the stable-regime result gets stronger. The four attribution views give a coherent picture. All of this is solid.\n\nThe soft spot is the one the stress-test note flags: the attribution codebook is title-driven, authored for this corpus, and validated on a 30-QSB post hoc subset of the same corpus. That is a real limitation, and the paper admits it. But I do not think it is as damaging as the stress-test suggests. The XSA-tracker statistic (113/464) is independent of the codebook and points the same direction. The incidence and weighted views, while sharing the same token rules, at least show that the result is not an artifact of one tie-breaking scheme. Moving five bulletins from upstream to Qubes-core drops the share from 79.8% to 75.2% — still upstream dominance. So the concern is about precision, not about the qualitative conclusion. For a revision, I would push for a fully independent dual-coder audit or a pre-registered codebook, and for release of the secondary event-level series if provenance allows.\n\nMinor caveats: the Diebold–Mariano tests are explicitly low-power, which the paper acknowledges; the secondary series is aggregate-only; latency analysis is a lower bound, not a full reconstruction. None of these undercut the main claims.\n\nWho gets value: security measurement researchers, anyone doing advisory-based longitudinal studies, and Qubes/Xen practitioners who want a quantitative basis for treating upstream advisory streams as first-class inputs. It deserves a serious referee. I would accept it into peer review with a request for an independent attribution audit.","headline":"A careful, reproducible study of the Qubes advisory record whose central claims hold up; main caveat is the post hoc attribution codebook, which shifts the headline numbers a bit but not the qualitative picture.","tokens_in":12583,"tokens_out":1707,"would_cite":true,"duration_ms":19310,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that roughly 80% of fourteen years of public Qubes advisories are attributable to upstream Xen, CPU, and integration components rather than Qubes-core logic, with statistically flat disclosure since 2018.","keywords":["Qubes OS","Xen","security advisories","upstream dependence","vulnerability discovery models","change-point analysis","Poisson regression","microarchitectural vulnerabilities"],"falsifier":"Have two independent coders read the full text of all 109 QSBs and assign primary components without seeing the codebook; if the resulting upstream share falls below 50%, or if a codebook-precedence reversal drops the share below a majority, the central claim is a measurement artifact. Separately, the stable-regime claim is falsified if complete 2025-2026 data produce a significantly positive post-2018 slope rather than the flat slope estimated here.","tokens_in":11656,"feed_emoji":"🐛","tokens_out":6065,"duration_ms":58211,"temperature":0.7,"pith_summary":"The paper tries to establish that Qubes OS's public security advisory record is dominated by upstream components—Xen, CPU/microarchitecture, and integrations—rather than by Qubes' own core logic. Across 109 Qubes Security Bulletins from 2011 to 2025, 79.8% are attributed upstream, a share that survives all four weighting schemes. It also claims the quarterly disclosure series broke around 2015Q1 and that post-2018 annual rates are statistically flat, meaning the record is stable but not quiet. A careful reader cares because this is a rare longitudinal case where component boundaries are security-relevant and the advisory record is clean enough to test whether upstream dependence is visible in public data.","feed_headline":"Four of five Qubes security bulletins trace to upstream bugs","feed_subtitle":"A 14-year audit of 109 bulletins shows the disclosure record is stable—and dominated by Xen, CPU, and microcode flaws.","key_machinery":"The load-bearing mechanism is a deterministic, title-driven attribution codebook: a fixed precedence order (CPU/microarchitecture > Xen/hypervisor > Qubes-core > upstream integration) maps each bulletin's title tokens to one primary and multiple incidence labels, turning free-text QSB titles into a reproducible four-category ledger. This ledger produces the 79.8% upstream figure and feeds the change-point and count-model tests. Supporting machinery includes Bayesian single-change-point and BIC piecewise Poisson segmentation to locate regime breaks, overdispersion diagnostics (Pearson dispersion, negative-binomial profile-likelihood bounds) to validate Poisson assumptions, and rolling one-ste","core_discovery":"On its own terms, the paper's central discovery is that the public Qubes advisory stream is shaped far more by the hypervisor, processor, and upstream integration layers than by Qubes-maintained control logic. Using a deterministic, title-driven attribution codebook, 87 of 109 QSBs (79.8%) receive an upstream primary label; the share stays at 80.9% under weighted multi-label attribution and 82.5% under identifier weighting, while Qubes-core never exceeds 20.2%. Regime analysis consistently locates a single dominant break at 2015Q1, after which advisory volume enters a higher plateau; a piecewise Poisson model with an architecture-informed 2018 break yields a net post-2018 slope statistically","pith_inferences":["If the upstream share is as stable as the paper finds, Qubes' real assurance bottleneck sits outside its own codebase; a natural next step is to test whether future QSB rates track Xen advisory rates or microcode release waves, which the paper measures only at the aggregate level.","The same protocol could be applied to other isolation-heavy systems with explicit trust anchors—other Xen-based desktops, microkernel or unikernel systems—to test whether upstream dominance is a general property of compressed-TCB architectures rather than a Qubes quirk.","The negative VDM forecast result suggests operational teams should budget advisory-response effort using rolling historical averages rather than fitted growth curves; the paper states the statistical result but leaves the budgeting consequence implicit.","One testable extension is to correlate the quarterly series against external indicators such as Xen advisory volume or new microcode releases, to check whether the 2015Q1 break and the post-2018 plateau are driven by upstream disclosure waves rather than Qubes' own release process."],"forward_implications":["Monitoring the Qubes-maintained Xen Security Advisory (XSA) tracker and the Xen advisory stream should be a first-class input to operational security planning, since upstream issues account for the large majority of public bulletins.","The documented security-testing-to-stable cadence of roughly two weeks is the operative exposure window; update staging decisions should be made against it.","Post-2018 advisory volume is statistically flat, so near-term jumps in QSB count are more likely to come from exogenous microcode or transient-execution waves (31.5% of post-2018 bulletins) than from Qubes-core code changes.","S-shaped vulnerability discovery models fit the cumulative record descriptively but do not beat a rolling three-year mean in short-horizon forecasts, so simple baselines should be favored for annual planning.","Qubes-core advisories remain a minority under every attribution scheme, so the public record is not evidence of a Qubes-core quality problem."],"fun_headline_variants":["Upstream bugs fuel 8 in 10 Qubes security advisories","Qubes' security record is stable—and upstream-dominated","Fourteen-year Qubes audit finds Xen, CPU flaws dominate","Most Qubes bulletins trace to Xen, microcode, upstream","Qubes disclosures steady, but upstream loads stay high"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline upstream share rests on the premise that the hand-written, title-only attribution codebook with fixed precedence (CPU > Xen > Qubes-core > upstream) correctly identifies each bulletin's true primary component, and that the 30-bulletin post hoc audit—drawn from the same corpus the rules were designed on—proves the codebook right.","fun_headline_variants_meta":{"raw":{"variants":["Upstream bugs fuel 8 in 10 Qubes security advisories","Qubes' security record is stable—and upstream-dominated","Fourteen-year Qubes audit finds Xen, CPU flaws dominate","Most Qubes bulletins trace to Xen, microcode, upstream","Qubes disclosures steady, but upstream loads stay high"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1277,"prompt_tokens":849,"completion_tokens":428,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":342}},"tokens_in":593,"tokens_out":428,"duration_ms":4528,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T01:38:38.496259+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent coders read the full text of all 109 QSBs and assign primary components without seeing the codebook; if the resulting upstream share falls below 50%, or if a codebook-precedence reversal drops the share below a majority, the central claim is a measurement artifact. Separately, the stable-regime claim is falsified if complete 2025-2026 data produce a significantly positive post-2018 slope rather than the flat slope estimated here.","supporting_citations":[],"review_version":1}