{"id":"f1caf01f-e798-4e5a-9d00-42245338b96a","arxiv_id":"2501.10359","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A systematic NNLO data-theory comparison shows that the main PDF sets generalise to unseen LHC and HERA data about equally well once PDF, alpha_s, and missing higher order uncertainties are included.","lead":"This paper checks how well nine leading maps of the proton's internal structure, called PDFs, predict high-precision LHC and HERA measurements that were not used to build those maps. After including all theory uncertainties, the main PDF sets predict the data about equally well, and one ATLAS measurement is identified that some sets describe poorly.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Regularisation threshold Z=4 controls whether PDF differences are significant; stability to Z is not demonstrated.","rationale":"The reader identified the regularisation of ill-conditioned covariance matrices as the weakest assumption, and I agree that this is the most load-bearing aspect of the central claim. The paper's Appendix A provides Table A.1 which partially substantiates the assertion of pattern preservation: the relative ordering of PDF sets by chi2 is largely stable between unregularised and regularised values, and Table A.2 shows the covariance modifications are moderate (variances up to ~5%, correlations up to 0.06). However, the more relevant quantity for the conclusion is not just the ranking but the statistical significance of the differences, and here the regularisation has a huge effect: it compresses chi2 spreads from many sigma to sub-sigma or few-sigma levels. For example, CMS W± data would, without regularisation, discriminate between PDF sets at the 10-sigma level, which would contradict the paper's headline conclusion. The paper's defence that the unregularised values are spurious due to ill-conditioning is plausible, but the threshold Z=4 is empirical, and no stability test is provided. A concrete check is to vary Z and use alternative decorrelation models. This does not require rejecting the paper; it is a condition that can be satisfied with additional analysis. The reader's verdict of CONDITIONAL is therefore appropriate, and my concern reinforces that condition rather than changing the verdict. I set agreement_with_reader to 'partial' because my focus is not on whether regularisation distorts some PDF sets more than others (the ordering is preserved) but on whether the regularisation choice determines the statistical significance of the comparative statement.","tokens_in":57709,"tokens_out":11682,"duration_ms":116991,"concrete_test":"Recompute chi2_exp+th and the Delta n_sigma estimator (Eq. 3.18) for all datasets with ZL>4 (CMS W±, ATLAS/CMS single-inclusive jets, ATLAS dijets, H1 low-Q2 jets) using clipping thresholds Z = 2, 3, 5, and 8, and also using an independent decorrelation model such as the one applied to jet data in Ref. [111]. If the maximal |Delta n_sigma| between CT18, MSHT20, NNPDF4.0, and PDF4LHC21 exceeds about 2 for any dataset at any Z, or if the PDF ranking changes materially, the 'comparable predictive power' claim is not robust to the regularisation prescription; if the conclusion is stable across thresholds and models, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of comparable predictive power rests on chi2 values computed after singular-value clipping of ill-conditioned experimental covariance matrices (Sect. 3.2, threshold Z=4). For the most precise datasets (CMS W±, ATLAS/CMS jets, ATLAS dijets, H1 low-Q2), Table A.1 shows this regularisation changes chi2 by many sigma: e.g., CMS W+ unregularised spread across all PDF sets is 10.2–14.6 (about 13σ), while the regularised spread is 0.85–1.56 (about 2σ). Thus the statement that no PDF set is significantly better on these datasets is determined by the choice Z=4. The paper asserts (Appendix A) that regularisation does not alter the relative pattern of chi2 across PDF sets, and the ranking of best/worst PDFs is indeed roughly preserved in Table A.1. However, the magnitudes of the differences—and hence their statistical significance—are drastically reduced. If a different clipping threshold (e.g., Z=3 or Z=8) or an alternative physically motivated decorrelation model (e.g., Ref. [111] for jets) were used, the PDF differences on these high-precision datasets could become significant, overturning the conclusion that the PDF sets generalise similarly well. The paper does not test the robustness of its conclusions to this regularisation choice, so the load-bearing assumption is the invariance of the comparative ranking and its significance to the regularisation prescription.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks nine modern PDF sets (ABMP16, CT18/CT18A/CT18Z, MSHT20, NNPDF3.1, NNPDF4.0, PDF4LHC15, PDF4LHC21) against high-precision LHC Run II and HERA datasets for Drell-Yan, top-quark pair, inclusive jet, dijet, and HERA jet production. Theoretical predictions are computed at NNLO QCD using PineAPPL interpolation grids, and the data-theory agreement is quantified through reduced chi-square statistics that include experimental, PDF, alpha_s, and missing-higher-order uncertainties. The central claim is that, once all theory uncertainties are properly accounted for, the CT18, MSHT20, NNPDF4.0, and PDF4LHC21 sets provide a comparable description of the data and generalise similarly well to unseen measurements.","tokens_in":57951,"tokens_out":8839,"duration_ms":86757,"significance":"If the conclusions hold, this is a useful and timely quantitative comparison. The paper provides a transparent methodology, makes the NNLO interpolation grids publicly available, and explicitly documents the effect of covariance-matrix regularisation in Appendix A. The most nontrivial finding, that NNPDF4.0 with its markedly smaller PDF uncertainties does not describe the data worse on average than less precise sets, is an important input for the PDF-user community. The study of the ATLAS 8 TeV Z rapidity measurement, including dedicated refits and comparisons with earlier versions of the data, is careful and informative. The main caveat is that the central conclusion of 'comparable predictive power' rests on the regularisation of ill-conditioned experimental covariance matrices, whose threshold is not varied or validated.","major_comments":[{"comment":"The conclusion that no PDF set is significantly better on the CMS W and LHC jet datasets is not robustly established because it depends on the singular-value clipping threshold Z=4. Table A.1 shows that regularisation changes chi2_exp+th for CMS W+ from an unregularised spread of 10.2-14.6 (about 13 sigma) to a regularised spread of 0.85-1.56 (about 2 sigma), with analogous reductions for the jet datasets. The paper asserts in Appendix A that the regularisation does not alter the relative pattern of chi2 values across PDF sets, but it does not demonstrate that the statistical significance of the PDF-to-PDF differences is stable to the choice of Z or to alternative physically motivated decorrelation models (e.g., the model of Ref. [111] for jets). Please provide a scan over the clipping threshold (for instance Z=3, 4, 6, 8) or an alternative decorrelation prescription, and show how the Delta-n_sigma estimators and the final conclusions change.","section":"Sect. 3.2, Appendix A"},{"comment":"The 7-point MHO prescription is mis-specified: the definition of Delta0- is repeated with the same arguments (1,1/2), so the independent variation (mu_R=1, mu_F=2) is missing. Since the MHO covariance matrix Eq. (3.5) is a key ingredient of the theory covariance matrix that drives the central conclusion, please correct this definition. If the numerical results were obtained with the correct set of scale variations, state so explicitly; if they were obtained using Eq. (3.7) as written, the affected chi2 values must be recomputed and the conclusions checked.","section":"Eq. (3.7)"},{"comment":"The estimator Delta-n_sigma uses sqrt(2/ndat) as the standard deviation of the spread of chi2 values across PDF sets. This is the expected fluctuation of a single chi2 statistic under the null hypothesis, but the chi2 values for different PDF sets are not independent: they are computed from the same data and from theory predictions that are correlated through the underlying PDFs and the same theoretical framework. The statement that PDF-to-PDF differences are 'almost always within Delta-n_sigma = 1' is therefore not a statistically rigorous significance statement. Please replace this heuristic with a more appropriate test, for example by constructing the distribution of chi2 differences under a bootstrap or by using the covariance of theory predictions across PDF sets.","section":"Eq. (3.18), Sect. 4.6"}],"minor_comments":[{"comment":"The number of data points for LHCb 13 TeV Z is listed as 17 in Table 3.1 but as 18 in Tables 2.1 and 4.1; similarly, the H1 low-Q2 single-inclusive jet and dijet datasets are listed with 48 points in Table 3.1 but with 37 points in Tables 2.1 and 4.6. Please correct these inconsistencies.","section":"Table 3.1"},{"comment":"The caption states sqrt(2/ndat)=0.23 for the CMS W datasets with ndat=18; the correct value is sqrt(2/18)=0.33, as stated in Table 4.1.","section":"Fig. 4.1 caption"},{"comment":"The sentence in Sect. 3.2 that the regularisation 'does not alter our judgement' on the relative ability of PDF sets is an assertion rather than a demonstrated result; after the requested stability test is added, this statement should be either substantiated or qualified.","section":"Sect. 3.2"},{"comment":"The phrase 'similar predictive power' may overstate the case because including PDF uncertainties in the covariance matrix makes agreement easier for sets with large uncertainties; the paper does show chi2_exp separately, but the summary would benefit from explicitly noting that the comparable description holds only after the PDF uncertainties of each set are folded into the figure of merit.","section":"Sect. 4.4, Sect. 5"},{"comment":"For LHC and HERA jet production, the computations are at NNLO in the leading-color approximation and do not include NLO electroweak corrections; the possible impact of these missing terms on the chi2 values, especially at high pT, should be commented on explicitly, even if the effect is expected to be PDF-independent.","section":"Sect. 2.3-2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is strong in its scope and execution, but the central claim is more sensitive to the covariance regularisation than the text admits. The authors are part of the NNPDF collaboration and use their own tools (PineAPPL, the regularisation method of Ref. [109]); this is not a problem per se, but an independent check of the regularisation robustness, or at minimum a clear scan over the threshold, would substantially increase confidence in the conclusions. I would not reject the paper: the issues raised are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know this: it is the most complete test yet of how well the major PDF sets predict LHC Run II and HERA data they were not fitted to, computed at exact NNLO with PDF, alpha_s, and MHO uncertainties folded in. The headline result - CT18, MSHT20, NNPDF4.0, and PDF4LHC21 describe the unseen data comparably well once theory uncertainties are included - is, as far as I can tell, supported by the tables. That is a useful and non-obvious statement for the precision-physics community.\n\nWhat's genuinely new: the breadth of processes (DY, ttbar, jets, HERA jets) at exact NNLO, the PineAPPL grids, and the careful breakdown of chi2 into experimental, MHO, and PDF+alpha_s components. The special study of the ATLAS 8 TeV Z rapidity data is the best part. They show the new measurement is consistent with its older version but pulls against other Drell-Yan data, and that forcing it into NNPDF4.0 degrades the global fit by ~4 sigma. That is a concrete, reproducible finding.\n\nThe soft spot is exactly where the reader put it. The conclusion about 'comparable predictive power' is largely determined by the covariance regularisation at Z=4. For CMS W and the jet datasets, the unregularised chi2 values are enormous and the spread across PDF sets collapses after clipping. The paper asserts in Appendix A that regularisation does not change the relative pattern of chi2, and the ranking of best/worst sets is roughly stable, but it does not demonstrate that the significance of the PDF differences is insensitive to Z. The stress-test note is right: change the clipping threshold and the statement 'no PDF set is significantly better' could break. This is not a fatal flaw - the authors are transparent, they provide the unregularised numbers, and the regularisation is grounded in Ref. [109] - but it is a real gap. A robustness scan over Z, or an alternative decorrelation model for jets, would settle it. The LHCb correlation speculation is minor; the absence of a full code release is a minor inconvenience given grids are public.\n\nBottom line: this deserves a serious referee. I would recommend conditional acceptance, asking for the Z-robustness test. Anyone working on PDF choice for mW, sin2theta_eff, or alpha_s extractions should read it.","headline":"A comprehensive, transparent PDF benchmark whose main claim is plausible but rests on a regularisation choice that needs a robustness test.","tokens_in":58571,"tokens_out":2660,"would_cite":true,"duration_ms":25799,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that, once PDF, strong-coupling and missing-higher-order uncertainties are all included, the major PDF sets describe LHC Run II and HERA data equally well and none generalises better to unseen data.","keywords":["parton distribution functions","PDF benchmarking","NNLO QCD predictions","LHC Run II data","Drell-Yan production","top quark pair production","jet cross sections","theoretical uncertainties"],"falsifier":"Recompute the $\\chi^2_{\\rm exp+th}$ spread for the CMS W rapidity and ATLAS/CMS single-inclusive jet datasets using the experiments' own alternative correlation or decorrelation models instead of the universal $Z=4$ clipping. If any PDF set moves by more than about one standard deviation of the $\\chi^2$ distribution relative to the others, the claim that all PDF sets generalise equally well to those data would fail.","tokens_in":57483,"feed_emoji":"⚛️","tokens_out":6973,"duration_ms":62306,"temperature":0.7,"pith_summary":"The paper asks whether the leading parton distribution function (PDF) sets can be told apart by how well they predict high-precision measurements that were, with one exception, not included in their fits. It compares NNLO QCD predictions for LHC Run II Drell-Yan, top-quark pair, single-inclusive jet and di-jet cross sections, and HERA jet cross sections, using ABMP16, CT18 (and variants), MSHT20, NNPDF3.1, NNPDF4.0, PDF4LHC15 and PDF4LHC21. The central finding is that CT18, MSHT20, NNPDF4.0 and PDF4LHC21 describe all the datasets comparably well once PDF, strong-coupling and missing-higher-order uncertainties are all included in the comparison. The paper concludes that no major PDF set generalises better to unseen data. This matters because PDF choice is often the dominant uncertainty in LHC determinations of Standard Model parameters such as the W mass, the strong coupling and the weak mixing angle.","feed_headline":"LHC data cannot tell major PDF sets apart","feed_subtitle":"Counting PDF, coupling and missing-order errors, CT18, MSHT20, NNPDF4.0 and PDF4LHC21 predict unseen data equally well.","key_machinery":"The engine of the comparison is the reduced $\\chi^2$ figure of merit of Eq. (3.1), evaluated with a total covariance matrix that adds four independent sources of uncertainty: experimental, missing higher orders (estimated by 7-point renormalisation/factorisation scale variations), PDF (from Hessian eigenvectors or Monte Carlo replicas), and $\\alpha_s(m_Z)=0.118\\pm0.001$. Predictions are computed at NNLO QCD accuracy and stored as PineAPPL interpolation grids, which lets every PDF set be evaluated at no extra cost. Because several experimental covariance matrices are ill-conditioned, the paper regularises them by clipping singular values below a threshold $Z=4$, following Ref. [109]. The $\\Delta\\chi^2$ and $\\Delta n_\\sigma$ estimators then convert raw $\\chi^2$ differences into units of the expected statistical fluctuation of the $\\chi^2$ distribution.","core_discovery":"On the paper's own terms, the discovery is that the apparent differences in predictive power between modern PDF sets largely disappear when all sources of theoretical uncertainty are treated on the same footing. Computed with the $\\chi^2$ per data point and a covariance matrix that adds experimental, missing-higher-order, PDF and $\\alpha_s$ uncertainties, the CT18, MSHT20, NNPDF4.0 and PDF4LHC21 sets give statistically equivalent descriptions of every dataset examined. The one exception is the ATLAS 8 TeV inclusive Z rapidity distribution, which NNPDF4.0 describes poorly even though an earlier version of the same measurement was used in its fit; the paper traces this to a tension between that dataset and other Drell-Yan data, and shows that including it in a fit with missing-higher-order uncertainties restores an acceptable description. From this the paper concludes that all major PDF sets have similar predictive power and that the spread between sets should not be read as evidence that one set is systematically better.","pith_inferences":["Beyond the paper: the same methodology could be applied to the newer aN3LO and MHOU-equipped PDF sets to test whether comparable predictive power survives another order in perturbation theory.","Beyond the paper: the finding that PDF set choice rarely moves $\\chi^2_{\\rm exp+th}$ by more than one standard deviation suggests that future PDF benchmarking should report agreement including theory covariance matrices as standard, rather than comparing experimental $\\chi^2$ only.","Beyond the paper: a direct test of the weakest assumption would be to recompute the CMS W and jet conclusions using experiment-specific decorrelation models; if the regularisation changes relative PDF rankings in any of those datasets, the comparable-predictive-power claim would need to be restricted to the remaining datasets.","Beyond the paper: if the conclusion holds, the spread between PDF sets quoted in $\\alpha_s(m_Z)$, $m_W$, and $\\sin^2\\theta_{\\rm eff}$ analyses should be interpreted as methodological spread rather than as evidence that one set is wrong."],"forward_implications":["If the claim holds, precision Standard Model measurements at the LHC do not need to identify a single best PDF set; any of the major sets yields equally reliable predictions once its uncertainties are propagated.","A PDF set with small uncertainties, such as NNPDF4.0, does not thereby predict unseen data better; precision and predictive power are separate properties.","The ATLAS 8 TeV Z rapidity measurement, which underlies a precise $\\alpha_s(m_Z)$ extraction, is in tension with other Drell-Yan data in NNPDF4.0, so PDF uncertainties quoted from a single baseline set may understate the spread.","LHC jet and top-quark pair measurements, not HERA jets or Drell-Yan rapidity shapes, are the datasets most able to discriminate PDF sets and to constrain the large-$x$ gluon."],"supporting_citations":[{"why":"Supplies the CT18/CT18A/CT18Z PDF sets whose predictions are compared and whose uncertainties enter the covariance matrix.","marker":"[6]"},{"why":"Supplies the MSHT20 PDF set, one of the four central sets in the comparable-predictive-power conclusion.","marker":"[7]"},{"why":"Supplies the NNPDF4.0 PDF set, the main precision/accuracy case study and the exception analysis.","marker":"[8]"},{"why":"Supplies the PDF4LHC21 combination set, whose near-zero $\\Delta\\chi^2$ behaviour is explained by construction as the average of other sets.","marker":"[26]"},{"why":"Provides the formalism for including missing-higher-order, PDF and $\\alpha_s$ uncertainties in the covariance matrix.","marker":"[105]"},{"why":"Provides the theory-uncertainty formalism variant and the MHOU-equipped NNPDF4.0 sets used in the ATLAS 8 TeV Z investigation.","marker":"[106]"},{"why":"Provides the ill-conditioning diagnostic ($Z$ threshold) and the singular-value clipping regularisation on which the CMS W and jet conclusions depend.","marker":"[109]"},{"why":"Provides the PineAPPL interpolation grids that allow fast NNLO predictions for many PDF sets and scale variations.","marker":"[69]"},{"why":"The ATLAS 8 TeV Z rapidity measurement that is the single exception and the subject of the dedicated tension analysis.","marker":"[42]"}],"fun_headline_variants":["PDF sets tied in LHC precision test","No best PDF set: all match LHC data equally","Uncertainties erase PDF differences at LHC","Modern PDFs show equal predictive power at LHC","LHC data can't pick a favorite PDF set"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that cutting the smallest singular values of the experimental covariance matrices (the same threshold $Z=4$ for all datasets) does not change which PDF set looks best, even though for the CMS W and LHC jet data this procedure shifts the $\\chi^2$ by several standard deviations, and the paper asserts this rather than demonstrating it.","fun_headline_variants_meta":{"raw":{"variants":["PDF sets tied in LHC precision test","No best PDF set: all match LHC data equally","Uncertainties erase PDF differences at LHC","Modern PDFs show equal predictive power at LHC","LHC data can't pick a favorite PDF set"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000501,"raw_usage":{"total_tokens":2433,"prompt_tokens":912,"completion_tokens":1521,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":1447}},"tokens_in":528,"tokens_out":1521,"duration_ms":11039,"temperature":1.0,"reasoning_tokens":1447,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:09:50.636755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the $\\chi^2_{\\rm exp+th}$ spread for the CMS W rapidity and ATLAS/CMS single-inclusive jet datasets using the experiments' own alternative correlation or decorrelation models instead of the universal $Z=4$ clipping. If any PDF set moves by more than about one standard deviation of the $\\chi^2$ distribution relative to the others, the claim that all PDF sets generalise equally well to those data would fail.","supporting_citations":[{"cited_title":"Regularising experimental correlations in LHC data: theory and application to a global analysis of parton distributions","cited_arxiv_id":"2207.00690","evidence_quote":"Provides the ill-conditioning diagnostic ($Z$ threshold) and the singular-value clipping regularisation on which the CMS W and jet conclusions depend."}],"review_version":1}