{"id":"ba577938-2e62-4001-8a91-cbe860176a3a","arxiv_id":"1909.00001","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"New CT18-family NNLO PDFs are presented, with a fitted x-dependent factorization scale that lowers the HERA inclusive DIS chi-square by more than 50 units and mimics some low-x resummation effects.","lead":"The CTEQ-TEA collaboration presents four new NNLO parton distribution function families (CT18, CT18A, CT18X, CT18Z) from a global QCD fit to HERA, Tevatron, and LHC data, plus an x-dependent scale that improves the HERA fit. Generalists should care because these PDFs are standard inputs for LHC predictions, but this is a proceedings summary whose main results are validated in a companion paper.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The x-dependent scale is tuned to HERA chi-square, so the claim that it mimics resummation needs an out-of-sample check before CT18Z's gluon can be trusted.","rationale":"The reader identified the same weakest assumption: the x-dependent scale is tuned to the data whose improvement is reported, so the resummation-mimicry claim lacks an out-of-sample check. The proceedings text is a report on a larger program (CT18, companion paper, future code release), so ACCEPT is premature and REJECT is too strong. The paper is transparent about the tuning, and the CTEQ-TEA framework is credible, including published companion work and cross-validated APPLgrid tables. However, the quantitative claim that the scale absorbs the small-x logarithms that resummation treats explicitly is not established by an in-sample chi-square reduction: with three fitted coefficients and no uncertainty or cross-validation, the >50 unit improvement could be overfitting. The suggested hold-out test directly settles this. Thus the verdict CONDITIONAL is appropriate, and the condition should be that the mu_F,x scale survives an out-of-sample or independent-code test and that the CT18Z grids/code be made public.","tokens_in":5881,"tokens_out":1634,"duration_ms":14531,"concrete_test":"Reproduce the CT18Z fit with the x-dependent scale but hold out one high-precision kinematic region, e.g., fit without H1 FL and without Q < 3.5 GeV HERA inclusive data, then compute the chi-square of the held-out data using the refitted mu_F,x coefficients. If the held-out chi-square does not improve (or worsens) relative to the nominal-scale fit, the >50 unit HERA improvement is in-sample overfitting and the resummation-mimicry claim fails. Alternatively, release the CT18Z fit code and grids so an independent group can recompute the HERA chi-square with the stated scale, and additionally compare CT18Z predictions for yet-unmeasured LHC observables (e.g., forward D-meson or Z production) to assess out-of-sample behavior.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract and body claim that using mu_F,x^2 = 0.82(Q^2 + 0.3 GeV^2/x^0.3) reduces HERA I+II inclusive DIS chi-square by >50 units at NNLO, matching the improvement attributed to low-x resummation. The load-bearing problem is that the coefficients 0.82 and 0.3 and the exponent 0.3 are chosen to minimize chi-square on the very same HERA data whose improvement is reported ('the numerical coefficients in mu_F,x are chosen to minimize chi-square for the HERA DIS data'). With three parameters free, an in-sample chi-square gain of >50 units on a large high-precision data set is not evidence of physics equivalence to resummation; it is overfitting unless the same scale also improves genuine predictions for out-of-sample kinematics (e.g., predictions for the HERA data at different Q, x or for LHC observables sensitive to the small-x gluon). The paper further states that CT18X/Z alter the gluon and quark PDFs at small x (Figure 1), yet no uncertainty or goodness-of-fit information is given to show the change is demanded by data rather than an artifact of an ad hoc scale. The distinction matters because the companion CT18 paper presumably uses the nominal scale; if the mu_F,x fit is discarded, the CT18Z gluon and the 1% Higgs cross-section reduction rest on an in-sample tuning. This is not an internal inconsistency and the paper is transparent about the tuning; but 'comparable quality of improvement' is not established out of sample, so the central claim of resummation mimicry is conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings contribution from the CTEQ-TEA collaboration reports on the implementation of HERA I+II and LHC data in the new CT18 global QCD analysis. The paper presents four NNLO (and corresponding NLO) PDF families: CT18, CT18A, CT18X, and CT18Z, which differ in data selection and in the factorization scale used for DIS. The main physics claim is that using an x-dependent factorization scale mu_F,x^2 = 0.82(Q^2 + 0.3 GeV^2/x^0.3) in fixed-order NNLO DIS calculations reduces the HERA I+II inclusive DIS chi-square by more than 50 units, an improvement described as comparable to that obtained with low-x resummation in Refs. [6,7]. The paper also discusses tensions among HERA, Tevatron, and LHC data sets, the use of PDFSense and ePump for data selection, and the impact of the alternative scales on PDFs and on the NNLO Higgs cross section.","tokens_in":6221,"tokens_out":3575,"duration_ms":40297,"significance":"If the central claim is correct, CT18Z and its siblings provide a systematic family of NNLO PDFs that span a range of plausible scale and data-selection choices, and the x-dependent scale idea could offer a practical fixed-order alternative to small-x resummation for global fits. The paper is valuable for its candid discussion of data tensions, its explicit statement that the x-dependent scale coefficients are tuned to minimize the HERA chi-square, and its use of modern fast-interfacing and reweighting tools. However, the key quantitative evidence for the resummation-mimicry claim is currently in-sample: the scale is fitted to the same HERA data whose chi-square improvement is reported, and no out-of-sample test or uncertainty analysis is provided. The significance of CT18X/Z therefore rests on validation that is not presented in this manuscript.","major_comments":[{"comment":"The sentence defining mu_F,x^2 states that the numerical coefficients are chosen to minimize chi-square for the HERA DIS data. The subsequent claim that this scale produces a reduction of more than 50 chi-square units and a 'comparable quality of improvement' to low-x resummation is therefore an in-sample fit statistic, not independent evidence of physical equivalence. The manuscript should provide an out-of-sample validation: for example, fix the coefficients using a training subset of HERA data or a HERA-only fit, then evaluate the held-out HERA bins and LHC observables sensitive to the small-x gluon. The improved chi-square reported for charm SIDIS and H1 FL in Fig. 1(right) does not resolve this issue because these data sets are part of the same fit.","section":"Combined HERA I+II DIS data and an x-dependent factorization scale"},{"comment":"Figure 1(left) shows ratios of CT18 PDFs obtained with the x-dependent and standard scales, but no PDF uncertainty bands or chi-square decomposition are provided. Without error bands and per-data-set chi-square values, the reader cannot determine whether the enhanced small-x gluon and altered quark distributions are statistically required or are an artifact of the tuned scale. The manuscript should show the uncertainties of CT18 and CT18Z and report the chi-square per data set for both scale choices.","section":"Figure 1"},{"comment":"The comparison with low-x resummation in Refs. [6,7] is not apples-to-apples as presented: the data selections, kinematic cuts, chi-square definitions, and perturbative orders differ between the CT18 fit and the resummed fits. The claim that the improvement is 'comparable' requires a table specifying the data sets, Q and x cuts, and chi-square definitions used in each comparison. Without this, the >50 unit improvement cannot be interpreted as equivalent to the resummation effect.","section":"Combined HERA I+II DIS data and an x-dependent factorization scale"},{"comment":"The abstract and body present CT18, A, X, and Z as four new PDF families, but this proceedings contribution does not provide the defining criteria for each family beyond brief descriptions: no LHAPDF names, no complete input data lists, no central chi-square values, and no error sets are given. Since these definitions are load-bearing for the paper's central claim, the manuscript should either include a compact table summarizing the data sets and scale choices for each family or explicitly refer to a publicly available companion document where these details are given.","section":"Selection of new LHC experiments"}],"minor_comments":[{"comment":"There are typographical errors in this section: 'thee+p and e−p' should be 'the e+p and e−p', 'with e HERA' should be 'with HERA', and 'multi-prone' should be 'multi-pronged'.","section":"Combined HERA I+II DIS data and an x-dependent factorization scale"},{"comment":"The left panel's caption contains a garbled formula, 'a (x,Q)/f (2)a f', which appears to be an attempt to write the ratio of PDFs; it should be cleaned up to something like f_a^{(1)}(x,Q)/f_a^{(2)}(x,Q).","section":"Figure 1"},{"comment":"The phrase 'unconstrained region x > 0.5' is imprecise: the valence quark region is constrained by sum rules, so it would be clearer to say 'a region less tightly constrained by the fitted data' or to specify what is meant by 'unconstrained'.","section":"Combined HERA I+II DIS data and an x-dependent factorization scale"},{"comment":"The caption labels 'CT18 NNLO' and 'CT18Z NNLO' but the text refers to 'CT14HERA2' in the upper panel; the notation and the relationship between the two panels should be clarified.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings-style contribution, so some technical details are expected to reside in the companion CT18 paper. The main referee concern—the in-sample tuning of the x-dependent scale—should be addressed either with an explicit out-of-sample test in this paper or by a clear reference to such a test in the companion publication. If the companion paper already contains this validation, a brief statement and citation would be sufficient; if not, the claim of 'comparable quality of improvement' to resummation should be softened until such validation exists."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this is a competent proceedings writeup of the CT18 family, and the four new PDF ensembles are real new objects. But the claim that the x-dependent scale reproduces the effect of low-x resummation is not backed by an out-of-sample check, because the scale coefficients were tuned to the same HERA data whose chi-square improvement is quoted. That specific claim should be treated as conditional until tested on unmeasured kinematics or on LHC observables.\n\nWhat's genuinely good: the paper is transparent about tensions. The SE distribution figure and the discussion of why the ATLAS W/Z data elevate the dimuon chi-square is honest. The use of PDFSense and ePump to preselect data, the cross-validation of APPLgrid tables, and the acknowledgment that the CDHSW data were removed because they prefer a harder gluon are all the right kind of methodological candor. The paper also does a service by pointing to the companion full CT18 paper for the actual fit details.\n\nThe soft spots, in order. First, the x-dependent scale. The text says explicitly the coefficients 0.82, 0.3 GeV^2, and 0.3 are \"chosen to minimize chi-square for the HERA DIS data.\" With three parameters free, an in-sample improvement of more than 50 chi-square units on a large correlated data set is exactly what you'd expect from tuning; it is not evidence that the scale captures the physics of resummation. The comparison to resummed fits is also indirect: they cite the resummation papers but don't show a quantitative side-by-side. Second, the PDF changes at small x from the scale choice have no uncertainty bands shown, so it's unclear whether the gluon shift is demanded by data or is an artifact of the ansatz. That's a presentation gap, minor in a proceedings but relevant to anyone who wants to use CT18Z. Third, the global fit itself has the usual selection-after-inspection issue: data sets are added or removed after looking at fit quality. They are open about it, and it's standard practice, but it means the quoted chi-squares are not a rigorous measure of global fitness.\n\nNone of this sinks the paper. The CT18 family is a credible incremental update of an established program, and the four ensembles are useful tools for exploring scale and data-selection sensitivity. The resummation-mimicry claim is the part I'd want to see pinned down with a genuine prediction before I'd bank on it.\n\nWho gets value: anyone doing NNLO phenomenology at the LHC who wants up-to-date PDFs with controlled variations; PDF fitters will go to the companion paper. I'd take this to my reading group and cite it, but I'd also flag the scale-tuning issue to my students.\n\nRecommendation: yes, send it to peer review. It's a legitimate contribution, and referees should push for an out-of-sample check of the x-dependent scale.","headline":"A solid proceedings writeup of the CT18 PDF family, but the resummation-mimicry claim rests on a scale tuned in-sample and needs an out-of-sample check.","tokens_in":6813,"tokens_out":2670,"would_cite":true,"duration_ms":26060,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CT18Z global QCD analysis shows that a tuned x-dependent factorization scale, $\\mu_{F,x}^2=0.82(Q^2+0.3\\,\\mathrm{GeV}^2/x^{0.3})$, cuts the NNLO HERA inclusive-DIS $\\chi^2$ by more than 50 units and reproduces the improvement…","keywords":["CT18 PDFs","parton distribution functions","NNLO QCD global analysis","HERA I+II inclusive DIS","x-dependent factorization scale","low-x resummation","LHC data","strangeness PDF"],"falsifier":"Take the fitted $\\mu_{F,x}$ scale into a predictor outside the fitted HERA kinematics, for example forward charm or Drell-Yan production at the LHC at small $x$, or HERA data with $Q$ below 2 GeV, and compare CT18Z predictions with data and with a resummed PDF set. If the CT18Z predictions agree with data at resummation quality, the scale is a valid proxy; if the >50-unit improvement disappears or the predictions deviate once kinematics change, the ansatz is tuned rather than physical.","tokens_in":5674,"feed_emoji":"📐","tokens_out":8045,"duration_ms":76952,"temperature":0.7,"pith_summary":"The paper presents four new families of next-to-next-to-leading-order parton distribution functions---CT18, CT18A, CT18X, and CT18Z---produced in a global QCD analysis that adds eleven LHC data sets. Its central claim is that a fitted x-dependent factorization scale, $\\mu_{F,x}^2 = 0.82(Q^2 + 0.3\\,\\mathrm{GeV}^2/x^{0.3})$, lowers the NNLO $\\chi^2$ of the combined HERA I+II inclusive deep-inelastic-scattering data by more than 50 units, matching the gain previously attributed to low-$x$ resummation. A sympathetic reader would care because the scale choice changes the small-$x$ gluon and strangeness PDFs and therefore shifts NNLO predictions for LHC observables, while also offering a fixed-order alternative to resummation. The four ensembles are designed to bracket the PDF uncertainty coming from data selection and scale choice rather than to give a single answer.","feed_headline":"A tuned scale cuts HERA chi-square by 50 units","feed_subtitle":"CT18Z PDFs absorb small-x logarithms with one x-dependent scale, matching resummation quality.","key_machinery":"The load-bearing object is the x-dependent factorization scale $\\mu_{F,x}^2 = 0.82(Q^2+0.3\\,\\mathrm{GeV}^2/x^{0.3})$, a saturation-inspired one-scale ansatz whose coefficients were chosen to minimize the HERA DIS $\\chi^2$. At small $x$, the added $0.3/x^{0.3}$ term raises the scale above $Q^2$, effectively absorbing a class of enhanced power logarithms that would otherwise be treated by low-$x$ resummation; at larger $x$ it reduces to a constant factor $0.82$ times $Q^2$. The same machinery also includes fast impact-analysis programs to select eleven new LHC data sets and a parallelized fitting code with fast grid interfaces, but the scale ansatz is what carries the central claim about matching resummation.","core_discovery":"On its own terms, the paper establishes that the conventional fixed-order NNLO description of HERA data can be substantially improved without resummation by evaluating DIS cross sections with the tuned scale $\\mu_{F,x}^2=0.82(Q^2+0.3\\,\\mathrm{GeV}^2/x^{0.3})$ instead of $\\mu_F^2=Q^2$. In the region $Q>2$ GeV, $x>10^{-5}$, this reduces $\\chi^2(\\mathrm{HERA\\ I+II})$ by more than 50 units, a quality of improvement the paper describes as comparable to that reported by resummed analyses. The resulting CT18X and CT18Z fits show reduced $u$ and $d$ (anti-)quark PDFs and increased gluon and strangeness at $x<10^{-2}$, with compensating changes at $x>0.5$ to preserve valence and momentum sum rules. In CT18Z, removal of CDHSW nuclear data and inclusion of ATLAS 7 TeV W/Z data combine to reduce the NNLO gluon-fusion Higgs cross section by about 1% relative to CT14 and CT18.","pith_inferences":["A natural extension not drawn by the paper: the same scale ansatz could be lifted into nuclear PDF fits or future electron-ion collider projections, where the small-$x$ regime is scanned directly; if it holds there, the saturation-inspired wording gains predictive content.","An explicit overlap test would fit with both the tuned scale and a resummed theory simultaneously and see whether the scale loses its benefit; if it does, the two mechanisms are not independent alternatives but overlapping ways of absorbing the same logarithms.","The tuned coefficients 0.82 and 0.3 depend on the data set; one could attempt to derive them from a first-principles saturation momentum and re-fit newer datasets to see whether the same functional form recurs, rather than being a one-off adjustment."],"forward_implications":["At NNLO, the small-$x$ behavior of the gluon and strangeness PDFs in CT18X/Z differs from CT18; quantities such as the gluon-fusion Higgs cross section shift by about 1%, altering the central value that precision electroweak and Higgs analyses compare with.","The reported >50-unit $\\chi^2$ drop means the combined HERA data do not, by themselves, force low-$x$ resummation: a fixed-order fit with the tuned scale is an equally good description, so resummed PDFs and scale-shifted PDFs should both be used for uncertainty bands.","Because CT18, A, X, and Z are generated under different data-selection and scale assumptions, the spread among them gives a practical systematic uncertainty on NNLO predictions for LHC processes beyond the one-sigma Hessian error.","Including the ATLAS 7 TeV W/Z rapidity data raises the tension with the strangeness-sensitive dimuon data from CCFR and NuTeV; the elevated goodness-of-fit variables in CT18Z show that the strangeness PDF is constrained by competing data sets, not uniquely by either."],"supporting_citations":[{"why":"Supplies the combined HERA I+II inclusive DIS data whose NNLO fit quality is the paper's central testbed.","marker":"[3]"},{"why":"Provides the low-x resummed fit that sets the benchmark: the paper's scale choice reproduces a comparable chi-square improvement over fixed-order NNLO.","marker":"[6]"},{"why":"Corroborates the resummed benchmark through an independent low-x resummation analysis of HERA data.","marker":"[7]"},{"why":"Supplies the qualitative saturation-inspired argument that motivates the functional form of the x-dependent scale.","marker":"[8]"},{"why":"Inclusion of the ATLAS 7 TeV W/Z rapidity data drives the CT18A/Z variants and raises the reported strangeness tension.","marker":"[12]"},{"why":"Describes the impact-analysis program used to predict which data sets will constrain the fit, guiding selection of the eleven new LHC sets.","marker":"[14]"},{"why":"Describes the reweighting program used to quickly estimate the impact of candidate LHC data before the global fit.","marker":"[15]"}],"fun_headline_variants":["Tuned scale cuts HERA chi-square by 50","CT18Z: scale shift beats resummation","HERA fit improved 50 chi-square with one scale","New CT18Z PDFs: smarter scale, better HERA fit","Scale trick slashes HERA chi-square by 50"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The premise is that a scale whose two coefficients were tuned to the HERA data can genuinely reproduce the small-$x$ logarithms that resummation computes explicitly, so its benefit should persist for unmeasured kinematics rather than being an in-sample artifact.","fun_headline_variants_meta":{"raw":{"variants":["Tuned scale cuts HERA chi-square by 50","CT18Z: scale shift beats resummation","HERA fit improved 50 chi-square with one scale","New CT18Z PDFs: smarter scale, better HERA fit","Scale trick slashes HERA chi-square by 50"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1279,"prompt_tokens":865,"completion_tokens":414,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":331}},"tokens_in":481,"tokens_out":414,"duration_ms":4235,"temperature":1.0,"reasoning_tokens":331,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:15:23.892995+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the fitted $\\mu_{F,x}$ scale into a predictor outside the fitted HERA kinematics, for example forward charm or Drell-Yan production at the LHC at small $x$, or HERA data with $Q$ below 2 GeV, and compare CT18Z predictions with data and with a resummed PDF set. If the CT18Z predictions agree with data at resummation quality, the scale is a valid proxy; if the >50-unit improvement disappears or the predictions deviate once kinematics change, the ansatz is tuned rather than physical.","supporting_citations":[],"review_version":1}