{"id":"87f2349a-9678-4a96-b13c-b8e380cf91de","arxiv_id":"2504.13289","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Symbolic regression fits to lattice QCD and model GPDs show that the isovector GPD H_{u-d} approximately factorizes in x and t in the trained kinematic region, and a new Taylor-coefficient clustering criterion (ECC) groups the many fits into a few stable solution families.","lead":"This paper uses symbolic regression to fit the proton's quark spatial-distribution function (GPD) to lattice QCD data and three models, and finds that in the measured range the x and momentum-transfer dependence approximately factorize. It also introduces a clustering method, ECC, to assess when such symbolic fits are converging.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Low-x localization of factorization breaking is not established by the integrated moment ratios; A30/A10 emphasizes high x, so the Sec. IV C conclusion overreaches.","rationale":"The reader's weakest assumption was generic extrapolation beyond the 13×5 grid; I agree that extrapolation is the danger zone, but the sharper, testable flaw is the inference that untrained factorization breaking is located at low x. Section IV C uses full-x lattice moments to draw that localization. Because A30 weights x², the observed A30/A10 t-dependence, if real, points at least as much to high-x as to low-x breaking, and no x-localized statistic is provided. This is a correctness risk internal to the argument, not a dispute over physics consensus. I do not object to the ECC method or to the supported in-region factorization result. I also note a secondary pathology worth checking separately: the FF exemplar in Eq. (38) has a denominator 1.67 t - 0.638 with a zero at |t|≈0.382 GeV², inside the training region, so the Fourier-transform densities in Figs. 18–20 may be ill-defined for that exemplar; this reinforces a conditional verdict but is not the main logical flaw. The reader's CONDITIONAL verdict remains appropriate: the methodological contribution stands, but the low-x localization claim needs an explicit test or should be softened to 'breaking outside the trained region.'","tokens_in":28317,"tokens_out":12228,"duration_ms":117639,"concrete_test":"Fit to the lattice moment data of Refs. [54–56] a model composed of the factorized trained-region form times a non-factorized correction supported only on x<0.255 (e.g., a Regge-inspired x^{α(t)} factor), and check whether the resulting A20/A10 and A30/A10 reproduce the observed t-dependence. If a low-x-only correction cannot reproduce the A30/A10 trend without additional breaking at x>0.255, then the 'entirely focused in the low x region' conclusion is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"At the end of Sec. IV C the paper concludes that LQCD data allow a breaking of factorization 'entirely focused in the low x region.' The in-region support for approximate factorization (0.255≤x≤0.855, |t|≤1 GeV²) is reasonable: the fixed-t ratios in Fig. 13d are flat at the ~10% level, and the BF/FF MSE comparison in Fig. 16 is consistent with that. The problematic step is the localization. It is inferred from the t-dependence of the full-x moment ratios A20/A10 and A30/A10 in Fig. 17. These integrals do not localize the breaking in x; moreover A30 = ∫ x² H dx weights large x more heavily than small x, so a given t-dependence in A30/A10 is more readily produced by breaking at moderate/high x than by breaking confined to x<0.255. The displayed data are equally compatible with breaking at x>0.855 or with breaking at |t|>1 GeV² for all x, and the paper performs no test that isolates the x location of the breaking. Since the subsequent bRMS(x) interpretation and the comparison with Reggeized GGL/GK models rest on this localization, the headline physics conclusion is not established by the presented evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript applies the symbolic regression package PySR to the lattice QCD results for the isovector GPD H_{u-d}(x,t,ξ=0) from Ref. [29], with the goal of extracting interpretable analytic expressions and using them to infer physics. The authors introduce an 'Expansion Coefficient Clustering' (ECC) criterion to assess the consistency of the many symbolic solutions produced by independent SR runs. They compare three selection criteria: unconstrained Best-Fit (BF), Force-Factorized (FF), and a Semi-Reggeized form. They test the factorization of x and t dependence using fixed-t ratios, MSE comparisons between BF and FF fits, KL divergences of the MSE distributions, and moment ratios A20/A10, A30/A10. The paper claims that, in the trained region (0.255≤x≤0.855, |t|≤1 GeV²), the LQCD data approximately factorize, and that any factorization breaking is confined to low x. It then uses the fitted forms to compute transverse densities ρ(x,b_T) and the average transverse radius b_RMS(x), and compares these with phenomenological models. The central physics conclusions are that LQCD data favor approximate factorization in the measured region, and that the extracted radii show model-dependent trends in the extrapolation region.","tokens_in":28630,"tokens_out":5635,"duration_ms":52242,"significance":"The paper is a serious methodological contribution: it brings symbolic regression with systematic replica analysis to GPD phenomenology, and it proposes a concrete criterion (ECC) for assessing the consistency of SR solutions. The inventory of hyperparameters in Appendix B and the explicit discussion of unaccounted LQCD correlations are useful for reproducibility. If the factorization claims survive scrutiny, the result that the LQCD isovector GPD approximately factorizes in the measured x,t region would be a notable input to GPD phenomenology and to the design of future extractions. The extrapolated densities and radii are potentially interesting, but their credibility hinges on the validity of extrapolation far outside the training grid, which is not established by the presented evidence. Overall, the in-region factorization claim is defensible, but the paper's headline conclusions about low-x localization and the spatial radii overreach what the current tests demonstrate.","major_comments":[{"comment":"The conclusion that 'LQCD data allow for a breaking of factorization in the x,t behavior of GPDs that is entirely focused in the low x region' is not supported by the analysis presented. The evidence cited for this localization is the t-dependence of the integrated moment ratios A20/A10 and A30/A10 in Fig. 17. These are integrals over the full x range, and A30 weights large x more heavily than small x. A non-constant A30/A10 could equally arise from factorization breaking at moderate or high x, or even from effects at |t|>1 GeV², since the integrals receive contributions from all x and t. The paper performs no test that isolates the x location of the breaking, so the phrase 'entirely focused in the low x region' is not a demonstrated result. This matters because the subsequent interpretation of b_RMS(x) and the comparison with Reggeized GGL/GK models rest on this localization. I recommend either removing the localization claim or adding a direct x-resolved test, such as evaluating the fixed-t ratio R(x,|t|) separately in x bins and testing whether deviations from unity occur only for x<0.255.","section":"Sec. IV C"},{"comment":"The extraction of the transverse densities ρ(x,b_T) and the average radii b_RMS(x) treats the symbolic expressions as reliable predictions outside the training region, particularly at x<0.255 and |t|>1 GeV². This is an extrapolation from a 13×5 lattice grid, and the ECC clusters diverge precisely in this extrapolation region, as shown in Figs. 6 and 10. The claim that 'all four of these distributions are a prediction of the combined SR & LQCD framework' (Fig. 20 caption) overstates the status of these curves: they are model-dependent extrapolations with no physical constraints (e.g., positivity, matching to PDFs at x→0 and x→1, or Regge behavior at small x) imposed in the BF fits. In particular, the upward turn of BF Cluster 3 as x→1 would, if taken at face value, conflict with the QCD expectation of point-like configurations, but this behavior occurs in the region where no data exist and where the different clusters disagree. I recommend reframing these as illustrative extrapolations, or adding validation checks such as training on a subset of x and testing on the held-out region, or imposing physical endpoint constraints in the loss function.","section":"Sec. IV D"},{"comment":"The KL-divergence comparison KL(MSE_BF||MSE_FF) lacks a BF-versus-BF baseline. The authors themselves note (Sec. IV C) that 'we are unable to establish hierarchy of source factorization breaking' and propose generating two sets of BF replicas as a baseline in future work. Without this baseline, the numerical values in Table II cannot be interpreted as an absolute measure of compatibility with factorization; a large KL value could also arise from differences in the convergence of BF versus FF runs that are unrelated to the factorization hypothesis. The paper does correctly use the overlapping histograms in Fig. 16 to support the in-region factorization claim for VGG and LQCD, and the strong separation for GGL and GK. My concern is only with the quantitative use of the KL values themselves; I recommend presenting the KL numbers as qualitative indicators and adding the BF-baseline comparison before drawing any conclusion about the relative degree of factorization breaking.","section":"Sec. IV C"},{"comment":"The treatment of uncertainties is insufficient for the strength of the factorization claim. The paper states that 'the LQCD errors are highly correlated but unaccounted for' and that the MSE and WMSE 'do not carry statistical significance.' Yet the conclusion that LQCD factorizes 'to within ≈10%' in the training region is stated without an uncertainty on that percentage. The 10% figure is a scatter over the fixed-t ratio points, not a confidence interval. Since the central physics claim is a quantitative statement about the degree of factorization, the paper should either propagate the LQCD point-to-point uncertainties into the R(x,|t|) ratios, or explicitly state that the 10% is a measure of deviation from a constant in the plotted points and is not statistically quantified.","section":"Sec. IV A 4"}],"minor_comments":[{"comment":"The ratio in Eq. (26) is written as H_{u-d}(x,t)/H_{u-d}(x,t), which is identically 1; the denominator should involve a different t value, presumably t_j. Please fix the typo.","section":"Eq. (26)"},{"comment":"The definition of the dimensionless variable t≡−t/Λ² combined with the expression e^{−t} in Eq. (33) appears to give a t-dependence that grows with |t|, which would be unphysical for a form factor or GPD. Please check the sign convention and ensure the expression actually decreases with |t|.","section":"Eq. (33)"},{"comment":"The 'Semi-Reggeized' exemplar is selected as 'the first SR result obtained' with the required form, not as a best-fit or an ensemble average. This makes the comparison with the BF and FF exemplars in Table I and Fig. 14 potentially biased. Please clarify that this is a single, non-representative example, or run multiple RCA searches and report the distribution.","section":"Sec. IV A 3"},{"comment":"The phrase 'systematic convergence' overstates what ECC demonstrates. ECC clusters replicas into families based on Taylor coefficients; it does not show convergence to a unique symbolic form. Please either rename the criterion (e.g., 'systematic consistency') or qualify that the clusters represent distinct but equally valid solutions from the SR search.","section":"Sec. III C"},{"comment":"The sentence 'In Fig. I, we present a typical result for each cluster' appears to refer to Table I, which lists the functional forms, not to a figure. Please correct the cross-reference.","section":"Sec. IV B"},{"comment":"There are several typographical issues, including 'Relevent', 'farther develops', 'It has been convergence is not particularly well-defined', and inconsistent use of 't' for both the physical and dimensionless variable. A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a promising methodology but the published version should be more conservative in its claims. The in-region factorization evidence is reasonable, but the low-x localization and the extrapolated radius predictions are not established. I would encourage the authors to either add direct tests of the x-location of factorization breaking and validate the extrapolation, or reframe those parts as model-dependent illustrations. The lack of a BF-versus-BF baseline in the KL analysis is a known limitation that should be acknowledged in the abstract or conclusions, not only in the body. The paper's 'first time' claims regarding consistency of symbolic regression are difficult to verify and should be phrased more modestly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Fair to say: the interesting new thing here is ECC, and the factorization claim is only half earned. If you read it for the method, the ECC idea is worth taking seriously: expand each PySR replica as Taylor coefficients around the midpoint, cluster in coefficient space, and use the spread of clusters as a convergence/extrapolation diagnostic. That is a real contribution, clearly described and tested on 1000 replicas. The application to GPDs is also new, and benchmarking against VGG, GGL, GK is sensible. The in-region evidence for approximate factorization (0.255 ≤ x ≤ 0.855, |t| ≤ 1 GeV²) is about as good as you can expect from 13×5 points: fixed-t ratios flat at ~10%, BF vs FF MSE overlap. Credit where due: they are unusually explicit that the MSEs are fitness metrics, not statistical errors, and that LQCD correlations are unaccounted for. They also flag the missing BF-vs-BF baseline in the KL test themselves.\n\nThe soft spots are in the extrapolation and the low-x conclusion. The Sec IV C claim that factorization breaking is 'entirely focused in the low x region' is not established. The A30/A10 ratio weights large x more heavily, and integrated moments cannot localize breaking in x; the paper performs no test that isolates where in x the breaking sits. The stress-test note has it right. The data are equally compatible with breaking at high x or at |t|>1 GeV². Since the bRMS(x) interpretations in Fig 20 and the comparison with Reggeized models rest on this localization, that part overreaches. Same for the densities/radii in Figs 18-20: the ECC clusters diverge precisely in the extrapolation region, and the paper itself says these distributions are 'a prediction of the combined SR & LQCD framework', but there is no uncertainty attached to that prediction. The semi-Reggeized exemplar is also just one SR answer with the required form, not a distribution over such forms, so it cannot carry the weight of a Regge-behavior claim.\n\nNone of this kills the paper. The methodological core—ECC and the force-factorized penalty—is sound enough to merit serious review, and the in-region factorization statement is honestly supported. What needs revision is the scope of the physics conclusions: stop localizing the breaking to low x, present the extrapolated densities as illustrative scenarios, and either release code/data or show the replicas. I would send it to review, with the expectation of a major revision. Citation pattern is fine; the failure to release artifacts is the main reproducibility gap.","headline":"New method (ECC) is worth a close look, but the low-x localization of factorization breaking is not supported by the moment-ratio evidence and should be dialed back in revision.","tokens_in":29205,"tokens_out":3322,"would_cite":false,"duration_ms":29554,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that lattice QCD data for the isovector GPD $H_{u-d}(x,t,\\xi=0)$ approximately factorize in the measured kinematic region, with any factorization breaking concentrated at low $x$.","keywords":["generalized parton distributions","symbolic regression","lattice QCD","factorization","transverse spatial densities","expansion coefficient clustering","Dirac form factor","proton tomography"],"falsifier":"Obtain lattice QCD points for $H_{u-d}(x,t,\\xi=0)$ at $x<0.255$ for the same $|t|$ values up to about $1\\,\\mathrm{GeV}^2$. If the fixed-$t$ ratio $R(x,|t|)=H_{u-d}(x,0)/H_{u-d}(x,|t|)$ remains within about 10 percent of constant in that new low-$x$ region, or if it deviates by substantially more than 10 percent within the originally trained $0.255\\le x\\le 0.855$ range, the paper's claim that factorization breaking is confined to low $x$ is falsified.","tokens_in":28146,"feed_emoji":"⚛️","tokens_out":9037,"duration_ms":77217,"temperature":0.7,"pith_summary":"The paper tries to establish that symbolic regression can turn a small 13-by-5 grid of lattice QCD points for the quark GPD $H_{u-d}(x,t,\\xi=0)$ into interpretable analytic expressions, and that those expressions reveal an approximate separation of $x$ and $t$ dependence in the trained region $0.255\\le x\\le 0.855$, $|t|\\le 1\\,\\mathrm{GeV}^2$. It matters because the Fourier transform of the GPD in $t$ gives the quark's transverse spatial distribution: a factorized GPD means the quark transverse radius is nearly independent of $x$, whereas QCD-based intuition expects high-$x$ quarks to be more point-like. The paper argues that the apparent factorization breaking seen in integrated lattice moment ratios is not contradictory, because it can be entirely focused in the low-$x$ region outside current data. A reader should care because the paper offers a concrete route to extracting proton tomography from sparse lattice data while quantifying how much of the result is assumption rather than data.","feed_headline":"GPDs factorize in the measured x-t region, regression finds","feed_subtitle":"Symbolic-regression fits to lattice data map where quark transverse size changes, with breaking pinned to low x.","key_machinery":"The load-bearing machinery is Expansion Coefficient Clustering: each of roughly a thousand independent symbolic-regression replicas is expanded in Taylor polynomials around the midpoint of the training region, $x=0.555$, and the resulting coefficient vectors are clustered so that replicas with the same extrapolation behavior form bands. This is paired with two custom training criteria: a force-factorization penalty that nudges replicas toward the form $f_1(x)f_2(t)$, and a semi-Reggeized form $x^\\alpha(1-x)^\\beta P(x)g(t)$ imposed through a redundant-variable trick. A finite-moment filter keeps only replicas whose integral over $x$ at $t=0$ is well defined, which both removes poles and, without being enforced, yields a Dirac form factor $A_{10}(0)$ near the expected value of 1.","core_discovery":"On the paper's own terms, the central discovery is that unconstrained symbolic-regression fits to the lattice data still approximately factorize numerically: the fixed-$t$ ratio $R(x,|t|)=H_{u-d}(x,0)/H_{u-d}(x,|t|)$ stays within about 10 percent of constant in the trained region, and the moment ratios $A_{20}(t)/A_{10}(t)$ and $A_{30}(t)/A_{10}(t)$ reconstructed from the fitted replicas are nearly $t$-independent. The paper interprets this as evidence that the lattice data allow factorization breaking only in the low-$x$ region, and that this is compatible with existing integrated lattice moment results. It also claims a methodological discovery: comparing the distributions of best-fit versus forced-factorized mean-squared errors, through a Kullback-Leibler divergence, separates factorizing sources from non-factorizing ones, showing that symbolic regression can act as a hypothesis-testing tool rather than just a fitting tool.","pith_inferences":["Going beyond the paper, one could test whether the factorization pattern is special to the isovector unpolarized GPD: applying the same pipeline to the gluon GPD or the helicity-flip GPD $E$ would show whether constant radii are a general lattice-data feature or an artifact of this one observable.","The clusters that diverge at low $x$ imply a concrete prediction that future data can settle: once lattice calculations reach $x\\approx 0.1$, the transverse radius should either grow steeply, confirming low-$x$ factorization breaking, or remain flat, contradicting the paper's low-$x$ interpretation.","The ECC Taylor basis centered at $x=0.555$ may hide endpoint behavior; an expansion basis better adapted to $x\\to 0$ or $x\\to 1$, such as a Gegenbauer or Bernoulli basis, would reveal whether the cluster structure is an artifact of the midpoint expansion."],"forward_implications":["If the factorization claim survives, the quark transverse RMS radius $b_{\\mathrm{RMS}}(x)$ is approximately constant across $0.255\\le x\\le 0.855$, meaning current lattice data do not yet resolve the expected shrinking of high-$x$ quark configurations.","The observed $t$-dependence of integrated moment ratios does not refute approximate local factorization; the tension dissolves if the breaking is confined to $x<0.255$ and $|t|\\ge 1\\,\\mathrm{GeV}^2$.","Comparing best-fit and forced-factorized loss distributions gives a quantitative, reproducible test of whether any future GPD dataset is compatible with factorization.","The Taylor-coefficient clustering criterion offers a general convergence check for symbolic regression on small physics datasets, beyond this particular GPD extraction.","A finite-moment filter alone eliminates pole instabilities and biases replicas toward the correct charge sum rule, suggesting physics constraints can be injected cheaply into symbolic regression workflows."],"supporting_citations":[{"why":"Supplies the 13-by-5 lattice QCD grid of $H_{u-d}(x,t,\\xi=0)$ values that all symbolic-regression replicas are trained on.","marker":"[29]"},{"why":"Documents the symbolic-regression algorithm and its hyperparameters, giving the method its mutation, selection, and parsimony behavior.","marker":"[24]"},{"why":"Provides the spectator-model GPD parametrization used as a non-factorizing benchmark source for the regression and the KL-divergence test.","marker":"[41, 42]"},{"why":"Provides the double-distribution GPD parametrization used as a second non-factorizing benchmark source.","marker":"[43]"},{"why":"Provides the phenomenological GPD with built-in factorized $x$-$t$ dependence used as the factorizing-benchmark baseline for the KL-divergence test.","marker":"[44, 45]"},{"why":"Supplies the experimental Dirac form-factor data used to check the extrapolated $A_{10}(t)$ from the fitted replicas.","marker":"[51]"},{"why":"Provides lattice values of $A_{20}(t)/A_{10}(t)$ against which the reconstructed moment ratio from symbolic-regression replicas is compared.","marker":"[54, 55]"},{"why":"Provides the $A_{30}(t)/A_{10}(t)$ lattice ratio at heavier pion mass used to test whether the reconstructed replicas reproduce global factorization breaking.","marker":"[56]"}],"fun_headline_variants":["AI regression finds GPDs nearly factorize in trained x-t region","Symbolic regression pins GPD factorization to mid-x, low-x breaking","KL divergence separates factorizing from non-factorizing GPD fits","Lattice GPDs stay factorized until low x, regression shows","Symbolic regression exposes approximate GPD factorization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The symbolic expressions trained on the 13-by-5 grid are treated as reliable predictions outside that grid, especially below $x=0.255$ and above $|t|=1\\,\\mathrm{GeV}^2$, when computing the spatial densities, radii, and form-factor extrapolation.","fun_headline_variants_meta":{"raw":{"variants":["AI regression finds GPDs nearly factorize in trained x-t region","Symbolic regression pins GPD factorization to mid-x, low-x breaking","KL divergence separates factorizing from non-factorizing GPD fits","Lattice GPDs stay factorized until low x, regression shows","Symbolic regression exposes approximate GPD factorization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1544,"prompt_tokens":959,"completion_tokens":585,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":575,"tokens_out":585,"duration_ms":4894,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:11:59.521016+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain lattice QCD points for $H_{u-d}(x,t,\\xi=0)$ at $x<0.255$ for the same $|t|$ values up to about $1\\,\\mathrm{GeV}^2$. If the fixed-$t$ ratio $R(x,|t|)=H_{u-d}(x,0)/H_{u-d}(x,|t|)$ remains within about 10 percent of constant in that new low-$x$ region, or if it deviates by substantially more than 10 percent within the originally trained $0.255\\le x\\le 0.855$ range, the paper's claim that factorization breaking is confined to low $x$ is falsified.","supporting_citations":[{"cited_title":"Generalized Parton Distributions and Color Transparency","cited_arxiv_id":"hep-ph/0405014","evidence_quote":"Supplies the 13-by-5 lattice QCD grid of $H_{u-d}(x,t,\\xi=0)$ values that all symbolic-regression replicas are trained on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the $A_{30}(t)/A_{10}(t)$ lattice ratio at heavier pion mass used to test whether the reconstructed replicas reproduce global factorization breaking."}],"review_version":1}