{"id":"4996b585-e0c7-4d70-931a-5d4822f4fb04","arxiv_id":"2501.07123","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A transformer-based symbolic regression tool rediscovers a Lund-like formula for fragmentation functions from COMPASS multiplicity data, but the discovery is a post-hoc selection among many candidate equations.","lead":"This paper uses symbolic regression, a machine learning technique, to find a mathematical formula for quark fragmentation functions from charged hadron data measured at CERN's COMPASS experiment. The formula resembles the Lund string model, but the strongest claims are weakened by post-hoc model selection and per-bin refitting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The final f_SR is a human-selected generalization (merging f1 and f3, freeing the exponent) rather than a direct SR output, so the central 'no pre-assumed functional form' claim is not yet established.","rationale":"I read the paper as an honest proof-of-concept, and the full table of raw SR outputs is a useful admission of diversity. However, the novelty claim rests on inferring a functional form without a pre-assumed ansatz, and the protocol visibly reintroduces human model selection: g4 is constructed by merging f1 and f3 and freeing an exponent, not output by NeSymReS. The reader's factorization concern about Eq.6 is legitimate but is a standard assumption in LO SIDIS analyses and is not the main threat to the central claim; the more load-bearing issue is that the final form's structure is not actually learned. A bootstrap recovery test would settle whether g4 is a stable inference or a post-hoc choice. Because the paper is already conditional and this concern strengthens that condition rather than refuting the entire approach, I keep the verdict unchanged.","tokens_in":25,"tokens_out":6364,"duration_ms":135833,"concrete_test":"Rerun NeSymReS on 100 bootstrap resamples of each of the 38 h+ and h- bins and tally how often the raw SR output is f3 or g4 (or an equivalent form) rather than one of the alternative expressions in Table 1. If g4 is not the dominant minimal-complexity expression recovered across resamples, then the final form is a human selection rather than a robust symbolic-regression inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the Abstract and Eq.14 is that SR infers f_SR(z)=a(1-z)^c exp(-bz) directly from data without a pre-assumed form. Tracing the pipeline in Sec.5.1, this is not quite what happened. Table 1 lists 38 raw SR outputs per charge; among them many trigonometric and rational forms appear, and the form a exp(-bz)/(z-c) occurs in only one y-bin (0.2<y<0.3). The authors then define f3 as a special case of that one output, and construct g4(z)=a(1-z)^c exp(-bz) 'by freeing the exponent' in Eq.5 as a merge of f1 and f3. That step is human model selection, not symbolic-regression inference. The generalisation checks in Table 2 are fits of this hand-chosen g4 to pion and kaon multiplicities, even though the paper states that for pions and kaons the majority of SR-learned equations were trigonometric and 'do not describe' the data. Similarly, in Sec.5.3 the purely learned f(z,x) fails at high z, and the authors manually multiply by (1-z)^-0.2 before fitting Eq.13. Thus the statement that 'no constraints or assumptions are made' and the abstract phrasing 'the function learned by symbolic regression' overstate what the protocol establishes: the final functional family is inferred only after substantial human modification, and its predictive tests are per-bin refits rather than fixed-parameter out-of-sample predictions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the transformer-based symbolic regression tool NeSymReS to charged-hadron, pion, and kaon multiplicities from COMPASS, and claims to infer, without a pre-assumed functional form, the FF-like function f_SR(z)=a(1-z)^c exp(-bz), which it says resembles the Lund string function and is a candidate for use in global FF fits. The analysis proceeds in three steps: one-dimensional SR on multiplicities in individual (x,y) bins (Sec. 5.1), SR on leading-order extracted FFs (Sec. 5.2), and two-dimensional SR on M(z,x) (Sec. 5.3). The paper reports chi2/ndf values for per-bin fits of g4 to h, pi, and K multiplicities and interprets the factorization of the learned bivariate function as experimental evidence for the factorization theorem.","tokens_in":21909,"tokens_out":3337,"duration_ms":35714,"significance":"If the central claim were established, this would be a useful proof-of-concept that symbolic regression can propose an interpretable, physically plausible FF parameterization directly from noisy experimental data, complementing traditional global-fit methodology. The paper is transparent in reporting the full set of 38 SR outputs (Table 1), the per-bin fit qualities (Table 2), and the explicit human-in-the-loop selection step for the lowest-loss cosine expression. However, the significance is currently conditional: the final functional family is substantially human-selected, and the 'test' evaluations are per-bin refits rather than fixed-parameter predictions, so the paper's headline claim that the function was learned directly from data is not yet supported.","major_comments":[{"comment":"The central claim that f_SR(z)=a(1-z)^c exp(-bz) was inferred directly by symbolic regression is not supported by the reported pipeline. Table 1 shows 38 per-bin SR outputs, the majority of which are trigonometric or rational forms; the form a exp(-bz)/(z-c) appears in only one y-bin, and the lowest-loss output a(1-b cos(3z))^c is explicitly rejected on physical grounds. Equation (5) then defines g4 by merging f1 and f3 and freeing the exponent, which is a human model-selection step, not an SR output. The abstract and Sec. 2 'no constraints or assumptions are made' therefore overstate what the protocol establishes. The authors should either present g4 as a human-in-the-loop generalization of SR outputs, or demonstrate that g4 is itself recovered when SR is run on the merged data.","section":"Sec. 5.1, Table 1, Eq. (5), Eq. (14)"},{"comment":"The 'test' evaluation is not an out-of-sample test of the functional form, because the parameters a, b, and c are refit independently in every kinematic bin. Table 2 therefore measures the flexibility of a three-parameter family rather than the predictive power of the specific form selected in the training bin. The paper's claim that the function 'maintained its predictive power even in bins where it wasn't directly trained' (Sec. 6) is not justified by fits with per-bin parameters. A fixed-parameter evaluation on held-out bins, or at least a clear statement that the generalization claim refers to the functional form and not the parameter values, is needed.","section":"Sec. 5.1, Table 2, Figs. 4-6"},{"comment":"The claim that Eq. (12) provides 'the first experimental evidence of the factorization theorem' is circular in the context of the paper. Equation (6) already assumes a factorized leading-order SIDIS cross section, and the multiplicity definition in Eq. (7) inherits that factorization. The SR output in Eq. (12) is also not a direct test: the z dependence fails at large z, and Eq. (13) is obtained by manually multiplying by (1-z)^gamma before fitting. The strong y-dependence of the fitted parameters in Table 3 further shows that a single two-dimensional function with common parameters is not obtained, so the factorization claim should be substantially weakened or removed.","section":"Sec. 5.3, Eqs. (6), (12), (13)"},{"comment":"The interpretation of the LO FF extraction as verifying the functional form learned from multiplicities relies on assumptions that are not discussed: factorization at LO, use of MSTW08 PDFs, neglect of target-mass and higher-twist effects, and the identification of the measured multiplicity shape with the z-dependence of FFs. Additionally, the extracted function h(z) in Eq. (10) does not vanish at z=1, so the similarity to g4 is only structural. The paper should either add explicit caveats about these assumptions or present the LO extraction as a consistency check rather than independent confirmation.","section":"Sec. 5.2, Eqs. (6)-(10)"},{"comment":"The reported numerical constants for f2 are internally suspicious: the text gives a approximately -9.8, b = -7.7, c approximately 1, which, with c approximated by 1 and z in (0.2,0.8), would make exp(-bz) = exp(+7.7 z) a rapidly increasing function of z, contrary to the stated description of the data at high z. The sign convention for b should be clarified or corrected, since this numerical example is used to motivate f3.","section":"Sec. 5.1, text near Eq. (4)"}],"minor_comments":[{"comment":"'compliments them' should be 'complements them'; also 'Colider' should be 'Collider' in the first paragraph.","section":"Introduction"},{"comment":"'direclty' should be 'directly' in the concluding paragraph.","section":"Conclusion"},{"comment":"'evaluated' in 'where the PDFs are evaluyated' should be 'evaluated'; there is also an 'od' typo in 'the FF od the quark'.","section":"Sec. 5.2"},{"comment":"Reference [24] is an empty placeholder (OpenAI ChatGPT-4) and should either be completed with a proper citation or removed.","section":"References"},{"comment":"The terms 'out-of-distribution' and 'out-of-sample' are used without definitions; since the evaluation is per-bin refits, these terms should be replaced or explicitly defined to avoid implying fixed-parameter prediction.","section":"Sec. 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's main novelty over prior SR applications in HEP is modest, and the current text overclaims by presenting a human-selected generalization as a direct SR output. The per-bin fits themselves are reasonable, but the load-bearing claims about direct inference and out-of-sample generalization need to be reframed or supported by additional fixed-parameter tests. The topic fits the journal's scope, and the underlying data presentation is transparent enough that a revision could address the issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful, honest proof-of-concept that symbolic regression can land on a Lund-like functional form from COMPASS multiplicities, and the g4 fits are genuinely good. But the abstract's claim that SR infers the form directly from data without a pre-assumed function does not survive reading the pipeline. The final f_SR is assembled by hand from several SR outputs, and the 'test' is a per-bin refit rather than a fixed-parameter prediction.\n\nWhat is actually new: the paper applies a pretrained transformer-based symbolic regression to real SIDIS multiplicity data and proposes a candidate FF parameterization. It also reports all 38 raw SR outputs, including the rejected trigonometric ones, which is more transparent than most SR papers. The fits of g4(z)=a(1-z)^c exp(-bz) to h±, π±, and K± multiplicities across the x-y bins give χ2/ndf values mostly near 1. That is a real empirical result for this dataset.\n\nSoft spots, in proportion:\n\n1. The central inference is human-selected. Table 1 lists many trigonometric outputs; the lowest-loss candidate is a cosine function and is discarded on physical grounds. f3 appears in only one bin, and g4 is then defined by merging f1 and f3 and freeing the exponent. That is model building with SR as a suggestion engine, not SR learning the function with no assumptions. The statement that 'no constraints or assumptions are made' is inaccurate.\n\n2. The generalization test is weaker than claimed. 'Test' in Sec. 5.1 means fitting g4 separately in each bin with its own a, b, c. A fixed-parameter out-of-sample prediction would be much stronger evidence than per-bin fits, no matter how reasonable those fits look.\n\n3. The factorization claim in Sec. 5.3 is overreach. Equation 6 already assumes factorization at LO; finding a product form in one y bin does not constitute experimental evidence for the factorization theorem. The learned f(z,x) also fails at high z, and the authors manually add (1-z)^-0.2 before fitting Eq. 13.\n\n4. The LO extraction in Sec. 5.2 is not well-posed as written: two pion multiplicity equations for three independent FFs. That needs explanation or revision.\n\nBottom line: this is a legitimate proof-of-concept and the authors have been transparent enough that a referee can engage with the actual pipeline. It is not a new physical function—f_SR is a special case of existing Lund/DSS-style parameterizations. Worth sending to peer review because the methodology question is timely, but the authors should be pushed to reframe the claims, make the human-in-the-loop selection explicit, and add a true out-of-sample check.","headline":"A transparent proof-of-concept that symbolic regression can propose a Lund-like form from COMPASS data, but the final form is a human-selected generalization, not a direct SR output, and the paper's stronger claims overstate what is established.","tokens_in":22408,"tokens_out":3847,"would_cite":false,"duration_ms":42959,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Symbolic regression recovers a data-driven functional form for fragmentation functions","keywords":["fragmentation functions","symbolic regression","semi-inclusive deep inelastic scattering","charged hadron multiplicities","Lund string function","global QCD fits","interpretable machine learning","transformer-based symbolic regression"],"falsifier":"Fit $f_{\\rm SR}(z)$ separately in fine bins of $Q^2$ (or $y$) and check whether the fitted parameters $a$, $b$, and $c$ stay constant after DGLAP evolution; a systematic drift or visible $Q^2$ dependence would indicate that the $z$-shape is contaminated by non-factorizing contributions. Alternatively, evolve $f_{\\rm SR}$ with DGLAP and compare its predictions with $e^+e^-$ annihilation or proton-proton hadron-production data: if the evolved form fails to describe those independent measurements while standard parameterizations succeed, the claim that it is a viable global-fit candidate would be refuted.","tokens_in":21338,"feed_emoji":"⚛️","tokens_out":8142,"duration_ms":71117,"temperature":0.7,"pith_summary":"The paper claims that symbolic regression, applied directly to measured charged-hadron multiplicities from semi-inclusive deep inelastic scattering, can infer a functional form of fragmentation functions without assuming one in advance. The function it recovers, $f_{\\rm SR}(z)=a(1-z)^c\\exp(-bz)$, resembles the Lund string fragmentation function and fits the measured $z$-dependence of unidentified hadrons, pions, and kaons across the surveyed kinematic bins. Because fragmentation functions cannot be computed from perturbative QCD and are normally fixed by global fits to a pre-assumed template, a data-driven candidate like this could reduce model bias in future global analyses. The paper also finds a similar structure when fragmentation functions are extracted directly from pion multiplicities at leading order, which it reads as consistency between the multiplicity-level and fragmentation-function-level inferences.","feed_headline":"Data alone yields a Lund-like fragmentation formula","feed_subtitle":"A symbolic-regression search on noisy SIDIS multiplicities returns a (1-z)^c e^{-bz} form to test in global QCD fits.","key_machinery":"The load-bearing mechanism is symbolic regression as implemented by a pretrained transformer that maps a set of $(z,M_h)$ points to an equation skeleton, with constants filled by nonlinear optimization; expressions are represented as unary-binary trees, so the search runs over a discrete space of mathematical formulas rather than over fixed template parameters. The univariate analysis on $\\{z,M_h\\}$ yields the candidate family $g_4(z)=a(1-z)^c\\exp(-bz)$, and the leading-order factorization formula, which expresses the multiplicity as a ratio of SIDIS to DIS cross sections through parton distribution functions and fragmentation functions, is what lets the paper attach the learned $z$-shape to fragmentation functions rather than to the full cross section. The same machinery is then run on two-dimensional data $\\{z,x,M_h\\}$ and on leading-order-extracted fragmentation-function distributions, producing the corroborating structure $a\\exp(-bz)/(z-c)^2$.","core_discovery":"The central claim is that the $z$-shape of semi-inclusive deep inelastic scattering multiplicities contains enough information for a transformer-based symbolic regression model to output an interpretable, compact formula, $f_{\\rm SR}(z)=a(1-z)^c\\exp(-bz)$, before any fragmentation-function template is imposed. Fitting this form to charged hadron, pion, and kaon multiplicities in individual $(x,y)$ bins gives reduced $\\chi^2$ values of order one across nearly all bins, including bins not used for the symbolic-regression training, and the functional form is close to but distinct from the Lund symmetric fragmentation function $f(z)\\propto (1/z)(1-z)^\\alpha\\exp(-\\beta m_h^2/z)$. When the same pipeline is applied to fragmentation functions extracted point-by-point from pion multiplicities at leading order, the surviving forms share the structure $a\\exp(-bz)/(z-c)^2$, which the paper takes as corroboration. On this basis the paper proposes $f_{\\rm SR}(z)$ as a candidate parameterization for global QCD fits, arguing that both the model and its parameters would then originate from data.","pith_inferences":["If the same symbolic-regression exercise is run on $e^+e^-$ or proton-proton data and the same functional family emerges after DGLAP evolution, the result would generalize from SIDIS factorization to a universal quark-to-hadron fragmentation shape.","The family $a(1-z)^c\\exp(-bz)$ is flexible enough to interpolate between the Lund-like limit and the $1/z^2$ falloff seen in some bins, so nested fits could quantify how much of the discovered form is genuinely data-driven versus an artifact of the pretrained model's preference for short expressions.","A direct test would be to seed the symbolic-regression search with the Lund form and see whether the model returns a simpler or different expression, separating the influence of expression-tree-length bias from the information actually present in the data."],"forward_implications":["Future global QCD fits could use the data-derived form $f_{\\rm SR}(z)$ instead of a pre-assumed template, so that both the model and its parameters come from data.","The same compact form fits unidentified hadrons, pions, and kaons, including bins outside the training set, indicating that one $z$-shape may serve across hadron species in the measured kinematic range.","The two-dimensional inference yields a factorized dependence $e^{-\\alpha z}\\cdot e^{2.3(1\\pm\\beta x)^2}$, which the paper reads as direct experimental evidence for the factorization assumption usually imposed by hand.","The structural similarity between the multiplicity-level result and the leading-order-extracted fragmentation functions supports the internal consistency of that extraction procedure."],"supporting_citations":[{"why":"Supplies the charged-hadron and charged-pion SIDIS multiplicities that form the training data for the univariate and bivariate symbolic-regression runs.","marker":"[35]"},{"why":"Supplies the charged-kaon multiplicities used to test the learned form out of sample.","marker":"[36]"},{"why":"Defines the Lund symmetric fragmentation function whose $(1-z)^\\alpha \\exp(-\\beta m_h^2/z)$ structure the learned function resembles.","marker":"[16]"},{"why":"Represents the established global-fit parameterization approach that the paper positions its data-derived form against.","marker":"[6]"},{"why":"Represents a newer global fit of pion fragmentation functions, used as the phenomenological baseline for functional forms.","marker":"[7]"},{"why":"Provides the pretrained transformer-based symbolic regression model that performs the equation search.","marker":"[43]"},{"why":"Provides the leading-order parton distribution functions used in the direct point-by-point fragmentation-function extraction from pion multiplicities.","marker":"[45]"},{"why":"Grounds the definition of fragmentation functions and their role in SIDIS cross sections, supporting the factorization interpretation.","marker":"[15]"}],"fun_headline_variants":["Symbolic regression derives a fragmentation formula from data","Data-only symbolic regression finds a Lund-like fragmentation law","Machine-learned fragmentation function emerges without pre-set form","First symbolic regression on multiplicities yields interpretable fragmentation term"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the measured multiplicity factorizes at leading order into parton distribution functions and fragmentation functions, so that the $z$-dependence of the data is attributable to fragmentation alone; if target remnants, higher-twist effects, or unaccounted $Q^2$ dependence mix into the $z$-shape, the learned function is not really a fragmentation function.","fun_headline_variants_meta":{"raw":{"variants":["Symbolic regression derives a fragmentation formula from data","Data-only symbolic regression finds a Lund-like fragmentation law","Machine-learned fragmentation function emerges without pre-set form","First symbolic regression on multiplicities yields interpretable fragmentation term"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1607,"prompt_tokens":947,"completion_tokens":660,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":597}},"tokens_in":563,"tokens_out":660,"duration_ms":6730,"temperature":1.0,"reasoning_tokens":597,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:49:10.025178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit $f_{\\rm SR}(z)$ separately in fine bins of $Q^2$ (or $y$) and check whether the fitted parameters $a$, $b$, and $c$ stay constant after DGLAP evolution; a systematic drift or visible $Q^2$ dependence would indicate that the $z$-shape is contaminated by non-factorizing contributions. Alternatively, evolve $f_{\\rm SR}$ with DGLAP and compare its predictions with $e^+e^-$ annihilation or proton-proton hadron-production data: if the evolved form fails to describe those independent measurements while standard parameterizations succeed, the claim that it is a viable global-fit candidate would be refuted.","supporting_citations":[{"cited_title":"Multiplicities of charged pions and charged hadrons from deep-inelastic scattering of muons off an isoscalar target","cited_arxiv_id":null,"evidence_quote":"Supplies the charged-hadron and charged-pion SIDIS multiplicities that form the training data for the univariate and bivariate symbolic-regression runs."},{"cited_title":"Multiplicities of charged kaons from deep-inelastic muon scattering off an isoscalar target","cited_arxiv_id":null,"evidence_quote":"Supplies the charged-kaon multiplicities used to test the learned form out of sample."},{"cited_title":"Global analysis of fragmentation functions for pions and kaons and their uncertainties","cited_arxiv_id":null,"evidence_quote":"Represents the established global-fit parameterization approach that the paper positions its data-derived form against."}],"review_version":1}