Pith. sign in

REVIEW 5 major objections 5 minor 2 cited by

Inferring Interpretable Models of Fragmentation Functions using Symbolic Regression

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Symbolic regression recovers a data-driven functional form for fragmentation functions

desk verdict A transparent proof-of-concept that symbolic regression can propose a Lund-like form from COMPASS data, but the final form is a human-selected generalization, not a direct SR output, and the paper's stronger claims overstate what is established. read the letter →

arxiv 2501.07123 v1 pith:Z4OBIRFV submitted 2025-01-13 hep-ph cs.LGcs.SChep-th

classification hep-phcs.LGcs.SChep-th
keywords fragmentationfunctionssymbolicregressionsemi-inclusivedeepinelasticscatteringchargedhadronmultiplicitiesLundstringfunctionglobalQCDfitsinterpretablemachinelearningtransformer-based
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that symbolic regression, applied directly to measured charged-hadron multiplicities from semi-inclusive deep inelastic scattering, can infer a functional form of fragmentation functions without assuming one in advance. The function it recovers, $f_{\rm SR}(z)=a(1-z)^c\exp(-bz)$, resembles the Lund string fragmentation function and fits the measured $z$-dependence of unidentified hadrons, pions, and kaons across the surveyed kinematic bins. Because fragmentation functions cannot be computed from perturbative QCD and are normally fixed by global fits to a pre-assumed template, a data-driven candidate like this could reduce model bias in future global analyses. The paper also finds a similar structure when fragmentation functions are extracted directly from pion multiplicities at leading order, which it reads as consistency between the multiplicity-level and fragmentation-function-level inferences.

What carries the argument

The load-bearing mechanism is symbolic regression as implemented by a pretrained transformer that maps a set of $(z,M_h)$ points to an equation skeleton, with constants filled by nonlinear optimization; expressions are represented as unary-binary trees, so the search runs over a discrete space of mathematical formulas rather than over fixed template parameters. The univariate analysis on $\{z,M_h\}$ yields the candidate family $g_4(z)=a(1-z)^c\exp(-bz)$, and the leading-order factorization formula, which expresses the multiplicity as a ratio of SIDIS to DIS cross sections through parton distribution functions and fragmentation functions, is what lets the paper attach the learned $z$-shape to fragmentation functions rather than to the full cross section. The same machinery is then run on two-dimensional data $\{z,x,M_h\}$ and on leading-order-extracted fragmentation-function distributions, producing the corroborating structure $a\exp(-bz)/(z-c)^2$.

What would settle it

Fit $f_{\rm SR}(z)$ separately in fine bins of $Q^2$ (or $y$) and check whether the fitted parameters $a$, $b$, and $c$ stay constant after DGLAP evolution; a systematic drift or visible $Q^2$ dependence would indicate that the $z$-shape is contaminated by non-factorizing contributions. Alternatively, evolve $f_{\rm SR}$ with DGLAP and compare its predictions with $e^+e^-$ annihilation or proton-proton hadron-production data: if the evolved form fails to describe those independent measurements while standard parameterizations succeed, the claim that it is a viable global-fit candidate would be refuted.

Watch

Extended reading notes

Core claim

The central claim is that the $z$-shape of semi-inclusive deep inelastic scattering multiplicities contains enough information for a transformer-based symbolic regression model to output an interpretable, compact formula, $f_{\rm SR}(z)=a(1-z)^c\exp(-bz)$, before any fragmentation-function template is imposed. Fitting this form to charged hadron, pion, and kaon multiplicities in individual $(x,y)$ bins gives reduced $\chi^2$ values of order one across nearly all bins, including bins not used for the symbolic-regression training, and the functional form is close to but distinct from the Lund symmetric fragmentation function $f(z)\propto (1/z)(1-z)^\alpha\exp(-\beta m_h^2/z)$. When the same pipeline is applied to fragmentation functions extracted point-by-point from pion multiplicities at leading order, the surviving forms share the structure $a\exp(-bz)/(z-c)^2$, which the paper takes as corroboration. On this basis the paper proposes $f_{\rm SR}(z)$ as a candidate parameterization for global QCD fits, arguing that both the model and its parameters would then originate from data.

Load-bearing premise

The central assumption is that the measured multiplicity factorizes at leading order into parton distribution functions and fragmentation functions, so that the $z$-dependence of the data is attributable to fragmentation alone; if target remnants, higher-twist effects, or unaccounted $Q^2$ dependence mix into the $z$-shape, the learned function is not really a fragmentation function.

Editorial extensions

If this is right

  • Future global QCD fits could use the data-derived form $f_{\rm SR}(z)$ instead of a pre-assumed template, so that both the model and its parameters come from data.
  • The same compact form fits unidentified hadrons, pions, and kaons, including bins outside the training set, indicating that one $z$-shape may serve across hadron species in the measured kinematic range.
  • The two-dimensional inference yields a factorized dependence $e^{-\alpha z}\cdot e^{2.3(1\pm\beta x)^2}$, which the paper reads as direct experimental evidence for the factorization assumption usually imposed by hand.
  • The structural similarity between the multiplicity-level result and the leading-order-extracted fragmentation functions supports the internal consistency of that extraction procedure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the same symbolic-regression exercise is run on $e^+e^-$ or proton-proton data and the same functional family emerges after DGLAP evolution, the result would generalize from SIDIS factorization to a universal quark-to-hadron fragmentation shape.
  • The family $a(1-z)^c\exp(-bz)$ is flexible enough to interpolate between the Lund-like limit and the $1/z^2$ falloff seen in some bins, so nested fits could quantify how much of the discovered form is genuinely data-driven versus an artifact of the pretrained model's preference for short expressions.
  • A direct test would be to seed the symbolic-regression search with the Lund form and see whether the model returns a simpler or different expression, separating the influence of expression-tree-length bias from the information actually present in the data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper applies the transformer-based symbolic regression tool NeSymReS to charged-hadron, pion, and kaon multiplicities from COMPASS, and claims to infer, without a pre-assumed functional form, the FF-like function f_SR(z)=a(1-z)^c exp(-bz), which it says resembles the Lund string function and is a candidate for use in global FF fits. The analysis proceeds in three steps: one-dimensional SR on multiplicities in individual (x,y) bins (Sec. 5.1), SR on leading-order extracted FFs (Sec. 5.2), and two-dimensional SR on M(z,x) (Sec. 5.3). The paper reports chi2/ndf values for per-bin fits of g4 to h, pi, and K multiplicities and interprets the factorization of the learned bivariate function as experimental evidence for the factorization theorem.

Significance. If the central claim were established, this would be a useful proof-of-concept that symbolic regression can propose an interpretable, physically plausible FF parameterization directly from noisy experimental data, complementing traditional global-fit methodology. The paper is transparent in reporting the full set of 38 SR outputs (Table 1), the per-bin fit qualities (Table 2), and the explicit human-in-the-loop selection step for the lowest-loss cosine expression. However, the significance is currently conditional: the final functional family is substantially human-selected, and the 'test' evaluations are per-bin refits rather than fixed-parameter predictions, so the paper's headline claim that the function was learned directly from data is not yet supported.

major comments (5)
  1. [Sec. 5.1, Table 1, Eq. (5), Eq. (14)] The central claim that f_SR(z)=a(1-z)^c exp(-bz) was inferred directly by symbolic regression is not supported by the reported pipeline. Table 1 shows 38 per-bin SR outputs, the majority of which are trigonometric or rational forms; the form a exp(-bz)/(z-c) appears in only one y-bin, and the lowest-loss output a(1-b cos(3z))^c is explicitly rejected on physical grounds. Equation (5) then defines g4 by merging f1 and f3 and freeing the exponent, which is a human model-selection step, not an SR output. The abstract and Sec. 2 'no constraints or assumptions are made' therefore overstate what the protocol establishes. The authors should either present g4 as a human-in-the-loop generalization of SR outputs, or demonstrate that g4 is itself recovered when SR is run on the merged data.
  2. [Sec. 5.1, Table 2, Figs. 4-6] The 'test' evaluation is not an out-of-sample test of the functional form, because the parameters a, b, and c are refit independently in every kinematic bin. Table 2 therefore measures the flexibility of a three-parameter family rather than the predictive power of the specific form selected in the training bin. The paper's claim that the function 'maintained its predictive power even in bins where it wasn't directly trained' (Sec. 6) is not justified by fits with per-bin parameters. A fixed-parameter evaluation on held-out bins, or at least a clear statement that the generalization claim refers to the functional form and not the parameter values, is needed.
  3. [Sec. 5.3, Eqs. (6), (12), (13)] The claim that Eq. (12) provides 'the first experimental evidence of the factorization theorem' is circular in the context of the paper. Equation (6) already assumes a factorized leading-order SIDIS cross section, and the multiplicity definition in Eq. (7) inherits that factorization. The SR output in Eq. (12) is also not a direct test: the z dependence fails at large z, and Eq. (13) is obtained by manually multiplying by (1-z)^gamma before fitting. The strong y-dependence of the fitted parameters in Table 3 further shows that a single two-dimensional function with common parameters is not obtained, so the factorization claim should be substantially weakened or removed.
  4. [Sec. 5.2, Eqs. (6)-(10)] The interpretation of the LO FF extraction as verifying the functional form learned from multiplicities relies on assumptions that are not discussed: factorization at LO, use of MSTW08 PDFs, neglect of target-mass and higher-twist effects, and the identification of the measured multiplicity shape with the z-dependence of FFs. Additionally, the extracted function h(z) in Eq. (10) does not vanish at z=1, so the similarity to g4 is only structural. The paper should either add explicit caveats about these assumptions or present the LO extraction as a consistency check rather than independent confirmation.
  5. [Sec. 5.1, text near Eq. (4)] The reported numerical constants for f2 are internally suspicious: the text gives a approximately -9.8, b = -7.7, c approximately 1, which, with c approximated by 1 and z in (0.2,0.8), would make exp(-bz) = exp(+7.7 z) a rapidly increasing function of z, contrary to the stated description of the data at high z. The sign convention for b should be clarified or corrected, since this numerical example is used to motivate f3.
minor comments (5)
  1. [Introduction] 'compliments them' should be 'complements them'; also 'Colider' should be 'Collider' in the first paragraph.
  2. [Conclusion] 'direclty' should be 'directly' in the concluding paragraph.
  3. [Sec. 5.2] 'evaluated' in 'where the PDFs are evaluyated' should be 'evaluated'; there is also an 'od' typo in 'the FF od the quark'.
  4. [References] Reference [24] is an empty placeholder (OpenAI ChatGPT-4) and should either be completed with a proper citation or removed.
  5. [Sec. 5.1] The terms 'out-of-distribution' and 'out-of-sample' are used without definitions; since the evaluation is per-bin refits, these terms should be replaced or explicitly defined to avoid implying fixed-parameter prediction.

Circularity Check

3 steps flagged · score 6.0 of 10

The 'test' predictions are per-bin refits of a hand-merged function, and the factorization 'evidence' reads the assumed factorization back from the fit.

  1. fitted input called prediction [Sec. 5.1, Eq. (5), Table 2, Sec. 6]
    "to evaluate the performance of the learned models by SR on “test” data, we performed fits of M h(z) in individual kinematic bins. ... In addition, we consider a general form of f3 by taking the power exponent in the term (1 − z) as a free fit parameter, referred to as g4. ... We thus use g4(z) to fit multiplicities of charged hadrons, pions, and kaons in individual (x, y) bins to check the generalizability of the learned model."

    The generalizability tests in Table 2 and Figures 4–6 are not predictions with fixed SR-learned coefficients; the parameters a, b, c are re-fit independently in every (x, y) bin. A three-parameter function refit to a handful of points in each bin will accommodate the local shape, so the reported χ2/ndf measures fitting flexibility, not out-of-sample transfer. The later claim that the function 'maintained its predictive power even in bins where it wasn't directly trained' converts these per-bin fits into predictions; the predictive power is built into the fitting step, making the validation circular.

  2. self definitional [Sec. 5.3, Eqs. (6) and (12)]
    "d3σh/dx dQ2 dz = C(x,Q2) Σ q e2 q fq(x,Q2)Dh q (z,Q2). ... A first observation is a factorization in the dependence of f (z, x) (Eq. 12) upon z and x through the exponential, providing the first experimental evidence of the factorization theorem that is usually assumed in phenomenological studies of FFs."

    Equation (6) already defines the SIDIS multiplicity as a sum of products of an x-dependent PDF factor and a z-dependent FF factor. The multiplicities fed to SR are interpreted through this factorized expression, and the factorized exponential form f(z,x)=exp(−6.4z)·exp(2.3(1+0.27x)^2) is then announced as 'first experimental evidence' of factorization. The evidence is not independent: the assumed factorized structure is an input to the analysis, and no non-factorized alternative was tested, so the conclusion is read back from the assumed ansatz rather than derived from the data.

1 more flagged steps
  1. other [Sec. 5.1 and Conclusion, Eq. (14)]
    "we consider a general form of f3 by taking the power exponent in the term (1 − z) as a free fit parameter, referred to as g4. This choice is mainly driven by the existence of a power exponent “2” in the learned function f1. Thus, merging f1 and f3 into a general form requires the freeing of the exponent parameter. ... The resulting function is: f SR = a(1 − z)c exp(−bz) (14)"

    The headline result f_SR is not a direct symbolic-regression output: it is the hand-defined g4 obtained by merging two SR outputs and freeing an exponent before fitting. Eq. (14) is therefore g4 by construction, and presenting it as 'the function learned by symbolic regression' (Abstract) makes the final answer equivalent to the authors' own post-selected fit function. This is a self-definitional step in the derivation chain: the claimed machine inference is completed by human model selection, so the 'no pre-assumed functional form' claim is not established by the pipeline.

full rationale

The paper's underlying experimental input is external COMPASS data and the pretrained NeSymReS model is an independent tool, so there is no load-bearing self-citation loop. However, the derivation chain contains two places where a claimed output is equivalent to an input by construction. First, the generalizability tests are per-bin fits of g4 with free parameters, so the 'predictions' are statistically forced by refitting rather than by fixed SR-learned constants; the good χ2 values cannot support the claim of predictive power in untrained bins. Second, the 'first experimental evidence of factorization' in Sec 5.3 is read from a function fitted to multiplicities whose interpretation already assumes the factorized Eq. (6); the separable form is not an independent test. In addition, the final f_SR is the hand-merged, exponent-freed g4 rather than a direct SR output, which undermines the central 'no pre-assumed functional form' claim even though it is more an overstatement than a strict identity. These issues make the central validation partially circular, so a score of 6 is appropriate rather than a higher score that would require a fully forced self-citation or definitional equivalence.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper's quantitative content rests on fit parameters a,b,c and the 2D parameters N,gamma,alpha,beta, all determined from the same COMPASS data. It assumes standard QCD factorization and isospin symmetries, and it imports a pretrained symbolic regression model. No new entities are introduced.

free parameters (4)
  • a (normalization in g4) = varies per (x,y) bin, not tabulated
    Fit to COMPASS multiplicities in each kinematic bin via g4(z)=a(1-z)^c exp(-bz); the paper does not quote the resulting a,b,c values.
  • b (exponential slope in g4) = varies per (x,y) bin, not tabulated
    Fitted per bin; in the 2D analysis b is replaced by alpha in exp(-alpha z), and alpha varies from 2.98 to 10.72 across y bins (Table 3).
  • c (power of (1-z) in g4) = varies per (x,y) bin; one bin gives c approximately 1
    Fitted per bin; in the 2D analysis the equivalent gamma varies from -2.93 to 0.55 across y bins (Table 3), so the exponent is not stable.
  • N, gamma, alpha, beta in f'(z,x) = tabulated per y bin in Table 3
    These four parameters are fit to M_h(z,x) in each y bin and change strongly with y, showing that no single set of constants predicts across the phase space.
assumptions (5)
  • domain assumption SIDIS cross section factorizes into PDFs and FFs at leading order (Eq.6).
    Used in Sec 5.2 to extract FFs from multiplicities and to justify interpreting the learned function as an FF.
  • domain assumption Isospin and charge symmetry reduce pion FFs to three independent functions D_fav, D_unf, D_str (Eq.8).
    Required to set up the linear system in Sec 5.2; the system is underdetermined with two equations and three unknowns.
  • domain assumption MSTW08 LO PDFs provide the quark distributions.
    External phenomenological input to the FF extraction in Sec 5.2.
  • domain assumption NeSymReS pretrained transformer, trained on 100M random equations, produces a valid search over skeletons.
    Tooling assumption; no fine-tuning or domain-specific training is described.
  • ad hoc to paper Physical FFs are smooth and non-periodic in z, justifying rejection of trigonometric SR outputs.
    Applied in Sec 5.1 to discard lower-loss functions such as a(1-b cos(3z))^c, making the selection partly prior-guided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Inferring Interpretable Models of Fragmentation Functions using Symbolic Regression." pith.science (2026). https://pith.science/paper/Z4OBIRFV

@misc{pith2026250107123,
  author       = {Pith},
  title        = {Pith review of: Inferring Interpretable Models of Fragmentation Functions using Symbolic Regression},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z4OBIRFV}},
  note         = {Machine review of arXiv:2501.07123}
}
read the original abstract

Machine learning is rapidly making its path into natural sciences, including high-energy physics. We present the first study that infers, directly from experimental data, a functional form of fragmentation functions. The latter represent a key ingredient to describe physical observables measured in high-energy physics processes that involve hadron production, and predict their values at different energy. Fragmentation functions can not be calculated in theory and have to be determined instead from data. Traditional approaches rely on global fits of experimental data using a pre-assumed functional form inspired from phenomenological models to learn its parameters. This novel approach uses a ML technique, namely symbolic regression, to learn an analytical model from measured charged hadron multiplicities. The function learned by symbolic regression resembles the Lund string function and describes the data well, thus representing a potential candidate for use in global FFs fits. This study represents an approach to follow in such QCD-related phenomenology studies and more generally in sciences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic Extraction of Non-Perturbative Transverse-Momentum-Dependent Distributions from Drell-Yan Data

    hep-ph 2026-07 conditional novelty 6.0 of 10

    Symbolic regression on a factorized NN fit to 482 Drell–Yan points yields a 9-constant analytical non-perturbative TMD with χ²/ndf≈1.04 and a retained x–b_T cross term.

  2. Generalized Parton Distributions from Symbolic Regression

    hep-ph 2025-04 conditional novelty 6.0 of 10

    Symbolic regression fits to lattice QCD and model GPDs show that the isovector GPD H_{u-d} approximately factorizes in x and t in the trained kinematic region, and a new Taylor-coefficient clustering criterion (ECC) g...

Reference graph

Works this paper leans on

44 extracted references · 29 canonical work pages · cited by 2 Pith papers

  1. [1]

    Hadronization

    Webber BR. Hadronization. In: Summer School on Hadronic Aspects of Collider Physics; 1994. p. 49-77

  2. [2]

    The ALICE experiment at the CERN LHC

    Aamodt K, et al. The ALICE experiment at the CERN LHC. JINST. 2008;3:S08002

  3. [3]

    The COMPASS experiment

    Schmitt L. The COMPASS experiment. In: 29th International Conference on High-Energy Physics

  4. [4]

    The CMS Experiment at the CERN LHC

    Chatrchyan S, et al. The CMS Experiment at the CERN LHC. JINST. 2008;3:S08004

  5. [5]

    The ATLAS Experiment at the CERN Large Hadron Collider

    Aad G, et al. The ATLAS Experiment at the CERN Large Hadron Collider. JINST. 2008;3:S08003

  6. [6]

    Global analysis of fragmentation functions for pions and kaons and their uncertainties

    de Florian D, Sassot R, Stratmann M. Global analysis of fragmentation functions for pions and kaons and their uncertainties. Phys Rev D. 2007 Jun;75:114010. Available from: https://link.aps.org/ doi/10.1103/PhysRevD.75.114010

  7. [7]

    Pion fragmentation functions at high energy colliders

    Borsa I, Sassot R, de Florian D, Stratmann M. Pion fragmentation functions at high energy colliders. Phys Rev D. 2022 Feb;105:L031502. Available from: https://link.aps.org/doi/10.1103/PhysRevD. 105.L031502

  8. [8]

    Asymptotic freedom in parton language

    Altarelli G, Parisi G. Asymptotic freedom in parton language. Nuclear Physics B. 1977;126(2):298-318. Available from: https://www.sciencedirect.com/science/article/pii/0550321377903844

Show all 44 references
  1. [9]

    Interpretable scientific discovery with symbolic regression: a review

    Makke N, Chawla S. Interpretable scientific discovery with symbolic regression: a review. Artificial Intelligence Review. 2024 Jan;57. Available from: https://doi.org/10.1007/s10462-023-10622-0

  2. [10]

    Symbolic Regression: A Pathway to Interpretability Towards Automated Scientific Discovery

    Makke N, Chawla S. Symbolic Regression: A Pathway to Interpretability Towards Automated Scientific Discovery. KDD ’24. New York, NY, USA: Association for Computing Machinery; 2024. p. 6588–6596. Available from: https://doi.org/10.1145/3637528.3671464

  3. [11]

    Rediscovering orbital mechanics with machine learning

    Lemos P, Jeffrey N, Cranmer M, Ho S, Battaglia P. Rediscovering orbital mechanics with machine learning. arXiv; 2022. Available from: https://arxiv.org/abs/2202.02306

  4. [12]

    Robust learning from noisy, incomplete, high- dimensional experimental data via physically constrained symbolic regression

    Reinbold PAK, Kageorge LM, Schatz MF, Grigoriev RO. Robust learning from noisy, incomplete, high- dimensional experimental data via physically constrained symbolic regression. Nature Communications. 2021 May;12:3219. Available from: https://doi.org/10.1038/s41467-021-23479-0

  5. [13]

    Data-driven discovery of Tsallis-like distribution using symbolic regression in high-energy physics

    Makke N, Chawla S. Data-driven discovery of Tsallis-like distribution using symbolic regression in high-energy physics. PNAS Nexus. 2024 10:pgae467. Available from: https://doi.org/10.1093/ pnasnexus/pgae467

  6. [14]

    Symbolic Regression on FPGAs for Fast Machine Learning Inference

    Tsoi HF, Pol AA, Loncar V, Govorkova E, Cranmer M, Dasu S, et al. Symbolic Regression on FPGAs for Fast Machine Learning Inference. EPJ Web of Conferences. 2024;295:09036. Available from: http: //dx.doi.org/10.1051/epjconf/202429509036

  7. [15]

    Parton fragmentation functions

    Metz A, Vossen A. Parton fragmentation functions. Progress in Particle and Nuclear Physics. 2016 Nov;91:136–202. Available from: http://dx.doi.org/10.1016/j.ppnp.2016.08.003

  8. [16]

    Parton fragmentation and string dynamics

    Andersson B, Gustafson G, Ingelman G, Sj¨ ostrand T. Parton fragmentation and string dynamics. Physics Reports. 1983;97(2):31-145. Available from: https://www.sciencedirect.com/science/article/ pii/0370157383900807

  9. [17]

    Pion and kaon fragmentation functions at next- to-next-to-leading order

    Abdul Khalek R, Bertone V, Khoudli A, Nocera ER. Pion and kaon fragmentation functions at next- to-next-to-leading order. Phys Lett B. 2022;834:137456

  10. [18]

    Helicity-dependent parton distribution functions at next-to-next-to- leading order accuracy from inclusive and semi-inclusive deep-inelastic scattering data

    Bertone V, Chiefa A, Nocera ER. Helicity-dependent parton distribution functions at next-to-next-to- leading order accuracy from inclusive and semi-inclusive deep-inelastic scattering data. 2024 4

  11. [19]

    The path to proton structure at 1% accuracy

    Ball RD, et al. The path to proton structure at 1% accuracy. Eur Phys J C. 2022;82(5):428. 16

  12. [20]

    Convolutional Networks for Images, Speech, and Time-Series; 1995

    Lecun Y, Bengio Y. Convolutional Networks for Images, Speech, and Time-Series; 1995

  13. [21]

    Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network

    Sherstinsky A. Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network. Physica D: Nonlinear Phenomena. 2020 Mar;404:132306. Available from: http: //dx.doi.org/10.1016/j.physd.2019.132306

  14. [22]

    Playing Atari with Deep Reinforcement Learning; 2013

    Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, et al.. Playing Atari with Deep Reinforcement Learning; 2013

  15. [23]

    Attention Is All You Need

    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention Is All You Need. CoRR. 2017;abs/1706.03762. Available from: http://arxiv.org/abs/1706.03762

  16. [25]

    Is machine learning good or bad for the natural sciences?; 2024

    Hogg DW, Villar S. Is machine learning good or bad for the natural sciences?; 2024. Available from: https://arxiv.org/abs/2405.18095

  17. [26]

    A perspective on symbolic machine learning in physical science; 2024

    Makke N, Chawla S. A perspective on symbolic machine learning in physical science; 2024. NeurIPS workshop on machine learning and the physical sciences, Vancouver. Available from: https:// ml4physicalsciences.github.io/2024/

  18. [27]

    Machine learning and the physical sciences

    Carleo G, Cirac I, Cranmer K, Daudet L, Schuld M, Tishby N, et al. Machine learning and the physical sciences. Rev Mod Phys. 2019 Dec;91:045002. Available from: https://link.aps.org/doi/10.1103/ RevModPhys.91.045002

  19. [28]

    BACON: A PRODUCTION SYSTEM THAT DISCOVERS EMPIRICAL LA WS; 1979

    W LP. BACON: A PRODUCTION SYSTEM THAT DISCOVERS EMPIRICAL LA WS; 1979. Available from: https://www.ijcai.org/Proceedings/77-1/Papers/057.pdf

  20. [29]

    Scientific Discovery: Computational Explorations of the Creative Process

    Langley P, Simon HA, Bradshaw GL, Zytkow JM. Scientific Discovery: Computational Explorations of the Creative Process. Cambridge, MA, USA: MIT Press; 1987

  21. [30]

    Determining Arguments of Invariant Functional Descriptions

    Kokar MM. Determining Arguments of Invariant Functional Descriptions. Machine Learning. 1986 Dec;1:403-22

  22. [31]

    Data-driven approaches to empirical discovery

    Langley P, Zytkow JM. Data-driven approaches to empirical discovery. Artificial Intelli- gence. 1989;40(1):283-312. Available from: https://www.sciencedirect.com/science/article/pii/ 0004370289900519

  23. [32]

    Discovery of equations: experimental evaluation of convergence

    Zembowicz R, ˙ Zytkow JM. Discovery of equations: experimental evaluation of convergence. In: Proceedings of the Tenth National Conference on Artificial Intelligence. AAAI’92. AAAI Press; 1992. p. 70–75

  24. [33]

    Declarative Bias in Equation Discovery

    Todorovski L, Dzeroski S. Declarative Bias in Equation Discovery. In: Proceedings of the Fourteenth International Conference on Machine Learning. ICML ’97. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc.; 1997. p. 376–384

  25. [34]

    Genetic programming as a means for programming computers by natural selection

    Koza JR. Genetic programming as a means for programming computers by natural selection. Satistics and Computing. 1994 Jun;4(2):87-112. Available from: https://doi.org/10.1007/BF00175355

  26. [35]

    Multiplicities of charged pions and charged hadrons from deep-inelastic scattering of muons off an isoscalar target

    Adolph C, et al. Multiplicities of charged pions and charged hadrons from deep-inelastic scattering of muons off an isoscalar target. Phys Lett B. 2017;764:1-10

  27. [36]

    Multiplicities of charged kaons from deep-inelastic muon scattering off an isoscalar target

    Adolph C, et al. Multiplicities of charged kaons from deep-inelastic muon scattering off an isoscalar target. Phys Lett B. 2017;767:133-41

  28. [37]

    The COMPASS experiment at CERN

    Abbon P, et al. The COMPASS experiment at CERN. Nucl Instrum Meth A. 2007;577:455-518

  29. [38]

    Charged hadron fragmentation functions at high energy colliders

    Borsa I, Stratmann M, de Florian D, Sassot R. Charged hadron fragmentation functions at high energy colliders. Phys Rev D. 2024 Mar;109:052004. Available from: https://link.aps.org/doi/10.1103/ PhysRevD.109.052004. 17

  30. [39]

    Hierarchical Genetic Algorithms Operating on Populations of Computer Programs

    Koza JR. Hierarchical Genetic Algorithms Operating on Populations of Computer Programs. In: Pro- ceedings of the 11th International Joint Conference on Artificial Intelligence - Volume 1. IJCAI’89. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc.; 1989. p. 768–774

  31. [40]

    Jan Lukasiewicz: Aristotle’s Syllogistic from the Standpoint of Modern Formal Logic

    Robinson R. Jan Lukasiewicz: Aristotle’s Syllogistic from the Standpoint of Modern Formal Logic. Second edition enlarged. Pp. xvi 222. Oxford: Clarendon Press, 1957. Cloth, 305. net. The Classical Review. 1958;8(3-4):282–282

  32. [41]

    Symbolic Regression is NP-hard; 2022

    Virgolin M, Pissis SP. Symbolic Regression is NP-hard; 2022

  33. [42]

    A living Review of Symbolic Regression; 2022

    Makke N, Chawla S. A living Review of Symbolic Regression; 2022

  34. [43]

    Neural Symbolic Regression that Scales

    Biggio L, Bendinelli T, Neitz A, Lucchi A, Parascandolo G. Neural Symbolic Regression that Scales. CoRR. 2021;abs/2106.06427. Available from: https://arxiv.org/abs/2106.06427

  35. [44]

    Attention is All You Need

    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is All You Need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS’17. Red Hook, NY, USA: Curran Associates Inc.; 2017. p. 6000–6010

  36. [45]

    Parton distributions for the LHC

    Martin AD, Stirling WJ, Thorne RS, Watt G. Parton distributions for the LHC. The European Physical Journal C. 2009 Jul;63(2):189–285. Available from: http://dx.doi.org/10.1140/epjc/ s10052-009-1072-5 . 18

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.