Pith. sign in

REVIEW 4 major objections 5 minor 36 references

A single number derived from the attractor's invariant measure sets the ceiling on what any equation-discovery algorithm can recover from trajectory data.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:17 UTC pith:CAYPRKKE

load-bearing objection The regime-ordering result is real and the formal core is careful, but the paper advertises a full-moment-matrix ceiling while actually proving a Schur-complement version, and the abstract needs rewording more than the framework needs rebuilding. the 4 major comments →

arxiv 2607.18490 v1 pith:CAYPRKKE submitted 2026-07-20 cs.LG nlin.CD

Attractor Geometry Determines the Identifiability Limits of System Discovery

classification cs.LG nlin.CD MSC 37M1037A30
keywords identifiability ceilingsymbolic equation discoverySINDyevolutionary symbolic regressioninvariant measuremoment matrixattractor coverageLorenz system
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the difficulty of discovering governing equations from data is not primarily an algorithmic problem but an information-geometric one: the attractor itself determines how identifiable the system is. Its central claim is that one scalar, the smallest eigenvalue of the moment matrix built from the candidate dictionary and the invariant measure, sets a pre-computable ceiling for both sparse regression (SINDy) and evolutionary symbolic regression (PySR). Where that eigenvalue vanishes, as at a fixed point, recovery is impossible for any algorithm, sparse or combinatorial; where it is positive but small, it fixes how much noise and data volume the problem can tolerate. The paper shows that chaos usually raises this eigenvalue by spreading the attractor, but also amplifies noise, and because noise enters the two algorithms at different powers, deeper chaos can help SINDy while hurting PySR. This reframes the practical question of equation discovery: not which algorithm, but what the attractor permits.

Core claim

The central discovery is that the identifiability of a symbolic-discovery problem is governed by the smallest eigenvalue of the invariant-measure moment matrix, M = integral of Phi(x)Phi(x)^T d mu(x). Using a within-system design on Lorenz-84, where a single forcing parameter runs the system through fixed-point, limit-cycle, and chaotic regimes while the equations, libraries, and protocols stay fixed, the authors show that this single number orders recovery difficulty for both SINDy and PySR across data volume, noise, and structural prior. At a fixed point the moment matrix is rank one, the smallest eigenvalue is zero, and the design matrix has spark at most two, so no sparse or combinatoria

What carries the argument

The central object is the invariant-measure moment matrix M = integral of Phi(x)Phi(x)^T d mu(x), where Phi is the dictionary of candidate functions and mu is the long-run probability distribution over the attractor; its smallest eigenvalue is lambda_min(M). Its finite-sample counterpart is (1/N) Theta^T Theta, and the Birkhoff ergodic theorem guarantees convergence of the empirical matrix to M. This eigenvalue measures how fully the attractor covers dictionary function space, and the paper uses it as a pre-computable, algorithm-independent identifiability ceiling. A Schur complement of the overcomplete moment matrix plays the same role for PySR's conditioning bottleneck, and a spark argumen

Load-bearing premise

The load-bearing premise is that a short reference trajectory already samples the invariant measure well enough that the empirical smallest eigenvalue is a stable, regime-level property rather than a trajectory-length artifact; the paper's finite-time guarantee for chaos assumes exponential decay of correlations, which it explicitly notes has not been independently verified for the Lorenz-84/96 chaotic regimes studied.

What would settle it

Take the Lorenz-84 system in a slow-mixing chaotic regime (for example, F just past the onset of chaos) and compute the smallest eigenvalue of the empirical moment matrix from reference trajectories of length 10, 100, and 1000 time units; if it varies by more than an order of magnitude between 100 and 1000 units, the short-reference-trajectory pre-run diagnostic collapses. Alternatively, construct a fixed-point regime where the smallest eigenvalue is zero and run any equation-discovery algorithm on noiseless data with a dictionary containing at least two nonzero terms: any successful coefficie

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the smallest eigenvalue of the moment matrix vanishes, no algorithm, sparse or combinatorial, can recover a single coefficient from that attractor, so running discovery there is pointless at any data volume.
  • A short reference trajectory plus one SVD yields the smallest eigenvalue before any run, telling the experimenter whether recovery is possible and how much noise or data the problem tolerates.
  • Deepening chaos improves conditioning but also amplifies noise; SINDy benefits when the conditioning gain outweighs a linear noise cost, while PySR requires a conditioning gain more than an order of magnitude larger, so deeper chaos is not universally beneficial.
  • A sharper structural prior, which reduces wrong-term contamination, raises the relevant Schur-complement eigenvalue for PySR and grows more valuable as noise worsens, making prior quality a noise-independent lever.
  • The parameter-free mechanistic scores built entirely from Lorenz-84 transfer without refitting to Lorenz-96, indicating the mechanism is general rather than curve-fitting.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The ceiling argument implies that adding transient or perturbation data, not just settled attractor data, can lift the smallest eigenvalue above the attractor floor, since transients sample regions the invariant measure under-weights; this is a testable extension of the paper's own Proposition S1.
  • If the empirical smallest eigenvalue proves unstable across short trajectory lengths in slow-mixing regimes, for example near bifurcation onset, the pre-run diagnostic would need explicit mixing-time checks; the paper's finite-time bound relies on exponential decay of correlations, which it notes has not been independently verified for the Lorenz-84/96 chaotic regimes.
  • The same moment-matrix ceiling should constrain black-box neural discovery methods as well, since any unbiased estimator must respect the same Cramer-Rao floor; testing neural ODEs against the smallest eigenvalue would be a natural next step.
  • The framework suggests a design rule: when a system can be steered, choose operating points that maximize the smallest eigenvalue while keeping state amplitude in check, turning experimental design into an optimization that could be automated with ergodic perturbation theory.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper studies equation discovery from trajectory data using a within-system design on Lorenz-84, where the forcing parameter F moves the system through fixed-point, limit-cycle, and chaotic regimes while the equations and candidate-term library stay fixed. It proposes that the smallest eigenvalue λmin(M) of the invariant-measure moment matrix M (Eq. 6) is a universal, algorithm-independent identifiability ceiling: where λmin(M)=0 recovery is allegedly impossible for any algorithm, and as λmin(M) grows both SINDy and PySR improve. From this matrix the authors derive two 'mechanistic scores,' FSINDy and FPySR, and validate them on a held-out Lorenz-96 system and on non-polynomial jerk systems. They also introduce Soft F1, a coefficient-weighted structural metric. The formal core includes a spark argument for fixed-point impossibility (Prop. S2), a Cramér–Rao-style floor (Prop. S4), and a finite-time concentration result (Prop. S3). The empirical sections show the expected regime ordering across data volume, noise, and prior quality, with the claimed noise-channel asymmetry between SINDy and PySR.

Significance. If the central claim were established, this would be an important conceptual contribution: replacing algorithm-centric benchmarking with a pre-computable, attractor-geometric identifiability ceiling. The paper has real strengths: the fixed-point spark argument (Prop. S2) is self-contained and rigorous; the Schur-complement quantity in Prop. S4 is the right type of object; the held-out L96 and jerk-system transfers are genuine zero-refit checks; SI 6 directly addresses the configuration-dependence of the limit-cycle ceiling; and SI 5 includes several honest robustness controls (normalization, clean-reference σmin, partial correlations). The Soft F1 metric is useful. However, the advertised single-number ceiling is not the quantity proven: the formal floor uses λmin(Mnc|aw), after projecting out wrong terms, while the abstract and main text repeatedly cite λmin(M). This mismatch is load-bearing rather than cosmetic, because wrong-term-only degeneracies can make λmin(M)=0 without making the ground-truth coefficients unidentifiable. The practical pre-run diagnostic also rests on an unverified mixing assumption. These issues are corrigible, but they change the message and the required c

major comments (4)
  1. [Abstract; §Connection to the moment matrix, Eq. (6)–(7); SI 3, Prop. S4, Eq. (S10)] The headline claim is that λmin(M), the smallest eigenvalue of the full moment matrix, is the algorithm-independent ceiling and that 'where it vanishes, recovery is impossible for any algorithm.' The formal floor proven in Prop. S4/Eq. (S10) is λmin(Mnc|aw), the Schur complement after eliminating wrong dictionary terms, not λmin(M). These are different objects. λmin(M)=0 can be caused entirely by degeneracies among wrong terms—for example, a duplicated or nearly constant wrong term—while the ground-truth subspace remains well conditioned and the active coefficients are identifiable by least squares or sparse regression. Prop. S2 covers only the fixed-point full-dictionary rank collapse, and Prop. S6 covers exact collinearity involving ground-truth terms; neither supports the general 'λmin(M)=0 ⇒ impossibility' statement. The framework survives replacing λmin(M) by λmin(Mnc|aw) (or a mini
  2. [SI 3, Prop. S3(iii) and its Remark; Abstract/Eq. (6) 'short reference trajectory'] The practical promise that λmin(M) can be read 'from a short reference trajectory before any run' depends on finite-time concentration of the empirical moment matrix. For chaotic regimes, Prop. S3(iii) assumes exponential decay of correlations, and the SI explicitly states this hypothesis 'has not been independently verified for the L84/L96 chaotic regimes studied here.' Without verified mixing, a short trajectory can misestimate λmin(M) by an amount that is uncontrolled—especially for L84, where lobe-switching timescales are slow relative to typical recording windows. This does not invalidate the asymptotic ceiling, but it undermines the pre-run diagnostic and the within-chaos comparisons built on five regimes per system. The authors should either verify the mixing assumption numerically (e.g., autocorrelation-time and block-bootstrap estimates) or restrict the pre-run claims to the asy
  3. [SI 2, Functional-form selection; Main text 'Parameter-free mechanistic indicators'] The paper calls FSINDy and FPySR 'parameter-free mechanistic scores,' but SI 2 states that FPySR's outer 1/4 power and geometric-mean combination 'were selected by maximizing mean Spearman correlation on L84 across all three experiments.' Thus the scores are free of fitted constants but not free of form selection tuned to the outcomes. The sensitivity grid shows that all members of the considered family behave similarly, and the L96 transfer is a genuine held-out check, which mitigates the concern. Still, the main text's claim that the scores 'contain no fitted parameters' and 'confirm mechanism rather than curve-fitting' is overstated. The distinction between derived channels and empirically selected functional form should be stated in the main text, not only in the SI.
  4. [Methods/PySR; SI 2, Eq. (7)] The PySR conditioning score σmin,partial and the associated ceiling λmin(Mnc|aw) require a concrete finite dictionary of 'wrong terms' against which each ground-truth term is projected. PySR's actual search space is the operator set {+,−,×} with a complexity cap, which is unbounded in expression count. The paper does not specify the wrong-term dictionary used to compute σmin,partial for PySR, nor how higher-complexity or non-polynomial candidate expressions are excluded or accounted for. If σmin,partial is computed over a degree-3 polynomial dictionary (as the examples in the main text suggest), the resulting ceiling does not necessarily govern PySR's full search space. This is a reproducibility gap and a limitation of the claimed PySR-specific ceiling. Please specify the construction and discuss the restriction.
minor comments (5)
  1. [Results, first paragraph] Typo: 'the fixed-point regime remains show a low for SINDy' should be 'remains low' or 'remains shown to be low.'
  2. [Fig. 2 caption] The caption labels the prior-quality panels as '(C, prior quality)' although the main text refers to Fig. 2D and 2H for prior quality. The panel letters should be checked for consistency.
  3. [Methods, Derivative estimation] 'noise-to-signal ratio' should be 'signal-to-noise ratio' to match η's definition and standard terminology.
  4. [Table S1 / SI 5 consistency] The text around Table S1 mentions 'ground-truth The natural bimodality...' with a stray 'ground-truth.' Remove the artifact.
  5. [Eq. (1)] The definition of δj uses |cj|+|ĉj| in the denominator; if both coefficients are zero the term is absent from both expressions and the weight is presumably not computed. Clarify the convention for terms absent from both.

Circularity Check

1 steps flagged

Mostly self-contained derivation; one disclosed functional-form fit makes the L84 PySR validation partially in-sample, while L96 and the λmin theorems remain independent.

specific steps
  1. fitted input called prediction [SI 2, Functional-form selection; main text Eq. (5) and Table 1]
    "Two further choices are not derived and are stated as such: the outer 1/4 power applied to σmin,partial, and the geometric-mean combination with Qnoise. Both were selected by maximizing mean Spearman correlation on L84 across all three experiments, before L96 was ever consulted"

    The PySR score's functional form was selected by maximizing the same per-dimension-averaged Spearman correlations that Table 1 reports for L84 (0.78–0.86). Thus the L84 FPySR row is an in-sample fit to the validation target, not an independent prediction; the 'mechanism rather than curve-fitting' claim for L84 is partly circular by construction. The L96 transfer is a genuine held-out test (the form was frozen before L96), so the circularity is confined to the L84 leg and does not infect the central regime-ordering/λmin theorem.

full rationale

The derivation chain is largely self-contained: λmin(M) is defined via Birkhoff (Eq. 6), and the impossibility/floor results (Props S2, S4, S6) are proven from linear algebra and ergodic theory rather than fitted to outcomes. The mechanistic scores are built from data-matrix singular values and dictionary residuals, not from soft-F1 labels; correlating them with soft F1 is a legitimate mechanistic consistency check. The one construction that reduces to its own validation target is the PySR functional form: SI 2 explicitly states the outer 1/4 power and geometric mean were selected on L84 by maximizing the same Spearman correlations later reported. This is disclosed and is followed by a real held-out L96 test, so I score it moderate rather than severe. The abstract's λmin(M)=0 universal-impossibility phrasing is sharper than the propositions, which use λmin(Mnc|aw) or the fixed-point spark argument; that mismatch is a correctness/scope issue, not circularity, because the paper does prove the Schur-complement floor at Eq. S10/Eq. 7. No load-bearing self-citation or ansatz-via-citation was found.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

Everything the central claim rests on that the reader didn't pay for upstream. No new physical entities are introduced: λmin(M), FSINDy, FPySR, and Soft F1 are derived quantities/metrics, not postulated entities. The main pulled-in items are standard ergodic theory and the domain assumptions of an adequate dictionary, full-state observation, and an unverified mixing hypothesis for the chaotic finite-time bounds (self-flagged in SI 3). The single most consequential fitted element is the FPySR functional form, selected on L84 outcomes.

free parameters (4)
  • FPySR outer functional form (exponent and combination rule) = 1/4 power on σmin,partial; geometric mean with Qnoise; outer square root
    Selected on L84 by maximizing mean Spearman correlation before L96 was consulted (SI 2); sensitivity grid shows plateau Δ|ρ| ≤ 0.013 on L84, 0.042 on L96. Disclosed, but it is a fit to discovery outcomes on the training system.
  • Soft F1 coefficient-fidelity scale α = 3
    Set by the convention that a 2× coefficient mismatch costs one nat (δ=1/3, w=e^-1). A metric-design choice, not fitted to outcomes; regime ordering is insensitive to α ∈ {1.5, 3, 6} (SI 5).
  • STLSQ hyperparameters (λ, α) = λ = 0.05, ridge α = 0.05
    Fixed a priori across all regimes (Methods); enters FSINDy explicitly. SI 6 shows the R1 ceiling is configuration-dependent (oracle tuning recovers soft F1 = 1.0) and that the score tracks outcomes across a 2,244-cell grid.
  • Regime classification thresholds and LC amplitude split = λ1 band ±0.01/0.02; LC split at largest gap in σx
    Hand-chosen grid binnings structure all experiments; the R1/R2 limit-cycle split is data-driven (bimodality of σx). Disclosed in Methods and Table S1.
axioms (6)
  • standard math Birkhoff ergodic theorem — time averages along a generic trajectory converge to invariant-measure expectations
    Basis for (1/N)ΘᵀΘ → M (Eq. 6) and all moment-matrix asymptotics. Standard, but its application to numerically integrated trajectories assumes a unique physical invariant measure is being sampled.
  • domain assumption Ergodicity of the simulated flows with a unique physical invariant measure per regime
    Regime-stratified slots are treated as samples from one invariant measure per F; multi-stability or multiple basins would break the single-M picture. Not verified beyond the Lyapunov-based classification.
  • domain assumption Chaotic regimes support SRB measures with exponential decay of correlations
    Required for the finite-time concentration bound of Prop S3(iii); the paper explicitly states this has not been independently verified for L84/L96 (SI 3, Remark).
  • domain assumption Dictionary contains the ground-truth terms (default/overcomplete prior)
    The ceiling claim presupposes an adequate library; the null-prior experiment collapses all regimes to baseline, and the paper notes this is 'not explicable by any conditioning floor' (Results).
  • domain assumption Full state observation; additive i.i.d. Gaussian measurement noise
    Stated in limitations (Discussion); noise model underlies TFP, Qnoise, and the Cramér-Rao floor (Prop S4). Finite-difference derivative noise scaling σε ≈ ησx/√(2 dt²) is assumed.
  • standard math Weyl's inequality, Cramér-Rao bound, Donoho-Elad spark theorem
    Used for σmin(Θ) → √N λmin(M), the necessary-condition floor, and the spark-collapse argument (SI 1, SI 3).

pith-pipeline@v1.3.0-alltime-deepseek · 45005 in / 30050 out tokens · 285600 ms · 2026-08-01T15:17:15.629009+00:00 · methodology

0 comments
read the original abstract

Symbolic discovery of governing equations from data is limited not only by algorithm design and data volume, but by the geometry of the attractor: what the long-run dynamics allow to be recovered. Using a within-system design on Lorenz-84, where one forcing parameter drives fixed-point, limit-cycle, and chaotic regimes while the governing equations and library stay fixed, we show that a single number, $\lambda_{\min}(M)$, the smallest eigenvalue of the invariant-measure moment matrix, sets the identifiability ceiling for both sparse regression (SINDy) and evolutionary symbolic regression (PySR). Derived from the Birkhoff ergodic theorem and obtained from a short reference trajectory before any run, $\lambda_{\min}(M)$ measures how fully the attractor covers function space: where it vanishes, recovery is impossible for any algorithm, sparse or combinatorial alike; as it grows, both algorithms improve. Chaos raises $\lambda_{\min}(M)$ by spreading the attractor, but also enlarges it and amplifies noise; because noise enters SINDy's regression bottleneck linearly and PySR's discrimination channel superlinearly, the same transition can push the two methods in opposite directions, so deeper chaos is not uniformly better. Parameter-free mechanistic scores from this framework transfer without refitting to a held-out Lorenz-96 system, confirming mechanism rather than curve-fitting; a criterion read from the equations predicts when added chaos will not improve conditioning. We also introduce Soft F1, a coefficient-weighted structural metric that resolves performance differences invisible to binary-success and predictive scores. The first question of discovery is then not which algorithm, but what the attractor permits.

Figures

Figures reproduced from arXiv: 2607.18490 by Fabio Anselmi, Matteo Gallo, Paolo Lazzari.

Figure 1
Figure 1. Figure 1: Attractor coverage controls the conditioning and outcome of symbolic discovery. Five dynamical regimes of the same governing equations are shown from bottom to top in order of increasing attractor coverage of state space. As coverage grows, state-space trajectories (first column) fill phase space more thoroughly; the identification problem becomes better conditioned (second column), with spurious symmetric… view at source ↗
Figure 2
Figure 2. Figure 2: The dynamical regime ordering is broadly preserved across algorithms and experi￾mental conditions, with systematic inversions at data extremes and high noise. Rows: SINDy (top) and PySR (bottom). Background shading identifies regime families: grey = fixed point (R0), blue = limit cycles (R1–R2), red = chaos (R3–R7). (A, starvation) Soft F1 vs. training-set size N (log-spaced 20–10,000; dark: large N; light… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 4 canonical work pages

  1. [1]

    SRBench++: Principled benchmarking of symbolic regression with domain- expert interpretation.IEEE Transactions on Evolutionary Computation, 29:1127–1134, 2024

    Fabrício Olivetti de França, Marco Virgolin, Michael Kommenda, Manzur Majumder, Miles Cranmer, et al. SRBench++: Principled benchmarking of symbolic regression with domain- expert interpretation.IEEE Transactions on Evolutionary Computation, 29:1127–1134, 2024. doi: 10.1109/tevc.2024.3423681

  2. [2]

    Brunton, Joshua L

    Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Discovering governing equations from data by sparse identification of nonlinear dynamical systems.Proceedings of the National Academy of Sciences, 113(15):3932–3937, 2016. doi: 10.1073/pnas.1517384113. 15

  3. [3]

    Interpretable machine learning for science with PySR and SymbolicRegres- sion.jl.arXiv preprint arXiv:2305.01582, 2023

    Miles Cranmer. Interpretable machine learning for science with PySR and SymbolicRegres- sion.jl.arXiv preprint arXiv:2305.01582, 2023

  4. [4]

    Distilling free-form natural laws from experimental data

    Michael Schmidt and Hod Lipson. Distilling free-form natural laws from experimental data. Science, 324(5923):81–85, 2009

  5. [5]

    Nathan Kutz, and Steven L

    Kathleen Champion, Bethany Lusch, J. Nathan Kutz, and Steven L. Brunton. Data-driven discovery of coordinates and governing equations.Proceedings of the National Academy of Sciences, 116(45):22445–22451, 2019. doi: 10.1073/pnas.1906995116

  6. [6]

    Mangan, Steven L

    Niall M. Mangan, Steven L. Brunton, Joshua L. Proctor, and J. Nathan Kutz. Inferring biological networks by sparse identification of nonlinear dynamics.IEEE Transactions on Molecular, Biological and Multi-Scale Communications, 2(1):52–63, 2016. doi: 10.1109/TMBMC.2016. 2633265

  7. [7]

    Discovering symbolic models from deep learning with inductive biases

    Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho. Discovering symbolic models from deep learning with inductive biases. InAdvances in Neural Information Processing Systems, volume 33, pages 17429–17442,

  8. [8]

    Champion, Steven L

    Kathleen P. Champion, Steven L. Brunton, and J. Nathan Kutz. Discovery of nonlinear multiscale systems: Sampling strategies and embeddings.SIAM Journal on Applied Dynamical Systems, 18(1):312–347, 2019. doi: 10.1137/17M115477X

  9. [9]

    Multi-objective SINDy for parameterized model discovery from single transient trajectory data.Nonlinear Dynamics, 113:10911–10927, 2024

    José Antonio Lemus and Björn Herrmann. Multi-objective SINDy for parameterized model discovery from single transient trajectory data.Nonlinear Dynamics, 113:10911–10927, 2024. doi: 10.1007/s11071-024-10825-2

  10. [10]

    On the persistency of excitation

    Ivan Markovsky, Eduardo Prieto-Araujo, and Florian Dörfler. On the persistency of excitation. Automatica, 147:110657, 2022. doi: 10.1016/j.automatica.2022.110657

  11. [11]

    Dover Publications, 2nd edition, 2008

    Karl Johan Åström and Björn Wittenmark.Adaptive Control. Dover Publications, 2nd edition, 2008

  12. [12]

    Nathan Kutz

    Yuxuan Bao and J. Nathan Kutz. Information theory and discriminative sampling for model discovery.arXiv preprint arXiv:2512.16000, 2025. doi: 10.48550/arxiv.2512.16000

  13. [13]

    When is a system discoverable from data? Discovery requires chaos.arXiv preprint arXiv:2511.08860, 2025

    Zakhar Shumaylov, Peter Zaika, Philipp Scholl, Gitta Kutyniok, Lior Horesh, and Carola- Bibiane Schönlieb. When is a system discoverable from data? Discovery requires chaos.arXiv preprint arXiv:2511.08860, 2025

  14. [14]

    Exact recovery of chaotic systems from highly corrupted data.Multiscale Modeling & Simulation, 15(3):1108–1129, 2017

    Giang Tran and Rachel Ward. Exact recovery of chaotic systems from highly corrupted data.Multiscale Modeling & Simulation, 15(3):1108–1129, 2017. doi: 10.1137/16M1086637. arXiv:1607.01067

  15. [15]

    Extreme sampling in model identification via L1 optimization: The sparse regression cases.SIAM Journal on Applied Mathematics, 78(6): 3279–3295, 2018

    Hayden Schaeffer, Giang Tran, and Rachel Ward. Extreme sampling in model identification via L1 optimization: The sparse regression cases.SIAM Journal on Applied Mathematics, 78(6): 3279–3295, 2018. doi: 10.1137/17M1120792. arXiv:1707.08528

  16. [16]

    Extracting structured dynamical systems using sparse optimization with very few samples.Multiscale Modeling & Simulation, 18(4):1435–1461, 2020

    Hayden Schaeffer, Giang Tran, Rachel Ward, and Linan Zhang. Extracting structured dynamical systems using sparse optimization with very few samples.Multiscale Modeling & Simulation, 18(4):1435–1461, 2020. doi: 10.1137/18M1194730. arXiv:1805.04158

  17. [17]

    Recovery guarantees for polynomial approximation from dependent data with outliers.arXiv preprint arXiv:1811.10115, 2018

    Lam Si Tung Ho, Hayden Schaeffer, Giang Tran, and Rachel Ward. Recovery guarantees for polynomial approximation from dependent data with outliers.arXiv preprint arXiv:1811.10115, 2018

  18. [18]

    Kaptanoglu, Linan Zhang, Zachary G

    Alan A. Kaptanoglu, Linan Zhang, Zachary G. Nicolaou, Urban Fasel, and Steven L. Brunton. Benchmarking sparse system identification with low-dimensional chaos.Nonlinear Dynamics, 111:13143–13164, 2023. doi: 10.1007/s11071-023-08525-4

  19. [19]

    Chaos as an interpretable benchmark for forecasting and data-driven modelling

    William Gilpin. Chaos as an interpretable benchmark for forecasting and data-driven modelling. arXiv preprint arXiv:2110.05266, 2021. 16

  20. [20]

    Edward N. Lorenz. Irregularity: A fundamental property of the atmosphere.Tellus A, 36(2): 98–110, 1984

  21. [21]

    Edward N. Lorenz. Predictability: A problem partly solved. InProc. ECMWF Seminar on Predictability, Vol. 1, pages 1–18. ECMWF, 1996

  22. [22]

    Birkhoff

    George D. Birkhoff. Proof of the ergodic theorem.Proceedings of the National Academy of Sciences, 17(12):656–660, 1931

  23. [23]

    C. G. Broyden. The convergence of a class of double-rank minimization algorithms 1. general considerations.IMA Journal of Applied Mathematics, 6(1):76–90, 1970. doi: 10.1093/imamat/ 6.1.76

  24. [24]

    Delahunt and J

    Charles B. Delahunt and J. Nathan Kutz. A toolkit for data-driven discovery of governing equations in high-noise regimes.IEEE Access, 10:31210–31234, 2022. doi: 10.1109/ACCESS. 2022.3159335

  25. [25]

    Sparse identification of nonlinear dynamical systems via reweighted ℓ1-regularized least squares.Computer Methods in Applied Mechanics and Engineering, 376:113620, 2021

    Alexandre Cortiella, Kwang-Chun Park, and Alireza Doostan. Sparse identification of nonlinear dynamical systems via reweighted ℓ1-regularized least squares.Computer Methods in Applied Mechanics and Engineering, 376:113620, 2021. doi: 10.1016/j.cma.2020.113620

  26. [26]

    Rethinking symbolic regression datasets and benchmarks for scientific discovery.arXiv preprint arXiv:2206.10540, 2022

    Yoshitomo Matsubara, Naoya Chiba, Ryo Igarashi, Tatsunori Taniai, and Yoshitaka Ushiku. Rethinking symbolic regression datasets and benchmarks for scientific discovery.arXiv preprint arXiv:2206.10540, 2022. doi: 10.48550/arxiv.2206.10540

  27. [27]

    L. G. A. dos Reis, V . L. P. S. Caminha, and T. J. P. Penna. Benchmarking symbolic regression constant optimization schemes.arXiv preprint arXiv:2412.02126, 2024. doi: 10.48550/arxiv. 2412.02126

  28. [28]

    Nathan Kutz, and Steven L

    Kadierdan Kaheman, J. Nathan Kutz, and Steven L. Brunton. SINDy-PI: a robust algorithm for parallel implicit sparse identification of nonlinear dynamics.Proceedings of the Royal Society A, 476(2242):20200279, 2020. doi: 10.1098/rspa.2020.0279

  29. [29]

    Nathan Mundhenk, Mikel Landajuela, Ruben Glatt, Claudio P

    T. Nathan Mundhenk, Mikel Landajuela, Ruben Glatt, Claudio P. Santiago, Daniel M. Fais- sol, and Brenden K. Petersen. Symbolic regression via neural-guided genetic programming population seeding.arXiv preprint arXiv:2111.00053, 2021

  30. [30]

    Radhakrishna Rao

    C. Radhakrishna Rao. Information and accuracy attainable in the estimation of statistical parameters.Bulletin of the Calcutta Mathematical Society, 37:81–91, 1945

  31. [31]

    Jonathan Botvinick-Greenhouse, Robert T. W. Martin, and Yunan Yang. Invariant measures in time-delay coordinates for unique dynamical system identification.Physical Review Letters, 135(16):167202, 2024. doi: 10.1103/ppys-lx68

  32. [32]

    Lyapunov characteristic exponents for smooth dynamical systems and for Hamiltonian systems; a method for computing all of them

    Giancarlo Benettin, Luigi Galgani, Antonio Giorgilli, and Jean-Marie Strelcyn. Lyapunov characteristic exponents for smooth dynamical systems and for Hamiltonian systems; a method for computing all of them. Part 1: Theory.Meccanica, 15(1):9–20, 1980. doi: 10.1007/ BF02128236

  33. [33]

    Kaptanoglu, Brian M

    Alan A. Kaptanoglu, Brian M. de Silva, Urban Fasel, Kadierdan Kaheman, Andy J. Goldschmidt, Jared L. Callaham, Charles B. Delahunt, Zachary G. Zheng, Joshua Mann, J. Nathan Kutz, and Steven L. Brunton. PySINDy: A comprehensive Python package for robust sparse system identification.Journal of Open Source Software, 7(69):3994, 2022. doi: 10.21105/joss.03994

  34. [34]

    combined

    David L. Donoho and Michael Elad. Optimally sparse representation in general (nonorthogonal) dictionaries via ℓ1 minimization.Proceedings of the National Academy of Sciences, 100(5): 2197–2202, 2003. doi: 10.1073/pnas.0437847100. 17 Supporting Information Attractor Geometry Determines the Identifiability Limits of System Discovery Table S1 lists, for both...

  35. [36]

    In practice

    are the harder structural target for the overcomplete search. This is additional texture on top of, not a revision of, the system-level regime ordering in Fig. 2 and Table 1: the per-dimension means reported here are exactly what is averaged into every system-level number in the main text. 40 SI 6. R1 Hyperparameter Control: Noiseless Recovery and Noise C...

  36. [2020]

    doi: 10.48550/arxiv.2006.11287