REVIEW 3 major objections 6 minor 42 references
Deficiency of equation-finding approach to data-driven modeling of dynamical systems
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read For chaotic systems with imperfect data, sparse-optimization equation discovery is non-unique: many different recovered equations all reproduce the same attractor and coincide in their dominant Koopman eigenvalues.
desk verdict Real SINDy degeneracy under imputation, but 'virtually identical attractors' overstates the evidence, and the imputation pipeline may be loading the dice. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by three components working in sequence: machine-learning imputation of missing data that reconstructs a continuous time series; sparse optimization (SINDy) over a small library of monomials that selects a minimal set of active terms; and Koopman analysis via Ulam's discretization of the Perron–Frobenius operator, which yields the eigenvalue spectrum used to compare models. The load-bearing identity is the agreement of the leading Koopman eigenvalues across recovered equations, explained by on-off intermittency in the velocity-field difference between recovered and true dynamics.
What would settle it
Repeat the pipeline without imputation: run sparse optimization directly on the observed data segments (or on data repaired by a method trained on a different system) and compare the recovered attractors. If the equations no longer all reproduce the target attractor, the imputation step is doing the work.
Extended reading notes
Core claim
For the Lorenz system at 20%, 30%, and 50% random missing data, the recovered equation sets differ drastically from the true Lorenz equations and from each other, yet their attractors have KL distances near 0.4–0.5 to the true attractor and Lyapunov exponents close to (0.9, 0, -12). The Koopman eigenvalue spectra of recovered and original systems agree for several dozen dominant eigenvalues, with divergence only below an equation-dependent threshold. The mechanism is on-off intermittent disagreement of the velocity fields: for most points the recovered vector field nearly matches the true one, with rare large bursts, so long-term statistics are preserved. The paper's central claim is that th
Load-bearing premise
The machine-learning imputation that fills missing data points is trained on the same system and therefore may already encode the true attractor, so the agreement across recovered equations could be partly an artifact of the reconstruction rather than a general property of equation-finding.
Editorial extensions
If this is right
- Different recovered equations for the same chaotic system are not a sign of algorithmic failure; they are the expected outcome, so no single discovered equation should be interpreted as the physical law.
- Attractor statistics (Lyapunov exponents, KL divergence, leading Koopman eigenvalues) are too coarse to certify a model: distinct equations pass the same statistical tests.
- If subdominant Koopman modes matter for a question—responses where small eigenvalues play a role—models that agree on the dominant spectrum can still disagree.
- The paper's recommendation is to shift from equation extraction to direct, equation-free data-driven modeling for such systems.
Reading between the lines
- Because the imputation step uses the same system's data, it may already constrain the reconstruction to the target attractor; testing with cross-system imputation or no imputation would separate the imputation bias from a genuine property of equation-finding.
- The equivalence class of 'models that share an attractor' suggests a formal notion: two vector fields are observationally equivalent for a chaotic system if their leading Koopman spectra coincide; the equation-dependent threshold below which spectra split might be linked to the spectral gap or noise floor.
- The on-off intermittent velocity-field discrepancy implies that any fitting procedure that only penalises statistical attractor error—rather than pointwise vector-field error—will exhibit the same underdetermination, which may extend beyond SINDy to neural-network-based equation discovery.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that sparse-optimization equation discovery (SINDy) is fundamentally non-unique for chaotic systems when data contain missing segments. Using a transformer-based ML imputation method to reconstruct missing data and then applying SINDy, the authors recover structurally different equations for different missing-data ratios (20%, 30%, 50%) in the Lorenz system. They report that these different equations nevertheless generate similar attractors (KL distances 0.42–0.54, similar Lyapunov exponents) and agree in their dominant Koopman eigenvalues, with differences appearing only below an equation-dependent threshold. They attribute this to on–off intermittency in the velocity-field discrepancies. The paper concludes that equation-based modeling from imperfect chaotic data can be misleading and that direct data-driven, e.g., machine-learning, approaches may be preferable.
Significance. If the central claim were established, the result would be a significant caveat for the equation-discovery literature: it would show that sparse optimization can return multiple structurally distinct models with indistinguishable statistical behavior, undermining the interpretability of recovered equations. The paper usefully combines SINDy with Koopman analysis and tests several missing-data ratios, and the velocity-field intermittency diagnostic is a creative way to visualize the discrepancy. However, the current demonstration is vulnerable to a circularity concern: the ML imputation step is trained on the same system's data and may constrain the reconstructed trajectory to the true attractor, so the subsequent attractor agreement may be inherited from the reconstruction rather than being a property of equation finding. The reported KL distances and Lyapunov exponent variations also do not obviously support the 'virtually identical' wording. The paper's potential significance is high, but the evidence as presented is not yet conclusive.
major comments (3)
- [Fig. 1(b) and Sec. I of SI (imputation step)] The central claim is that structurally different recovered equations generate statistically identical attractors. However, the only pipeline considered imputes missing observations with the ML method of ref. [25], which is trained on the same system's data and therefore encodes the target attractor. SINDy is then fitted to a trajectory that is pointwise constrained to that attractor, so the observed agreement of attractors and Koopman spectra may be inherited from the reconstruction rather than being a property of equation finding. The manuscript does not report a control without imputation (e.g., SINDy applied to the observed gap-containing series) or with an imputation method blind to the true dynamics. This control is necessary to support the abstract's claim of a general deficiency of the equation-finding approach.
- [Fig. 2 and Abstract] The abstract and summary state that the recovered equations generate 'virtually identical chaotic attractors,' but the three reported KL distances are 0.52, 0.42, and 0.54, and the Lyapunov exponents vary between cases (e.g., contracting exponents -12.139 and -13.404). The ground-truth Lyapunov exponents are not stated in the caption, so 'agreeing with the ground truth' cannot be verified. With only three realizations of the missing-data pattern and no error bars or repeated-realization statistics in the main text, the claim of virtual identity is not supported by the numbers shown. The paper should either temper the wording or report the statistical distribution (mean/standard deviation) of KL distances and exponents over many realizations.
- [Koopman analysis, Fig. 2(b,d,f)] The Koopman spectra comparison is computed from long trajectories of the recovered equations and the original Lorenz system. Because the recovered trajectories are close to the true attractor by construction (see the imputation concern above), the agreement of the dominant eigenvalues may simply reflect the agreement of the invariant density and slow mixing modes, and carries no additional information about equation-level fidelity. To make the Koopman result informative, the authors should compare the recovered spectra against those of an ensemble of random dynamical systems constrained to the same invariant density, or show that the Ulam approximation is converged with respect to grid resolution. As written, the Koopman agreement is not independent evidence for the paper's thesis.
minor comments (6)
- [Introduction, first sentence] The phrase 'biblical importance' is informal and not appropriate for a journal report; please replace with a more neutral phrasing.
- [Introduction, paragraph 2] There is a typo: 'measurement of the measurement of the time series' should be 'measurement of the time series'.
- [Koopman section] The notation 'P = M^⊺' is potentially confusing. Please clarify whether M is row-stochastic or column-stochastic and define the orientation explicitly in the main text or in the SI.
- [Sec. III and Sec. IV of SI] The main text relies heavily on the SI for statistical analysis and for tests on additional systems, but the SI is not included in the submitted manuscript package. Please ensure the SI is available to the reviewers and readers.
- [Fig. 2 caption] The caption reports Lyapunov exponents for the recovered equations but does not provide the ground-truth Lorenz values. Please include them so the reader can judge the agreement quantitatively.
- [Sec. IVB] There is a formatting inconsistency: 'Sec. IVB' should be 'Sec. IV B' or 'Sec. IV-B'.
Circularity Check
Velocity-field agreement is partly manufactured by best-match selection; the central non-uniqueness finding still rests on independent diagnostics, but an imputation-based confound is not controlled.
-
fitted input called prediction
[Section 3 (unnumbered), paragraph beginning 'How is it that drastically different equations can produce the same chaotic attractor?' and Fig. 3]
"For point k on the reference attractor with velocity f k, we search and identify n points on the other attractor that are close to and compute their velocities {gj}n j=1. ... We thus select the point on those candidates that best aligns with f by maximizing the cosine similarity cos θjk = f k · gj/(∥f k∥∥gj∥), while also keeping the relative norm mismatch |∥gk∥ − ∥f j∥|/∥f k∥ as small as possible."
The velocity-field discrepancy is not measured at corresponding points; instead, among n nearby candidates, the one with maximum cosine similarity is chosen. On a fractal chaotic attractor, the maximum over many candidates is biased high even for uncorrelated fields, so the long 'off' phases and sharp zero peak in Fig. 3 are partly manufactured by the selection rule. The paper then uses this manufactured agreement as the explanation for why distinct equations share an attractor ('the velocity fields coincide to high accuracy across most of the attractor'), making the supporting evidence small by construction rather than an independent confirmation.
full rationale
The paper is an empirical demonstration, not a formal derivation, so most of its claims are not circular in the equation-level sense. The central observation that SINDy returns structurally different equations for different missing-data realizations is supported by the displayed equations themselves and does not reduce to an input. The claim that these equations generate nearly identical attractors is separately supported by KL distances, Lyapunov exponents, and Koopman spectra, which are independent diagnostics. However, two concerns limit the circularity score. First, the missing-data imputation step (ref. [25], by the same authors) may be trained on the same target system; if so, the imputed time series already lies on the true attractor, and the later attractor agreement is partly inherited from the reconstruction prior. The main text gives no no-imputation control or control with an imputer not informed by the same system, so this confound is not excluded. Second, the velocity-field comparison explicitly selects the best-matching point among n candidates, which makes the reported agreement and on-off intermittency partly an artifact of the matching rule. These are correctness risks and partial construction of one supporting diagnostic, but they do not make the central non-uniqueness result itself reduce to its inputs, so the score is moderate rather than high.
Assumptions & free parameters
free parameters (4)
- missing data ratio =
20%, 30%, 50%
- SINDy sparsity threshold and library
- Ulam grid resolution
- ML imputation hyperparameters
assumptions (4)
- domain assumption The target systems are deterministic chaotic systems with polynomial vector fields (Lorenz system and similar).
- domain assumption The ML imputation method preserves the invariant statistics of the true system.
- domain assumption SINDy with the chosen library can approximate the dynamics on the attractor from reconstructed data.
- standard math Ulam discretization approximates the Koopman/Perron-Frobenius operator sufficiently for eigenvalue comparison.
Cite this review
Pith. "Pith review of Deficiency of equation-finding approach to data-driven modeling of dynamical systems." pith.science (2026). https://pith.science/paper/O5WEYOHU
@misc{pith2026250903769,
author = {Pith},
title = {Pith review of: Deficiency of equation-finding approach to data-driven modeling of dynamical systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/O5WEYOHU}},
note = {Machine review of arXiv:2509.03769}
}
read the original abstract
Finding the governing equations from data by sparse optimization has become a popular approach to deterministic modeling of dynamical systems. Considering the physical situations where the data can be imperfect due to disturbances and measurement errors, we show that for many chaotic systems, widely used sparse-optimization methods for discovering governing equations produce models that depend sensitively on the measurement procedure, yet all such models generate virtually identical chaotic attractors, leading to a striking limitation that challenges the conventional notion of equation-based modeling in complex dynamical systems. Calculating the Koopman spectra, we find that the different sets of equations agree in their large eigenvalues and the differences begin to appear when the eigenvalues are smaller than an equation-dependent threshold. The results suggest that finding the governing equations of the system and attempting to interpret them physically may lead to misleading conclusions. It would be more useful to work directly with the available data using, e.g., machine-learning methods.
Figures
Reference graph
Works this paper leans on
-
[25]
Z.-M. Zhai, B. D. Stern, and Y.-C. Lai, Bridging known and unknown dynamics by transformer-based machine- learning inference from sparse observations, Nat. Com- mun. 16, 8053 (2025)
work page 2025
-
[1]
C. Grebogi, S. M. Hammel, J. A. Yorke, and T. Sauer, Shadowing of physical trajectories in chaotic dynamics: Containment and refinement, Phys. Rev. Lett. 65, 1527 (1990)
work page 1990
- [2]
-
[3]
F. Cecconi, M. Cencini, M. Falcioni, and A. Vulpiani, Predicting the future from the past: An old problem from a modern perspective, Ame. J. Phys. 80, 1001 (2012)
work page 2012
-
[4]
Ott, Chaos in Dynamical Systems, 2nd ed
E. Ott, Chaos in Dynamical Systems, 2nd ed. (Cambridge University Press, Cambridge, UK, 2002)
work page 2002
-
[5]
E. N. Lorenz, Deterministic nonperiodic flow, J. Atmos. Sci. 20, 130 (1963)
work page 1963
-
[6]
W.-X. Wang, R. Yang, Y.-C. Lai, V. Kovanis, and C. Grebogi, Predicting catastrophes in nonlinear dynami- cal systems by compressive sensing, Phys. Rev. Lett.106, 154101 (2011)
2011
-
[7]
Y.-C. Lai, Finding nonlinear system equations and com- plex network structures from data: A sparse optimization approach, Chaos 31, 082101 (2021)
work page 2021
Show all 42 references
-
[8]
S. L. Brunton, J. L. Proctor, and J. N. Kutz, Discovering governing equations from data by sparse identification of nonlinear dynamical systems, Proc. Nat. Acad. Sci. (USA) 113, 3932 (2016)
2016
-
[9]
Boffetta, M
G. Boffetta, M. Cencini, M. Falcioni, and A. Vulpiani, Predictability: a way to characterize complexity, Phys. Rep. 356, 367 (2002)
2002
-
[10]
Pardo, Statistical Inference Based on Divergence Mea- sures, Statistics: A Series of Textbooks and Monographs (CRC Press, 2018)
L. Pardo, Statistical Inference Based on Divergence Mea- sures, Statistics: A Series of Textbooks and Monographs (CRC Press, 2018)
2018
-
[11]
B. O. Koopman, Hamiltonian systems and transforma- tions in Hilbert space, Proc. Nat. Acad. Sci. (USA) 17, 315 (1931)
1931
-
[12]
Mezi´ c, Spectral properties of dynamical systems, model reduction and decompositions, Nonlinear Dynam
I. Mezi´ c, Spectral properties of dynamical systems, model reduction and decompositions, Nonlinear Dynam. 41, 309 (2005)
2005
-
[13]
Budiˇ si´ c, R
M. Budiˇ si´ c, R. Mohr, and I. Mezi´ c, Applied Koopman- ism, Chaos 22, 047510 (2012)
2012
-
[14]
S. L. Brunton and J. N. Kutz, Data-Driven Science and Engineering: Machine Learning, Dynamical Sys- tems, and Control (Cambridge University Press, Cam- bridge, UK, 2022)
2022
-
[15]
Zagli, M
N. Zagli, M. Colbrook, V. Lucarini, I. Mezi´ c, and J. Moroney, Bridging the gap between Koopmanism and response theory: Using natural variability to predict forced response, SIAM J. Applied Dyn. Sys. (accepted) , arXiv:2410.01622 (2025)
2025 arXiv
-
[16]
Lucarini, Interpretable and equation-free response theory for complex systems, Phil
V. Lucarini, Interpretable and equation-free response theory for complex systems, Phil. Trans. R. Soc. A , doi: 10.1098/rsta.2025.0081, arXiv:2502.07908 (2025)
2025
-
[17]
M. J. Colbrook, Chapter 4 - the multiverse of dynamic mode decomposition algorithms, in Numerical Analysis Meets Machine Learning, Handbook of Numerical Anal- ysis, Vol. 25, edited by S. Mishra and A. Townsend (El- sevier, 2024) pp. 127–230
2024
-
[18]
Kong, H.-W
L.-W. Kong, H.-W. Fan, C. Grebogi, and Y.-C. Lai, Ma- chine learning prediction of critical transition and system collapse, Phys. Rev. Res. 3, 013090 (2021)
2021
-
[19]
D. J. Gauthier, E. Bollt, A. Griffith, and W. A. Barbosa, Next generation reservoir computing, Nat. Commun. 12, 5564 (2021)
2021
-
[20]
Patel, D
D. Patel, D. Canaday, M. Girvan, A. Pomerance, and E. Ott, Using machine learning to predict statistical properties of non-stationary dynamical processes: Sys- tem climate, regime transitions, and the effect of stochas- ticity, Chaos 31, 033149 (2021)
2021
-
[21]
J. Z. Kim and D. S. Bassett, A neural machine code and programming framework for the reservoir computer, Nat. Mach. Intell. 5, 622 (2023)
2023
-
[22]
Flynn, V
A. Flynn, V. A. Tsachouridis, and A. Amann, Seeing double with a multifunctional reservoir computer, Chaos 33, 113115 (2023)
2023
-
[23]
T. J. Liu, N. Boull´ e, R. Sarfati, and C. J. Earls, LLMs learn governing principles of dynamical systems, reveal- ing an in-context neural scaling law, arXiv preprint arXiv:2402.00795 (2024)
2024 arXiv
-
[24]
Zhang and W
Y. Zhang and W. Gilpin, Zero-shot forecasting of chaotic systems, arXiv preprint arXiv:2409.15771 (2024)
2024 arXiv
-
[26]
Lasota and M
A. Lasota and M. C. Mackey, Chaos, Fractals, and Noise: Stochastic Aspects of Dynamics, Vol. 97 (Springer Science & Business Media, 2013)
2013
-
[27]
S. Klus, P. Koltai, and C. Sch¨ utte, On the numerical ap- proximation of the Perron-Frobenius and Koopman op- erator, J. Comp. Dyn. 3, 51 (2016)
2016
-
[28]
Ikeda, I
M. Ikeda, I. Ishikawa, and C. Schlosser, Koopman and Perron–Frobenius operators on reproducing kernel Ba- nach spaces, Chaos 32 (2022)
2022
-
[29]
Froyland, On Ulam approximation of the isolated spectrum and eigenfunctions of hyperbolic maps, Dis- crete Contin
G. Froyland, On Ulam approximation of the isolated spectrum and eigenfunctions of hyperbolic maps, Dis- crete Contin. Dyn. Syst. Ser. A 17, 203 (2007)
2007
-
[30]
C. Bose, G. Froyland, C. Gonz´ alez-Tokman, and R. Mur- ray, Ulam’s method for Lasota–Yorke maps with holes, SIAM J. Appl. Dyn. Syst. 13, 1010 (2014)
2014
-
[31]
M. J. Molina, T. A. O’Brien, G. Anderson, M. Ashfaq, K. E. Bennett, W. D. Collins, K. Dagon, J. M. Restrepo, and P. A. Ullrich, A review of recent and emerging ma- chine learning applications for climate variability and weather phenomena, Artifi. Intel. Earth Sys. 2, 220086 (2023)
2023
-
[32]
L.-W. Kong, Y. Weng, B. Glaz, M. Haile, and Y.-C. Lai, Reservoir computing as digital twins for nonlinear dy- namical systems, Chaos 33, 033111 (2023)
2023
-
[33]
McCann and P
K. McCann and P. Yodzis, Nonlinear dynamics and pop- ulation disappearances, Am. Nat. 144, 873 (1994)
1994
-
[34]
Z. Lin, Z. Lu, Z. Di, and Y. Tang, Learning noise-induced transitions by multi-scaling reservoir computing, Nat. Commun. 15, 6584 (2024)
2024
-
[35]
Z.-M. Zhai, M. Moradi, L.-W. Kong, and Y.-C. Lai, Detecting weak physical signal from noise: A machine- learning approach with applications to magnetic- anomaly-guided navigation, Phys. Rev. Appl. 19, 034030 (2023)
2023
-
[36]
Canaday, A
D. Canaday, A. Pomerance, and D. J. Gauthier, Model- free control of dynamical systems with deep reservoir computing, J. Phys. Complex. 2, 035025 (2021)
2021
-
[37]
Z.-M. Zhai, M. Moradi, L.-W. Kong, B. Glaz, M. Haile, and Y.-C. Lai, Model-free tracking control of complex 6 dynamical trajectories with machine learning, Nat. Com- mun. 14, 5698 (2023)
2023
-
[38]
Liu and M
Z. Liu and M. Tegmark, Machine learning conservation laws from trajectories, Phys. Rev. Lett. 126, 180604 (2021)
2021
-
[39]
Tenachi, R
W. Tenachi, R. Ibata, and F. I. Diakogiannis, Deep sym- bolic regression for physics guided by units constraints: Toward the automated discovery of physical laws, Astro- phys. J. 959, 99 (2023)
2023
-
[40]
Z. Liu, Y. Wang, S. Vaidya, F. Ruehle, J. Halver- son, M. Soljaˇ ci´ c, T. Y. Hou, and M. Tegmark, KAN: Kolmogorov-arnold networks, arXiv preprint arXiv:2404.19756 (2024)
2024 arXiv
-
[41]
Panahi, M
S. Panahi, M. Moradi, E. M. Bollt, and Y.-C. Lai, Data- driven model discovery with Kolmogorov-Arnold net- works, Phys. Rev. Res. 7, 023037 (2025)
2025
-
[42]
Udrescu and M
S.-M. Udrescu and M. Tegmark, AI Feynman: A physics- inspired method for symbolic regression, Sci. Adv. 6, eaay2631 (2020)
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.