REVIEW 3 major objections 5 minor 42 references
This paper establishes that the parameters of a noisy dynamical model can be uniquely recovered from 2p+1 randomly chosen features of the observed time series, where p is the number of unknown parameters.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 21:32 UTC pith:3VJOWTKF
load-bearing objection The central genericity step doesn't go through—prevalence over C^1 maps doesn't transfer to the finite-dimensional random cosine family—so the title claim isn't proven, though the idea is worth a serious revision. the 3 major comments →
Dynamic models with p parameters are identified by 2p+1 random features
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that identification of a p-dimensional parameter θ in a noisy dynamical model can be achieved through the expectation values of k=2p+1 random features φ_i(x)=cos(Ω_i·x+α_i), with Gaussian frequencies and uniform phases. Defining Φ(θ)=E_θ[φ(X)], the authors argue that for almost every draw of the random features, Φ is one-to-one on the compact parameter space Θ and is a C^1 diffeomorphism between Θ and its image, provided the map from parameters to stationary or local distributions is smooth and injective. Therefore Φ(θ1)=Φ(θ2) only when θ1=θ2, and the population discrepancy Q(θ)=∥Φ(θ0)-Φ(θ)∥ is nonzero away from the truth. This converts parameter identification i
What carries the argument
The load-bearing object is the random Fourier feature map φ_i(x)=cos(Σ_j Ω_{ij} x_j + α_i) and its expectation Φ(θ)=E_θ[φ(X)]. The dimension count k≥2p+1 comes from the fractal embedding prevalence theorem, which says that almost every smooth map from R^p to R^k is one-to-one on a compact p-dimensional set and immersed on its smooth pieces. The paper applies this theorem to the expected-feature map, transferring genericity from the space of all smooth maps to the specific random-feature family; this transfer is what makes the identification claim an almost-sure property of the random draws rather than a statement about specially designed measurements.
Load-bearing premise
The decisive premise is that embeddings are not just prevalent among all smooth maps but are almost sure under the specific distribution of random Fourier features; the paper asserts this transfer but its proof only covers the space of all C^1 maps, so the identification claim depends on that unproven inheritance.
What would settle it
Take a parametric family satisfying the paper's assumptions and, for a fixed k=2p+1, numerically estimate the probability (over draws of Ω and α) that the feature-expectation map Φ is non-injective on a fine grid covering Θ; if for some family this probability does not approach zero, the almost-sure embedding claim is false. Equivalently, exhibit two parameters θ1≠θ2 such that E_{θ1}cos(Ω·X+α)=E_{θ2}cos(Ω·X+α) for a set of (Ω,α) of positive measure.
If this is right
- For any parametric time-series model satisfying the smoothness and injectivity assumptions, parameter estimation reduces to minimizing a discrepancy between observed and simulated random-feature averages, with no user-chosen summary statistics or neural-network training.
- The 2p+1 count gives a concrete prescription for the number of features to simulate, mirroring the classical 2d+1 delays in state-space reconstruction and 2r+1 experiments for noiseless systems.
- Because the noise can be non-iid, non-Gaussian, and state-dependent, the result covers realistic observational and process noise that breaks standard likelihood-based identifiability arguments.
- The rolling-window version extends identification to nonstationary trajectories, including noisily observed differential equations on a fixed horizon, where time averages do not converge.
- The same feature-matching principle could be applied to other data structures, such as spatiotemporal fields or networks, wherever a class of random features for the dependence structure is available.
Where Pith is reading between the lines
- The proof establishes prevalence of embeddings among all C^1 maps but does not explicitly prove that the random cosine family inherits this prevalence; verifying that the pushforward measure of the random (Ω,α) puts full mass on embeddings would close the gap between the conditional result ('if the features embed') and the almost-sure statement.
- If the almost-sure embedding holds, the same machinery suggests a general recipe for automatic summary statistics in simulation-based inference: draw random features once, fix them, and match expectations; this could be tested on models outside dynamics, such as spatial point processes or random graphs.
- The identification result is non-asymptotic in the sense that it concerns population expectations; the finite-sample question is whether the empirical discrepancies concentrate quickly enough, which the companion work addresses with asymptotics. A natural testable extension is to measure how the required sample size scales with p and with the feature count.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that for a p-parameter dynamic model observed with non-iid noise, the expectation values of 2p+1 random Fourier features of length-(m+1) blocks identify the parameter vector. Theorems 2 and 3 are stated for stationary and nonstationary cases, respectively, and are supposed to follow from the Sauer–Yorke–Casdagli fractal Whitney embedding theorem: since almost every C^1 map from R^p to R^k is an embedding for k≥2p+1, and the random-feature expectation map is C^1, the random features should almost surely embed the parameter space. Applications to a noisily observed Hénon map and Lorenz-63 system are presented as illustrations. The formal theorems only prove prevalence among all C^1 maps; the transfer to the specific finite-dimensional random cosine family is not proved.
Significance. The intended contribution is valuable if correct: it would replace user-chosen summary statistics or trained network summaries by simple random Fourier features and give a clean identification principle for nonlinear dynamic models with dependent, non-Gaussian noise. The paper also has genuine strengths: it is explicit about the genericity axioms, the smoothness computation in Step 2 of the proof of Theorem 2 is coherent under the stated Fréchet differentiability assumption, and the focus on non-iid noise is a useful extension of noiseless embedding results. However, the core claim is not established. The paper's central step—that a random draw from the finite-dimensional family Eq. (1) produces an embedding almost surely—is asserted rather than proved, and the formal theorems do not even state this claim. The simulations are two examples and cannot substitute for the missing proof. As it stands, the rigorously established part of the paper is a conditional finite-feature reduction under full-distribution identifiability, not the headline identification theorem.
major comments (3)
- [§II.D, Theorem 2; Appendix A Step 3] The step from prevalence in C^1 to the random Fourier feature family is unproved and is load-bearing. Step 3 of the proof of Theorem 2 only establishes that embeddings are prevalent in C^1(Θ,R^k), a restatement of Sauer et al. The feature maps generated by Eq. (1) form a finite-dimensional family, parameterized by (Ω_i, α_i), i=1,...,k, of dimension k(d(m+1)+1). A finite-dimensional subset of the infinite-dimensional space C^1(R^p,R^k) is shy, not prevalent, so the prevalence statement about C^1 gives no probability statement for random draws from this family. The sentence in §II.D, 'Indeed, by Theorem 2, they will, almost surely,' is therefore not a consequence of Theorem 2. One needs a direct proof that, under Assumption 2, the distribution over (Ω_i, α_i) makes θ ↦ (E_θ[cos(Ω_1·X+α_1)], ..., E_θ[cos(Ω_k·X+α_k)]) injective on Θ with probability one. No such argument appears; Theorem 3
- [§II.D, Theorem 2 statement] The formal statement of Theorem 2 does not contain the headline claim. It concludes only that C⊂C^1(R^p,R^k) and that almost every C^1 map is one-to-one on Θ, with no assertion about the specific random feature map of Eq. (2). The identification claim is then added in the prose ('Suppose the randomly chosen functions ... Indeed, by Theorem 2, they will, almost surely'). Thus the stated theorem cannot support the abstract's claim. This is not a minor presentational issue: a reader cannot verify the main theorem from the formal result. The same structural problem applies to Theorem 3.
- [Assumption 2.2, §II.D] Assumption 2.2 assumes θ ↦ P_θ is a bijection, which is essentially full distributional identifiability from the (m+1)-block. This is a strong primitive: it already grants the kind of identifiability the paper aims to establish, albeit at the level of the full distribution rather than 2p+1 features. The finite-feature reduction is a meaningful goal, but the paper should present it as such and should not imply identification is obtained from scratch. As written, the rigorously proved part of Theorem 2 is conditional both on full-distribution identifiability and on the unproved random-feature embedding, so the title overclaims.
minor comments (5)
- [§II.D, Assumption 2.3] The map p takes values in the space of probability measures, which is not a vector space. Fréchet differentiability should be defined with respect to the ambient vector space of signed measures with total-variation norm. Please state this embedding explicitly.
- [§II.E, Step 3 of Theorem 3 proof] The proof uses the fact that Lipschitz maps do not increase box-counting dimension. Please specify which box-counting dimension (upper/lower) is used and give the inequality explicitly, since the manuscript does not define the notion.
- [§III] The simulation section lacks implementation details: the optimization method, the number of simulations s, the lag m, the number of Monte Carlo replications behind the densities, and the handling of the rolling-window choices are not described. This limits reproducibility and makes the empirical illustrations hard to assess.
- [Abstract and §IV] The abstract promises practical estimation procedures, but the consistency of the time-average and rolling-window estimators is entirely deferred to the companion paper [21]. If the present paper is meant to stand alone, the relevant consistency statements should be restated or precisely referenced.
- [Figure 1 caption] The caption refers to color for distinguishing parameter ranges; please ensure the figure is also readable in black-and-white or add markers.
Circularity Check
The central bridge from prevalence in C^1 to the random-feature family is asserted, not proved, and key technical content is imported from the authors' own companion paper.
specific steps
-
other
[Section II.D, paragraph after Theorem 2; Appendix A, Proof of Theorem 2, Step 3]
"Suppose the randomly chosen functions from Eq. (2) yield an embedding. Indeed, by Theorem 2, they will, almost surely. Then we can distinguish parameters θ1≠θ2 by looking at the expectation values of the random features Φ(θ1) and Φ(θ2)."
The paper's identification claim requires that the specific finite-dimensional family of random Fourier features (Eq. 1) almost surely gives an injective map θ↦Eθ[φ(X)]. The proof's Step 3 only establishes that embeddings are prevalent in C^1(Θ,R^k), i.e. that almost every C^1 map is one-to-one on Θ. Since the random-feature family generated by Eq. (1) is a restricted finite-dimensional subset of C^1, prevalence in C^1 does not imply prevalence on draws from Eq. (1). The sentence 'Indeed, by Theorem 2, they will, almost surely' asserts the missing conclusion; the subsequent identification argument therefore reduces to the very embedding property that needed proof, making the central result conditional on an unverified premise.
-
self citation load bearing
[Section II.D; Section II.E; Appendix A.2; Appendix A.3]
"The proof is a streamlined version of the proof of Theorem 2 in the companion paper [21]."
The consistency of the estimators and the technical DGP setup for nonstationary processes are not proved in this manuscript; they are imported from the authors' own unpublished companion paper [21]. The proofs of the two main theorems are explicitly said to be 'streamlined versions' of theorems in that same companion paper. Thus parts of the derivation chain are completed only by a self-citation whose content is not independently established in the present paper. This is load-bearing for the claim that the estimators are consistent, although the identification statement itself is not entirely reduced to [21].
full rationale
The paper is not globally circular: Theorem 2 honestly derives C^1-smoothness of Φ from Assumption 2, and the identification statement is explicitly conditional on the feature map being an embedding. However, the crucial sentence claiming that the random features 'will, almost surely' be embeddings is not supported by the proof, which only establishes prevalence in the larger space C^1(Θ,R^k); this is a non-sequitur rather than a derivation. In addition, the consistency results and the full proofs for the nonstationary case are deferred to the authors' own companion paper. These features make the central claim depend on an unverified transfer and on self-citation, but they do not make the result a tautology or a renamed fit. Score 4 reflects partial reliance on self-citation plus an asserted, rather than derived, load-bearing premise.
Axiom & Free-Parameter Ledger
free parameters (3)
- Lag order m
- Rolling-window size w_n =
minimizer within [n^{1/2}, n^{3/4}] of sum of squared distances
- Number of simulations s =
s=10 in figures
axioms (6)
- standard math Fractal Whitney Embedding Prevalence Theorem (Sauer et al., Theorem 1)
- domain assumption Assumption 1: uniform weak convergence of time-averaged block distributions
- domain assumption Assumption 2.2: θ↦P_θ is a bijection
- domain assumption Assumption 2.3: Fréchet C1 extension in total variation norm
- domain assumption Assumptions 3–5: local reparametrization, local bijections, and positive-measure distinguishability
- ad hoc to paper Random Fourier feature family is prevalent/generic
invented entities (1)
-
None
no independent evidence
read the original abstract
A foundational principle in nonlinear dynamics is that the structure of a dynamical system can be recovered from a small number of generic measurements or coordinates. We develop an analogous principle for the identification of dynamic models for time series {\em with noise}, which builds on previous identification results for noiseless dynamical systems. The noise is allowed to be non-iid, non-Gaussian, and dependent on the state. Our results cover noisily observed differential equations and discrete-time dynamical systems, as well as stochastic models with process noise. We illustrate the utility of this identification principle using a Lorenz-63 model and a H\'{e}non map model, both with observational noise.
Figures
Reference graph
Works this paper leans on
-
[1]
A generic subset ofVis dense inV
-
[2]
IfF⊃EandEis generic, thenFis generic
-
[3]
A countable intersection of generic sets is generic
-
[4]
Every translate of a generic set is generic
-
[5]
probability one
A subsetEof finite-dimensional Euclidean space is generic if and only ifEis a set of full Lebesgue measure. 4 (a) (b) 0.16 0.08 0.00 0.08 0.16 1 0.15 0.00 0.15 2 0.2 0.1 0.0 0.1 0.2 3 2.00 1.75 1.50 1.25 1.00 0.75 0.50 0.25 0.00 0.4 0.2 0.0 0.2 0.4 1 0.4 0.2 0.0 0.2 0.4 2 0.50 0.25 0.00 0.25 0.50 3 100 102 104 106 108 110 112 114 FIG. 1. The samek= 2p+ 1 ...
-
[6]
statistic
In statistical theory, any (measurable) function of the data is called a “statistic”. We find this term leads to confusion among non-statisticians, so we will speak of “features”, as is common in data mining and machine learning
-
[7]
Similarly, Bara´ nskiet al
show that almost every time-delay map is an embedding. Similarly, Bara´ nskiet al
-
[8]
Kantz and T
H. Kantz and T. Schreiber,Nonlinear Time Series Analysis(Cambridge University Press, Cambridge, England, 2003)
2003
-
[9]
Starket al.[10, 11] extend the Takens [2] embedding theorem to forced systems and stochastic systems
show that the time-delay map formed by a polynomial perturbation of a given locally Lipschitz observation function is almost surely an embedding. Starket al.[10, 11] extend the Takens [2] embedding theorem to forced systems and stochastic systems. Casdagliet al
-
[10]
0< µ(C)<∞for some compact subsetCofV
-
[11]
almost every
The setE+vhas fullµ-measure (complement ofE+vhas measure zero) for allv∈V. IfE⊂Vcontains a prevalent Borel set, then it is common to say that “almost every” element ofVlies inE. Next, we state a classical embedding result for smooth maps from Saueret al.[7], which involves the notion of animmersion. Definition 2(Immersion [20]).A smooth map Φ :R p − →Rk i...
-
[12]
Recently, Botvinick-Greenhouseet al.[13] established a measure- theoretic generalization using optimal transport theory
use statistical methods to account for the effect of iid observational noise on the state space reconstruction. Recently, Botvinick-Greenhouseet al.[13] established a measure- theoretic generalization using optimal transport theory. These results all trace back to the classical Whitney [14] embedding theorem from differential topology. Related results als...
-
[13]
The mapθ7→P θ is a bijection. 8
-
[14]
In other words, there exists an open setO⊂R p with Θ⊂Oand a mapp:O− → P R(m+1)×d , such thatp(θ) =P θ for allθ∈Θ
For some open setO⊃Θ, the mapθ7→P θ admits an extensionp:O− → P R(m+1)×d that is Fr´ echetC1 onOwith respect to the total variation norm. In other words, there exists an open setO⊂R p with Θ⊂Oand a mapp:O− → P R(m+1)×d , such thatp(θ) =P θ for allθ∈Θ. Furthermore,pis Fr´ echetC 1 onO with respect to the total variation norm. We linkpfrom Assumption 2 to Φ...
-
[15]
˜Θu is a compact subset ofR ˜p, andΘis a compact subset ofR p with non-empty interior
-
[16]
The map ˜θu 7→ ˜P˜θu,u is a bijection
-
[17]
In other words, for allu∈[0,1], there exists an open set ˜Ou ⊂R ˜pwith ˜Θu ⊂ ˜Ou and a map ˜pu : ˜Ou − → P R(m+1)×d such that ˜pu(˜θu) = ˜P˜θu,u
For some open set ˜Ou ⊃ ˜Θu, the map ˜θu 7→ ˜P˜θu,u admits an extension ˜pu : ˜Ou − → P R(m+1)×d that is Fr´ echetC1 on ˜Ou with respect to the total variation norm. In other words, for allu∈[0,1], there exists an open set ˜Ou ⊂R ˜pwith ˜Θu ⊂ ˜Ou and a map ˜pu : ˜Ou − → P R(m+1)×d such that ˜pu(˜θu) = ˜P˜θu,u. Also, ˜pu is Fr´ echetC1 on ˜Ou with respect ...
-
[18]
Proof of Theorem 1 See the proof of Theorem 2.3 in [7]. □
-
[19]
almost every
Proof of Theorem 2 The proof is a streamlined version of the proof of Theorem 2 in the companion paper [21]. Step 1: Limiting expectations.We claim that, for allθ∈Θ, the limiting expectation values of thekrandom features are given by Φ(θ) = Z R(m+1)×d φ(x)dP θ(x). Observe that, for allθ∈Θ and allr∈ {0,1, . . . , s}, asn− → ∞, we have 1 n−m nX t=m+1 Eθ φ h...
-
[20]
almost every
Proof of Theorem 3 The proof is a streamlined version of the proof of Theorem 5 in the companion paper [21]. Step 1: Local expectations.We claim that∀ ˜θu ∈ ˜Θu the local expectation values of the random features are given by ˜Φu(˜θu) = Z R(m+1)×d φ(x)d ˜P˜θu,u(x), u∈[0,1]. Observe that, for allθ∈Θ, allu∈[0,1], and allr∈ {0,1, . . . , s}, we have ˜Φu(˜θu)...
-
[21]
N. H. Packard, J. P. Crutchfield, J. D. Farmer, and R. S. Shaw, Geometry from a time series, Physical Review Letters45, 712 (1980)
1980
-
[22]
Takens, Detecting strange attractors in turbulence, Symposium on Dynamical Systems and Turbulence, Warwick 1980 , 366 (1981)
F. Takens, Detecting strange attractors in turbulence, Symposium on Dynamical Systems and Turbulence, Warwick 1980 , 366 (1981)
1980
-
[23]
estimation
Terminology differs confusingly across related fields. Among statisticians, inferring values of parameters from data is called “estimation”, while, e.g., climatologists and some economists are more apt to speak of “calibration”. We will generally follow conventions of statistics, but will try to define terms clearly
-
[24]
system identification
By contrast, what control engineering and signal processing calls “system identification” is closer to statistics’ “estimation”
-
[25]
There are rare exceptions with genuine redundancies or symmetries in the model parameteri- zation, but these are usually accidental, and/or can be eliminated by, e.g., picking a canonical value from each congruence class
-
[26]
Sauer, J
T. Sauer, J. A. Yorke, and M. Casdagli, Embedology, Journal of Statistical Physics65, 579 (1991)
1991
-
[27]
Bara´ nski, Y
K. Bara´ nski, Y. Gutman, and A. ´Spiewak, A probabilistic takens theorem, Nonlinearity33, 4940 (2020)
2020
-
[28]
Stark, D
J. Stark, D. S. Broomhead, M. E. Davies, and J. Huke, Takens embedding theorems for forced 18 and stochastic systems, Nonlinear Analysis30, 5303–5314 (1997)
1997
-
[29]
Stark, D
J. Stark, D. S. Broomhead, M. E. Davies, and J. Huke, Delay embeddings for forced systems. ii. stochastic forcing, Journal of Nonlinear Science13, 519 (2003)
2003
-
[30]
Casdagli, S
M. Casdagli, S. Eubank, J. D. Farmer, and J. Gibson, State space reconstruction in the presence of noise, Physica D: Nonlinear Phenomena51, 52 (1991)
1991
-
[31]
Botvinick-Greenhouse, M
J. Botvinick-Greenhouse, M. Oprea, R. Maulik, and Y. Yang, Measure-theoretic time-delay embedding, Journal of Statistical Physics192(2025)
2025
-
[32]
Whitney, Differentiable manifolds, Annals of Mathematics37, 645 (1936)
H. Whitney, Differentiable manifolds, Annals of Mathematics37, 645 (1936)
1936
-
[33]
Aeyels, Generic observability of differentiable systems, SIAM Journal on Control and Op- timization19, 595 (1981)
D. Aeyels, Generic observability of differentiable systems, SIAM Journal on Control and Op- timization19, 595 (1981)
1981
-
[34]
Aeyels, On the number of samples necessary to achieve observability, Systems and Control Letters1, 92 (1981)
D. Aeyels, On the number of samples necessary to achieve observability, Systems and Control Letters1, 92 (1981)
1981
-
[35]
E. D. Sontag, For differential equations withrparameters, 2r+ 1 experiments are enough for identification, Journal of Nonlinear Science12, 553 (2003)
2003
-
[36]
Yorke and W
J. Yorke and W. Ott, Prevalence, Bulletin of the American Mathematical Society42, 263 (2005)
2005
-
[37]
B. R. Hunt, T. Sauer, and J. A. Yorke, Prevalence: a translation-invariant almost every on infinite-dimensional spaces, Bulletin of the American Mathematical Society27, 217 (1992)
1992
-
[38]
Gallot, D
S. Gallot, D. Hulin, and J. Lafontaine,Riemannian Geometry, Vol. 2 (Springer-Verlag, Berlin, 1990)
1990
-
[39]
Wieck-Sosa and C
M. Wieck-Sosa and C. R. Shalizi, Estimating dynamic models by matching random features (2026), manuscript submitted for publication
2026
-
[40]
H´ enon, A two-dimensional mapping with a strange attractor, Communications in Mathe- matical Physics50, 69 (1976)
M. H´ enon, A two-dimensional mapping with a strange attractor, Communications in Mathe- matical Physics50, 69 (1976)
1976
-
[41]
E. N. Lorenz, Deterministic nonperiodic flow, Journal of the Atmospheric Sciences20, 130–141 (1963)
1963
-
[42]
Whitney, Analytic extensions of differentiable functions defined in closed sets, Transactions of the American Mathematical Society36, 63 (1934)
H. Whitney, Analytic extensions of differentiable functions defined in closed sets, Transactions of the American Mathematical Society36, 63 (1934). 19
1934
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.