REVIEW 2 major objections 5 minor 1 cited by
Non-Asymptotic Bounds for Closed-Loop Identification of Unstable Nonlinear Stochastic Systems
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper derives non-asymptotic, high-probability error bounds for least-squares identification of a class of unstable nonlinear closed-loop stochastic systems, under a regional excitation condition.
desk verdict Genuine extension of non-asymptotic identification to sub-exponentially unstable nonlinear closed-loop systems, with a checkable regional excitation condition; the title overreaches and condition (11) can be vacuous, but the core argument holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Regional excitation (Definition 2): for every direction ζ of the regressor space, the projected regressor ζ⊤ψ(x+W, α(x+W,S,ϑ)) has probability at least pPE of having magnitude at least cPE, uniformly over the region X and over all parameter guesses ϑ. Because it is defined only through the known basis functions ψ, the control policy α, and the noise distributions µs, µw, it can be verified without knowing the true parameter θ*. It supports a single-direction persistency-of-excitation lemma (Lemma 7), and an ε-covering argument over the unit sphere converts that into a high-probability, linearly growing lower bound on the minimum eigenvalue of the regularized Gramian G(t). The other load-bearing piece is the sub-exponential input-to-state bound (Assumption 3), which limits the growth of the closed-loop trajectory and ensures the burn-in time is finite.
What would settle it
Run many Monte-Carlo trials of the double integrator example with δ=0.1 and check the empirical frequency of the event that |θ̂(t)−θ*| ≤ e(t,δ,x0) for every t ≥ T_burn-in; a frequency below 0.9 would directly refute Theorem 1. Alternatively, construct a system satisfying the assumptions with polynomial growth of high degree and check numerically whether the bound e(t) truly behaves as O(√(ln t/t)).
Extended reading notes
Core claim
Theorem 1 is the central claim. Under Assumptions 1–6 (measurability, independent sub-Gaussian process noise, a sub-exponential input-to-state bound on the closed-loop trajectory, bounded controls, polynomially growing basis functions, and regional excitation over a set XPE), the paper defines two offline-computable times: T_burn-in, when persistency of excitation begins, and T_excited, a conservative lower bound on how long the one-step predicted state remains inside XPE. If T_burn-in ≤ T_excited, then with probability at least 1−δ the estimation error |θ̂(t)−θ*| is bounded by the explicit, time-dependent quantity e(t,δ,x0) for every t in the interval. Corollary 2 adds global excitation and polynomial instability, giving e(t,δ,x0) = O(√(ln t/t)) and convergence to zero for all times; the authors state this matches the rate for linear systems with spectral radius at most one, and that to their knowledge no comparable non-asymptotic guarantee existed for this class of unstable nonlinear closed-loop systems.
Load-bearing premise
The load-bearing premise is that the closed-loop trajectory grows at most sub-exponentially in time (Assumption 3); if the system is exponentially unstable or otherwise grows faster, the burn-in time can be infinite or the condition T_burn-in ≤ T_excited can fail, in which case Theorem 1 issues no guarantee.
Editorial extensions
If this is right
- Under global excitation, the RLS estimate enters and remains inside any arbitrarily small ball around the true parameter with probability at least 1−δ, for every initial state.
- The rate O(√(ln t/t)) matches the benchmark for linear systems with spectral radius at most one, so the identified nonlinear class is certified at the same speed as that known linear case.
- The bounds and the condition T_burn-in ≤ T_excited are verifiable offline from known objects, so a designer can decide before running an experiment whether the planned controller and noise will yield informative data.
- In the merely regional case the bound decreases over the PE interval and stops improving after the trajectory leaves the exciting region, matching the simulated behavior of the piecewise affine example.
- The regional excitation condition can be verified for the PWA example even though the block martingale small-ball condition (which requires future regressors to be non-degenerate on average given the past) fails, enlarging the set of systems with certified finite-sample identification.
Reading between the lines
- A testable extension is to weaken Assumption 3 to exponentially growing trajectories; the burn-in time would likely diverge, suggesting that segment-based or restart-based identification would be needed for genuinely unstable plants.
- Regional excitation can be read as a design constraint: shaping the policy α or the exploratory noise µs to maximize cPE and pPE would directly shrink the certified error e(t), though the paper does not address this optimization.
- The confidence region from Theorem 1 could feed a robust adaptive controller with finite-time guarantees, a use the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies regularised least-squares identification of the linearly parameterised discrete-time nonlinear stochastic system (1) in closed loop with a known, not necessarily stabilising policy. It introduces a regional excitation condition (Def. 2) that is checkable from the basis functions, policy, and noise distributions, and proves in Thm. 1 that, whenever the burn-in time in (10) does not exceed the excited time in (9), the RLS estimation error is bounded by e(t, δ, x0) in (13) uniformly on the PE interval with probability at least 1−δ. Under global excitation, Cor. 1 extends this to all times past burn-in and gives asymptotic decay; under polynomial reachability growth, Cor. 2 gives the O(sqrt(ln t / t)) rate. Two examples are analysed: a PWA system that fails the BMSB condition but satisfies regional excitation, and a double integrator controlled by an arbitrary bounded policy.
Significance. If the proofs are correct, this is a genuinely new finite-sample guarantee for closed-loop identification of a class of unstable nonlinear systems, avoiding mixing or boundedness assumptions. The regional excitation notion is a useful, verifiable alternative to BMSB, and the PWA example makes the advantage concrete. The proofs are detailed and combine standard self-normalized martingale inequalities, Chernoff bounds, and covering arguments; the examples verify the assumptions explicitly. The main caveats are that Assumption 3 restricts instability to sub-exponential growth and that condition (11) can be vacuous, so the actual scope is narrower than the word ‘unstable’ in the title suggests.
major comments (2)
- [Assumption 3, Theorem 1] Assumption 3 rules out exponentially unstable systems, including the canonical linear system X(t+1)=ρX(t)+W(t) with ρ>1: the minimal comparison function is χ1(t)=ρ^t, for which ln χ1(t)=Θ(t), not o(t), so χ1 is not K1-SE. Consequently Theorem 1 and Corollary 2 do not apply to the most basic unstable linear plants, and the title’s ‘unstable nonlinear systems’ overstates the scope. The abstract is careful (‘sub-exponentially unstable’), but the introduction and conclusions should state this restriction prominently and explain why it is inherent; otherwise readers will likely misapply the theorem.
- [Condition (11), Section 3.1] Theorem 1 is vacuous when Tburn-in(δ,x0) > Texcited(δ,x0), and no general sufficient condition for (11) is provided; the paper only verifies it case-by-case in two examples. Since Texcited is finite for any bounded XPE and Tburn-in grows with d and 1/pPE, there are natural systems for which the PE interval is empty and the claimed non-asymptotic bound does not exist. The discussion after (11) acknowledges this qualitatively, but a result intended as a non-asymptotic guarantee needs either a quantitative sufficient condition for (11) or an explicit statement that the theorem covers only systems for which the PE interval is nonempty.
minor comments (5)
- [Equation (13), after Lemma 3] The error bound e(t,δ,x0) contains γ^{1/2}|θ*|_F, so it is not fully data-independent in the sense claimed in the paragraph after Lemma 3; to compute a numerical confidence interval one must know a bound on |θ*|. The paper should replace |θ*| by a known bound B in the statement, or state |θ*| ≤ B as an assumption, and correct the word ‘data-independent’.
- [Section 4.1] The notation in Example 1 is confusing because x denotes both the state variable and the threshold parameter; for instance XPE = (−∞, 0.9x] and the constants bw, bs in Prop. 3 mix the two roles. Renaming the threshold, say x̄ or c, would make the example much easier to check.
- [Notation, Section 1] The definition of little-o is misstated: after defining f(r)=O(g(r)), the text says ‘h(r)=o(r) if lim f(r)/g(r)=0’, which should read h(r)=o(g(r)) with lim h(r)/g(r)=0.
- [Throughout] There are several small typos: ‘simualtions’ in Sec. 4.1.2, ‘inequlity’ in the proof of Thm. 1, and an extra comma in Cor. 2’s statement ‘e(t, , x0)’. These should be fixed in revision.
- [Equation (13)] The displayed equation for e(t,δ,x0) has a malformed line break with the equation number inserted mid-formula; the formatting should be corrected so that the bound is legible as a single expression.
Circularity Check
No significant circularity: the non-asymptotic bounds follow from explicitly stated assumptions, external martingale inequalities, and a regional excitation condition defined without the unknown parameter; no fitted value is renamed as a prediction.
full rationale
The derivation chain is self-contained relative to its assumptions. Regional excitation (Def. 2) is a property of the known basis functions, control policy, and noise distributions only; it does not involve the true parameter theta*. Lemma 1 converts moment lower/upper bounds into c_PE and p_PE, and these constants are computed from the system model in the examples, not fitted to the estimation error. Theorem 1 combines a data-dependent least-squares bound (Lemma 3, proved via the external self-normalized martingale inequality of Abbasi-Yadkori et al. [1]) with high-probability state/regressor bounds (Lemma 4, from Assumptions 2-5) and a regional persistency-of-excitation result (Lemma 5, from Assumption 6). The error bound e(t,delta,x0) in Eq. (13) is assembled from these ingredients rather than being defined to match the conclusion. The paper also states explicitly in Sec. 3.1.1 that condition (11) may fail and that it cannot be guaranteed without a particular system, so the conditional nature of Theorem 1 is an acknowledged scope limitation rather than a masked assumption. The cited works [1], [2], [15], and [27] are external and are used for standard concentration and PE arguments, not as a self-citation chain that forces the main result. The claim that Assumption 3 excludes exponentially unstable systems is a scope restriction, not circularity. No fitted input is called a prediction, and no known result is merely renamed.
Assumptions & free parameters
free parameters (1)
- regularization parameter γ =
e.g., 0.0001 in Example 1
assumptions (9)
- standard math Assumption 1: f, ψ, and α are Borel measurable.
- domain assumption Assumption 2: W(t) i.i.d., zero-mean, σ_w^2-sub-Gaussian, independent of S(t).
- domain assumption Assumption 3: Sub-exponential input-to-state bound on reachable states with comparison functions in K1-SE/K2-SE classes.
- domain assumption Assumption 4: Controls are magnitude-bounded by u_max.
- domain assumption Assumption 5: Basis functions grow at most polynomially in (state, control).
- domain assumption Assumption 6: Regional excitation over X_PE with constants c_PE, p_PE.
- domain assumption Assumption 7: Global excitation over the whole state space.
- domain assumption Assumption 8: Polynomial ISS bound (χ1, χ3, χ4, σ2 are APB).
- standard math Background: standard concentration and covering-number results (Lemmas 8-16) are used without proof.
Cite this review
Pith. "Pith review of Non-Asymptotic Bounds for Closed-Loop Identification of Unstable Nonlinear Stochastic Systems." pith.science (2026). https://pith.science/paper/43Y3HTRE
@misc{pith2026241204157,
author = {Pith},
title = {Pith review of: Non-Asymptotic Bounds for Closed-Loop Identification of Unstable Nonlinear Stochastic Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/43Y3HTRE}},
note = {Machine review of arXiv:2412.04157}
}
read the original abstract
We consider the problem of least squares parameter estimation from single-trajectory data for discrete-time, unstable, closed-loop nonlinear stochastic systems, with linearly parameterised uncertainty. Assuming a region of the state space produces informative data, and the system is sub-exponentially unstable, we establish non-asymptotic guarantees on the estimation error at times where the state trajectory evolves in this region. If the whole state space is informative, high probability guarantees on the error hold for all times. Examples are provided where our results are useful for analysis, but existing results are not.
Figures
Forward citations
Cited by 1 Pith paper
-
Non-asymptotic Bounds of Learning-based Linear MPC With Input Constraints and Unbounded Stochastic Noise
A certainty-equivalence MPC scheme with online least-squares identification and a deadbeat fallback attains a high-probability non-asymptotic practical stability bound for unknown input-constrained linear systems with...
Reference graph
Works this paper leans on
-
[1]
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári. Improved algorithms for linear stochastic bandits.Advances in neural information processing systems , 24, 2011
work page 2011
-
[2]
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári. Regret bounds for the adaptive control of linear quadratic systems. In Proceedings of the 24th Annual Conference on Learning Theory, pages 1–26. JMLR Workshop and Conference Proceedings, 2011
work page 2011
-
[3]
David Angeli and Eduardo D Sontag. Forward completeness, unboundedness observability, and their lyapunov characterizations.Systems & Control Letters, 38(4- 5):209–217, 1999
work page 1999
-
[4]
Finite sample properties of system identification methods
Marco C Campi and Erik Weyer. Finite sample properties of system identification methods. IEEE Transactions on Automatic Control, 47(8):1329–1334, 2002
work page 2002
-
[5]
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu. Regret bounds for robust adaptive control of the linear quadratic regulator. Advances in Neural Information Processing Systems, 31, 2018
work page 2018
-
[6]
Finite time identification in unstable linear systems
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis. Finite time identification in unstable linear systems. Automatica, 96:342–353, 2018
work page 2018
-
[7]
Learning nonlinear dynamical systems from a single trajectory
Dylan Foster, Tuhin Sarkar, and Alexander Rakhlin. Learning nonlinear dynamical systems from a single trajectory. In Learning for Dynamics and Control , pages 851–861. PMLR, 2020
work page 2020
-
[8]
Sergio Grammatico, Anantharaman Subbaraman, and Andrew R Teel. Discrete-time stochastic control systems: A continuous lyapunov function implies robustness to strictly causal perturbations. Automatica, 49(10):2939–2952, 2013
work page 2013
Show all 37 references
-
[9]
Self-convergence of weighted least-squares with applications to stochastic adaptive control
Lei Guo. Self-convergence of weighted least-squares with applications to stochastic adaptive control. IEEE Trans. Autom. Control, 41(1):79–89, 1996
1996
-
[10]
Foundations of modern probability, volume 2
Olav Kallenberg and Olav Kallenberg. Foundations of modern probability, volume 2. Springer, 1997
1997
-
[11]
Rates of uniform convergence of empirical means with mixing processes
Rajeeva L Karandikar and Mathukumalli Vidyasagar. Rates of uniform convergence of empirical means with mixing processes. Statistics & probability letters , 58(3):297–307, 2002
2002
-
[12]
Near-optimal offline and streaming algorithms for learning non-linear dynamical systems
SuhasKowshik,DheerajNagaraj,PrateekJain,andPraneeth Netrapalli. Near-optimal offline and streaming algorithms for learning non-linear dynamical systems. Advances in Neural Information Processing Systems, 34:8518–8531, 2021
2021
-
[13]
Asymptotic properties of general autoregressive models and strong consistency of least-squares estimates of their parameters
TL Lai and CZ Wei. Asymptotic properties of general autoregressive models and strong consistency of least-squares estimates of their parameters. Journal of multivariate analysis, 13(1):1–23, 1983. 16
1983
-
[14]
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems.The Annals of Statistics, 10(1):154–166, 1982
Tze Leung Lai and Ching Zong Wei. Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems.The Annals of Statistics, 10(1):154–166, 1982
1982
-
[15]
Reinforcement learning with fast stabilization in linear dynamical systems
Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, and Animashree Anandkumar. Reinforcement learning with fast stabilization in linear dynamical systems. InInt. Conf. Artif. Intell. Statist., pages 5354–5390. PMLR, 2022
2022
-
[16]
Bandit algorithms
Tor Lattimore and Csaba Szepesvári. Bandit algorithms . Cambridge University Press, 2020
2020
-
[17]
Non-asymptotic system identification for linear systems with nonlinear policies
Yingying Li, Tianpeng Zhang, Subhro Das, Jeff Shamma, and Na Li. Non-asymptotic system identification for linear systems with nonlinear policies. arXiv preprint arXiv:2306.10369, 2023
2023 arXiv
-
[18]
On the convergence of least squares estimator for nonlinear autoregressive models
Zhaobo Liu and Chanying Li. On the convergence of least squares estimator for nonlinear autoregressive models. In 2021 40th Chinese Control Conference (CCC) , pages 1389–
2021
-
[19]
Active learning for nonlinear system identification with guarantees
Horia Mania, Michael I Jordan, and Benjamin Recht. Active learning for nonlinear system identification with guarantees. arXiv preprint arXiv:2006.10277 , 2020
2006 arXiv
-
[20]
A tutorial on concentration bounds for system identification
Nikolai Matni and Stephen Tu. A tutorial on concentration bounds for system identification. In 2019 IEEE 58th Conference on Decision and Control (CDC) , pages 3741–
2019
-
[21]
Revisiting ho–kalman- based system identification: Robustness and finite-sample analysis
Samet Oymak and Necmiye Ozay. Revisiting ho–kalman- based system identification: Robustness and finite-sample analysis. IEEE Transactions on Automatic Control , 67(4):1914–1928, 2021
1914
-
[22]
Near optimal finite time identification of arbitrary linear dynamical systems
Tuhin Sarkar and Alexander Rakhlin. Near optimal finite time identification of arbitrary linear dynamical systems. In International Conference on Machine Learning , pages 5610–
-
[23]
Finite time lti system identification.The Journal of Machine Learning Research, 22(1):1186–1246, 2021
Tuhin Sarkar, Alexander Rakhlin, and Munther A Dahleh. Finite time lti system identification.The Journal of Machine Learning Research, 22(1):1186–1246, 2021
2021
-
[24]
Non-asymptotic and accurate learning of nonlinear dynamical systems
Yahya Sattar and Samet Oymak. Non-asymptotic and accurate learning of nonlinear dynamical systems. The Journal of Machine Learning Research , 23(1):6248–6296, 2022
2022
-
[25]
Finite sample identification of bilinear dynamical systems
Yahya Sattar, Samet Oymak, and Necmiye Ozay. Finite sample identification of bilinear dynamical systems. In2022 IEEE 61st Conference on Decision and Control (CDC),pages 6705–6711. IEEE, 2022
2022
-
[26]
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan Foster. Naive exploration is optimal for online lqr. In Int. Conf. Mach. Learn. , pages 8937–8948. PMLR, 2020
2020
-
[27]
Learning without mixing: Towards a sharp analysis of linear system identification
MaxSimchowitz,HoriaMania,StephenTu,MichaelIJordan, and Benjamin Recht. Learning without mixing: Towards a sharp analysis of linear system identification. InConf. Learn. Theory, pages 439–473. PMLR, 2018
2018
-
[28]
On the continuity of the generalized inverse
GW Stewart. On the continuity of the generalized inverse. SIAM Journal on Applied Mathematics , 17(1):33–45, 1969
1969
-
[29]
Finite sample analysis of stochastic system identification
Anastasios Tsiamis and George J Pappas. Finite sample analysis of stochastic system identification. In 2019 IEEE 58th Conference on Decision and Control (CDC) , pages 3648–3654. IEEE, 2019
2019
-
[30]
Statistical learning theory for control: A finite sample perspective
Anastasios Tsiamis, Ingvar Ziemann, Nikolai Matni, and George J Pappas. Statistical learning theory for control: A finite sample perspective. arXiv preprint arXiv:2209.05423 , 2022
2022 arXiv
-
[31]
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin. High-dimensional probability: An introduction with applications in data science , volume 47. Cambridge university press, 2018
2018
-
[32]
A learning theory approach to system identification and stochastic adaptive control
Mathukumalli Vidyasagar and Rajeeva L Karandikar. A learning theory approach to system identification and stochastic adaptive control. Probabilistic and randomized methods for design under uncertainty , pages 265–302, 2006
2006
-
[33]
Finite sample properties of linear model identification.IEEE Transactions on Automatic Control , 44(7):1370–1383, 1999
Erik Weyer, Robert C Williamson, and Iven MY Mareels. Finite sample properties of linear model identification.IEEE Transactions on Automatic Control , 44(7):1370–1383, 1999
1999
-
[34]
Real analysis: theory of measure and integration
James J Yeh. Real analysis: theory of measure and integration. World Scientific Publishing Company, 2014
2014
-
[35]
Learning with little mixing
Ingvar Ziemann and Stephen Tu. Learning with little mixing. Advances in Neural Information Processing Systems, 35:4626–4637, 2022
2022
-
[36]
Single trajectory nonparametric learning of nonlinear dynamics
Ingvar M Ziemann, Henrik Sandberg, and Nikolai Matni. Single trajectory nonparametric learning of nonlinear dynamics. In conference on Learning Theory , pages 3333–
-
[37]
Trigonometric series , volume 1
Antoni Zygmund. Trigonometric series , volume 1. Cambridge university press, 2002. A Technical Lemmas Lemma 9(Boundonsub-Gaussiansequenceuniformly over time) Consider an Rd-valued random sequence {Y (t)}t∈N such that Y (t) is σ2 y-sub-Gaussian for all t ∈ N. Then, for anyδ ∈ (...
2002
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.