{"id":"11f59dbf-e27c-4b45-a8b1-fc12253a7fe5","arxiv_id":"2607.13513","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A certainty-equivalence MPC scheme with online least-squares identification and a deadbeat fallback attains a high-probability non-asymptotic practical stability bound for unknown input-constrained linear systems with unbounded sub-Gaussian noise.","lead":"An adaptive model-predictive controller that learns an unknown linear system online while enforcing input limits is proven to keep states bounded with high probability, despite unbounded noise. The proof combines recursive least squares with a saturated deadbeat fallback controller and a hysteresis switching rule.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniform PE lemma (Lemma 1) is only referenced, not proven; Proposition 1 and Theorem 1 inherit an unverified anti-concentration bound.","rationale":"The reader's weakest-assumption analysis correctly identifies Lemma 1 as the load-bearing unverified step. My independent reading confirms that Proposition 1, Proposition 3, and Theorem 1 all rely on this lemma, and Remark 6 explicitly defers the proof to a companion paper. No other issue is as fundamental: the horizon condition's dependence on unknown matrices is a non-constructivity concern but Corollary 1 shows it can be made conservative using σ_B; the bounded/noise inconsistency in the conclusion is a scope wording issue; the per-step probabilistic argument in Lemma 10 is delicate but appears to use uniform-in-time events from Proposition 1 and [35], so it is less clearly broken. The missing PE proof is the single point where the entire global stability claim could fail. My recommendation is UNCHANGED because the reader's CONDITIONAL verdict already appropriately captures that the paper should be revised to include or verify Lemma 1. I am not arguing for rejection: no error has been demonstrated, and the lemma is plausible. I am arguing that the paper is not self-contained at its core and that the conditional verdict should explicitly require the PE proof before the central claim is accepted.","tokens_in":29057,"tokens_out":12356,"duration_ms":130246,"concrete_test":"For the scalar case n=m=κ=1 with A=B=1, u_max=6, C=3, w~N(0,1), v~Unif[-3,3], numerically search a fine grid over x∈[-100,100], θ∈[-2,2]^2 (or θ∈R^2), unit vectors ζ∈S^1, and s∈{MPC,DB} to compute the minimum over (x,θ,ζ,s) of P(|ζ^T[(x+w);(π_s(x+w,θ)+v)]|^2 ≥ c) for c=0.01,0.1,1. If this minimum is zero or decays to zero as |x| grows, Lemma 1's uniform constants do not exist in this simple setting, invalidating Proposition 1 for a basic case. If the minimum stays bounded below, try to write a standalone proof for this scalar case; the proof will reveal whether uniformity across x and θ is actually obtainable or requires extra assumptions (e.g., bounded x or θ near θ*).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1's high-probability stability bound depends on Proposition 1's RLS error bound, which is stated to follow from Lemma 1 plus results in [35]. Lemma 1 (Section IV-A) claims uniform constants c_PE,p_PE for every x, every unit vector ζ, every mode s∈{MPC,DB}, and every θ, with no compactness or closeness conditions on the state or parameter. Remark 6 explicitly says 'The proof can be found in [46]' and no proof is given here. This is not a cosmetic omission: the MPC policy is not globally Lipschitz in the state, and for large |x| it saturates; the uniformity over all x, θ, and both modes is precisely what guarantees persistent excitation along unstable closed-loop trajectories before burn-in. If Lemma 1 fails, the RLS error bound ϵ_t in (25) may not be valid, T_burn-in may be infinite, and Proposition 3, Lemma 10, and Theorem 1 collapse. The paper also imports Lemmas 4, 6, and the final ISS step from same-group preprints, but those are at least stated; Lemma 1 is the core new identification ingredient and its proof is essential for the claimed global, non-asymptotic guarantee. A secondary but related issue is that the horizon condition (12) is expressed in terms of the unknown true system matrices; Corollary 1 partially addresses this via conservative bounds, but the theorem as stated remains non-constructive. These are addressable gaps rather than demonstrated contradictions, so the appropriate disposition is conditional acceptance pending the missing PE proof.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a certainty-equivalence learning-based MPC scheme for unknown discrete-time LTI systems with hard input constraints and unbounded sub-Gaussian disturbances. The controller switches between a saturated deadbeat law and an MPC law, with hysteresis, and injects bounded excitation; system parameters are updated online by regularized least squares. The main result, Theorem 1, asserts a non-asymptotic high-probability practical stability bound after a finite transient, under a horizon condition (12). The proof proceeds through a uniform persistent-excitation lemma, a non-asymptotic RLS error bound, Lyapunov-like bounds for the deadbeat and MPC value functions, a convex-decay stochastic ISS argument, and a conversion from κ-step to per-step bounds. A simulation illustrates the behavior.","tokens_in":29488,"tokens_out":12131,"duration_ms":126754,"significance":"If the result is correct, it is a notable advance: simultaneous online identification and global stabilization of an unknown linear system under hard input constraints and unbounded noise, without pre-stabilizing controllers, terminal ingredients, or a known compact uncertainty set. The switching construction and the attempt to propagate finite-time identification errors through the MPC value function are interesting and nontrivial. However, the paper is heavily dependent on companion works: the uniform PE lemma (Lemma 1) is not proved here, the RLS bound is imported from [35], the final stochastic-ISS step imports [40, Theorem 2], and the DB Lipschitz lemma is from [47]. The algebraic core (Lemmas 2, 3, 5, 7, 8 and Proposition 2) is presented in detail, but one key inequality in Lemma 5 appears invalid. The contribution is therefore promising but not yet established in the manuscript as written.","major_comments":[{"comment":"Lemma 1 is the sole source of the uniform PE constants c_PE, p_PE used by Proposition 1 to define T_burn-in and the RLS error bound ϵ_t. Its proof is not in this manuscript: Remark 6 explicitly sends the reader to the companion paper [46]. The lemma quantifies over all x, ζ, s, and θ with no compactness or closeness assumptions, so this is not a routine detail. Moreover, Section VI states ‘We proved the probabilistic PE condition…’, which contradicts Remark 6. Since T_burn-in, Eq. (25), Proposition 3, and Theorem 1 all collapse if the PE constants are not available, the full proof (or a self-contained statement with explicit c_PE, p_PE) must appear in this paper before the main claim can be assessed.","section":"Section IV-A, Lemma 1 and Remark 6"},{"comment":"The displayed inequality E[σ(|˜g1(x)+R*(π diff)+R*¯v+Rκ(A,I)¯w|)] ≤ E[σ(|˜g1(x)|+D_MPC∥R*∥∥θ̂−θ*∥+|h|)] = σ(...) is not valid. Since σ(s)=(e^s−1)^2 is convex and increasing, an upper bound on E[σ(|Y+N|)] requires control of E[e^{2|N|}], whereas h = ln E e^{|R*¯v|} + ln E e^{|Rκ(A,I)¯w|} only controls E[e^{|N|}] (and in fact E[e^{2|N|}] ≥ (E[e^{|N|}])^2, typically strictly). For a concrete failure mode, take N∼N(0,1); then E[(e^{|N|}−1)^2] ≈ 9.9, while (e^{ln E e^{|N|}}−1)^2 ≈ 3.1. The same issue propagates to Lemma 7, Proposition 2, and Theorem 1 through the terms E_w1, E_w2, and E_w. The proof needs a corrected bound, e.g. using a bound on E[e^{2|N|}] or a different concentration argument, and the downstream conditions (12), (50) must be re-derived accordingly.","section":"Section IV-B, Lemma 5, Eq. (43)"}],"minor_comments":[{"comment":"Remark 4 says the theorem gives probability at least 1−δ, but Theorem 1 states 1−3δ. This mismatch should be corrected.","section":"Section III-C, Remark 4"},{"comment":"The conclusion describes ‘additive i.i.d bounded stochastic disturbances’, but Assumption 1 and the abstract allow unbounded sub-Gaussian noise. Please fix the wording.","section":"Section VI"},{"comment":"Both lemmas refer to ‘p being defined in (19)’, but p is defined in Eq. (16); Eq. (19) defines q_3^max. This typo should be corrected in the final version.","section":"Section IV-B, Lemmas 2 and 7"},{"comment":"Condition (12) and Corollary 1 are expressed in terms of the true unknown matrices R_* and A, e.g. σ_min(R_*), ∥R_*^{-1} A^κ∥, and ¯E_w. As a result, the horizon/excitation conditions are not verifiable from prior knowledge. The theorem is still a valid existence statement, but the paper should clarify this and, if possible, provide computable sufficient conditions or bounds.","section":"Section III-C, Theorem 1 and Corollary 1"},{"comment":"There are several typos and small errors: ‘asic algebraic manipulations’, ‘taht’, ‘PN’ in Sec. IV-B, and an undefined H_1 in the proof of Lemma 4 (the constraint matrix M is defined but H_1 is not). A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main concern for the editor is dependency: Lemma 1, Proposition 1, [40, Theorem 2], and [47] are all from the same author group and the first is not included. Even setting that aside, the inequality in Eq. (43) of Lemma 5 is a concrete mathematical error in the central derivation, not a mere presentation issue. I would encourage the editor to insist on a corrected proof of Lemma 5 and a self-contained proof (or a publicly available, verifiable proof) of Lemma 1 before any acceptance decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this paper proves a genuinely new result — a high-probability, non-asymptotic global practical stability bound for certainty-equivalence MPC with online RLS identification, hard input constraints, and unbounded sub-Gaussian noise, without terminal constraints, a pre-stabilizing controller, or a known compact parameter set. That combination is not in the literature. The architecture (excited RLS, saturated deadbeat fallback, hysteresis switching to keep the closed-loop map globally Lipschitz in the parameters) is sensible, and the proof skeleton is coherent.\n\nWhat is actually new and good: Lemmas 2 and 3, giving Lyapunov decreases for the known-parameter deadbeat and MPC policies, are derived in the paper; the exponential stage cost trick to get a state-independent horizon is clean; and the switching analysis in Proposition 2 is honest, with explicit case splits. The paper also does not hide its debts — it repeatedly flags which results come from companion papers.\n\nThe soft spot is exactly where the stress-test points: Lemma 1, the uniform persistent-excitation / anti-concentration lemma, is the load-bearing new identification ingredient, and its proof is only referenced to the companion paper [46]. This is not cosmetic. The MPC policy is not globally Lipschitz in the state; uniformity over all x, theta, and both modes is what guarantees excitation along unstable closed-loop trajectories before burn-in. If that lemma fails, Proposition 1 and Theorem 1 collapse. The RLS bound in Proposition 1 is imported from [35], and the final stochastic-ISS step comes from [40] — all same-group preprints. That is acceptable if those results hold, but it makes this paper a dependent application rather than a self-contained proof. Also, condition (12) is expressed in terms of the unknown true system matrices; Corollary 1 gives conservative sufficient bounds, so the theorem as stated is non-constructive. Minor but real: the conclusion says \"bounded stochastic disturbances\", while the abstract and body say unbounded sub-Gaussian noise. That inconsistency should be fixed.\n\nWho is this for: people in adaptive MPC, learning-based control, and finite-time system identification. It deserves a serious referee — not a desk reject — and I would send it out, with a clear request to include or attach the proof of Lemma 1 and to restate the horizon condition with known constants. As it stands, I would not cite it as a self-contained guarantee.","headline":"Genuinely new CE-MPC stability theorem whose main identification lemma lives in a companion paper; send it to referees, but require the missing proof before acceptance.","tokens_in":29953,"tokens_out":1944,"would_cite":false,"duration_ms":22840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C55","93E35","93D30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a certainty-equivalence MPC-deadbeat switching law with online least-squares estimation renders unknown input-constrained linear systems globally practically stable in high probability, with explicit non-asymptotic bo","keywords":["learning-based MPC","certainty-equivalence control","recursive least squares","input constraints","sub-Gaussian noise","non-asymptotic bounds","persistent excitation","hysteresis switching"],"falsifier":"For the numerical example in Section V, compute the empirical probability that ζ^⊤[(x+w);(π_s(x+w,θ)+v)]^2 ≥ c_PE over many noise draws, sweeping x over a large grid, ζ over the unit sphere, and θ in a neighborhood of θ̂; if for some (x,ζ,θ,s) this probability falls below the asserted p_PE, Lemma 1 fails and the burn-in bound of Proposition 1 no longer follows. More directly, run long closed-loop simulations and check whether the RLS error bound (25) holds after T_burn-in; a single episode in which ∥θ̂_t−θ*∥ exceeds ϵ_t falsifies the claim.","tokens_in":28917,"feed_emoji":"🎛️","tokens_out":4860,"duration_ms":49226,"temperature":0.7,"pith_summary":"The paper tries to prove that a controller can learn and stabilize an unknown linear system at the same time, even when the input is hard-saturated and the noise is unbounded. It proposes a certainty-equivalence scheme that estimates the system matrices online with regularized least squares and switches between two feedback laws: model predictive control and a saturated deadbeat controller, with dither added for excitation. The central result is a high-probability bound: after a finite transient, the state stays inside a tube that shrinks to a noise-dependent offset, with failure probability at most 3δ. This matters because existing adaptive MPC guarantees usually require a known uncertainty set, a pre-stabilizing controller, or terminal ingredients; here none of those are assumed. If the proof is correct, it turns simultaneous learning and constrained stabilization into a quantifiable, global property rather than a regional one.","feed_headline":"One controller learns and stabilizes unknown linear systems","feed_subtitle":"Switching between MPC and a deadbeat law gives global high-probability stability without terminal costs or a known model set.","key_machinery":"The central object is the four-part controller assembled in Algorithm 1: an RLS estimator, a dither signal v uniform on [−C,C]^κm, a saturated deadbeat law π_DB, and an MPC law π_MPC, glued together by a hysteresis switching rule with width c_3. The hysteresis is load-bearing because it restores global Lipschitz continuity of the control law with respect to the estimated parameters—a property the MPC law alone lacks. The exponential stage cost ℓ(x,u)=Q(e^{|x|}−1)^2+u^⊤Ru supplies a state-independent horizon bound that makes the finite-horizon value function decrease without terminal constraints. These pieces let the authors propagate finite-time RLS error bounds through the value functions a","core_discovery":"The central claim is Theorem 1: under sub-Gaussian disturbances and κ-step reachability, the switching law of Algorithm 1 with a prediction horizon N satisfying condition (12) gives P(|x_t| ≤ η̃(t) + c̃) ≥ 1 − 3δ for all t ≥ t_1, where η̃ ∈ L decays and c̃ is a noise-dependent offset. The proof chain is: RLS with injected dither ensures persistent excitation (Lemma 1), so the estimation error is bounded non-asymptotically after a finite burn-in (Proposition 1); the exponential stage cost makes the finite-horizon MPC value function a Lyapunov function without terminal constraints (Lemma 3); hysteresis switching makes the closed-loop law globally Lipschitz in the parameter estimate; together t","pith_inferences":["The hysteresis-switching construction suggests a general recipe: when a certainty-equivalence controller is only locally Lipschitz in the parameters, adding hysteresis between two controllers can recover global Lipschitz continuity and make finite-time identification errors propagatable through the value function—this could transfer to other adaptive MPC designs.","The exponential stage cost is doing heavy lifting; the same trick may yield terminal-cost-free stability bounds for other constrained nonlinear MPC problems, though it makes the optimization more nonlinear.","The uniform persistency-of-excitation lemma is the unverified load-bearing block; a natural next step is to compute or bound c_PE and p_PE explicitly for structured examples, since the burn-in time and all downstream bounds scale inversely with them.","A systematic numerical sweep over dither amplitude C, horizon N, and noise variance could map the conservatism of the bound and test whether the ensemble behavior observed in the simulation reflects the worst-case theory."],"forward_implications":["Global practical stability: the high-probability bound holds for any initial state, so learning can begin while the plant is far from the origin.","No terminal constraints or known compact parameter set: the only requirements are κ-step reachability, a known noise-variance lower bound, and ∥A∥ ≤ 1.","Explicit controller tuning: Corollary 1 gives concrete sufficient conditions on N, R/Q, and u_max that satisfy the main stability condition (12).","Quantified failure probability: the 3δ term decomposes into estimation, block-level, and per-step events, so a user can trade off confidence against horizon and excitation levels.","Hard input constraints are respected at every time step by reserving C of u_max for dither and saturating the deadbeat law."],"fun_headline_variants":["Switching MPC achieves global stability under unbounded noise","Online learning in MPC: non-asymptotic stability guarantees","No model? MPC learns and stabilizes with high probability","Unbounded noise? Learning-based MPC still guarantees stability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole result rests on Lemma 1's claim that the closed-loop regressor is uniformly persistently exciting for every state, estimate, and controller mode with fixed constants c_PE and p_PE—a uniform anti-concentration bound whose proof is deferred to a companion paper and is not verified here.","fun_headline_variants_meta":{"raw":{"variants":["Switching MPC achieves global stability under unbounded noise","Online learning in MPC: non-asymptotic stability guarantees","No model? MPC learns and stabilizes with high probability","Unbounded noise? Learning-based MPC still guarantees stability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000587,"raw_usage":{"total_tokens":2551,"prompt_tokens":659,"completion_tokens":1892,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":403,"completion_tokens_details":{"reasoning_tokens":1826}},"tokens_in":403,"tokens_out":1892,"duration_ms":12758,"temperature":1.0,"reasoning_tokens":1826,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T04:59:03.326318+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For the numerical example in Section V, compute the empirical probability that ζ^⊤[(x+w);(π_s(x+w,θ)+v)]^2 ≥ c_PE over many noise draws, sweeping x over a large grid, ζ over the unit sphere, and θ in a neighborhood of θ̂; if for some (x,ζ,θ,s) this probability falls below the asserted p_PE, Lemma 1 fails and the burn-in bound of Proposition 1 no longer follows. More directly, run long closed-loop simulations and check whether the RLS error bound (25) holds after T_burn-in; a single episode in which ∥θ̂_t−θ*∥ exceeds ϵ_t falsifies the claim.","supporting_citations":[],"review_version":1}