{"id":"44f89a13-ed8f-4909-99ed-461f18e4cb7c","arxiv_id":"2507.02131","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Perturbed gradient descent is small-disturbance input-to-state stable under a generalized nonlinear PL condition, and LQR policy gradient methods inherit this guarantee.","lead":"This paper shows that gradient descent with noisy gradient estimates still settles near the optimum if the objective satisfies a nonlinear version of the Polyak-Lojasiewicz condition. It then proves that policy gradient, natural policy gradient, and Gauss-Newton methods for linear quadratic regulator control all have this robustness property.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1 is false as stated: step sizes with no lower bound (e.g., η(k)=2^{-k} for J(z)=z^2) satisfy 0<η≤1/L but fail to converge, so the theorem's Lyapunov argument has no valid decay rate.","rationale":"The reader identified the lack of a fixed lower bound on the step size as a fragile premise, but treated the overall situation as 'correctable gaps, not contradictions.' My analysis shows the gap is stronger: Theorem 4.1 is false as written, with an elementary counterexample satisfying every assumption. This is therefore the single most load-bearing concern. The imported K-PL estimate Lemma 5.4 is also a serious gap for the LQR section, since Theorems 5.1–5.3 depend on it and the proof is not included; however, the step-size issue is decisive because it invalidates the central theorem even when Lemma 5.4 is granted. The counterexample is not an artifact of an unusual perturbation: it uses zero disturbance and a legitimate step-size sequence, so the claimed ISS guarantee fails in the noise-free case. A repair is conceivable—for example, requiring η(k) to be bounded below by a state-dependent positive function and re-proving Definition 3.3 with a genuine positive definite α2—but the manuscript does not supply such a condition, and Theorems 5.2 and 5.3 have the same structural omission. For this reason, I recommend rejecting the current formulation, while acknowledging the framework and the LQR Lipschitz estimates (Lemmas 5.1–5.3) appear useful and largely self-contained.","tokens_in":21958,"tokens_out":9364,"duration_ms":116278,"concrete_test":"Run the one-dimensional example J(z)=z^2 with η(k)=2^{-(k+2)} and e(k)=0. Verify that z(k)=z0∏_{j=0}^{k−1}(1−2^{-(j+1)}) and that, since ∑2^{-(j+1)}<∞, this product converges to a positive constant, so z(k) does not tend to 0 for z0≠0. This directly contradicts the claimed small-disturbance ISS property in Theorem 4.1. If the authors intend a different step-size rule, the manuscript should state a lower bound on η(k) and exhibit the corresponding α2 in Definition 3.3.","verdict_should_be":"REJECT","load_bearing_attack":"The main theorem is false as stated because it imposes only 0<η(k)≤1/L(J(z(k))) with no lower bound. The proof's inequality (27) gives decay proportional to η(k), so the dissipation rate can vanish. A concrete counterexample: take J(z)=z^2 on R. Then J is coercive, ∇J is 2-Lipschitz on every sublevel set, and ∥∇J(z)∥=2|z|=2√(J(z)−J(z*)), so the K-PL condition holds with α5(r)=2√r. Choose η(k)=2^{-(k+2)} and set the disturbance e(k)=0. Then 0<η(k)≤1/2=1/L, but z(k+1)=(1−2^{-(k+1)})z(k), so z(k)=z0∏_{j=0}^{k−1}(1−2^{-(j+1)}) converges to a positive constant times z0, not to zero, whenever z0≠0. Hence the unperturbed system is not asymptotically stable and cannot be small-disturbance ISS. The appeal to Theorem 3.1 requires a positive definite α2 independent of k; under the stated step-size condition no such α2 exists. The same missing lower bound infects Theorems 5.1–5.3, which rely on Theorem 4.1 and on analogous step-size conditions with no lower bound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a discrete-time small-disturbance ISS framework with a Lyapunov characterization (Theorem 3.1), generalizes the linear Polyak–Łojasiewicz condition to a K-PL condition, and claims that gradient descent with step sizes 0<η(k)≤1/L(J(z(k))) is small-disturbance ISS under coercivity, sublevel-set Lipschitz gradients, and K-PL (Theorem 4.1). It then applies this result to the LQR cost, claiming that standard policy gradient, natural policy gradient, and Gauss-Newton methods are small-disturbance ISS (Theorems 5.1–5.3).","tokens_in":22263,"tokens_out":8423,"duration_ms":93548,"significance":"If the main theorem were correct, the proposed K-PL framework would provide a useful robustness guarantee for nonconvex gradient methods and would give a clean control-theoretic interpretation of perturbed policy optimization for LQR. The discrete-time Lyapunov characterization in Theorem 3.1 is a valuable contribution in itself. However, the central step-size condition in Theorem 4.1 is missing a uniform lower bound, and the theorem is false as stated; the same defect propagates to the LQR theorems. The LQR results also rely on a global K-PL estimate imported without proof from a same-author paper. With a corrected statement and a complete proof of the imported lemma, the framework could be significant, but the current version is not publishable in this form.","major_comments":[{"comment":"Theorem 4.1 is false as stated. The proof obtains the estimate J(z(k+1))-J(z(k)) ≤ -(3η(k)/8)α5(J(z(k))-J(z*))² in Eq. (27), but the hypothesis 0<η(k)≤1/L(J(z(k))) imposes no lower bound on η(k), so the decay rate can vanish along the trajectory. This prevents application of Theorem 3.1, which requires a K∞ function α2 in Definition 3.3 independent of k. A concrete counterexample is J(z)=z² on Z=R with L=2, α5(r)=2√r, and η(k)=2^{-(k+2)}; all hypotheses of Definition 4.1 hold, but with e≡0 the iteration is z(k+1)=(1-2^{-(k+1)})z(k), which converges to a nonzero multiple of z(0) whenever z(0)≠0. This violates the 0-input asymptotic stability required by Definition 3.2. The theorem must be repaired by adding a uniform lower bound on the step size, such as η(k)≥η_min>0, and the same correction must be propagated through Corollaries 4.1–4.2 and Theorems 5.1–5.3.","section":"Theorem 4.1, Eq. (27)"},{"comment":"The global K-PL estimate for the LQR cost, ∥∇J2(K)∥F ≥ r/(b1r+b2) with r=J2(K)-J2(K*), is stated without proof and cited as Lemma 5.7 of [12], a paper by the same first three authors. This estimate is load-bearing: without it, Theorems 5.1–5.3 do not follow from the arguments presented. The authors should include a complete proof or at least a detailed derivation of the constants b1 and b2 so that the reader can verify the estimate independently, rather than relying on a same-author citation for a central ingredient.","section":"Lemma 5.4, Eq. (55)"},{"comment":"The natural-gradient and Gauss-Newton theorems inherit the step-size lower-bound problem. In Theorem 5.2, inequality (80) gives V5(K(k+1))-V5(K(k)) ≤ -η(k)λmin(R)V5(K(k))/4, and the stated condition 0<η(k)≤min{1/(2∥R∥),1/(6∥R∥c(K(k)))} has no lower bound. Hence V5 is not a small-disturbance ISS-Lyapunov function under Definition 3.3 unless η(k) is uniformly bounded below. Theorem 5.3 has the same issue with the condition 0<η(k)≤min{1,1/(4c(K(k)))}, and its proof is only a one-sentence sketch. These theorems need the corrected step-size hypothesis and a fully written proof for the Gauss-Newton case.","section":"Theorems 5.2 and 5.3"}],"minor_comments":[{"comment":"The displayed update for standard gradient descent appears to be missing the K(k) term: it should read K(k+1)=K(k)-η(k)(∇J2(K(k))+W(k)), and the argument of ∇J2 is written as k(k) instead of K(k).","section":"Eq. (54)"},{"comment":"The step-size condition is written as 0<η(k)≤1/L(J(K(k))), but L is defined for J2 in Lemma 5.3; the notation should be L(J2(K(k))) for consistency.","section":"Theorem 5.1 and its proof"},{"comment":"The name Łojasiewicz is typeset with a spurious space in the abstract; this should be corrected throughout for consistency with the reference list.","section":"Abstract and Introduction"},{"comment":"The K-PL estimate is displayed as ∥∇J(z)∥ ≥ J(z)/(√2/2+2J(z)); it would be helpful to show the intermediate algebra, as the expression is not immediately transparent and is used to illustrate the framework.","section":"Example 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's LQR half leans heavily on [12] and [10], both by overlapping author sets, for coercivity and the global K-PL estimate. This is not improper citation, but it means the genuinely new content is Theorem 4.1 and the discrete-time Lyapunov characterization. Since Theorem 4.1 is false as stated, the manuscript needs substantial revision before it can be considered. If the authors add a uniform step-size lower bound and provide a self-contained proof of Lemma 5.4, the corrected framework would be a reasonable contribution to Automatica."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the bad news: Theorem 4.1 is false as stated. The step-size condition 0 < η(k) ≤ 1/L(J(z(k))) has no uniform lower bound, and the proof's inequality (27) gives decay proportional to η(k). Take J(z)=z^2, η(k)=2^{-(k+2)}, e(k)=0. Then z(k+1)=(1−2^{-(k+1)})z(k), so z(k) converges to a nonzero multiple of z0, not to zero. The unforced system is not asymptotically stable, so small-disturbance ISS cannot hold. The appeal to Theorem 3.1 requires a K∞ decay rate independent of k, and that rate does not exist. The same gap infects Theorems 5.1–5.3.\n\nThat said, the paper is not empty. The discrete-time small-disturbance ISS Lyapunov characterization in Section 3 is a genuine extension of the continuous-time work, and the converse proof is nontrivial. The LQR analysis—coercivity, Lipschitz gradient over sublevel sets, and the natural-gradient Lyapunov function—is substantial and mostly careful. Lemma 5.3's Lipschitz bound is worked out in detail. The natural gradient and Gauss-Newton applications are new relative to [12], though they inherit the step-size issue, and Theorem 5.3 is asserted rather than proved.\n\nThe soft spots, in proportion: the step-size flaw is load-bearing. Adding a uniform lower bound on η(k), or taking a constant step size, would likely fix it; the counterexample is a vanishing step size, not a deeper defect. Lemma 5.4 is imported from the authors' own prior paper without proof. That is not circularity, but a referee should verify the estimate because Theorems 5.1–5.3 rest on it. Theorem 5.3 needs a real proof, not a one-sentence reference to Theorem 5.2.\n\nWho gets value from this paper? People working on robustness certificates for policy gradient methods, and control theorists applying ISS tools to optimization. The K-PL framework is worth knowing. But in current form the main claim is wrong, so it needs major revision before it is citable.\n\nIf this lands on your desk, send it to review—it is fixable and the framework is useful—but tell the referee to check the step-size condition carefully before accepting.","headline":"A useful ISS framework for gradient methods, but the main theorem is false as stated because vanishing step sizes break the Lyapunov decay.","tokens_in":22800,"tokens_out":3325,"would_cite":false,"duration_ms":35437,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93D25","93D30","93C55","90C30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Perturbed gradient descent is small-disturbance ISS when the objective satisfies a nonlinear Polyak-Lojasiewicz condition; LQR policy gradient, natural gradient, and Gauss-Newton all qualify.","keywords":["input-to-state stability","small-disturbance ISS","gradient descent","Polyak-Lojasiewicz condition","policy gradient","linear quadratic regulator","Lyapunov function","discrete-time nonlinear systems"],"falsifier":"For a concrete LQR instance, such as $A=0$, $B=Q=R=1$ where $J_2(K)=(K^2+1)/(2K)$ on $K>0$, evaluate whether $\\|\\nabla J_2(K)\\|_F \\ge (J_2(K)-J_2(K^*))/(b_1(J_2(K)-J_2(K^*))+b_2)$ holds for all stabilizing gains with the constants $b_1,b_2$ of equation (55); a single violating gain would invalidate Theorem 5.1.","tokens_in":21759,"feed_emoji":"📉","tokens_out":4663,"duration_ms":47558,"temperature":0.7,"pith_summary":"This paper establishes that gradient descent remains robust when the gradient is corrupted by a bounded disturbance, provided the objective satisfies a nonlinear generalization of the Polyak-Lojasiewicz condition, called the K-PL condition. The notion of small-disturbance input-to-state stability (ISS) is introduced for discrete-time systems, with a Lyapunov characterization, and it is shown that perturbed gradient descent with step size bounded by the inverse of the local Lipschitz constant is small-disturbance ISS. A direct application shows that the LQR cost satisfies the K-PL condition, so the standard policy gradient algorithm for LQR, as well as natural gradient and Gauss-Newton variants, are all small-disturbance ISS under gradient estimation errors. The paper's significance is that it replaces convexity or linearly decaying perturbation assumptions with a more flexible gradient-dominance condition, giving quantitative robustness guarantees for optimization algorithms that are widely used in reinforcement learning.","feed_headline":"Noisy gradient descent still lands near the optimum","feed_subtitle":"A nonlinear Polyak-Lojasiewicz condition turns bounded gradient errors into bounded suboptimality—proved for LQR policy gradient.","key_machinery":"The load-bearing concept is the K-PL condition, a nonlinear replacement for the classical Polyak-Lojasiewicz inequality: $\\|\\nabla J(z)\\| \\ge \\alpha_5(J(z)-J(z^*))$ with $\\alpha_5$ in class K, together with coercivity and Lipschitz continuity of the gradient on sublevel sets. The proof uses the cost gap $J(z)-J(z^*)$ as a small-disturbance ISS-Lyapunov function and shows that its decrease is proportional to $\\alpha_5(J(z)-J(z^*))^2$ plus a term quadratic in the disturbance. For the LQR application, the paper supplies a Lipschitz constant $L(h)$ for the gradient over sublevel sets and invokes a global K-PL estimate for the LQR cost, with the saturating function $\\alpha_6$ from Lemma 5.4, so that Theorem 4.1 applies directly to policy gradient updates.","core_discovery":"The central claim is Theorem 4.1: if the objective function J is coercive, its gradient is Lipschitz on sublevel sets, and it satisfies the K-PL estimate $\\|\\nabla J(z)\\| \\ge \\alpha_5(J(z)-J(z^*))$ for a class-K function $\\alpha_5$, then the perturbed gradient descent $z(k+1)=z(k)-\\eta(k)(\\nabla J(z(k))+e(k))$ is small-disturbance ISS whenever $0<\\eta(k)\\le 1/L(J(z(k)))$. The ultimate bound on the cost gap is $\\alpha_5^{-1}(2\\|e\\|_\\infty)$, a nonlinear function of the disturbance amplitude. The paper then verifies the required properties for the LQR cost $J_2(K)$, including a global K-PL estimate with the saturating function $\\alpha_6(r)=r/(b_1 r+b_2)$, and thereby proves that the standard policy gradient, natural gradient, and Gauss-Newton methods for LQR are small-disturbance ISS.","pith_inferences":["The same Lyapunov-template argument could be attempted for other policy-gradient variants such as TRPO or PPO, provided an analogous K-PL estimate is established; the paper does not attempt this extension.","The saturating K-PL function $\\alpha_6(r)=r/(b_1 r+b_2)$ implies that the steady-state LQR cost error grows sublinearly with the gradient-noise amplitude, a quantitative prediction that numerical simulations could verify.","The LQR robustness theorems depend on Lemma 5.4, whose proof is deferred to a companion paper; a self-contained proof would make the LQR claims independently verifiable from this manuscript alone."],"forward_implications":["Bounded gradient noise from round-off, sampling, or estimation errors no longer causes divergence; the iterates eventually stay inside the set $\\{z : J(z)-J(z^*) \\le \\alpha_5^{-1}(2\\|e\\|_\\infty)\\}$, whose size is controlled by the noise bound.","For LQR policy optimization, model-free implementations with finite-difference or adaptive-dynamic-programming gradient estimators inherit a quantitative robustness guarantee against estimation errors.","The natural gradient and Gauss-Newton updates for LQR are also small-disturbance ISS under explicit step-size constraints, so the robustness conclusion is not tied to one update rule.","If the K-PL function is strengthened to class K-infinity, the gradient descent is ISS; if relaxed to a positive definite function, it is integral ISS, with the classical linear-PL case recovering exponential ISS as a special case."],"supporting_citations":[{"why":"Supplies the continuous-time small-disturbance ISS concept and the global K-PL estimate for the LQR cost that Lemma 5.4 uses.","marker":"[12]"},{"why":"Provides the discrete-time ISS and ISS-Lyapunov theory, including the comparison lemma, used in Theorem 3.1.","marker":"[25]"},{"why":"Introduces size functions and the ISS framework for perturbed gradient flows, including Lemma 2.6 and coercivity of sublevel sets.","marker":"[56]"},{"why":"Supplies the descent lemma for smooth functions used to bound the cost decrease along a gradient step.","marker":"[42]"},{"why":"Establishes coercivity of the LQR cost over the set of stabilizing gains, invoked in the proof of Theorem 5.1.","marker":"[10]"},{"why":"Gives the Lyapunov equation solution formula that underpins the LQR cost, its gradient, and the associated matrix bounds.","marker":"[54]"}],"fun_headline_variants":["Small-noise gradient descent stays near the optimum","Perturbed gradient descent is input-to-state stable","Gradient descent resists small disturbances, proves new bound","K-PL condition makes gradient descent robust to noise","LQR policy gradient tolerates small perturbations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The crucial premise is that the LQR cost's gradient never becomes too small compared with how far the current cost is from optimal, for every stabilizing controller; the paper cites this global K-PL estimate from a companion paper without proving it here.","fun_headline_variants_meta":{"raw":{"variants":["Small-noise gradient descent stays near the optimum","Perturbed gradient descent is input-to-state stable","Gradient descent resists small disturbances, proves new bound","K-PL condition makes gradient descent robust to noise","LQR policy gradient tolerates small perturbations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1733,"prompt_tokens":924,"completion_tokens":809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":736}},"tokens_in":540,"tokens_out":809,"duration_ms":8292,"temperature":1.0,"reasoning_tokens":736,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:38:47.023597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a concrete LQR instance, such as $A=0$, $B=Q=R=1$ where $J_2(K)=(K^2+1)/(2K)$ on $K>0$, evaluate whether $\\|\\nabla J_2(K)\\|_F \\ge (J_2(K)-J_2(K^*))/(b_1(J_2(K)-J_2(K^*))+b_2)$ holds for all stabilizing gains with the constants $b_1,b_2$ of equation (55); a single violating gain would invalidate Theorem 5.1.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the continuous-time small-disturbance ISS concept and the global K-PL estimate for the LQR cost that Lemma 5.4 uses."},{"cited_title":"Input-to-state stability for discrete-time nonlinear systems","cited_arxiv_id":null,"evidence_quote":"Provides the discrete-time ISS and ISS-Lyapunov theory, including the comparison lemma, used in Theorem 3.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces size functions and the ISS framework for perturbed gradient flows, including Lemma 2.6 and coercivity of sublevel sets."},{"cited_title":"Introductory Lectures on Convex Optimization: A Basic Course , volume 87","cited_arxiv_id":null,"evidence_quote":"Supplies the descent lemma for smooth functions used to bound the cost decrease along a gradient step."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Lyapunov equation solution formula that underpins the LQR cost, its gradient, and the associated matrix bounds."}],"review_version":1}