{"id":"2ff586c5-2f8e-44ec-ba6e-74bc4aaae783","arxiv_id":"1908.02447","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An optimization-based adaptive iterative learning control algorithm achieves perfect tracking for unknown nonlinear time-varying systems from measured data only, with bounded parameter estimates and robustness to nonrepetitive disturbances.","lead":"This paper designs a learning controller that lets a repetitive nonlinear system learn perfect tracking from measured input-output data, with no explicit model. It proves all signals stay bounded and the tracking error converges, even when disturbances change from run to run.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Projection step (16) is omitted from the proof of Theorem 1, so the stated β_hat bound and condition (41) are not established.","rationale":"The reader's weakest_assumption focuses on the global regularity condition (A1). My review finds the core convergence argument under A1 to be largely coherent, but identifies a different load-bearing gap: the projection (16) in Algorithm 1 is not included in the analysis of Theorem 1. The bound β_hat in (29) is used directly in the selection condition (41) and in Lemma 8; if the reset can make the implemented estimates larger than that bound, the stated sufficient condition is not proven. This is an internal proof gap rather than a criticism of the assumption. It supports the existing CONDITIONAL verdict (hence UNCHANGED), because the issue is concrete and requires a proof revision, but the main idea appears salvageable. I disagree with the reader's choice of weakest assumption; A1 is a strong but standard scope condition, while the projection/bound mismatch is a correctness risk in the submitted argument.","tokens_in":37904,"tokens_out":24321,"duration_ms":260652,"concrete_test":"Prove Theorem 1 for the actual projected recursion by deriving a corrected bound: with ρ=µ1/(µ1+µ2), a reset at step k gives ||θhat_k|| ≤ ρ||θhat_{k-1}|| + √Tβ_θ + ||θhat_{1,0}||, so the steady-state bound increases by ||θhat_{1,0}||/(1−ρ) beyond (29). Then run Algorithm 1 on the scalar-order system y_k(t+1)=a y_k(t)+b u_k(t)+sin(y_k(t)) with t=1, µ2=0.001, and an initial estimate with large off-diagonal entry; record max_k ||θhat_k|| and check whether it exceeds the value in (29) whenever the reset (16) is active. If it does, verify whether the λ from (41) still guarantees the row sums in (35) and (40) with the corrected bound; if not, the stated design condition must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The boundedness proof of Theorem 1 analyzes the unprojected update (13)/(15), deriving the contraction bound (29) for the estimate vector. Algorithm 1's actual update includes the reset (16): whenever the diagonal estimate θhat_{k,k-1,t}(t) drops below ε, that component is replaced by the initial estimate. Because only one component is reset while the other t components retain the candidate values, the l2 norm of the implemented estimate can exceed the bound (29); the recurrence (28) no longer describes the actual sequence. Since Theorem 2's sufficient condition (41) and Lemma 8's verification of the contraction row-sum conditions (35) and (40) both use the specific β_hat from (29), the displayed λ condition may be insufficient for the algorithm as stated. The gap is repairable by enlarging β_hat to account for resets, but the proof as written is incomplete.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an optimization-based adaptive iterative learning control (ILC) scheme (Algorithm 1) for single-input single-output nonlinear time-varying discrete-time systems described by (1). The design relies on a dynamical linearization (Lemma 1) under a global differentiability and sign condition (Assumption (A1)), an optimization-based parameter estimator (12) with a projection/reset mechanism (16), and an input update (17) obtained from a high-order quadratic index (3). The main convergence result (Theorem 2) claims boundedness of the input and output sequences and asymptotic perfect tracking under the selection condition (41). Section V extends the scheme to systems with bounded nonrepetitive disturbances and initial shifts (51) and states a robust tracking result (Theorem 3). Simulation tests in Section VI illustrate the behavior on a nonlinear example.","tokens_in":38080,"tokens_out":8739,"duration_ms":100551,"significance":"If the main results are correct, the paper would offer a data-driven ILC law that does not require an explicit system model, handles nonlinear time-varying dynamics with bounded and sign-fixed input-output coupling, and avoids the usual eigenvalue-based contraction-mapping analysis by using a double-dynamics argument with nonnegative matrices. The proof of Theorem 2 is genuinely inductive and quite detailed, and the optimization-based estimator with the projection (16) is a plausible mechanism for keeping parameter estimates bounded. The paper also makes a useful robustness claim for nonrepetitive disturbances, which is not well covered by earlier optimization-based adaptive ILC results. However, several load-bearing issues remain: the proof of Theorem 1 does not account for the projection step (16), the proof of Theorem 3 is omitted for the main boundedness and tracking claims, and the claimed T-independence of condition (41) is contradicted by the definition of beta_hat in (29). These are repairable but require substantive revision.","major_comments":[{"comment":"The proof of Theorem 1 analyzes only the unprojected update (13)/(15) and derives the bound (29) for the sequence that never experiences the reset (16). The actual Algorithm 1 resets the diagonal component to the initial estimate whenever it drops below epsilon. Since the reset can increase the Euclidean norm relative to the unprojected candidate, the recurrence (28) and the bound (29) do not apply to the implemented sequence. Because beta_hat from (29) is used in Lemma 8 and in the sufficient condition (41), the stated lambda condition may not be valid for the algorithm as written. The gap appears repairable by an induction that explicitly handles the reset, or by enlarging beta_hat to cover the reset value, but the proof must be supplied.","section":"Section IV, Theorem 1 proof; Algorithm 1 step (S2), Eq. (16)"},{"comment":"After deriving the perturbed error dynamics (59)-(60) and input dynamics (61)-(62), the proof states that the boundedness and robust/perfect tracking results follow 'by following the same steps as the proof of Theorem 2, which is thus omitted here.' These are central claims of the paper, and the omission is not acceptable: the additional disturbance and initial-shift terms in (60) and (62) affect the driving signals in the induction over t, and their boundedness and convergence must be verified explicitly. A proof (or at least a rigorous sketch covering all the new terms) is needed.","section":"Section V, Theorem 3 proof, parts 2) and 3)"},{"comment":"Remark 7 claims that the use of the double-dynamics approach makes the selection condition (41) independent of the learning time horizon T. This is not supported by the manuscript: the quantity beta_hat in Theorem 1 is defined in (29) as max_t ||theta_hat_{1,0,t}(t)||_2 + (mu_1+mu_2)/mu_2 sqrt(T) beta_theta, which depends on T. Since condition (41) contains beta_hat, it depends on T through beta_hat. The claimed improvement over the T-dependent conditions in, e.g., [27] is therefore not established and the remark should be revised or qualified.","section":"Remark 7 versus Eq. (29) and Eq. (41)"},{"comment":"The same symbol beta_f is used in (4) for the uniform upper bound on the partial derivatives and in (5) for the strictly positive lower bound on the input-output coupling derivative. This makes the interval in (6) degenerate, [beta_f, beta_f], which cannot hold for a general nonlinearity. Correspondingly, the diagonal interval in Lemma 1, Eq. (9), appears as [beta_f, beta_f]. If distinct lower and upper bounds, say underline{beta}_f and overline{beta}_f, were intended, the notation must be corrected consistently, because Lemma 8 and condition (41) rely on the lower bound beta_f via the product theta_{k,k-1,t}(t) theta_hat_{k,k-1,t}(t).","section":"Assumption (A1), Eqs. (4)-(6) and Lemma 1, Eq. (9)"}],"minor_comments":[{"comment":"The simulation sets the initial estimate as theta_hat_{0,-1,t}(i), while step (S1) defines the initial value as theta_hat_{1,0,t}(i). Either the algorithm or the simulation notation should be made consistent.","section":"Section III, Algorithm 1 and Section VI"},{"comment":"The statement that (14) guarantees theta_hat_{1,0,t}(t) has the same sign as the coupling derivative assumes without loss of generality that the sign is positive. This is fine, but the sign convention should be stated explicitly when discussing the projection in (16).","section":"Remark 3"},{"comment":"The caption and surrounding text of Figure 4 are corrupted in places (e.g., stray 'max 50' fragments), and the figure should be cleaned up before publication.","section":"Section VI, Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript leans on the authors' earlier results [33], [34] for the nonrepetitive ILC lemmas. This is legitimate, but the authors should state precisely which parts of Lemmas 5 and 7 are imported, since the current proof of Theorem 2 depends on those external results. The projection-gap in Theorem 1 and the omitted proof of Theorem 3 are the main technical obstacles; both are fixable but need to be addressed in full before the paper can be considered for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid incremental contribution to data-driven ILC, but there is a real gap in the proof of Theorem 1 and a false claim in Remark 7. I'd send it to review, but only with major revision, not accept.\n\nThe genuinely new piece is the µ2-regularized parameter update in (12), which gives bounded estimates without a step-size factor, plus the double-dynamics analysis that avoids eigenvalue conditions for iteration-dependent parameters. The induction over t in the proof of Theorem 2 is careful, and the robustness setup in Theorem 3 targets the right problem. If the main result holds, it is a useful extension of [24]-[30] within the subfield.\n\nNow the soft spots, in order of importance.\n\nFirst, the stress-test note about the projection step is correct. The proof of Theorem 1 analyzes the unprojected recurrence (13)/(24) and derives the bound (29) for that recurrence. But Algorithm 1 includes reset (16), which replaces a diagonal estimate by the initial value whenever it drops below ε. That replacement can only increase the 2-norm of the vector. The stated β_hat is therefore not a bound for the actual estimates as written. Since condition (41) and Lemma 8 use β_hat, the λ selection is not justified. This is repairable—e.g., add a term for the initial estimate size—but it has to be done.\n\nSecond, Theorem 3 is the robustness headline, and the proof is explicitly omitted. Relying on [33, Lemma 2] and saying 'same steps as Theorem 2' is not enough for a journal paper. The reader is being asked to take the main advertised contribution on faith.\n\nThird, Remark 7 says condition (41) is independent of the time horizon T. Looking at (29), β_hat contains sqrt(T), so (41) is not T-independent. The claim is contradicted by the paper's own bound. This is minor in the sense that the result may still hold, but the remark should be fixed or qualified.\n\nAssumption (A1) is a global Lipschitz-type condition; that is standard for this ILC family and I wouldn't call it a defect. The self-citation of [33] is not a problem—it is the natural tool for nonrepetitive convergence—but the dependence is heavy. The simulations are a single academic example; they confirm but don't make the theory.\n\nBottom line: for someone working in data-driven ILC, this is worth reading. The core strategy is promising and the proof skeleton is mostly coherent. With the projection gap closed and Theorem 3 actually proven, I'd be comfortable citing it. It deserves a serious referee.","headline":"A competent extension of data-driven ILC with a real proof gap in Theorem 1 and an overstated T-independence claim; worth review after major revision.","tokens_in":38587,"tokens_out":3950,"would_cite":false,"duration_ms":43881,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A purely data-driven iterative learning law can drive an unknown nonlinear time-varying system to perfect tracking using only measured input/output data.","keywords":["iterative learning control","data-driven control","nonlinear time-varying systems","optimization-based design","adaptive parameter estimation","nonrepetitive uncertainties","double-dynamics analysis","nonnegative matrices"],"falsifier":"Run Algorithm 1, with gains satisfying (41), on a smooth nonlinear plant that obeys (A1) and see whether the tracking error fails to converge to zero or any estimated parameter exceeds the predicted bound $\\beta_{\\hat\\theta}$; a single such plant would refute Theorem 2. A sharper test: replace the input-output coupling term $t/(t+2)u_k(t)$ in the paper's simulation with a smooth function that changes sign at some time step while keeping the function smooth; if the tracking error still converges, then Assumption (A1) is not necessary, and if it diverges, the sign-fixity premise is confirmed as load-bearing.","tokens_in":37702,"feed_emoji":"🎯","tokens_out":6981,"duration_ms":66989,"temperature":0.7,"pith_summary":"The paper claims that a purely data-driven iterative learning control (ILC) law can drive the output of an unknown nonlinear time-varying system to perfect tracking of a desired trajectory as the number of repetitions grows, using only measured input/output data. It proposes Algorithm 1, which at each repetition solves two optimization problems: one to update the input and one to estimate the parameters of a dynamical linearization of the unknown plant. The main result (Theorem 2) states that under a global smoothness/sign condition on the plant and a simple gain condition ($\\gamma_1+\\gamma_2 > \\sum_{i=3}^m \\gamma_i$ and $\\lambda > (\\gamma_1^2+\\gamma_1\\gamma_2)\\bar{\\beta}_f\\beta_{\\hat\\theta}$), the input, output, and all estimated parameters remain bounded and the tracking error converges to zero. Theorem 3 extends the same guarantees to robust bounded tracking when the plant is hit by nonrepetitive disturbances and initial shifts. If correct, the work gives a model-free route to perfect repetition tasks for a broad class of nonlinear plants.","feed_headline":"Data-driven learning hits perfect tracking without a plant model","feed_subtitle":"Learns unknown nonlinear time-varying plants from input/output data alone, with bounded signals.","key_machinery":"The central object is the extended dynamical linearization (Lemma 1), which expresses the output difference between two iterations as $\\Theta_{i,j}\\Delta u$, a lower-triangular matrix whose entries are uniformly bounded and whose diagonal entries are the input-output coupling derivatives, kept positive and bounded away from zero by Assumption (A1). Around it the paper builds an optimization-based adaptive estimator (Lemma 3/step (S2)) whose cost includes the previous estimate's deviation and an $\\ell^2$ penalty that automatically keeps the estimate bounded, and an input update (17) derived from minimizing the weighted high-order error index (3). The convergence machinery is the double-dynamics analysis: the lifted error recursion (31) and the input recursion (38) are coupled through driving signals $\\kappa_k(t)$ and $\\psi_k(t)$; nonnegative-matrix arguments (Lemma 10) turn the scalar condition (35) into the contraction condition (C), and induction over time steps proves boundedness and convergence. The binding condition (41) is what makes both the error matrix and the input coefficient contract simultaneously, independent of the horizon length $T$.","core_discovery":"On its own terms, the paper establishes that the unknown nonlinear time-varying system (1) can be made to track a prescribed trajectory perfectly in the limit of iterations, without any explicit model knowledge. The proof rests on an extended dynamical linearization (Lemma 1) that represents the difference of outputs between two iterations as a lower-triangular matrix $\\Theta_{i,j}$ times the input difference, with all entries bounded and diagonal entries lying in $[\\underline{\\beta}_f,\\bar{\\beta}_f]$. Algorithm 1 estimates these entries by optimizing an index that penalizes the one-step prediction error, the change of the estimate along the iteration axis, and the size of the estimate; the last two terms make the estimates uniformly bounded by construction. The convergence analysis follows a double-dynamics approach: the tracking-error dynamics and the input dynamics are coupled nonrepetitive linear-like systems driven by signals that vanish when earlier time steps have converged, and induction over the time horizon completes the argument. The central parametric condition (41) is shown to be sufficient for both the error contraction and the input boundedness, and the same framework yields robust tracking (Theorem 3) under bounded nonrepetitive disturbances and initial shifts.","pith_inferences":["The analysis gives no quantitative convergence rate; a natural extension would be to bound the contraction factor $\\zeta$ in terms of the gains and plant bounds and compare it with classical contraction-mapping ILC rates.","Since the estimator is bounded via the $\\mu_2$ penalty, the method may also tolerate linearization parameters that drift slowly along the iteration axis, as the proofs appear to handle iteration-dependent parameters as long as they stay bounded.","One could test the sign-fixity assumption empirically by monitoring the estimated diagonal $\\hat\\theta_{k,k-1,t}(t)$; if it repeatedly hits the floor $\\varepsilon$, the plant likely violates Assumption (A1), and a sign-switching extension would be needed.","The robust result points toward a practical stopping rule: terminate when the maximal tracking error stops decreasing and scale the terminal error by the disturbance bounds."],"forward_implications":["A user can implement Algorithm 1 without knowing the plant equations or their parameters; only input/output measurements from previous repetitions are required.","The gain condition (41) is independent of the horizon length $T$, so longer tasks do not force tighter tuning of $\\lambda$ and $\\gamma_i$.","Estimated linearization parameters are guaranteed bounded by the optimization itself, without an extra step-size or projection, for any input update gains.","Under bounded nonrepetitive disturbances and initial shifts, tracking error converges to a small neighborhood of zero that shrinks with the disturbance/initial-shift bounds; perfect tracking is recovered when these uncertainties tend to zero.","Special cases covered include first-order ILC ($m=1$), where the condition reduces to $\\lambda > \\gamma_1^2 \\bar{\\beta}_f \\beta_{\\hat\\theta}$, and time-invariant nonlinear plants."],"supporting_citations":[{"why":"Defines the iterative learning control problem of bettering an operation by repetition, the framework this paper extends to unknown nonlinear time-varying plants.","marker":"[1]"},{"why":"Supplies the dynamic-linearization-based data-driven control machinery that underlies Lemma 1's extended linearization.","marker":"[2]"},{"why":"One of the earlier optimization-based adaptive ILC designs for nonlinear non-affine systems that Algorithm 1 generalizes.","marker":"[24]"},{"why":"Provides the high-order optimization index and baseline ILC structure whose estimator the paper modifies with a boundedness penalty.","marker":"[26]"},{"why":"The higher-order optimal ILC baseline whose horizon-dependent eigenvalue-based convergence condition the paper improves on.","marker":"[27]"},{"why":"Introduces the double-dynamics analysis and the nonrepetitive-ILC lemmas (its Lemma 2) that Lemmas 5-7 and the proof of Theorem 2 rely on.","marker":"[33]"},{"why":"Extends the robust double-dynamics treatment to nonrepetitive references and initial shifts, providing the template for Theorem 3.","marker":"[34]"},{"why":"The mean value theorem reference used in the proofs of Lemmas 1 and 9 to derive the extended linearizations.","marker":"[36]"},{"why":"Supplies the nonnegative-matrix properties used in Lemma 10 to turn the scalar contraction condition into the product-matrix condition.","marker":"[37]"}],"fun_headline_variants":["No-model ILC: perfect tracking for unknown nonlinear time-varying plants","Data-only ILC achieves perfect tracking for nonlinear time-varying systems","Optimization-based ILC: perfect tracking with zero model knowledge","Perfect tracking from data only: no plant model needed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument depends on Assumption (A1): the unknown plant is smooth everywhere, all its partial derivatives are globally bounded, and its input-output coupling derivative is always positive and bounded away from zero.","fun_headline_variants_meta":{"raw":{"variants":["No-model ILC: perfect tracking for unknown nonlinear time-varying plants","Data-only ILC achieves perfect tracking for nonlinear time-varying systems","Optimization-based ILC: perfect tracking with zero model knowledge","Perfect tracking from data only: no plant model needed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000963,"raw_usage":{"total_tokens":4087,"prompt_tokens":922,"completion_tokens":3165,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":3095}},"tokens_in":538,"tokens_out":3165,"duration_ms":20589,"temperature":1.0,"reasoning_tokens":3095,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:44:36.489605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Algorithm 1, with gains satisfying (41), on a smooth nonlinear plant that obeys (A1) and see whether the tracking error fails to converge to zero or any estimated parameter exceeds the predicted bound $\\beta_{\\hat\\theta}$; a single such plant would refute Theorem 2. A sharper test: replace the input-output coupling term $t/(t+2)u_k(t)$ in the paper's simulation with a smooth function that changes sign at some time step while keeping the function smooth; if the tracking error still converges, then Assumption (A1) is not necessary, and if it diverges, the sign-fixity premise is confirmed as load-bearing.","supporting_citations":[{"cited_title":"Bettering ope ration of robots by learning,","cited_arxiv_id":null,"evidence_quote":"Defines the iterative learning control problem of bettering an operation by repetition, the framework this paper extends to unknown nonlinear time-varying plants."},{"cited_title":"An overview of dynamic-linear ization-based data-driven control and applications,","cited_arxiv_id":null,"evidence_quote":"Supplies the dynamic-linearization-based data-driven control machinery that underlies Lemma 1's extended linearization."},{"cited_title":"Dual-stage optimal iterative learning control for nonlinear non-afﬁne discrete-time sy stems,","cited_arxiv_id":null,"evidence_quote":"One of the earlier optimization-based adaptive ILC designs for nonlinear non-affine systems that Algorithm 1 generalizes."},{"cited_title":"A uniﬁed data-drive n design framework of optimality-based generalized iterat ive learning control,","cited_arxiv_id":null,"evidence_quote":"Provides the high-order optimization index and baseline ILC structure whose estimator the paper modifies with a boundedness penalty."},{"cited_title":"Computationally efﬁ cient data-driven higher order optimal iterative learning control,","cited_arxiv_id":null,"evidence_quote":"The higher-order optimal ILC baseline whose horizon-dependent eigenvalue-based convergence condition the paper improves on."},{"cited_title":"Robust iterative learning cont rol for nonrepetitive uncertain systems,","cited_arxiv_id":null,"evidence_quote":"Introduces the double-dynamics analysis and the nonrepetitive-ILC lemmas (its Lemma 2) that Lemmas 5-7 and the proof of Theorem 2 rely on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The mean value theorem reference used in the proofs of Lemmas 1 and 9 to derive the extended linearizations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the nonnegative-matrix properties used in Lemma 10 to turn the scalar contraction condition into the product-matrix condition."}],"review_version":1}