{"id":"de88c960-2eb2-4ce1-8fbb-fe8634cd5dde","arxiv_id":"2506.15465","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DATA-DRIVEN PRONTO iteratively estimates local linearized dynamics from perturbed closed-loop experiments, solves an LQR subproblem with those estimates, and provably converges to a neighborhood of the optimal solution whose size shrinks with the exploration dither amplitude.","lead":"This paper introduces a data-driven version of the PRONTO optimal control method that solves nonlinear control problems without a model of the system's dynamics, using repeated experiments with small random perturbations. It proves that the algorithm converges close to the true optimum, with the final error shrinking as the perturbations shrink, and demonstrates the result on a swing-up pendulum robot.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 4 is the load-bearing premise: the proof needs a uniform error bound over all well-conditioned dithers down to zero, which the stated compactness argument does not deliver, and the pendubot experiment never verifies (22).","rationale":"The reader's weakest assumption identifies Assumption 4, and my reading agrees that this is the central load-bearing premise. The theorem is explicitly conditional on Assumption 4, and the entire practical-stability argument reduces the problem to a bounded descent-direction error whose size is controlled by the dither amplitude exactly through the well-conditioned identification bound. The proof as written has a genuine uniformity gap: the constants in (42), (37), and (46) are obtained on compact subsets of a set that excludes zero and uses strict inequality, while the theorem's quantification allows dithers arbitrarily close to zero and condition (22) with non-strict κ≤M. Thus the claim that b(δ_x,δ_u)→0 is not fully justified without an additional argument showing the relevant Lipschitz constants do not blow up as the data batches scale down. This is a correctness risk in the theorem's proof, not merely a missing verification. The numerical experiment is qualitatively consistent with the intended behavior, and no contradiction with the theorem was found, so the appropriate disposition remains conditional rather than rejection. The reader's CONDITIONAL verdict therefore stands; my concern refines and reinforces it but does not change it.","tokens_in":1877,"tokens_out":2693,"duration_ms":176151,"concrete_test":"Instrument the Section V pendubot setup to compute, at every t in every iteration and for all 15 dither-amplitude runs, the condition number κ([ΔX_t;ΔU_t][ΔX_t;ΔU_t]^T), the identification error norm ∥(ΔA,ΔB)∥, and the ratio ∥(ΔA,ΔB)∥/∥(d_x,d_u)∥. If the maximum κ and this ratio remain bounded independent of the run index as δ is halved, the uniformity needed in (37)/(46) holds for the example, weakening the concern. If κ exceeds the M used in the proof or the ratio grows as δ→0, the experiment fails to instantiate Theorem 1, and the paper should supply a sufficient condition or a verification procedure for Assumption 4 before the theorem can be regarded as operational.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1's conclusion (23) depends on the error-propagation chain: Lemma 3 maps dithers to data with constant r in (42), Lemma 5 maps data to identification error with constant g in (37), and Lemma 6 maps identification error to descent-error with constant p in (46). The composite bound (50), δ_ζ = pgrT(δ_u+δ_x), is what makes the ultimate ball b(δ_x,δ_u) shrink to zero. Every one of these constants is obtained by a compactness argument, but the admissible set in (45) is not compact: Lemma 4 defines F_M by strict κ<M; Lemma 5 quantifies over compact K⊂F_M; yet Assumption 4 and Theorem 1 allow κ≤M and allow dithers of norm arbitrarily close to zero, whose data matrices accumulate at 0∉F_M. Nothing in the proof shows that p, g, r can be chosen uniformly over these noncompact domains. This is more than a technical detail: without a scale-invariant bound on κ and a uniform Lipschitz constant as δ→0, the vanishing-error conclusion of Theorem 1 is not established by the written argument. Remark 3 concedes that Assumption 4 is only well characterized for LTI systems, and the Section V experiment draws d_u∼U(0,δ_u), d_x∼U(−δ_x,δ_x) but never reports κ of the batches (22), so the claimed corroboration does not demonstrate that the theorem's hypotheses were satisfied.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes DATA-DRIVEN PRONTO, an extension of the model-based PRONTO algorithm for finite-horizon nonlinear optimal control when the dynamics are unknown but experiments can be performed from a fixed initial condition. At each iteration, L closed-loop trajectories are collected by adding small dithers around the current nominal trajectory through a tracking controller; the resulting data batches are used to least-squares identify time-varying linearizations (16), which replace the exact Jacobians in the LQR subproblem (17). The main result, Theorem 1, claims that under smoothness assumptions (Assumptions 1-3) and a well-posed identification assumption (Assumption 4), the iterates are locally uniformly ultimately bounded around an isolated optimum, with an ultimate bound b(δ_x,δ_u) that is strictly increasing in the dither amplitudes and vanishes as they tend to zero. The proof proceeds by viewing the algorithm as a perturbed PRONTO iteration and bounding the descent-direction error through a chain of lemmas. A pendubot swing-up example compares the data-driven algorithm with model-based PRONTO and shows that the distance to the optimum decreases as the dither amplitude is halved.","tokens_in":21277,"tokens_out":11731,"duration_ms":126448,"significance":"If the main theorem and its proof can be repaired, this is a worthwhile contribution to data-driven numerical optimal control. The paper gives a clean algorithmic template, a self-contained proof structure, and a nontrivial numerical demonstration on an underactuated mechanical system. The central idea of estimating trajectory-dependent linearizations from closed-loop dither experiments, rather than fitting a global model, is attractive and is clearly contrasted with iterative learning control and derivative-free optimization. The claimed result is falsifiable: the numerical experiment tests the qualitative prediction that suboptimality decreases with dither amplitude, and no fitted constants enter the statement of Theorem 1. However, as detailed below, the proof currently has a load-bearing gap in the uniformity argument, and the numerical section does not verify the theorem's standing assumptions; both need to be addressed before the conclusions can be accepted.","major_comments":[{"comment":"The proof's uniformity step is not valid because the admissible set H×D' defined in (45) is not compact. The set Δ^{-1}_{XU}(F^T_M) is open since F_M is open by Lemma 4, and the balls B_{σ'}(η*) and B_{δ'_{xu}} are open; moreover, under Assumption 4 this set contains dithers of arbitrarily small norm, whose data matrices accumulate at Δ=0, while 0∉F^T_M. Lemma 5 only supplies a constant g(H,K) for compact K⊂F^T_M, and Lemma 6's δ_AB(η) is used only on a compact neighborhood of η*. Consequently, the constants r, g, p in (42), (43), (46), and (50) are not shown to exist uniformly over all dithers satisfying (22), and the vanishing ultimate bound b(δ_x,δ_u) in (23) is not established by the written argument. This is the load-bearing step of Theorem 1 and must be repaired, for example by a separate small-dither argument exploiting the conical structure of F_M and the bound ∥Δt^†∥=1/σ_min(Δt) with κ(ΔtΔt^T)≤M.","section":"Section IV-C, Eqs. (41)-(50)"},{"comment":"The identity o_η(Δt)Δt^T(ΔtΔt^T)^{-1}=o_η(ΔtΔt^T)(ΔtΔt^T)^{-1} is dimensionally inconsistent: o_η(Δt) is an n×L matrix while ΔtΔt^T is (n+m)×(n+m), and the little-o of a product is not defined by this expression. The desired estimate can be obtained directly from (83) using ∥Δt^†∥=1/σ_min(Δt) and the conditioning condition κ(ΔtΔt^T)≤M, but the derivation as written is invalid. This matters because Lemma 5's bound (37) feeds directly into the descent-error bound (46) and hence into the final bound (50).","section":"Appendix VII-E, Eqs. (86)-(90)"},{"comment":"The numerical experiment does not verify the hypotheses of Theorem 1. Assumption 4 requires that for every dither bound there exist L dither sequences such that (22) holds; the experiment draws d_u∼U(0,δ_u) and d_x∼U(−δ_x,δ_x) but does not report the condition numbers of [ΔX_t;ΔU_t][ΔX_t;ΔU_t]^T. In addition, Section V-A states that the controller π is redesigned at each iteration from the estimated linearizations, so π is not the fixed C^2 map required by Assumption 3 and the projection operator P changes between iterations, whereas the stability analysis in Lemmas 1-2 treats P as fixed. Thus the experiment demonstrates a plausible qualitative trend but does not instantiate the theorem's setting.","section":"Section V, Steps L1 and numerical setup"}],"minor_comments":[{"comment":"The inequality ∥(d^k_x,d^k_u)∥≤T(δ_u+δ_x) is not correct for L>1. From the definition in (29), the stack has nTL+mTL entries, so its Euclidean norm is bounded by √(TL)(√n δ_x+√m δ_u), not by T(δ_u+δ_x). The subsequent choice (51) can be repaired by replacing the factor T with a dimension-dependent constant, so this is not fatal, but the displayed bound should be corrected.","section":"Section IV-C, Eq. (48)"},{"comment":"The notation for open versus closed balls is confusing: the text defines B_p(¯x) as the open ball but then says 'we use B_p(¯x) when the ball is closed.' This ambiguity directly affects the compactness claims in (40), (41), and (45), where open balls are intersected with open preimages and declared compact.","section":"Notation, Section I"},{"comment":"There are several typographical issues: 'demostrate' in Section V, 'novanishing' in reference [42], and the repeated 'subj.to' formatting in the optimization problems. These should be corrected in a final revision.","section":"Throughout"},{"comment":"Assumption 4 is very strong: it postulates well-conditioned identification for arbitrarily small dither amplitudes, not merely for sufficiently rich excitation, and Remark 3 concedes that it is only well characterized for linear time-invariant systems. The paper would be strengthened by a discussion of how the condition might be checked online or relaxed for nonlinear systems, even if only in a local approximate sense.","section":"Section III-A, Assumption 4 and Remark 3"}],"recommendation":"major_revision","confidential_remarks":"The main theorem is plausible and the algorithmic idea is valuable, but the proof as written does not close the small-dither uniformity argument, and the numerical section does not instantiate the theorem's assumptions. If the authors can supply a correct uniformity argument, likely using the conical structure of F_M and precise little-o estimates rather than compactness over noncompact sets, the paper would be a solid contribution. The mismatch between the adaptive controller used in the experiments and the fixed-controller assumption in the theory should also be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good to have a concrete extension of PRONTO to the model-free setting. The authors combine PRONTO's projection-operator descent with closed-loop dither exploration and per-trajectory least-squares LTV identification, then prove a local practical-stability theorem: with dither bounds small enough and a conditioning assumption on the data, iterates stay eventually in a ball about the isolated optimum whose radius vanishes as the dither amplitude goes to zero. That theorem is the main new thing, and it is not dressed up. The algorithm is sensible, the closed-loop tracking controller for exploration is a nice practical touch, and no fitted constants enter the theorem. The pendubot experiment matches the qualitative prediction: smaller dithers give smaller suboptimality, without tuning the bound to the data. Citation pattern is fair: [31] is accurately placed as the closest comparison, and the ILC/RL overview is not padded with self-citations.\n\nThe soft spots are in the proof, and they are real. Assumption 4 is load-bearing: it postulates well-conditioned data batches for all dither bounds down to zero. Remark 3 concedes the condition is only well characterized for LTI systems, and the experiment never checks it. More seriously, the compactness chain does not close over the dither-to-zero limit. Lemma 4 defines F_M with strict κ<M as an open cone; Lemma 5 quantifies over a compact K⊂F_M. But Assumption 4 and Theorem 1 only require κ≤M, and they let dithers shrink to zero, where the data matrices accumulate at zero—which is not in F_M. So the uniform constants g, r, p in Lemmas 3–6 are not obtained on the domain actually used in Theorem 1. Equation (86) also hand-waves the little-o: it rewrites o_η(Δt)Δt^T(ΔtΔt^T)^{-1} as o_η(ΔtΔt^T)(ΔtΔt^T)^{-1}, absorbing the column factor into a scalar little-o. That is not legitimate on its face, and the linear bound in (91) depends on it.\n\nNone of this destroys the paper. The theorem is probably repairable—one would need a scale-invariant persistency-of-excitation lower bound and a uniform Lipschitz constant that survives as δ→0. But as written the vanishing-error conclusion is not established, and Assumption 4 remains a hypothesis rather than a verified condition. A referee should push on exactly this.\n\nVerdict: send it to peer review. It is a serious, honest contribution that extends a known method to a harder setting, and the gap is specific enough to be fixable. The authors should also report condition numbers or a persistence-of-excitation check in the experiments.","headline":"A useful model-free extension of PRONTO with a serious uniformity gap in the proof, but absolutely worth a real referee.","tokens_in":21805,"tokens_out":3571,"would_cite":true,"duration_ms":36916,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DATA-DRIVEN PRONTO solves nonlinear optimal control with no model by probing trajectories and estimating local linearizations from data, then proves the iterates converge to a neighbourhood of the optimum whose radius shrinks with the…","keywords":["data-driven optimal control","model-free optimal control","PRONTO","projection operator","trajectory linearization","least-squares identification","ultimate boundedness","iterative learning control"],"falsifier":"Run the pendubot experiment with the uniform dithers used in the paper and, at each iteration, compute the condition number of $[\\Delta X_t;\\Delta U_t][\\Delta X_t;\\Delta U_t]^\\top$ for each $t$; if convergence to the claimed ball persists even when the data batches never satisfy the bound $M$, then Assumption 4 is not necessary for the observed behaviour. Conversely, exhibit any system, initial trajectory, and dither schedule satisfying Assumption 4 whose iterates leave the ball of radius $b(\\delta_x,\\delta_u)$ infinitely often, which would refute Theorem 1.","tokens_in":20717,"feed_emoji":"🎛️","tokens_out":7607,"duration_ms":75612,"temperature":0.7,"pith_summary":"This paper claims that a finite-horizon optimal control problem for an unknown nonlinear dynamical system can be solved without a model, by repeatedly probing the system and estimating local linearizations from data. The proposed algorithm, DATA-DRIVEN PRONTO, wraps the projection-operator method PRONTO: at each iteration it perturbs the current trajectory with small dither signals in closed loop, fits time-varying affine dynamics to the measured responses, and solves an LQR subproblem with those estimated dynamics to get a descent direction. The main theorem states that if the system and cost are twice differentiable, the cost Hessian is positive definite, and an identifiability condition holds near the optimum, then for small enough dithers and a sufficiently good initial trajectory the iterates are uniformly ultimately bounded in a ball around an isolated optimum, with the ultimate radius a strictly increasing function of the dither amplitudes that goes to zero as the dithers shrink. A careful reader would care because this turns a model-free repeated-task setting, usually handled by iterative learning control or ad-hoc exploration, into an optimization algorithm with explicit local convergence guarantees and a tunable accuracy-exploration trade-off.","feed_headline":"Probing trajectories solves optimal control with no model","feed_subtitle":"DATA-DRIVEN PRONTO estimates local linearizations from dither experiments; the optimum is approached as exploration shrinks.","key_machinery":"The load-bearing mechanism is the projection operator $P$ from PRONTO, implemented by a $C^2$ tracking controller $\\pi(\\alpha_t,\\mu_t,x_t,t)$ that satisfies $\\pi(\\alpha,\\mu,\\alpha,t)=\\mu$ and therefore leaves true trajectories unchanged; it maps any tentative curve onto the trajectory manifold $\\mathcal{T}$. Around this operator, the algorithm builds three smooth maps: $\\Delta XU$ (closed-loop dither response, zero at zero dithers), $\\Delta AB$ (Jacobian estimation error, linearly bounded on the well-conditioned set $F_M$), and $\\Delta\\zeta$ (descent error, zero for exact Jacobians). Their composition $\\Delta\\zeta(\\eta,\\Delta AB(\\eta,\\Delta XU(\\eta,d_x,d_u)))$ is what turns dither amplitude into a bounded perturbation of the known-stable PRONTO update, and Assumption 4—the existence of well-conditioned dither batches near the optimum—is what keeps that composition defined. The named conclusion is Theorem 1, with the least-squares estimator of the time-varying linearizations as the concrete identification step.","core_discovery":"On the paper's own terms, the discovery is that the model-based PRONTO descent step can be replaced by a data-driven one whose error vanishes smoothly with the exploration amplitude. The proof identifies the algorithm as a perturbed version of the projection-operator update: with $\\eta_k=(x_k,u_k)$ the current trajectory, the true descent direction solves the LQR subproblem using the exact Jacobians $A_t=\\nabla_1 f^\\top$, $B_t=\\nabla_2 f^\\top$, while DATA-DRIVEN PRONTO solves the same subproblem with the least-squares estimates $\\hat A_t,\\hat B_t$ obtained from dither experiments. Lemma 5 bounds the Jacobian estimation error linearly by the size of the data batches when the batches are well conditioned, and Lemma 6 shows the resulting descent error $\\hat\\zeta-\\zeta$ is a $C^1$ function that is zero when the Jacobians are exact; composing these bounds gives $\\|\\hat\\zeta-\\zeta\\|\\le pgr\\|(d_x,d_u)\\|$. Theorem 1 then follows from a practical-stability result for discrete-time systems under bounded nonvanishing perturbations: starting near an isolated optimum, the iterates satisfy $\\|\\eta_k-\\eta^\\star\\|\\le b(\\delta_x,\\delta_u)$ for all $k\\ge N$, with $b$ strictly increasing and $b(0,0)=0$.","pith_inferences":["A practical way to test whether Assumption 4 is necessary is to record the condition number of $[\\Delta X_t;\\Delta U_t][\\Delta X_t;\\Delta U_t]^\\top$ during the pendubot runs; the paper's experiment uses uniform dithers without checking it, so persistent convergence despite violations would show the assumption is sufficient but not necessary.","The same error-composition argument suggests a data-driven constrained variant could be built by estimating constraint Jacobians alongside the dynamics, extending the result beyond unconstrained stage costs.","When the tracking controller itself is learned from the same estimated linearizations, as in the numerical example, the exploration policy and the descent direction are coupled; the theory does not explicitly model that coupling, so controller quality could affect the effective conditioning more than the dither size alone."],"forward_implications":["Any task that can be repeated from the same initial condition becomes solvable without a model: run at least $n+m$ perturbed closed-loop experiments per iteration, fit local linearizations, and iterate.","The ultimate suboptimality is user-tunable: shrinking the dither bounds $\\delta_x,\\delta_u$ drives the asymptotic error bound $b(\\delta_x,\\delta_u)$ to zero, so accuracy is limited only by how well-conditioned tiny explorations remain.","Because only local first-order approximations are identified, systematic global model-parametrization errors are avoided; the method is model-free by construction and can use partial model knowledge only to improve the tracking controller.","The analysis unifies iterative learning control with numerical optimization, since the closed-loop exploration is guided by a tracking controller and the result is a descent method on the trajectory manifold."],"supporting_citations":[{"why":"Supplies the projection-operator PRONTO algorithm that the data-driven version extends.","marker":"[35]"},{"why":"Characterizes the trajectory manifold and the projection operator used to define the unconstrained problem.","marker":"[38]"},{"why":"Provides the stable approximate-Newton constrained optimization result used in Lemma 1 for PRONTO's local stability.","marker":"[1]"},{"why":"Provides the practical-stability theorem under bounded nonvanishing perturbations used for the ultimate-boundedness conclusion.","marker":"[42]"},{"why":"Establishes persistency of excitation, the LTI counterpart the paper cites for well-posed identification.","marker":"[36]"},{"why":"Gives sufficient-richness conditions for LTI systems that support the characterization of Assumption 4.","marker":"[37]"},{"why":"Provides the pendubot model used in the numerical swing-up experiment.","marker":"[39]"}],"fun_headline_variants":["Model-free PRONTO: dither trajectories reveal descent","Data-driven PRONTO converges via perturbed trajectories","PRONTO without a model: closed-loop experiments do the job","Dither-driven PRONTO steers optimal control without a model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that near the optimum, arbitrarily small exploration dithers can always be chosen so that the resulting data batches are well-conditioned at every time step, a condition the paper characterizes rigorously only for linear time-invariant systems and does not verify numerically.","fun_headline_variants_meta":{"raw":{"variants":["Model-free PRONTO: dither trajectories reveal descent","Data-driven PRONTO converges via perturbed trajectories","PRONTO without a model: closed-loop experiments do the job","Dither-driven PRONTO steers optimal control without a model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1425,"prompt_tokens":1013,"completion_tokens":412,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":342}},"tokens_in":629,"tokens_out":412,"duration_ms":4741,"temperature":1.0,"reasoning_tokens":342,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:35:08.385619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pendubot experiment with the uniform dithers used in the paper and, at each iteration, compute the condition number of $[\\Delta X_t;\\Delta U_t][\\Delta X_t;\\Delta U_t]^\\top$ for each $t$; if convergence to the claimed ball persists even when the data batches never satisfy the bound $M$, then Assumption 4 is not necessary for the observed behaviour. Conversely, exhibit any system, initial trajectory, and dither schedule satisfying Assumption 4 whose iterates leave the ball of radius $b(\\delta_x,\\delta_u)$ infinitely often, which would refute Theorem 1.","supporting_citations":[{"cited_title":"A projection operator approach to the optimization of trajectory functionals,","cited_arxiv_id":null,"evidence_quote":"Supplies the projection-operator PRONTO algorithm that the data-driven version extends."},{"cited_title":"The trajectory manifold of a nonlinear control system,","cited_arxiv_id":null,"evidence_quote":"Characterizes the trajectory manifold and the projection operator used to define the unconstrained problem."},{"cited_title":"Numerical optimal control,","cited_arxiv_id":null,"evidence_quote":"Provides the stable approximate-Newton constrained optimization result used in Lemma 1 for PRONTO's local stability."},{"cited_title":"Stability of discrete nonlinear systems under novanishing perturbations: application to a nonlinear model-matching problem,","cited_arxiv_id":null,"evidence_quote":"Provides the practical-stability theorem under bounded nonvanishing perturbations used for the ultimate-boundedness conclusion."},{"cited_title":"On Sufficient Richness for Linear Time-Invariant Systems","cited_arxiv_id":"2502.04062","evidence_quote":"Gives sufficient-richness conditions for LTI systems that support the characterization of Assumption 4."},{"cited_title":"Siciliano, L","cited_arxiv_id":null,"evidence_quote":"Provides the pendubot model used in the numerical swing-up experiment."}],"review_version":2}