{"id":"5c9a57ce-1478-46ec-8116-192646536dee","arxiv_id":"2607.04006","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Finite-sample MPPI on unconstrained LTI/quadratic systems is a high-probability perturbation of LQR and is practically exponentially stable in expectation above an explicit sample threshold.","lead":"This paper proves that sampling-based MPPI control stays practically stable on linear systems when you take enough samples, with an explicit sample count formula. It gives robotics and control engineers a first rigorous way to choose how many MPPI samples are enough for LTI plants.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies that the load-bearing structural assumption is the DARE terminal cost plus unconstrained LTI/quadratic structure, which is scope rather than a correctness flaw. Within that scope the argument is self-contained: exact first-action coincidence with −Kx removes horizon truncation from the nominal law, the two-component error (Monte Carlo + temperature bias) is controlled by concentration and κλ → 0 as λ → 0, and the stopped-process Lyapunov bound yields the three residual floors with an explicit M*. Public code and the simulation suite further support reproducibility. No stronger internal concern (e.g., circular use of compactness, unjustified interchange of limits, or failure of the small-gain condition) is present. Therefore the ACCEPT verdict with low correctness risk stands; no adjustment is warranted.","tokens_in":17967,"tokens_out":518,"duration_ms":5522,"concrete_test":"Independently recompute the double-integrator DARE quantities in Table I (P, K, αP, α, C(0)w) and the Corollary-1 threshold M* with the paper’s C1=0.22; then re-run the noise-free median Lyapunov-ratio sweep of Experiment 2 for M ∈ {100,153,200} and verify that ˆρ(M) ≤ 0.943 precisely when M ≥ M*, confirming the certificate is consistent with the published constants.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Theorem 1 / Corollary 1 is internally coherent within the stated scope. The reduction of finite-horizon MPPI to an LQR perturbation rests on Assumption 3 (DARE terminal cost) and unconstrained LTI/quadratic structure, which the paper states explicitly (Remark 5, §II). Concentration (Lemmas 1–3), closed-form bias via completing-the-square (Proposition 1), high-probability finite-horizon invariance via the stopped supermartingale (Lemma 5), and the Lyapunov perturbation under the small-gain condition Φ(β∞) ≤ αP/2 form a non-circular chain. Bounded-control Assumption 6 keeps the bad-event residual linear in η rather than √η. Finite-horizon localization under unbounded Gaussian noise is acknowledged honestly (Remark 10). Conservatism of ρ and M* is documented in simulation rather than hidden. No derivation break or hidden circularity is apparent that would undermine the strongest claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper establishes finite-sample closed-loop practical exponential stability for Model Predictive Path Integral (MPPI) control on discrete-time LTI systems with quadratic costs and additive Gaussian process noise. Under the DARE terminal cost, the exact finite-horizon MPC first action coincides with infinite-horizon LQR for every horizon N, so finite-sample MPPI is treated as a stochastic perturbation of LQR. Lemmas 1–3 give high-probability approximation of the LQR feedback by a Monte Carlo term O(M^{-1/2}) plus an infinite-sample temperature bias characterized in closed form via completing-the-square (Proposition 1). Lemma 5 establishes high-probability finite-horizon invariance of a Lyapunov sublevel set via a stopped supermartingale, resolving the compact-set circularity. Theorem 1 then yields an unconditional Lyapunov bound on the stopped process and, on paths that remain in Ω_R over [0,T] (probability ≥1−δ), the bound E[∥x_k∥ 1_{τ_R>T}] ≤ c ρ^k ∥x_0∥ + γ_w √tr(Σ_w) + γ_M e_M(η) + γ_η √η. Corollary 1 supplies an explicit sufficient sample threshold M* computable from the DARE solution, LQR margin, temperature, and horizon; the joint limit recovers the stochastic LQR certificate. Simulations on a double integrator illustrate the certificate’s conservatism and qualitative consistency.","tokens_in":18158,"tokens_out":863,"duration_ms":9119,"significance":"Closed-loop stability of finite-sample MPPI under receding-horizon execution has been identified as an open problem; this manuscript supplies the first explicit finite-sample certificate for the unconstrained LTI/quadratic case. The reduction via DARE terminal cost is clean, the bias formula is closed-form, the sample threshold is computable rather than existential, and the three residual floors (process noise, MPPI approximation, confidence) are transparent. Simulation code is released and the conservatism of ρ and M* is documented rather than hidden. Within its stated scope the result is a solid foundation that recovers classical stochastic LQR in the appropriate limit and connects sampling-based MPC to the Mayne et al. Lyapunov framework. The restriction to unconstrained LTI/quadratic systems is a genuine limitation of scope, not a flaw in the argument as written.","major_comments":[],"minor_comments":[{"comment":"In the abstract and §I the phrase “finite-sample certificate is parametrized by the selected planning horizon” is accurate (Remark 5), but a short explicit pointer in Corollary 1 to the N-dependence of C_{X,U}, H_N and F_N would help readers who might otherwise expect horizon-independent constants.","section":null},{"comment":"Table I lists ρ(certificate bound)=(1−α/2)^{1/2}=0.9429; the same quantity appears as ρ in (36). A single consistent symbol (or a parenthetical note that the square-root form is used for the state-norm bound) would avoid a momentary notational mismatch.","section":null},{"comment":"Experiment 1 caption states that the plotted envelope omits the calibrated MPPI approximation floor. Adding a second curve that includes a representative γ_M e_M term (even if only for illustration) would make the comparison with Theorem 1 more direct.","section":null},{"comment":"Assumption 6 (bounded MPPI update) is essential for the linear-in-η bad-event residual. A one-sentence remark that actuator saturation or truncated sampling already implements this bound in every practical MPPI code base would further reassure readers.","section":null},{"comment":"The arXiv identifier of the companion nonlinear paper [15] is still “in preparation”; once available, a forward reference with the arXiv number would improve traceability.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is carefully written, the technical chain is coherent, and the authors are transparent about scope and conservatism. I see no load-bearing gaps that would justify major revision or rejection. Fit for a control/optimization journal is good; the LTI restriction is appropriate for a first certificate paper."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper actually closes a recognized gap: finite-sample closed-loop stability for MPPI, at least for unconstrained LTI/quadratic systems with DARE terminal cost. That is the punchline. Prior work had open-loop sampling rates, optimizer contraction, and free-energy bounds; this one gives a receding-horizon Lyapunov certificate with an explicit sample threshold M*.\n\nWhat is new is the reduction: under the DARE terminal cost the exact finite-horizon first action equals LQR for every horizon, so finite-M MPPI is just a stochastic perturbation of LQR. They split the error into Monte Carlo (O(M^{-1/2})) plus temperature bias (closed-form via completing-the-square, vanishes as λ→0), then run a stopped-supermartingale argument to get high-probability invariance of a Lyapunov sublevel set and practical exponential stability in expectation with three residual floors (noise, approximation, confidence). Constants are written in terms of the stacked horizon-N cost matrices, so the certificate is parametrized by N. Simulations on the double integrator, public code, and honest discussion of the 5× conservatism of M* and the loose ρ all help.\n\nSoft spots are mostly scope, not hidden breaks. Everything rides on unconstrained LTI + DARE terminal cost; drop either and the LQR coincidence disappears. Finite-horizon localization under unbounded Gaussian noise is the right statement (they say so), not a dodge. Assumptions on bounded warm-start and saturated control keep the bad-event residual linear in η; practical, but required. The math chain itself looks coherent: concentration, bias formula, invariance, small-gain Lyapunov perturbation. No circular construction. Citations are appropriate; self-cites are background or a companion nonlinear paper still in prep.\n\nThis is for people who care about sampling-based MPC theory and want a foundation case before the nonlinear/constrained extensions. It deserves a serious referee. I would engage with it, cite the LTI certificate when I need a baseline, and watch for the nonlinear follow-up.","headline":"First explicit finite-sample closed-loop stability certificate for MPPI on LTI systems; scoped tightly, math holds, useful M* formula.","tokens_in":18830,"tokens_out":510,"would_cite":true,"duration_ms":5539,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93D15","93E15","49N10"],"pacs":[],"model":"grok-4.5","headline":"Finite-sample MPPI on linear systems is a controllable perturbation of LQR, with explicit sample counts that guarantee practical exponential closed-loop stability.","keywords":["MPPI","model predictive control","finite-sample stability","LQR","Lyapunov perturbation","sampling-based MPC","DARE terminal cost","practical exponential stability"],"falsifier":"On the double-integrator benchmark (or any LTI plant meeting the assumptions), compute the analytical M-star from the paper’s Corollary 1; if for every M at or above that threshold the empirical median Lyapunov ratio remains strictly larger than the certified decay rate while process noise and temperature stay at the paper’s values, the practical-stability claim is false.","tokens_in":18805,"feed_emoji":"⚙️","tokens_out":1136,"duration_ms":18586,"temperature":0.7,"pith_summary":"This paper gives the first finite-sample closed-loop stability guarantees for Model Predictive Path Integral (MPPI) control on discrete-time linear systems with additive Gaussian noise. The key observation is that, with the DARE terminal cost, the exact finite-horizon optimal first action equals the infinite-horizon LQR law for every planning horizon, so finite-sample MPPI can be treated as a stochastic perturbation of classical LQR. The authors prove that the control error splits into a Monte Carlo term that shrinks like one over square-root of the sample count and a temperature bias that vanishes as temperature goes to zero; once both are small enough relative to the LQR Lyapunov margin, the closed-loop state decays exponentially in expectation on high-probability paths that stay inside a compact sublevel set over any finite operating horizon. Three residual floors remain—process noise, MPPI approximation, and per-step sampling failure—and the sufficient sample threshold is written explicitly in terms of the DARE solution, the LQR margin, temperature, and horizon. A sympathetic reader cares because the result finally answers, with a computable formula, how many MPPI samples are enough for certified stability, and recovers the classical stochastic LQR bound in the infinite-sample, zero-temperature limit.","feed_headline":"How many MPPI samples guarantee LQR-like stability","feed_subtitle":"An explicit threshold turns sampling error into three residual floors and a computable certificate.","key_machinery":"The exact coincidence of the first finite-horizon optimal action with the infinite-horizon LQR gain for every planning horizon when the terminal cost is the DARE solution. That identity lets the authors view MPPI as a stochastic perturbation of LQR, decompose the error into Monte Carlo concentration plus closed-form temperature bias (via Gaussian completing-the-square), and absorb both into the classical LQR Lyapunov decrease through a stopped-process supermartingale argument.","core_discovery":"For unconstrained LTI systems with quadratic costs and DARE terminal cost, finite-sample MPPI approximates the LQR feedback with high probability: the error decomposes into a Monte Carlo term of order M to the minus one-half and an infinite-sample temperature bias that vanishes as temperature tends to zero. Under a small-gain condition on that bias, a Lyapunov perturbation argument yields practical exponential stability in expectation on sample paths that remain in a compact Lyapunov sublevel set over a finite horizon, with three explicit residual floors and an explicit sufficient sample threshold M-star computable from the DARE solution and LQR stability margin.","pith_inferences":["The same first-action coincidence may serve as a template for sample-complexity certificates of other sampling-based MPC methods that still lack closed-loop guarantees.","Bounded or truncated process noise would lift the finite-horizon localization to infinite-horizon high-probability invariance, exactly as the paper’s own future-work section anticipates.","Because the certificate is deliberately conservative, tighter problem-specific concentration constants could convert the sufficient M-star into a sharper design knob without changing the controller architecture.","The closed-form temperature-bias gain suggests an adaptive schedule that lowers temperature once samples have concentrated, shrinking the residual floor at no extra online cost."],"forward_implications":["With sample count above an explicit threshold built from the DARE solution and LQR margin, the closed loop is practically exponentially stable in expectation on high-probability finite-horizon paths.","In the joint limit of infinite samples and vanishing temperature the bound recovers the classical stochastic LQR stability certificate.","The sample threshold and residual floors are parametrized by planning horizon, temperature, and sampling covariance, so designers can trade samples against temperature and horizon.","The same bound admits a practical input-to-state stability reading with three explicit gains for process noise, MPPI approximation, and confidence loss.","On the double-integrator the analytical threshold is numerically computable and qualitatively marks the onset of certified decay rates."],"fun_headline_variants":["Explicit sample threshold for MPPI closed-loop LQR stability","Finite-sample MPPI approximates LQR feedback with high probability","Three residual floors certify practical exponential stability of MPPI","How sampling and temperature set MPPI stability margins vs LQR","Computable M-star for MPPI as stochastic perturbation of LQR"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The whole argument rests on the system being unconstrained, linear, and quadratic with the exact infinite-horizon cost-to-go used as the terminal cost, so that every finite-horizon plan has the same first move as classical LQR.","fun_headline_variants_meta":{"raw":{"variants":["Explicit sample threshold for MPPI closed-loop LQR stability","Finite-sample MPPI approximates LQR feedback with high probability","Three residual floors certify practical exponential stability of MPPI","How sampling and temperature set MPPI stability margins vs LQR","Computable M-star for MPPI as stochastic perturbation of LQR"]},"model":"grok-4.5","effort":"low","cost_usd":0.00311,"raw_usage":{"total_tokens":1175,"prompt_tokens":889,"num_sources_used":0,"completion_tokens":88,"cost_in_usd_ticks":31100000,"prompt_tokens_details":{"text_tokens":889,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":198,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":889,"tokens_out":88,"duration_ms":2454,"temperature":1.0,"reasoning_tokens":198,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T22:21:50.777427+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the double-integrator benchmark (or any LTI plant meeting the assumptions), compute the analytical M-star from the paper’s Corollary 1; if for every M at or above that threshold the empirical median Lyapunov ratio remains strictly larger than the certified decay rate while process noise and temperature stay at the paper’s values, the practical-stability claim is false.","supporting_citations":[],"review_version":1}