{"id":"d6d73eee-7d5b-4648-abe3-1fa31b6d7d72","arxiv_id":"1908.01404","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"OPmin, a modified optimistic planning algorithm, provably computes near-optimal inputs and, used in receding horizon, stabilizes nonlinear switched discrete-time systems under stabilizability and detectability assumptions.","lead":"An artificial-intelligence planning algorithm, optimistic planning, is adapted into OPmin for computing near-optimal switching signals in nonlinear discrete-time control systems, with formal stability guarantees. The paper matters because it offers a path to near-optimal, stable control for switched nonlinear systems where standard linear-quadratic methods may not apply.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central near-optimality and stability theorems are not proved in the manuscript: Theorem 1 and Corollary 2 are explicitly deferred to the unpublished companion [6], and Theorem 2's appendix imports [6]. The paper's main claims therefore hinge on an external proof that readers cannot check.","rationale":"Read in good faith, the paper has a coherent algorithmic idea: OPmin is a minimization variant of optimistic planning, Proposition 1 and Proposition 2 are self-contained and plausible, and the stability framework via Assumption 1 follows standard MPC dissipation arguments. The numerical example illustrates rather than proves the general claims. The single most load-bearing weakness is that the two headline results — the near-optimality guarantee and the global exponential stability guarantee — are not proved in the manuscript; they are imported from [6], which is not publicly available and is itself only submitted. The appendix to Theorem 2 is a sketch that explicitly depends on [6, Theorem 2] and [6, Remark 3], so it does not cure the problem. If [6] is correct and has exactly the hypotheses used here, the central claims may well hold; if [6] is unavailable, incorrect, or uses different assumptions, the main contributions of this paper are unsupported. This is not an accusation of misconduct; it is a structural gap in the argument as presented. The reader's CONDITIONAL verdict is therefore appropriate: the paper should not be accepted until the companion results are public or the derivations are included in full. I would not move the verdict to REJECT, because the cited companion may be sound and the framework is plausible; I also would not move it to ACCEPT. The existing conditional status stands unchanged.","tokens_in":15690,"tokens_out":14411,"duration_ms":147086,"concrete_test":"Obtain [6] and independently re-derive inequality (8) from Assumption 1 without invoking [6, Theorem 3] as a black box. The check succeeds only if the derivation holds for every gamma in (0,1], all nonnegative stage costs (not restricted to [0,1]), and arbitrary state-dependent horizon d(x); if the derivation requires an extra hypothesis such as fixed horizon, bounded costs, or additional lower/upper bounds on W, then Theorem 1 and Corollary 2 are not established as stated. Equivalently, if [6] cannot be retrieved, the manuscript fails this check because the central proofs are not publicly verifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is not the OPmin mechanism itself but the support for the quantitative claims. Theorem 1, which gives the near-optimality inequality (8), is stated with \"Theorem 1 follows from [6, Theorem 3], therefore the proof is omitted.\" Corollary 2, the global exponential stability result, says the proof is omitted and is said to follow from [6, Corollary 2]. Theorem 2's appendix states that it follows the proof of [6, Theorem 2], with only a sketch of the modification via Proposition 3. Thus every central result rests on an unpublished, non-archival companion [6], described only as \"submitted for journal publication\" and available through a reviewer URL. This is not a cosmetic citation: the near-optimality error bound (8), the exponential-stability condition (12), and the running-cost bound in Theorem 3 all inherit the correctness and hypotheses of [6]. The text does not let a reader verify whether [6] assumes, for example, a fixed horizon, bounded stage costs, or a different Lyapunov function; any such mismatch would invalidate the transfer to OPmin's state-dependent horizons and generic nonnegative costs. Since the paper's advertised contribution is precisely to remove the [0,1]-stage-cost restriction and to handle gamma = 1 and state-dependent horizons, the exact difference between [6]'s setting and this one is where a full derivation is needed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes OPmin, a modification of the optimistic planning (OP) algorithm of Hren and Munos, adapted to the minimization of discounted or undiscounted infinite-horizon costs for nonlinear switched discrete-time systems with finite input sets and arbitrary nonnegative stage costs. The algorithm is described in Algorithm 1, and the authors state that it computes, for a given budget, the exact finite-horizon value function V_{γ,d(x)}(x) for a state-dependent horizon d(x). Under a standing assumption on existence of optimal infinite-horizon sequences and under Assumption 1, which combines a stabilizability bound on V_{γ,∞} and a detectability/dissipation inequality involving a Lyapunov-like function W, the paper claims: Theorem 1, a two-sided bound between V_{γ,d(x)} and V_{γ,∞} with an explicit state- and horizon-dependent error v_{γ,d(x)}; Corollary 1, a linear-growth version of the bound that decays exponentially in d(x); Theorem 2, semiglobal practical stability of the receding-horizon closed-loop system; Corollary 2, global exponential stability under strengthened linear assumptions; Theorem 3 and Corollary 3, bounds on the running cost relative to V_{γ,∞}. A numerical example with the cubic integrator is reported. The main proofs of Theorem 1 and Corollary 2 are omitted and deferred to the authors' unpublished companion manuscript [6], and the appendix proof of Theorem 2 and Proposition 3 also import key steps from [6].","tokens_in":15954,"tokens_out":5581,"duration_ms":58794,"significance":"If the results are correct, the paper makes a useful contribution by adapting optimistic planning to a control-oriented setting: it removes the stage-cost normalization to [0,1], allows the undiscounted case γ=1, treats cost minimization directly, and provides stability guarantees for a receding-horizon implementation, which are absent from the original OP analysis. The state-dependent horizon appearing both in the algorithm and in the error bounds is a genuine conceptual step beyond the fixed-horizon analysis in the authors' companion work. The result relating the running cost to the infinite-horizon value function, with exponential decay in the horizon, is also valuable. However, the significance is conditional: each of these claims depends on proofs that are not present in this manuscript and are instead deferred to an unpublished companion paper. The reader cannot verify that the companion's assumptions, which may involve fixed horizons and a specific discounted setting, transfer to the state-dependent horizons, γ=1, and generic nonnegative costs treated here.","major_comments":[{"comment":"The central near-optimality inequality is not proved in this manuscript; the text states that 'Theorem 1 follows from [6, Theorem 3], therefore the proof is omitted.' Reference [6] is described only as 'submitted for journal publication' and is not an archival, publicly citable source. A reader therefore cannot check whether [6] covers γ=1, nonnegative stage costs outside [0,1], and state-dependent horizons, all of which are explicit claimed improvements over OP. Please provide a self-contained proof of Theorem 1, or make the companion paper publicly verifiable and show explicitly that its assumptions and proof steps extend to OPmin with state-dependent horizons.","section":"Section IV-B, Theorem 1 and Eq. (8)"},{"comment":"The global exponential stability result is stated with 'the proof is omitted' and with justification that it follows from [6, Corollary 2] plus 'the modifications given in the appendix.' Because the receding-horizon implementation uses a state-dependent horizon d(x), uniformity of the Lyapunov decay over all horizon sequences generated by OPmin is a load-bearing issue. The full proof must appear in this paper, including verification that the Lyapunov decrease in Proposition 3 holds uniformly over the possible state-dependent horizons.","section":"Section IV-C, Corollary 2 and Eq. (12)"},{"comment":"The appendix proof of Theorem 2 is not self-contained: it explicitly imports 'item (i) of Theorem 1 in [6]', 'the same manipulations as in the proof of Theorem 1 in [6]', and 'the steps of [6]'. Proposition 3, which is the key Lyapunov property used for Theorem 2, is therefore not verifiable from the manuscript alone. Please provide a complete proof of Proposition 3 and of Theorem 2 that does not rely on unpublished companion results.","section":"Appendix, proof of Theorem 2 and Proposition 3"}],"minor_comments":[{"comment":"The definitions of the class-K functions in the error term are garbled: the text reads 'αY = αY := αW, αY := αV + αW', which defines αY twice and leaves the intended upper and lower functions unclear. Please introduce αY and αY with distinct, consistent definitions.","section":"Section IV-B, Theorem 1"},{"comment":"In the paragraph defining leaves, there is a typo: 'the the set of all leaves' should be 'the set of all leaves'.","section":"Section III-B, Algorithm 1"},{"comment":"The verb 'chosing' should be 'choosing'.","section":"Remark 1"},{"comment":"Table I reports 'estimated running cost' computed over 200 simulation steps, whereas Theorem 3 concerns the infinite-horizon running cost. A sentence on the truncation error or on why 200 steps is sufficient would help the reader interpret the numerical comparison.","section":"Section V, Table I"},{"comment":"The remark notes that the error bound v_{γ,d(x)} is not monotonic in γ because of competing terms; the discussion would be clearer if the two terms were displayed explicitly.","section":"Section IV-B, Remark 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's main theorems are deferred to the authors' own unpublished companion manuscript [6], which is not available in a verifiable archival form. This is the principal obstacle to acceptance. The gap is fillable: the authors could either include complete proofs in this paper or make the companion paper publicly available and state precisely how its assumptions specialize to the OPmin setting. I therefore recommend major revision rather than rejection. I also note that the paper's contribution is largely incremental relative to [6]; the new state-dependent horizon analysis and the running-cost bound are the elements that would justify separate publication, and those are precisely the parts that need to be fully proved here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the OPmin paper. Bottom line: the main results do not stand alone. Theorem 1 and Corollary 2 are explicitly stated as consequences of the authors' companion [6], which is 'submitted for journal publication' and available only through a reviewer URL. Theorem 2's appendix follows the proof of [6, Theorem 2] and only sketches the state-dependent-horizon adaptation. Since Theorem 3 and Corollary 3 inherit Corollary 2, essentially the whole quantitative core is on loan from an unpublished manuscript. That is a load-bearing dependency, not a stylistic choice.\n\nCredit where credit is due: the algorithm modification is clean. OP for cost minimization, generic nonnegative stage costs, and the undiscounted case is a useful step beyond [10], and Proposition 1 and 2 (what the algorithm returns, and the budget-horizon relation) are proved in the text. The state-dependent near-optimality bound in Theorem 1 is a genuine concept: it doesn't blow up at gamma=1 and it shrinks near the attractor. The running-cost analysis in Theorem 3 is sketched enough to follow, conditional on Corollary 2. The cubic-integrator example is illustrative only—a single academic case with no code and no baseline—but it does show the expected trend.\n\nSoft spots, in proportion: the missing proofs are the big one. The authors are transparent about it, but the paper cannot be independently verified in its current form. A serious referee would need to check [6] carefully, and the transfer from fixed to state-dependent horizons in Theorem 2 is exactly where subtle errors like to hide. Assumption 1(ii) is a strong uniform dissipation inequality across all modes; it is standard in the MPC literature but not constructive, and the paper does not discuss how to verify it beyond the example.\n\nMy recommendation: send it to peer review, but with a firm request that [6] be made publicly available (e.g., posted on arXiv) and that the reviewers explicitly verify the horizon-dependent modifications. The idea is worth engaging, and the authors are honest about what they borrowed. Desk rejection would be too harsh. If I were citing the quantitative guarantees, I'd wait until the companion is checkable. Reading group: maybe, if you want to discuss when deferring proofs to a companion is acceptable.","headline":"The algorithm and concepts are worthwhile, but the main theorems are deferred to an unpublished companion and the paper cannot be fully verified as submitted.","tokens_in":16534,"tokens_out":5024,"would_cite":false,"duration_ms":53899,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C55","93D15","93D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A reworked optimistic-planning algorithm controls nonlinear switched systems near-optimally and with stability guarantees.","keywords":["optimistic planning","OPmin","nonlinear switched systems","near-optimal control","receding-horizon control","practical stability","detectability","discounted cost"],"falsifier":"For a two-mode scalar switched linear system with quadratic stage costs, compute $V_{\\gamma,\\infty}$ exactly by value iteration and run OPmin with a known budget; if the measured gap $V_{\\gamma,\\infty}-V_{\\gamma,d(x)}$ exceeds the bound $v_{\\gamma,d(x)}(x)$ from Theorem 1, or the receding-horizon trajectory fails to enter the predicted $\\delta$-neighborhood within the budget prescribed by Theorem 2, the central claims are false.","tokens_in":15446,"feed_emoji":"⚙️","tokens_out":8888,"duration_ms":80045,"temperature":0.7,"pith_summary":"Optimistic planning is a tree-search method from artificial intelligence that computes near-optimal inputs for systems with finitely many control choices. The paper rewrites it as OPmin so that it applies to control problems as normally posed: stage costs may be arbitrary nonnegative functions such as quadratics, discounting is optional, and the goal is cost minimization rather than reward maximization. The main claim is that, under a stabilizability-and-detectability condition, the truncated cost OPmin returns sits within an explicit state-dependent error of the infinite-horizon optimum, and that the same algorithm run in receding-horizon mode makes the closed-loop system practically stable. A strengthened version of the condition gives global exponential stability and shows that the running cost approaches the infinite-horizon value exponentially as the computational budget grows.","feed_headline":"Reworked optimistic planning stabilizes nonlinear switched systems","feed_subtitle":"It now handles quadratic costs and undiscounted horizons, while certifying near-optimal performance.","key_machinery":"The engine is Assumption 1, a uniform detectability/dissipation condition on every mode, together with the optimistic tree search inherited from OP. In the tree, each leaf carries the cost of the switching sequence that reaches it; OPmin always expands the leaf with the smallest such cost, so the first fully explored level yields an exact finite-horizon optimum, and the resulting horizon $d(x)$ is as large as the budget allows. The dissipation inequality is what turns the tail of the infinite-horizon cost into a bounded, state-dependent error term $v_{\\gamma,d(x)}(x)$, and the same inequality serves as the Lyapunov decrement that produces the stability bound.","core_discovery":"Under Assumption 1 — a Lyapunov-like dissipation inequality $W(f_u(x))-W(x) \\le -\\alpha_W(\\sigma(x))+\\ell_u(x)$ together with a stabilizability bound $V_{\\gamma,\\infty}(x)\\le \\alpha_V(\\sigma(x))$ — the paper proves for every state $x$, discount factor $\\gamma\\in(0,1]$, and state-dependent horizon $d(x)$ that $V_{\\gamma,d(x)}(x)\\le V_{\\gamma,\\infty}(x)\\le V_{\\gamma,d(x)}(x)+v_{\\gamma,d(x)}(x)$, where $v_{\\gamma,d(x)}(x)$ is an explicit function built from the comparison functions. The same assumptions imply that the receding-horizon closed loop satisfies $\\sigma(\\varphi(k,x))\\le \\max\\{\\beta(\\sigma(x),k),\\delta\\}$ for arbitrary $\\delta,\\Delta>0$ once the discount factor is close enough to $1$ and the budget large enough; with linear comparison functions this becomes global exponential stability. The paper further shows that the running cost of the receding-horizon scheme lies within $w_{\\gamma,\\bar d}\\,\\sigma(x)$ of the infinite-horizon optimal value, with $w_{\\gamma,\\bar d}$ decaying exponentially in the minimum horizon.","pith_inferences":["Because Theorem 3 requires no terminal cost or terminal constraint, the paper suggests a model-predictive-control formulation whose only tuning knobs are budget and discount factor; a testable extension is whether the state-dependent horizon can be shortened near the attractor without losing the guarantees.","The exponential factor $\\left(1-\\frac{a_W}{\\bar a_V+\\bar a_W}\\right)^{\\bar d}$ in Corollary 1 quantifies how 'detectable' the cost is; one could turn this into a design guideline for choosing stage costs that make the error small quickly.","The bound in Theorem 1 vanishes as $\\sigma(x)\\to 0$, which the authors note; a natural extension is to use OPmin as a local stabilizer after a coarse global schedule, or to certify performance in a neighborhood of the attractor."],"forward_implications":["OPmin can be applied to optimal control problems with stage costs that are arbitrary nonnegative functions, including quadratic costs, and with discount factor $\\gamma=1$; the error bound remains finite and decays as the horizon grows.","For any prescribed initial-state bound $\\Delta$ and final tolerance $\\delta$, a budget and discount factor can be chosen so that every receding-horizon trajectory converges to the $\\delta$-neighborhood of the attractor, so the scheme gives certified practical stability.","Under linear comparison functions, the closed loop is globally exponentially stable and the mismatch between the running cost and $V_{\\gamma,\\infty}$ decays exponentially in the minimum horizon $\\bar d$ (equivalently, in the budget).","The algorithm provides both an open-loop near-optimal switching sequence and a receding-horizon feedback law from the same computation, so the same tool covers planning and feedback stabilization."],"supporting_citations":[{"why":"Supplies the original optimistic-planning algorithm and the near-optimality bound that OPmin modifies, along with the branching-factor viewpoint.","marker":"[10]"},{"why":"Companion paper whose Theorem 3 yields Theorem 1 and whose Theorem 2 proof is adapted to prove Theorem 2; the error term and Lyapunov decrement come from it.","marker":"[6]"},{"why":"Provides the stabilizability/detectability conditions used in Assumption 1 and the cubic-integrator example used to illustrate the results.","marker":"[8]"},{"why":"Source of the same stabilizability/detectability framework and of the stability notion based on a generic measuring function $\\sigma$.","marker":"[14]"},{"why":"Defines the running cost of receding-horizon controllers and the relative-performance comparison that Theorem 3 and Corollary 3 position against.","marker":"[9]"},{"why":"Supports the standing assumption that an optimal infinite-horizon input sequence exists.","marker":"[11]"}],"fun_headline_variants":["Optimistic planning gains stability for switched systems","Near-optimal control with stability via optimistic planning","OPmin: stable near-optimal control for switched systems","Undiscounted costs tamed in switched-system control","Stability-certified optimistic planning for switched dynamics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guarantees rest on Assumption 1(ii), the detectability/dissipation inequality that must hold for every mode and every stage cost, and on the correctness of the companion-paper results [6] that Theorem 1 and Theorem 2 are derived from; if that inequality fails or those results are wrong, the claims collapse.","fun_headline_variants_meta":{"raw":{"variants":["Optimistic planning gains stability for switched systems","Near-optimal control with stability via optimistic planning","OPmin: stable near-optimal control for switched systems","Undiscounted costs tamed in switched-system control","Stability-certified optimistic planning for switched dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2510,"prompt_tokens":1032,"completion_tokens":1478,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":1405}},"tokens_in":648,"tokens_out":1478,"duration_ms":11972,"temperature":1.0,"reasoning_tokens":1405,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:14:38.429607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a two-mode scalar switched linear system with quadratic stage costs, compute $V_{\\gamma,\\infty}$ exactly by value iteration and run OPmin with a known budget; if the measured gap $V_{\\gamma,\\infty}-V_{\\gamma,d(x)}$ exceeds the bound $v_{\\gamma,d(x)}(x)$ from Theorem 1, or the receding-horizon trajectory fails to enter the predicted $\\delta$-neighborhood within the budget prescribed by Theorem 2, the central claims are false.","supporting_citations":[{"cited_title":"Hren and R","cited_arxiv_id":null,"evidence_quote":"Supplies the original optimistic-planning algorithm and the near-optimality bound that OPmin modifies, along with the branching-factor viewpoint."},{"cited_title":"Granzotto, R","cited_arxiv_id":null,"evidence_quote":"Companion paper whose Theorem 3 yields Theorem 1 and whose Theorem 2 proof is adapted to prove Theorem 2; the error term and Lyapunov decrement come from it."},{"cited_title":"Grimm, M","cited_arxiv_id":null,"evidence_quote":"Provides the stabilizability/detectability conditions used in Assumption 1 and the cubic-integrator example used to illustrate the results."},{"cited_title":"Postoyan, L","cited_arxiv_id":null,"evidence_quote":"Source of the same stabilizability/detectability framework and of the stability notion based on a generic measuring function $\\sigma$."},{"cited_title":"Grüne and A","cited_arxiv_id":null,"evidence_quote":"Defines the running cost of receding-horizon controllers and the relative-performance comparison that Theorem 3 and Corollary 3 position against."},{"cited_title":"Keerthi and E","cited_arxiv_id":null,"evidence_quote":"Supports the standing assumption that an optimal infinite-horizon input sequence exists."}],"review_version":1}