{"id":"dff18b9f-d381-4b29-aa99-2aba90bc79fd","arxiv_id":"1908.01747","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"For Caputo fractional-order systems, the optimal value must depend on the entire trajectory history; with that history dependence, the dynamic programming principle and a Hamilton-Jacobi-Bellman equation hold, and smooth solutions yield optimal feedback controls.","lead":"This paper proves that the optimal-control recursion known as dynamic programming remains valid for systems governed by fractional derivatives, provided the 'state' is replaced by the whole history of the trajectory. It then derives the corresponding Hamilton-Jacobi-Bellman equation and shows how to build optimal feedback controls for such fractional-order systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HJB results depend on an unverified ci-smoothness assumption that can fail for admissible problems; this is an explicit limitation, not an internal error.","rationale":"The reader's weakest_assumption identifies exactly the same concern: the value functional is not shown to be ci-smooth, and the paper itself acknowledges this limitation. My stress-test confirms the conditional theorems are internally consistent; the proof of Theorem 10.1 follows the standard DPP-to-HJB argument using Lemma 9.2, and Theorem 11.1 is a correct verification theorem under its stated assumptions. The concern is about applicability, not correctness. A simple zero-dynamics problem with a nondifferentiable terminal cost shows that ci-smoothness can fail for problems satisfying all standing assumptions, so the HJB results do not cover the entire stated problem class. Since the authors explicitly frame the HJB results as conditional and flag the smoothness gap in Section 13, this does not undermine the paper's claims. The DPP part (Theorem 6.1) is unconditional and appears correct. The example in Section 12 provides a nontrivial check that the conditional theory is not vacuous. Therefore the reader's ACCEPT verdict remains appropriate, and no correction to the manuscript's stated claims is required, though the limitation should be kept prominent.","tokens_in":29137,"tokens_out":12367,"duration_ms":141040,"concrete_test":"For the problem (CDαx)(τ) = 0, χ ≡ 0, T = 1, σ(x) = |x|, compute ρ(t,w) = |w(0) + (1/Γ(α))∫_0^t (CDαw)(ξ)/(1−ξ)^{1−α} dξ|. Fix (t,w) with this linear expression equal to 0 and choose two extensions with CDαx = l1 and CDαx = l2 on (t,T), with l1 and l2 not collinear (or l1 = 1, l2 = −1 in scalar case). Evaluate [ρ(τ,xτ) − ρ(t,w)]/(τ − t) as τ ↓ t. If the limit is not a fixed linear function of l with a single gradient vector independent of l, ci-differentiability fails, confirming the HJB theorems' smoothness assumption is not automatically satisfied.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main HJB theorems (Theorem 10.1 and Theorem 11.1) are conditional on the value functional being ci-smooth of order α. The paper proves only continuity of ρ (Theorem 8.3), and Section 13 explicitly concedes that ρ may lack the required smoothness in general. This is not a flaw in the conditional proofs, but it means the HJB connection is not established for the full problem class stated in Section 3. The assumption is also not merely unverified. Consider the admissible problem with f ≡ 0, χ ≡ 0, an arbitrary compact control set P, and terminal cost σ(x) = |x|. Then the value functional is ρ(t,w) = |ρ*(t,w)|, where ρ* is a linear functional of the history, and at any position with ρ*(t,w) = 0, the required ci-expansion would have to satisfy |⟨s,l⟩| = ⟨s,l⟩ for all l ∈ R^n, which is impossible for a fixed vector s unless n = 1 and s is chosen in an l-dependent way. Thus ci-differentiability fails, and the HJB verification theorem does not apply to this otherwise admissible problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a Bolza-type optimal control problem for systems governed by a Caputo fractional differential equation of order α ∈ (0,1). Because the Caputo derivative is nonlocal, the author argues in Section 4 that the value of the problem must be viewed as a functional of the entire history w(·) on [0,t], not merely of the current state w(t). The main results are: (i) the dynamic programming principle (Theorem 6.1), proved directly for the history-dependent value functional; (ii) continuity of the value functional on the space of positions with a Hausdorff-type metric (Theorem 8.3); (iii) a new notion of fractional coinvariant (ci-) derivatives and a corresponding Hamilton-Jacobi-Bellman equation (Section 9 and Eq. (10.1)); (iv) a verification theorem stating that a ci-smooth value functional satisfies the HJB equation (Theorem 10.1) and that a ci-smooth solution of the HJB equation with the natural terminal condition coincides with the value and yields an optimal feedback strategy (Theorem 11.1). The paper closes with a fully worked example that is solved using the HJB framework.","tokens_in":29400,"tokens_out":9630,"duration_ms":98203,"significance":"If correct, the paper gives a clean and mathematically rigorous extension of dynamic programming and HJB methods to fractional-order control systems, and it is transparent about the history-dependent nature of the problem, which is an essential and non-obvious point. The direct proof of the dynamic programming principle is unconditional and does not rely on smoothness assumptions. The introduction of fractional ci-derivatives is natural and, as shown in Lemma 9.2, yields a simple formula for the total derivative of functionals along fractional trajectories, which is a useful tool for future work. The HJB verification results are explicitly conditional on ci-smoothness, and Section 13 correctly states that the value functional may fail to satisfy this assumption in general. This is a scope limitation rather than an internal inconsistency: the paper does not claim the HJB equivalence for the entire non-smooth problem class. The worked example in Section 12 is internally consistent and illustrates the theory well.","major_comments":[],"minor_comments":[{"comment":"The notation ρ(θ,x(·)) is slightly ambiguous because x(·) denotes the motion on [0,θ] in the theorem statement, while elsewhere x(·) denotes a full trajectory on [0,T]; please clarify explicitly that ρ(θ,x(·)) means ρ(θ,x_θ(·)), the value functional evaluated at the history of x on [0,θ].","section":"Theorem 6.1, Eq. (6.2)"},{"comment":"The application of Dini's theorem is very terse: the proof bounds the upper right derivative of ω, but to conclude two-sided Lipschitz continuity one should note that the same argument applied to −ω gives the lower bound; adding one sentence would make the argument fully transparent.","section":"Lemma 9.2, proof"},{"comment":"The caveat that the value functional may not be ci-smooth is stated only in the concluding section; because it is essential for interpreting Theorems 10.1 and 11.1, this limitation should also be signaled in the introduction or in the abstract's discussion of the HJB results.","section":"Section 13"},{"comment":"The phrase \"Formally extending the motion x(·) up to T\" is vague; please specify that the extension is obtained by continuing with the same constant control on [t+δ,T], after which Lemma 9.2 is applied on [t,t+δ].","section":"Proof of Theorem 10.1"},{"comment":"In the displayed formulas for ∂_t^α φ and ∇^α φ in the first case, the absence of parentheses around \"2ρ*(t,w(·)) + (T−t)^α\" makes the formulas ambiguous; adding parentheses will improve readability.","section":"Section 12, formulas for ci-derivatives"}],"recommendation":"minor_revision","confidential_remarks":"The paper relies on the author's own prior results, especially [15, Proposition 2], for existence and uniqueness of motions; this reliance is appropriate and clearly cited. The conditional nature of the HJB results is explicitly acknowledged, and I see no basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: Theorem 6.1, the dynamic programming principle for history-dependent values, is the real result and it is correct. The HJB part is conditional, and the smoothness assumption is not a minor technicality: it fails for an elementary admissible problem. The author admits this in Section 13, so it is a limitation, not a hidden error.\n\nWhat is new: earlier fractional HJB work used state-only value functions, which are wrong for Caputo dynamics because of memory. The history-dependent value functional ρ(t,w(·)) fixes that, and the Section 4 example convincingly shows why a state-only position fails. The DPP proof is a standard argument, but the setup is careful: positions are histories, the semigroup property is used explicitly, and Proposition 7.1 supplies the uniform Hölder and Lipschitz bounds needed later. Continuity of ρ (Theorem 8.3) is a genuine piece of work. The fractional coinvariant derivative is a natural adaptation of ci-calculus to the fractional setting, and Lemma 9.2's total derivative formula is the useful payoff.\n\nWhere the soft spots are: The HJB verification theorems (10.1 and 11.1) require ci-smoothness of the value functional. This is not merely unverified. Take f ≡ 0, χ ≡ 0, and terminal cost σ(x) = |x|. Then ρ(t,w) = |ρ*(t,w)|, where ρ* is a linear functional of the history. At a history with ρ* = 0, the ci-expansion would force the map l ↦ |⟨s,l⟩| to be linear in l, which is impossible; even in one dimension the increment behaves like |l|(τ−t)^α, not linearly in τ−t. So the HJB equation does not apply to a perfectly admissible problem. The author explicitly acknowledges that the value may lack smoothness and defers to future viscosity-type theory, so the theorems are honestly stated; but anyone citing the HJB result should know it covers a restricted class. Lemma 9.2's Dini theorem step is terse and could use expansion. The citation pattern leans on the author's own prior result for existence, uniqueness, and semigroup properties, but that is a published adjacent result, not circular.\n\nWho it is for: people working on fractional optimal control, and anyone connecting fractional dynamics to the functional-differential HJB tradition. It deserves a serious referee; I would recommend accept with minor revisions, mainly asking the authors to make the non-genericity of ci-smoothness visible earlier in the abstract and introduction.","headline":"History-dependent DPP for Caputo fractional control is correct and unconditional; the HJB half is honest but applies only to a restricted, non-generic smooth class.","tokens_in":29856,"tokens_out":3444,"would_cite":true,"duration_ms":38486,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["26A33","34A08","49L20","35F21"],"pacs":[],"model":"deepseek-v4-flash","headline":"Dynamic programming holds for fractional-order optimal control when the value functional is defined on trajectory histories, and a new calculus of coinvariant derivatives turns this principle into a Hamilton–Jacobi–Bellman equation.","keywords":["optimal control","fractional derivatives","Caputo derivative","dynamic programming principle","Hamilton-Jacobi-Bellman equation","coinvariant derivatives","feedback control","history-dependent value"],"falsifier":"Use the scalar system $(C D^\\alpha x)(\\tau)=\\Gamma(\\alpha+1)u(\\tau)$, $|u(\\tau)|\\leq 1$, $\\tau\\in[0,T]$, with terminal cost $\\sigma(x(T))=|x(T)|$ instead of $x^2$. The value becomes $\\rho(t,w(\\cdot)) = \\max\\{0,\\, |\\rho_*(t,w(\\cdot))| - (T-t)^\\alpha\\}$, where $\\rho_*$ is the ci-smooth functional computed in Section 12; at any history with $|\\rho_*(t,w(\\cdot))| = (T-t)^\\alpha$ this functional has a kink and is not ci-differentiable of order $\\alpha$. This directly violates the hypothesis of Theorems 10.1 and 11.1, showing that the dynamic programming principle alone does not produce the HJB equation.","tokens_in":28939,"feed_emoji":"⚙️","tokens_out":9029,"duration_ms":79078,"temperature":0.7,"pith_summary":"The paper shows that dynamic programming works for fractional-order optimal control—but only if the 'position' of the system is taken to be the entire history of the motion up to the current time, not just the current state. The reason is that the Caputo derivative is nonlocal: its value at time $t$ depends on all earlier values, so any sub-problem starting at $t$ must be fed the past trajectory. With the value as a functional $\\rho(t,w(\\cdot))$ on histories, the paper proves the dynamic programming principle in full generality (Theorem 6.1), and then introduces a tailored calculus of 'coinvariant derivatives' of order $\\alpha$ to write an associated Hamilton–Jacobi–Bellman equation. Under a ci-smoothness assumption, the value functional satisfies this equation, and any solution of the HJB Cauchy problem with the right terminal condition equals the value and yields an optimal feedback control strategy. The paper's concluding section notes the main limitation: the value may not actually be ci-smooth, so the HJB connection is conditional.","feed_headline":"History-based value yields Bellman principle for fractional control","feed_subtitle":"Coinvariant derivatives convert this principle into an HJB equation whose solutions build optimal feedback.","key_machinery":"The load-bearing object is the fractional coinvariant derivative. For a functional $\\varphi(t,w(\\cdot))$ on histories, $\\varphi$ is ci-differentiable of order $\\alpha$ at $(t,w(\\cdot))$ if there exist $\\partial^\\alpha_t \\varphi(t,w(\\cdot)) \\in \\mathbb{R}$ and $\\nabla^\\alpha \\varphi(t,w(\\cdot)) \\in \\mathbb{R}^n$ such that for every extension $x(\\cdot)$ of the history $w(\\cdot)$ and every $\\tau>t$, $\\varphi(\\tau,x_\\tau(\\cdot)) - \\varphi(t,w(\\cdot)) = \\partial^\\alpha_t \\varphi\\,(\\tau-t) + \\langle\\nabla^\\alpha \\varphi,\\, (I^{1-\\alpha}(x(\\cdot)-x(0)))(\\tau) - (I^{1-\\alpha}(w(\\cdot)-w(0)))(t)\\rangle + o(\\tau-t)$. Interpreting the difference of fractional integrals as $\\int_t^\\tau (C D^\\alpha x)(\\xi)\\,d\\xi$, this yields the clean total-derivative formula $\\frac{d}{d\\tau}\\varphi(\\tau,x_\\tau(\\cdot)) = \\partial^\\alpha_t \\varphi + \\langle\\nabla^\\alpha \\varphi,\\, (C D^\\alpha x)(\\tau)\\rangle$ along motions (Lemma 9.2). This formula is what converts the dynamic programming principle into the infinitesimal HJB equation, and it also makes the extremal-shift feedback construction possible.","core_discovery":"For the Bolza problem with Caputo dynamics $(C D^\\alpha x)(\\tau)=f(\\tau,x(\\tau),u(\\tau))$, the value functional $\\rho(t,w(\\cdot)) = \\inf_{u\\in U(t,T)} \\big[\\sigma(x(T))+\\int_t^T \\chi(\\tau,x(\\tau),u(\\tau))\\,d\\tau\\big]$ over histories $w(\\cdot) \\in AC^\\alpha([0,t],\\mathbb{R}^n)$ satisfies, for every intermediate time $\\theta$, the dynamic programming principle $\\rho(t,w(\\cdot)) = \\inf_{u\\in U(t,\\theta)} \\big(\\rho(\\theta,x(\\cdot)) + \\int_t^\\theta \\chi(\\tau,x(\\tau),u(\\tau))\\,d\\tau\\big)$, where $x(\\cdot)$ is the motion that continues the history $w(\\cdot)$ under $u$. This principle is proved without extra assumptions. Using a new notion of coinvariant (ci-) differentiation of order $\\alpha$, the paper associates the problem with the Hamilton–Jacobi–Bellman equation $\\partial^\\alpha_t \\varphi(t,w(\\cdot)) + H(t,w(t),\\nabla^\\alpha \\varphi(t,w(\\cdot))) = 0$, with the Hamiltonian $H(\\tau,x,s)=\\min_{u\\in P}(\\langle s,f(\\tau,x,u)\\rangle + \\chi(\\tau,x,u))$. If the value functional is ci-smooth of order $\\alpha$, it solves this equation (Theorem 10.1); conversely, any ci-smooth solution with terminal condition $\\varphi(T,w(\\cdot))=\\sigma(w(T))$ coincides with the value functional, and the strategy $U^\\circ(t,w(\\cdot)) \\in \\arg\\min_{u\\in P}\\big(\\langle\\nabla^\\alpha \\varphi(t,w(\\cdot)), f(t,w(t),u)\\rangle + \\chi(t,w(t),u)\\big)$ is optimal (Theorem 11.1).","pith_inferences":["The gap the paper leaves open is exactly the differentiability of $\\rho$; a natural next step is a viscosity or minimax theory for equation (10.1) on the history space $\\mathcal{G}$, and the compactness results in Section 10 (Proposition 10.2) provide the infrastructure for such a theory.","Because the DPP is unconditional, it can be used as a correctness test for numerical schemes: any discretization that ignores the history dependence of the value will fail Bellman's recursion, so practical algorithms for fractional optimal control should feed the whole past trajectory into the value update.","The same ci-derivative calculus can be imported into stability analysis: Lemma 9.2 offers a total-derivative formula for Lyapunov–Krasovskii functionals on histories, which sidesteps the chain-rule difficulties that arise with functions $V(t,x(t))$ in fractional systems.","A testable extension: compute $\\rho(t,w(\\cdot))$ for the Section 12 example with terminal cost $|x(T)|$ instead of $x^2(T)$; the value has a kink at $|\\rho_*| = (T-t)^\\alpha$, localizing exactly where viscosity-type solutions would be needed and providing a benchmark problem for a nonsmooth HJB theory."],"forward_implications":["The dynamic programming principle (Theorem 6.1) is unconditional: for any admissible control split at an intermediate time $\\theta$, the value from $t$ equals the infimum of the running cost on $[t,\\theta]$ plus the value from the resulting history at $\\theta$.","If the value functional is ci-smooth of order $\\alpha$, it solves the Hamilton–Jacobi–Bellman equation (10.1), so in that case optimal control problems reduce to solving a single functional PDE.","Conversely, a ci-smooth solution of the HJB Cauchy problem with terminal condition $\\sigma(w(T))$ is exactly the value functional, making the HJB equation equivalent to the optimal control problem in the smooth case.","Given such a solution $\\varphi$, the feedback strategy $U^\\circ(t,w(\\cdot)) = \\arg\\min_{u\\in P}\\big(\\langle\\nabla^\\alpha \\varphi(t,w(\\cdot)), f(t,w(t),u)\\rangle + \\chi(t,w(t),u)\\big)$ is optimal in the positional sense: for any sufficiently fine partition, the stepwise control law $\\{U^\\circ,\\Delta\\}$ produces $\\varepsilon$-optimal controls.","Corollary 11.4: when the value functional itself is ci-smooth, the strategy built from its ci-gradient by the same extremal shift is optimal."],"supporting_citations":[{"why":"Establishes the optimality principle that the dynamic programming principle formalizes and supplies the standard proof pattern used in Theorem 6.1.","marker":"[6]"},{"why":"Provides the neutral-type functional-differential representation of Caputo systems, including existence, uniqueness, and the semigroup property of motions used to define sub-problems.","marker":"[15]"},{"why":"Develops the ci-differentiation framework and functional Hamilton–Jacobi theory for hereditary systems that this paper extends to fractional order.","marker":"[31]"},{"why":"Supplies the fractional integral estimates, Hölder-continuity bound (2.1), and the integration scheme behind Lemma 7.2.","marker":"[38]"},{"why":"Provides the fractional Gronwall-type inequality (Lemma 7.3) and the basic theory of Caputo differential equations used throughout.","marker":"[10]"},{"why":"Gives the standard two-sided proof of the dynamic programming principle that Theorem 6.1 follows.","marker":"[44]"},{"why":"Provides the extremal-shift positional strategy and the compactness assertion used for the feedback law in Section 11 and Proposition 10.2.","marker":"[13]"}],"fun_headline_variants":["Fractional control: history-based DPP and coinvariant HJB","New coinvariant derivative yields HJB equation for fractional systems","DPP for Caputo control proven, leading to optimal feedback","History-valued fractional control: Bellman principle to HJB","Coinvariant calculus connects fractional DPP to HJB feedback"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The HJB and feedback theorems assume the value functional is ci-smooth of order $\\alpha$, meaning it is ci-differentiable at every interior position with continuous ci-derivative and ci-gradient; the paper proves continuity of the value but not this differentiability, and Section 13 explicitly concedes that the value functional may fail to be ci-smooth.","fun_headline_variants_meta":{"raw":{"variants":["Fractional control: history-based DPP and coinvariant HJB","New coinvariant derivative yields HJB equation for fractional systems","DPP for Caputo control proven, leading to optimal feedback","History-valued fractional control: Bellman principle to HJB","Coinvariant calculus connects fractional DPP to HJB feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000765,"raw_usage":{"total_tokens":3450,"prompt_tokens":1058,"completion_tokens":2392,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":2307}},"tokens_in":674,"tokens_out":2392,"duration_ms":16490,"temperature":1.0,"reasoning_tokens":2307,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:04:40.999872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the scalar system $(C D^\\alpha x)(\\tau)=\\Gamma(\\alpha+1)u(\\tau)$, $|u(\\tau)|\\leq 1$, $\\tau\\in[0,T]$, with terminal cost $\\sigma(x(T))=|x(T)|$ instead of $x^2$. The value becomes $\\rho(t,w(\\cdot)) = \\max\\{0,\\, |\\rho_*(t,w(\\cdot))| - (T-t)^\\alpha\\}$, where $\\rho_*$ is the ci-smooth functional computed in Section 12; at any history with $|\\rho_*(t,w(\\cdot))| = (T-t)^\\alpha$ this functional has a kink and is not ci-differentiable of order $\\alpha$. This directly violates the hypothesis of Theorems 10.1 and 11.1, showing that the dynamic programming principle alone does not produce the HJB equation.","supporting_citations":[{"cited_title":"Bellman , Dynamic programming, Princeton University Press, 1957","cited_arxiv_id":null,"evidence_quote":"Establishes the optimality principle that the dynamic programming principle formalizes and supplies the standard proof pattern used in Theorem 6.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Develops the ci-differentiation framework and functional Hamilton–Jacobi theory for hereditary systems that this paper extends to fractional order."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the fractional integral estimates, Hölder-continuity bound (2.1), and the integration scheme behind Lemma 7.2."},{"cited_title":"Yong , Diﬀerential games: a concise introduction , W orld scientiﬁc, 2015, https://doi.org/ 10.1142/9121","cited_arxiv_id":null,"evidence_quote":"Gives the standard two-sided proof of the dynamic programming principle that Theorem 6.1 follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the extremal-shift positional strategy and the compactness assertion used for the feedback law in Section 11 and Proposition 10.2."}],"review_version":1}