{"id":"71b2a8c5-94fd-4d0b-b22a-dfb1a3af590f","arxiv_id":"2607.03986","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Non-Markovian path-dependent control problems with signature rewards admit optimal feedback controls and value processes as local linear signature expansions whose coefficients solve infinite-dimensional Riccati equations on the extended tensor algebra, globalized by dynamic recentering.","lead":"The paper gives a semi-explicit feedback solution for a class of non-Markovian path-dependent stochastic control problems by reducing them to Riccati equations on the tensor algebra of path signatures. This yields implementable optimal controls beyond the classical linear-quadratic setting, including signature tracking and Volterra lifts.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates class B as the load-bearing premise and correctly notes that the paper’s contribution is the reduction of the control problem to that class together with the dynamic-recentering construction. Because the local expansions and the existence of an optimal weak control are rigorously established once p∈B, and because the numerics corroborate both the local formula and the necessity of recentering, the ACCEPT verdict with high confidence is appropriate. The concrete check proposed above is a low-cost algebraic verification that would further increase confidence but is not expected to alter the claim.","tokens_in":23686,"tokens_out":431,"duration_ms":4563,"concrete_test":"Independently re-derive the identification of α*_t with ⟨ψ^{X_σ}_t |1 , X_{σ,t}⟩ on a single recentering interval by applying the Clark–Ocone formula (Remark 4.8) to the Brownian signature and verifying that the resulting Malliavin derivative coincides with the Riccati feedback; agreement confirms that the Itô-plus-Riccati argument of the proof of Theorem 4.3 contains no algebraic gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Theorem 4.3 rests on the admissible class B (Definition 3.1) and the local existence/radius results of Theorem 3.2, both imported from the authors’ concurrent analytic paper. Within that scope the derivation is tight: Boué–Dupuis reduces the control problem to a conditional log-Laplace transform, Chen’s identity plus left-shift recentering (Proposition 3.3) keep the terminal condition inside B, and Itô’s formula on the tensor algebra identifies the feedback maps on each stochastic interval [σ,τ_σ]. The only residual analytic gap—the a.s. finiteness of the number of recentering times—is acknowledged by the authors and does not invalidate the local representation that is actually proved. No internal inconsistency or hidden assumption that would falsify the stated theorem was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper solves a class of non-Markovian stochastic control problems with path-dependent rewards of the form F(X) = ⟨p, bX_T⟩, where p belongs to an admissible class B of signature coefficients. Using the Boué–Dupuis variational representation, the problem is reduced to conditional log-Laplace transforms of linear signature functionals. These transforms are expanded via an infinite-dimensional Riccati equation on the extended tensor algebra (imported from the authors’ concurrent analytic work). The main result (Theorem 4.3) constructs a weak optimal control α* whose value process and feedback map admit local signature expansions V*_t = ⟨ψ^{bX_σ}_t , bX_{σ,t}⟩ and α*_t = ⟨ψ^{bX_σ}_t |1 , bX_{σ,t}⟩ on stochastic intervals between recentering times, with coefficients solving the recentered Riccati system (3.5). A dynamic recentering algorithm restores a global representation. Applications include tracking of signature functionals and signature lifts of nonlinear Volterra control problems, with numerical illustrations against a Monte-Carlo Clark–Ocone benchmark.","tokens_in":23880,"tokens_out":996,"duration_ms":7565,"significance":"If the result holds, the paper supplies the first rigorous closed-loop feedback representation of optimal controls for a nontrivial class of genuinely path-dependent, non-quadratic rewards, expressed as time-dependent linear functionals of the controlled signature. This goes beyond both classical LQ theory and existing signature-parametrization approaches that optimize only over restricted control classes. The reduction via Boué–Dupuis plus the recentering device cleanly separates the analytic existence theory (class B, radius of convergence) from the control-theoretic construction, and the numerical examples demonstrate that the method is implementable beyond the linear-quadratic regime. The work therefore opens a concrete route from signature algebra to non-Markovian stochastic control.","major_comments":[{"comment":"The load-bearing existence and radius statements (Theorem 3.2) and the preservation of class B under left shifts (Proposition 3.3) are imported wholesale from the concurrent arXiv 2606.29622. While the control derivation itself is self-contained once those results are granted, the present manuscript never states the precise hypotheses under which the imported theorems apply (e.g., the quantitative bounds on the shuffle exponential). A short self-contained appendix or a precise citation of the exact statements used would make the paper independently readable and would clarify the scope of Theorem 4.3.","section":null},{"comment":"Remark 4.7 and the numerical experiment of Figure 2 treat time-dependent running rewards ⟨f_t , bX_t⟩ as if the same Riccati theory applies, yet the authors explicitly note that the required extension of the log-Laplace results of Abi Jaber–Attal–Sotnikov (2026a) is only conjectural. Either the time-dependent case should be removed from the main claims or a precise statement of what is proved versus what is numerically observed should be added, so that the reader can distinguish the theorem from the conjecture.","section":null}],"minor_comments":[{"comment":"Definition 3.1 of class B is dense; a short paragraph explaining why the leading even-degree negative terms are necessary (and why lower-order terms can be absorbed) would help non-specialists.","section":null},{"comment":"In the proof of Theorem 4.3 the application of Itô’s formula on the tensor algebra cites Theorem 3.5 of the companion paper; a one-line reminder of the precise integrability condition used would improve readability.","section":null},{"comment":"Figures 1–2 would benefit from a clearer legend distinguishing the three methods (Riccati with recentering / without / Monte-Carlo) and from an explicit statement of the truncation level N_trunc = 10 and the number of Monte-Carlo paths.","section":null},{"comment":"The a.s. finiteness of the number of recentering times is left open; a brief remark that the local representation remains valid pathwise even if infinitely many recenterings accumulate would remove any residual ambiguity.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is tightly linked to the authors’ concurrent analytic work on the Brownian signature. The control contribution is genuine and the derivation is clean, but the journal may wish to ensure that the companion paper is either already accepted or made available in a stable form so that referees of the present manuscript can verify the imported analytic statements. Fit for a strong applied-probability / stochastic-control journal is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper gives a clean, usable feedback form for a genuine class of path-dependent control problems. The move is to reduce the problem via Boué–Dupuis to conditional log-Laplace transforms of linear signature functionals, then import the authors’ concurrent analytic results on the tensor algebra to get an infinite Riccati system whose local solutions yield both the value process and the optimal control as signature expansions. That is new relative to their earlier verification-style signature-control paper and relative to the signature-parametrization literature that only optimizes over restricted control classes.\n\nWhat works: Theorem 4.3 is carefully derived. Lemma 4.12 constructs the Föllmer drift in the weak formulation; Chen + left-shift recentering keeps the terminal condition inside class B; Itô on the tensor algebra identifies the feedback maps on each stochastic interval between recentering times. The numerics (tracking signature functionals, time-dependent targets, Volterra lifts) match the Monte-Carlo Clark–Ocone benchmark and make the necessity of recentering visible. When the reward collapses to classical LQ the formulas recover the usual finite Riccati, which is a useful sanity check.\n\nSoft spots are real but proportional. Everything rests on class B (and the radius estimates) from the concurrent arXiv 2606.29622; if a natural reward sits outside B the representation is unavailable. The time-dependent-source extension is only conjectured (though numerically supported). Global a.s. finiteness of the number of recentering times is left open; the authors prove only the local representation. None of these break the stated theorem.\n\nThis is for people who already work with signatures or infinite-dimensional Riccati systems and want a rigorous closed-loop solution beyond LQ/Volterra-LQ. The math and citation pattern look solid; the dependence on the companion paper is transparent rather than circular. I would send it to referees and I would cite the feedback representation and the recentering algorithm.","headline":"Solid closed-loop signature control via Boué–Dupuis + tensor Riccati, with recentering that actually works for a nontrivial class beyond LQ.","tokens_in":24491,"tokens_out":494,"would_cite":true,"duration_ms":5512,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93E20","60L10","34G20"],"pacs":[],"model":"grok-4.5","headline":"Path-dependent stochastic control reduces to Riccati equations on the tensor algebra of signatures, with an explicit feedback law recovered by dynamic recentering.","keywords":["path signatures","stochastic optimal control","Riccati equation","tensor algebra","Boué-Dupuis formula","non-Markovian control","dynamic recentering"],"falsifier":"Take a concrete reward in class B (for instance the quartic tracking functional used in Section 5), solve the truncated Riccati system with dynamic recentering, and compare the resulting closed-loop trajectories and value process against an independent Monte-Carlo evaluation of the Clark-Ocone formula; systematic discrepancy outside numerical truncation error would refute the claimed representation.","tokens_in":24576,"feed_emoji":"∫","tokens_out":646,"duration_ms":5381,"temperature":0.7,"pith_summary":"The paper shows how to solve a large family of non-Markovian control problems whose rewards depend on the whole path of a controlled Brownian motion. By writing the reward as a linear functional of the path's time-augmented signature and invoking the Boué-Dupuis variational formula, the authors convert the original problem into the evaluation of a conditional log-Laplace transform. That transform is known to solve an infinite-dimensional Riccati equation living on the extended tensor algebra. The solution of the Riccati equation supplies the coefficients of an explicit feedback map: both the value process and the optimal control become infinite linear combinations of the signature of the controlled path. Because the series converge only locally, the authors introduce a dynamic recentering procedure that restarts the expansion whenever the current signature leaves the disk of convergence, thereby obtaining a globally valid representation. The method recovers classical linear-quadratic formulae as special cases and is illustrated on genuinely nonlinear, path-dependent examples that previously lacked closed-form solutions.","feed_headline":"Signature Riccati equations solve path-dependent control","feed_subtitle":"Local expansions plus dynamic recentering give global feedback beyond linear-quadratic cases","key_machinery":"The recentered Riccati equation (3.5) on the extended tensor algebra, whose solution ψ supplies the time-dependent coefficients of the local signature expansions of the value process and the optimal control.","core_discovery":"For terminal rewards whose signature coefficients lie in the admissible class B, there exists a weak optimal control whose value process and feedback law admit local signature expansions whose coefficients solve a recentered Riccati equation on the extended tensor algebra; a dynamic recentering algorithm stitches these local expansions into a global representation over the whole horizon.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Riccati systems on signatures solve non-Markovian path-dependent control","Local signature expansions give global feedback via dynamic recentering","Tensor algebra Riccati equations yield optimal control for signature rewards","Path signatures reformulate control as infinite-dimensional Riccati systems","Signature lifts handle Volterra and path-dependent problems beyond LQ"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The reward coefficient must belong to a special algebraic class that guarantees both existence of a Riccati solution and a positive radius of convergence for the signature series; if a natural path functional falls outside that class, the feedback representation fails.","fun_headline_variants_meta":{"raw":{"variants":["Riccati systems on signatures solve non-Markovian path-dependent control","Local signature expansions give global feedback via dynamic recentering","Tensor algebra Riccati equations yield optimal control for signature rewards","Path signatures reformulate control as infinite-dimensional Riccati systems","Signature lifts handle Volterra and path-dependent problems beyond LQ"]},"model":"grok-4.5","effort":"low","cost_usd":0.006884,"raw_usage":{"total_tokens":1678,"prompt_tokens":700,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":68840000,"prompt_tokens_details":{"text_tokens":700,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":888,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":700,"tokens_out":90,"duration_ms":6895,"temperature":1.0,"reasoning_tokens":888,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T22:30:09.786603+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Take a concrete reward in class B (for instance the quartic tracking functional used in Section 5), solve the truncated Riccati system with dynamic recentering, and compare the resulting closed-loop trajectories and value process against an independent Monte-Carlo evaluation of the Clark-Ocone formula; systematic discrepancy outside numerical truncation error would refute the claimed representation.","supporting_citations":[],"review_version":1}