{"id":"69e6b3e4-5940-4d51-8f82-2a9060017959","arxiv_id":"2412.03294","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"For an averaged ensemble of linear stochastic systems, the Schrodinger bridge is steered optimally by a non-Markovian stochastic feedforward control that integrates past noise.","lead":"This paper derives the optimal way to steer a family of related random systems, each with slightly different dynamics, from one probability distribution to another. The resulting control is a stochastic feedforward rule that uses the whole history of noise, unlike the standard memoryless Schrodinger bridge controllers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1's optimal-prior formula conditions on the passive noise path instead of the controlled state; in the A=0 Markov limit it gives a nonzero singular control when u=0 is optimal, so the central formula is false as stated.","rationale":"The reader's weakest_assumption correctly flags the transfer of the Markov h-transform form to the non-Markov case as unproved. The stress-test pass sharpens this into a concrete inconsistency: the proof of Theorem 4.1 replaces the optimal prior conditional on the controlled process, defined in (4.14), by the passive process conditional on the noise path, defined in (4.18), without justification. In the Markov limit A(theta)=A, the two conditionals differ exactly by the accumulated control, so the claimed formula does not reduce to the known and correct Theorem 2.1 result. The scalar passive-target example gives an explicit contradiction: u=0 is feasible and optimal, while (4.1) returns a nonzero, even infinite-norm, control. Therefore the explicit central formula of the paper is not merely lacking a proof; it is false as written. The pinned-control characterization in Theorem 3.1 may still be correct, and a repaired theorem may exist with the prior conditioned on the controlled state or history, but that is a substantive revision rather than a minor gap.","tokens_in":19819,"tokens_out":23444,"duration_ms":235922,"concrete_test":"Run the closed-form check: solve the Schrodinger system (2.9) for A=0, B=1, eps=1, tf=1, x0=0, rho0=delta_0, rhof=N(0,1). Here phi_f=1 is valid and u=0 attains cost 0, so any nonzero control of the form (4.1) cannot be optimal. Evaluate the right-hand side of (4.1): with (4.3) equal to N(W_t,1-t), it becomes -int_0^t dW/(1-tau)+W_t, which is not zero and has infinite L2 norm. If the authors instead intend p* in (4.2) to condition on the controlled state, replace (4.2) by q(t,x*(t),tf,xf)phi_f(xf)/int q phi_f and verify whether the singular term cancels; the published formula does not do so.","verdict_should_be":"REJECT","load_bearing_attack":"The proof of Theorem 4.1 defines the optimal prior density correctly in (4.14) as a ratio of controlled densities, but then, in characterizing p*, switches to the conditional density (4.18) of the passive process y given x0 and {W(s):0<=s<=t}. This is a different object: the controlled state x*(t) in (4.6) differs from the passive state x0+sqrt(eps)*int_0^t Phi(tf,tau)dW(tau) by the accumulated control integral. For A(theta)=A, the optimal prior is q(t,x*(t),tf,·)phi_f / int q phi_f, as in Theorem 2.1, not the passive density. Consequently, (4.1) retains the singular stochastic-integral term of the pinned control unconditionally. A direct test: take A=0, B=1, eps=1, tf=1, x0=0, mu0=delta_0, muf=N(0,1). Then phi_f=1 solves (2.9), u=0 is feasible with cost 0, and hence is optimal. Formula (4.1)-(4.3) gives u*(t|0)=-int_0^t dW_tau/(1-tau)+W_t, which is not zero and has infinite L2 norm. Thus the central formula is false as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a Schrödinger bridge problem for a parameter-averaged stochastic system. Individual systems evolve as dX(t,θ)=(A(θ)X(t,θ)+B(θ)u(t))dt+√εB(θ)dW(t), and the object to be controlled is the averaged process x(t)=∫_0^1 X(t,θ)dθ. The authors argue that the optimal control for this non-Markov averaged problem is a stochastic feedforward control. Section 2 reviews the single Markov case and gives a path-integral representation. Section 3 derives a pinned local control u*(t|x0,xf) through a penalized discrete optimal control problem, and Section 4 proposes a general formula u*(t|x0)=∫u*(t|x0,xf)p*(tf,xf|0,x0,t)dxf, where p* is built from a conditional transition density of the passive averaged process. Numerical examples in Section 5 simulate the proposed controller.","tokens_in":20142,"tokens_out":9301,"duration_ms":92667,"significance":"If Theorem 4.1 were valid, the paper would be the first to present a path-integral, non-Markov stochastic feedforward control for an averaged Schrödinger bridge problem, and the absence of fitted parameters in the control law would be a strength. The discrete calculation in Proposition 3.1 is explicit and appears to be carried out in detail. However, the central theorem is not merely incompletely proved: as stated it is false in a simple Markov limit. This prevents the paper from achieving its advertised significance in its current form.","major_comments":[{"comment":"Theorem 4.1 is false as stated. Consider A(θ)=0, B(θ)=1, ε=1, tf=1, x0=0, μ0=δ0, μf=N(0,1). Then φf=1 solves (2.9), and u=0 is feasible with cost 0, hence optimal. However, substituting the data into (4.1)-(4.3) gives u*(t|0)=-∫_0^t (1-τ)^{-1}dW(τ)+W(t), which is nonzero and has infinite L2 norm on [0,1]. The root cause is that (4.18) conditions on the passive noise path {W(s):0≤s≤t}, whereas the optimal prior defined in (4.14) is a ratio of densities of the controlled process (4.6) and (4.8). These are different objects unless the control is zero, and the proof never establishes that (4.14) equals (4.18).","section":"Section 4, Theorem 4.1"},{"comment":"The passage from the discrete penalized control (3.15) to the continuous hard-constrained control (3.4) is asserted without proof. Immediately after (3.15), the paper states that [22, Definition 3.1.6 and Corollary 3.1.8] imply the double limit lim_{a→∞}lim_{k→∞}||u*_{a,k,i}-u*||^2_{L2(P)}=0, but no argument is given for the convergence of the discrete adapted controls to the Itô integral in (3.4), for the removal of the 1/(2a) regularization, or for the interchange of the limits. Since Theorem 3.1 is the foundation for the pinned controls used in Theorem 4.1, this gap is load-bearing.","section":"Section 3, proof of Theorem 3.1"},{"comment":"The argument from (4.13) to (4.16) does not establish the claimed equality even if the prior distribution were correct. A process Z(t) with E[Z(t)]=0 can be nonzero, and the statement that the 'extra sample path' ∫_0^t Φ(t,τ)Z(τ)dτ contradicts optimality is not a rigorous uniqueness proof for the non-Markov stochastic feedforward setting. This step is presented as if it were a contradiction, but no uniqueness theorem for the optimal control is stated or proved.","section":"Section 4, proof of Theorem 4.1"}],"minor_comments":[{"comment":"Equation (5.7) states ∫_0^1 e^{-θ}dθ = 1 - e; the correct value is 1 - e^{-1}, and this appears to be a typo.","section":"Section 5, Eq. (5.7)"},{"comment":"The notation for the conditional density is inconsistent: (4.2) writes q-hat with arguments (0,x0,t,tf,y), while (4.3) uses the same symbol with a different ordering; the meaning of the five arguments should be stated explicitly.","section":"Section 4, Eq. (4.2) and (4.3)"},{"comment":"The numerical section is illustrative rather than quantitative: no cost values, no comparison with a baseline, and no convergence assessment are reported for the Monte Carlo estimate in Figure 7.","section":"Section 5"},{"comment":"The conclusion describes the solution as 'formally solved' and identifies the computation of the path integral as future work; the abstract and Theorem 4.1, by contrast, assert the result without this caveat. The qualification should be reflected in the main claims.","section":"Conclusion"}],"recommendation":"reject","confidential_remarks":"The central formula of the paper fails an elementary Markov-limit test, so I recommend rejection. The discrete calculation in Proposition 3.1 may be salvageable as a separate contribution if the continuous limit is proved carefully and the optimal prior is derived with respect to the controlled state, but that is a substantial reworking rather than a local revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The problem is real: averaged Schrodinger bridge over a parameter ensemble with a parameter-independent control, which produces a non-Markov averaged process, is worth working on. The pinned-control formula in Theorem 3.1, stochastic feedforward for a fixed endpoint, is a plausible new object and the A=0 check is a good sanity test. Section 2 is a known Markov result, but it is cleanly rederived and sets up the path-integral style. The discrete Proposition 3.1 is worked out in detail and the linear-in-noise structure of the local control is believable.\n\nThe bad news is that Theorem 4.1 does not hold as stated. The stress-test example is decisive: take A=0, B=1, eps=1, tf=1, x0=0, mu0=delta_0, muf=N(0,1). Then phi_f=1, u=0 is feasible with cost 0 and is optimal. Plugging into (4.1)-(4.3) gives u*(t|0) = -∫_0^t dW_tau/(1-tau) + W_t, which is not zero and has divergent L2 norm. The paper's claim that qhat reduces to the Markov density in this limit does not fix the problem: the feedforward pinned control is not a function of the current state, and the singular stochastic integral term does not cancel when averaged over the passive conditional density. The proof of Theorem 4.1 switches from the controlled-state density in (4.14) to the passive noise-path density in (4.18) without justification, and that is the load-bearing step.\n\nThere are also softer issues. The proof of Theorem 3.1 relies on an unproved double limit in a and the time step; the discrete-to-continuous passage may be fixable but is not established. The numerics are qualitative, with no code or reproducible details, so they do not add much evidence. The citation pattern is fine, and the reliance on the authors' prior work for the averaged non-Markov process is reasonable.\n\nVerdict: this deserves a serious referee, not a desk reject, because the problem is meaningful and Section 3 may survive as a contribution. But the central claim of the abstract - the characterization of the optimal control for the general Schrodinger bridge problem - is currently false. The authors need either a corrected prior density, likely conditioning on the controlled state and rewriting the pinned control in state-feedback form, or a substantial narrowing of the claims. I would tell the editor to send it out with a clear expectation of major revision or rejection.","headline":"The pinned-control result in Section 3 is a genuine new idea, but Theorem 4.1 is false as stated: in the A=0 limit it gives a nonzero singular control where u=0 is optimal.","tokens_in":20623,"tokens_out":11066,"would_cite":false,"duration_ms":106195,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93E20","49Q22","60G15"],"pacs":[],"model":"deepseek-v4-flash","headline":"For a Schrödinger bridge over an averaged linear ensemble, the optimal control is non-Markovian: a stochastic feedforward law driven by past and present noise.","keywords":["Schrödinger bridge","stochastic feedforward control","parameter-averaged systems","ensemble control","path integral optimal control","non-Markov Gaussian process","optimal transport","averaged controllability"],"falsifier":"Take the scalar ensemble with $A(\\theta)=-\\theta$, $B(\\theta)=1$, solve the penalized discrete problem with $a\\to\\infty$ and $\\Delta t\\to 0$, and compare the limiting control to (3.29) and the empirical final density to $\\rho_f$; a persistent gap would refute the double-limit premise on which the theorem rests.","tokens_in":19596,"feed_emoji":"🌉","tokens_out":13636,"duration_ms":115939,"temperature":0.7,"pith_summary":"This paper studies a Schrödinger bridge problem in which the dynamics are not a single Markov process but an ensemble of linear systems labeled by a parameter $\\theta$, and the quantity whose distribution matters is the average of the ensemble over $\\theta$. The paper claims that the optimal control steering this averaged ensemble from an initial distribution $\\mu_0$ to a final distribution $\\mu_f$ is a stochastic feedforward control: it depends on the initial condition and on past and present noise, in contrast with the Markov state-feedback laws of classical Schrödinger bridges. The argument expresses the optimal control as an expectation of pinned bridge controls over an optimal prior distribution, a path-integral viewpoint that can accommodate non-Markov averaged dynamics. If the theorem is correct, it gives a principled way to robustly deform probability distributions under parameter uncertainty, with direct consequences for ensemble control and optimal transport.","feed_headline":"Bridge control over parameter ensembles remembers past noise","feed_subtitle":"One control steers a whole parameter family; optimality requires noise-history and initial-state memory.","key_machinery":"The engine of the argument is the pinned-bridge averaging identity: the global optimal control is a path integral of local controls for bridges pinned at each possible final state, weighted by an optimal prior distribution. In the averaged setting the local pinned control (3.29) is explicitly non-Markovian, $$u^*(t|x_0,x_f)=-\\sqrt{\\epsilon}\\int_0^t \\Phi(t_f,t)^T G_{t_f,\\tau}^{-1}\\Phi(t_f,\\tau)\\,dW(\\tau)+\\Phi(t_f,t)^T G_{t_f,0}^{-1}\\left(x_f-\\left(\\$int_0^{1}$ $e^{{A(\\theta)t_f}}$d\\$\\theta$\\right)x_0\\right),$$ and the prior $p^*(t_f,x_f|0,x_0,t)$ is formed from the conditional transition density of the passive averaged process, which conditions on the history of the noise rather than only on the current state. Invertibility of the averaged Gramian $G_{t_f,t}$ is the condition that makes the pinned bridges nondegenerate and supplies the averaged observability needed for the construction.","core_discovery":"The paper sets out to establish that, for a Schrödinger bridge over an averaged linear ensemble, the optimal control is a stochastic feedforward law rather than a Markov state feedback. The central statement is Theorem 4.1: if the averaged controllability Gramian $G_{t_f,t}=\\int_t^{t_f}\\Phi(t_f,\\tau)\\Phi(t_f,\\tau)^T d\\tau$ is invertible for every $0\\le t<t_f$, where $\\Phi(t_f,\\tau)=\\int_0^1 e^{A(\\theta)(t_f-\\tau)}B(\\theta)\\,d\\theta$, then the optimal control for problem (1.5)-(1.6) is $$u^*(t|x_0)=\\int_{\\mathbb{R}^d}u^*(t|x_0,x_f)\\,p^*(t_f,x_f|0,x_0,t)\\,dx_f.$$ Here $u^*(t|x_0,x_f)$ is the pinned bridge control (3.29), a feedforward functional of the noise history, and $p^*$ is the optimal prior distribution built from the conditional transition density $\\hat q_{\\epsilon,G}$, which depends on the past noise $\\sqrt{\\epsilon}\\int_0^t\\Phi(t_f,\\tau)\\,dW(\\tau)$. The paper argues that this representation reduces to the classical Markov bridge when the parameters are absent, and that the noise-history dependence is forced by the fact that the averaged process is non-Markov.","pith_inferences":["Extension: the paper leaves the numerical evaluation of the path integral to future work; a direct test is to approximate $p^*$ by Monte Carlo sampling over noise trajectories and compare the resulting histogram to $\\rho_f$.","Extension: classical Schrödinger bridges are solved by forward-backward equations in the current state, but the noise-history dependence in $p^*$ suggests that online deployment would require smoothing or particle-filter methods over the noise path rather than a state feedback law.","Extension: the result implies that memory-based or recurrent policies are natural for ensemble reinforcement learning problems, since the optimal action at time $t$ is not a function of the current averaged state alone.","Extension: the authors explicitly restrict to the case where control and noise enter through the same matrix $B(\\theta)$; a natural follow-up is to test whether the pinned-bridge averaging formula persists when the control and noise channels differ."],"forward_implications":["If Theorem 4.1 is correct, Schrödinger bridges for parameter-averaged ensembles cannot be solved by Markov state feedback; the optimal controller must track the noise history because the control depends on $\\int_0^t\\Phi(t_f,\\tau)\\,dW(\\tau)$.","The path-integral formula yields a concrete algorithm: compute $\\Phi$ and $G$, solve the Schrödinger system for $\\varphi_0,\\varphi_f$, build pinned feedforward controls, then average them over the optimal prior; the numerical examples indicate that this steers the initial density to the target density.","The control cost is proportional to the relative entropy between the controlled averaged law and the passive averaged law, so the bridge can be read as the most likely robust deformation of $\\mu_0$ into $\\mu_f$ under parameter perturbation.","Invertibility of $G_{t_f,t}$ plays the role of controllability for the averaged system; without it, neither the pinned controls nor the optimal prior are well defined.","Because the control is parameter-independent and uses only the averaged process's noise realization, the same input is consistent across all members of the ensemble, which is what makes the steering robust to parameter perturbations."],"supporting_citations":[{"why":"Defines the averaged non-Markov process and its controllability Gramian, which the bridge construction starts from.","marker":"[18]"},{"why":"Provides the optimal pinned bridge for a linear Markov system, the object generalized to the averaged ensemble.","marker":"[8]"},{"why":"Gives the Markov optimal-control representation used to verify the path-integral formula in the single-process case.","marker":"[10]"},{"why":"Supplies the form of the optimal controlled prior distribution that Theorem 4.1 relies on.","marker":"[45]"},{"why":"Reduces the Schrödinger bridge to the static problem and the $\\varphi_0,\\varphi_f$ factorization used throughout.","marker":"[4]"},{"why":"Introduces the path-integral method for optimal control that the paper extends to the non-Markov averaged setting.","marker":"[25]"},{"why":"Establishes the averaged observability condition equivalent to invertibility of $G_{t_f,t}$.","marker":"[15]"},{"why":"Underlies the Girsanov change of measure and the convergence argument for the discretized controls.","marker":"[22]"}],"fun_headline_variants":["Optimal bridge control uses noise history, not just state","Stochastic feedforward control for Schrödinger bridge ensembles","Ensemble bridge control demands past noise memory","Non-Markovian steering for averaged Schrödinger bridges"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the discrete penalized controls converge to the exact constrained bridge as the penalty $a$ tends to infinity and the time step tends to zero, and that the standard formula for the optimal prior distribution, known for memoryless bridges, also holds for the non-Markov averaged process; if either fails, the theorem's control formula is not established.","fun_headline_variants_meta":{"raw":{"variants":["Optimal bridge control uses noise history, not just state","Stochastic feedforward control for Schrödinger bridge ensembles","Ensemble bridge control demands past noise memory","Non-Markovian steering for averaged Schrödinger bridges"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1334,"prompt_tokens":966,"completion_tokens":368,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":305}},"tokens_in":582,"tokens_out":368,"duration_ms":3908,"temperature":1.0,"reasoning_tokens":305,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:34:13.769551+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the scalar ensemble with $A(\\theta)=-\\theta$, $B(\\theta)=1$, solve the penalized discrete problem with $a\\to\\infty$ and $\\Delta t\\to 0$, and compare the limiting control to (3.29) and the empirical final density to $\\rho_f$; a persistent gap would refute the double-limit premise on which the theorem rests.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the averaged non-Markov process and its controllability Gramian, which the bridge construction starts from."},{"cited_title":"Chen, G Tryphon, ”Stochastic bridges of linear systems,” IEEE Transactions on Automatic Con- trol, vol","cited_arxiv_id":null,"evidence_quote":"Provides the optimal pinned bridge for a linear Markov system, the object generalized to the averaged ensemble."},{"cited_title":"Optimal transport ov er a linear dynamical system,","cited_arxiv_id":null,"evidence_quote":"Gives the Markov optimal-control representation used to verify the path-integral formula in the single-process case."},{"cited_title":"Todorov, ”Eﬃcient computation of optimal actions,” Proceedings of the national academy of sci- ences, vol 106, no 28, pp 11478–11483, 2009","cited_arxiv_id":null,"evidence_quote":"Supplies the form of the optimal controlled prior distribution that Theorem 4.1 relies on."},{"cited_title":"J Kappen, ”Path integrals and symmetry breaking for optima l control theory,”Journal of statistical mechanics: theory and experiment , vol 2005, no","cited_arxiv_id":null,"evidence_quote":"Introduces the path-integral method for optimal control that the paper extends to the non-Markov averaged setting."},{"cited_title":"Optimal transport for averaged control,","cited_arxiv_id":null,"evidence_quote":"Establishes the averaged observability condition equivalent to invertibility of $G_{t_f,t}$."},{"cited_title":"Øksendal, ”Stochastic diﬀerential equations,” Springer, 20 03","cited_arxiv_id":null,"evidence_quote":"Underlies the Girsanov change of measure and the convergence argument for the discretized controls."}],"review_version":1}