Pith. sign in

REVIEW 3 major objections 4 minor 48 references

Schrodinger Bridge over Averaged Systems

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read For a Schrödinger bridge over an averaged linear ensemble, the optimal control is non-Markovian: a stochastic feedforward law driven by past and present noise.

desk verdict The pinned-control result in Section 3 is a genuine new idea, but Theorem 4.1 is false as stated: in the A=0 limit it gives a nonzero singular control where u=0 is optimal. read the letter →

arxiv 2412.03294 v1 pith:QCCT5472 submitted 2024-12-04 math.OC

classification math.OC MSC 93E2049Q2260G15
keywords Schrödingerbridgestochasticfeedforwardcontrolparameter-averagedsystemsensemblepathintegraloptimalnon-MarkovGaussianprocesstransportaveragedcontrollability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies a Schrödinger bridge problem in which the dynamics are not a single Markov process but an ensemble of linear systems labeled by a parameter $\theta$, and the quantity whose distribution matters is the average of the ensemble over $\theta$. The paper claims that the optimal control steering this averaged ensemble from an initial distribution $\mu_0$ to a final distribution $\mu_f$ is a stochastic feedforward control: it depends on the initial condition and on past and present noise, in contrast with the Markov state-feedback laws of classical Schrödinger bridges. The argument expresses the optimal control as an expectation of pinned bridge controls over an optimal prior distribution, a path-integral viewpoint that can accommodate non-Markov averaged dynamics. If the theorem is correct, it gives a principled way to robustly deform probability distributions under parameter uncertainty, with direct consequences for ensemble control and optimal transport.

What carries the argument

The engine of the argument is the pinned-bridge averaging identity: the global optimal control is a path integral of local controls for bridges pinned at each possible final state, weighted by an optimal prior distribution. In the averaged setting the local pinned control (3.29) is explicitly non-Markovian, $$u^*(t|x_0,x_f)=-\sqrt{\epsilon}\int_0^t \Phi(t_f,t)^T G_{t_f,\tau}^{-1}\Phi(t_f,\tau)\,dW(\tau)+\Phi(t_f,t)^T G_{t_f,0}^{-1}\left(x_f-\left(\$int_0^{1}$ $e^{{A(\theta)t_f}}$d\$\theta$\right)x_0\right),$$ and the prior $p^*(t_f,x_f|0,x_0,t)$ is formed from the conditional transition density of the passive averaged process, which conditions on the history of the noise rather than only on the current state. Invertibility of the averaged Gramian $G_{t_f,t}$ is the condition that makes the pinned bridges nondegenerate and supplies the averaged observability needed for the construction.

What would settle it

Take the scalar ensemble with $A(\theta)=-\theta$, $B(\theta)=1$, solve the penalized discrete problem with $a\to\infty$ and $\Delta t\to 0$, and compare the limiting control to (3.29) and the empirical final density to $\rho_f$; a persistent gap would refute the double-limit premise on which the theorem rests.

Watch

Extended reading notes

Core claim

The paper sets out to establish that, for a Schrödinger bridge over an averaged linear ensemble, the optimal control is a stochastic feedforward law rather than a Markov state feedback. The central statement is Theorem 4.1: if the averaged controllability Gramian $G_{t_f,t}=\int_t^{t_f}\Phi(t_f,\tau)\Phi(t_f,\tau)^T d\tau$ is invertible for every $0\le t<t_f$, where $\Phi(t_f,\tau)=\int_0^1 e^{A(\theta)(t_f-\tau)}B(\theta)\,d\theta$, then the optimal control for problem (1.5)-(1.6) is $$u^*(t|x_0)=\int_{\mathbb{R}^d}u^*(t|x_0,x_f)\,p^*(t_f,x_f|0,x_0,t)\,dx_f.$$ Here $u^*(t|x_0,x_f)$ is the pinned bridge control (3.29), a feedforward functional of the noise history, and $p^*$ is the optimal prior distribution built from the conditional transition density $\hat q_{\epsilon,G}$, which depends on the past noise $\sqrt{\epsilon}\int_0^t\Phi(t_f,\tau)\,dW(\tau)$. The paper argues that this representation reduces to the classical Markov bridge when the parameters are absent, and that the noise-history dependence is forced by the fact that the averaged process is non-Markov.

Load-bearing premise

The load-bearing premise is that the discrete penalized controls converge to the exact constrained bridge as the penalty $a$ tends to infinity and the time step tends to zero, and that the standard formula for the optimal prior distribution, known for memoryless bridges, also holds for the non-Markov averaged process; if either fails, the theorem's control formula is not established.

Editorial extensions

If this is right

  • If Theorem 4.1 is correct, Schrödinger bridges for parameter-averaged ensembles cannot be solved by Markov state feedback; the optimal controller must track the noise history because the control depends on $\int_0^t\Phi(t_f,\tau)\,dW(\tau)$.
  • The path-integral formula yields a concrete algorithm: compute $\Phi$ and $G$, solve the Schrödinger system for $\varphi_0,\varphi_f$, build pinned feedforward controls, then average them over the optimal prior; the numerical examples indicate that this steers the initial density to the target density.
  • The control cost is proportional to the relative entropy between the controlled averaged law and the passive averaged law, so the bridge can be read as the most likely robust deformation of $\mu_0$ into $\mu_f$ under parameter perturbation.
  • Invertibility of $G_{t_f,t}$ plays the role of controllability for the averaged system; without it, neither the pinned controls nor the optimal prior are well defined.
  • Because the control is parameter-independent and uses only the averaged process's noise realization, the same input is consistent across all members of the ensemble, which is what makes the steering robust to parameter perturbations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: the paper leaves the numerical evaluation of the path integral to future work; a direct test is to approximate $p^*$ by Monte Carlo sampling over noise trajectories and compare the resulting histogram to $\rho_f$.
  • Extension: classical Schrödinger bridges are solved by forward-backward equations in the current state, but the noise-history dependence in $p^*$ suggests that online deployment would require smoothing or particle-filter methods over the noise path rather than a state feedback law.
  • Extension: the result implies that memory-based or recurrent policies are natural for ensemble reinforcement learning problems, since the optimal action at time $t$ is not a function of the current averaged state alone.
  • Extension: the authors explicitly restrict to the case where control and noise enter through the same matrix $B(\theta)$; a natural follow-up is to test whether the pinned-bridge averaging formula persists when the control and noise channels differ.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies a Schrödinger bridge problem for a parameter-averaged stochastic system. Individual systems evolve as dX(t,θ)=(A(θ)X(t,θ)+B(θ)u(t))dt+√εB(θ)dW(t), and the object to be controlled is the averaged process x(t)=∫_0^1 X(t,θ)dθ. The authors argue that the optimal control for this non-Markov averaged problem is a stochastic feedforward control. Section 2 reviews the single Markov case and gives a path-integral representation. Section 3 derives a pinned local control u*(t|x0,xf) through a penalized discrete optimal control problem, and Section 4 proposes a general formula u*(t|x0)=∫u*(t|x0,xf)p*(tf,xf|0,x0,t)dxf, where p* is built from a conditional transition density of the passive averaged process. Numerical examples in Section 5 simulate the proposed controller.

Significance. If Theorem 4.1 were valid, the paper would be the first to present a path-integral, non-Markov stochastic feedforward control for an averaged Schrödinger bridge problem, and the absence of fitted parameters in the control law would be a strength. The discrete calculation in Proposition 3.1 is explicit and appears to be carried out in detail. However, the central theorem is not merely incompletely proved: as stated it is false in a simple Markov limit. This prevents the paper from achieving its advertised significance in its current form.

major comments (3)
  1. [Section 4, Theorem 4.1] Theorem 4.1 is false as stated. Consider A(θ)=0, B(θ)=1, ε=1, tf=1, x0=0, μ0=δ0, μf=N(0,1). Then φf=1 solves (2.9), and u=0 is feasible with cost 0, hence optimal. However, substituting the data into (4.1)-(4.3) gives u*(t|0)=-∫_0^t (1-τ)^{-1}dW(τ)+W(t), which is nonzero and has infinite L2 norm on [0,1]. The root cause is that (4.18) conditions on the passive noise path {W(s):0≤s≤t}, whereas the optimal prior defined in (4.14) is a ratio of densities of the controlled process (4.6) and (4.8). These are different objects unless the control is zero, and the proof never establishes that (4.14) equals (4.18).
  2. [Section 3, proof of Theorem 3.1] The passage from the discrete penalized control (3.15) to the continuous hard-constrained control (3.4) is asserted without proof. Immediately after (3.15), the paper states that [22, Definition 3.1.6 and Corollary 3.1.8] imply the double limit lim_{a→∞}lim_{k→∞}||u*_{a,k,i}-u*||^2_{L2(P)}=0, but no argument is given for the convergence of the discrete adapted controls to the Itô integral in (3.4), for the removal of the 1/(2a) regularization, or for the interchange of the limits. Since Theorem 3.1 is the foundation for the pinned controls used in Theorem 4.1, this gap is load-bearing.
  3. [Section 4, proof of Theorem 4.1] The argument from (4.13) to (4.16) does not establish the claimed equality even if the prior distribution were correct. A process Z(t) with E[Z(t)]=0 can be nonzero, and the statement that the 'extra sample path' ∫_0^t Φ(t,τ)Z(τ)dτ contradicts optimality is not a rigorous uniqueness proof for the non-Markov stochastic feedforward setting. This step is presented as if it were a contradiction, but no uniqueness theorem for the optimal control is stated or proved.
minor comments (4)
  1. [Section 5, Eq. (5.7)] Equation (5.7) states ∫_0^1 e^{-θ}dθ = 1 - e; the correct value is 1 - e^{-1}, and this appears to be a typo.
  2. [Section 4, Eq. (4.2) and (4.3)] The notation for the conditional density is inconsistent: (4.2) writes q-hat with arguments (0,x0,t,tf,y), while (4.3) uses the same symbol with a different ordering; the meaning of the five arguments should be stated explicitly.
  3. [Section 5] The numerical section is illustrative rather than quantitative: no cost values, no comparison with a baseline, and no convergence assessment are reported for the Monte Carlo estimate in Figure 7.
  4. [Conclusion] The conclusion describes the solution as 'formally solved' and identifies the computation of the path integral as future work; the abstract and Theorem 4.1, by contrast, assert the result without this caveat. The qualification should be reflected in the main claims.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: Theorem 4.1's control formula is an explicit expression in the problem data, not a fit; the paper's self-citations support background conditions but do not force the result.

full rationale

The claimed derivation is not circular in the sense of the analysis. The main formula (4.1) is an integral over the explicitly computed pinned control (3.29) and the conditional density (4.18); both are constructed from A(theta), B(theta), the Gramian G, and the solutions phi0, phi_f of the integral equations (2.9). No fitted parameter is renamed as a prediction. The proof does import the Markov h-transform characterization of the optimal prior from [45] for the Markov case (Section 2, 'Following from [45], ... characterized as (2.8)') and then extends it to the non-Markov averaged process without proof by writing the passive conditional density (4.18). This is an unproved, load-bearing assumption and a likely source of the mathematical error identified in the skeptic's counterexample, but it is not a circularity: the output is not equal to its input by construction. The double limit in Theorem 3.1 ('we conclude that lim_{a->infinity} lim_{k->infinity} ... = 0') is likewise an asserted convergence rather than a derived limit; this is an omitted proof, not a circular step. The self-citations [15] and [18] are used to state the averaged observability equivalence and the non-Markov property; these are background facts for the assumptions and do not define the target control in terms of itself. The paper itself flags that the path integral remains to be computed (Section 6), which is a limitation of tractability, not evidence of circularity. Accordingly, score 2 reflects minor self-citation with independent central content.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The main theorem rests on structural assumptions about the averaged system and on two unproved analytic steps. No numerical constants are fitted to data; epsilon is the noise intensity and the densities rho0, rhof are problem data, not free parameters.

assumptions (5)
  • domain assumption G_{tf,t} = integral from t to tf of Phi(tf,tau)Phi(tf,tau)^T dtau is invertible for all 0 <= t < tf.
    Required for Eq. (1.11) and for the pinned control formula in Theorem 3.1; inherited from the averaged observability inequality literature [15,34-37].
  • domain assumption The ensemble state is the uniform average x(t) = integral over theta in [0,1] of X(t,theta) dtheta, and theta is only a parameter index.
    This defines the problem in (1.6) and is what makes the averaged process non-Markovian.
  • standard math Girsanov's theorem and the path integral cost identity D(Px || Ry) = E[integral (1/(2 epsilon)) ||u||^2 dt] apply to the averaged process in (1.12)-(1.14).
    Used to connect the control problem to the Schrodinger bridge formulation in Section 1.
  • ad hoc to paper The optimal controlled prior distribution p* in the non-Markov case takes the same h-transform form as the Markov result in [45].
    Used without proof in Theorem 4.1, Eq. (4.2); this transfers a Markov result to a non-Markov Gaussian process.
  • ad hoc to paper The double limit lim_{a to infinity} lim_{k to infinity} of the penalized discrete optimal controls recovers the hard-constrained continuous optimal control.
    Used in the proof of Theorem 3.1; the convergence is asserted with a citation to Oksendal definitions, not proven.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Schrodinger Bridge over Averaged Systems." pith.science (2026). https://pith.science/paper/QCCT5472

@misc{pith2026241203294,
  author       = {Pith},
  title        = {Pith review of: Schrodinger Bridge over Averaged Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QCCT5472}},
  note         = {Machine review of arXiv:2412.03294}
}
read the original abstract

We consider a Schr\"odinger bridge problem where the Markov process is subject to parameter perturbations, forming an ensemble of systems. Our objective is to steer this ensemble from the initial distribution to the final distribution using controls robust to the parameter perturbations. Utilizing the path integral formalism, we demonstrate that the optimal control is a non-Markovian strategy, specifically a stochastic feedforward control, which depends on past and present noise. This unexpected deviation from established strategies for Schr\"odinger bridge problems highlights the intricate interrelationships present in the system's dynamics. From the perspective of optimal transport, a significant by-product of our work is the demonstration that, when the evolution of a distribution is subject to parameter perturbations, it is possible to robustly deform the distribution to a desired final state using stochastic feedforward controls.

Figures

Figures reproduced from arXiv: 2412.03294 by the authors.

Figure 1
Figure 1. This is the plot of initial and final marginal non-Gaussian densities ρ0 and ρf defined in (5.2)- (5.3). This is useful for computing φ0 and φf in (2.9). hence G −1 1,τ = ˆ 1 τ Φ(1, t)Φ(1, t) T dt−1 = ˆ 1 τ 2 − 2 cos(1 − t) (1 − t) 2 dt−1 " 1 0 0 1# (5.10) is well-defined for all τ ∈ [0, 1). Therefore, from Theorem (4.1), we have that the optimal control for problem (1.5)-(1.6) where (5.8) is (4.1) where u ∗ (t|… view at source ↗
Figure 2
Figure 2. This is the plot of φ0 and φf computed from (2.9), where ρ0 and ρf are given in (5.2)-(5.3) with [PITH_FULL_IMAGE:figures/full_fig_p023_2.png] view at source ↗
Figure 3
Figure 3. This is the plot of 10 sample paths characterizing the one￾dimensional optimal local control process u ∗ (t|0, 1) in (3.29) with parameters in (5.4)-(5.7). such as existence of barrier, non-linearity and extend other data-driven problems associated to the classical Schrodinger bridge problem to our case etc. References [1] E. Schr¨odinger, Uber die Umkehrung der Naturgesetze ¨ . Verlag der Akademie der Wissenschafte… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: This is the plot of 10 sample paths characterizing the one￾dimensional optimal control process u ∗ (t|0) in (4.1) using the pinned control process in [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 5
Figure 5. Figure 5: This is the plot of 10 sample paths characterizing the one￾dimensional optimal averaged pinned state process x ∗ (t|0, 1) in (4.8) induced by the optimal local control process in [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]
Figure 6
Figure 6. Figure 6: This is the plot of 10 sample paths characterizing the one dimen￾sional optimal averaged state process x ∗ (t|0) in 4.6 induced by the optimal pinned control in [PITH_FULL_IMAGE:figures/full_fig_p027_6.png]
Figure 7
Figure 7. Figure 7: This is the Monte Carlo simulation that estimates ρf in [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: This is the plot 10 sample paths characterizing the two-dimensional optimal local control process u ∗ (t|x0, xf ) in (3.29) with parameters in (5.9)- (5.10) and (5.11). [32] A. J. Bray, A. J. McKane and T. J. Newman, ”Path integrals and non-Markov processes. II. Escape…
Figure 9
Figure 9. Figure 9: This is the plot of 10 sample paths characterizing the optimal averaged pinned state process x ∗ (t|x0, xf ) in (4.8) induced by the optimal local control in Figure (8). The lower black star indicated the initial point [1, 0]T and the upper black star indicates the poi…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 47 canonical work pages

  1. [1]

    Schr¨ odinger, ¨Uber die Umkehrung der Naturgesetze

    E. Schr¨ odinger, ¨Uber die Umkehrung der Naturgesetze . Verlag der Akademie der Wissenschaften in Kommission bei Walter De Gruyter u . . . , 1931

  2. [2]

    F¨ ollmer, ”Random fields and diffusion processes,” Lect

    H. F¨ ollmer, ”Random fields and diffusion processes,” Lect. Notes Math , vol. 1362, pp. 101–204, 1988

  3. [3]

    Y. Chen, T. T. Georgiou and M. Pavon, ”Optimal transport in sys tems and control,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 4, no. 1, pp. 89–113, 2021

  4. [4]

    L´ eonard, ”A survey of the Schrodinger problem and some of its connections with optimal trans- port,”arXiv preprint arXiv:1308.0215 , 2013

    C. L´ eonard, ”A survey of the Schrodinger problem and some of its connections with optimal trans- port,”arXiv preprint arXiv:1308.0215 , 2013

  5. [5]

    Bernstein, ”On the connections between random quantities,” Verh

    S. Bernstein, ”On the connections between random quantities,” Verh. Boarding school. Math.-Kongr., Zurich, vol. 1, pp. 288–309, 1932

  6. [6]

    Dai Pra, ”A stochastic control approach to reciprocal diffu sion processes,” Applied Mathematics and Optimization , vol

    P. Dai Pra, ”A stochastic control approach to reciprocal diffu sion processes,” Applied Mathematics and Optimization , vol. 23, no. 1, pp. 313–329, 1991. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 -3 -2 -1 0 1 2 3 4 5 6 Figure 4. This is the plot of 10 sample paths characterizing the one- dimensional optimal control process u∗(t|0) in (4.1) using the pinned c...

  7. [7]

    D. P. Pra., M. Pavon,” On the Markov processes of Schr¨ odinger , the Feynman-Kac formula and stochastic control”, Realization and Modelling in System Theory: Proceedings of the International Symposium MTNS-89, Volume I , no. 4, pp. 497–504, 1990

  8. [8]

    Chen, G Tryphon, ”Stochastic bridges of linear systems,” IEEE Transactions on Automatic Con- trol, vol

    Y. Chen, G Tryphon, ”Stochastic bridges of linear systems,” IEEE Transactions on Automatic Con- trol, vol. 61, no. 2, pp. 526–531, 2015

Show all 48 references
  1. [9]

    Optimal steering of a line ar stochastic system to a final probability distribution, part i,

    Y. Chen, T. T. Georgiou, and M. Pavon, “Optimal steering of a line ar stochastic system to a final probability distribution, part i,” IEEE Transactions on Automatic Control , vol. 61, no. 5, pp. 1158– 1169, 2015

  2. [10]

    Optimal transport ov er a linear dynamical system,

    Y. Chen, T. T. Georgiou, and M. Pavon, “Optimal transport ov er a linear dynamical system,” IEEE Transactions on Automatic Control , vol. 62, no. 5, pp. 2137–2152, 2016

  3. [11]

    Cuay´ ahuitl, D

    H. Cuay´ ahuitl, D. Lee, S. Ryu, Y. Cho, S. Choi, S. Indurthi, an d S. Yu, H. Choi, I. Hwang, and J. Kim, ”Ensemble-based deep reinforcement learning for chat bots,” Neurocomputing,vol 366, pp. 118–130,2019. 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 -0.2 0 0.2 0.4 0.6 0.8 1 1.2 ...

  4. [12]

    Nauta, Y

    J. Nauta, Y. Khaluf, and P. Simoens,”Using the Ornstein-Uhlenb eck process for random explo- ration”4th International Conference on Complexity, Future Inform ation Systems and Risk (COM- PLEXIS2019), pp 59–66, 2019

  5. [13]

    Ensemble control of linear systems,

    J.-S. Li and N. Khaneja, “Ensemble control of linear systems,” in 2007 46th IEEE Conference on Decision and Control , pp. 3768–3773, IEEE, 2007

  6. [14]

    Li, ”Ensemble control of finite-dimensional time-varying lin ear systems,” IEEE Transactions on Automatic Control, vol

    J.-S. Li, ”Ensemble control of finite-dimensional time-varying lin ear systems,” IEEE Transactions on Automatic Control, vol. 56, no. 2, pp. 345–357, 2010

  7. [15]

    Optimal transport for averaged control,

    D. O. Adu, “Optimal transport for averaged control,” IEEE Control Systems Letters , vol. 7, pp. 727– 732, 2022

  8. [16]

    Qi, Ji and A

    J. Qi, Ji and A. Zlotnik, Anatoly and J.-S. Li,”Optimal ensemble con trol of stochastic time-varying linear systems.” Systems & Control Letters , vol 62, no 11, pp. 1057–1064, 2013

  9. [17]

    Brockett and N

    R. Brockett and N. Khaneja, ”On the stochastic control of q uantum ensembles,” System Theory: Modeling, Analysis and Control , pp. 75–96, 2000

  10. [18]

    D. O. Adu and Y. Chen, ”Stochastic bridges over ensemble of line ar systems,” 2023 62nd IEEE Conference on Decision and Control (CDC) , pp. 2803–2808, 2023 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 -0.2 -0.1 0 0.1 0.2 0.3 0.4 Figure 6. This is the plot of 10 sample paths charact...

  11. [19]

    Nelson, ” Dynamical theories of Brownian motion, ” Princeton University Press, vol

    E. Nelson, ” Dynamical theories of Brownian motion, ” Princeton University Press, vol. 17, 1967

  12. [20]

    R. V. Handel, ”Stochastic calculus, filtering, and stochastic co ntrol,” Course notes., URL http://www. princeton. edu/rvan/acm217/ACM217. pdf , vol. 14, 2007

  13. [21]

    F. C. Klebaner, ”Introduction to stochastic calculus with applic ations”, World Scientific Publishing Company,2012

  14. [22]

    Øksendal, ”Stochastic differential equations,” Springer, 20 03

    B. Øksendal, ”Stochastic differential equations,” Springer, 20 03

  15. [23]

    P Feynman and Jr

    R. P Feynman and Jr. F. L Vernon, ”The theory of a general qu antum system interacting with a linear dissipative system,” Annals of physics , vol 281, no. 1-2 pp 547–607, 2000

  16. [24]

    P Feynman, A

    R. P Feynman, A. R. Hibbs and D. F. Styer, ”Quantum mechanics and path integrals,” Courier Corporation, 2010

  17. [25]

    J Kappen, ”Path integrals and symmetry breaking for optima l control theory,”Journal of statistical mechanics: theory and experiment , vol 2005, no

    H. J Kappen, ”Path integrals and symmetry breaking for optima l control theory,”Journal of statistical mechanics: theory and experiment , vol 2005, no. 11 pp 11011, 2005

  18. [26]

    J Kappen, ”Linear theory for control of nonlinear stochas tic systems,” Physical review letters , vol 95, no

    H. J Kappen, ”Linear theory for control of nonlinear stochas tic systems,” Physical review letters , vol 95, no. 20 pp 200201, 2005 Figure 7. This is the Monte Carlo simulation that estimates ρf in Figure 1 using 50 states of x0 randomly initialized from the support of ρ0 in 1...

  19. [27]

    J Kappen, ”An introduction to stochastic control theory, path integrals and reinforcement learn- ing,”AIP conference proceedings, vol 887, no

    H. J Kappen, ”An introduction to stochastic control theory, path integrals and reinforcement learn- ing,”AIP conference proceedings, vol 887, no. 1 pp 149–181, 2007

  20. [28]

    J Kappen, W

    H. J Kappen, W. Wiegerinck and B. Van Den Broek, ”A path integr al approach to agent plan- ning,”Autonomous Agents and Multi-Agent Systems , pp 41, 2007

  21. [29]

    Van Den Broek, W

    B. Van Den Broek, W. Wiegerinck and H. J Kappen, ”Graphical mo del inference in optimal control of stochastic multi-agent systems,” Journal of Artificial Intelligence Research , vol 32, pp 95–122, 2008

  22. [30]

    Wiegerinck, B

    W. Wiegerinck, B. Van Den Broekand H. J Kappen, ”Stochastic o ptimal control in continuous space- time multi-agent systems,” arXiv preprint arXiv:1206.6866 , 2012

  23. [31]

    A. J. McKane, H. C. Luckock and A. J. Bray, ”Path integrals an d non-Markov processes. I. General formalism,”Physical Review A , vol 41, no 2, pp 644, 1990. Figure 8. This is the plot 10 sample paths characterizing the two-dimensional optimal local control process u∗(t|x0, xf ...

  24. [32]

    A. J. Bray, A. J. McKane and T. J. Newman, ”Path integrals and non-Markov processes. II. Escape rates and stationary distributions in the weak-noise limit,” Physical Review A , vol 41, no 2, pp 657, 1990

  25. [33]

    H. C. Luckock and A. J. McKane, ”Path integrals and non-Mark ov processes. III. Calculation of the escape-rate prefactor in the weak-noise limit,” Physical Review A , vol 42, no 4, pp 1982, 1990

  26. [34]

    Averaged control and observation of parameter-depending wave equations,

    M. Lazar and E. Zuazua, “Averaged control and observation of parameter-depending wave equations,” Comptes Rendus Mathematique , vol. 352, no. 6, pp. 497–502, 2014

  27. [35]

    Averaged controllability of parameter dependent conservative semigroups,

    J. Loh´ eac and E. Zuazua, “Averaged controllability of parameter dependent conservative semigroups,” Journal of Differential Equations , vol. 262, no. 3, pp. 1540–1574, 2017

  28. [36]

    Averaged controllability for random evo lution partial differential equations,

    Q. L¨ u and E. Zuazua, “Averaged controllability for random evo lution partial differential equations,” Journal de Math´ ematiques Pures et Appliqu´ ees, vol. 105, no. 3, pp. 367–414, 2016

  29. [37]

    Averaged control,

    E. Zuazua, “Averaged control,” Automatica, vol. 50, no. 12, pp. 3077–3087, 2014

  30. [38]

    Halyo,”A combined stochastic feedforward and feedback co ntrol design methodology with appli- cation to autoland design,” 1987

    N. Halyo,”A combined stochastic feedforward and feedback co ntrol design methodology with appli- cation to autoland design,” 1987

  31. [39]

    P. S. Maybeck,”Stochastic models, estimation, and control”,19 82 Figure 9. This is the plot of 10 sample paths characterizing the optimal averaged pinned state process x∗(t|x0, xf ) in (4.8) induced by the optimal local control in Figure (8). The lower black star indicated the...

  32. [40]

    Halyo, H

    N. Halyo, H. Direskeneli, and D. B. Taylor, ”A stochastic optimal feedforward and feedback control methodology for superagility,” 1992

  33. [41]

    M. E. Halpen and Aeronautical Research Labs Melbourne (Aust ralia),”Application of Optimal Track- ing Methods to Aircraft Terrain Following,” 1989

  34. [42]

    Optimal Transport f or a Class of Linear Quadratic Dif- ferential Games,

    D. O. Adu, T. Basar and B. Gharesifard, “Optimal Transport f or a Class of Linear Quadratic Dif- ferential Games,” IEEE Transactions on Automatic Control , vol. 67, no. 11, pp. 6287-6294, 2022

  35. [43]

    Robust Matching for Teams,

    D. O. Adu, and B. Gharesifard, “Robust Matching for Teams,” Journal of Optimization Theory and Applications, vol. 200, no. 2, pp. 501–523, 2024

  36. [44]

    Y. Chen, T. T. Georgiou and M. Pavan, ”On the relation between optimal transport and Schr¨ odinger bridges: A stochastic control viewpoint,” Journal of Optimization Theory and Applications , vol 169, pp 671–691, 2016

  37. [45]

    Todorov, ”Efficient computation of optimal actions,” Proceedings of the national academy of sci- ences, vol 106, no 28, pp 11478–11483, 2009

    E. Todorov, ”Efficient computation of optimal actions,” Proceedings of the national academy of sci- ences, vol 106, no 28, pp 11478–11483, 2009

  38. [46]

    S. P. Boyd and L. Vandenberghe, Convex Optimization , 2004

  39. [47]

    M. A. Woodbury, Inverting Modified Matrices , 1950

  40. [48]

    H. V. Henderson and S. R. Searle, ”On deriving the inverse of a s um of matrices,” Siam Review , vol 23, no 1, pp 53–60, 1981

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.