REVIEW 2 major objections 5 minor 43 references
Stability properties of gradient flow dynamics for the symmetric low-rank matrix factorization problem
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read For gradient flow on symmetric low-rank factorization, global recovery occurs exactly when the signal block starts positive definite, and excess-parameter noise decays only at rate $O(1/t)$.
desk verdict A genuinely useful cascade decomposition and global stability theorem, but the O(1/t) noise bound rests on a false Riccati solution and needs correction before the paper is citable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the nonlinear change of variables $H_1=P_1^{-1}$, $H_0=P_1^{-1}P_0$, and $H_2=P_2-P_0^TP_1^{-1}P_0$, where $H_2$ is the Schur complement of $P_1$ in $P$. In these coordinates the lifted gradient flow becomes a cascade of three subsystems: $\dot H_2=-\Lambda_2H_2-H_2\Lambda_2-2H_2^2$ is autonomous and decays at $O(1/t)$; $\dot H_0=-\Lambda_1H_0-H_0(\Lambda_2+2H_2)$ is driven by $H_2$ and decays exponentially at rate $\lambda_{r^\star}$; and $\dot H_1=-\Lambda_1H_1-H_1\Lambda_1+2(I+H_0H_0^T)$ is driven by $H_0H_0^T$ and converges to $\Lambda_1^{-1}$. The cascade turns a non-convex matrix flow into a triangular system that can be analyzed one subsystem at a time, with the Schur-complement variable isolating exactly the excess-parameter dynamics.
What would settle it
Numerically integrate the lifted gradient flow $\dot P=(\Lambda-P)P+P(\Lambda-P)$ from a positive-definite $P_1(0)$ whose minimum eigenvalue is very small, and track the minimum eigenvalue of $P_1(t)$; if it reaches zero in finite time, Theorem 1's claim that $P_1(t)\succ 0$ for all times is false. Separately, set $\Lambda_2=0$, $P_1(0)=\Lambda_1$, $P_0(0)=0$, and $P_2(0)\ne 0$; the paper predicts every nonzero eigenvalue of $P_2$ follows $2\mu(0)/(1+t)$, so observing faster decay in that setup would disprove the claimed sharpness of the $O(1/t)$ rate.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a complete stability picture for the lifted gradient flow $\dot P=(\Lambda-P)P+P(\Lambda-P)$. The equilibrium set consists of matrices $\mathrm{diag}(\bar P_1,0)$ with $\bar P_1\Lambda_1=\bar P_1^2$; when the positive eigenvalues of $\Lambda_1$ are distinct, these are exactly the $2^{r^\star}$ diagonal matrices obtained by choosing any subset of those eigenvalues, and the only stable one is $\bar P=\mathrm{diag}(\Lambda_1,0)$. Theorem 1 proves that if $P_1(0)\succ 0$ then $\|P_1(t)-\Lambda_1\|_2\le 2\lambda_1^2\,\ell(t)$ and $\|P_0(t)\|_2\le(\lambda_1+2\lambda_1^2\ell(t))\|P_1^{-1}(0)P_0(0)\|_2\,e^{-\lambda_{r^\star}t}$, while $P_2(t)$ vanishes at the rate $O(1/t)$; conversely, if $P_1(0)$ is singular its null space is invariant, so the global optimum is unreachable. The rate separation is proved sharp: when $\Lambda_2=0$, $P_1(0)=\Lambda_1$, and $P_0(0)=0$, each nonzero eigenvalue of $P_2(t)$ decays exactly as $2\mu(0)/(1+t)$.
Load-bearing premise
The analysis runs in coordinates that require the signal block $P_1(t)$ to remain invertible, and the proof that $P_1(0)\succ 0$ guarantees this forever rests on the invariance of the null space of $P_1$; if such an invariance argument failed, the cascade coordinates and all of the rate bounds would collapse.
Editorial extensions
If this is right
- From any initialization with $P_1(0)\succ 0$, gradient flow recovers the positive part of $M$: $P_1(t)\to\Lambda_1$ and $P_0(t)\to 0$ at exponential rate $\lambda_{r^\star}$, with an error bound for $P_1$ that contains a transient algebraic-growth factor before the exponential decay dominates.
- In the over-parameterized regime $r>r^\star$, the excess-parameter block $P_2$ and the Schur complement $H_2$ cannot decay faster than $O(1/t)$ in general, because an explicit trajectory achieves $2\mu(0)/(1+t)$; the polynomial tail is an intrinsic feature of the dynamics, not an artifact of the proof.
- If $P_1(0)$ is singular, recovery is impossible: the null space of $P_1$ is invariant, so the trajectory is confined to a manifold that does not contain the global optimum.
- In the exactly parameterized case $r=r^\star$, the same analysis upgrades the noise decay to exponential, with $P_2(t)$ vanishing at rate $e^{-2\lambda_{r^\star}t}$.
- Recursively applying the same coordinate transformation separates the distinct eigenspaces of $\Lambda_1$, yielding convergence rates for the coupling blocks controlled by eigenvalue gaps, such as $e^{-(\hat\lambda_i-\hat\lambda_{i+1})t}$.
Reading between the lines
- An implication the paper leaves implicit is that initializations with $P_1(0)$ nearly singular pay a measurable alignment cost: the transient factor $t$ in the $\tilde H_1$ bound delays the onset of the exponential phase, predicting a plateau in the loss that a practitioner could look for in experiments.
- A testable design principle follows: choosing an initialization that makes the minimum eigenvalue of $P_1(0)$ large should shrink both the exponential convergence time and the constants in front of the polynomial tail; the paper itself does not optimize over initializations.
- If the same Schur-complement cascade transfers to the asymmetric factorization and matrix-sensing problems listed as future directions, one would expect the same signal/excess split, with a polynomial slow mode whenever the fitted rank exceeds the target rank.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies continuous-time gradient flow for the symmetric low-rank matrix factorization problem minimize (1/4)||XX^T - M||_F^2. After rotating into the eigenbasis of M, the authors analyze the lifted variable P = ZZ^T, whose dynamics are given by (5). They partition P into signal/noise blocks and introduce a nonlinear change of variables (13), H1 = P1^{-1}, H0 = P1^{-1}P0, H2 = P2 - P0^T P1^{-1}P0, which brings the dynamics into a cascade of three subsystems: H2 evolves autonomously, H0 is driven by H2, and H1 is driven by H0H0^T. The main results are a complete characterization of the equilibrium points (Lemma 1 and Proposition 1), local asymptotic stability of the global minimum (Proposition 2), and a global convergence theorem (Theorem 1) stating that if P1(0) is positive definite, then P1(t) -> Λ1 and P0(t) -> 0 exponentially at rate λ_{r*}, while the excess-parameter component P2(t) decays only as O(1/t) (Lemma 3 and Theorem 2). The paper also sketches a refined eigenvalue-gap decomposition in Remark 5. The central qualitative claim is that over-parameterization causes a slow polynomial tail in the noise component, while the signal component converges exponentially.
Significance. If established, the cascade decomposition is an elegant and potentially reusable structural insight for non-convex matrix factorization dynamics. The paper is self-contained: the main theorems are backed by derivations in the appendices, the statements are explicit, and no constants are fitted to data. The qualitative result that excess parameters create an O(1/t) slow mode, while the signal subspace converges exponentially, is a meaningful and falsifiable contribution to the understanding of over-parameterized gradient dynamics. However, the quantitative O(1/t) bound contains a specific error in the comparison solution used in Lemma 3, and the proof of the exponential H0 decay in Lemma 4 relies on an unproved smoothness assumption. The main structural conclusions appear defensible after correction, but the paper as written overstates a key quantitative rate.
major comments (2)
- [Section III-D and Appendix B (Lemma 3)] The comparison solution used in Lemma 3 is incorrect. For the scalar equation \dot x = -2x^2, the solution is x(t) = x(t0)/(1 + 2x(t0)(t - t0)), not x(t) = 2x(t0)/(1 + t - t0). Substitution shows the latter satisfies \dot x = -2x0/(1+t)^2, whereas -2x^2 = -8x0^2/(1+t)^2, with equality only for x0 = 1/4. Consequently, the stated bound VN(P2(t)) ≤ 2VN(P2(τ))/(1 + t - τ) is false. For example, with Λ2 = 0, P0 = 0 and scalar P2(0) = 0.1, the exact solution is P2(t) = 0.1/(1 + 0.2t); taking τ = 0.001 and t = 10, the claimed upper bound is about 0.0182, while the exact value is about 0.0333. The same incorrect formula appears in the sharpness example in Section III-D, where 'μ(t) = 2μ(0)/(1+t)' is not a solution of \dot μ = -2μ^2. This invalidates the quantitative statement of Lemma 3 and the matching ∥H2(t)∥ bound in Theorem 2, as well as the ∥Z2(t)∥ rate in Corollary 1. The qualitative O(1/t) decay survives: the comparison principle gives VN(t) ≤ VN(τ)/(1 + 2VN(τ)(t - τ)), which is O(1/t) with a constant depending on the initial data. Lemma 3, the sharpness example, Theorem 2's third displayed inequality, and Corollary 1 need to be restated with the corrected bound.
- [Appendix C, Lemma 4] The proof of Lemma 4 asserts that d∥H0∥2/dt = u^T \dot H0 v for the principal singular vectors u and v, attributing this to a 'well-known property of the derivative of singular vectors'. This identity requires the principal singular value to be simple along the trajectory, which is not guaranteed for the flow (14); at points where singular values cross, the principal singular vectors and the derivative of the spectral norm may not be differentiable. Because the exponential rate of H0 in Theorem 2 rests on Lemma 4, the argument should be made rigorous, for example by using a right-derivative argument or the variational characterization of the spectral norm to establish d^+∥H0(t)∥2 ≤ -λ_{r*}∥H0(t)∥2 without assuming simplicity.
minor comments (5)
- [Throughout] There are two items labeled Remark 1, one in Section III-A and one after Theorem 1; the second should be renumbered.
- [Remark 5 and Appendix D] The eigenvalue-gap rates asserted in Remark 5, namely ∥\hat H_{i,0}(t)∥2 ≤ \hat c e^{-(\hat λ_i - \hat λ_{i+1})t}, are stated without proof; since the subsystems in (15) involve time-varying couplings, the claim that they follow by 'similar arguments' is not immediate, and a proof or a downgrade to a conjecture is needed.
- [Corollary 2 and Remark 4] Corollary 2 defines Ψ(t), and Remark 4 refers to 'Proposition 2' when discussing Ψ(t); this cross-reference should point to Corollary 2.
- [Proof of Lemma 4] The notation '2P/P1' appears before the Schur complement notation is defined; the manuscript should either define P/P1 in Section III or write 2H2 explicitly.
- [Section III-D, sharpness example] In the sharpness example, the displayed solution μ(t) = 2μ(0)/(1+t) is not a solution of \dot μ = -2μ^2; the correct solution is μ(t) = μ(0)/(1 + 2μ(0)t), and the surrounding sentence should be updated accordingly.
Circularity Check
No significant circularity: the cascade/H-coordinate analysis is derived from the defining ODEs and the convergence bounds follow from comparison and perturbation arguments without fitted parameters.
full rationale
The paper's central claims, Theorem 1 and Theorem 2, are derived in Appendix C directly from the gradient flow ODE (5) and the algebraic change of variables (13). Proposition 3 computes H1, H0, and H2 dynamics from the definitions H1 = P1^(-1), H0 = P1^(-1)P0, and H2 = P2 - P0^T P1^(-1)P0 using the chain rule; no assumption equivalent to the conclusion is imported. Theorem 2 solves the resulting linear and exponentially forced dynamics for H1 and H0, and Lemma 4 bounds H0 via a singular-vector argument; these are ordinary ODE estimates. Theorem 1 then converts the H-coordinate bounds back to P through Lemma 5, an inverse-sensitivity inequality. The only self-citation, [22] on the rank-one case, is contextual motivation and is not used to prove any theorem. The O(1/t) decay of the excess-parameter component is not a fit: it is illustrated on the invariant manifold P1 = Lambda1, P0 = 0, where the dynamics reduce to dot(P2) = -2P2^2, and the sharpness example exhibits a genuine trajectory of that manifold. A referee-level concern is that Lemma 3's proof invokes x(t) = 2x(t0)/(1 + t - t0) as the solution of dot(x) = -2x^2, which is not a solution except for a special initial value; this affects the correctness of the stated constant in Lemma 3, since the correct comparison yields P2(t) <= 1/(2(t - t0)) asymptotically, but it is an error in a comparison function rather than a case of a prediction reducing to its inputs. No fitted constants are renamed as predictions, no load-bearing claim rests on a self-citation, and no hidden identifiability condition is imported from prior work.
Assumptions & free parameters
assumptions (5)
- domain assumption The target matrix M is symmetric and has eigen-decomposition with positive eigenvalues Λ1 and non-positive part -Λ2.
- domain assumption Initialization satisfies P1(0) = Z1(0)Z1(0)^T ≻ 0 (full-rank signal alignment).
- standard math Existence and uniqueness of solutions to the ODE and standard ODE comparison principle.
- standard math Schur complement quotient identity and LU factorization identities (Crabtree-Haynsworth).
- standard math Neumann series / small-gain bound for matrix inverse perturbation (Lemma 5).
Cite this review
Pith. "Pith review of Stability properties of gradient flow dynamics for the symmetric low-rank matrix factorization problem." pith.science (2026). https://pith.science/paper/MGDVNDYU
@misc{pith2026241115972,
author = {Pith},
title = {Pith review of: Stability properties of gradient flow dynamics for the symmetric low-rank matrix factorization problem},
year = {2026},
howpublished = {\url{https://pith.science/paper/MGDVNDYU}},
note = {Machine review of arXiv:2411.15972}
}
abstract
The symmetric low-rank matrix factorization serves as a building block in many learning tasks, including matrix recovery and training of neural networks. However, despite a flurry of recent research, the dynamics of its training via non-convex factorized gradient-descent-type methods is not fully understood especially in the over-parameterized regime where the fitted rank is higher than the true rank of the target matrix. To overcome this challenge, we characterize equilibrium points of the gradient flow dynamics and examine their local and global stability properties. To facilitate a precise global analysis, we introduce a nonlinear change of variables that brings the dynamics into a cascade connection of three subsystems whose structure is simpler than the structure of the original system. We demonstrate that the Schur complement to a principal eigenspace of the target matrix is governed by an autonomous system that is decoupled from the rest of the dynamics. In the over-parameterized regime, we show that this Schur complement vanishes at an $O(1/t)$ rate, thereby capturing the slow dynamics that arises from excess parameters. We utilize a Lyapunov-based approach to establish exponential convergence of the other two subsystems. By decoupling the fast and slow parts of the dynamics, we offer new insight into the shape of the trajectories associated with local search algorithms and provide a complete characterization of the equilibrium points and their global stability properties. Such an analysis via nonlinear control techniques may prove useful in several related over-parameterized problems.
Figures
Reference graph
Works this paper leans on
-
[1]
Phase retrieval via Wirtinger flow: Theory and algorithms,
E. J. Candes, X. Li, and M. Soltanolkotabi, “Phase retrieval via Wirtinger flow: Theory and algorithms,” IEEE Trans. Inform. Theory , vol. 61, no. 4, pp. 1985–2007, 2015
work page 1985
-
[2]
Solving random quadratic systems of equations is nearly as easy as solving linear systems,
Y . Chen and E. J. Candes, “Solving random quadratic systems of equations is nearly as easy as solving linear systems,” Comm. Pure Appl. Math. , vol. 70, no. 5, pp. 822–883, 2017
work page 2017
-
[3]
C. Ma, K. Wang, Y . Chi, and Y . Chen, “Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion,” in Proc. Int. Conf. Mach. Learn., pp. 3345–3354, 2018
work page 2018
-
[4]
Low-rank solutions of linear matrix equations via Procrustes Flow,
S. Tu, R. Boczar, M. Simchowitz, M. Soltanolkotabi, and B. Recht, “Low-rank solutions of linear matrix equations via Procrustes Flow,” in Proc. Int. Conf. Mach. Learn. , pp. 964–973, 2016
work page 2016
-
[5]
Rapid, robust, and reliable blind deconvolution via nonconvex optimization,
X. Li, S. Ling, T. Strohmer, and K. Wei, “Rapid, robust, and reliable blind deconvolution via nonconvex optimization,” Appl. Comput. Harmon. Anal., vol. 47, no. 3, pp. 893–934, 2019
work page 2019
-
[6]
Regularized gradient descent: A non-convex recipe for fast joint blind deconvolution and demixing,
S. Ling and T. Strohmer, “Regularized gradient descent: A non-convex recipe for fast joint blind deconvolution and demixing,” Inf. Inference, vol. 8, no. 1, pp. 1–49, 2019
work page 2019
-
[7]
Phase retrieval using alternating minimization,
P. Netrapalli, P. Jain, and S. Sanghavi, “Phase retrieval using alternating minimization,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 26, 2013
work page 2013
-
[8]
Phase retrieval with random Gaussian sensing vectors by alternating projections,
I. Waldspurger, “Phase retrieval with random Gaussian sensing vectors by alternating projections,” IEEE Trans. Inform. Theory , vol. 64, no. 5, pp. 3301–3312, 2018
work page 2018
Show all 43 references
-
[9]
Non-convex matrix sensing: Breaking the quadratic rank barrier in the sample complexity,
D. St ¨oger and Y . Zhu, “Non-convex matrix sensing: Breaking the quadratic rank barrier in the sample complexity,” arXiv preprint arXiv:2408.13276, 2024
2024 arXiv
-
[10]
When are nonconvex problems not scary?,
J. Sun, Q. Qu, and J. Wright, “When are nonconvex problems not scary?,” arXiv preprint arXiv:1510.06096 , 2015
2015 arXiv
-
[11]
Cubic regularization of Newton method and its global performance,
Y . Nesterov and B. T. Polyak, “Cubic regularization of Newton method and its global performance,” Math. Program., vol. 108, no. 1, pp. 177–205, 2006
2006
-
[12]
Trust-region methods,
J. Nocedal and S. J. Wright, “Trust-region methods,” Numerical opti- mization, pp. 66–100, 2006
2006
-
[13]
How to escape saddle points efficiently,
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan, “How to escape saddle points efficiently,” in Proc. Int. Conf. Mach. Learn. , pp. 1724–1732, 2017
2017
-
[14]
Escaping from saddle points—online stochastic gradient for tensor decomposition,
R. Ge, F. Huang, C. Jin, and Y . Yuan, “Escaping from saddle points—online stochastic gradient for tensor decomposition,” in Proc. Conf. Learn. Theory , pp. 797–842, 2015
2015
-
[15]
Non-convex learning via stochastic gradient Langevin dynamics: A nonasymptotic analysis,
M. Raginsky, A. Rakhlin, and M. Telgarsky, “Non-convex learning via stochastic gradient Langevin dynamics: A nonasymptotic analysis,” in Proc. Conf. Learn. Theory , pp. 1674–1703, 2017
2017
-
[16]
A hitting time analysis of stochastic gradient Langevin dynamics,
Y . Zhang, P. Liang, and M. Charikar, “A hitting time analysis of stochastic gradient Langevin dynamics,” in Proc. Conf. Learn. Theory , pp. 1980– 2022, 2017
1980
-
[17]
Implicit regularization in deep matrix factorization,
S. Arora, N. Cohen, W. Hu, and Y . Luo, “Implicit regularization in deep matrix factorization,” in Advances in Neural Information Processing Systems, vol. 32, 2019
2019
-
[18]
Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval,
Y . Chen, Y . Chi, J. Fan, and C. Ma, “Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval,” Math. Program., vol. 176, pp. 5–37, 2019
2019
-
[19]
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction,
D. St ¨oger and M. Soltanolkotabi, “Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 34, pp. 23831–23843, 2021
2021
-
[20]
Implicit balancing and regular- ization: Generalization and convergence guarantees for overparameterized asymmetric matrix sensing,
M. Soltanolkotabi, D. St ¨oger, and C. Xie, “Implicit balancing and regular- ization: Generalization and convergence guarantees for overparameterized asymmetric matrix sensing,” in Proc. Conf. Learn. Theory, pp. 5140–5142, 2023
2023
-
[21]
Global convergence of gradient descent for asymmetric low-rank matrix factorization,
T. Ye and S. S. Du, “Global convergence of gradient descent for asymmetric low-rank matrix factorization,” in Proc. Adv. Neural Inf. Process. Syst., vol. 34, pp. 1429–1439, 2021
2021
-
[22]
On the stability of gradient flow dynamics for a rank-one matrix approximation problem,
H. Mohammadi, M. Razaviyayn, and M. R. Jovanovi ´c, “On the stability of gradient flow dynamics for a rank-one matrix approximation problem,” in Proc. Amer . Control Conf., pp. 4533–4538, 2018
2018
-
[23]
Matrix completion has no spurious local minimum,
R. Ge, J. D. Lee, and T. Ma, “Matrix completion has no spurious local minimum,” 2016
2016
-
[24]
Deep learning without poor local minima,
K. Kawaguchi, “Deep learning without poor local minima,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 29, 2016
2016
-
[25]
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks,
S. Arora, S. Du, W. Hu, Z. Li, and R. Wang, “Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks,” in Proc. Int. Conf. Mach. Learn. , pp. 322–332, 2019
2019
-
[26]
Simplified neuron model as a principal component analyzer,
E. Oja, “Simplified neuron model as a principal component analyzer,” J. Math. Biol. , vol. 15, pp. 267–273, 1982
1982
-
[27]
Neural networks, principal components, and subspaces,
E. Oja, “Neural networks, principal components, and subspaces,” Int. J. Neural Syst. , vol. 1, no. 01, pp. 61–68, 1989
1989
-
[28]
Helmke and J
U. Helmke and J. B. Moore, Optimization and Dynamical Systems . Springer Science & Business Media, 2012
2012
-
[29]
Weighted low-rank approximations,
N. Srebro and T. Jaakkola, “Weighted low-rank approximations,” in Proc. Int. Conf. Mach. Learn. , pp. 720–727, 2003
2003
-
[30]
Understanding the dynamics of gradient flow in overparameterized linear models,
S. Tarmoun, G. Franca, B. D. Haeffele, and R. Vidal, “Understanding the dynamics of gradient flow in overparameterized linear models,” in Proc. Int. Conf. Mach. Learn. , pp. 10153–10161, 2021
2021
-
[31]
Constrained nlp via gradient flow penalty continuation: Towards self-tuning robust penalty schemes,
F. Scott, R. Conejeros, and V . S. Vassiliadis, “Constrained nlp via gradient flow penalty continuation: Towards self-tuning robust penalty schemes,” Comput. Chem. Eng. , vol. 101, pp. 243–258, 2017
2017
-
[32]
On the solution of differential-algebraic equations through gradient flow embedding,
E. A. del Rio-Chanona, C. Bakker, F. Fiorelli, M. Paraskevopoulos, F. Scott, R. Conejeros, and V . S. Vassiliadis, “On the solution of differential-algebraic equations through gradient flow embedding,” Com- put. Chem. Eng. , vol. 103, pp. 165–175, 2017
2017
-
[33]
C. A. Desoer and M. Vidyasagar, Feedback Systems: Input-Output Properties. Philadelphia, PA, USA: SIAM, 2009
2009
-
[34]
G. B. Folland, Real Analysis: Modern Techniques and Their Applications , vol. 40. John Wiley & Sons, 1999
1999
-
[35]
An identity for the Schur complement of a matrix,
D. E. Crabtree and E. V . Haynsworth, “An identity for the Schur complement of a matrix,” Proc. Am. Math. Soc. , vol. 22, no. 2, pp. 364– 366, 1969. 7 APPENDIX A. Characterization of equilibrium points It is easy to verify that any ¯P ∈ His an equilibrium point of system (5). ...
1969
-
[36]
We observe that ¯P1Λ1 = ¯P1( ¯P1 + Λ 1 − ¯P1) = ¯P 2 1 + ¯P1(Λ1 − ¯P1) = ¯P 2 1 + ¯P1VoDoV T o (a) = ¯P 2 1 where (a) follow from the fact that Vi and Vo are orthogonal
Proof of Lemma 1: First, we show that the existence of such a matrix V and diagonal matrices Di and Do is sufficient for ¯P ∈ H. We observe that ¯P1Λ1 = ¯P1( ¯P1 + Λ 1 − ¯P1) = ¯P 2 1 + ¯P1(Λ1 − ¯P1) = ¯P 2 1 + ¯P1VoDoV T o (a) = ¯P 2 1 where (a) follow from the fact that Vi a...
-
[37]
Proof of Proposition 2: We consider the objective function in optimization problem (1) as a Lyapunov function candidate, VF (P ) = (1 /4)∥P − Λ∥2 F . The derivative of VF along the trajectories of (5) satisfies ˙VF = (1 /4) trace [(P − Λ) ˙P ] + (1/4) trace [ ˙P (P − Λ)] (a) =...
-
[38]
Proof of Lemma 3: Let (λ(t), w(t)) ∈ (R+, Rn−r⋆ ) be the principal eigenpair of the matrix P2(t) with wT (t)w(t) = 1. Since ˙wT (t)P2(t)w(t)+ wT (t)P2(t) ˙w(t) = λ(t) d(wT (t)w(t)) dt = 0 8 the derivative of VN (t) := λ(t) = wT (t)P2(t)w(t) along the solutions of (5) satisfies...
-
[39]
Proof of Proposition 3: We can write ˙H1 = d( P −1 1 )/dt = −H1 ˙P1H1 = −H1 P1Λ1 + Λ1P1 − 2(P 2 1 + P0P T 0 ) H1 = −H1 H −1 1 Λ1 + Λ1H −1 1 − 2(H −2 1 + P0P T 0 ) H1 = −Λ1H1 − H1Λ1 + 2 I + 2 H1P0P T 0 H1 = −Λ1H1 − H1Λ1 + 2 I + 2 H0H T 0 where the last equality follows from H1P...
-
[40]
The next lemma establishes the exponential decay of H0
Proof of Theorem 2: We first present a technical result. The next lemma establishes the exponential decay of H0. Lemma 4: For the matrix H0 governed by (14), the deriva- tive of the spectral norm satisfies d∥H0∥2 dt ≤ −λr⋆ ∥H0∥2. Proof: Let u(t) and v(t) be the principal left ...
-
[41]
Lemma 5: Let the matrices A ≻ 0 and B be such that ∥A − B∥2 < σ := σmin(A)
Proof of Theorem 1: We first present a lemma that we use to establish an upper bound on the error ∥P1 − Λ1∥2 as a 9 function of ∥P −1 1 − Λ−1 1 ∥2. Lemma 5: Let the matrices A ≻ 0 and B be such that ∥A − B∥2 < σ := σmin(A). Then, the matrix B is invertible, and it satisfies ∥A...
-
[42]
Thus, the convergence results in Theorem 1 hold
Proof of Corollary 1: We begin by noting that the condition on Z(0) is equivalent to the condition P1(0) ≻ 0 in Theorem 1. Thus, the convergence results in Theorem 1 hold. For some c1 > 0, Lemma 3 implies, ∥P2(t)∥2 = ∥Z2(t)∥2 2 ≤ c1/t and thus ∥Z2(t)∥2 ≤ p c1/t. Moreover, for ...
-
[43]
Equa- tion (36) for i = 1 follows from (31) and the fact that ˆP1 = ˆP1,1
Proof of Proposition 4: We use induction to show that P/ ˆPi = Ui+1 (36) where ˆPi ∈ R(n−mi)×(n−mi) is the 11-block of P . Equa- tion (36) for i = 1 follows from (31) and the fact that ˆP1 = ˆP1,1. To prove the case i + 1, we write P/ ˆPi+1 = ( P/ ˆPi)/((P/ ˆPi))11 = Ui+1/(Ui+...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.