REVIEW 3 major objections 6 minor 1 cited by
Characterizations of Strong Variational Sufficiency in General Models of Composite Optimization
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read In composite optimization, strong variational sufficiency at a KKT point is equivalent to a generalized SSOSC and to positive definiteness of the generalized Hessian of the augmented Lagrangian.
desk verdict A serious and novel characterization with a genuine proof gap in the core (ii)=>(iii) direction, plus heavy reliance on an unpublished preprint; deserves peer review but needs major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the second-order variational function $\Gamma_r(x,u)(v)$, defined as the minimum of $\langle v,d-v\rangle$ over all $d$ and all matrices $W$ in the limiting Jacobian $J\mathrm{Prox}_r(x+u)$ satisfying $v=Wd$, and infinite outside the union of the ranges of those matrices. Here $J\mathrm{Prox}_r(x+u)$ is the set of all limits of classical derivatives of the proximal mapping at nearby differentiable points. This function encodes the nonsmooth second-order behavior of $g$; under Assumption 2.2 it collapses to the point-based quadratic form $\Upsilon(v)=\langle v,(W^\dagger-I)v\rangle$ on the range of $W$. The proof also uses a norm estimate for the limiting coderivative of the proximal mapping (Theorem 3.1), the identification of $\Gamma_r$ as a member of the quadratic bundle of $r$ (the collection of epi-limits of second-order subderivatives), and parabolic regularity with twice epi-differentiability to justify the required epi-convergence. The generalized Hessian of the augmented Lagrangian, $\partial^2_{xx}L_\sigma$, is the set of matrices $\nabla^2_{xx}L(\bar x,\bar u)+\sigma F'(\bar x)^T U F'(\bar x)$ as $U$ ranges over the generalized Jacobian of the proximal mapping of $\sigma g^*$; its positive definiteness is shown to be equivalent to the other two conditions.
What would settle it
Run the paper's own formulas on a nuclear-norm composite problem with a degenerate $F'(\bar x)$: compute the generalized SSOSC from Example 4.1 and the minimum eigenvalue of every matrix in $\partial^2_{xx}L_\sigma(\bar x,\bar u)$ as $\sigma$ grows. If the SSOSC is strictly positive but some generalized Hessian matrix is indefinite for every $\sigma$, Theorem 5.1 would be false; the Appendix provides exactly the proximal-mapping data needed for this calculation.
Extended reading notes
Core claim
At a KKT point $(\bar x,\bar u)$ of the composite problem, if $g$ satisfies Assumptions 2.1 and 2.2, then the following are equivalent: (i) strong variational sufficiency for local optimality holds at $\bar x$ relative to $\bar u$; (ii) the generalized SSOSC $\langle\nabla^2_{xx}L(\bar x,\bar u)d,d\rangle+\Gamma_g(F(\bar x),\bar u)(F'(\bar x)d)>0$ holds for all $d\neq 0$; and (iii) there exists $\sigma>0$ large enough that all matrices in the generalized Hessian $\partial^2_{xx}L_\sigma(\bar x,\bar u)$ are positive definite. The second-order variational function $\Gamma_g$ is defined point-based through the limiting Jacobian of the proximal mapping of $g$, and under Assumption 2.2 it reduces to the explicit quadratic form $\Upsilon(v)=\langle v,(W^\dagger-I)v\rangle$ on the range of a single matrix $W$. The theorem requires no constraint qualification, nondegeneracy, or LICQ, and it includes the first explicit SSOSC characterization for nuclear norm-based composite optimization.
Load-bearing premise
The entire theorem rests on Assumption 2.2, which postulates that at the KKT point the set of limiting derivative matrices of the proximal mapping of $g$ has a single matrix $W$ whose range is the union of all their ranges and which minimizes a certain inner product; this is verified only for the semidefinite cone and the nuclear norm, not for general convex $g$.
Editorial extensions
If this is right
- For any composite problem in the covered class, checking strong variational sufficiency is reduced to a point-based inequality (the generalized SSOSC) or to positive definiteness of matrices, with no constraint qualification needed.
- The same theorem supplies the first explicit second-order characterizations for composite problems involving the nuclear norm, and it recovers the classical SSOSC characterization for nonlinear programs in the nondegenerate full-rank case.
- The equivalence with positive definiteness of the generalized Hessian of the augmented Lagrangian provides numerical methods with a certificate: for sufficiently large $\sigma$, the augmented Lagrangian is locally strongly convex in the variational sense, supporting local convergence-rate results for proximal and augmented Lagrangian algorithms.
- Because the characterization is point-based, it can be evaluated from the data at one KKT point and does not require inspecting the whole quadratic bundle of $g$.
Reading between the lines
- If Assumption 2.2 is verified for other $C^2$-cone reducible functions, such as the second-order cone, the $\ell_\infty$ norm, or the spectral norm, the same theorem would immediately yield explicit SSOSC characterizations for those models; the paper leaves this as future work.
- The second-order variational function may be usable as a computational surrogate in algorithms even outside the fully verified cases, since it can be evaluated from proximal-mapping Jacobians wherever those are available, although the paper does not propose such an algorithm.
- The range-of-a-single-matrix condition in Assumption 2.2 marks the boundary between point-based and bundle-based second-order theory: where it fails, the quadratic bundle characterization may still hold, but no point-based SSOSC of the form (23) is guaranteed by this paper.
- A natural testable extension is to check whether Assumption 2.2 holds for all convex functions that are $C^2$-cone reducible; if it does, the general model treated here would cover essentially every standard nonpolyhedral composite problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies composite optimization problems of the form min f0(x)+g(F(x)) with C^2-smooth f0,F and l.s.c. convex g, and investigates the recently introduced notion of strong variational sufficiency for local optimality. Its main result, Theorem 5.1, asserts that, under Assumptions 2.1 and 2.2, strong variational sufficiency at a KKT point is equivalent to: (ii) a generalized strong second-order sufficient condition (SSOSC) expressed through a newly defined second-order variational function, and (iii) positive definiteness of all matrices in a generalized Hessian of the augmented Lagrangian for all sufficiently large penalty parameter σ. The paper also derives an estimate for the limiting coderivative of proximal mappings (Theorem 3.1), constructs a quadratic-bundle approximation of the second-order variational function (Lemma 5.4), and verifies the main assumptions and gives explicit SSOSC formulas for two nonpolyhedral examples: the indicator of the positive-semidefinite cone and the nuclear norm.
Significance. If Theorem 5.1 is correct, the paper supplies a genuinely point-based second-order characterization of strong variational sufficiency for a broad class of nonpolyhedral composite optimization problems, without constraint qualifications and without nondegeneracy assumptions. The explicit treatment of the nuclear norm case is new and potentially useful for algorithms based on strong variational sufficiency. The constructive verification of Assumption 2.2 for the SDP cone and the nuclear norm, and the closed-form formulas in Examples 4.1, A.1, and A.2, are concrete contributions. However, the central proof of Theorem 5.1 currently has a serious uniformity gap, and the paper relies at load-bearing points on an unpublished preprint by two of the authors. The result is therefore plausible but not yet fully established as written.
major comments (3)
- [§5, Proof of Theorem 5.1, Eqs. (35)–(37)] The proof of (ii)=>(iii) does not justify the passage from a parameter σ that depends on the direction d to the single uniform σ required by conclusion (iii). Inequality (35) chooses σ so that (σ/2 − ||W†−I||)||d2|| ≥ 2||W†−I|| ||d1||, where d1 = UU†d and d2 = (I−UU†)d depend on d. The subsequent sentence asserts that (36) and (37) hold 'uniformly for all σ > 0 sufficiently large', but no compactness, continuity, or uniform bound on the ratio ||d1||/||d2|| is supplied. As d approaches rge U with d2 ≠ 0, the required σ in (35) can blow up, so the written argument only yields, for each d, a possibly direction-dependent threshold. Since assertion (iii) requires one σ for which all matrices in ∂²_xxL_σ(¯x,¯u) are positive definite, this gap is load-bearing. A more careful use of the uniform margin in (29) may repair the argument, but that repair is not present in the manuscript.
- [§4, Eqs. (24)–(25); §5, Lemma 5.4] The central representation Γ_g(F(x),u)(v) = min_{−d∈D*∂g(F(x),u)(−v)} ⟨v,d⟩ is imported from the authors' unpublished preprint [25, Lemma 5.3], and the simplification to the quadratic form ⟨v,(W†−I)v⟩ in (25) is stated without proof. Lemma 5.4, which is used to connect the quadratic bundle to Γ_g for the implications (iii)=>(i) and (i)=>(ii), also appeals to [25, Proposition 3.3]. Because these representations underlie the SSOSC formulation and the main equivalence, the manuscript is not self-contained at a load-bearing point. The authors should either prove these facts in the present paper or state them as explicit lemmas with full hypotheses and proofs; citing an unpublished preprint of the same authors is not sufficient for a central ingredient of the main theorem.
- [§2, Assumption 2.2; Appendix A] The scope of the claimed 'complete characterization' is substantially narrower than the general framing suggests. Assumption 2.2 is the key structural premise: it fixes a matrix W in ∇Prox_g(¯x+¯u) whose range equals the union of the ranges of all limiting Jacobians, and imposes a minimality condition on the equation y = Wz. Yet the paper verifies this assumption only for the SDP indicator and the nuclear norm in Appendix A. No general criterion is given for when Assumption 2.2 holds, and no verification is provided for other convex functions in the stated class (e.g., polyhedral norms, second-order cones, or ℓ_p norms). Since Theorem 5.1 is conditional on this assumption, the title and abstract overstate the generality of the characterization unless the authors either prove Assumption 2.2 for a broader class or explicitly frame the contribution as valid for the two verified examples and any future cases where the assumption is checked.
minor comments (6)
- [§2, Assumption 2.2] The notation in Assumption 2.2 uses W both for the distinguished matrix and as a dummy variable in the union and in the minimization, which makes the statement difficult to parse; please rename the distinguished matrix (e.g., ar W) and make the scope of the minimization variables explicit.
- [§4, after Eq. (25)] The statement that Γ_g(F(x),u) is 'further simplified' to the quadratic form Υ_g should explicitly note that this simplification requires Assumption 2.2, as the text later suggests; otherwise the reader may take (25) to hold unconditionally for all convex g.
- [Appendix A, Example A.2] The first sentence of Example A.2 contains a typo: 'In whet follows' should read 'In what follows'.
- [References, [18]] The publisher city of Rockafellar's Convex Analysis is misspelled as 'Princeron'; it should be 'Princeton'.
- [§5, Lemma 5.1] The proof of Lemma 5.1 asserts that θ tends to ∞ as d approaches the boundary of {||d||=1, Md≠0}; this is correct because the numerator is nonnegative on the kernel by assumption, but the argument would be clearer if this nonnegativity were stated explicitly.
- [§5, Theorem 5.1 statement] In the sentence 'Suppose that g satisfies Assumptions 2.1 at F(x) for all u...' the phrase 'Assumptions 2.1' should be the singular 'Assumption 2.1'.
Circularity Check
The main characterization is not definitionally circular, but the proof's bridge from the quadratic bundle criterion to the pointbased SSOSC delegates a load-bearing equality to a same-author unpublished preprint.
-
self citation load bearing
[Section 4, equations (24)-(25)]
"Under the fulfillment of Assumption 2.2, the second-order variational function (22) can be equivalently expressed at the KKT point (x, u) as Γg(F(x), u)(v) = min_{−d∈D*∂g(F(x),u)(−v)} ⟨v, d⟩ (24) (see [25, Lemma 5.3] for more details)."
This equivalence is the bridge between the paper's newly introduced SOVF and the coderivative expression used later in Lemma 5.4. The paper does not prove it; it cites [25], an arXiv preprint by P. P. Tang and C. J. Wang, who are the second and third authors of the present paper. Since Theorem 5.1 replaces the quadratic bundle criterion by the SOVF-based SSOSC, the characterization is not self-contained at this key reduction and rests on a self-citation that is not independently verified here.
-
self citation load bearing
[Section 5, proof of Lemma 5.4, equation (28)]
"Moreover, we deduce from [25, Proposition 3.3] that D(∂r)(x, u)(0) = {d | dr(x)(d) = ⟨u, d⟩} ◦. (28)"
Lemma 5.4 is the pivotal step showing that Γ belongs to the quadratic bundle, which then yields the implications (iii)=>(i) and (i)=>(ii) of Theorem 5.1 through [22, Theorems 3 and 5]. The proof of Lemma 5.4 invokes [25, Proposition 3.3] for the domain equality (28). Reference [25] is an unpublished preprint by Tang and Wang, the second and third authors of this paper, so a load-bearing part of the main equivalence is imported from the authors' own prior work rather than established in the current manuscript.
full rationale
The paper does not assume its target theorem: Theorem 5.1 is a genuinely new equivalence, and much of the proof is carried out with independent estimates (e.g., Theorem 3.1 and the algebraic reduction of the SOVF to the quadratic form (25), which follows from the projection properties of W under Assumption 2.2). Assumption 2.2 is a stated hypothesis, not a hidden input, and it is constructively checked for the SDP cone and the nuclear norm in Appendix A. The only load-bearing circularity concern is the reliance on [25] (Tang and Wang, same authors) for the coderivative representation (24) and the domain equality (28) used in Lemma 5.4; these connect the quadratic bundle criterion to the pointbased SSOSC, so the central characterization inherits an unverified same-author preprint at a key step. No fitted-value or definitional circularity was found. The skeptic's observation about the σ-uniformity gap in inequality (35) is a mathematical proof gap, not a circular step, and therefore does not by itself raise the circularity score.
Assumptions & free parameters
assumptions (5)
- domain assumption Standing assumptions: f0 and F are C2-smooth, g is l.s.c. and convex.
- domain assumption Assumption 2.1: parabolic regularity and parabolic epi-differentiability of g at the relevant points.
- ad hoc to paper Assumption 2.2: existence of W in ∇Prox_g(xbar+ubar) with the range representation and the minimality condition.
- ad hoc to paper Representation (24): the second-order variational function equals the coderivative minimum from [25, Lemma 5.3].
- standard math Rockafellar's quadratic bundle characterization theorems [22, Theorems 3 and 5].
invented entities (1)
-
Second-order variational function Γ_r(x,u)
independent evidence
Cite this review
Pith. "Pith review of Characterizations of Strong Variational Sufficiency in General Models of Composite Optimization." pith.science (2026). https://pith.science/paper/IILJP2AV
@misc{pith2026250709522,
author = {Pith},
title = {Pith review of: Characterizations of Strong Variational Sufficiency in General Models of Composite Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/IILJP2AV}},
note = {Machine review of arXiv:2507.09522}
}
read the original abstract
This paper investigates a recently introduced notion of strong variational sufficiency in optimization problems whose importance has been highly recognized in optimization theory, numerical methods, and applications. We address a general class of composite optimization problems and establish complete characterizations of strong variational sufficiency for their local minimizers in terms of a generalized version of the strong second-order sufficient condition (SSOSC) and the positive-definiteness of an appropriate generalized Hessian of the augmented Lagrangian calculated at the point in question. The generalized SSOSC is expressed via a novel second-order variational function, which reflects specific features of nonconvex composite models. The imposed assumptions describe the spectrum of composite optimization problems covered by our approach while being constructively implemented for nonpolyhedral problems that involve the nuclear norm function and the indicator function of the positive-semidefinite cone without any constraint qualifications.
Forward citations
Cited by 1 Pith paper
-
Second-Order Characterizations of Tilt Stability in Composite Optimization
The paper proves no-gap pointbased second-order characterizations of tilt-stable local minimizers for composite optimization with convex parabolically regular functions, under metric subregularity constraint qualification.
Reference graph
Works this paper leans on
-
[25]
P. P. Tang and C. J. Wang,Perturbation analysis of a class of composite optimization problems, arXiv:2401.10728v2 (2025)
arXiv 2025
-
[1]
J. F. Bonnans and A. Shapiro,Perturbation Analysis of Optimization Problems, Springer, New York, 2000. 22
work page 2000
-
[2]
C. H. Chen, Y.-J. Liu, D. F. Sun, and K.-C. Toh,A semismooth Newton-CG dual proximal point algorithm for matrix spectral norm approximation problems, Math. Program.155 (2016), 435–470
work page 2016
-
[3]
F. H. Clarke,Optimization and Nonsmooth Analysis, 2nd edition, Philadelphia, PA, 1990
work page 1990
-
[4]
H. Gfrerer,On second-order variational analysis of variational convexity of prox-regular functions, Set-Valued Var. Anal.35(2025), https://doi.org/10.1007/s11228–025–00744–8
doi:10.1007/s11228 2025
-
[5]
K. F. Jiang, D. F. Sun, and K.-C. Toh,Solving nuclear norm regularized and semidefi- nite matrix least squares problems with linear equality constraints, pp. 133–162, Springer International Publishing, Heidelberg, 2013
work page 2013
-
[6]
P. D. Khanh, B. S. Mordukhovich, and V. T. Phat,Variational convexity of functions and variational sufficiency in optimization, SIAM J. Optim.33(2023), 1121–1158
work page 2023
-
[7]
,Coderivative-based Newton methods in structured nonconvex and nonsmooth opti- mization, arXiv:2403.04262 (2025)
arXiv 2025
Show all 26 references
-
[8]
P. D. Khanh, B. S. Mordukhovich, V. T. Phat, and L. D. Viet,Characterizations of variational convexity and tilt stability via quadratic bundles, J. Convex Anal., to appear, arXiv:2501.04629 (2025)
2025 arXiv
-
[9]
S. Ma, D. Goldfarb, and L. Chen,Fixed point and Bregman iterative methods for matrix rank minimization, Math. Program.128(2011), 321–353
2011
-
[10]
Mohammadi, B
A. Mohammadi, B. S. Mordukhovich, and M. E. Sarabi,Parabolic regularity in geometric variational analysis, Trans. Amer. Math. Soc.374(2021), 1711–1763
2021
-
[11]
,Variational analysis of composite models with applications to continuous optimiza- tion, Math. Oper. Res.47(2022), 397–426
2022
-
[12]
Mohammadi and M
A. Mohammadi and M. E. Sarabi,Twice epi-differentiability of extended-real-valued func- tions with applications in composite optimization, SIAM J. Optim.30(2020), 2379–2409
2020
-
[13]
B. S. Mordukhovich,Maximum principle in problems of time optimal control with nons- mooth constraints, J. Appl. Math. Mech.40(1976), 960–969
1976
-
[14]
,Variational Analysis and Generalized Differentiation, I: Basic Theory, II: Appli- cations, Springer, Berlin, 2006
2006
-
[15]
,Variational Analysis and Applications, Springer, Cham, Switzerland, 2018
2018
-
[16]
,Second-Order Variational Analysis in Optimization, Variational Stability, and Con- trol Theory, Algorithms, Applications, Springer, Cham, Switzerland, 2024
2024
-
[17]
J.-S. Pang, D. F. Sun, and J. Sun,Semismooth homeomorphisms and strong stability of semidefinite and lorentz complementarity problems, Math. Oper. Res.28(2003), 39–63
2003
-
[18]
R. T. Rockafellar,Convex Analysis, Princeton University Press, Princeron, NJ, 1970. 23
1970
-
[19]
Anal.27(2019), 863–893
,Progressive decoupling of linkages in optimization and variational inequalities with elicitable convexity or monotonicity, Set-Valued Var. Anal.27(2019), 863–893
2019
-
[20]
Math.47(2019), 547–561
,Variational convexity and the local monotonicity of subgradient mappings, Vietnam J. Math.47(2019), 547–561
2019
-
[21]
Program.199(2022), 375–420
,Convergence of augmented Lagrangian methods in extensions beyond nonlinear programming, Math. Program.199(2022), 375–420
2022
-
[22]
Program.198(2023), 159–194
,Augmented Lagrangians and hidden convexity in sufficient conditions for local op- timality, Math. Program.198(2023), 159–194
2023
-
[23]
R. T. Rockafellar and R. J.-B. Wets,Variational Analysis, Springer, Berlin, 1998
1998
-
[24]
Sun,The strong second order sufficient condition and constraint nondegeneracy in non- linear semidefinite programming and their applications, Math
D. Sun,The strong second order sufficient condition and constraint nondegeneracy in non- linear semidefinite programming and their applications, Math. Oper. Res.31(2006), 761– 776
2006
-
[26]
S. W. Wang, C. Ding, Y. J. Zhang, and X. Y. Zhao,Strong variational sufficiency for nonlinear semidefinite programming and its implications, SIAM J. Optim.33(2023), 2988– 3011. 24
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.