REVIEW 3 major objections 4 minor 29 references
Semiconcave Lipschitz functions admit smooth parametrized approximants that preserve semiconcavity and approximate both values and gradients in L^p.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 03:28 UTC pith:6T6ZLJHD
load-bearing objection Promising original construction, but the central W^{1,p} convergence proof rests on a false interpolation inequality; referee it for the construction, not for the proof. the 3 major comments →
Structure Preserving Approximation of Semiconcave Functions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Theorem 4.2 is the central universality statement: given any C-semiconcave, L-Lipschitz v and any tolerance δ>0, one can choose a regularization strength ε, a number of atoms n, a parameter space dimension m, and parameters θ so that the approximation v_{n,m,ε}(θ) is within δ of v in the C norm, its gradient is within δ in L^p(Ω), and the approximation is itself (C+δ)-semiconcave and (L+δ)-Lipschitz. The route: represent v as the infimum of countably many C² functions, select finitely many atoms, replace the hard minimum by the recursive smooth minimum ψ_{n,ε}, and parameterize the atoms through a universal C²-approximating family. The proof transfers C convergence of the atoms into gradient
What carries the argument
The construct is ψ_{n,ε}, a smooth stand-in for the minimum of n numbers, defined recursively as a_{i+1} − g_ε(a_{i+1} − ψ_{i,ε}) where g_ε is a C^{1,1} smoothing of the positive part that is nonnegative, has derivative in [0,1], has nonnegative second derivative, and converges to the positive part uniformly. Its gradient components are nonnegative and sum to one — a probability distribution over the indices — and its Hessian is negative semidefinite with an explicit lower bound. Composing ψ_{n,ε} with parameterized C² atoms yields the approximants v_{n,m,ε}; the probability interpretation identifies which atoms are active and, under mild conditions on g_ε, the limit ε→0 assigns all probabil
Load-bearing premise
The construction works only if the chosen finite-dimensional parameterization can approximate, in the C² norm, every one of the smooth building blocks that represent a semiconcave function; if that density property fails, the main universality result does not follow.
What would settle it
Take v(x)=min{exp(-|x-e|²/2), exp(-|x+e|²/2)} on (-1,1)^d, use tensor-product polynomial interpolation as the setting, and set g_ε to the piecewise-quadratic smoothing defined in the paper. On the hyperplane where the two atoms cross, evaluate the maximum gradient error and the Hamilton–Jacobi residual for decreasing ε (e.g., 10^-2, 10^-3, 10^-4) and increasing polynomial degree m. The paper predicts uniform convergence to zero on Ω_δ; if the gradient error or Hamiltonian residual fails to decrease as m grows (with ε held below δ/2), the central claim is falsified.
If this is right
- Any Lipschitz semiconcave value function of an optimal control problem can be approximated by smooth semiconcave functions whose gradients are close in L^p, making them usable for smooth feedback synthesis.
- The approximants' gradients converge pointwise to the gradient of a single active C² atom, not to an average, so for viscosity solutions of Hamilton–Jacobi equations the Hamiltonian residual tends to zero even at points where the true gradient is discontinuous.
- The error estimates convert C-norm approximation of smooth atoms into W^{1,p} gradient errors with constants that depend on the domain and dimension only through explicit factors, not exponentially, which matters for higher-dimensional control problems.
- On sets Ω_δ where the active index is separated by a margin δ, the gradient converges uniformly in W^{1,∞}; these sets can include parts of the discontinuity set, so the approximation captures discontinuities rather than smoothing them away.
- The numerical example with the exponential distance function shows the proposed regularized-min approximation achieving uniform gradient and Hamiltonian convergence on Ω_δ for small ε, while a soft-min (log-sum-exp) smoothing leaves the Hamiltonian residual bounded away from zero.
Where Pith is reading between the lines
- Beyond the paper: this architecture is effectively a smooth, trainable 'argmin layer'; plugging it into neural networks would give a universal approximator class for semiconcave functions that is also structure-preserving, extending convex-constrained networks to value functions of optimal control.
- Beyond the paper: the probability-distribution interpretation suggests an attention-like mechanism in which the smoothing parameter acts as a temperature interpolating between hard argmin (winner-take-all) and uniform averaging; this could be explored for uncertainty quantification in feedback design.
- Beyond the paper: the pointwise Hamiltonian consistency should hold for any smoothing g_ε satisfying the paper's assumptions (3.7)–(3.8), not just the two constructed examples; benchmarking different smoothings on the same test problem would be a natural, testable extension.
- Beyond the paper: combined with adaptive or data-driven parameter spaces satisfying Hypothesis 4.1, the method could become a meshless solver for Hamilton–Jacobi equations whose error depends mildly on dimension, though the paper does not itself demonstrate high-dimensional performance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a smooth, structure-preserving parametrization for approximating semiconcave functions. The construction combines a smooth approximation of the minimum of a finite family of functions (via a regularized positive part) with a parameterized family of C^2 functions. The authors prove that the approximants remain semiconcave and Lipschitz with controlled constants, derive a probabilistic interpretation of the gradient weights, analyze active sets, and establish approximation results in C(Ω̄) and for gradients in L^p (and W^{1,p} in some statements). Numerical experiments on a 2D example illustrate the behavior of the method and compare it with the log-sum-exp approximation, arguing that the proposed Moreau-based smoothing respects the viscosity-solution structure at gradient discontinuities.
Significance. If the main theorems held, the paper would contribute a novel universal approximation scheme for semiconcave functions with explicit structural preservation, of potential interest for optimal control and Hamilton-Jacobi-Bellman equations. The recursive construction leading to a probability-distribution interpretation of the gradient weights is elegant, and the numerical comparison with log-sum-exp is informative. The paper is also explicit about the required density hypothesis (Hypothesis 4.1) and verifies it in the Chebyshev setting. However, the central gradient-convergence proof relies on a technical inequality (Proposition 4.1) that is false, and several statements use the wrong norm (W^{1,p} instead of L^p). These issues undermine the proof of the main universality theorem as written.
major comments (3)
- [§4, Proposition 4.1] Inequality (4.2) is false. Counterexample: Ω=(0,1), p=1, v_N(x)=N^{-2} sin(Nx). Then ‖v_N‖_C=N^{-2}, ‖∇v_N‖_∞=N^{-1}, but ‖∇v_N‖_{L^1}≈(2/π)N^{-1}; the right-hand side of (4.2) is O(N^{-3/2}), asymptotically smaller by a factor N^{1/2}. The function is 1-semiconcave, so this is not rescued by semiconcavity. The proof's assertion that the parameter sets V_i(s) have measure bounded independently of v fails for oscillatory functions. Since Theorem 4.1 derives (4.9) by applying (4.2) to v_{n,m,ε}-v_n, and Theorem 4.2's proof invokes Theorem 4.1, the proof of the central gradient-convergence claim is invalid as written.
- [§4, Theorem 4.1] The statement claims convergence in W^{1,p}(Ω), but the proof and estimate (4.9) only control ‖∇v_{n,m,ε}-∇v_n‖_{L^p(Ω;R^d)}. The W^{1,p} norm also includes the L^p norm of the function, which is not bounded by the displayed right-hand side. As written, the theorem is false in general; it should be restated as an L^p gradient estimate or the estimate must be supplemented.
- [§2, Proposition 2.1] Equation (2.3) is written as lim_{n→∞} ‖v_n-v‖_{C(Ω)} + ‖∇v_n-∇v‖_{W^{1,p}(Ω)} = 0. Since ∇v_n is generally not in W^{1,p}(Ω), this norm is not applicable; the proof establishes L^p convergence of gradients. This is a statement-level error that should be corrected to L^p(Ω;R^d).
minor comments (4)
- [§4, Theorem 4.3, (4.13)] The bound (4.13) is missing the exponent on g'_ε(δ): the proof derives (1 - g'_ε(δ/2)^{n-1}), while the statement writes g'_ε(δ)(n-1). The same issue appears in (4.14). Please correct the formula and align the δ/2 vs δ in the argument.
- [§4, Theorem 4.2, proof] In the last line of the proof, the inequality "∥∇v_{n,m,ε}(θ_m)∥_{C^1(Ω)} ≤ L+δ" should be "v_{n,m,ε}(θ_m) is (L+δ)-Lipschitz" (or equivalently ‖∇v_{n,m,ε}‖_{L∞} ≤ L+δ).
- [Abstract and §6] The abstract and conclusions claim approximation in W^{1,p}(Ω) for p∈[1,∞] and p=∞, but the proved results are L^p gradient convergence plus uniform gradient convergence on the sets Ω_δ, not on all of Ω. Please temper these claims.
- [Throughout] There are several typos and notation slips: e.g., "Ω" vs "Ω̄" in a few places, "MoreauRegMin" capitalized inconsistently, and the index "θ_n" in Theorem 4.2 is used ambiguously. A careful proofreading pass is recommended.
Circularity Check
No significant circularity: the central universality theorem is a composition of an external semiconcavity representation, an explicit density hypothesis on C^2 atoms, and an explicitly derived smooth minimum; self-citations are motivational only.
full rationale
The derivation chain is self-contained and does not reduce to its own inputs. Theorem 2.1 uses the external Cannarsa–Sinestrari characterization of semiconcave functions as infima of C^2 functions with uniformly bounded second derivatives. Proposition 2.1 extracts a finite-minimum approximation v_n from that representation with C and W^{1,p} convergence. Proposition 2.2 derives the key properties of the smooth minimum ψ_{n,ε} — including the probability-distribution property (2.17) — directly from the stated assumptions on g_ε, rather than imposing them. Proposition 2.3 shows that composing ψ_{n,ε} with C^2 atoms preserves semiconcavity and Lipschitz bounds. Theorem 4.1 transfers C-convergence of the atoms to C and L^p gradient convergence using Proposition 4.1, and Theorem 4.2 assembles these ingredients under Hypothesis 4.1, a density assumption on the parameterized C^2 atom family. No parameter is fitted to the target function's values or gradients, and no quantity called a prediction is an output of an optimization that already used that quantity as data. The authors' prior works [19,20] are cited in the introduction and conclusion as background/motivation but are not used in the proofs of Theorems 4.1–4.3. A separate concern raised in the review brief is that Proposition 4.1's inequality (4.2) may be invalid, which would be a correctness gap in the gradient-transfer step; however, an incorrect estimate is not a circular reduction, so it does not affect the circularity score.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Representation theorem: every Lipschitz semiconcave function is the infimum of a family of C^2 functions with uniformly bounded Hessians (Theorem 2.1, from [6]).
- domain assumption Density of the parametrization setting: for each φ∈C^2(Ω̄), ξ_m(θ_m)→φ in C^2 (Hypothesis 4.1).
- ad hoc to paper The regularization g_ε satisfies properties (2.12), (3.7), (3.8).
- standard math Standard results on a.e. differentiability and upper-differentials of semiconcave functions from [6] (Theorem 3.3.3, Prop 3.3.4).
- standard math Coarea and area formulas from [21] and smooth approximation of Lipschitz functions from [22] in Proposition 4.1.
read the original abstract
This article addresses structure-preserving smooth approximation of semiconcave functions. semiconcave functions are of particular interest because they naturally arise in a variety of variational problems, including {optimal feedback control, game theory, and optimal transport}. We leverage the fact that any semiconcave function can be represented as the {infimum of a countable family of \(C^2\) functions}. This infimum is expressed in a form that allows {approximation by finitely many functions}, combined with {smoothing operations}, such that each element of the approximating sequence remains semiconcave. The {active sets of indices} contributing to the representation of the semiconcave function and its approximations are analyzed in detail. Moreover, we show that the {gradients of the elements in the expansion of the approximating functions form a probability distribution}, a property of particular interest for the {value function in optimal control}. Approximation results are established in \(C(\bar \Omega)\) and in \(W^{1,p}(\Omega)\) for \(p \in [1,\infty)\) and \(p = \infty\). Finally, {numerical results} are presented to illustrate the approach on a test example.
Figures
Reference graph
Works this paper leans on
-
[1]
Akian, S
M. Akian, S. Gaubert, and A. Lakhoua. The max-plus finite element method for solving deterministic optimal control problems: Basic properties and convergence analysis.SIAM J. Control Optim., 47(2):817–848, 2008
2008
-
[2]
B. Amos, L. Xu, and J. Z. Kolter. Input convex neural networks. In D. Precup and Y. W. Teh, editors,Proceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 146–155, New York, 06–11 Aug 2017. PMLR. URLhttps://proceedings.mlr.press/v70/amos17b.html
2017
-
[3]
Bardi and I
M. Bardi and I. Capuzzo-Dolcetta.Optimal Control and Viscosity Solutions of Hamilton- Jacobi-Bellman Equations. Birkh¨ auser, Boston, 1997
1997
-
[4]
S. Boyd and A. Magnani. Convex piecewise-linear fitting.Optim. Eng., 10:1–17, 2009. doi: 10.1007/s11081-008-9045-3
-
[5]
G. C. Calafiore, S. Gaubert, and C. Possieri. Log-sum-exp neural networks and posynomial models for convex and log-log-convex data.IEEE Transactions on Neural Networks and Learning Systems, 31(3):827–838, 2020. doi: 10.1109/TNNLS.2019.2910417
arXiv 2020
-
[6]
Cannarsa and C
P. Cannarsa and C. Sinestrari.Semiconcave functions, Hamilton-Jacobi equations, and optimal control. Progress in Nonlinear Differential Equations and Their Appli. Birkhauser, Boston, Feb. 2006
2006
-
[7]
Y. Chen, Y. Shi, and B. Zhang. Optimal control via neural networks: A convex approach. In International Conference on Learning Representations, 2019. URLhttps://openreview. net/forum?id=H1MW72AcK7. 32
2019
-
[8]
Darbon, P
J. Darbon, P. M. Dower, and T. Meng. Neural network architectures using min-plus algebra for solving certain high-dimensional optimal control problems and Hamilton–Jacobi PDEs. Math. Control Signals Systems, 35(1):1–44, Mar. 2023
2023
-
[9]
A. Douglis. Solutions in the large for multi-dimensional non linear partial differential equa- tions of first order.Ann. I. Fourier, 15(2):1–35, 1965. URLhttp://eudml.org/doc/73873
1965
-
[10]
P. M. Dower, W. M. McEneaney, and H. Zhang. Max-plus fundamental solution semigroups for optimal control problems. In2015 Proceedings of the Conference on Control and its Applications, pages 368—-375, 2015
2015
-
[11]
W. H. Fleming. The cauchy problem for a nonlinear first order partial differential equa- tion.J. Differ. Equ., 5(3):515–530, 1969. ISSN 0022-0396. doi: https://doi.org/10.1016/ 0022-0396(69)90091-6. URLhttps://www.sciencedirect.com/science/article/pii/ 0022039669900916
1969
-
[12]
B. Gao and L. Pavel. On the properties of the softmax function with application in game theory and reinforcement learning, 2018. URLhttps://arxiv.org/abs/1704.00805
Pith/arXiv arXiv 2018
-
[13]
Gaubert, W
S. Gaubert, W. McEneaney, and Z. Qu. Curse of dimensionality reduction in max-plus based approximation methods: Theoretical estimates and improved pruning algorithms. In 2011 50th IEEE Conference on Decision and Control and European Control Conference, pages 1054–1061, 2011
2011
-
[14]
F. G´ en´ erau,´E. Oudet, and B. Velichkov. Cut locus on compact manifolds and uniform semiconcavity estimates for a variational inequality.Archive for Rational Mechanics and Analysis, 246(2–3):561–602, 2022. doi: 10.1007/s00205-022-01821-0
- [15]
-
[16]
A. Goujon, S. Neumayer, and M. Unser. Learning weakly convex regularizers for convergent image-reconstruction algorithms.SIAM J. Imaging Sci., 17(1):91–115, 2024. doi: 10.1137/ 23M1565243. URLhttps://doi.org/10.1137/23M1565243
-
[17]
B. Hanin. Universal function approximation by deep neural nets with bounded width and relu activations.Mathematics, 7(10), 2019. ISSN 2227-7390. doi: 10.3390/math7100992. URLhttps://www.mdpi.com/2227-7390/7/10/992
-
[18]
S. N. Kruˇ zkov. Generalized solutions of the hamilton-jacobi equations of eikonal type. i. formulation of the problems; existence, uniqueness and stability theo- rems; some properties of the solutions.Math. USSR-Sb., 27(3):406, apr 1975. doi: 10.1070/SM1975v027n03ABEH002522. URLhttps://dx.doi.org/10.1070/ SM1975v027n03ABEH002522
-
[19]
K. Kunisch and D. V´ asquez-Varas. Convergence of machine learning methods for feedback control laws: averaged feedback learning scheme and data driven methods, 2025. URL https://arxiv.org/abs/2407.18403
Pith/arXiv arXiv 2025
-
[20]
K. Kunisch and D. V´ asquez-Varas. Consistent smooth approximation of feedback laws for infinite horizon control problems with non-smooth value functions.J. Differ. Equ., 411:438–477, 2024. ISSN 0022–0396. doi: https://doi.org/10.1016/j.jde.2024.08.010. URL https://www.sciencedirect.com/science/article/pii/S0022039624004923. 33
-
[21]
E. L.C. and R. Gariepy.Measure Theory and Fine Properties of Functions. Chapman and Hall/CRC, New York, revised edition (1st ed.) edition, 2015. doi: 10.1201/b18333
doi:10.1201/b18333 2015
-
[22]
Leoni.A First Course in Sobolev Spaces
G. Leoni.A First Course in Sobolev Spaces. Graduate studies in mathematics. American Mathematical Soc., Providence, Rhode Island, 2009. ISBN 9780821884157. URLhttps: //books.google.at/books?id=W3RLWwnY0RkC
2009
-
[23]
M¨ akel¨ a and P
M. M¨ akel¨ a and P. Neittaanm¨ aki.Nonsmooth Optimization. World Scientific, New Jersey,
-
[24]
Quarteroni and A
A. Quarteroni and A. Valli.Numerical approximation of partial differential equations. Springer series in computational mathematics. Springer, Berlin, Germany, 1 edition, Sept. 2008
2008
-
[25]
L. Rifford. Existence of Lipschitz and semiconcave control-Lyapunov functions.SIAM J. Control Optim., 39(4):1043–1064, 2000. doi: 10.1137/S0363012999356039
-
[26]
L. Rifford. Semiconcave control-Lyapunov functions and stabilizing feedbacks.SIAM J. Control Optim., 41(3):659–681, 2002. doi: 10.1137/S0363012900375342
-
[27]
Santambrogio.Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
F. Santambrogio.Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling. Birkh¨ auser, Cham, 2015. ISBN 978-3-319-20828-2. doi: 10.1007/ 978-3-319-20828-2
2015
-
[28]
X. Warin. The groupmax neural network approximation of convex functions.IEEE Transactions on Neural Networks and Learning Systems, 35(8):11608–11612, 2024. doi: 10.1109/TNNLS.2023.3240183. 34
arXiv 2024
-
[1992]
URLhttps://www.worldscientific.com/doi/abs/10.1142/ 1493
doi: 10.1142/1493. URLhttps://www.worldscientific.com/doi/abs/10.1142/ 1493
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.