Pith. sign in

REVIEW 3 major objections 4 minor 29 references

Semiconcave Lipschitz functions admit smooth parametrized approximants that preserve semiconcavity and approximate both values and gradients in L^p.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 03:28 UTC pith:6T6ZLJHD

load-bearing objection Promising original construction, but the central W^{1,p} convergence proof rests on a false interpolation inequality; referee it for the construction, not for the proof. the 3 major comments →

arxiv 2602.07770 v2 pith:6T6ZLJHD submitted 2026-02-08 math.OC

Structure Preserving Approximation of Semiconcave Functions

classification math.OC MSC 26B2549J5241A3049L25
keywords semiconcavitystructure-preserving approximationregularized minimumactive setsHamilton-Jacobi-Bellman equationsvalue functiongradient approximationviscosity solutions
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper builds a family of smooth, parametrized functions that can approximate any Lipschitz semiconcave function without destroying the property that makes semiconcave functions useful: the curvature bound that underlies their regularity. The approximation replaces the minimum of a finite set of smooth functions by a recursively regularized minimum, and the paper proves that as the regularization and the number of parameters are tuned, the approximants converge uniformly in value and in L^p gradient while staying (C+δ)-semiconcave and (L+δ)-Lipschitz. The authors care because semiconcave functions appear as value functions of optimal control and viscosity equations, where gradients encode optimal feedback laws but are discontinuous. A key by-product is that the gradient of the approximation behaves like a probability-weighted average; in the zero-regularization limit all weight concentrates on one active smooth function, which lets the approximation respect the underlying Hamilton–Jacobi equation even across gradient discontinuities.

Core claim

Theorem 4.2 is the central universality statement: given any C-semiconcave, L-Lipschitz v and any tolerance δ>0, one can choose a regularization strength ε, a number of atoms n, a parameter space dimension m, and parameters θ so that the approximation v_{n,m,ε}(θ) is within δ of v in the C norm, its gradient is within δ in L^p(Ω), and the approximation is itself (C+δ)-semiconcave and (L+δ)-Lipschitz. The route: represent v as the infimum of countably many C² functions, select finitely many atoms, replace the hard minimum by the recursive smooth minimum ψ_{n,ε}, and parameterize the atoms through a universal C²-approximating family. The proof transfers C convergence of the atoms into gradient

What carries the argument

The construct is ψ_{n,ε}, a smooth stand-in for the minimum of n numbers, defined recursively as a_{i+1} − g_ε(a_{i+1} − ψ_{i,ε}) where g_ε is a C^{1,1} smoothing of the positive part that is nonnegative, has derivative in [0,1], has nonnegative second derivative, and converges to the positive part uniformly. Its gradient components are nonnegative and sum to one — a probability distribution over the indices — and its Hessian is negative semidefinite with an explicit lower bound. Composing ψ_{n,ε} with parameterized C² atoms yields the approximants v_{n,m,ε}; the probability interpretation identifies which atoms are active and, under mild conditions on g_ε, the limit ε→0 assigns all probabil

Load-bearing premise

The construction works only if the chosen finite-dimensional parameterization can approximate, in the C² norm, every one of the smooth building blocks that represent a semiconcave function; if that density property fails, the main universality result does not follow.

What would settle it

Take v(x)=min{exp(-|x-e|²/2), exp(-|x+e|²/2)} on (-1,1)^d, use tensor-product polynomial interpolation as the setting, and set g_ε to the piecewise-quadratic smoothing defined in the paper. On the hyperplane where the two atoms cross, evaluate the maximum gradient error and the Hamilton–Jacobi residual for decreasing ε (e.g., 10^-2, 10^-3, 10^-4) and increasing polynomial degree m. The paper predicts uniform convergence to zero on Ω_δ; if the gradient error or Hamiltonian residual fails to decrease as m grows (with ε held below δ/2), the central claim is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any Lipschitz semiconcave value function of an optimal control problem can be approximated by smooth semiconcave functions whose gradients are close in L^p, making them usable for smooth feedback synthesis.
  • The approximants' gradients converge pointwise to the gradient of a single active C² atom, not to an average, so for viscosity solutions of Hamilton–Jacobi equations the Hamiltonian residual tends to zero even at points where the true gradient is discontinuous.
  • The error estimates convert C-norm approximation of smooth atoms into W^{1,p} gradient errors with constants that depend on the domain and dimension only through explicit factors, not exponentially, which matters for higher-dimensional control problems.
  • On sets Ω_δ where the active index is separated by a margin δ, the gradient converges uniformly in W^{1,∞}; these sets can include parts of the discontinuity set, so the approximation captures discontinuities rather than smoothing them away.
  • The numerical example with the exponential distance function shows the proposed regularized-min approximation achieving uniform gradient and Hamiltonian convergence on Ω_δ for small ε, while a soft-min (log-sum-exp) smoothing leaves the Hamiltonian residual bounded away from zero.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: this architecture is effectively a smooth, trainable 'argmin layer'; plugging it into neural networks would give a universal approximator class for semiconcave functions that is also structure-preserving, extending convex-constrained networks to value functions of optimal control.
  • Beyond the paper: the probability-distribution interpretation suggests an attention-like mechanism in which the smoothing parameter acts as a temperature interpolating between hard argmin (winner-take-all) and uniform averaging; this could be explored for uncertainty quantification in feedback design.
  • Beyond the paper: the pointwise Hamiltonian consistency should hold for any smoothing g_ε satisfying the paper's assumptions (3.7)–(3.8), not just the two constructed examples; benchmarking different smoothings on the same test problem would be a natural, testable extension.
  • Beyond the paper: combined with adaptive or data-driven parameter spaces satisfying Hypothesis 4.1, the method could become a meshless solver for Hamilton–Jacobi equations whose error depends mildly on dimension, though the paper does not itself demonstrate high-dimensional performance.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces a smooth, structure-preserving parametrization for approximating semiconcave functions. The construction combines a smooth approximation of the minimum of a finite family of functions (via a regularized positive part) with a parameterized family of C^2 functions. The authors prove that the approximants remain semiconcave and Lipschitz with controlled constants, derive a probabilistic interpretation of the gradient weights, analyze active sets, and establish approximation results in C(Ω̄) and for gradients in L^p (and W^{1,p} in some statements). Numerical experiments on a 2D example illustrate the behavior of the method and compare it with the log-sum-exp approximation, arguing that the proposed Moreau-based smoothing respects the viscosity-solution structure at gradient discontinuities.

Significance. If the main theorems held, the paper would contribute a novel universal approximation scheme for semiconcave functions with explicit structural preservation, of potential interest for optimal control and Hamilton-Jacobi-Bellman equations. The recursive construction leading to a probability-distribution interpretation of the gradient weights is elegant, and the numerical comparison with log-sum-exp is informative. The paper is also explicit about the required density hypothesis (Hypothesis 4.1) and verifies it in the Chebyshev setting. However, the central gradient-convergence proof relies on a technical inequality (Proposition 4.1) that is false, and several statements use the wrong norm (W^{1,p} instead of L^p). These issues undermine the proof of the main universality theorem as written.

major comments (3)
  1. [§4, Proposition 4.1] Inequality (4.2) is false. Counterexample: Ω=(0,1), p=1, v_N(x)=N^{-2} sin(Nx). Then ‖v_N‖_C=N^{-2}, ‖∇v_N‖_∞=N^{-1}, but ‖∇v_N‖_{L^1}≈(2/π)N^{-1}; the right-hand side of (4.2) is O(N^{-3/2}), asymptotically smaller by a factor N^{1/2}. The function is 1-semiconcave, so this is not rescued by semiconcavity. The proof's assertion that the parameter sets V_i(s) have measure bounded independently of v fails for oscillatory functions. Since Theorem 4.1 derives (4.9) by applying (4.2) to v_{n,m,ε}-v_n, and Theorem 4.2's proof invokes Theorem 4.1, the proof of the central gradient-convergence claim is invalid as written.
  2. [§4, Theorem 4.1] The statement claims convergence in W^{1,p}(Ω), but the proof and estimate (4.9) only control ‖∇v_{n,m,ε}-∇v_n‖_{L^p(Ω;R^d)}. The W^{1,p} norm also includes the L^p norm of the function, which is not bounded by the displayed right-hand side. As written, the theorem is false in general; it should be restated as an L^p gradient estimate or the estimate must be supplemented.
  3. [§2, Proposition 2.1] Equation (2.3) is written as lim_{n→∞} ‖v_n-v‖_{C(Ω)} + ‖∇v_n-∇v‖_{W^{1,p}(Ω)} = 0. Since ∇v_n is generally not in W^{1,p}(Ω), this norm is not applicable; the proof establishes L^p convergence of gradients. This is a statement-level error that should be corrected to L^p(Ω;R^d).
minor comments (4)
  1. [§4, Theorem 4.3, (4.13)] The bound (4.13) is missing the exponent on g'_ε(δ): the proof derives (1 - g'_ε(δ/2)^{n-1}), while the statement writes g'_ε(δ)(n-1). The same issue appears in (4.14). Please correct the formula and align the δ/2 vs δ in the argument.
  2. [§4, Theorem 4.2, proof] In the last line of the proof, the inequality "∥∇v_{n,m,ε}(θ_m)∥_{C^1(Ω)} ≤ L+δ" should be "v_{n,m,ε}(θ_m) is (L+δ)-Lipschitz" (or equivalently ‖∇v_{n,m,ε}‖_{L∞} ≤ L+δ).
  3. [Abstract and §6] The abstract and conclusions claim approximation in W^{1,p}(Ω) for p∈[1,∞] and p=∞, but the proved results are L^p gradient convergence plus uniform gradient convergence on the sets Ω_δ, not on all of Ω. Please temper these claims.
  4. [Throughout] There are several typos and notation slips: e.g., "Ω" vs "Ω̄" in a few places, "MoreauRegMin" capitalized inconsistently, and the index "θ_n" in Theorem 4.2 is used ambiguously. A careful proofreading pass is recommended.

Circularity Check

0 steps flagged

No significant circularity: the central universality theorem is a composition of an external semiconcavity representation, an explicit density hypothesis on C^2 atoms, and an explicitly derived smooth minimum; self-citations are motivational only.

full rationale

The derivation chain is self-contained and does not reduce to its own inputs. Theorem 2.1 uses the external Cannarsa–Sinestrari characterization of semiconcave functions as infima of C^2 functions with uniformly bounded second derivatives. Proposition 2.1 extracts a finite-minimum approximation v_n from that representation with C and W^{1,p} convergence. Proposition 2.2 derives the key properties of the smooth minimum ψ_{n,ε} — including the probability-distribution property (2.17) — directly from the stated assumptions on g_ε, rather than imposing them. Proposition 2.3 shows that composing ψ_{n,ε} with C^2 atoms preserves semiconcavity and Lipschitz bounds. Theorem 4.1 transfers C-convergence of the atoms to C and L^p gradient convergence using Proposition 4.1, and Theorem 4.2 assembles these ingredients under Hypothesis 4.1, a density assumption on the parameterized C^2 atom family. No parameter is fitted to the target function's values or gradients, and no quantity called a prediction is an output of an optimization that already used that quantity as data. The authors' prior works [19,20] are cited in the introduction and conclusion as background/motivation but are not used in the proofs of Theorems 4.1–4.3. A separate concern raised in the review brief is that Proposition 4.1's inequality (4.2) may be invalid, which would be a correctness gap in the gradient-transfer step; however, an incorrect estimate is not a circular reduction, so it does not affect the circularity score.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical or mathematical entities. It relies on the standard representation of semiconcave functions, a density hypothesis on the parametrizing family, and explicit conditions on the regularizer g_ε. The approximation parameters ε, n, m are not fitted to data; they are standard discretization/regularization parameters. No free parameters are fitted to force the central convergence claims.

axioms (5)
  • domain assumption Representation theorem: every Lipschitz semiconcave function is the infimum of a family of C^2 functions with uniformly bounded Hessians (Theorem 2.1, from [6]).
    This is the starting point of the whole construction; it is quoted from Cannarsa-Sinestrari and assumed to hold for the target semiconcave functions.
  • domain assumption Density of the parametrization setting: for each φ∈C^2(Ω̄), ξ_m(θ_m)→φ in C^2 (Hypothesis 4.1).
    Load-bearing for Theorem 4.2; in the example it is satisfied by Chebyshev interpolation, but the general theorem requires the user to supply such a setting.
  • ad hoc to paper The regularization g_ε satisfies properties (2.12), (3.7), (3.8).
    These conditions are used to control the smooth approximation of the positive part and the limiting gradient concentration. Explicit examples exist, but (3.8) is asserted incorrectly for g_{ε,A}.
  • standard math Standard results on a.e. differentiability and upper-differentials of semiconcave functions from [6] (Theorem 3.3.3, Prop 3.3.4).
    Used in Proposition 2.1 and Remark 3.1 to justify gradient convergence and the reachability of D^*v_n.
  • standard math Coarea and area formulas from [21] and smooth approximation of Lipschitz functions from [22] in Proposition 4.1.
    Needed for the interpolation-type estimate connecting C and L^p gradient errors.

pith-pipeline@v1.3.0-alltime-deepseek · 30129 in / 14922 out tokens · 136370 ms · 2026-08-03T03:28:34.230736+00:00 · methodology

0 comments
read the original abstract

This article addresses structure-preserving smooth approximation of semiconcave functions. semiconcave functions are of particular interest because they naturally arise in a variety of variational problems, including {optimal feedback control, game theory, and optimal transport}. We leverage the fact that any semiconcave function can be represented as the {infimum of a countable family of \(C^2\) functions}. This infimum is expressed in a form that allows {approximation by finitely many functions}, combined with {smoothing operations}, such that each element of the approximating sequence remains semiconcave. The {active sets of indices} contributing to the representation of the semiconcave function and its approximations are analyzed in detail. Moreover, we show that the {gradients of the elements in the expansion of the approximating functions form a probability distribution}, a property of particular interest for the {value function in optimal control}. Approximation results are established in \(C(\bar \Omega)\) and in \(W^{1,p}(\Omega)\) for \(p \in [1,\infty)\) and \(p = \infty\). Finally, {numerical results} are presented to illustrate the approach on a test example.

Figures

Figures reproduced from arXiv: 2602.07770 by Donato V\'asquez-Varas, Karl Kunisch.

Figure 1
Figure 1. Figure 1: Illustrative example for (3.20). At x = 0 the three functions {ϕi} 3 i=1 are equal to v3(0). Additionally, we have that D∗ v3(0) = {−1, 1}. (3.24) To prove this, we note that, for x < 0, v3 is differentiable and v ′ 3 (x) = exp(x). Then clearly we have exp(0) = 1 is an element of D∗v3(0). For proving that −1 is an element of D∗v3(0), we note that for x ∈ (0, 1) we have that vn = −x and consequently v ′ n (… view at source ↗
Figure 2
Figure 2. Figure 2: Active sets (a) δ = 10−1 (b) δ = 10−2 [PITH_FULL_IMAGE:figures/full_fig_p027_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Ωδ sets of vd. In the remainder of this section, we use d = 2 for the ease of the exposition, since it allows to depict the active sets and the set Ωδ. In [PITH_FULL_IMAGE:figures/full_fig_p027_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Error for δ = 0. The results are summarized in [PITH_FULL_IMAGE:figures/full_fig_p030_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Error for δ = 10−2 (a) DW∞(·, δ) error for δ = 10−3 . (b) DH∞(·, δ) error for δ = 10−3 [PITH_FULL_IMAGE:figures/full_fig_p031_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Error for δ = 10−3 represent in a better manner the gradient of vd and viscosity solutions of HJB equations. For the cases δ = 10−2 and δ = 10−3 , DC, DW1 and DH1 do not differ from what was already shown in [PITH_FULL_IMAGE:figures/full_fig_p031_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

29 extracted references · 8 canonical work pages

  1. [1]

    Akian, S

    M. Akian, S. Gaubert, and A. Lakhoua. The max-plus finite element method for solving deterministic optimal control problems: Basic properties and convergence analysis.SIAM J. Control Optim., 47(2):817–848, 2008

  2. [2]

    B. Amos, L. Xu, and J. Z. Kolter. Input convex neural networks. In D. Precup and Y. W. Teh, editors,Proceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 146–155, New York, 06–11 Aug 2017. PMLR. URLhttps://proceedings.mlr.press/v70/amos17b.html

  3. [3]

    Bardi and I

    M. Bardi and I. Capuzzo-Dolcetta.Optimal Control and Viscosity Solutions of Hamilton- Jacobi-Bellman Equations. Birkh¨ auser, Boston, 1997

  4. [4]

    Boyd and A

    S. Boyd and A. Magnani. Convex piecewise-linear fitting.Optim. Eng., 10:1–17, 2009. doi: 10.1007/s11081-008-9045-3

  5. [5]

    G. C. Calafiore, S. Gaubert, and C. Possieri. Log-sum-exp neural networks and posynomial models for convex and log-log-convex data.IEEE Transactions on Neural Networks and Learning Systems, 31(3):827–838, 2020. doi: 10.1109/TNNLS.2019.2910417

  6. [6]

    Cannarsa and C

    P. Cannarsa and C. Sinestrari.Semiconcave functions, Hamilton-Jacobi equations, and optimal control. Progress in Nonlinear Differential Equations and Their Appli. Birkhauser, Boston, Feb. 2006

  7. [7]

    Y. Chen, Y. Shi, and B. Zhang. Optimal control via neural networks: A convex approach. In International Conference on Learning Representations, 2019. URLhttps://openreview. net/forum?id=H1MW72AcK7. 32

  8. [8]

    Darbon, P

    J. Darbon, P. M. Dower, and T. Meng. Neural network architectures using min-plus algebra for solving certain high-dimensional optimal control problems and Hamilton–Jacobi PDEs. Math. Control Signals Systems, 35(1):1–44, Mar. 2023

  9. [9]

    A. Douglis. Solutions in the large for multi-dimensional non linear partial differential equa- tions of first order.Ann. I. Fourier, 15(2):1–35, 1965. URLhttp://eudml.org/doc/73873

  10. [10]

    P. M. Dower, W. M. McEneaney, and H. Zhang. Max-plus fundamental solution semigroups for optimal control problems. In2015 Proceedings of the Conference on Control and its Applications, pages 368—-375, 2015

  11. [11]

    W. H. Fleming. The cauchy problem for a nonlinear first order partial differential equa- tion.J. Differ. Equ., 5(3):515–530, 1969. ISSN 0022-0396. doi: https://doi.org/10.1016/ 0022-0396(69)90091-6. URLhttps://www.sciencedirect.com/science/article/pii/ 0022039669900916

  12. [12]

    Gao and L

    B. Gao and L. Pavel. On the properties of the softmax function with application in game theory and reinforcement learning, 2018. URLhttps://arxiv.org/abs/1704.00805

  13. [13]

    Gaubert, W

    S. Gaubert, W. McEneaney, and Z. Qu. Curse of dimensionality reduction in max-plus based approximation methods: Theoretical estimates and improved pruning algorithms. In 2011 50th IEEE Conference on Decision and Control and European Control Conference, pages 1054–1061, 2011

  14. [14]

    G´ en´ erau,´E

    F. G´ en´ erau,´E. Oudet, and B. Velichkov. Cut locus on compact manifolds and uniform semiconcavity estimates for a variational inequality.Archive for Rational Mechanics and Analysis, 246(2–3):561–602, 2022. doi: 10.1007/s00205-022-01821-0

  15. [15]

    Ghosh, A

    A. Ghosh, A. Pananjady, A. Guntuboyina, and K. Ramchandran. Max-affine regression with universal parameter estimation for small-ball designs. In2020 IEEE International Symposium on Information Theory (ISIT), pages 2706–2710, 2020. doi: 10.1109/ISIT44484. 2020.9174116

  16. [16]

    Goujon, S

    A. Goujon, S. Neumayer, and M. Unser. Learning weakly convex regularizers for convergent image-reconstruction algorithms.SIAM J. Imaging Sci., 17(1):91–115, 2024. doi: 10.1137/ 23M1565243. URLhttps://doi.org/10.1137/23M1565243

  17. [17]

    B. Hanin. Universal function approximation by deep neural nets with bounded width and relu activations.Mathematics, 7(10), 2019. ISSN 2227-7390. doi: 10.3390/math7100992. URLhttps://www.mdpi.com/2227-7390/7/10/992

  18. [18]

    S. N. Kruˇ zkov. Generalized solutions of the hamilton-jacobi equations of eikonal type. i. formulation of the problems; existence, uniqueness and stability theo- rems; some properties of the solutions.Math. USSR-Sb., 27(3):406, apr 1975. doi: 10.1070/SM1975v027n03ABEH002522. URLhttps://dx.doi.org/10.1070/ SM1975v027n03ABEH002522

  19. [19]

    Kunisch and D

    K. Kunisch and D. V´ asquez-Varas. Convergence of machine learning methods for feedback control laws: averaged feedback learning scheme and data driven methods, 2025. URL https://arxiv.org/abs/2407.18403

  20. [20]

    Kunisch and D

    K. Kunisch and D. V´ asquez-Varas. Consistent smooth approximation of feedback laws for infinite horizon control problems with non-smooth value functions.J. Differ. Equ., 411:438–477, 2024. ISSN 0022–0396. doi: https://doi.org/10.1016/j.jde.2024.08.010. URL https://www.sciencedirect.com/science/article/pii/S0022039624004923. 33

  21. [21]

    E. L.C. and R. Gariepy.Measure Theory and Fine Properties of Functions. Chapman and Hall/CRC, New York, revised edition (1st ed.) edition, 2015. doi: 10.1201/b18333

  22. [22]

    Leoni.A First Course in Sobolev Spaces

    G. Leoni.A First Course in Sobolev Spaces. Graduate studies in mathematics. American Mathematical Soc., Providence, Rhode Island, 2009. ISBN 9780821884157. URLhttps: //books.google.at/books?id=W3RLWwnY0RkC

  23. [23]

    M¨ akel¨ a and P

    M. M¨ akel¨ a and P. Neittaanm¨ aki.Nonsmooth Optimization. World Scientific, New Jersey,

  24. [24]

    Quarteroni and A

    A. Quarteroni and A. Valli.Numerical approximation of partial differential equations. Springer series in computational mathematics. Springer, Berlin, Germany, 1 edition, Sept. 2008

  25. [25]

    L. Rifford. Existence of Lipschitz and semiconcave control-Lyapunov functions.SIAM J. Control Optim., 39(4):1043–1064, 2000. doi: 10.1137/S0363012999356039

  26. [26]

    L. Rifford. Semiconcave control-Lyapunov functions and stabilizing feedbacks.SIAM J. Control Optim., 41(3):659–681, 2002. doi: 10.1137/S0363012900375342

  27. [27]

    Santambrogio.Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling

    F. Santambrogio.Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling. Birkh¨ auser, Cham, 2015. ISBN 978-3-319-20828-2. doi: 10.1007/ 978-3-319-20828-2

  28. [28]

    X. Warin. The groupmax neural network approximation of convex functions.IEEE Transactions on Neural Networks and Learning Systems, 35(8):11608–11612, 2024. doi: 10.1109/TNNLS.2023.3240183. 34

  29. [1992]

    URLhttps://www.worldscientific.com/doi/abs/10.1142/ 1493

    doi: 10.1142/1493. URLhttps://www.worldscientific.com/doi/abs/10.1142/ 1493