Pith. sign in

REVIEW 2 major objections 6 minor 2 cited by

Deep Learning as the Disciplined Construction of Tame Objects

T0 review · 2 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This expository note argues that tame geometry—the study of sets and functions definable in o-minimal structures—is the natural home for deep learning, where nearly every practical activation and loss is definable and that definability is w

desk verdict A faithful exposition of tame geometry for DL, with a realism claim that needs either more evidence or a softer statement. read the letter →

arxiv 2509.18025 v2 pith:OBLIM2QG submitted 2025-09-22 math.OC cs.AIcs.LGmath.LOstat.ML

classification math.OCcs.AIcs.LGmath.LOstat.ML MSC 03C6490C2668T0749J52
keywords tamegeometryo-minimalstructuresdeeplearningstochasticsubgradientmethodClarkesubdifferentialnonsmoothnonconvexoptimizationdefinablefunctionsstratification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This expository note makes a two-part case. First, it argues that tame geometry—the study of sets and functions definable in o-minimal structures—is realistic for deep learning: the standard activations, losses, compositions, and partial minimizations used in neural networks are definable in one common such structure, so the networks themselves are tame. Second, it argues that this tameness is prolific, because it supplies exactly the finiteness and stratification properties needed to prove convergence of the stochastic subgradient method on the nonsmooth, nonconvex objectives that actually arise in training. The load-bearing result is Proposition 4.8: a definable locally Lipschitz objective with bounded iterates and square-summable, non-summable step sizes has every limit point Clarke critical and the objective values converge. A sympathetic reader should take the paper as a unified exposition of an existing proof chain, not as a claim of brand-new theorems.

What carries the argument

The central objects are o-minimal structures: collections of subsets of R^n closed under Boolean operations, products, and projections, in which one-dimensional definable sets are finite unions of intervals and points. A function is tame (definable) when its graph belongs to such a structure. The argument is carried by four pieces: the Projection formula, which stratifies any definable locally Lipschitz function into finitely many C^1 manifolds and confines each Clarke subgradient to the sum of the Riemannian gradient and the normal space of the stratum; the resulting Chain rule, which differentiates f along absolutely continuous curves stratum by stratum; the Descent identity expressing f(x

What would settle it

Exhibit a widely used neural-network training objective that is provably not definable in any o-minimal structure—say, a network using an unrestricted sine or cosine activation on an unbounded domain, or a composition that joins functions from two structures whose amalgamation is not o-minimal—and show that the stochastic subgradient method either fails to converge or converges to a point that is not Clarke critical. Alternatively, test the coverage claim directly by carefully classifying all 400 surveyed activation functions and checking whether the 89% figure survives; a substantially lower

Watch

Extended reading notes

Core claim

The paper's central claim is that o-minimality is both a realistic and a prolific mathematical framework for deep learning. Realism: nearly all activation and loss functions that appear in practice—ReLU, logistic, tanh, softplus, swish, mish, ELU, GELU, arctan, and standard losses—are definable in a common o-minimal structure, with the Pfaffian closure, a structure built from semialgebraic sets plus antiderivatives, sufficing for all of them; composition with definable linear maps preserves definability. Prolific: definability implies finite stratifications into smooth pieces, a projection formula for Clarke subdifferentials, a chain rule along absolutely continuous curves, and a weak Sard p

Load-bearing premise

The whole argument leans on the premise that real neural-network building blocks are mostly definable in one common o-minimal structure; the paper supports this with a rough count (about 89% of 400 activations 'by direct inspection', about 5% unclear, about 6% excluded) and with the caveat that composition preserves tameness only inside a single structure.

Editorial extensions

If this is right

  • Any definable locally Lipschitz training objective—including those built from ReLU, logistic, tanh, softplus, swish, mish, ELU, GELU, arctan, and standard losses—has the property that every bounded stochastic subgradient run with square-summable, non-summable step sizes converges to Clarke critical points.
  • The convergence guarantee extends to automatic-differentiation-based implementations, because definable conservative fields agree with Clarke subdifferentials almost everywhere; this covers the practical tools used for deep learning.
  • The same finiteness principles imply that central paths in definable convex semidefinite programs converge, and that definable curves have one-sided limits everywhere—a concrete manifestation of the 'no infinite oscillation' principle.
  • Because definable hypothesis classes have finite VC dimension, tame hypothesis spaces are PAC learnable under the fundamental theorem of statistical learning.
  • Optimization over tame functions is at least first-order tractable in a strong sense, in contrast to optimization over unrestricted trigonometric functions, which is undecidable even in the box-constrained case.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's rough estimate holds, the framework reaches far beyond the Table 1 list: most of the 400 surveyed activations are tame, so the convergence theorem should apply to most current architectures out of the box; the 5% fractional-derivative and 6% trigonometric examples are precisely where failures would surface.
  • An implicit tension: quantitative rate-of-convergence arguments often rely on polynomial boundedness and a field-of-exponents parameter, whereas many practical tame networks require the exponential function and thus live in structures that are not polynomially bounded; closing that gap would require a separate argument the paper does not provide.
  • A testable extension is to take a non-tame activation such as cos(z) on an unbounded domain, compose it with a simple linear layer, and check numerically whether stochastic subgradient descent still finds Clarke critical points; the theory predicts convergence can fail, with oscillation rather than divergence of the objective as the likely failure mode.
  • The same tools suggest a design rule for practitioners: keep every component of a network definable in a single o-minimal structure and avoid unrestricted periodic functions if the convergence guarantee is wanted—a far more permissive rule than convexity or smoothness.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper argues that o-minimal/tame geometry provides a unified and 'realistic' framework for deep learning, and that it is 'prolific' in yielding convergence guarantees. After introducing semialgebraic geometry and o-minimal structures (R_alg, R_exp, R_an, R_an,exp, R_Pfaff), the paper presents tame properties such as dimension, stratification, and definable choice. The central mathematical section is Section 4, which exposits the convergence of the stochastic subgradient method (SSM) on definable locally Lipschitz functions. The route is: Proposition 4.4 gives a continuous-time descent identity via the Projection formula 4.5 and Chain rule 4.7; Proposition 4.8 then asserts that every limit point of the bounded, square-summable-but-non-summable SSM iteration is Clarke critical, using a Weak Sard + Descent Lyapunov verification attributed to Davis et al. The manuscript closes with extensions to small-batch SSM, automatic differentiation, and conservative fields. The mathematical exposition is largely a careful condensation of known results, and the paper is transparent about relying on external benchmarks.

Significance. If accepted as a survey/expository note, the paper is a useful and largely reliable bridge between tame geometry, nonsmooth optimization, and deep learning. The conditional mathematical result—Proposition 4.8, that definable locally Lipschitz objectives are well-behaved for SSM—is correctly presented as a theorem of Davis et al., and the derivation of the descent identity through Projection formula 4.5, Chain rule 4.7, and Proposition 4.4 is coherent and faithful to the cited sources. The paper deserves credit for its explicit caveats about composition across structures (Remark 3.15), its honest admission of the heuristic nature of the 400-activation-function count (Remark 3.19), and its clear separation of the proven conditional statement from the broader 'realistic framework' claim. However, the paper's central scope claim—that tame geometry covers nearly all deep-learning objects—is not established by the evidence provided. The mathematical core is sound, but the transfer of Proposition 4.8 to 'current deep-learning architectures' requires a stronger, checkable inventory of definability in a common o-minimal structure, including non-activation components. As it stands, the reali

major comments (2)
  1. [§1.2, §5, Remark 3.19, Non-example 3.4, Remark 3.15] The central 'realistic framework' claim is not supported by a checkable inventory. Remark 3.19 estimates that 'around 89%' of the 400 surveyed activation functions are definable 'by direct inspection', while admitting 'a few liberties' with the count; about 5% are unclear and about 6% use unrestricted sine/cosine. This is not an itemized or reproducible audit. More importantly, the audit covers only activation and loss functions. Common deep-learning components such as sinusoidal positional encodings, rotary position embeddings, and periodic activation networks (e.g., SIREN) use unrestricted sin/cos and are therefore non-tame by Non-example 3.4; the manuscript does not address them. Since Proposition 4.8 requires the entire training objective to be definable in a single o-minimal structure—and Remark 3.15 correctly notes that composition is tame only within one common structure—Remark 3.
  2. [Remark 3.19 (items (1)–(3))] The counting conventions further undermine the 'nearly all' claim. Functions with stochastic parameters (e.g., Noisy ReLU) are counted as definable 'since it is semialgebraic in the two variables z and a'; however, as the footnote itself acknowledges, random variables on an abstract sample space fall outside the definable framework. If such objects are counted as definable, the 89% figure overstates the fraction of practical activation mechanisms that are actually covered by Proposition 4.8. Similarly, the 5% fractional-derivative class is left as 'a bit unclear', and the 6% sine/cosine class is explicitly non-tame. The paper should either present a precise breakdown with examples, or replace 'nearly all' with a more limited statement such as 'the common activations and losses listed in Table 1'. This is not a purely stylistic point: it determines whether the convergence guarantee in Pro
minor comments (6)
  1. [Title page] Typo: '*Eqal contribution' should read '*Equal contribution'.
  2. [§4.1] Typo: 'relevence' should be 'relevance'.
  3. [§4.4.1] Typo: 'Lispchitz' should be 'Lipschitz'.
  4. [§3.3, Lemma 3.23] Typo: 'Painlevéve-Kuratowski' should be 'Painlevé-Kuratowski'.
  5. [Proposition 4.3] Typo: 'exercice' should be 'exercise'.
  6. [Table 1] The notation (✓) is explained in the caption, but the parenthetical marks may confuse readers because the same symbol appears in multiple columns; a short example of how to read the table would help.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the convergence results are expositions of external theorems, and the only self-citations are illustrative.

full rationale

The paper is explicitly expository ('The ideas presented in this document are not new'), and its derivation chain is anchored in external benchmarks: Proposition 4.8 is quoted from Davis et al. [32], Projection formula 4.5 from Bolte et al. [17], Verdier stratification from Loi [75], and o-minimality of R_exp, R_an,exp, R_Pfaff from Wilkie, van den Dries-Miller, and Speissegger. The proof of Proposition 4.4 uses Projection formula 4.5 and Chain rule 4.7, which are proved from those external results plus elementary analysis; no equation reduces to a fitted or estimated value. The one self-citation ([7], used only for the illustrative right pane of Figure 5) is not load-bearing, and [8] is likewise a side remark about acceleration. The 'realistic' half of the thesis rests on an empirical estimate (Remark 3.19: 'we roughly estimate' 89% of 400 activations definable 'by direct inspection', with 'a few liberties', 5% unclear, 6% non-tame). That is a correctness/evidence limitation, not circularity: the estimate is not used to define the mathematical objects that are then 'predicted.' If the scope claim were challenged (e.g., sinusoidal positional encodings, or the requirement of a single common structure per Remark 3.15), the correct verdict would be that the realism claim is under-supported, not that the derivation is circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

As an expository survey, the ledger is dominated by imported background: the o-minimality theorems, the Projection formula, the stochastic approximation correspondence, and the Weak Sard property are lifted from the cited literature (Wilkie, Speissegger, van den Dries-Miller, Loi, Bolte et al., Davis et al.). The paper's own load-bearing addition is the ad hoc estimate in Remark 3.19 that about 89% of practical activation functions are tame, which is explicitly approximate. There are no fitted free parameters and no invented entities; the 'disciplined construction' framing is a naming of known facts, not a new postulate.

assumptions (5)
  • domain assumption The o-minimality of ℝalg, ℝan, ℝexp, ℝan,exp, ℝPfaff, and ℝG (Theorems 3.9, 3.11, 3.13, 3.16), cited to Gabrielov, Wilkie, van den Dries-Miller, and Speissegger but not proved in the paper.
    The paper's 'realistic framework' claim presupposes these external o-minimality theorems; they are imported background, not proved in the note.
  • domain assumption Projection formula 4.5 (Bolte et al. [17, Proposition 4]): for a definable locally Lipschitz f, ∂_c f(x) ⊂ ∇_R f(x) + N_M(x) on each stratum M of a C1-definable Whitney stratification.
    This imported structural result is what makes Chain rule 4.7, Corollary 4.6, and hence the descent argument of Proposition 4.4 work; Section 4.2 cites [17] rather than proving it.
  • domain assumption The stochastic approximation correspondence between the discrete SSM iteration and the continuous differential inclusion (SSMcont), per Benaïm-Hofbauer-Sorin [10] and Duchi-Ruan [46].
    Section 4.3's proof sketch of Proposition 4.8 replaces the discrete iteration by continuous trajectories using this classical correspondence, stated with citations and not proved.
  • domain assumption Weak Sard property for definable functions: the set of Clarke critical values is finite, via stratification plus Sard's theorem (cited to [17, Coro. 9(ii)] and [30, Rem. 3.1.5]).
    The Lyapunov verification in Section 4.3 requires this property to rule out wandering among critical levels; the paper gives the finite-strata intuition and defers to prior work.
  • ad hoc to paper About 89% of the 400 activation functions surveyed by Kunc and Kléma are definable in one of the listed o-minimal structures, by direct inspection.
    Remark 3.19 bases the 'realistic' half of the thesis on this count, produced by 'direct inspection' with the authors' own caveat 'we take a few liberties with this count'; it is not itemized or independently checkable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning as the Disciplined Construction of Tame Objects." pith.science (2026). https://pith.science/paper/OBLIM2QG

@misc{pith2026250918025,
  author       = {Pith},
  title        = {Pith review of: Deep Learning as the Disciplined Construction of Tame Objects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OBLIM2QG}},
  note         = {Machine review of arXiv:2509.18025}
}
read the original abstract

One can see deep-learning models as compositions of functions within the so-called tame geometry. In this expository note, we give an overview of some topics at the interface of tame geometry (also known as o-minimality), optimization theory, and deep learning theory and practice. To do so, we gradually introduce the concepts and tools used to build convergence guarantees for stochastic gradient descent in a general nonsmooth nonconvex, but tame, setting. This illustrates some ways in which tame geometry is a natural mathematical framework for the study of AI systems, especially within Deep Learning.

Figures

Figures reproduced from arXiv: 2509.18025 by the authors.

Figure 1
Figure 1. Illustration of the set 𝑆 in Eq. (2.1). Example 2.1 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Topologist’s sine curve 𝑡 ↦→ sin𝑡 −1 over (0,1]. This function exhibits infinite oscillations, it is not definable in any o-minimal structure; see also Non-Example 3.4. We note that Lemma 2.8 for semialgebraic functions extends to the more general setting of functions definable in any structure. Lemma 3.3 (Composability). In any structure, the composition of definable functions is definable. Proof. The same as the p… view at source ↗
Figure 3
Figure 3. Inclusion relations between the various o-minimal structures we consider. Notable separations: • arctan is definable in ℝarctan := (ℝalg, arctan) ⊆ ℝPfaff but not in ℝexp [14]; see also [82] • exp |[0,1] is definable in ℝan but not in ℝarctan [14] • sin |[0,2𝜋] is definable in ℝan but not in ℝexp [14] • 𝑥 √ 2 : (0, +∞) → ℝ is definable in ℝℝ alg but not in ℝG [80, 44] • 𝜙 (𝑥) := log Γ(𝑥) − (𝑥 − 1 2 ) log 𝑥 + 𝑥 − 1 2… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Dimension theory: the limit of a tame curve has smaller dimension than the curve. The two plots show non-tame examples, where the dimension of the limit of the curve is equal (a) or higher (b) than the dimension of the curve. Non-example 3.35 (Topologist’s sine curve).…
Figure 5
Figure 5. Figure 5: Left pane: a generic 2-dimensional DNN, with level lines (white) and non￾smooth points (black). This is neither a convex function, nor a smooth function, although the “cells” indicate regions where the function is smooth. Right pane: piece￾wise polynomial approximation…
Figure 6
Figure 6. Figure 6: Illustration of Projection formula 4.5 for two definable functions. In both cases, the stratum 𝑀 is smooth, 𝑓 |𝑀 is smooth, and there holds 𝜕 𝑐 𝑓 (𝑥) ⊂ ∇𝑅 𝑓 (𝑥) + 𝑁𝑀 (𝑥). In the context of a curve satisfying Eq. (SSMcont), the situation (a) is generic, while the situat…
Figure 7
Figure 7. Figure 7: The behavior of optimization algorithms is driven by the optimal stratum 𝑀★ = {(0, 𝑢) ⊤ : 𝑢 > 0}. Here, we show the iterates of noiseless full-batch SSM (SSMdiscrete), the nonsmooth-BFGS [71] and the Proximal Gradient [9, Chap. 10.4] algorithms, on a LASSO-type objecti…
Figure 8
Figure 8. Figure 8: Chances of evaluating an activation function at a nonsmooth point for feedforward ReLU Neural Networks of varying number of layers and layer width, for random weights and evaluation points. Probabilities are computed with 105 samples. differentiable, and their Jacobian…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lipschitzian SLLNs for random functions

    math.OC 2026-07 accept novelty 8.0 of 10

    Empirical averages of random locally Lipschitz functions converge to their expectation in the Lipschitz pseudometric under separability or NIP-definability conditions, giving uniform convergence of subdifferentials.

  2. Fast approximation and learning of binary classification tasks in o-minimal structures using ReLU neural networks

    math.LO 2026-06 unverdicted novelty 7.0 of 10

    ReLU networks approximate traceable definable subsets of the unit cube in L^p with size O(ε^{-p(n-1)/m}) and yield ERM learning rates of order N^{-m/(m+pn-p)} for hinge loss under uniform component bounds.

Reference graph

Works this paper leans on

109 extracted references · 2 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al.,{TensorFlow}: a system for{Large-Scale} machine learning, 12th USENIX symposium on operating systems design and implementation (OSDI 16), 2016, pp. 265–283

  2. [2]

    Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner, Machine bias, Ethics of data and analytics, Auerbach Publications, 2022, pp. 254–264

  3. [3]

    195, Princeton University Press, Princeton, NJ, 2017

    Matthias Aschenbrenner, Lou van den Dries, and Joris van der Hoeven, Asymptotic differential algebra and model theory of transseries , Annals of Mathematics Studies, vol. 195, Princeton University Press, Princeton, NJ, 2017. MR 3585498

  4. [4]

    Matthias Aschenbrenner, Lou van den Dries, and Joris van der Hoeven, Maximal hardy fields, 2025

  5. [5]

    Francis Bach, Learning theory from first principles , Adaptive Computation and Machine Learning, The MIT Press, Cambridge, Massachusetts, 2024

  6. [6]

    Bachute and Javed M

    Mrinal R. Bachute and Javed M. Subhedar, Autonomous Driving Architectures: Insights of Machine Learning and Deep Learning Algorithms , Machine Learning with Applications 6 (2021), 100164

  7. [7]

    Gilles Bareilles, Johannes Aspman, Jiri Nemecek, and Jakub Marecek, Piecewise polynomial regression of tame functions via integer programming , arXiv preprint arXiv:2311.13544 (2023)

  8. [8]

    Gilles Bareilles, Franck Iutzeler, and Jérôme Malick, Newton acceleration on manifolds identified by proximal gradient methods, Mathematical Programming (2022)

Show all 109 references
  1. [9]

    Amir Beck, First-Order Methods in Optimization , Society for Industrial and Applied Mathematics, Philadel- phia, PA, October 2017

  2. [10]

    4, 673–695

    Michel Benaïm, Josef Hofbauer, and Sylvain Sorin, Stochastic approximations and differential inclusions, part ii: Applications, Mathematics of Operations Research 31 (2006), no. 4, 673–695

  3. [11]

    G. C. Bento, B. S. Mordukhovich, T. S. Mota, and Yu Nesterov,Convergence of Descent Optimization Algorithms under Polyak-Ł ojasiewicz-Kurdyka Conditions , February 2025

  4. [12]

    Raghu Reddy, Machine learning techniques for credit risk evaluation: A systematic literature review, Journal of Banking and Financial Technology 4 (2020), no

    Siddharth Bhatore, Lalit Mohan, and Y. Raghu Reddy, Machine learning techniques for credit risk evaluation: A systematic literature review, Journal of Banking and Financial Technology 4 (2020), no. 1, 111–138

  5. [13]

    3, 1117–1147

    Pascal Bianchi, Walid Hachem, and Sholom Schechtman, Convergence of Constant Step Stochastic Gradient Descent for Non-Smooth Non-Convex Functions, Set-Valued and Variational Analysis30 (2022), no. 3, 1117–1147

  6. [14]

    4, 1173–1178

    Ricardo Bianconi, Nondefinability results for expansions of the field of real numbers by the exponential function and by the restricted sine function , The Journal of Symbolic Logic 62 (1997), no. 4, 1173–1178

  7. [15]

    Bochnak, M

    J. Bochnak, M. Coste, and M.F. Roy, Real algebraic geometry , Ergebnisse der Mathematik und ihrer Grenzge- biete. 3. Folge / A Series of Modern Surveys in Mathematics, Springer Berlin Heidelberg, 2013

  8. [16]

    Jérôme Bolte, Aris Daniilidis, and Adrian Lewis, Tame functions are semismooth, Math. Program. 117 (2009), no. 1-2, 5–19. MR 2421297

  9. [17]

    Jérôme Bolte, Aris Daniilidis, Adrian Lewis, and Masahiro Shiota,Clarke subgradients of stratifiable functions, SIAM J. Optim. 18 (2007), no. 2, 556–572. MR 2338451

  10. [18]

    Jérôme Bolte and Edouard Pauwels, A mathematical model for automatic differentiation in machine learning , Advances in Neural Information Processing Systems 33 (2020), 10809–10819

  11. [19]

    1, 19–51

    , Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning, Mathematical Programming 188 (2021), no. 1, 19–51

  12. [20]

    1, 553–603

    , Curiosities and counterexamples in smooth convex optimization , Mathematical Programming 195 (2022), no. 1, 553–603

  13. [21]

    48, Springer, 2009

    Vivek S Borkar, Stochastic approximation: a dynamical systems viewpoint , vol. 48, Springer, 2009

  14. [22]

    1, 258–280

    Michael Boshernitzan, Hardy fields and existence of transexponential functions , Aequationes mathematicae 30 (1986), no. 1, 258–280

  15. [23]

    Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning , SIAM Review 60 (2018), no

    Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning , SIAM Review 60 (2018), no. 2, 223–311

  16. [24]

    Nicolas Boumal, An introduction to optimization on smooth manifolds , Cambridge University Press, 2023

  17. [25]

    bradley- williams, i

    David Bradley-Williams and Immanuel Halupczok, Riso-stratifications and a tree invariant: D. bradley- williams, i. halupczok, Selecta Mathematica 31 (2025), no. 3, 52

  18. [26]

    3, 319–332

    Hunter Chase and James Freitag, Model theory and machine learning , Bulletin of Symbolic Logic 25 (2019), no. 3, 319–332

  19. [27]

    Clarke, Optimization and Nonsmooth Analysis, Society for Industrial and Applied Mathematics, January 1990

    Frank H. Clarke, Optimization and Nonsmooth Analysis, Society for Industrial and Applied Mathematics, January 1990

  20. [28]

    Gabriel Conant, Forking and dividing, https://www.forkinganddividing.com/, Accessed: 2025-09-22

  21. [29]

    DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 33

    Michel Coste, An introduction to o-minimal geometry, Istituti editoriali e poligrafici internazionali Pisa, 2000. DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 33

  22. [30]

    thesis, Migration - université en cours d’affectation, December 2001

    Didier d’Acunto, Sur les courbes intégrales du champ de gradient , Ph.D. thesis, Migration - université en cours d’affectation, December 2001

  23. [31]

    Damek Davis, Dmitriy Drusvyatskiy, and Liwei Jiang, Active manifolds, stratifications, and convergence to local minima in nonsmooth optimization , Foundations of Computational Mathematics (2025), 1–83

  24. [32]

    Lee, Stochastic Subgradient Method Converges on Tame Functions, Foundations of Computational Mathematics 20 (2020), no

    Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee, Stochastic Subgradient Method Converges on Tame Functions, Foundations of Computational Mathematics 20 (2020), no. 1, 119–154

  25. [33]

    1, 79–138

    Jan Denef and Lou van den Dries, P-adic and real subanalytic sets , Annals of Mathematics 128 (1988), no. 1, 79–138

  26. [34]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova,Bert: Pre-training of deep bidirectional transformers for language understanding , 2019

  27. [35]

    Martel, Joëlle Pineau, Peter Railton, Christine Tappolet, and Nathalie Voarino, Montréal Declaration for a responsible development of artificial intelligence , (2018)

    Marc-Antoine Dilhac, Catherine Régis, Christophe Abrassart, Yoshua Bengio, Guillaume Chicoisne, Nathalie De Marcellis-Warin, Sébastien Gambs, Vincent Gautrais, Martin Gibert, Lyse Langlois, François Laviolette, Pascale Lehoux, Jocelyn Maclure, Marie D. Martel, Joëlle Pineau, P...

  28. [36]

    Logic Found

    Lou van den Dries, Remarks on Tarski’s problem concerning(R,+,·, exp), Logic colloquium ’82 (Florence, 1982), Stud. Logic Found. Math., vol. 112, North-Holland, Amsterdam, 1984, pp. 97–121. MR 762106

  29. [37]

    , A generalization of the Tarski-Seidenberg theorem, and some nondefinability results , Bull. Amer. Math. Soc. (N.S.) 15 (1986), no. 2, 189–193. MR 854552

  30. [38]

    248, Cambridge University Press, Cambridge, 1998

    , Tame topology and o-minimal structures , London Mathematical Society Lecture Note Series, vol. 248, Cambridge University Press, Cambridge, 1998. MR 1633348

  31. [39]

    Press, Somerville, MA, 1999, pp

    , O-minimal structures and real analytic geometry , Current developments in mathematics, 1998 (Cambridge, MA), Int. Press, Somerville, MA, 1999, pp. 105–152

  32. [40]

    Lou van den Dries, Angus Macintyre, and David Marker, The elementary theory of restricted analytic fields with exponentiation, Ann. of Math. (2) 140 (1994), no. 1, 183–205. MR 1289495

  33. [41]

    London Math

    , Logarithmic-exponential power series, J. London Math. Soc. (2) 56 (1997), no. 3, 417–434. MR 1610431

  34. [42]

    Lou van den Dries and Chris Miller, On the real exponential field with restricted analytic functions , Israel Journal of Mathematics 85 (1994), 19–56

  35. [43]

    2, 497–540

    , Geometric categories and o-minimal structures, Duke Mathematical Journal 84 (1996), no. 2, 497–540

  36. [44]

    3, 513–565

    Lou van den Dries and Patrick Speissegger, The field of reals with multisummable series and the exponential function, Proceedings of the London Mathematical Society 81 (2000), no. 3, 513–565

  37. [45]

    Shiv Ram Dubey, Satish Kumar Singh, and Bidyut Baran Chaudhuri, Activation functions in deep learning: A comprehensive survey and benchmark , Neurocomputing 503 (2022), 92–108

  38. [46]

    4, 3229–3259

    John C Duchi and Feng Ruan, Stochastic methods for composite and weakly convex optimization problems , SIAM Journal on Optimization 28 (2018), no. 4, 3229–3259

  39. [47]

    6, Heldermann, Berlin, 1989

    Ryszard Engelking, General topology, Sigma series in pure mathematics, vol. 6, Heldermann, Berlin, 1989

  40. [48]

    Haishuo Fang, Ji-Ung Lee, Nafise Sadat Moosavi, and Iryna Gurevych,Transformers with learnable activation functions, 2023

  41. [49]

    4, 282–291

    Andrei M Gabrielov, Projections of semi-analytic sets , Functional Analysis and its applications 2 (1968), no. 4, 282–291

  42. [50]

    L. M. Graña Drummond and Y. Peterzil,The central path in smooth convex semidefinite programs, Optimization 51 (2002), no. 2, 207–233. MR 1928037

  43. [51]

    155– 210

    Michael Grant, Stephen Boyd, and Yinyu Ye, Disciplined Convex Programming, Global Optimization: From Theory to Implementation (Leo Liberti and Nelson Maculan, eds.), Springer US, Boston, MA, 2006, pp. 155– 210

  44. [52]

    Andreas Griewank and Andrea Walther, Evaluating derivatives: Principles and techniques of algorithmic differentiation, 2nd ed ed., Society for Industrial and Applied Mathematics, Philadelphia, PA, 2008

  45. [53]

    A Grothendieck, Esquisse d’un programme, London Math. Soc. Lect. Note Ser. 1 (1984), 7–48

  46. [54]

    Halická, E

    M. Halická, E. de Klerk, and C. Roos, On the convergence of the central path in semidefinite optimization , SIAM J. Optim. 12 (2002), no. 4, 1090–1099. MR 1922510

  47. [55]

    5, 1745–1780

    Martin Helmer and Vidit Nanda,Conormal Spaces and Whitney Stratifications, Foundations of Computational Mathematics 23 (2023), no. 5, 1745–1780

  48. [56]

    Dan Hendrycks and Kevin Gimpel, Gaussian error linear units (gelus) , 2016

  49. [57]

    Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal, Fundamentals of Convex Analysis, Springer Berlin Heidelberg, Berlin, Heidelberg, 2001

  50. [58]

    Lê Nguyên Hoang and El Mahdi El Mhamdi,Le fabuleux chantier: Rendre l’intelligence artificielle robustement bénéfique, EDP Sciences, Les Ulis, 2021. 34 G. BAREILLES, A. GEHRET, J. ASPMAN, J. LEPŠOVÁ, AND J. MAREČEK

  51. [59]

    A. D. Ioffe, An invitation to tame optimization , SIAM J. Optim. 19 (2008), no. 4, 1894–1917. MR 2486055

  52. [60]

    31, Curran Associates, Inc., 2018

    Sham M Kakade and Jason D Lee, Provably Correct Automatic Sub-Differentiation for Qualified Programs , Advances in Neural Information Processing Systems, vol. 31, Curran Associates, Inc., 2018

  53. [61]

    Kechris, Classical descriptive set theory , Graduate Texts in Mathematics, vol

    Alexander S. Kechris, Classical descriptive set theory , Graduate Texts in Mathematics, vol. 156, Springer- Verlag, New York, 1995. MR 1321597

  54. [62]

    Knight, Anand Pillay, and Charles Steinhorn, Definable sets in ordered structures

    Julia F. Knight, Anand Pillay, and Charles Steinhorn, Definable sets in ordered structures. II , Trans. Amer. Math. Soc. 295 (1986), no. 2, 593–605. MR 833698

  55. [63]

    Julian Kranz, Davide Gallon, Steffen Dereich, and Arnulf Jentzen, Sad neural networks: Divergent gradient flows and asymptotic optimality via o-minimal structures , 2025

  56. [64]

    Lothar Sebastian Krapp and Laura Wirth, Measurability in the fundamental theorem of statistical learning , arXiv preprint arXiv:2410.10243 (2024)

  57. [65]

    Vladimír Kunc and Jiří Kléma,Three decades of activations: A comprehensive survey of 400 activation functions for neural networks, 2024

  58. [66]

    Krzysztof Kurdyka, Olivier Le Gal, and Xuan Viet Nhan Nguyen,Tangent cones and𝐶1 regularity of definable sets, J. Math. Anal. Appl. 457 (2018), no. 1, 978–990. MR 3702738

  59. [67]

    Lasserre, Global Optimization with Polynomials and the Problem of Moments , SIAM Journal on Optimization 11 (2001), no

    Jean B. Lasserre, Global Optimization with Polynomials and the Problem of Moments , SIAM Journal on Optimization 11 (2001), no. 3, 796–817

  60. [68]

    Jean Bernard Lasserre, An introduction to polynomial and semi-algebraic optimization , Cambridge Texts in Applied Mathematics, Cambridge University Press, 2015

  61. [69]

    Olivier Le Gal, A generic condition implying o-minimality for restricted 𝐶∞-functions, Ann. Fac. Sci. Toulouse Math. (6) 19 (2010), no. 3-4, 479–492. MR 2790804

  62. [70]

    Wonyeol Lee, Hangyeol Yu, Xavier Rival, and Hongseok Yang,On correctness of automatic differentiation for non-differentiable functions, Advances in Neural Information Processing Systems 33 (2020), 6719–6730

  63. [71]

    Lewis and Michael L

    Adrian S. Lewis and Michael L. Overton, Nonsmooth optimization via quasi-Newton methods , Mathematical Programming 141 (2013), no. 1, 135–163

  64. [72]

    Lewis and Tonghua Tian, The Structure of Conservative Gradient Fields , SIAM Journal on Opti- mization 31 (2021), no

    Adrian S. Lewis and Tonghua Tian, The Structure of Conservative Gradient Fields , SIAM Journal on Opti- mization 31 (2021), no. 3, 2080–2083

  65. [73]

    1, 81–109

    Leo Liberti, Undecidability and hardness in mixed-integer nonlinear programming , RAIRO-Operations Re- search 53 (2019), no. 1, 81–109

  66. [74]

    1, 401–409

    Ta Loi, Whitney stratification of sets definable in the structure ℝexp, Banach Center Publications 33 (1996), no. 1, 401–409

  67. [75]

    Ta Lê Loi, Verdier and strict Thom stratifications in o-minimal structures , Illinois J. Math. 42 (1998), no. 2, 347–356. MR 1612771

  68. [76]

    Angus Macintyre and Eduardo D Sontag, Finiteness results for sigmoidal “neural” networks , Proceedings of the twenty-fifth annual ACM symposium on Theory of computing, 1993, pp. 325–334

  69. [77]

    Angus Macintyre and A. J. Wilkie, On the decidability of the real exponential field , Kreiseliana, A K Peters, Wellesley, MA, 1996, pp. 441–467. MR 1435773

  70. [78]

    05, WORLD SCIENTIFIC (EUROPE), May 2023

    Victor Magron and Jie Wang, Sparse Polynomial Optimization: Theory and Practice , Series on Optimization and Its Applications, vol. 05, WORLD SCIENTIFIC (EUROPE), May 2023

  71. [79]

    Pure Appl

    Chris Miller, Expansions of the real field with power functions , Ann. Pure Appl. Logic 68 (1994), no. 1, 79–94. MR 1278550

  72. [80]

    1, 79–94

    , Expansions of the real field with power functions , Annals of Pure and Applied Logic 68 (1994), no. 1, 79–94

  73. [81]

    , Exponentiation is hard to avoid , Proc. Amer. Math. Soc. 122 (1994), no. 1, 257–259. MR 1195484

  74. [82]

    Commun., vol

    , Basics of o-minimality and Hardy fields , Lecture notes on o-minimal structures and real analytic geometry, Fields Inst. Commun., vol. 62, Springer, New York, 2012, pp. 43–69. MR 2976990

  75. [83]

    4, 337–341

    Aleš Nekvinda and Luděk Zajíček, A simple proof of the Rademacher theorem , Časopis pro pěstování matematiky 113 (1988), no. 4, 337–341

  76. [84]

    137, Springer International Publishing, Cham, 2018

    Yurii Nesterov, Lectures on Convex Optimization , Springer Optimization and Its Applications, vol. 137, Springer International Publishing, Cham, 2018

  77. [85]

    2, 381–389

    Nhan Nguyen, Saurabh Trivedi, and David Trotman, A geometric proof of the existence of definable whitney stratifications, Illinois Journal of Mathematics 58 (2014), no. 2, 381–389

  78. [86]

    Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall, Activation functions: Comparison of trends in practice and research for deep learning , CoRR abs/1811.03378 (2018)

  79. [87]

    DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 35

    Adele Padgett and Patrick Speissegger, Definability of complex functions in o-minimal structures , 2025. DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 35

  80. [88]

    MR 4511519

    Adele Lee Padgett, Sublogarithmic-Transexponential Series, ProQuest LLC, Ann Arbor, MI, 2022, Thesis (Ph.D.)–University of California, Berkeley. MR 4511519

  81. [89]

    European Parliament, Artificial Intelligence Act, 2024

  82. [90]

    Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, Automatic differentiation in PyTorch, (2017)

  83. [91]

    Anand Pillay and Charles Steinhorn, Definable sets in ordered structures , Bull. Amer. Math. Soc. (N.S.) 11 (1984), no. 1, 159–162. MR 741730

  84. [92]

    II, Trans

    , Definable sets in ordered structures. II, Trans. Amer. Math. Soc.295 (1986), no. 2, 565–592. MR 833697

  85. [93]

    III, Trans

    , Definable sets in ordered structures. III, Trans. Amer. Math. Soc.309 (1988), no. 2, 469–476. MR 943306

  86. [94]

    317, Springer Science & Business Media, 2009

    R Tyrrell Rockafellar and Roger J-B Wets, Variational analysis, vol. 317, Springer Science & Business Media, 2009

  87. [95]

    Jean-Philippe Rolin, Establishing the o-minimality for expansions of the real field , Model theory with applications to algebra and analysis. Vol. 1, London Math. Soc. Lecture Note Ser., vol. 349, Cambridge Univ. Press, Cambridge, 2008, pp. 249–282. MR 2441383

  88. [96]

    , A survey on o-minimal structures , Real Algebraic Geometry (2011), 105

  89. [97]

    , Construction of o-minimal structures from quasianalytic classes , Lecture Notes on O-minimal Structures and Real Analytic Geometry, Springer, 2012, pp. 71–109

  90. [98]

    1, 258–266

    Maxwell Rosenlicht, On the value group of a differential valuation , American Journal of Mathematics 101 (1979), no. 1, 258–266

  91. [99]

    Stuart Jonathan Russell, Human compatible: Artificial intelligence and the problem of control , Viking, New York, 2019

  92. [100]

    12, 883–890

    Arthur Sard, The measure of the critical values of differentiable maps , Bulletin of the American Mathematical Society 48 (1942), no. 12, 883–890

  93. [101]

    2, 365–374

    Abraham Seidenberg, A new decision method for elementary algebra , Annals of Mathematics 60 (1954), no. 2, 365–374

  94. [102]

    92, Elsevier, 1990

    Saharon Shelah, Classification theory: and the number of non-isomorphic models , vol. 92, Elsevier, 1990

  95. [103]

    Reine Angew

    Patrick Speissegger, The Pfaffian closure of an o-minimal structure , J. Reine Angew. Math.508 (1999), 189–211. MR 1676876

  96. [104]

    MR 28796

    Alfred Tarski, A Decision Method for Elementary Algebra and Geometry , The Rand Corporation, Santa Monica, CA, 1948. MR 28796

  97. [105]

    Marcus Tressl, Introduction to o-minimal structures and an application to neural network learning , Course outline for LMS and EPSRC Short Instructional Course on Model Theory at the University of Leeds (2010)

  98. [106]

    David Trotman, Stratification theory, Handbook of geometry and topology of singularities I, Springer, 2020, pp. 243–273

  99. [107]

    Lasserre, Certifying global optimality of AC-OPF solutions via sparse polynomial optimization, Electric Power Systems Research 213 (2022), 108683

    Jie Wang, Victor Magron, and Jean B. Lasserre, Certifying global optimality of AC-OPF solutions via sparse polynomial optimization, Electric Power Systems Research 213 (2022), 108683

  100. [108]

    A. J. Wilkie, Model completeness results for expansions of the ordered field of real numbers by restricted Pfaffian functions and the exponential function , J. Amer. Math. Soc. 9 (1996), no. 4, 1051–1094. MR 1398816

  101. [109]

    Julio Zamora Esquivel, Adan Cruz Vargas, Rodrigo Camacho Perez, Paulo Lopez Meyer, Hector Cordourier, and Omesh Tickoo, Adaptive activation functions using fractional calculus , Proceedings of the IEEE/CVF international conference on computer vision workshops, 2019, pp. 0–0

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.