REVIEW 2 major objections 6 minor 2 cited by
Deep Learning as the Disciplined Construction of Tame Objects
T0 review · 2 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This expository note argues that tame geometry—the study of sets and functions definable in o-minimal structures—is the natural home for deep learning, where nearly every practical activation and loss is definable and that definability is w
desk verdict A faithful exposition of tame geometry for DL, with a realism claim that needs either more evidence or a softer statement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are o-minimal structures: collections of subsets of R^n closed under Boolean operations, products, and projections, in which one-dimensional definable sets are finite unions of intervals and points. A function is tame (definable) when its graph belongs to such a structure. The argument is carried by four pieces: the Projection formula, which stratifies any definable locally Lipschitz function into finitely many C^1 manifolds and confines each Clarke subgradient to the sum of the Riemannian gradient and the normal space of the stratum; the resulting Chain rule, which differentiates f along absolutely continuous curves stratum by stratum; the Descent identity expressing f(x
What would settle it
Exhibit a widely used neural-network training objective that is provably not definable in any o-minimal structure—say, a network using an unrestricted sine or cosine activation on an unbounded domain, or a composition that joins functions from two structures whose amalgamation is not o-minimal—and show that the stochastic subgradient method either fails to converge or converges to a point that is not Clarke critical. Alternatively, test the coverage claim directly by carefully classifying all 400 surveyed activation functions and checking whether the 89% figure survives; a substantially lower
Extended reading notes
Core claim
The paper's central claim is that o-minimality is both a realistic and a prolific mathematical framework for deep learning. Realism: nearly all activation and loss functions that appear in practice—ReLU, logistic, tanh, softplus, swish, mish, ELU, GELU, arctan, and standard losses—are definable in a common o-minimal structure, with the Pfaffian closure, a structure built from semialgebraic sets plus antiderivatives, sufficing for all of them; composition with definable linear maps preserves definability. Prolific: definability implies finite stratifications into smooth pieces, a projection formula for Clarke subdifferentials, a chain rule along absolutely continuous curves, and a weak Sard p
Load-bearing premise
The whole argument leans on the premise that real neural-network building blocks are mostly definable in one common o-minimal structure; the paper supports this with a rough count (about 89% of 400 activations 'by direct inspection', about 5% unclear, about 6% excluded) and with the caveat that composition preserves tameness only inside a single structure.
Editorial extensions
If this is right
- Any definable locally Lipschitz training objective—including those built from ReLU, logistic, tanh, softplus, swish, mish, ELU, GELU, arctan, and standard losses—has the property that every bounded stochastic subgradient run with square-summable, non-summable step sizes converges to Clarke critical points.
- The convergence guarantee extends to automatic-differentiation-based implementations, because definable conservative fields agree with Clarke subdifferentials almost everywhere; this covers the practical tools used for deep learning.
- The same finiteness principles imply that central paths in definable convex semidefinite programs converge, and that definable curves have one-sided limits everywhere—a concrete manifestation of the 'no infinite oscillation' principle.
- Because definable hypothesis classes have finite VC dimension, tame hypothesis spaces are PAC learnable under the fundamental theorem of statistical learning.
- Optimization over tame functions is at least first-order tractable in a strong sense, in contrast to optimization over unrestricted trigonometric functions, which is undecidable even in the box-constrained case.
Reading between the lines
- If the paper's rough estimate holds, the framework reaches far beyond the Table 1 list: most of the 400 surveyed activations are tame, so the convergence theorem should apply to most current architectures out of the box; the 5% fractional-derivative and 6% trigonometric examples are precisely where failures would surface.
- An implicit tension: quantitative rate-of-convergence arguments often rely on polynomial boundedness and a field-of-exponents parameter, whereas many practical tame networks require the exponential function and thus live in structures that are not polynomially bounded; closing that gap would require a separate argument the paper does not provide.
- A testable extension is to take a non-tame activation such as cos(z) on an unbounded domain, compose it with a simple linear layer, and check numerically whether stochastic subgradient descent still finds Clarke critical points; the theory predicts convergence can fail, with oscillation rather than divergence of the objective as the likely failure mode.
- The same tools suggest a design rule for practitioners: keep every component of a network definable in a single o-minimal structure and avoid unrestricted periodic functions if the convergence guarantee is wanted—a far more permissive rule than convexity or smoothness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that o-minimal/tame geometry provides a unified and 'realistic' framework for deep learning, and that it is 'prolific' in yielding convergence guarantees. After introducing semialgebraic geometry and o-minimal structures (R_alg, R_exp, R_an, R_an,exp, R_Pfaff), the paper presents tame properties such as dimension, stratification, and definable choice. The central mathematical section is Section 4, which exposits the convergence of the stochastic subgradient method (SSM) on definable locally Lipschitz functions. The route is: Proposition 4.4 gives a continuous-time descent identity via the Projection formula 4.5 and Chain rule 4.7; Proposition 4.8 then asserts that every limit point of the bounded, square-summable-but-non-summable SSM iteration is Clarke critical, using a Weak Sard + Descent Lyapunov verification attributed to Davis et al. The manuscript closes with extensions to small-batch SSM, automatic differentiation, and conservative fields. The mathematical exposition is largely a careful condensation of known results, and the paper is transparent about relying on external benchmarks.
Significance. If accepted as a survey/expository note, the paper is a useful and largely reliable bridge between tame geometry, nonsmooth optimization, and deep learning. The conditional mathematical result—Proposition 4.8, that definable locally Lipschitz objectives are well-behaved for SSM—is correctly presented as a theorem of Davis et al., and the derivation of the descent identity through Projection formula 4.5, Chain rule 4.7, and Proposition 4.4 is coherent and faithful to the cited sources. The paper deserves credit for its explicit caveats about composition across structures (Remark 3.15), its honest admission of the heuristic nature of the 400-activation-function count (Remark 3.19), and its clear separation of the proven conditional statement from the broader 'realistic framework' claim. However, the paper's central scope claim—that tame geometry covers nearly all deep-learning objects—is not established by the evidence provided. The mathematical core is sound, but the transfer of Proposition 4.8 to 'current deep-learning architectures' requires a stronger, checkable inventory of definability in a common o-minimal structure, including non-activation components. As it stands, the reali
major comments (2)
- [§1.2, §5, Remark 3.19, Non-example 3.4, Remark 3.15] The central 'realistic framework' claim is not supported by a checkable inventory. Remark 3.19 estimates that 'around 89%' of the 400 surveyed activation functions are definable 'by direct inspection', while admitting 'a few liberties' with the count; about 5% are unclear and about 6% use unrestricted sine/cosine. This is not an itemized or reproducible audit. More importantly, the audit covers only activation and loss functions. Common deep-learning components such as sinusoidal positional encodings, rotary position embeddings, and periodic activation networks (e.g., SIREN) use unrestricted sin/cos and are therefore non-tame by Non-example 3.4; the manuscript does not address them. Since Proposition 4.8 requires the entire training objective to be definable in a single o-minimal structure—and Remark 3.15 correctly notes that composition is tame only within one common structure—Remark 3.
- [Remark 3.19 (items (1)–(3))] The counting conventions further undermine the 'nearly all' claim. Functions with stochastic parameters (e.g., Noisy ReLU) are counted as definable 'since it is semialgebraic in the two variables z and a'; however, as the footnote itself acknowledges, random variables on an abstract sample space fall outside the definable framework. If such objects are counted as definable, the 89% figure overstates the fraction of practical activation mechanisms that are actually covered by Proposition 4.8. Similarly, the 5% fractional-derivative class is left as 'a bit unclear', and the 6% sine/cosine class is explicitly non-tame. The paper should either present a precise breakdown with examples, or replace 'nearly all' with a more limited statement such as 'the common activations and losses listed in Table 1'. This is not a purely stylistic point: it determines whether the convergence guarantee in Pro
minor comments (6)
- [Title page] Typo: '*Eqal contribution' should read '*Equal contribution'.
- [§4.1] Typo: 'relevence' should be 'relevance'.
- [§4.4.1] Typo: 'Lispchitz' should be 'Lipschitz'.
- [§3.3, Lemma 3.23] Typo: 'Painlevéve-Kuratowski' should be 'Painlevé-Kuratowski'.
- [Proposition 4.3] Typo: 'exercice' should be 'exercise'.
- [Table 1] The notation (✓) is explained in the caption, but the parenthetical marks may confuse readers because the same symbol appears in multiple columns; a short example of how to read the table would help.
Circularity Check
No significant circularity: the convergence results are expositions of external theorems, and the only self-citations are illustrative.
full rationale
The paper is explicitly expository ('The ideas presented in this document are not new'), and its derivation chain is anchored in external benchmarks: Proposition 4.8 is quoted from Davis et al. [32], Projection formula 4.5 from Bolte et al. [17], Verdier stratification from Loi [75], and o-minimality of R_exp, R_an,exp, R_Pfaff from Wilkie, van den Dries-Miller, and Speissegger. The proof of Proposition 4.4 uses Projection formula 4.5 and Chain rule 4.7, which are proved from those external results plus elementary analysis; no equation reduces to a fitted or estimated value. The one self-citation ([7], used only for the illustrative right pane of Figure 5) is not load-bearing, and [8] is likewise a side remark about acceleration. The 'realistic' half of the thesis rests on an empirical estimate (Remark 3.19: 'we roughly estimate' 89% of 400 activations definable 'by direct inspection', with 'a few liberties', 5% unclear, 6% non-tame). That is a correctness/evidence limitation, not circularity: the estimate is not used to define the mathematical objects that are then 'predicted.' If the scope claim were challenged (e.g., sinusoidal positional encodings, or the requirement of a single common structure per Remark 3.15), the correct verdict would be that the realism claim is under-supported, not that the derivation is circular.
Assumptions & free parameters
assumptions (5)
- domain assumption The o-minimality of ℝalg, ℝan, ℝexp, ℝan,exp, ℝPfaff, and ℝG (Theorems 3.9, 3.11, 3.13, 3.16), cited to Gabrielov, Wilkie, van den Dries-Miller, and Speissegger but not proved in the paper.
- domain assumption Projection formula 4.5 (Bolte et al. [17, Proposition 4]): for a definable locally Lipschitz f, ∂_c f(x) ⊂ ∇_R f(x) + N_M(x) on each stratum M of a C1-definable Whitney stratification.
- domain assumption The stochastic approximation correspondence between the discrete SSM iteration and the continuous differential inclusion (SSMcont), per Benaïm-Hofbauer-Sorin [10] and Duchi-Ruan [46].
- domain assumption Weak Sard property for definable functions: the set of Clarke critical values is finite, via stratification plus Sard's theorem (cited to [17, Coro. 9(ii)] and [30, Rem. 3.1.5]).
- ad hoc to paper About 89% of the 400 activation functions surveyed by Kunc and Kléma are definable in one of the listed o-minimal structures, by direct inspection.
Cite this review
Pith. "Pith review of Deep Learning as the Disciplined Construction of Tame Objects." pith.science (2026). https://pith.science/paper/OBLIM2QG
@misc{pith2026250918025,
author = {Pith},
title = {Pith review of: Deep Learning as the Disciplined Construction of Tame Objects},
year = {2026},
howpublished = {\url{https://pith.science/paper/OBLIM2QG}},
note = {Machine review of arXiv:2509.18025}
}
read the original abstract
One can see deep-learning models as compositions of functions within the so-called tame geometry. In this expository note, we give an overview of some topics at the interface of tame geometry (also known as o-minimality), optimization theory, and deep learning theory and practice. To do so, we gradually introduce the concepts and tools used to build convergence guarantees for stochastic gradient descent in a general nonsmooth nonconvex, but tame, setting. This illustrates some ways in which tame geometry is a natural mathematical framework for the study of AI systems, especially within Deep Learning.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Lipschitzian SLLNs for random functions
Empirical averages of random locally Lipschitz functions converge to their expectation in the Lipschitz pseudometric under separability or NIP-definability conditions, giving uniform convergence of subdifferentials.
-
Fast approximation and learning of binary classification tasks in o-minimal structures using ReLU neural networks
ReLU networks approximate traceable definable subsets of the unit cube in L^p with size O(ε^{-p(n-1)/m}) and yield ERM learning rates of order N^{-m/(m+pn-p)} for hinge loss under uniform component bounds.
Reference graph
Works this paper leans on
-
[1]
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al.,{TensorFlow}: a system for{Large-Scale} machine learning, 12th USENIX symposium on operating systems design and implementation (OSDI 16), 2016, pp. 265–283
2016
-
[2]
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner, Machine bias, Ethics of data and analytics, Auerbach Publications, 2022, pp. 254–264
2022
-
[3]
195, Princeton University Press, Princeton, NJ, 2017
Matthias Aschenbrenner, Lou van den Dries, and Joris van der Hoeven, Asymptotic differential algebra and model theory of transseries , Annals of Mathematics Studies, vol. 195, Princeton University Press, Princeton, NJ, 2017. MR 3585498
2017
-
[4]
Matthias Aschenbrenner, Lou van den Dries, and Joris van der Hoeven, Maximal hardy fields, 2025
2025
-
[5]
Francis Bach, Learning theory from first principles , Adaptive Computation and Machine Learning, The MIT Press, Cambridge, Massachusetts, 2024
2024
-
[6]
Bachute and Javed M
Mrinal R. Bachute and Javed M. Subhedar, Autonomous Driving Architectures: Insights of Machine Learning and Deep Learning Algorithms , Machine Learning with Applications 6 (2021), 100164
2021
-
[7]
Gilles Bareilles, Johannes Aspman, Jiri Nemecek, and Jakub Marecek, Piecewise polynomial regression of tame functions via integer programming , arXiv preprint arXiv:2311.13544 (2023)
arXiv 2023
-
[8]
Gilles Bareilles, Franck Iutzeler, and Jérôme Malick, Newton acceleration on manifolds identified by proximal gradient methods, Mathematical Programming (2022)
2022
Show all 109 references
-
[9]
Amir Beck, First-Order Methods in Optimization , Society for Industrial and Applied Mathematics, Philadel- phia, PA, October 2017
2017
-
[10]
4, 673–695
Michel Benaïm, Josef Hofbauer, and Sylvain Sorin, Stochastic approximations and differential inclusions, part ii: Applications, Mathematics of Operations Research 31 (2006), no. 4, 673–695
2006
-
[11]
G. C. Bento, B. S. Mordukhovich, T. S. Mota, and Yu Nesterov,Convergence of Descent Optimization Algorithms under Polyak-Ł ojasiewicz-Kurdyka Conditions , February 2025
2025
-
[12]
Raghu Reddy, Machine learning techniques for credit risk evaluation: A systematic literature review, Journal of Banking and Financial Technology 4 (2020), no
Siddharth Bhatore, Lalit Mohan, and Y. Raghu Reddy, Machine learning techniques for credit risk evaluation: A systematic literature review, Journal of Banking and Financial Technology 4 (2020), no. 1, 111–138
2020
-
[13]
3, 1117–1147
Pascal Bianchi, Walid Hachem, and Sholom Schechtman, Convergence of Constant Step Stochastic Gradient Descent for Non-Smooth Non-Convex Functions, Set-Valued and Variational Analysis30 (2022), no. 3, 1117–1147
2022
-
[14]
4, 1173–1178
Ricardo Bianconi, Nondefinability results for expansions of the field of real numbers by the exponential function and by the restricted sine function , The Journal of Symbolic Logic 62 (1997), no. 4, 1173–1178
1997
-
[15]
Bochnak, M
J. Bochnak, M. Coste, and M.F. Roy, Real algebraic geometry , Ergebnisse der Mathematik und ihrer Grenzge- biete. 3. Folge / A Series of Modern Surveys in Mathematics, Springer Berlin Heidelberg, 2013
2013
-
[16]
Jérôme Bolte, Aris Daniilidis, and Adrian Lewis, Tame functions are semismooth, Math. Program. 117 (2009), no. 1-2, 5–19. MR 2421297
2009
-
[17]
Jérôme Bolte, Aris Daniilidis, Adrian Lewis, and Masahiro Shiota,Clarke subgradients of stratifiable functions, SIAM J. Optim. 18 (2007), no. 2, 556–572. MR 2338451
2007
-
[18]
Jérôme Bolte and Edouard Pauwels, A mathematical model for automatic differentiation in machine learning , Advances in Neural Information Processing Systems 33 (2020), 10809–10819
2020
-
[19]
1, 19–51
, Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning, Mathematical Programming 188 (2021), no. 1, 19–51
2021
-
[20]
1, 553–603
, Curiosities and counterexamples in smooth convex optimization , Mathematical Programming 195 (2022), no. 1, 553–603
2022
-
[21]
48, Springer, 2009
Vivek S Borkar, Stochastic approximation: a dynamical systems viewpoint , vol. 48, Springer, 2009
2009
-
[22]
1, 258–280
Michael Boshernitzan, Hardy fields and existence of transexponential functions , Aequationes mathematicae 30 (1986), no. 1, 258–280
1986
-
[23]
Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning , SIAM Review 60 (2018), no
Léon Bottou, Frank E. Curtis, and Jorge Nocedal, Optimization Methods for Large-Scale Machine Learning , SIAM Review 60 (2018), no. 2, 223–311
2018
-
[24]
Nicolas Boumal, An introduction to optimization on smooth manifolds , Cambridge University Press, 2023
2023
-
[25]
bradley- williams, i
David Bradley-Williams and Immanuel Halupczok, Riso-stratifications and a tree invariant: D. bradley- williams, i. halupczok, Selecta Mathematica 31 (2025), no. 3, 52
2025
-
[26]
3, 319–332
Hunter Chase and James Freitag, Model theory and machine learning , Bulletin of Symbolic Logic 25 (2019), no. 3, 319–332
2019
-
[27]
Clarke, Optimization and Nonsmooth Analysis, Society for Industrial and Applied Mathematics, January 1990
Frank H. Clarke, Optimization and Nonsmooth Analysis, Society for Industrial and Applied Mathematics, January 1990
1990
-
[28]
Gabriel Conant, Forking and dividing, https://www.forkinganddividing.com/, Accessed: 2025-09-22
2025
-
[29]
DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 33
Michel Coste, An introduction to o-minimal geometry, Istituti editoriali e poligrafici internazionali Pisa, 2000. DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 33
2000
-
[30]
thesis, Migration - université en cours d’affectation, December 2001
Didier d’Acunto, Sur les courbes intégrales du champ de gradient , Ph.D. thesis, Migration - université en cours d’affectation, December 2001
2001
-
[31]
Damek Davis, Dmitriy Drusvyatskiy, and Liwei Jiang, Active manifolds, stratifications, and convergence to local minima in nonsmooth optimization , Foundations of Computational Mathematics (2025), 1–83
2025
-
[32]
Lee, Stochastic Subgradient Method Converges on Tame Functions, Foundations of Computational Mathematics 20 (2020), no
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee, Stochastic Subgradient Method Converges on Tame Functions, Foundations of Computational Mathematics 20 (2020), no. 1, 119–154
2020
-
[33]
1, 79–138
Jan Denef and Lou van den Dries, P-adic and real subanalytic sets , Annals of Mathematics 128 (1988), no. 1, 79–138
1988
-
[34]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova,Bert: Pre-training of deep bidirectional transformers for language understanding , 2019
2019
-
[35]
Martel, Joëlle Pineau, Peter Railton, Christine Tappolet, and Nathalie Voarino, Montréal Declaration for a responsible development of artificial intelligence , (2018)
Marc-Antoine Dilhac, Catherine Régis, Christophe Abrassart, Yoshua Bengio, Guillaume Chicoisne, Nathalie De Marcellis-Warin, Sébastien Gambs, Vincent Gautrais, Martin Gibert, Lyse Langlois, François Laviolette, Pascale Lehoux, Jocelyn Maclure, Marie D. Martel, Joëlle Pineau, P...
2018
-
[36]
Logic Found
Lou van den Dries, Remarks on Tarski’s problem concerning(R,+,·, exp), Logic colloquium ’82 (Florence, 1982), Stud. Logic Found. Math., vol. 112, North-Holland, Amsterdam, 1984, pp. 97–121. MR 762106
1982
-
[37]
, A generalization of the Tarski-Seidenberg theorem, and some nondefinability results , Bull. Amer. Math. Soc. (N.S.) 15 (1986), no. 2, 189–193. MR 854552
1986
-
[38]
248, Cambridge University Press, Cambridge, 1998
, Tame topology and o-minimal structures , London Mathematical Society Lecture Note Series, vol. 248, Cambridge University Press, Cambridge, 1998. MR 1633348
1998
-
[39]
Press, Somerville, MA, 1999, pp
, O-minimal structures and real analytic geometry , Current developments in mathematics, 1998 (Cambridge, MA), Int. Press, Somerville, MA, 1999, pp. 105–152
1998
-
[40]
Lou van den Dries, Angus Macintyre, and David Marker, The elementary theory of restricted analytic fields with exponentiation, Ann. of Math. (2) 140 (1994), no. 1, 183–205. MR 1289495
1994
-
[41]
London Math
, Logarithmic-exponential power series, J. London Math. Soc. (2) 56 (1997), no. 3, 417–434. MR 1610431
1997
-
[42]
Lou van den Dries and Chris Miller, On the real exponential field with restricted analytic functions , Israel Journal of Mathematics 85 (1994), 19–56
1994
-
[43]
2, 497–540
, Geometric categories and o-minimal structures, Duke Mathematical Journal 84 (1996), no. 2, 497–540
1996
-
[44]
3, 513–565
Lou van den Dries and Patrick Speissegger, The field of reals with multisummable series and the exponential function, Proceedings of the London Mathematical Society 81 (2000), no. 3, 513–565
2000
-
[45]
Shiv Ram Dubey, Satish Kumar Singh, and Bidyut Baran Chaudhuri, Activation functions in deep learning: A comprehensive survey and benchmark , Neurocomputing 503 (2022), 92–108
2022
-
[46]
4, 3229–3259
John C Duchi and Feng Ruan, Stochastic methods for composite and weakly convex optimization problems , SIAM Journal on Optimization 28 (2018), no. 4, 3229–3259
2018
-
[47]
6, Heldermann, Berlin, 1989
Ryszard Engelking, General topology, Sigma series in pure mathematics, vol. 6, Heldermann, Berlin, 1989
1989
-
[48]
Haishuo Fang, Ji-Ung Lee, Nafise Sadat Moosavi, and Iryna Gurevych,Transformers with learnable activation functions, 2023
2023
-
[49]
4, 282–291
Andrei M Gabrielov, Projections of semi-analytic sets , Functional Analysis and its applications 2 (1968), no. 4, 282–291
1968
-
[50]
L. M. Graña Drummond and Y. Peterzil,The central path in smooth convex semidefinite programs, Optimization 51 (2002), no. 2, 207–233. MR 1928037
2002
-
[51]
155– 210
Michael Grant, Stephen Boyd, and Yinyu Ye, Disciplined Convex Programming, Global Optimization: From Theory to Implementation (Leo Liberti and Nelson Maculan, eds.), Springer US, Boston, MA, 2006, pp. 155– 210
2006
-
[52]
Andreas Griewank and Andrea Walther, Evaluating derivatives: Principles and techniques of algorithmic differentiation, 2nd ed ed., Society for Industrial and Applied Mathematics, Philadelphia, PA, 2008
2008
-
[53]
A Grothendieck, Esquisse d’un programme, London Math. Soc. Lect. Note Ser. 1 (1984), 7–48
1984
-
[54]
Halická, E
M. Halická, E. de Klerk, and C. Roos, On the convergence of the central path in semidefinite optimization , SIAM J. Optim. 12 (2002), no. 4, 1090–1099. MR 1922510
2002
-
[55]
5, 1745–1780
Martin Helmer and Vidit Nanda,Conormal Spaces and Whitney Stratifications, Foundations of Computational Mathematics 23 (2023), no. 5, 1745–1780
2023
-
[56]
Dan Hendrycks and Kevin Gimpel, Gaussian error linear units (gelus) , 2016
2016
-
[57]
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal, Fundamentals of Convex Analysis, Springer Berlin Heidelberg, Berlin, Heidelberg, 2001
2001
-
[58]
Lê Nguyên Hoang and El Mahdi El Mhamdi,Le fabuleux chantier: Rendre l’intelligence artificielle robustement bénéfique, EDP Sciences, Les Ulis, 2021. 34 G. BAREILLES, A. GEHRET, J. ASPMAN, J. LEPŠOVÁ, AND J. MAREČEK
2021
-
[59]
A. D. Ioffe, An invitation to tame optimization , SIAM J. Optim. 19 (2008), no. 4, 1894–1917. MR 2486055
2008
-
[60]
31, Curran Associates, Inc., 2018
Sham M Kakade and Jason D Lee, Provably Correct Automatic Sub-Differentiation for Qualified Programs , Advances in Neural Information Processing Systems, vol. 31, Curran Associates, Inc., 2018
2018
-
[61]
Kechris, Classical descriptive set theory , Graduate Texts in Mathematics, vol
Alexander S. Kechris, Classical descriptive set theory , Graduate Texts in Mathematics, vol. 156, Springer- Verlag, New York, 1995. MR 1321597
1995
-
[62]
Knight, Anand Pillay, and Charles Steinhorn, Definable sets in ordered structures
Julia F. Knight, Anand Pillay, and Charles Steinhorn, Definable sets in ordered structures. II , Trans. Amer. Math. Soc. 295 (1986), no. 2, 593–605. MR 833698
1986
-
[63]
Julian Kranz, Davide Gallon, Steffen Dereich, and Arnulf Jentzen, Sad neural networks: Divergent gradient flows and asymptotic optimality via o-minimal structures , 2025
2025
-
[64]
Lothar Sebastian Krapp and Laura Wirth, Measurability in the fundamental theorem of statistical learning , arXiv preprint arXiv:2410.10243 (2024)
2024
-
[65]
Vladimír Kunc and Jiří Kléma,Three decades of activations: A comprehensive survey of 400 activation functions for neural networks, 2024
2024
-
[66]
Krzysztof Kurdyka, Olivier Le Gal, and Xuan Viet Nhan Nguyen,Tangent cones and𝐶1 regularity of definable sets, J. Math. Anal. Appl. 457 (2018), no. 1, 978–990. MR 3702738
2018
-
[67]
Lasserre, Global Optimization with Polynomials and the Problem of Moments , SIAM Journal on Optimization 11 (2001), no
Jean B. Lasserre, Global Optimization with Polynomials and the Problem of Moments , SIAM Journal on Optimization 11 (2001), no. 3, 796–817
2001
-
[68]
Jean Bernard Lasserre, An introduction to polynomial and semi-algebraic optimization , Cambridge Texts in Applied Mathematics, Cambridge University Press, 2015
2015
-
[69]
Olivier Le Gal, A generic condition implying o-minimality for restricted 𝐶∞-functions, Ann. Fac. Sci. Toulouse Math. (6) 19 (2010), no. 3-4, 479–492. MR 2790804
2010
-
[70]
Wonyeol Lee, Hangyeol Yu, Xavier Rival, and Hongseok Yang,On correctness of automatic differentiation for non-differentiable functions, Advances in Neural Information Processing Systems 33 (2020), 6719–6730
2020
-
[71]
Lewis and Michael L
Adrian S. Lewis and Michael L. Overton, Nonsmooth optimization via quasi-Newton methods , Mathematical Programming 141 (2013), no. 1, 135–163
2013
-
[72]
Lewis and Tonghua Tian, The Structure of Conservative Gradient Fields , SIAM Journal on Opti- mization 31 (2021), no
Adrian S. Lewis and Tonghua Tian, The Structure of Conservative Gradient Fields , SIAM Journal on Opti- mization 31 (2021), no. 3, 2080–2083
2021
-
[73]
1, 81–109
Leo Liberti, Undecidability and hardness in mixed-integer nonlinear programming , RAIRO-Operations Re- search 53 (2019), no. 1, 81–109
2019
-
[74]
1, 401–409
Ta Loi, Whitney stratification of sets definable in the structure ℝexp, Banach Center Publications 33 (1996), no. 1, 401–409
1996
-
[75]
Ta Lê Loi, Verdier and strict Thom stratifications in o-minimal structures , Illinois J. Math. 42 (1998), no. 2, 347–356. MR 1612771
1998
-
[76]
Angus Macintyre and Eduardo D Sontag, Finiteness results for sigmoidal “neural” networks , Proceedings of the twenty-fifth annual ACM symposium on Theory of computing, 1993, pp. 325–334
1993
-
[77]
Angus Macintyre and A. J. Wilkie, On the decidability of the real exponential field , Kreiseliana, A K Peters, Wellesley, MA, 1996, pp. 441–467. MR 1435773
1996
-
[78]
05, WORLD SCIENTIFIC (EUROPE), May 2023
Victor Magron and Jie Wang, Sparse Polynomial Optimization: Theory and Practice , Series on Optimization and Its Applications, vol. 05, WORLD SCIENTIFIC (EUROPE), May 2023
2023
-
[79]
Pure Appl
Chris Miller, Expansions of the real field with power functions , Ann. Pure Appl. Logic 68 (1994), no. 1, 79–94. MR 1278550
1994
-
[80]
1, 79–94
, Expansions of the real field with power functions , Annals of Pure and Applied Logic 68 (1994), no. 1, 79–94
1994
-
[81]
, Exponentiation is hard to avoid , Proc. Amer. Math. Soc. 122 (1994), no. 1, 257–259. MR 1195484
1994
-
[82]
Commun., vol
, Basics of o-minimality and Hardy fields , Lecture notes on o-minimal structures and real analytic geometry, Fields Inst. Commun., vol. 62, Springer, New York, 2012, pp. 43–69. MR 2976990
2012
-
[83]
4, 337–341
Aleš Nekvinda and Luděk Zajíček, A simple proof of the Rademacher theorem , Časopis pro pěstování matematiky 113 (1988), no. 4, 337–341
1988
-
[84]
137, Springer International Publishing, Cham, 2018
Yurii Nesterov, Lectures on Convex Optimization , Springer Optimization and Its Applications, vol. 137, Springer International Publishing, Cham, 2018
2018
-
[85]
2, 381–389
Nhan Nguyen, Saurabh Trivedi, and David Trotman, A geometric proof of the existence of definable whitney stratifications, Illinois Journal of Mathematics 58 (2014), no. 2, 381–389
2014
-
[86]
Chigozie Nwankpa, Winifred Ijomah, Anthony Gachagan, and Stephen Marshall, Activation functions: Comparison of trends in practice and research for deep learning , CoRR abs/1811.03378 (2018)
2018 arXiv
-
[87]
DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 35
Adele Padgett and Patrick Speissegger, Definability of complex functions in o-minimal structures , 2025. DEEP LEARNING AS THE DISCIPLINED CONSTRUCTION OF TAME OBJECTS 35
2025
-
[88]
MR 4511519
Adele Lee Padgett, Sublogarithmic-Transexponential Series, ProQuest LLC, Ann Arbor, MI, 2022, Thesis (Ph.D.)–University of California, Berkeley. MR 4511519
2022
-
[89]
European Parliament, Artificial Intelligence Act, 2024
2024
-
[90]
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, Automatic differentiation in PyTorch, (2017)
2017
-
[91]
Anand Pillay and Charles Steinhorn, Definable sets in ordered structures , Bull. Amer. Math. Soc. (N.S.) 11 (1984), no. 1, 159–162. MR 741730
1984
-
[92]
II, Trans
, Definable sets in ordered structures. II, Trans. Amer. Math. Soc.295 (1986), no. 2, 565–592. MR 833697
1986
-
[93]
III, Trans
, Definable sets in ordered structures. III, Trans. Amer. Math. Soc.309 (1988), no. 2, 469–476. MR 943306
1988
-
[94]
317, Springer Science & Business Media, 2009
R Tyrrell Rockafellar and Roger J-B Wets, Variational analysis, vol. 317, Springer Science & Business Media, 2009
2009
-
[95]
Jean-Philippe Rolin, Establishing the o-minimality for expansions of the real field , Model theory with applications to algebra and analysis. Vol. 1, London Math. Soc. Lecture Note Ser., vol. 349, Cambridge Univ. Press, Cambridge, 2008, pp. 249–282. MR 2441383
2008
-
[96]
, A survey on o-minimal structures , Real Algebraic Geometry (2011), 105
2011
-
[97]
, Construction of o-minimal structures from quasianalytic classes , Lecture Notes on O-minimal Structures and Real Analytic Geometry, Springer, 2012, pp. 71–109
2012
-
[98]
1, 258–266
Maxwell Rosenlicht, On the value group of a differential valuation , American Journal of Mathematics 101 (1979), no. 1, 258–266
1979
-
[99]
Stuart Jonathan Russell, Human compatible: Artificial intelligence and the problem of control , Viking, New York, 2019
2019
-
[100]
12, 883–890
Arthur Sard, The measure of the critical values of differentiable maps , Bulletin of the American Mathematical Society 48 (1942), no. 12, 883–890
1942
-
[101]
2, 365–374
Abraham Seidenberg, A new decision method for elementary algebra , Annals of Mathematics 60 (1954), no. 2, 365–374
1954
-
[102]
92, Elsevier, 1990
Saharon Shelah, Classification theory: and the number of non-isomorphic models , vol. 92, Elsevier, 1990
1990
-
[103]
Reine Angew
Patrick Speissegger, The Pfaffian closure of an o-minimal structure , J. Reine Angew. Math.508 (1999), 189–211. MR 1676876
1999
-
[104]
MR 28796
Alfred Tarski, A Decision Method for Elementary Algebra and Geometry , The Rand Corporation, Santa Monica, CA, 1948. MR 28796
1948
-
[105]
Marcus Tressl, Introduction to o-minimal structures and an application to neural network learning , Course outline for LMS and EPSRC Short Instructional Course on Model Theory at the University of Leeds (2010)
2010
-
[106]
David Trotman, Stratification theory, Handbook of geometry and topology of singularities I, Springer, 2020, pp. 243–273
2020
-
[107]
Lasserre, Certifying global optimality of AC-OPF solutions via sparse polynomial optimization, Electric Power Systems Research 213 (2022), 108683
Jie Wang, Victor Magron, and Jean B. Lasserre, Certifying global optimality of AC-OPF solutions via sparse polynomial optimization, Electric Power Systems Research 213 (2022), 108683
2022
-
[108]
A. J. Wilkie, Model completeness results for expansions of the ordered field of real numbers by restricted Pfaffian functions and the exponential function , J. Amer. Math. Soc. 9 (1996), no. 4, 1051–1094. MR 1398816
1996
-
[109]
Julio Zamora Esquivel, Adan Cruz Vargas, Rodrigo Camacho Perez, Paulo Lopez Meyer, Hector Cordourier, and Omesh Tickoo, Adaptive activation functions using fractional calculus , Proceedings of the IEEE/CVF international conference on computer vision workshops, 2019, pp. 0–0
2019
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.