Pith. sign in

REVIEW 1 major objections 2 minor 1 cited by

The stochastic complexity of regular non-smooth models is well-posed and matches automatic differentiation outputs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Establishes measure-theoretic foundations for NML in regular non-smooth models and introduces the PDL-PPMH geometric MCMC sampler to compute stochastic complexity exactly.

T0 review reviewed 2026-06-30 challenge →

load-bearing objection This paper extends NML to non-smooth PDL estimators via geometric measure theory and adds a geometric MCMC sampler, but the coarea formula extension to conservative Jacobians needs tighter justification on uniqueness. the 1 major comments →

arxiv 2605.24477 v1 pith:XE7B4PKC submitted 2026-05-23 cs.LG cs.ITmath.ITmath.STstat.TH

The Normalized Maximum Likelihood for Regular Non-Smooth Models: Measure-Theoretic Foundations and Geometric Sampling

classification cs.LG cs.ITmath.ITmath.STstat.TH
keywords normalized maximum likelihoodstochastic complexitynon-smooth modelspath-differentiable Lipschitzgeometric measure theoryMetropolis-Hastings samplerLasso regressionautomatic differentiation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes a measure-theoretic definition of the normalized maximum likelihood codelength for non-smooth estimators that dominate modern machine learning. It applies the coarea formula to conservative Jacobians on path-differentiable Lipschitz models, proving that the resulting stochastic complexity integral is rigorously defined and consistent with automatic differentiation. A new geometric MCMC sampler called PDL-PPMH is introduced to compute the quantity exactly by traversing non-differentiable level sets. A reader would care because the method yields a data-efficient model selection criterion that matches cross-validation performance on high-dimensional Lasso without splitting the data.

Core claim

By bridging the coarea formula with conservative Jacobians for regular path-differentiable Lipschitz estimators, the stochastic complexity for non-smooth models is well-posed and theoretically consistent with the outputs of modern Automatic Differentiation. The Propose-and-Project Metropolis-Hastings sampler enables exact computation by using a stochastic tangent space proposal and a convergent non-smooth projection step, demonstrated on a P=2000 Lasso posterior while showing that the exact NML criterion achieves statistically indistinguishable predictive optima from cross-validation without data splitting.

What carries the argument

The coarea formula applied together with conservative Jacobians on regular path-differentiable Lipschitz estimators, which defines the NML integral over non-smooth level sets.

Load-bearing premise

The estimators must be regular path-differentiable Lipschitz models so that the coarea formula can be applied together with conservative Jacobians.

What would settle it

A direct comparison on a simple Lasso or Sparse SVM example where the NML value obtained from the PDL-PPMH sampler differs from the value produced by automatic differentiation on the same model.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The NML codelength becomes computable for non-smooth estimators such as Lasso and Sparse SVMs.
  • Exact NML supplies a data-efficient alternative to cross-validation that requires no data splitting.
  • The sampler scales to high-dimensional problems such as P=2000 while quantifying the exactness-versus-mixing-time trade-off.
  • The framework supports theoretical analysis of the NML codelength for any regular non-smooth model.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same measure-theoretic construction may apply to other non-smooth optimization routines used in deep learning.
  • The consistency result could justify replacing cross-validation with NML in settings where data are scarce.
  • The mixing-time scaling observed on Lasso may generalize to other PDL models and guide practical sampler tuning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. The paper develops a measure-theoretic framework extending the coarea formula to conservative Jacobians of path-differentiable Lipschitz (PDL) estimators, proving that the Normalized Maximum Likelihood (NML) stochastic complexity is well-posed for regular non-smooth models and consistent with automatic differentiation outputs. It introduces the PDL-PPMH geometric MCMC sampler with stochastic tangent proposals and a convergent non-smooth projection solver, and demonstrates exact NML computation on high-dimensional Lasso (P=2000) as a data-efficient alternative to cross-validation.

Significance. If the central extension holds, the result would enable exact NML computation for non-smooth estimators common in modern ML (Lasso, sparse SVMs), providing a theoretically grounded alternative to cross-validation without data splitting. The geometric sampler and its justification constitute a technical contribution to sampling on non-differentiable level sets.

major comments (1)
  1. [measure-theoretic foundations section bridging coarea formula with conservative Jacobians] The extension of the coarea formula to set-valued conservative Jacobians for PDL models asserts a well-posed NML via a unique Hausdorff measure on level sets, but provides no explicit invariance argument or selection theorem showing independence from the choice of measurable selection of the Jacobian. Different selections could yield different measures, directly undermining the well-posedness claim and the asserted consistency with AD outputs.
minor comments (2)
  1. [Introduction and preliminaries] Clarify the precise definition of 'regular' PDL models and the conditions under which the conservative Jacobian is applied, ideally with a dedicated preliminary subsection.
  2. [empirical evaluation] The empirical section on P=2000 Lasso should include explicit mixing-time diagnostics and error bounds for the sampler to support the scaling claims.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the detailed and constructive report. The single major comment raises a substantive point about invariance under measurable selections of the conservative Jacobian. We address it directly below and commit to a targeted revision that strengthens the measure-theoretic foundations without altering the paper's core claims.

read point-by-point responses
  1. Referee: [measure-theoretic foundations section bridging coarea formula with conservative Jacobians] The extension of the coarea formula to set-valued conservative Jacobians for PDL models asserts a well-posed NML via a unique Hausdorff measure on level sets, but provides no explicit invariance argument or selection theorem showing independence from the choice of measurable selection of the Jacobian. Different selections could yield different measures, directly undermining the well-posedness claim and the asserted consistency with AD outputs.

    Authors: We agree that an explicit invariance argument is not stated as a standalone lemma in the current manuscript, even though the well-posedness claim rests on the uniqueness of the induced Hausdorff measure. In the revised version we will insert a new short lemma (placed immediately after the statement of the extended coarea formula) showing that, for path-differentiable Lipschitz functions, any two measurable selections of the conservative Jacobian differ on a set of Lebesgue measure zero in the parameter space; consequently they generate identical (n-1)-Hausdorff measures on almost every level set. The proof relies on the upper semicontinuity of the conservative Jacobian together with the standard coarea formula applied to the Clarke subdifferential. This addition will also make the consistency with automatic-differentiation outputs fully rigorous, since every AD output lies inside the conservative Jacobian. We view the revision as a clarification rather than a change of substance. revision: yes

Circularity Check

0 steps flagged

No circularity; derivation applies external geometric measure theory results independently.

full rationale

The paper extends NML computation to PDL models by invoking the classical coarea formula and conservative Jacobians from geometric measure theory, then introduces a separately justified MCMC sampler (PDL-PPMH) with stochastic tangent proposals and non-smooth projections. No equations reduce a claimed prediction to a fitted input by construction, no self-citation chains bear the central well-posedness argument, and no uniqueness theorems or ansatzes are smuggled via prior author work. The framework is self-contained against external GMT benchmarks rather than internally referential.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 1 invented entities

The central claim rests on the applicability of the coarea formula to PDL estimators and the convergence properties of the proposed non-smooth projection solver; no free parameters are mentioned.

axioms (2)
  • domain assumption Coarea formula applies to the level sets of regular path-differentiable Lipschitz estimators when paired with conservative Jacobians
    Invoked to prove that stochastic complexity remains well-posed for non-smooth models.
  • domain assumption The non-smooth projection solver converges for the level sets arising in PDL estimators
    Required to guarantee that the PDL-PPMH sampler produces valid samples.
invented entities (1)
  • PDL-PPMH sampler no independent evidence
    purpose: Geometric MCMC algorithm to traverse non-differentiable level sets of the maximum likelihood estimator
    New algorithm introduced to compute the NML exactly; no independent evidence outside the paper is provided.

reviewed 2026-06-30 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Normalized Maximum Likelihood for Regular Non-Smooth Models: Measure-Theoretic Foundations and Geometric Sampling." pith.science (2026). https://pith.science/paper/XE7B4PKC

@misc{pith2026260524477,
  author       = {Pith},
  title        = {Pith review of: The Normalized Maximum Likelihood for Regular Non-Smooth Models: Measure-Theoretic Foundations and Geometric Sampling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XE7B4PKC}},
  note         = {Machine review of arXiv:2605.24477}
}
Share X Bluesky LinkedIn Reddit HN
abstract

The Normalized Maximum Likelihood (NML) codelength, or stochastic complexity, represents a principled criterion for universal coding. While recent coarea-based formulations provided a calculation method for smooth models, this framework collapses for the non-smooth estimators ubiquitous in modern machine learning (e.g., Lasso, Sparse SVMs). In this work, we provide a rigorous framework for computing the NML for regular path-differentiable Lipschitz (PDL) estimators. By applying classical geometric measure theory and bridging the coarea formula with conservative Jacobians, we prove that the stochastic complexity for non-smooth models is well-posed and theoretically consistent with the outputs of modern Automatic Differentiation. To compute this quantity exactly, we introduce the Propose-and-Project Metropolis-Hastings (PDL-PPMH) sampler, a geometric MCMC algorithm capable of traversing the non-differentiable level sets of the maximum likelihood estimator. We theoretically justify its components, including a stochastic tangent space proposal and a provably convergent non-smooth projection solver. We demonstrate the method's robustness by sampling from a high-dimensional Lasso posterior ($P=2000$), while simultaneously quantifying the computational scaling that governs the trade-off between exactness and mixing time. Crucially, we empirically demonstrate that our exact NML criterion provides a highly data-efficient alternative to cross-validation, achieving statistically indistinguishable predictive optima without requiring data splitting. Altogether, our work paves the way for the theoretical analysis of the NML codelength for regular non-smooth models.

Figures

Figures reproduced from arXiv: 2605.24477 by Gary P. T. Choi, Trenton Lau.

Figure 1
Figure 1. Figure 1: Algorithmic Scaling (Time vs. k). Cost is stratified by N and scales with the active set size [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 4
Figure 4. Figure 4: displays the discrete trace of the sampler cor￾rectly stepping down the faces of the polytope to locate the posterior mean. To rigorously prove that the sampler is not trapped in local minima, which is a common failure mode in correlated designs, we present the Autocorrelation Function (ACF) in [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Autocorrelation Function. Rapid decay to zero indicates highly efficient sampling and independence [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 8
Figure 8. Figure 8: Truth Recovery. Probability of selecting exactly k = 5. exactly at the λ corresponding to the peak truth recovery and the stabilization of the NML complexity. Altogether, these results confirm that the generalized measure-theoretic NML framework operates exactly as de￾sired: the information-theoretic optimum coincides accurately with both the structural truth of the data and the predictive optimum, validat… view at source ↗
Figure 11
Figure 11. Figure 11: Empirical validation of non-smooth NML asymptotics. (Left) The exact computed complexity scales perfectly linearly with log N. The empirical slope matches the theoretical k 2 penalty, demonstrating that classical dimensional penalties hold for non-smooth regular models. (Right) The residual (Complexity − k 2 log N) converges to a constant. This explicitly bounds the non-smooth geometric penalty, proving t… view at source ↗
Figure 10
Figure 10. Figure 10: Validation. Generalization error (MSE) minimizes at the NML optimum. for regular non-smooth models on the active manifold. Further￾more, by plotting the residual (Total Complexity − k 2 log N) in [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 12
Figure 12. Figure 12: Model Complexity Penalties for N = 100 (Left) and N = 1000 (Right). Standard criteria (BIC, AIC) behave as rigid step functions of k. The Exact PDL-NML computes a continuous geometric volume. At small N, the exact geometry vastly diverges from asymptotic approximations; at large N, it converges towards the theoretical bounds [PITH_FULL_IMAGE:figures/full_fig_p018_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Predictive Generalization and Statistical Significance for N = 100 (Left) and N = 250 (Right). At N = 250, Exact NML matches the 5-Fold CV predictive optimum perfectly without holding out data (p > 0.05). At N = 100, data splitting harms CV, allowing the analytically computed Exact NML to identify a model with statistically significantly lower MSE. work, extending its rigorous application to the non-smoot… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Exact Schur-Sylvester Dimensionality Reductions for Non-Smooth Stochastic Complexity and Manifold Sampling

    cs.LG 2026-06 unverdicted novelty 5.0

    Exact Schur-Sylvester reductions lower PPMH projection and volume costs for non-smooth NML from O(N^3) to O(k^3 + N^2 k), with reported 14,100x speedups on high-dimensional data while preserving double-precision equivalence.

Reference graph

Works this paper leans on

64 extracted references · 64 canonical work pages · cited by 1 Pith paper · 1 internal anchor

  1. [1]

    Rissanen,Stochastic Complexity in Statistical Inquiry

    J. Rissanen,Stochastic Complexity in Statistical Inquiry. World Scien- tific, 1989

  2. [2]

    P. D. Gr ¨unwald,The Minimum Description Length Principle. MIT Press, 2007

  3. [3]

    Minimum description length revisited,

    P. Gr ¨unwald and T. Roos, “Minimum description length revisited,” International Journal of Mathematics for Industry, vol. 11, no. 01, p. 1930001, 2019

  4. [4]

    Strong optimality of the normalized ML models as univer- sal codes and information in data,

    J. Rissanen, “Strong optimality of the normalized ML models as univer- sal codes and information in data,”IEEE Transactions on Information Theory, vol. 47, no. 5, pp. 1712–1717, 2001. 19

  5. [5]

    Model selection and the principle of minimum description length,

    M. H. Hansen and B. Yu, “Model selection and the principle of minimum description length,”Journal of the American Statistical Association, vol. 96, no. 454, pp. 746–774, 2001

  6. [6]

    Foundation of calculating nor- malized maximum likelihood for continuous probability models,

    A. Suzuki, K. Fukuzawa, and K. Yamanishi, “Foundation of calculating normalized maximum likelihood for continuous probability models,” arXiv preprint arxiv:2409.08387, 2024

  7. [7]

    Normalized maximum likelihood code-length on Riemannian manifold data spaces,

    K. Fukuzawa, A. Suzuki, and K. Yamanishi, “Normalized maximum likelihood code-length on Riemannian manifold data spaces,”IEEE Transactions on Information Theory, 2026

  8. [8]

    A wrapped normal distribution on hyperbolic space for gradient-based learning,

    Y . Nagano, S. Yamaguchi, Y . Fujita, and M. Koyama, “A wrapped normal distribution on hyperbolic space for gradient-based learning,” in Proceedings of the 36th International Conference on Machine Learning (K. Chaudhuri and R. Salakhutdinov, eds.), vol. 97 ofProceedings of Machine Learning Research, pp. 4693–4702, PMLR, 09–15 Jun 2019

  9. [9]

    The decomposed normalized maximum likelihood code-length criterion for selecting hier- archical latent variable models,

    K. Yamanishi, T. Wu, S. Sugawara, and M. Okada, “The decomposed normalized maximum likelihood code-length criterion for selecting hier- archical latent variable models,”Data Mining and Knowledge Discovery, vol. 33, no. 4, pp. 1017–1058, 2019

  10. [10]

    Dimensionality selection for hyper- bolic embeddings using decomposed normalized maximum likelihood code-length,

    R. Yuki, Y . Ike, and K. Yamanishi, “Dimensionality selection for hyper- bolic embeddings using decomposed normalized maximum likelihood code-length,”Knowledge and Information Systems, vol. 65, no. 12, pp. 5601–5634, 2023

  11. [11]

    An NML-based model selection criterion for general relational data modeling,

    Y . Sakai and K. Yamanishi, “An NML-based model selection criterion for general relational data modeling,” in2013 IEEE International Conference on Big Data, pp. 421–429, 2013

  12. [12]

    Watanabe,Algebraic Geometry and Statistical Learning Theory

    S. Watanabe,Algebraic Geometry and Statistical Learning Theory. Cambridge University Press, 2009

  13. [13]

    Deep learning is singular, and that’s good,

    S. Wei, D. Murfet, M. Gong, H. Li, J. Gell-Redman, and T. Quella, “Deep learning is singular, and that’s good,”IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 12, pp. 10473– 10486, 2022

  14. [14]

    Fisher information and stochastic complexity,

    J. Rissanen, “Fisher information and stochastic complexity,”IEEE Transactions on Information Theory, vol. 42, no. 1, pp. 40–47, 1996

  15. [15]

    Fourier-analysis-based form of nor- malized maximum likelihood: Exact formula and relation to complex Bayesian prior,

    A. Suzuki and K. Yamanishi, “Fourier-analysis-based form of nor- malized maximum likelihood: Exact formula and relation to complex Bayesian prior,”IEEE Transactions on Information Theory, vol. 67, no. 9, pp. 6164–6180, 2021

  16. [16]

    Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning,

    J. Bolte and E. Pauwels, “Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning,”Mathe- matical Programming, vol. 188, no. 1, pp. 19–51, 2021

  17. [17]

    Federer,Geometric Measure Theory

    H. Federer,Geometric Measure Theory. Berlin, Heidelberg, New York: Springer-Verlag, 1969

  18. [18]

    On the complexity of nonsmooth automatic differentiation,

    J. Bolte, R. Boustany, E. Pauwels, and B. Pesquet-Popescu, “On the complexity of nonsmooth automatic differentiation,” inInternational Conference on Learning Representations (ICLR), 2023

  19. [19]

    Universal sequential coding of single messages,

    Y . M. Shtar’kov, “Universal sequential coding of single messages,” Problems of Information Transmission, vol. 23, no. 3, pp. 3–17, 1987

  20. [20]

    MDL, Bayesian inference and the geometry of the space of probability distributions,

    V . Balasubramanian, “MDL, Bayesian inference and the geometry of the space of probability distributions,”Neural Computation, vol. 9, no. 2, pp. 349–368, 1997

  21. [21]

    R. T. Rockafellar and R. J.-B. Wets,Variational Analysis, vol. 317 of Grundlehren der mathematischen Wissenschaften. Springer Science & Business Media, 2009

  22. [22]

    van den Dries,Tame Topology and o-Minimal Structures

    L. van den Dries,Tame Topology and o-Minimal Structures. Cambridge University Press, 1998

  23. [23]

    Riemann manifold Langevin and Hamiltonian Monte Carlo methods,

    M. Girolami and B. Calderhead, “Riemann manifold Langevin and Hamiltonian Monte Carlo methods,”Journal of the Royal Statistical Society: Series B (Statistical Methodology), vol. 73, no. 2, pp. 123–214, 2011

  24. [24]

    The lasso problem and uniqueness,

    R. J. Tibshirani, “The lasso problem and uniqueness,”Electronic Journal of Statistics, vol. 7, pp. 1456–1490, 2013

  25. [25]

    Evidence slopes and effective dimension in singular linear models,

    K. Rao, “Evidence slopes and effective dimension in singular linear models,”arXiv preprint arXiv:2601.01238, 2026

  26. [26]

    Symmetry in neural network parameter spaces,

    B. Zhao, R. Walters, and R. Yu, “Symmetry in neural network parameter spaces,”Transactions on Machine Learning Research, 2025

  27. [27]

    A projected semismooth Newton method for a class of nonconvex composite programs with strong prox- regularity,

    J. Hu, K. Deng, J. Wu, and Q. Li, “A projected semismooth Newton method for a class of nonconvex composite programs with strong prox- regularity,”Journal of Machine Learning Research, vol. 25, no. 56, pp. 1–32, 2024

  28. [28]

    The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems,

    J. Bolte, A. Daniilidis, and A. Lewis, “The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems,”SIAM Journal on Optimization, vol. 17, no. 4, pp. 1205–1223, 2007

  29. [29]

    Strongly regular generalized equations,

    S. M. Robinson, “Strongly regular generalized equations,”Mathematics of Operations Research, vol. 5, no. 1, pp. 43–62, 1980

  30. [30]

    A nonsmooth version of Newton’s method,

    L. Qi and J. Sun, “A nonsmooth version of Newton’s method,”Mathe- matical Programming, vol. 58, no. 1, pp. 353–367, 1993

  31. [31]

    Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods,

    H. Attouch, J. Bolte, and B. F. Svaiter, “Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods,” Mathematical Programming, vol. 137, no. 1, pp. 91–129, 2013

  32. [32]

    A mathematical model for automatic differen- tiation in machine learning,

    J. Bolte and E. Pauwels, “A mathematical model for automatic differen- tiation in machine learning,”Advances in Neural Information Processing Systems, vol. 33, pp. 10809–10819, 2020

  33. [33]

    A robust gradient sampling algorithm for nonsmooth, nonconvex optimization,

    J. V . Burke, A. S. Lewis, and M. L. Overton, “A robust gradient sampling algorithm for nonsmooth, nonconvex optimization,”SIAM Journal on Optimization, vol. 15, no. 3, pp. 751–779, 2005

  34. [34]

    A survey of Monte Carlo methods for noisy and costly densities with application to reinforcement learning and ABC,

    F. Llorente, L. Martino, J. Read, and D. Delgado, “A survey of Monte Carlo methods for noisy and costly densities with application to reinforcement learning and ABC,”International Statistical Review, vol. 93, no. 1, pp. 18–61, 2025

  35. [35]

    Gradient sampling algorithm for subsmooth functions,

    D. Boskos, J. Cort ´es, and S. Mart ´ınez, “Gradient sampling algorithm for subsmooth functions,”arXiv preprint arXiv:2503.16638v1, 2025

  36. [36]

    Spectral gaps for a Metropolis–Hastings algorithm in infinite dimensions,

    M. Hairer, A. M. Stuart, and S. J. V ollmer, “Spectral gaps for a Metropolis–Hastings algorithm in infinite dimensions,”The Annals of Applied Probability, vol. 24, no. 6, pp. 2455–2490, 2014

  37. [37]

    On random- ized step sizes in Metropolis–Hastings algorithms,

    S. Grazzi, S. Livingstone, and L. Riou-Durand, “On random- ized step sizes in Metropolis–Hastings algorithms,”arXiv preprint arXiv:2601.19710v1, 2026

  38. [38]

    Learning capacity: A measure of the effective dimensionality of a model,

    D. Chen, W.-K. Chang, and P. Chaudhari, “Learning capacity: A measure of the effective dimensionality of a model,”arXiv preprint arXiv:2305.17332, 2024

  39. [39]

    Verzellesi,New and old sub-Riemannian challenges bridging analysis and geometry

    S. Verzellesi,New and old sub-Riemannian challenges bridging analysis and geometry. PhD thesis, Universit `a di Trento, 2024

  40. [40]

    Stability and exponential convergence of nonhomo- geneous Markov chains,

    A. Y . Mitrophanov, “Stability and exponential convergence of nonhomo- geneous Markov chains,”Journal of Applied Probability, vol. 42, no. 4, pp. 1043–1051, 2005

  41. [41]

    Explicit error bounds for Markov chain Monte Carlo,

    D. Rudolf, “Explicit error bounds for Markov chain Monte Carlo,” Dissertationes Mathematicae, vol. 485, pp. 1–93, 2012

  42. [42]

    Gradient sampling methods for nonsmooth optimization,

    J. V . Burke, F. E. Curtis, A. S. Lewis, M. L. Overton, and L. E. Sim ˜oes, “Gradient sampling methods for nonsmooth optimization,”Numerical nonsmooth optimization: State of the art algorithms, pp. 201–225, 2020. 20 APPENDIXA DERIVATION ANDSTABILITY OF THEPROJECTION DERIVATIVE The Metropolis-Hastings acceptance ratio in the PDL- PPMH sampler requires comp...

  43. [43]

    The distance between the ideal proposalyand the practical proposal˜yis bounded:∥y−˜y∥ ≤C pϵfeas for a constantC p. 21

  44. [44]

    The difference between the ideal acceptance probabil- ityα(x, y)and the practical one˜α(x, y)is bounded: |α(x, y)−˜α(x, y)| ≤C αϵfeas for a constantC α

  45. [45]

    The proposal densityq(x,·)is continuously differen- tiable with bounded derivatives. Proposition 3 (One-Step Error Bound):Under Assump- tion 3, there exists a constantK P >0such that for any state x, the total variation distance between the ideal and practical kernels is bounded by: ∥P(x,·)− ˜P(x,·)∥ TV ≤K P ϵfeas.(22) Proof Sketch:The TV distance can be ...

  46. [46]

    We remedy this by constructing a suitable extension of the function ˆθ

    Extension of the Lipschitz Function The central rectifiability result, Theorem 1, applies to Lips- chitz functions defined on open sets, but the domainXof our MLE is not necessarily open. We remedy this by constructing a suitable extension of the function ˆθ. The MLE is assumed to be Lipschitz continuous on its domainX. We now invoke a standard result for...

  47. [47]

    It therefore perfectly matches the hypotheses of Theorem 1

    Applying the Main Rectifiability Theorem The extended function ˆθext is Lipschitz onR N . It therefore perfectly matches the hypotheses of Theorem 1. Applying the theorem, we conclude that forL k-almost allθ ′ ∈R k, the level set of the extended function, ˆθ−1 ext ({θ′}), is a(N−k)-rectifiable set

  48. [48]

    The level set of the original function, ˆθ−1({θ′}), consists of all pointsxsuch that bothx∈ Xand ˆθ(x) =θ ′

    Restriction to the Original Domain and Conclusion The final step is to demonstrate that this geometric property of the extended function’s level sets is inherited by the level sets of our original MLE. The level set of the original function, ˆθ−1({θ′}), consists of all pointsxsuch that bothx∈ Xand ˆθ(x) =θ ′. Since ˆθext and ˆθare identical onX, we can ex...

  49. [49]

    For a PDL function ˆθand an integrable functionu(x), the formula states: Z X u(x)Jconsˆθ(x)dLN(x) = Z Θ Z ˆθ−1(θ′) u(x)dHN−k (x) dLK(θ′)

    The Bridge between Spaces: The Coarea Formula The core of the proof is the coarea formula for path- differentiable Lipschitz functions (Theorem 2). For a PDL function ˆθand an integrable functionu(x), the formula states: Z X u(x)Jconsˆθ(x)dLN(x) = Z Θ Z ˆθ−1(θ′) u(x)dHN−k (x) dLK(θ′). Our strategy is to algebraically manipulate the data-space integral int...

  50. [50]

    To align this with the coarea formula, we multiply and divide the integrand byJ consˆθ(x)

    Transforming the Data-Space Integral We begin with the data-space definition of the model complexity, slightly modified with an indicator function for full rigor: Cdata(µΘ) = Z X p[x|ˆθ(x)]v( ˆθ(x))1Jcons ˆθ(x)>0(x)dLN(x). To align this with the coarea formula, we multiply and divide the integrand byJ consˆθ(x). This is valid as the indicator function res...

  51. [51]

    On the domain ˆθ−1({θ′}), we have ˆθ(x) =θ ′ by definition

    Simplification within the Level Sets A crucial simplification occurs inside the inner integral. On the domain ˆθ−1({θ′}), we have ˆθ(x) =θ ′ by definition. Furthermore, Assumption 1(3) guarantees thatJ consˆθ(x)>0 forH N−k -almost allxon these sets, so the indicator becomes redundant. The termv( ˆθ(x))becomesv(θ ′)and can be factored out of the inner inte...

  52. [52]

    Since all functions in the integrands are non-negative, Tonelli’s theorem applies, stating that one integral is finite if and only if the other is finite

    Conclusion on Equality and Finiteness We have shown thatC data(µΘ) =C param(µΘ). Since all functions in the integrands are non-negative, Tonelli’s theorem applies, stating that one integral is finite if and only if the other is finite. This completes the proof. APPENDIXE ON THEUNIQUENESS OF THECONSERVATIVEJACOBIAN FACTOR While pathwise AD provides a valid...

  53. [53]

    Its determinant is the product of itsKstrictly positive eigenvalues,det(M) =QK j=1 λj(M)

    Analysis of the MatrixGG T The matrixM=GG T is aK×Ksymmetric, positive definite matrix under the full rank assumption. Its determinant is the product of itsKstrictly positive eigenvalues,det(M) =QK j=1 λj(M)

  54. [54]

    Thus,λ j(GGT ) =σ j(G)2 for theKpositive singular values ofG

    Connecting the Jacobian Factor to Singular Values A key identity from linear algebra states that the non-zero eigenvalues ofGG T are the squares of the non-zero singular values ofG. Thus,λ j(GGT ) =σ j(G)2 for theKpositive singular values ofG. We can now derive an explicit expression for the conservative Jacobian factor: Jcons(G) = q det(GGT ) = vuut KY j...

  55. [55]

    if and only if

    Proof of the Biconditional Statement The theorem’s “if and only if” statement follows directly from this identity. (⇒) If the product of singular values is constant for allG∈ S(x 0), then by the identity,J cons(G)must also be constant. (⇐) IfJ cons(G)is constant for allG∈ S(x 0), then by the identity, the product of singular values must also be constant. ...

  56. [56]

    One begins with the data-space integral and uses the classical coarea formula for Lipschitz maps to transform it into the nested parameter-space integral

    Equivalence of Data-Space and Parameter-Space Inte- grals This part of the proof follows the same logic as that of Proposition 1. One begins with the data-space integral and uses the classical coarea formula for Lipschitz maps to transform it into the nested parameter-space integral. Since all integrands are non-negative, Tonelli’s theorem ensures that on...

  57. [57]

    The set of such problematic values is the set of critical values, Θcrit

    The Measure of the Critical Values The parameter-space integrand can be infinite if the denom- inatorJ K ˆθ(x)is zero for pointsxon the level set ˆθ−1({θ′}). The set of such problematic values is the set of critical values, Θcrit. We show that this set is negligible in the parameter space. A pointxwhere the derivative exists but the Jacobian vanishes is a...

  58. [58]

    A defining property of the Lebesgue integral is that altering the integrand on a set of measure zero does not change the integral’s value

    Synthesis and Well-Posedness We have established that the parameter-space integral is numerically equal to the data-space integral and that the set of parameter valuesθ ′ where the integrand might be infinite is a set of measure zero. A defining property of the Lebesgue integral is that altering the integrand on a set of measure zero does not change the i...

  59. [59]

    terms: E[bfN(θ′)] =E X∼q H(X) JA(X)

    The Expectation of the AD-Based Estimator 25 The expectation of the AD-based NML estimator is the expectation of one of its i.i.d. terms: E[bfN(θ′)] =E X∼q H(X) JA(X) . By the Law of the Unconscious Statistician, this expectation can be written as an integral over the data space: E[bfN(θ′)] = Z X H(x) JA(x) q(x)dx

  60. [60]

    ideal” estimator that uses the true classical Jacobian,J K ˆθ(x), which exists almost everywhere. Assuming the estimator is designed to be unbiased, we have: f(θ ′) =E X∼q

    An Integral Representation for the True Value The true value,f(θ ′), is the expectation of an “ideal” estimator that uses the true classical Jacobian,J K ˆθ(x), which exists almost everywhere. Assuming the estimator is designed to be unbiased, we have: f(θ ′) =E X∼q " H(X) JK ˆθ(X) # = Z X H(x) JK ˆθ(x) q(x)dx

  61. [61]

    By the linearity of the integral, we can combine them: Bias(bfN) =E[ bfN]−f(θ ′) = Z X H(x) JA(x) q(x)dx− Z X H(x) JK ˆθ(x) q(x)dx = Z X H(x) 1 JA(x) − 1 JK ˆθ(x) ! q(x)dx

    Assembling the Bias Integral The bias is the difference between these two expectations. By the linearity of the integral, we can combine them: Bias(bfN) =E[ bfN]−f(θ ′) = Z X H(x) JA(x) q(x)dx− Z X H(x) JK ˆθ(x) q(x)dx = Z X H(x) 1 JA(x) − 1 JK ˆθ(x) ! q(x)dx. As established in Theorem 4, the term in the parentheses is non-zero only on a set ofL D-measure...

  62. [62]

    How- ever, the SJO-GS oracle (Algorithm 1) replaces the determin- istic Jacobian with a random variableG out sampled uniformly from anϵ-ballB(x, ϵ)

    Restoring Feller Continuity via Gradient Sampling The deterministic transition kernel is discontinuous where the AD selection map is discontinuous (Proposition 2). How- ever, the SJO-GS oracle (Algorithm 1) replaces the determin- istic Jacobian with a random variableG out sampled uniformly from anϵ-ballB(x, ϵ). By Rademacher’s theorem, the set of points w...

  63. [63]

    To ensure the chain does not get “stuck” (which would destroy the spectral gap), the variance of the randomized acceptance ratio must be strictly bounded

    The Foster-Lyapunov Drift Condition and Bounded Vari- ance Introducing a randomized oracle turns the method into a Pseudo-Marginal MCMC [34]. To ensure the chain does not get “stuck” (which would destroy the spectral gap), the variance of the randomized acceptance ratio must be strictly bounded. Becausem(x)is a path-differentiable Lipschitz (PDL) function...

  64. [64]

    small set,

    Synthesis via the Weak Harris Theorem We invoke the Weak Harris Theorem [36], which estab- lishes Wasserstein andL 2 spectral gaps for complex MCMC algorithms. 1) The chain is irreducible and aperiodic (due to full support on the tangent space and positive rejection probability). 2) The transition kernel is strongly Feller contin- uous (proven in Step 1)....

This paper was first reviewed by grok-4.3 on June 30, 2026.