{"id":"3fca1695-a141-48d6-ab5d-0b65f0ef6e19","arxiv_id":"2608.09840","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For real-analytic loss functions, Langevin dynamics restricted to the zero set converges, in a Dirichlet-form sense, to a hierarchy of processes that are attracted to more singular strata.","lead":"This math paper identifies a candidate limit for a diffusion that is pinned to the zero set of a smooth, real-analytic loss function when the noise-to-drift ratio is small. The limit is a hierarchy of random motions on increasingly singular parts of the zero set, which may explain why stochastic gradient methods favor simple, generalizing solutions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The advertised global evolution on the zero set is not a theorem: per-stratum Dirichlet forms do not determine how—or whether—the process continues at frontier points, and weak convergence of X^beta is left open.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the per-stratum Dirichlet forms are not patched into a global Markov process on the zero set, and weak convergence of X^beta is not established. My reading confirms this and sharpens it: the missing concatenation is not merely a technical detail but an underdetermination of the limiting dynamics at frontier points. In the paper's own two-axis example, the per-stratum form on the starting axis can force absorption at the origin, and whether the process subsequently evolves on the other axis is an extra modeling choice. Without a construction of a Dirichlet form on the full zero set whose restrictions to the strata are the E_{lambda,m}, and without Mosco convergence of E_beta to that global form, the advertised 'bias toward more singular strata' is not a theorem. The paper is honest about this gap, so this is not an internal inconsistency; it is a substantial limitation of the central interpretive claim. Since the reader already conditioned the verdict on this issue, no change to the verdict is needed.","tokens_in":26838,"tokens_out":14196,"duration_ms":143982,"concrete_test":"Work out the minimal two-stratum example V(x1,x2)=x1^{2k1}x2^{2k2}, k1>k2 (e.g. k1=2,k2=1). Prove or disprove that X^beta, started at (1,0), converges weakly as beta->infty to the concatenated process: Bessel flow with dimension delta1=1-k1/k2 on {x2=0} until absorption at 0, then Bessel flow with dimension delta2=1-k2/k1 on {x1=0}. If the true limit is instead absorption at the origin, or if different natural gluings give different limits, the strict 'descent to more singular strata' conclusion is not valid. A tractable first step is to compute the capacity of {0} for the limiting one-dimensional form with weight |x|^{-k1/k2}; if the point is nonpolar, the per-stratum form alone is silent about the post-hitting dynamics, confirming that concatenation is an extra assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1.1 and Theorem 3.1 only establish rescaled convergence of E_beta to a separate Dirichlet form E_{lambda,m} on each stratum, for test functions whose supports avoid deeper strata. The paper explicitly leaves both weak convergence of X^beta and the concatenation of the per-stratum forms into a single process on the whole zero set open (Section 1.1). This gap is load-bearing because the headline 'strongly biased toward more singular strata' is a statement about a global limiting process, not about a hierarchy of forms. The per-stratum forms do not determine behavior at frontier points: for example, in the two-axis model V=x1^{2k1}x2^{2k2}, the form on the starting axis can have a nonpolar absorbing point at the origin, and the claim that the process then continues on the other axis is an additional gluing prescription, not a consequence of the theorem. Remark 3.4 also notes that it is not even shown that mu_{lambda,m} assigns infinite mass to neighborhoods of frontier points. Thus the central advertised mechanism—that SGLD/SGD dynamics descend through progressively more singular strata—is not supported by the proved results without an additional construction and convergence argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the Langevin SDE dX_t = -β∇V(X_t)dt + √2 dB_t for a nonnegative real-analytic potential V in the large-β limit, starting on the zero set Z=V^{-1}(0). Using Hironaka resolution of singularities, the author decomposes Z into finitely many strata Z_{λ,m} indexed by the local learning coefficient λ and multiplicity m, ordered so that smaller λ (and larger m for equal λ) corresponds to more singular behavior. The main technical result, Theorem 3.1, establishes rescaled Laplace asymptotics: for test functions supported away from deeper strata, c_β ∫ f e^{-βV} dx converges to ∫_{Z_{λ,m}} f dμ_{λ,m}, where μ_{λ,m} is a Radon measure pushforward of an explicit density in resolved coordinates. Theorem 4.1 then shows that the corresponding gradient form E_{λ,m}(f,g)=∫_{Z_{λ,m}} ∇f·∇g dμ_{λ,m} is closable on a suitable domain and its closure is a strongly local regular Dirichlet form, hence corresponds to a Markov process with continuous trajectories on the one-point compactification of Z_{λ,m}. The paper is explicit that weak convergence of X^β to these processes is not proved, and that concatenating the per-stratum processes into a single process on the whole zero set is left open. The global narrative—that SGLD/SGD dynamics descend through progressively more singular strata—is presented as a heuristic consequence of the per-stratum forms.","tokens_in":26976,"tokens_out":10045,"duration_ms":100435,"significance":"If the per-stratum results are correct, the paper makes a substantial contribution: it gives a rigorous, resolution-based description of the limiting Dirichlet forms associated with a large family of degenerate Langevin diffusions, and it connects this to the singularity theory of statistical learning. The proof chain is unusually detailed and self-contained, and the paper is careful to label the global process as a candidate rather than a proved limit. The main theorems are genuinely new and go beyond the smooth-manifold results of Li et al. (2022). The paper also avoids any fitted parameters: the normalization c_β and the limiting measures are determined by the intrinsic invariants λ and m. The principal weakness is that the advertised global mechanism—the strong bias toward more singular strata—is not itself a theorem, and the manuscript acknowledges this. The significance of the paper is therefore conditional on the per-stratum convergence being regarded as the main contribution, with the global evolution treated as a well-motivated conjecture.","major_comments":[{"comment":"The abstract and Section 1.1 describe the result as a hierarchy of Dirichlet forms corresponding to a stochastic evolution that is 'strongly biased toward higher-dimensional, or more singular, strata.' As the paper itself notes in Section 1.1, rigorously concatenating the per-stratum processes into a single process on Z is left open, and Remark 3.4 states that it is not even proved that μ_{λ,m} assigns infinite mass to neighborhoods of frontier points. This gap is load-bearing for the global narrative: per-stratum Dirichlet forms do not determine whether the process continues at a frontier point, and the two-axis example V=x_1^{2k_1}x_2^{2k_2} shows that the form on the starting axis can have an absorbing point at the origin, so continuation on the other axis is an additional gluing prescription rather than a consequence of Theorem 1.1. The authors should reframe the global 'descent through strata' statement explicitly as a conjecture, and adjust the abstract and title so that the proved per-stratum convergence is the primary claim.","section":"§1.1 (Limitations bullet), Remark 3.4"},{"comment":"The heuristic that nonintegrable density blowup at frontier points 'corresponds to an exploding drift strong enough to force absorption' is not established by the proved results. Pointwise blowup of the density in (3.2), or even infinite total mass of μ_{λ,m} near a point, does not by itself imply that the associated Dirichlet form process is absorbed there; this depends on capacity and recurrence properties of the form. Since the absorption mechanism is the basis for the advertised bias toward deeper strata, it should be stated as a heuristic or proved directly. The current text in the introduction and Remark 3.4 is appropriately hedged, but the abstract's unconditional wording does not match this.","section":"§4, Eq. (4.1), Remark 3.4"},{"comment":"Theorem 1.1 states convergence of c_β E_β(f,g) to E_{λ,m}(f,g) for f,g in a domain that depends on a Whitney stratification of Z_{λ,m}. The theorem statement does not mention that the tangency condition depends on the specific stratification constructed in Lemma 4.2, and that different stratifications could give different domains. This is not a technical error, but it is a presentation issue that should be clarified in the statement of Theorem 1.1, since the domain D_{λ,m} is defined before the stratification is introduced.","section":"Theorem 1.1 and Theorem 4.1"}],"minor_comments":[{"comment":"The phrase 'strongly biased toward higher-dimensional, or \"more singular\", strata' is misleading: in the example V=x_1^{2k}x_2^{2k}, the deepest stratum is the origin, which is 0-dimensional, while the less singular strata are the punctured axes. The ordering by singularity type is not an ordering by stratum dimension, so the abstract should say only 'more singular'.","section":"Abstract"},{"comment":"There is a typo: 'using on argument based on Itô’s formula' should read 'using an argument based on Itô’s formula'.","section":"Page 3, paragraph after Eq. (1.2)"},{"comment":"The notation y_{-I} in the density formula (3.2) is used but not defined; it should be stated that this denotes the coordinates with indices outside I.","section":"Section 3, around Eq. (3.2)"},{"comment":"The statement 'whose gradient is tangent to Z_{λ,m}' depends on a Whitney stratification, but the theorem does not specify that the stratification is the one constructed in Theorem 4.1. This should be made explicit to avoid ambiguity.","section":"Theorem 1.1"},{"comment":"The notation c' for the shifted contour is close to the notation c_U for Laurent coefficients; renaming the contour parameter would improve readability.","section":"Section 3, proof of Theorem 3.1"},{"comment":"The paper cites Drusvyatskiy and Larsson (2015) for an auxiliary approximation result in Lemma B.2(iii). This self-citation is appropriate and should be kept, but a brief sentence noting the result's role would help readers unfamiliar with it.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is technically careful and unusually honest about its limitations. The main theorems appear sound, and I do not see a circularity or a fitting-to-data issue. The reason I recommend major revision rather than minor revision is that the gap between the proved per-stratum Dirichlet form convergence and the advertised global 'descent through strata' mechanism is substantial and affects the way the paper will be read and cited. The authors can address this by significantly reframing the abstract, title, and introductory discussion so that the per-stratum convergence is the central claim and the global evolution is clearly labeled as a conjecture. I would not reject the paper; the per-stratum results are valuable and should be published after this revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a serious look. It proves the first large-beta Dirichlet-form convergence on zero-set strata for general real-analytic potentials, using the singularity types (lambda,m) from singular learning theory. The main theorems (3.1 and 4.1) are detailed and appear sound: the Mellin-Barnes representation, the Laurent expansion in Lemma 3.6, and the coarea-based Lemma 4.2 all check out. The limiting measures are derived from intrinsic invariants, not fitted to anything. The treatment of nondegenerate Hessian smooth manifolds by Li et al. (2022) is genuinely extended to a much larger class.\n\nThe paper is honest about its limits. It does not prove weak convergence of X^beta, and it says the concatenation of per-stratum forms into a single process on Z is open. The abstract's 'strongly biased toward more singular strata' is a candidate interpretation, not a theorem. The stress-test concern about frontier behavior is real: the per-stratum forms don't determine how or whether the process continues at points in deeper strata. But this is not a hidden flaw; Section 1.1 says so in plain language. The paper is careful to call the global evolution a candidate.\n\nMinor quibbles: the heuristic bound beta V(X^beta) = O(1) is asserted without proof, and Remark 3.4 notes the infinite-mass question for mu_{lambda,m} is open. Neither shakes the main results. The self-citation to Drusvyatskiy-Larsson is for a technical approximation lemma and is appropriate.\n\nThis paper will be of real value to people working on singular learning theory and on diffusion limits for SGD/SGLD. It gives a rigorous per-stratum picture and a clear statement of what is missing for a full process convergence theorem. That is exactly what a good paper should do.\n\nVerdict: send it to peer review. The referee should ask for the abstract to be toned down (or the global statement separated from the theorems) and should push for more discussion of possible gluing constructions. But the core analysis is solid and the open problems are clearly identified. I'd cite this once it's published.","headline":"Solid per-stratum Dirichlet-form limits for Langevin near real-analytic zero sets; the advertised global 'bias' is explicitly left open, but the theorems are real and worth refereeing.","tokens_in":27597,"tokens_out":2217,"would_cite":true,"duration_ms":20745,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60J60","60J25","31C25","14E15","32B20"],"pacs":[],"model":"deepseek-v4-flash","headline":"For nonnegative real-analytic potentials, the large-β Langevin diffusion has a candidate limit that descends through progressively more singular strata and is biased toward them.","keywords":["large-β Langevin diffusion","Dirichlet form convergence","real-analytic potential","singular learning theory","local learning coefficient","zero set stratification","stochastic gradient Langevin dynamics","Laplace asymptotics"],"falsifier":"Take a potential with two singularity strata, such as V(x,y)=$x^{{2k_1}}$ $y^{{2k_2}}$, and compute the limit measure μ_{λ,m} in resolved coordinates on a small ball around a point that maps to the deeper stratum. The paper's absorption picture predicts the density is nonintegrable there, so the limiting process must hit the deeper stratum and stop; if the computation yields an integrable density, the descent mechanism would fail. The paper itself notes it is not proved that μ_{λ,m} assigns infinite mass to neighborhoods of frontier points, making this the concrete check.","tokens_in":26534,"feed_emoji":"🧲","tokens_out":13250,"duration_ms":117340,"temperature":0.7,"pith_summary":"The paper studies the Langevin diffusion dXt = -β∇V(Xt)dt + √2 dBt for a nonnegative real-analytic potential V, in the limit of large β, assuming the process starts on the zero set of V. It establishes that the Dirichlet form of the process converges, after rescaling by β^λ (log β)^(1-m), to a hierarchy of Dirichlet forms indexed by the local learning coefficient λ and multiplicity m of each singularity stratum. Each limiting form is a strongly local regular Dirichlet form, so it defines a Markov process with continuous paths on that stratum, and the structure of the limiting density indicates that paths are absorbed into deeper, more singular strata. The result gives a candidate mechanism for the empirical observation that stochastic gradient Langevin dynamics tends to find singular solutions that generalize well. The paper does not claim full weak convergence of the original processes to these limiting processes, and concatenating the per-stratum evolutions into one process on the whole zero set is left open.","feed_headline":"Large-β Langevin limit descends a hierarchy of singular strata","feed_subtitle":"The limit is biased toward deeper, more singular strata—a candidate mechanism for why SGLD finds generalizing solutions.","key_machinery":"The engine of the paper is the Dirichlet form of the diffusion, Eβ(f,g) = ∫ ∇f·∇g $e^{{-βV}}$ dx, viewed as a Laplace integral. Resolution of singularities monomializes V, so the Laplace asymptotics can be extracted by residues of localized zeta functions; the local learning coefficient λ and multiplicity m of each point of the zero set determine the leading decay $β^{{-λ}}$(log β)^{-(m-1)} and index the strata Z_{λ,m}. The limiting measures μ_{λ,m} are pushforwards of explicit densities in resolved coordinates, whose blowup rates encode the pull toward deeper strata. A Whitney stratification with a local integrability condition on each stratum makes the limiting forms closable, so standard Dirichlet-form theory yields the associated Markov processes.","core_discovery":"The central discovery is Theorem 1.1: for each singularity type (λ,m), with cβ = β^λ (log β)^(1-m), and for $C^{1}$ test functions whose gradient is tangent to the stratum Z_{λ,m} and whose support avoids deeper strata, the rescaled Dirichlet forms satisfy cβ Eβ(f,g) → ∫_{Z_{λ,m}} ∇f·∇g dμ_{λ,m}, where μ_{λ,m} is a Radon measure constructed through resolution of singularities. The closure of this limiting form is a strongly local regular Dirichlet form, which corresponds to a Markov process with continuous trajectories in the one-point compactification of the stratum. The dynamics on each stratum can leave only by reaching the frontier, and points on the frontier lie in deeper strata; the explicit density of μ_{λ,m} in resolved coordinates shows nonintegrable blowup toward those deeper strata, which is interpreted as absorption. The upshot is a candidate limiting evolution on the whole zero set that descends through progressively more singular strata and is strongly biased toward them.","pith_inferences":["If a future tightness and Mosco-convergence argument closes the gap, the limiting process would be the concatenation of the per-stratum processes, and the singularity-type hierarchy would then govern the long-run distribution of SGLD on the empirical-loss landscape—directly connecting the fixed-sample dynamics to Bayesian free-energy asymptotics of singular learning theory.","The explicit density of μ_{λ,m} yields a quantitative prediction that can be tested before any convergence theorem: near a stratum front, the effective one-dimensional drift should diverge like a power of the distance, with an exponent determined by the ratio of local learning coefficients, so simulated SGLD trajectories in toy potentials could be used to estimate that exponent.","The restriction to symmetric, isotropic-noise diffusions suggests the singularity bias may be special to SGLD-type algorithms; a natural test is whether anisotropic noise (e.g. noise covariance ∇²V) preserves the same hierarchy or shifts the preferred stratum, which would distinguish this mechanism from more generic flat-minima biases.","One could ask whether the candidate limiting occupation measure on the deepest reachable stratum matches the Bayesian posterior's concentration region for the same empirical loss; if it does, the 'generalization puzzle' would be explained by the same singularities that govern Bayesian free energy, now appearing dynamically."],"forward_implications":["Every singularity stratum carries a well-defined Markov process with continuous paths in its one-point compactification, so the zero set acquires a hierarchical family of candidate limit dynamics rather than a single SDE.","Because trajectories can leave a stratum only through its frontier, and frontier points lie in deeper strata, the limiting evolution is biased toward deeper (more singular) strata; in the two-axis example this reproduces Bessel-type absorption at the origin.","For the SGLD/SGD interpretation, the result supplies a mechanism for the observed preference for singular, well-generalizing solutions: the effective drift toward more singular strata is encoded in the limiting measure's density and diverges as the process approaches deeper strata.","The paper explicitly does not prove weak convergence of X^β or concatenation of the per-stratum processes into a single process on all of Z; these remain open, and the advertised descent should be read as a candidate limit unless one of them is established.","The theory extends immediately to a tilted version with weight decay, where the limiting measures are multiplied by e^{-γ|x|^2/2} and the limiting dynamics acquires a tangential drift -γ τ(x) dt."],"supporting_citations":[{"why":"Provides the resolution of singularities theorem used to monomialize V in the Laplace-asymptotics argument.","marker":"Hironaka (1964)"},{"why":"Supplies the classical division-of-distributions and residue calculus for the localized zeta functions.","marker":"Atiyah (1970)"},{"why":"Canonical desingularization result from which the global monomializing resolution with numerical data (kα,hα) is deduced.","marker":"Bierstone and Milman (1997)"},{"why":"Defines the local learning coefficient and multiplicity and establishes the volume asymptotics that underlie the stratification.","marker":"Watanabe (2009)"},{"why":"Source for the modern formulation of the local learning coefficient as a singularity-aware complexity measure.","marker":"Lau et al. (2024)"},{"why":"Supplies the subanalytic-set toolkit (Boolean properties, images, stratifications) used throughout Sections 2 and 4.","marker":"Bierstone and Milman (1988)"},{"why":"Provides the Whitney stratification theory, including existence of subanalytic Whitney stratifications and frontier-condition refinements.","marker":"Shiota (1997)"},{"why":"Coarea formula used to represent the pulled-back limiting measure as an absolutely continuous measure with explicit density.","marker":"Federer (1969)"},{"why":"Gives the correspondence between strongly local regular Dirichlet forms and Markov processes with continuous trajectories, yielding the candidate processes.","marker":"Fukushima et al. (2011)"},{"why":"Provides the closability criterion and direct-sum argument used to prove the limiting forms are closable.","marker":"Albeverio et al. (1989)"}],"fun_headline_variants":["Langevin limit dives through singular strata","Zero-set Langevin dynamics favor deeper singularities","Singular learning theory: Langevin descends strata hierarchy","Large-β limit biases Langevin to most singular strata","Langevin on zero set: a descent into singular strata"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The advertised single limiting process that descends through progressively more singular strata depends on gluing the separately constructed stratum processes into one process on the whole zero set, which the paper states is left open.","fun_headline_variants_meta":{"raw":{"variants":["Langevin limit dives through singular strata","Zero-set Langevin dynamics favor deeper singularities","Singular learning theory: Langevin descends strata hierarchy","Large-β limit biases Langevin to most singular strata","Langevin on zero set: a descent into singular strata"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1298,"prompt_tokens":960,"completion_tokens":338,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":260}},"tokens_in":576,"tokens_out":338,"duration_ms":4265,"temperature":1.0,"reasoning_tokens":260,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:40:29.303142+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a potential with two singularity strata, such as V(x,y)=$x^{{2k_1}}$ $y^{{2k_2}}$, and compute the limit measure μ_{λ,m} in resolved coordinates on a small ball around a point that maps to the deeper stratum. The paper's absorption picture predicts the density is nonintegrable there, so the limiting process must hit the deeper stratum and stop; if the computation yields an integrable density, the descent mechanism would fail. The paper itself notes it is not proved that μ_{λ,m} assigns infinite mass to neighborhoods of frontier points, making this the concrete check.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the subanalytic-set toolkit (Boolean properties, images, stratifications) used throughout Sections 2 and 4."},{"cited_title":"Geometry of subanalytic and semialgebraic sets, volume 150 of Progress in Mathematics","cited_arxiv_id":null,"evidence_quote":"Provides the Whitney stratification theory, including existence of subanalytic Whitney stratifications and frontier-condition refinements."},{"cited_title":"Geometric measure theory, volume Band 153 of Die Grundlehren der mathematischen Wissenschaften","cited_arxiv_id":null,"evidence_quote":"Coarea formula used to represent the pulled-back limiting measure as an absolutely continuous measure with explicit density."},{"cited_title":"Dirichlet forms and symmetric M arkov processes , volume 19 of De Gruyter Studies in Mathematics","cited_arxiv_id":null,"evidence_quote":"Gives the correspondence between strongly local regular Dirichlet forms and Markov processes with continuous trajectories, yielding the candidate processes."}],"review_version":1}