Pith. sign in

REVIEW 2 cited by

Stochastic Langevin Differential Inclusions with Applications to Machine Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.11533 v3 pith:FRFSXWFD submitted 2022-06-23 math.OC cs.LGcs.NAmath.NA

classification math.OCcs.LGcs.NAmath.NA
keywords stochasticdifferentialasymptoticdriftflowfoundationalgradientinclusions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Stochastic differential equations of Langevin-diffusion form have received significant attention, thanks to their foundational role in both Bayesian sampling algorithms and optimization in machine learning. In the latter, they serve as a conceptual model of the stochastic gradient flow in training over-parameterized models. However, the literature typically assumes smoothness of the potential, whose gradient is the drift term. Nevertheless, there are many problems for which the potential function is not continuously differentiable, and hence the drift is not Lipschitz continuous everywhere. This is exemplified by robust losses and Rectified Linear Units in regression problems. In this paper, we show some foundational results regarding the flow and asymptotic properties of Langevin-type Stochastic Differential Inclusions under assumptions appropriate to the machine-learning settings. In particular, we show strong existence of the solution, as well as an asymptotic minimization of the canonical free-energy functional.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Metropolis-adjusted Subdifferential Langevin Algorithm

    stat.ME 2025-07 conditional novelty 4.0 of 10

    MASLA replaces the gradient in MALA with an element of a conservative subdifferential field, yielding a Metropolis-Hastings sampler that is reversible for locally Lipschitz, non-convex targets when the potential is tw...

  2. Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Weight-matrix probability densities in a toy autoencoder are evolved with the Fokker-Planck equation driven by the ADAM update, and the resulting output distributions roughly match training at epoch 5.

Pith tools