{"id":"eec2d3a3-6f8e-45d9-a17a-7e987ee746cc","arxiv_id":"2505.14177","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PSGLA is proven to converge for non-convex composite potentials, up to a step-size bias, via a new drift-stability bound for inexact ULA.","lead":"Researchers prove, under an assumption of strong convexity at infinity, that the Proximal Stochastic Gradient Langevin Algorithm converges for non-convex sampling targets, a result previously available only in convex settings. The proof uses a new drift-stability bound for inexact Langevin algorithms combined with Moreau envelope properties, and is validated on Gaussian mixtures and image inpainting.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 3(ii) is the load-bearing hypothesis: it requires the Moreau envelope to be uniformly strongly convex at infinity, a condition that fails for standard non-smooth regularizers like L1/TV and is not verified for the DnCNN experiments.","rationale":"The reader and I converge on the same weakest point. Theorem 3's proof depends entirely on the uniform strong convexity at infinity of g_gamma; if that fails, the drift b_gamma need not be weakly dissipative, so the shadow chain may not be geometrically ergodic and Theorem 1 cannot be applied. The manuscript acknowledges Assumption 3(ii) is 'technical and hard to verify' (Appendix B), and the only sufficient condition supplied (Lemma 16) requires g to be globally smooth and strongly convex, which excludes the non-smooth setting that motivates PSGLA. I also noted two small proof issues: the chain defined with b0 has a sign inconsistency (the proof later compares with b0=grad f + grad g_gamma, not -grad g_gamma - grad f), and eq. (86) states a W1 bound where W_p is needed; both are readily fixable and do not affect the core argument if Assumption 3(ii) holds. The verdict CONDITIONAL remains appropriate: the theoretical claim is plausible under a very restrictive assumption, but its applicability to the paper's own experiments is not established.","tokens_in":57642,"tokens_out":24377,"duration_ms":204831,"concrete_test":"Take g(x)=||x||_1. Its Moreau envelope g_gamma is the separable Huber function: for each coordinate, g_gamma,i(t)=t^2/(2gamma) for |t|<=gamma and |t|-gamma/2 for |t|>gamma. The Hessian of g_gamma is 0 in any coordinate with |x_i|>gamma. Hence for any R>gamma, there exists x with ||x||>=R such that lambda_min(grad^2 g_gamma(x))=0, so Assumption 3(ii) fails. This analytical check settles that Theorem 3 does not apply to L1 or TV regularization. A complementary check for the paper's main experiments would be to estimate the smallest eigenvalue of the Jacobian of the residual map x -> x - D_sqrt(gamma)(x) for the DnCNN denoiser along directions of large ||x||; if this eigenvalue is not uniformly bounded below by 8Lf+4Lg, Assumption 3(ii) is also violated for the plug-and-play prior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 3(ii) is the only mechanism that makes the shadow drift b_gamma (Eq. 11) weakly dissipative (Lemma 21), which in turn gives geometric ergodicity of the shadow chain and enables the use of Theorem 1. The assumption requires grad^2 g_gamma ⪰ mu Id on R^d \\ B(0,R0) for all gamma in (0, gamma1] with mu >= 8Lf + 4Lg. This is a strong, uniform positive-definiteness condition on the Moreau envelope, not a consequence of the weak convexity of g (Lemma 6 only ensures weak convexity of g_gamma). Lemma 16 shows the condition holds if g itself is globally smooth and strongly convex at infinity, but then PSGLA is not needed for non-smooth problems. For standard non-smooth regularizers the condition fails: if g(x)=||x||_1, then g_gamma is a separable Huber function whose Hessian is zero in any coordinate with |x_i|>gamma, so inf_{||x||>=R} lambda_min(grad^2 g_gamma(x))=0 for every gamma. The same reasoning applies to total-variation type regularizers. The paper's claim that its theory applies to the DnCNN-based PnP-PSGLA (Appendix C.2) is not substantiated: the authors only note the denoiser is a proximal operator of a convex g, but do not verify the uniform lower bound on grad^2 g_gamma. Thus the central theorem, while conditionally correct, does not cover the paper's main non-smooth/non-convex examples, and the gap between theory and experiments is real.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:38:52.304179+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}