{"id":"a5766ebf-ea54-472d-a6d9-59812d767507","arxiv_id":"2605.14663","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Under the local PL condition with multiplicative noise for C² functions, (S)GD asymptotic rates match those of strongly convex quadratics via a geometric argument.","lead":"The paper proves that under a local Polyak-Lojasiewicz condition with multiplicative gradient noise, the asymptotic convergence rates of gradient descent and stochastic gradient descent equal those for strongly convex quadratic functions. A smart generalist might read it to understand why optimization in non-convex machine learning models can still achieve fast local rates similar to convex cases.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Central claim requires multiplicative noise to vanish asymptotically; additive noise breaks rate equivalence","rationale":"The reader's weakest assumption already isolates the multiplicative noise and local-PL region as the conditions under which the rate equivalence can hold. The geometric argument does not remove the need for those conditions, so the concern is correctly placed and the UNVERDICTED status is appropriate.","tokens_in":1620,"tokens_out":346,"duration_ms":25013,"concrete_test":"Take the 1-D C² function f(x)=x⁴/4 (satisfies local PL near 0 with μ=0 for small |x| after rescaling) and run both GD and SGD with multiplicative noise σ²∝|∇f|² versus additive noise σ²=const. Measure the observed contraction factor |x_{k+1}/x_k| for large k; check whether it equals the quadratic rate (1−μ/L) only under the multiplicative model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result equates the asymptotic (S)GD rate under local PL to the strongly convex quadratic rate. This equivalence is derived via a geometric view of the PL inequality that reduces the local dynamics to a linear contraction once inside the PL region. The reduction step uses the multiplicative noise assumption (variance scales with ||∇f||) so that the stochastic perturbation vanishes at the same rate as the deterministic gradient term. Under additive noise the perturbation remains O(1) while the gradient term →0, so the trajectory cannot inherit the deterministic linear rate. The C² + local-PL assumptions alone are therefore insufficient; the noise model is load-bearing for the claimed rate match.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that for C² functions satisfying the local Polyak-Łojasiewicz inequality under a multiplicative gradient noise model (motivated by overparameterized networks), the asymptotic convergence rates of GD and SGD match the linear rates known for strongly convex quadratics. The proof relies on a geometric interpretation of the PL condition that reduces local dynamics to a linear contraction once inside the PL region.","tokens_in":1717,"tokens_out":307,"duration_ms":21669,"significance":"If correct, the result supplies a clean geometric unification of local convergence rates across convex and non-convex regimes, with direct relevance to optimization landscapes in machine learning. The explicit multiplicative-noise assumption and geometric reduction are strengths that make the rate-equivalence claim falsifiable and potentially useful for algorithm design.","major_comments":[{"comment":"The multiplicative noise model (variance scaling with ||∇f||) is load-bearing for the claimed rate match. Under additive noise the stochastic perturbation remains O(1) while the gradient term vanishes, preventing inheritance of the deterministic linear rate. The manuscript must therefore contain an explicit step (in the local-analysis section following the geometric reduction) showing that the noise term vanishes at the same rate as the deterministic term inside the PL region; without this step the equivalence to the quadratic case does not follow from C² + local PL alone.","section":"noise-model assumption and local PL analysis"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment of the paper's significance and for the constructive major comment. We address it point-by-point below.","responses":[{"response":"We agree that the multiplicative structure is essential for inheriting the linear rate and that an explicit verification is needed. In the local-analysis section (following the geometric reduction to the PL region), the proof already contains this step: because the noise variance scales with ||∇f(x)||² and the local PL inequality gives ||∇f(x)||² ≥ μ(f(x)−f*), the stochastic term is bounded by a quantity that contracts at the same linear rate as the deterministic gradient term. This appears in the derivation of the one-step contraction for E[f(x_{k+1})] (immediately after the geometric reduction is invoked). To address the referee’s request for greater explicitness, we will insert a short dedicated paragraph and a displayed inequality that isolates the vanishing of the noise term.","revision_made":"partial","referee_comment":"[noise-model assumption and local PL analysis] The multiplicative noise model (variance scaling with ||∇f||) is load-bearing for the claimed rate match. Under additive noise the stochastic perturbation remains O(1) while the gradient term vanishes, preventing inheritance of the deterministic linear rate. The manuscript must therefore contain an explicit step (in the local-analysis section following the geometric reduction) showing that the noise term vanishes at the same rate as the deterministic term inside the PL region; without this step the equivalence to the quadratic case does not follow from C² + local PL alone."}],"tokens_in":1193,"tokens_out":343,"duration_ms":18444,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core result is that under C2 smoothness, a local PL inequality, and multiplicative gradient noise, both GD and SGD achieve the same asymptotic linear rate as in the strongly convex quadratic case. The argument rests on a geometric reading of PL that turns the local dynamics into a contraction once inside the region.\n\nThe geometric step is the main contribution. It gives a direct way to transfer the classical rate without extra assumptions or heavy machinery. The multiplicative noise model is handled cleanly so the stochastic term vanishes at the same pace as the gradient, preserving the deterministic contraction asymptotically. This matches the motivation from overparameterized networks.\n\nThe paper states its scope plainly: the claim is local and asymptotic, and the noise must be multiplicative. With additive noise the equivalence would fail because the perturbation stays order one while the gradient shrinks. That dependence is load-bearing but is not hidden. The C2 and local-PL conditions are standard and sufficient for the reduction they use.\n\nA minor limitation is that reaching the local PL region still needs separate argument in any global analysis, but the paper does not claim to solve that. The result is narrow but precise.\n\nThis is for readers working on optimization rates in machine learning who already accept local PL as a modeling assumption. It offers a shortcut for asymptotic analysis that could be cited when the noise model fits. The work is coherent on its own terms and deserves a serious referee.","headline":"Local PL with multiplicative noise lets (S)GD match the strongly convex quadratic asymptotic rate via a geometric reduction.","tokens_in":2197,"tokens_out":351,"would_cite":false,"duration_ms":16891,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Under local PL with multiplicative noise, (S)GD matches the asymptotic rate of strongly convex quadratics","keywords":["local Polyak-Lojasiewicz inequality","gradient descent","stochastic gradient descent","asymptotic convergence rates","non-convex optimization","multiplicative noise","geometric interpretation"],"falsifier":"Observe or simulate SGD on a C² function obeying local PL and multiplicative noise and find that its asymptotic rate is slower than the known quadratic rate.","tokens_in":2494,"feed_emoji":"","tokens_out":383,"duration_ms":33474,"temperature":0.7,"pith_summary":"The paper establishes that gradient descent and stochastic gradient descent achieve the same asymptotic convergence rates on C² functions satisfying the local Polyak-Lojasiewicz inequality as they do on strongly convex quadratics. This equivalence holds even though the functions may be non-convex globally, provided the gradient noise obeys a multiplicative model. The authors reach this conclusion through a geometric reading of the PL condition. A sympathetic reader would care because the result indicates that local PL behavior alone can deliver the optimal rates that are usually derived only under global strong convexity.","feed_headline":"SGD matches quadratic rates under local PL","feed_subtitle":"Asymptotic convergence equals the strongly convex quadratic case for non-convex functions when noise is multiplicative.","key_machinery":"Geometric interpretation of the local Polyak-Lojasiewicz inequality under the multiplicative gradient noise model","core_discovery":"Using a geometric interpretation of the PL-condition, the authors prove that in this possibly non-convex setting, the asymptotic convergence rate of (S)GD matches the rate obtained for strongly convex quadratics.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Local PL geometry gives SGD quadratic rates","SGD matches quadratic convergence under local PL","Geometric proof: local PL yields quadratic SGD rates","Asymptotic SGD rates equal quadratics in local PL"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The function must be C² smooth, satisfy the local PL inequality, and the gradient noise must be multiplicative.","fun_headline_variants_meta":{"raw":{"variants":["Local PL geometry gives SGD quadratic rates","SGD matches quadratic convergence under local PL","Geometric proof: local PL yields quadratic SGD rates","Asymptotic SGD rates equal quadratics in local PL"]},"model":"grok-4.3","cost_usd":0.003396,"raw_usage":{"total_tokens":1722,"prompt_tokens":512,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":33962000,"prompt_tokens_details":{"text_tokens":512,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1155,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":512,"tokens_out":55,"duration_ms":10086,"temperature":1.0,"reasoning_tokens":1155,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T20:23:26.147813+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Observe or simulate SGD on a C² function obeying local PL and multiplicative noise and find that its asymptotic rate is slower than the known quadratic rate.","supporting_citations":[],"review_version":1}