Pith. sign in

REVIEW 2 cited by

On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08041 v1 pith:ROUMBLMC submitted 2024-10-10 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords kansconvergenceglobaldescentgradientphysics-informedregressiontasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Kolmogorov--Arnold Networks (KANs), a recently proposed neural network architecture, have gained significant attention in the deep learning community, due to their potential as a viable alternative to multi-layer perceptrons (MLPs) and their broad applicability to various scientific tasks. Empirical investigations demonstrate that KANs optimized via stochastic gradient descent (SGD) are capable of achieving near-zero training loss in various machine learning (e.g., regression, classification, and time series forecasting, etc.) and scientific tasks (e.g., solving partial differential equations). In this paper, we provide a theoretical explanation for the empirical success by conducting a rigorous convergence analysis of gradient descent (GD) and SGD for two-layer KANs in solving both regression and physics-informed tasks. For regression problems, we establish using the neural tangent kernel perspective that GD achieves global linear convergence of the objective function when the hidden dimension of KANs is sufficiently large. We further extend these results to SGD, demonstrating a similar global convergence in expectation. Additionally, we analyze the global convergence of GD and SGD for physics-informed KANs, which unveils additional challenges due to the more complex loss structure. This is the first work establishing the global convergence guarantees for GD and SGD applied to optimize KANs and physics-informed KANs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Training in a B-spline KAN basis is equivalent to preconditioned gradient descent on a multichannel ReLU MLP, and geometric refinement plus trainable knots accelerate and improve training.

  2. Kolmogorov Arnold Networks (KANs) for Imbalanced Data -- An Empirical Perspective

    cs.LG 2025-07 conditional novelty 4.0 of 10

    On ten KEEL datasets, KANs outperform MLPs on raw imbalanced data but resampling and focal loss degrade KANs while MLPs with those techniques match KAN performance at far lower cost.

Pith tools