REVIEW 11 cited by
On the expressiveness and spectral bias of KANs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
On the expressiveness and spectral bias of KANs
read the original abstract
Kolmogorov-Arnold Networks (KAN) \cite{liu2024kan} were very recently proposed as a potential alternative to the prevalent architectural backbone of many deep learning models, the multi-layer perceptron (MLP). KANs have seen success in various tasks of AI for science, with their empirical efficiency and accuracy demostrated in function regression, PDE solving, and many more scientific problems. In this article, we revisit the comparison of KANs and MLPs, with emphasis on a theoretical perspective. On the one hand, we compare the representation and approximation capabilities of KANs and MLPs. We establish that MLPs can be represented using KANs of a comparable size. This shows that the approximation and representation capabilities of KANs are at least as good as MLPs. Conversely, we show that KANs can be represented using MLPs, but that in this representation the number of parameters increases by a factor of the KAN grid size. This suggests that KANs with a large grid size may be more efficient than MLPs at approximating certain functions. On the other hand, from the perspective of learning and optimization, we study the spectral bias of KANs compared with MLPs. We demonstrate that KANs are less biased toward low frequencies than MLPs. We highlight that the multi-level learning feature specific to KANs, i.e. grid extension of splines, improves the learning process for high-frequency components. Detailed comparisons with different choices of depth, width, and grid sizes of KANs are made, shedding some light on how to choose the hyperparameters in practice.
Forward citations
Cited by 11 Pith papers
-
Necessary and sufficient conditions for universality of Kolmogorov-Arnold networks
Deep KANs achieve universal approximation if and only if they include at least one non-affine edge function σ, while two-layer KANs require σ to be nonpolynomial.
-
Necessary and sufficient conditions for universality of Kolmogorov-Arnold networks
Deep KANs with edge functions restricted to affine maps plus one fixed non-affine continuous function σ are dense in C(K) for any compact K if and only if σ is non-affine.
-
Necessary and sufficient conditions for universality of Kolmogorov-Arnold networks
Deep KANs with edge functions from a finite affine family plus one fixed non-affine continuous function σ are dense in C(K) for compact K precisely when σ is non-affine.
-
KANEL\'E: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation
Quantized, pruned Kolmogorov-Arnold Networks can be compiled directly into FPGA lookup tables, achieving extreme latency/resource reductions and matching state-of-the-art LUT-based networks on several benchmarks.
-
QKAN: quantum Kolmogorov-Arnold networks with applications in machine learning and multivariate state preparation
QKAN is a quantum algorithmic framework using block-encodings and QSVT to implement wide-and-shallow networks for quantum learning and compositional state preparation.
-
Fast, accurate, and differentiable: a neural-network surrogate for NRSur7dq4 precessing binary black hole waveforms
A piecewise MLP surrogate emulates NRSur7dq4 over its full domain at NR-faithful accuracy with ~1 ms GPU latency and a fully differentiable JAX likelihood pipeline.
-
EML Trees Are Universal Approximators
EML trees are proven to be universal approximators for W^{k,∞} functions by mimicking polynomial representations and invoking classical neural network approximation results, with a proposed learning algorithm demonstr...
-
Fourier Feature Pyramids for Physics-Informed Neural Networks
beignet replaces random Fourier feature embeddings in PINNs with a trainable multi-resolution Fourier feature pyramid, achieving higher accuracy on PDE benchmarks with fewer parameters and near machine precision resid...
-
AR-KAN: Autoregressive-Weight-Enhanced Kolmogorov-Arnold Network for Time Series Forecasting
AR-KAN combines a pre-trained AR module with KAN to reduce redundancy while preserving temporal features, delivering lower probabilistic approximation error and stronger forecasting results on synthetic almost-periodi...
-
PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs
PG-KINN pairs a KAN trial space with a Petrov–Galerkin test space for forward and inverse PDEs, but the inverse benchmark data is inconsistent with the governing equation and the accuracy claims are not supported by t...
-
A Practitioner's Guide to Kolmogorov-Arnold Networks
A systematic review of Kolmogorov-Arnold Networks that maps their relation to Kolmogorov superposition theory, MLPs, and kernels, examines basis-function design choices, summarizes performance advances, and supplies a...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.