REVIEW 16 cited by
KAN or MLP: A Fairer Comparison
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper does not introduce a novel method. Instead, it offers a fairer and more comprehensive comparison of KAN and MLP models across various tasks, including machine learning, computer vision, audio processing, natural language processing, and symbolic formula representation. Specifically, we control the number of parameters and FLOPs to compare the performance of KAN and MLP. Our main observation is that, except for symbolic formula representation tasks, MLP generally outperforms KAN. We also conduct ablation studies on KAN and find that its advantage in symbolic formula representation mainly stems from its B-spline activation function. When B-spline is applied to MLP, performance in symbolic formula representation significantly improves, surpassing or matching that of KAN. However, in other tasks where MLP already excels over KAN, B-spline does not substantially enhance MLP's performance. Furthermore, we find that KAN's forgetting issue is more severe than that of MLP in a standard class-incremental continual learning setting, which differs from the findings reported in the KAN paper. We hope these results provide insights for future research on KAN and other MLP alternatives. Project link: https://github.com/yu-rp/KANbeFair
Forward citations
Cited by 16 Pith papers
-
KANEL\'E: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation
Quantized, pruned Kolmogorov-Arnold Networks can be compiled directly into FPGA lookup tables, achieving extreme latency/resource reductions and matching state-of-the-art LUT-based networks on several benchmarks.
-
Neural Discovery of Memory and Nonlocal Kernels in Integro-Differential Equations with Constrained Kolmogorov--Arnold Networks
Hard-constrained Bernstein MC-KANs recover positive, monotone, convex memory/nonlocal kernels from sparse noisy IDE data more robustly than soft-penalized Cheb-KANs, especially in 2D.
-
Kolmogorov-Arnold Network for Gene Regulatory Network Inference
scKAN uses Kolmogorov-Arnold networks in a one-vs-rest regression and treats model gradients as signed gene regulation strengths, outperforming baselines on several BEELINE benchmark tasks.
-
SPPSFormer: High-quality Superpoint-based Transformer for Roof Plane Instance Segmentation from Point Clouds
SPPSFormer uses boundary-accurate, uniform superpoints plus a Transformer-FourierKAN decoder and geometric postprocessing to achieve state-of-the-art roof plane segmentation on RoofN3D and Building3D.
-
KAN-LSTM-Transformer Neural Networks, MFV and Cosmological Parameters
KLT-Net reconstructs the SN Ia distance modulus non-parametrically; with MFV M_B and flat-ΛCDM Bayesian/Hessian inference it yields H0 ≈ 69.6 km s⁻¹ Mpc⁻¹ and Ωm ≈ 0.30.
-
Physical Analogue Kolmogorov-Arnold Networks based on Reconfigurable Nonlinear-Processing Units
A proposed analog KAN chip uses silicon RNPUs as physically programmable nonlinear edges, with estimated ~250 pJ per inference and ~10x smaller area than a digital MLP.
-
Quantum Variational Activation Functions Empower Kolmogorov-Arnold Networks
QKANs show strong empirical performance on regression, vision, and language tasks, but the claimed exponential parameter reduction is not rigorously established.
-
Sinusoidal Approximation Theorem for Kolmogorov-Arnold Networks
The paper states universal approximation theorems for sine-based Kolmogorov-Arnold networks, but the proof relies on an invalid lemma and an unproven matrix invertibility.
-
"KAN you hear me?" Exploring Kolmogorov-Arnold Networks for Spoken Language Understanding
Placing a KAN layer between two linear layers improves spoken language understanding accuracy over linear-only baselines on several speech-intent datasets.
-
Low Tensor-Rank Adaptation of Kolmogorov--Arnold Networks
A low tensor-rank adaptation (LoTRA) method and learning-rate guidance enable efficient fine-tuning of Kolmogorov-Arnold networks, validated on PDE solving and representation tasks.
-
Efficiency Bottlenecks of Convolutional Kolmogorov-Arnold Networks: A Comprehensive Scrutiny with ImageNet, AlexNet, LeNet and Tabular Classification
CKANs are measurably less efficient than standard CNNs, and on ImageNet the accuracy gap is large, but the paper's baseline and timing comparisons are not controlled.
-
SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
SechKAN combines sech basis functions with a 1D linear projection to build a KAN-style model whose parameter count matches MLPs and which is competitive or better than several KAN variants on tested benchmarks.
-
Clifford Kolmogorov-Arnold Networks
ClKAN extends complex-valued KANs to arbitrary Clifford algebras, and using scrambled Sobol-sequence grids cuts the parameter count in higher-dimensional spaces.
-
Kolmogorov Arnold Networks (KANs) for Imbalanced Data -- An Empirical Perspective
On ten KEEL datasets, KANs outperform MLPs on raw imbalanced data but resampling and focal loss degrade KANs while MLPs with those techniques match KAN performance at far lower cost.
-
Bridging KAN and MLP: MJKAN, a Hybrid Architecture with Both Efficiency and Expressiveness
MJKAN is a FiLM-modulated RBF layer that beats MLPs on some 1D regression tasks with carefully chosen basis counts, but underperforms MLPs on classification benchmarks.
-
Local Control Networks (LCNs): Optimizing Flexibility in Neural Network Data Pattern Capture
Local Control Networks put a separate learnable B-spline activation on every neuron and report small accuracy gains over MLPs and KANs on benchmark tasks.
Discussion (0). Continue with ORCID to comment.