Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

KKANs: Kurkova-Kolmogorov-Arnold Networks and Their Learning Dynamics

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read KKAN is a universal approximator, regardless of basis choice, and outperforms MLPs and KANs in function approximation and operator learning while matching optimized MLPs in PIML.

desk verdict A promising KART-inspired architecture with correct but standard theory, undermined by deconfounded empirical comparisons that leave the headline 'outperforms' claim unproven. read the letter →

arxiv 2412.16738 v1 pith:3XPZ5NBI submitted 2024-12-21 cs.LG cs.NAmath.NAstat.ML

classification cs.LGcs.NAmath.NAstat.ML MSC 68T0741A3065D15
keywords Kolmogorov-ArnoldrepresentationtheoremKKANuniversalapproximationphysics-informedmachinelearningneuraloperatorsinformationbottlenecktheorygeometriccomplexityself-scaledresidual-basedattention
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

KKAN is proposed as a two-block architecture—MLP-based inner functions feeding linear combinations of basis functions as outer functions—that adheres to Kurkova's approximate version of the Kolmogorov-Arnold representation theorem rather than to the deeply nested KAN stack. The paper's central claim is that this structure is a universal approximator for any continuous function on the unit cube, with no restriction on the choice of dense basis family, and that it beats MLPs and Chebyshev-KANs on function regression and operator learning while matching fully optimized MLPs in physics-informed problems. If correct, KKAN is a drop-in architecture that gives practitioners a principled reason to pick basis functions (Chebyshev, Legendre, RBFs, sin-series) without changing the outer training loop. The paper also claims that all these architectures learn through three universal stages—fitting, transition, and diffusion—and that diffusion, when the gradient signal-to-noise ratio is high, is where generalization is won; its self-scaled residual-based attention is designed to keep the SNR high and prolong that stage.

What carries the argument

The load-bearing construction is the approximate KART set $K^{m,d}_M$: sums of the form $\sum_{q=0}^m G_q(\sum_{p=1}^d \Psi_{p,q}(x_p))$, where $G_q$ and $\Psi_{p,q}$ are univariate functions chosen from dense subsets $A^g_M$ and $A^\psi_M$ of $C(I)$. This set is dense in $C([0,1]^d)$ whenever both ansatz classes are dense, so any universal one-dimensional approximator can fill the inner and outer blocks. The KKAN implementation uses an MLP-based inner block (with two trainable Chebyshev embedding layers) and a linear combination of basis functions—Chebyshev, Legendre, sin-series, Chebyshev-grid, or RBFs—as the outer block. The second mechanism is self-scaled residual-based attention, which raises the memory coefficient $\gamma$ in stages to keep the gradient signal-to-noise ratio high during the late diffusion phase.

What would settle it

A single controlled experiment would settle the empirical claim: train KKAN and MLP on Allen-Cahn with identical periodic hard-constraint embeddings, parameter budgets, optimizers, and collocation points, and check whether the MLP reaches or beats KKAN's reported relative L2 error of about $3\times10^{-5}$; for the universality claim, the analogous check is to find a continuous function and a finite dense basis family for which KKAN error does not drop below $\varepsilon$ as the number of features grows.

Watch

Extended reading notes

Core claim

The paper's central claim is that a two-block network built directly from Kurkova's approximate Kolmogorov-Arnold representation is a universal approximator and a practical drop-in architecture. Specifically, for any continuous function $f$ on the unit cube and any $\varepsilon>0$, there are univariate inner functions $\Psi_{p,q}$ and outer functions $G_q$ drawn from dense ansatz classes such that $\|f - \sum_{q=0}^m G_q(\sum_{p=1}^d \Psi_{p,q}(x_p))\|_\infty < \varepsilon$. The theorem does not depend on which dense basis the outer block uses. Empirically, the paper reports that KKANs outperform MLPs and Chebyshev-KANs on smooth and discontinuous function regression, reach high accuracy on the Allen-Cahn PIML benchmark, and beat the same baselines when embedded in DeepONet and QR-DeepONet for the Burgers operator.

Load-bearing premise

The empirical claim that KKAN matches or beats fully optimized MLPs in PIML rests on benchmark comparisons being apples-to-apples; in the Allen-Cahn comparison, KKAN enforces periodic boundary conditions as hard architectural constraints while several MLP baselines were trained with soft Dirichlet conditions, so part of the accuracy gap may come from the problem setup rather than the architecture.

Editorial extensions

If this is right

  • Because the universal approximation theorem holds for any dense one-dimensional ansatz, KKAN inherits the KART representation without committing to B-splines or Chebyshev polynomials; the same architecture can switch bases by changing the outer-block linear combination.
  • In function approximation benchmarks, KKAN reaches relative L2 errors of $5.86\times10^{-3}$ (discontinuous) and $1.74\times10^{-4}$ (smooth), outperforming MLPs and cKANs at comparable parameter counts and per-iteration cost.
  • In PIML, KKAN+ssRBA solves Allen-Cahn to $2.28\times10^{-5}$ relative L2 error, beats the KAN variants compared, and is competitive with optimized MLPs; it also trains faster than the larger MLP variant in the full-batch setting.
  • In operator learning, KKAN-based DeepONet and QR-DeepONet reach $3.07\times10^{-2}$ and $2.66\times10^{-2}$ relative L2 errors on the Burgers operator, the best among compared models, and QR reparameterization helps KKAN while hurting cKAN.
  • Across all models and tasks, training proceeds through fitting, transition, and diffusion stages; the diffusion stage coincides with high SNR and optimal generalization, and ssRBA prolongs it by maintaining SNR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the density assumption is the only requirement for universality, then any one-dimensional universal approximator—wavelets, splines, RBFs, polynomials—can be dropped into the outer block, making KKAN a modular front-end for spectral-style basis selection.
  • Beyond the paper: the observed three-stage learning curve and the SNR peak during diffusion suggest a practical early-stopping or learning-rate-scheduling rule based on SNR rather than validation loss; this is testable across architectures.
  • Beyond the paper: since ssRBA works by progressively raising the memory term $\gamma$ to sustain SNR, the same mechanism should transfer to plain MLPs and KAN variants; a simple test is to apply ssRBA to the MLP baselines in Table 5 under matched boundary conditions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript proposes Kurkova-Kolmogorov-Arnold Networks (KKANs), a two-block architecture in which MLP-based inner functions expand each input dimension and outer functions are linear combinations of basis functions, following Kurkova's approximate version of the Kolmogorov-Arnold representation theorem. It proves a universal approximation theorem for such two-block superpositions under a denseness assumption on the inner and outer ansatz spaces, extends the architecture to physics-informed machine learning and DeepONet/QR-DeepONet operator learning, and proposes a self-scaled residual-based attention (ssRBA) method. The paper benchmarks KKANs against MLPs and cKANs on discontinuous and smooth function approximation, the Allen-Cahn equation, and Burgers operator learning, and analyzes learning dynamics through signal-to-noise ratio and geometric complexity, identifying fitting, transition, and diffusion stages.

Significance. If the empirical claims held, KKAN would be an attractive drop-in architecture for scientific machine learning: it is simple, follows the KART structure more closely than stacked KANs, and the stated universal approximation result is proved under a clean denseness assumption. The strength of the paper is its theoretical core: Appendix A gives a correct and standard denseness argument, and the appendices contain substantial implementation detail, including hyperparameter tables and a concrete pseudo-code for ssRBA. However, several load-bearing empirical comparisons are confounded by unequal enhancements or boundary-condition treatments, and most benchmark tables report single runs without error bars. The universal approximation claim is also stated more broadly than the theorem supports. These issues make the headline 'KKANs outperform MLPs and KANs' claims currently under-supported, although they are plausibly fixable with additional experiments and careful qualification.

major comments (4)
  1. [Abstract; Theorem 1, Section 2.2; Appendix B.3.2] The paper claims universal approximation 'regardless of the choice of basis function' and 'for a general class of functions.' Theorem 1 (Eq. (7)) is correct only under the stated assumption that the ansatz classes A_M^g and A_M^psi are dense in C(I_g) and C(I_psi). That assumption is not satisfied by the fixed finite-dimensional implementations in Appendix B.3.2: for example, the RBF outer block uses D centers fixed on a grid with fixed width sigma, so its span is a finite-dimensional subspace and is not dense in C(I_g) for any fixed D; the same holds for a fixed polynomial degree in the Chebyshev, Legendre, or sine bases. The 'regardless of basis function' statement therefore needs to be qualified, for instance by requiring a family of bases that becomes dense as D grows, or by proving density for each implemented basis family with appropriate scaling of centers and widths.
  2. [Table 6; Appendix F.3.1] In the Burgers operator-learning comparison, only the KKAN rows are listed with the architecture enhancement WNadResNet; the DeepONet and DeepOcKAN baselines do not receive this or an equivalent enhancement. The resulting advantage of QR-DeepOKKAN over QR-DeepONet is small (2.66e-2 versus 2.72e-2), and no repeated-seed statistics are reported. The claim that KKANs outperform MLPs in operator learning is therefore not established as an architecture effect. Please either apply the same enhancement to the baselines or ablate WNadResNet from KKAN, and report mean and standard deviation over multiple seeds.
  3. [Table 5; Section 4.2.1] The external KAN baselines in Table 5 (AcNet, KAN [52] implementation, and KAN [41] implementation) are trained with Dirichlet boundary conditions, while KKAN+ssRBA enforces periodicity through a hard architectural constraint. For the Allen-Cahn equation, whose exact solution is periodic, this makes the KAN comparison mismatched and harder for the baselines. Consequently, the statement that KKAN 'outperforms all KAN-based formulations' is not an architecture-only conclusion. Please either implement KAN baselines with the same periodic embedding or explicitly restrict the claim to the periodic-constraint setup and list the boundary-condition mismatch as a limitation.
  4. [Section 5; Figures 11-14] The paper asserts a 'strong correlation' between geometric complexity and SNR and identifies three 'universal' learning stages across all architectures and tasks. The evidence is qualitative visual inspection of a small set of problems, without correlation coefficients, repeated seeds, or a statistical test, and the word 'universal' overreaches beyond the four benchmark settings. Please either quantify the claimed correlation and the staging across multiple runs or soften the claims to observations on the tested problems.
minor comments (5)
  1. [Abstract; Section 4] The abstract says KKANs outperform 'the original KANs,' but the benchmarks use cKANs (Chebyshev KANs) rather than the original B-spline KAN; the original KAN is not included in Tables 2-6. Please adjust the wording or add the original KAN baseline.
  2. [Tables 2, 3, 4, and 6] These tables report point estimates from presumably single runs, so the observed margins (for example, 5.86e-3 versus 1.26e-2 in Table 2, or 3.07e-5 versus 3.52e-5 in Table 4c) cannot be assessed for statistical significance. Adding repeated-seed statistics would materially strengthen the empirical claims.
  3. [Algorithm 1, line 11] The RBA update rule contains a typographical error: 'lambda_{alpha,i} <- gamma_k lambda_{alpha,i} + eta ||r_{alpha,i}|/max_j|r_{alpha,j}|' should read 'eta |r_{alpha,i}| / max_j |r_{alpha,j}|'.
  4. [Table 6 footnote] The comparison with the errors reported in [41] is indirect because those models predict only the final time u(x,1), while the present models predict the full solution history u(x,t). The text notes this, but the juxtaposition of numbers in the same table may still mislead; a separate comparison on the final time would be clearer.
  5. [Overall] No code or data availability statement is provided, which limits reproducibility of the benchmarks. Please include a statement or repository link.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation chain; the universality proof is a standard denseness argument over external KART, and the empirical comparisons, though containing fairness confounds, are not fitted inputs renamed as predictions.

full rationale

Theorem 1 (Section 2.2, Appendix A) is not circular: it assumes the exact KART representation f = sum_q g_q(sum_p psi_{p,q}) and uses density of the inner/outer ansatz classes A^z_M in C(I_z) to approximate psi_{p,q} and g_q separately. The proof is a standard epsilon-delta argument and depends on external results (KART, Cybenko, Kurkova), not on the paper's own fitted values. The phrase 'regardless of the basis function selection' overstates the density hypothesis, but that is a scope/correctness issue, not a definitional reduction. The empirical head-to-head results in Tables 5 and 6 are original benchmarks; the mismatched boundary conditions (Dirichlet KAN rows versus periodic KKAN rows) and the KKAN-only WNadResNet enhancement are experimental fairness confounds, not circular steps: no parameter was fitted to a subset and then reported as a prediction of the same subset. The IB/SNR stage framework and RBA are taken from prior work, including several self-citations, but they are used contextually, reproduced in Figures 11-14, and extended to new tasks; they do not carry the universal-approximation or benchmark claims by themselves. No equation in the paper reduces by construction to its own input, so there is no circularity to report.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the standard KART and denseness of the chosen basis families; for RBFs with fixed width the denseness is not fully justified in the paper. No physical entities are invented.

free parameters (5)
  • Number of KKAN features m = 32-64 depending on task
    Chosen by hand per benchmark; affects expressiveness and width of the inner expansion.
  • Outer basis polynomial degree D = 5-15 depending on task
    Chosen by hand; higher D can cause instability in cKAN but KKAN is stable.
  • ebMLP polynomial degree De = 2-15
    Chosen by hand for the inner block's Chebyshev embedding layers.
  • ssRBA hyperparameters (eta, lambda_max0, lambda_cap, Nstage, gamma_g, alpha, nu, c) = eta=0.01, lambda_max0=10, lambda_cap=20, Nstage=50000, gamma_g=0.99, alpha=0.99975, nu=2.0, c=0.5
    Set by hand based on prior RBA work; not systematically optimized in the paper.
  • Learning rates lr, lr_psi, lr_g = lr=1e-3/2e-4, lr_psi=1e-3, lr_g=2e-4 or 1e-3
    Chosen per architecture; KKAN uses separate inner/outer rates, and these choices affect convergence and final accuracy.
assumptions (4)
  • standard math KART holds: any continuous f on [0,1]^d can be represented as a sum of univariate inner/outer functions with m >= 2d+1.
    Theorem from Kolmogorov (1957) and variants; used in Section 2.2 to establish density of K_{m,d}^M.
  • domain assumption The parameterization classes A_M^z (MLPs, basis function expansions) are dense in C(I_z).
    Required for Theorem 1; holds for MLPs (Cybenko) and polynomials (Weierstrass), but questionable for RBFs with fixed width sigma.
  • domain assumption The discrete Dirichlet energy (geometric complexity) and SNR metrics are meaningful indicators of generalization.
    Assumed from prior work [61,62]; used to draw conclusions about learning stages and to motivate ssRBA.
  • domain assumption The training data and benchmarks are representative of scientific ML tasks.
    Used to claim general superiority of KKANs over MLPs and cKANs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of KKANs: Kurkova-Kolmogorov-Arnold Networks and Their Learning Dynamics." pith.science (2026). https://pith.science/paper/3XPZ5NBI

@misc{pith2026241216738,
  author       = {Pith},
  title        = {Pith review of: KKANs: Kurkova-Kolmogorov-Arnold Networks and Their Learning Dynamics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3XPZ5NBI}},
  note         = {Machine review of arXiv:2412.16738}
}
read the original abstract

Inspired by the Kolmogorov-Arnold representation theorem and Kurkova's principle of using approximate representations, we propose the Kurkova-Kolmogorov-Arnold Network (KKAN), a new two-block architecture that combines robust multi-layer perceptron (MLP) based inner functions with flexible linear combinations of basis functions as outer functions. We first prove that KKAN is a universal approximator, and then we demonstrate its versatility across scientific machine-learning applications, including function regression, physics-informed machine learning (PIML), and operator-learning frameworks. The benchmark results show that KKANs outperform MLPs and the original Kolmogorov-Arnold Networks (KANs) in function approximation and operator learning tasks and achieve performance comparable to fully optimized MLPs for PIML. To better understand the behavior of the new representation models, we analyze their geometric complexity and learning dynamics using information bottleneck theory, identifying three universal learning stages, fitting, transition, and diffusion, across all types of architectures. We find a strong correlation between geometric complexity and signal-to-noise ratio (SNR), with optimal generalization achieved during the diffusion stage. Additionally, we propose self-scaled residual-based attention weights to maintain high SNR dynamically, ensuring uniform convergence and prolonged learning.

Figures

Figures reproduced from arXiv: 2412.16738 by the authors.

Figure 1
Figure 1. KKAN-Inspired architecture. The inner block computes the inner functions by expand [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Enhanced-basis MLP (ebMLP). Each inner block expands its respective input dimension into [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Performance of KKANs for discontinuous function approximation. Columns show pre [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Results for discontinuous function approximation. (a) Relative [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Performance of KKANs+ssRBA for smooth function approximation. Columns show predic [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Results for smooth function approximation. (a) Relative [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Performance of KKAN+ssRBA for solving the Allen-Cahn equation. The columns display [PITH_FULL_IMAGE:figures/full_fig_p023_7.png]
Figure 8
Figure 8. Figure 8: Results for solving the Allen-Cahn Equation. (a) Relative [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: QR-DeepOKKAN predictions for three different initial conditions from the testing dataset. [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: Relative L 2 error convergence for operator learning the testing dataset. KKANs outperform cKANs and MLPs, while cKANs demonstrate faster initial convergence compared to the other models. The QR formulation enhances the performance of MLPs and KKANs but negatively imp…
Figure 11
Figure 11. Figure 11: Relative L 2 error (first row), SNR (second row), and geometric complexity (third row) con￾vergence for discontinuous function approximation using MLP (first column), cKAN (second column), and KKAN models (third column). The three stages of learning are observed acros…
Figure 12
Figure 12. Figure 12: Relative L 2 error (first row), SNR (second row), and geometric complexity (third row) convergence for smooth function approximation with MLP (first column), cKAN (second column), and KKAN models (third column). The three stages of learning are evident across all mode…
Figure 13
Figure 13. Figure 13: Relative L 2 error (first row), SNR (second row), and geometric complexity (third row) convergence for solving the Allen-Cahn Equation using MLP (first column), cKAN (second column), and KKAN models (third column). The three stages of learning are evident across all m…
Figure 14
Figure 14. Figure 14: Relative L 2 error (first row) and SNR (second row) for operator learning tasks using MLP (first column), cKAN (second column), and KKAN models (third column). Due to the high dimensionality of the branch net inputs (100), computing the geometric complexity is computa…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaled-cPIKANs: Domain Scaling in Chebyshev-based Physics-informed Kolmogorov-Arnold Networks

    math.NA 2025-01 conditional novelty 4.0 of 10

    Rescaling PDE spatial variables to [-1,1] before training Chebyshev-based physics-informed Kolmogorov-Arnold networks improves accuracy and convergence on wide oscillatory domains.

Reference graph

Works this paper leans on

153 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [52]

    L. F. Guilhoto, P. Perdikaris, Deep learning alternatives of the Kolmogorov su- perposition theorem, arXiv preprint arXiv:2410.01990 (2024)

  2. [41]

    Shukla, J

    K. Shukla, J. D. Toscano, Z. Wang, Z. Zou, G. E. Karniadakis, A comprehensive and F AIR comparison between MLP and KAN representations for differential equations and operator networks, Computer Methods in Applied Mechanics and Engineering 431 (2024) 117290. 40

  3. [1]

    J. D. Toscano, V. Oommen, A. J. Varghese, Z. Zou, N. A. Daryakenari, C. Wu, G. E. Karniadakis, From PINNs to PIKANs: Recent advances in physics- informed machine learning, arXiv preprint arXiv:2410.13228 (2024)

  4. [2]

    Raissi, H

    M. Raissi, H. Babaee, P. Givi, Deep learning of turbulent scalar mixing, Physical Review Fluids 4 (12) (2019) 124501

  5. [3]

    L. Lu, P. Jin, G. E. Karniadakis, DeepOnet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators, arXiv preprint arXiv:1910.03193 (2019)

  6. [4]

    Raissi, A

    M. Raissi, A. Yazdani, G. E. Karniadakis, Hidden fluid mechanics: Learning velocity and pressure fields from flow visualizations, Science 367 (6481) (2020) 1026–1030

  7. [5]

    S. Cai, H. Li, F. Zheng, F. Kong, M. Dao, G. E. Karniadakis, S. Suresh, Artifi- cial intelligence velocimetry and microaneurysm-on-a-chip for three-dimensional analysis of blood flow in physiology and disease, Proceedings of the National Academy of Sciences 118 (13) (2021). 37

  8. [6]

    K. A. Boster, S. Cai, A. Ladr´ on-de Guevara, J. Sun, X. Zheng, T. Du, J. H. Thomas, M. Nedergaard, G. E. Karniadakis, D. H. Kelley, Artificial intelligence velocimetry reveals in vivo flow rates, pressure gradients, and shear stresses in murine perivascular flows, Proceedings of the National Academy of Sciences 120 (14) (2023) e2217744120

Show all 153 references
  1. [7]

    J. D. Toscano, T. K¨ aufer, M. Maxey, C. Cierpka, G. E. Karniadakis, Inferring turbulent velocity and temperature fields and their statistics from Lagrangian ve- locity measurements using physics-informed Kolmogorov-Arnold Networks, arXiv preprint arXiv:2407.15727 (2024)

  2. [8]

    J. D. Toscano, C. Wu, A. Ladr´ on-de Guevara, T. Du, M. Nedergaard, D. H. Kelley, G. E. Karniadakis, K. A. Boster, Inferring in vivo murine cerebrospinal fluid flow using artificial intelligence velocimetry with moving boundaries and uncertainty quantification, Interface Focus...

  3. [9]

    G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, L. Yang, Physics-informed machine learning, Nature Reviews Physics 3 (6) (2021) 422– 440

  4. [10]

    Goodfellow, J

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, Y. Bengio, Generative adversarial networks, Communications of the ACM 63 (11) (2020) 139–144

  5. [11]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, Advances in Neural Infor- mation Processing Systems 30 (2017)

  6. [12]

    K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778

  7. [13]

    H. Li, Z. Xu, G. Taylor, C. Studer, T. Goldstein, Visualizing the loss landscape of neural nets, Advances in Neural Information Processing Systems 31 (2018)

  8. [14]

    J. D. Toscano, C. Zuniga-Navarrete, W. D. J. Siu, L. J. Segura, H. Sun, Teeth mold point cloud completion via data augmentation and hybrid RL-GAN, Jour- nal of Computing and Information Science in Engineering 23 (4) (2023) 041008

  9. [15]

    L. P. Kaelbling, M. L. Littman, A. W. Moore, Reinforcement learning: A survey, Journal of Artificial Intelligence Research 4 (1996) 237–285

  10. [16]

    J. Park, I. W. Sandberg, Universal approximation using radial-basis-function networks, Neural Computation 3 (2) (1991) 246–257. 38

  11. [17]

    Cranmer, Interpretable machine learning for science with PySR and Symbol- icRegression

    M. Cranmer, Interpretable machine learning for science with PySR and Symbol- icRegression. jl, arXiv preprint arXiv:2305.01582 (2023)

  12. [18]

    A. D. Jagtap, G. E. Karniadakis, Extended physics-informed neural networks (XPINNs): A generalized space-time domain decomposition based deep learn- ing framework for nonlinear partial differential equations, Communications in Computational Physics 28 (5) (2020)

  13. [19]

    A. D. Jagtap, E. Kharazmi, G. E. Karniadakis, Conservative physics-informed neural networks on discrete domains for conservation laws: Applications to for- ward and inverse problems, Computer Methods in Applied Mechanics and Engi- neering 365 (2020) 113028

  14. [20]

    A. D. Jagtap, K. Kawaguchi, G. E. Karniadakis, Adaptive activation functions accelerate convergence in deep and physics-informed neural networks, Journal of Computational Physics 404 (2020) 109136

  15. [21]

    S. Wang, Y. Teng, P. Perdikaris, Understanding and mitigating gradient flow pathologies in physics-informed neural networks, SIAM Journal on Scientific Computing 43 (5) (2021) A3055–A3081

  16. [22]

    L. D. McClenny, U. M. Braga-Neto, Self-adaptive physics-informed neural net- works, Journal of Computational Physics 474 (2023) 111722

  17. [23]

    Z. Liu, Y. Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljaˇ ci´ c, T. Y. Hou, M. Tegmark, KAN: Kolmogorov-Arnold Networks, arXiv preprint arXiv:2404.19756 (2024)

  18. [24]

    Z. Liu, P. Ma, Y. Wang, W. Matusik, M. Tegmark, KAN 2.0: Kolmogorov-Arnold networks meet science, arXiv preprint arXiv:2408.10205 (2024)

  19. [25]

    Y. Wang, J. W. Siegel, Z. Liu, T. Y. Hou, On the expressiveness and spectral bias of KANs, arXiv preprint arXiv:2410.01803 (2024)

  20. [26]

    C. J. Vaca-Rubio, L. Blanco, R. Pereira, M. Caus, Kolmogorov-Arnold Networks (KANs) for Time Series Analysis, arXiv preprint arXiv:2405.08790 (2024)

  21. [27]

    M. E. Samadi, Y. M¨ uller, A. Schuppert, Smooth Kolmogorov-Arnold networks enabling structural knowledge representation, arXiv preprint arXiv:2405.11318 (2024)

  22. [28]

    D. A. Sprecher, S. Draghici, Space-filling curves and Kolmogorov superposition- based neural networks, Neural Networks 15 (1) (2002) 57–67. 39

  23. [29]

    M. K¨ oppen, On the training of a Kolmogorov network, in: Artificial Neural Networks—ICANN 2002: International Conference Madrid, Spain, August 28– 30, 2002 Proceedings 12, Springer, 2002, pp. 474–479

  24. [30]

    Schmidhuber, Discovering neural nets with low Kolmogorov complexity and high generalization capability, Neural Networks 10 (5) (1997) 857–873

    J. Schmidhuber, Discovering neural nets with low Kolmogorov complexity and high generalization capability, Neural Networks 10 (5) (1997) 857–873

  25. [31]

    M.-J. Lai, Z. Shen, The Kolmogorov superposition theorem can break the curse of dimensionality when approximating high dimensional functions, arXiv preprint arXiv:2112.09963 (2021 (v1); 2024 (v5))

  26. [32]

    P.-E. Leni, Y. D. Fougerolle, F. Truchetet, The Kolmogorov spline network for image processing, in: Image Processing: Concepts, Methodologies, Tools, and Applications, IGI Global, 2013, pp. 54–78

  27. [33]

    He, On the optimal expressive power of ReLU DNNs and its applica- tion in approximation with Kolmogorov superposition theorem, arXiv preprint arXiv:2308.05509 (2023)

    J. He, On the optimal expressive power of ReLU DNNs and its applica- tion in approximation with Kolmogorov superposition theorem, arXiv preprint arXiv:2308.05509 (2023)

  28. [34]

    Somvanshi, S

    S. Somvanshi, S. A. Javed, M. M. Islam, D. Pandit, S. Das, A survey on Kolmogorov-Arnold network, arXiv preprint arXiv:2411.06078 (2024)

  29. [35]

    A. Pal, D. Das, Understanding the limitations of B-spline KANs: Convergence dynamics and computational efficiency, in: NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning

  30. [36]

    Li, Kolmogorov-Arnold networks are radial basis function networks, arXiv preprint arXiv:2405.06721 (2024)

    Z. Li, Kolmogorov-Arnold networks are radial basis function networks, arXiv preprint arXiv:2405.06721 (2024)

  31. [37]

    Bozorgasl, H

    Z. Bozorgasl, H. Chen, Wav-KAN: Wavelet Kolmogorov-Arnold networks (2024). arXiv:2405.12832

  32. [38]

    NLNR, JacobiKAN, https://github.com/mintisan/awesome-kan/ (2024)

  33. [39]

    Sidharth, A

    S. Sidharth, A. Keerthana, R. Gokul, K. Anas, Chebyshev polynomial-based Kolmogorov-Arnold networks: An efficient architecture for nonlinear function approximation, arXiv preprint arXiv:2405.07200 (2024)

  34. [40]

    S. S. Bhattacharjee, TorchKAN: Simplified KAN model with variations, https: //github.com/1ssb/torchkan/ (2024)

  35. [42]

    Y. Wang, J. Sun, J. Bai, C. Anitescu, M. S. Eshaghi, X. Zhuang, T. Rabczuk, Y. Liu, A physics-informed deep learning framework for solving forward and inverse problems based on Kolmogorov-Arnold networks, Computer Methods in Applied Mechanics and Engineering 433 (2025) 117518

  36. [43]

    D. W. Abueidda, P. Pantidis, M. E. Mobasher, DeepOKAN: Deep operator network based on Kolmogorov-Arnold networks for mechanics problems, arXiv preprint arXiv:2405.19143 (2024)

  37. [44]

    Zhang, T

    Z. Zhang, T. Shen, Y. Zhang, W. Zhang, Q. Wang, AL-PKAN: A hybrid GRU- KAN network with augmented lagrangian function for solving PDEs, Available at SSRN 4957859 (2024)

  38. [45]

    Jacob, A

    B. Jacob, A. A. Howard, P. Stinis, SPIKANs: Separable physics-informed Kolmogorov-Arnold networks, arXiv preprint arXiv:2411.06286 (2024)

  39. [46]

    A. A. Howard, B. Jacob, P. Stinis, Multifidelity Kolmogorov-Arnold networks, arXiv preprint arXiv: 2410.14764 (2024)

  40. [47]

    A. A. Howard, B. Jacob, S. H. Murphy, A. Heinlein, P. Stinis, Finite ba- sis Kolmogorov-Arnold networks: Domain decomposition for data-driven and physics-informed problems, arXiv preprint arXiv:2406.19662 (2024)

  41. [48]

    K. Ma, X. Lu, B. L. Nicola, B. Tang, Integrating Kolmogorov-Arnold networks with ordinary differential equations for efficient, interpretable and robust deep learning: A case study in the epidemiology of infectious diseases, medRxiv (2024) 2024–09

  42. [49]

    Fareaa, M

    A. Fareaa, M. S. Celebi, Learnable activation functions in physics-informed neural networks for solving partial differential equations, arXiv preprint arXiv:2411.15111 (2024)

  43. [50]

    Mostajeran, S

    F. Mostajeran, S. A. Faroughi, EPi-cKANs: Elasto-plasticity informed Kolmogorov-Arnold networks using Chebyshev polynomials, arXiv preprint arXiv:2410.10897 (2024)

  44. [51]

    Rigas, M

    S. Rigas, M. Papachristou, T. Papadopoulos, F. Anagnostopoulos, G. Alexan- dridis, Adaptive training of grid-dependent physics-informed Kolmogorov-Arnold networks, IEEE Access (2024)

  45. [53]

    Zhang, H

    X. Zhang, H. Zhou, Generalization bounds and model complexity for Kolmogorov-Arnold networks, arXiv preprint arXiv:2410.08026 (2024). 41

  46. [54]

    Igelnik, N

    B. Igelnik, N. Parikh, Kolmogorov’s spline network, IEEE Transactions on Neural Networks 14 (4) (2003) 725–733

  47. [55]

    K ˙ urkov´ a, Kolmogorov’s theorem is relevant, Neural Computation 3 (4) (1991) 617–622

    V. K ˙ urkov´ a, Kolmogorov’s theorem is relevant, Neural Computation 3 (4) (1991) 617–622

  48. [56]

    K ˙ urkov´ a, Kolmogorov’s theorem and multilayer neural networks, Neural net- works 5 (3) (1992) 501–506

    V. K ˙ urkov´ a, Kolmogorov’s theorem and multilayer neural networks, Neural net- works 5 (3) (1992) 501–506

  49. [57]

    Salimans, D

    T. Salimans, D. P. Kingma, Weight normalization: A simple reparameterization to accelerate training of deep neural networks, Advances in Neural Information Processing Systems 29 (2016)

  50. [58]

    S. Wang, H. Wang, P. Perdikaris, On the eigenvector bias of Fourier feature networks: From regression to solving multi-scale PDEs with physics-informed neural networks, Computer Methods in Applied Mechanics and Engineering 384 (2021) 113938

  51. [59]

    S. Wang, B. Li, Y. Chen, P. Perdikaris, PirateNets: Physics-informed Deep Learning with Residual Adaptive Networks, arXiv preprint arXiv:2402.00326 (2024)

  52. [60]

    Tishby, F

    N. Tishby, F. C. Pereira, W. Bialek, The information bottleneck method, arXiv preprint physics/0004057 (2000)

  53. [61]

    Dherin, M

    B. Dherin, M. Munn, M. Rosca, D. Barrett, Why neural networks find simple solutions: The many regularizers of geometric complexity, Advances in Neural Information Processing Systems 35 (2022) 2333–2349

  54. [62]

    S. J. Anagnostopoulos, J. D. Toscano, N. Stergiopulos, G. E. Karniadakis, Learn- ing in PINNs: Phase transition, total diffusion, and generalization, arXiv preprint arXiv:2403.18494 (2024)

  55. [63]

    S. J. Anagnostopoulos, J. D. Toscano, N. Stergiopulos, G. E. Karniadakis, Residual-based attention in physics-informed neural networks, Computer Meth- ods in Applied Mechanics and Engineering 421 (2024) 116805

  56. [64]

    Q. Liu, M. Chu, N. Thuerey, ConFIG: Towards conflict-free training of physics informed neural networks, arXiv preprint arXiv:2408.11104 (2024)

  57. [65]

    A. Kolmogorov, On the representation of continuous functions of several vari- ables by superpositions of continuous functions of a smaller number of variables, Proceedings of the USSR Academy of Sciences 108 (1956) 179–182, english trans- lation: Amer. Math. Soc. Transl., 17: ...

  58. [66]

    Arnold, On functions of three variables, Proceedings of the USSR Academy of Sciences 114 (1957) 679–681, english translation: Amer

    V. Arnold, On functions of three variables, Proceedings of the USSR Academy of Sciences 114 (1957) 679–681, english translation: Amer. Math. Soc. Transl., 28: Sixteen Papers on Analysis (1963), pp. 51–54

  59. [67]

    V. Arnold, On the representation of continuous functions of three variables as superpositions of continuous functions of two variables, Doklady Akademii Nauk SSSR 114 (4) (1957) 679–681, available on SpringerLink

  60. [68]

    Kolmogorov, On the representation of continuous functions of several variables as superpositions of continuous functions of one variable and additionEnglish translation: Amer

    A. Kolmogorov, On the representation of continuous functions of several variables as superpositions of continuous functions of one variable and additionEnglish translation: Amer. Math. Soc. Transl., 28: Sixteen Papers on Analysis (1963) (1957)

  61. [69]

    Girosi, T

    F. Girosi, T. Poggio, Representation properties of networks: Kolmogorov’s the- orem is irrelevant, Neural Computation 1 (4) (1989) 465–469

  62. [70]

    G. G. Lorentz, Metric entropy, widths, and superpositions of functions, The American Mathematical Monthly 69 (6) (1962) 469–485

  63. [71]

    Sprecher, Ph.d

    D. Sprecher, Ph.d. dissertation, Ph.D. thesis, University of Maryland (1963)

  64. [72]

    G. G. Lorentz, Approximation of Functions, Holt, Rinehart and Winston, Inc., 1966

  65. [73]

    Kahane, Sur le th´ eor` eme de superposition de Kolmogorov, Journal of Ap- proximation Theory 13 (1975) 229–234

    J.-P. Kahane, Sur le th´ eor` eme de superposition de Kolmogorov, Journal of Ap- proximation Theory 13 (1975) 229–234

  66. [74]

    G. G. Lorentz, M. Golitschek, Y. Makovoz, Constructive Approximation, Vol. 304 of Grundlehren der Mathematischen Wissenschaften, Springer, Berlin, 1996

  67. [75]

    D. A. Sprecher, On the structure of continuous functions of several variables, Transactions of the American Mathematical Society 115 (1965) 340–355

  68. [76]

    D. A. Sprecher, A numerical implementation of Kolmogorov’s superpositions, Neural Networks 9 (5) (1996) 765–772

  69. [77]

    D. A. Sprecher, A numerical implementation of Kolmogorov’s superpositions II, Neural Networks 10 (3) (1997) 447–457

  70. [78]

    Montanelli, H

    H. Montanelli, H. Yang, Error bounds for deep ReLU networks using the Kolmogorov–Arnold superposition theorem, Neural Networks 129 (2020) 1–6

  71. [79]

    Braun, An Application of Kolmogorov’s Superposition Theorem to Function Reconstruction in Higher Dimensions, Ph.D

    J. Braun, An Application of Kolmogorov’s Superposition Theorem to Function Reconstruction in Higher Dimensions, Ph.D. thesis, Universit¨ ats-und Landesbib- liothek Bonn (2009). 43

  72. [80]

    K¨ oppen, On the training of a kolmogorov network, in: ICANN 2002: Inter- national Conference on Artificial Neural Networks, Vol

    M. K¨ oppen, On the training of a kolmogorov network, in: ICANN 2002: Inter- national Conference on Artificial Neural Networks, Vol. 2415 of Lecture Notes in Computer Science, Springer, 2002, pp. 474–479

  73. [81]

    Braun, M

    J. Braun, M. Griebel, On a constructive proof of Kolmogorov’s superposition theorem, Constructive approximation 30 (2009) 653–675

  74. [82]

    D. A. Sprecher, From Algebra to Computational Algorithms: Kolmogorov and Hilbert’s Problem 13, Docent Press, 2017

  75. [83]

    Schmidt-Hieber, The Kolmogorov–Arnold representation theorem revisited, Neural Networks 137 (2021) 119–126

    J. Schmidt-Hieber, The Kolmogorov–Arnold representation theorem revisited, Neural Networks 137 (2021) 119–126

  76. [84]

    Bader, Space-Filling Curves: An Introduction with Applications in Scientific Computing, Springer, Berlin, Heidelberg, 2013

    M. Bader, Space-Filling Curves: An Introduction with Applications in Scientific Computing, Springer, Berlin, Heidelberg, 2013

  77. [85]

    Ismayilova, V

    A. Ismayilova, V. E. Ismailov, On the Kolmogorov neural networks, Neural Net- works 176 (2024) 106333

  78. [86]

    Ismailov, Addressing common misinterpretations of KART and UAT in neural network literature, arXiv preprint arXiv:2408.16389 (2024)

    V. Ismailov, Addressing common misinterpretations of KART and UAT in neural network literature, arXiv preprint arXiv:2408.16389 (2024)

  79. [87]

    Doss, A superposition theorem for unbounded continuous functions, Transac- tions of the American Mathematical Society 233 (1977) 197–203

    R. Doss, A superposition theorem for unbounded continuous functions, Transac- tions of the American Mathematical Society 233 (1977) 197–203

  80. [88]

    Demko, A superposition theorem for bounded continuous functions, Proceed- ings of the American Mathematical Society 66 (1) (1977) 75–78

    S. Demko, A superposition theorem for bounded continuous functions, Proceed- ings of the American Mathematical Society 66 (1) (1977) 75–78

  81. [89]

    Hattori, Dimension and superposition of bounded continuous functions on locally compact, separable metric spaces, Topology and its Applications 54 (1–3) (1993) 123–132

    Y. Hattori, Dimension and superposition of bounded continuous functions on locally compact, separable metric spaces, Topology and its Applications 54 (1–3) (1993) 123–132

  82. [90]

    Z. Feng, P. Gartside, Spaces with a finite family of basic functions, Bulletin of the London Mathematical Society 43 (1) (2011) 26–32

  83. [91]

    Sternfeld, Dimension, superposition of functions and separation of points, in compact metric spaces, Israel Journal of Mathematics 50 (1985) 13–53

    Y. Sternfeld, Dimension, superposition of functions and separation of points, in compact metric spaces, Israel Journal of Mathematics 50 (1985) 13–53

  84. [92]

    Laczkovich, A superposition theorem of Kolmogorov type for bounded con- tinuous functions, Journal of Approximation Theory 269 (2021) 105609

    M. Laczkovich, A superposition theorem of Kolmogorov type for bounded con- tinuous functions, Journal of Approximation Theory 269 (2021) 105609

  85. [93]

    Hecht-Nielsen, Kolmogorov’s mapping neural network existence theorem, in: Proceedings of the IEEE First International Conference on Neural Networks, Vol

    R. Hecht-Nielsen, Kolmogorov’s mapping neural network existence theorem, in: Proceedings of the IEEE First International Conference on Neural Networks, Vol. III, IEEE, Piscataway, NJ, 1987, pp. 11–13

  86. [94]

    A. G. Vitushkin, On Hilbert’s Thirteenth Problem, Doklady Akademii Nauk SSSR 95 (1954) 701–704. 44

  87. [95]

    Cybenko, Approximation by superpositions of a sigmoidal function, Mathe- matics of Control, Signals, and Systems 2 (4) (1989) 303–314

    G. Cybenko, Approximation by superpositions of a sigmoidal function, Mathe- matics of Control, Signals, and Systems 2 (4) (1989) 303–314

  88. [96]

    V. Brattka, From Hilbert’s 13th Problem to the theory of neural networks: con- structive aspects of Kolmogorov’s Superposition Theorem, Springer Berlin Hei- delberg, Berlin, Heidelberg, 2007, pp. 253–280

  89. [97]

    M. H. Freedman, The proof of Kolmogorov-Arnold may illuminate neural network learning, arXiv preprint arXiv:2410.08451 (2024)

  90. [98]

    Petersen, J

    P. Petersen, J. Zech, Mathematical Theory of Deep Learning, arXiv, 2024, arXiv:2407.18384 [cs.LG]

  91. [99]

    M.-J. Lai, Z. Shen, The optimal rate for linear KB-splines and LKB-splines ap- proximation of high dimensional continuous functions and its application, arXiv preprint arXiv:2401.03956 (2024)

  92. [100]

    Y. Wang, P. Jin, H. Xie, Tensor neural network and its numerical integration, arXiv preprint arXiv:2207.02754 (2022)

  93. [101]

    J. Cho, S. Nam, H. Yang, S.-B. Yun, Y. Hong, E. Park, Separable physics- informed neural networks, Advances in Neural Information Processing Systems 36 (2024)

  94. [102]

    E. C. Cyr, M. A. Gulian, R. G. Patel, M. Perego, N. A. Trask, Robust train- ing and initialization of deep neural networks: An adaptive basis viewpoint, in: Mathematical and Scientific Machine Learning, PMLR, 2020, pp. 512–536

  95. [103]

    Cai, Z.-Q

    W. Cai, Z.-Q. J. Xu, Multi-scale deep neural networks for solving high dimen- sional PDEs, arXiv preprint arXiv:1910.11710 (2019)

  96. [104]

    S. Wang, H. Wang, P. Perdikaris, On the eigenvector bias of Fourier feature networks: From regression to solving multi-scale PDEs with physics-informed neural networks, arXiv preprint arXiv:2012.10047 (2020)

  97. [105]

    L. Lu, P. Jin, G. Pang, Z. Zhang, G. E. Karniadakis, Learning nonlinear opera- tors via DeepONet based on the universal approximation theorem of operators, Nature machine intelligence 3 (3) (2021) 218–229

  98. [106]

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, A. Anandkumar, Fourier neural operator for parametric partial differential equa- tions, arXiv preprint arXiv:2010.08895 (2020)

  99. [107]

    S. Lee, Y. Shin, On the training and generalization of deep operator networks, SIAM Journal on Scientific Computing 46 (4) (2024) C273–C296. 45

  100. [108]

    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, arXiv preprint arXiv:1412.6980 (2014)

  101. [109]

    D. C. Liu, J. Nocedal, On the limited memory BFGS method for large scale optimization, Mathematical Programming 45 (1) (1989) 503–528

  102. [110]

    J. F. Urb´ an, P. Stefanou, J. A. Pons, Unveiling the optimization process of Physics Informed Neural Networks: How accurate and competitive can PINNs be?, arXiv preprint arXiv:2405.04230 (2024)

  103. [111]

    Tishby, N

    N. Tishby, N. Zaslavsky, Deep learning and the information bottleneck principle, in: 2015 IEEE Information Theory Workshop (ITW), IEEE, 2015, pp. 1–5

  104. [112]

    Shwartz-Ziv, N

    R. Shwartz-Ziv, N. Tishby, Opening the black box of deep neural networks via information, arXiv preprint arXiv:1703.00810 (2017)

  105. [113]

    Goldfeld, Y

    Z. Goldfeld, Y. Polyanskiy, The information bottleneck problem and its applica- tions in machine learning, IEEE Journal on Selected Areas in Information Theory 1 (1) (2020) 19–38

  106. [114]

    Shwartz-Ziv, Information flow in deep neural networks, arXiv preprint arXiv:2202.06749 (2022)

    R. Shwartz-Ziv, Information flow in deep neural networks, arXiv preprint arXiv:2202.06749 (2022)

  107. [115]

    B. Hao, U. Braga-Neto, C. Liu, L. Wang, M. Zhong, Structure preserving PINN for solving time dependent PDEs with periodic boundary, arXiv preprint arXiv:2404.16189 (2024)

  108. [116]

    Basir, I

    S. Basir, I. Senocak, Physics and equality constrained artificial neural networks: Application to forward and inverse problems with multi-fidelity data fusion, Jour- nal of Computational Physics 463 (2022) 111301

  109. [117]

    Basir, Investigating and mitigating failure modes in physics-informed neural networks (PINNs), arXiv preprint arXiv:2209.09988 (2022)

    S. Basir, Investigating and mitigating failure modes in physics-informed neural networks (PINNs), arXiv preprint arXiv:2209.09988 (2022)

  110. [118]

    Basir, I

    S. Basir, I. Senocak, An adaptive augmented Lagrangian method for train- ing physics and equality constrained artificial neural networks, arXiv preprint arXiv:2306.04904 (2023)

  111. [119]

    H. Son, S. W. Cho, H. J. Hwang, Enhanced physics-informed neural networks with augmented lagrangian relaxation method (AL-PINNs), Neurocomputing (2023) 126424

  112. [120]

    Y. Song, H. Wang, H. Yang, M. L. Taccari, X. Chen, Loss-attentional physics- informed neural networks, Journal of Computational Physics 501 (2024) 112781. 46

  113. [122]

    W. Chen, A. A. Howard, P. Stinis, Self-adaptive weights based on balanced resid- ual decay rate for physics-informed neural networks and deep operator networks, arXiv preprint arXiv:2407.01613 (2024)

  114. [123]

    Ramireza, J

    I. Ramireza, J. Pinoa, D. Pardob, M. Sanzc, L. del Rioe, A. Ortize, K. Moro- zovskaf, J. I. Aizpuruag, Residual-based attention physics-informed neural net- works for spatio-temporal ageing assessment of transformers operated in renew- able power plants, arXiv preprint arXiv:2...

  115. [124]

    S. Wang, P. Zhao, T. Song, Aspinn: An asymptotic strategy for solving singularly perturbed differential equations, arXiv preprint arXiv:2409.13185 (2024)

  116. [125]

    Ramirez, J

    I. Ramirez, J. Pino, D. Pardo, M. Sanz, L. del Rio, A. Ortiz, K. Morozovska, J. I. Aizpurua, Residual-based attention physics-informed neural networks for spatio-temporal ageing assessment of transformers operated in renewable power plants, Engineering Applications of Artifici...

  117. [126]

    S. Wang, P. Zhao, Q. Ma, T. Song, General-kindred physics-informed neural network to the solutions of singularly perturbed differential equations, Physics of Fluids 36 (11) (2024)

  118. [127]

    L. Lu, X. Meng, Z. Mao, G. E. Karniadakis, DeepXDE: A deep learning library for solving differential equations, SIAM Review 63 (1) (2021) 208–228

  119. [128]

    C. Wu, M. Zhu, Q. Tan, Y. Kartha, L. Lu, A comprehensive study of non- adaptive and residual-based adaptive sampling for physics-informed neural net- works, Computer Methods in Applied Mechanics and Engineering 403 (2023) 115671

  120. [129]

    X. Jin, S. Cai, H. Li, G. E. Karniadakis, NSFnets (Navier-Stokes flow nets): Physics-informed neural networks for the incompressible Navier-Stokes equations, Journal of Computational Physics 426 (2021) 109951

  121. [130]

    Xiang, W

    Z. Xiang, W. Peng, X. Liu, W. Yao, Self-adaptive loss balanced Physics-informed neural networks, Neurocomputing 496 (2022) 11–34

  122. [131]

    D. Liu, Y. Wang, A Dual-Dimer method for training physics-constrained neural networks with minimax architecture, Neural Networks 136 (2021) 112–125

  123. [132]

    S. Wang, S. Sankaran, P. Perdikaris, Respecting causality is all you need for train- ing physics-informed neural networks, arXiv preprint arXiv:2203.07404 (2022). 47

  124. [133]

    T. Zhou, X. Zhang, E. L. Droguett, A. Mosleh, A generic physics-informed neu- ral network-based framework for reliability assessment of multi-state systems, Reliability Engineering & System Safety 229 (2023) 108835

  125. [134]

    J. Yao, C. Su, Z. Hao, S. Liu, H. Su, J. Zhu, Multiadam: Parameter-wise scale- invariant optimizer for multiscale training of physics-informed neural networks, in: International Conference on Machine Learning, PMLR, 2023, pp. 39702– 39721

  126. [135]

    Ainsworth, J

    M. Ainsworth, J. Dong, Galerkin neural networks: A framework for approximat- ing variational equations with error control, SIAM Journal on Scientific Comput- ing 43 (4) (2021) A2474–A2501

  127. [136]

    Raissi, P

    M. Raissi, P. Perdikaris, G. E. Karniadakis, Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics 378 (2019) 686–707

  128. [137]

    S. Wang, H. Wang, P. Perdikaris, Learning the solution operator of paramet- ric partial differential equations with physics-informed DeepONets, Science Ad- vances 7 (40) (2021) eabi8605

  129. [138]

    Hornik, M

    K. Hornik, M. Stinchcombe, H. White, Multilayer feedforward networks are uni- versal approximators, Neural Networks 2 (5) (1989) 359–366

  130. [139]

    Glorot, Y

    X. Glorot, Y. Bengio, Understanding the difficulty of training deep feedforward neural networks, in: Proceedings of the thirteenth international conference on ar- tificial intelligence and statistics, JMLR Workshop and Conference Proceedings, 2010, pp. 249–256. 48 Appendix A. ...

  131. [140]

    Expand the Input Dimension: H 0 i = [C0, T0(xi), · · ·, CDeTDe(xi)], (B.5) where De is the polynomial degree, Tj denotes the Chebyshev polynomials, and Cj are trainable parameters

  132. [141]

    Apply an MLP with L Layers: Each layer is defined as: H l i = σ(W l−1 · H l−1 i + bl), (B.6) where θl = {W l, bl} are the weights and biases of the l−th layer, and σ is the activation function

  133. [142]

    Next, a combination layer aggregates the outputs along the input-dimension coordinate: ξq = dX i=1 Ψq,i(xi), (B.9) where d is the input dimension

    Apply a Second Polynomial Embedding: Expand the output of the MLP into an m-dimensional space: Ψi(xi) = [C L 0 , T0(H L i ), · · ·, CL DeTDe(H L i )], (B.7) Ψi(xi) = [Ψi,0, · · ·, Ψi,m], (B.8) where C L j are trainable parameters. Next, a combination layer aggregates the outpu...

  134. [143]

    Compute the transformed feature F from the input H using a weight-normalized layer and a nonlinear activation: F = σ(W · H + b), (C.7) where W and b are the weight and bias parameters of the layer, and σ is the hyperbolic tangent activation (tanh)

  135. [144]

    Apply a second transformation G to F : G = W ′ · F + b′, (C.8) where W ′ and b′ are parameters of another weight-normalized layer

  136. [145]

    The weights W and W ′ are reparameterized using weight normalization [57]

    Combine the original input H with the transformed feature G through an adap- tive residual connection: H ′ = σ(α · G + (1 − α) · H), (C.9) where α is a trainable scalar that adaptively balances the contributions of G and H. The weights W and W ′ are reparameterized using weigh...

  137. [146]

    First, the trunk network parameters µ and an auxiliary matrix A ∈ R(N +1)×K are optimized by solving: min µ,A ||Φ(µ)A − U||p p,p

    Optimize the Trunk Network. First, the trunk network parameters µ and an auxiliary matrix A ∈ R(N +1)×K are optimized by solving: min µ,A ||Φ(µ)A − U||p p,p. (D.1) After optimization, perform a QR decomposition of Φ(µ∗): Φ(µ∗) = Q∗R∗, where Q∗ is orthogonal, and R∗ is upper tr...

  138. [147]

    Using the precomputed R∗ and A∗, optimize the branch network parameters θ by solving: min θ ||C(θ) − R∗A∗||

    Optimize the Branch Network. Using the precomputed R∗ and A∗, optimize the branch network parameters θ by solving: min θ ||C(θ) − R∗A∗||. (D.2) This two-step process decouples the optimization of the trunk and branch networks, leveraging the QR decomposition to ensure numerica...

  139. [148]

    I denotes the input dimension

    Glorot [139] Learning rate lr 1e-3 2e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 2e-4 lr−Decay rate 0.9 0.9 0.9 lr−Decay step 5000 5000 5000 ssRBA:γ 0.999 0.999 0.999 ssRBA:η 0.01 0.01 0.01 ssRBA:λmax0 10 10 10 ssRBA:λcap 20 20 20 ssRBA:Nstage 50000 500...

  140. [149]

    I denotes the input dimension

    Glorot [139] Learning rate lr 1e-3 2e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 2e-4 lr−Decay rate 0.9 0.9 0.9 lr−Decay step 5000 5000 5000 ssRBA:γ 0.999 0.999 0.999 ssRBA:η 0.01 0.01 0.01 ssRBA:λmax0 10 10 10 ssRBA:λcap 20 20 20 ssRBA:Nstage 50000 500...

  141. [150]

    KKAN models include features of size m = 64, with polynomial degrees D for the outer blocks and De for the ebMLP inner blocks

    U − q 3 I , q 3 I Learning rate lr 1e-3 2e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 2e-4 lr−Decay rate 0.9 0.9 0.9 lr−Decay step 5000 5000 5000 Batch size 1e4 1e4 1e4 Table F.9: Implementation details for solving the Allen-Cahn Equation part (a). KKAN...

  142. [151]

    WNmMLP refers to the weight-normalized modified MLP architecture [21, 57]

    U − q 3 I , q 3 I Learning rate lr 1e-3 2e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 2e-4 lr−Decay rate 0.9 0.9 0.9 lr−Decay step 5000 5000 5000 Batch size 1e4 1e4 1e4 ssRBA-R:γ 0.999 0.999 0.999 ssRBA-R:η 0.01 0.01 0.01 ssRBA-R:λmax0 10 10 10 ssRBA-R:...

  143. [152]

    WNmMLP refers to the weight-normalized modified MLP architecture [21, 57], while WNadResNet refers to the weight- normalized adaptive residual network (see Section Appendix C.2)

    Glorot [139] Learning rate lr 1e-3 2e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 2e-4 lr−Decay rate 0.9 0.9 0.9 lr−Decay step 5000 5000 5000 ssRBA:γ 0.999 0.999 0.999 ssRBA:η 0.01 0.01 0.01 ssRBA:λmax0 10 10 10 ssRBA:λcap 20 20 20 ssRBA:γg 0.99 0.99 0.9...

  144. [153]

    The embedding dimension represents the number of neurons in the last layer of the branch and trunk networks

    [139] Learning rate lr 1e-3 3e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 1e-3 lr−Decay rate 0.9 0.9 0.99 lr−Decay step 2500 2500 5000 Table F.12: Implementation details for the Burgers equation using the DeepONet framework. The embedding dimension repr...

  145. [154]

    The embedding dimension refers to the number of neurons in the last layer of the branch and trunk networks

    [139] Learning rate lr 1e-3 3e-4 KKAN: Learning rate inner lrΨ 1e-3 KKAN: Learning rate outer lrG 1e-3 lr−Decay rate 0.9 0.9 0.99 lr−Decay step 2500 2500 5000 Table F.13: Implementation details for the Burgers equation using the QR-DeepONet framework. The embedding dimension r...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.