Pith. sign in

hub Canonical reference

Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks

Canonical reference. 100% of citing Pith papers cite this work as background.

24 Pith papers citing it
Background 100% of classified citations
abstract

We study the training process of Deep Neural Networks (DNNs) from the Fourier analysis perspective. We demonstrate a very universal Frequency Principle (F-Principle) -- DNNs often fit target functions from low to high frequencies -- on high-dimensional benchmark datasets such as MNIST/CIFAR10 and deep neural networks such as VGG16. This F-Principle of DNNs is opposite to the behavior of most conventional iterative numerical schemes (e.g., Jacobi method), which exhibit faster convergence for higher frequencies for various scientific computing problems. With a simple theory, we illustrate that this F-Principle results from the regularity of the commonly used activation functions. The F-Principle implies an implicit bias that DNNs tend to fit training data by a low-frequency function. This understanding provides an explanation of good generalization of DNNs on most real datasets and bad generalization of DNNs on parity function or randomized dataset.

hub tools

citation-role summary

background 5

citation-polarity summary

roles

background 5

polarities

background 5

representative citing papers

Fourier Feature Pyramids for Physics-Informed Neural Networks

cs.LG · 2026-05-22 · unverdicted · novelty 6.0

beignet replaces random Fourier feature embeddings in PINNs with a trainable multi-resolution Fourier feature pyramid, achieving higher accuracy on PDE benchmarks with fewer parameters and near machine precision residuals on the inviscid Burgers blowup using Adam.

Understanding Latent Diffusability via Fisher Geometry

cs.LG · 2026-04-03 · unverdicted · novelty 6.0

Latent diffusability is quantified by decomposing the MMSE rate along diffusion trajectories into Fisher Information and Fisher Information Rate, with three geometric penalties (dimensional compression, tangential distortion, curvature injection) identified as sources of failure.

Deep Learning for Subspace Regression

cs.LG · 2025-09-27 · unverdicted · novelty 6.0

Neural networks regress oversized subspaces for parametric problems using subspace-specific losses, with theory and experiments showing improved accuracy and smoother mappings.

WebSailor: Navigating Super-human Reasoning for Web Agent

cs.CL · 2025-07-03 · conditional · novelty 6.0

WebSailor trains open-source web agents to match proprietary performance on complex information-seeking tasks by generating high-uncertainty scenarios and using a new RL method called DUPO.

Conflict-Aware Harmonized Rotational Gradient for Multiscale Kinetic Regimes

cs.LG · 2026-04-27 · unverdicted · novelty 5.0

HRGrad resolves gradient conflicts in multi-task learning for asymptotic-preserving neural networks by encoding small parameters and using a gradient alignment metric, enabling stable training across all Knudsen numbers for BGK and linear transport equations.

A Practitioner's Guide to Kolmogorov-Arnold Networks

cs.LG · 2025-10-28 · accept · novelty 3.0

A systematic review of Kolmogorov-Arnold Networks that maps their relation to Kolmogorov superposition theory, MLPs, and kernels, examines basis-function design choices, summarizes performance advances, and supplies a practitioner's selection guide plus open challenges.

citing papers explorer

Showing 24 of 24 citing papers.