REVIEW 3 cited by
Convergence Analysis of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution at a linear convergence rate for the quadratic loss function. However, the learning rate of GD for training two-layer neural networks exhibits poor dependence on the sample size and the Gram matrix, leading to a slow training process. In this paper, we show that for training two-layer $\text{ReLU}^3$ Physics-Informed Neural Networks (PINNs), the learning rate can be improved from $\mathcal{O}(\lambda_0)$ to $\mathcal{O}(1/\|\bm{H}^{\infty}\|_2)$, implying that GD actually enjoys a faster convergence rate. Despite such improvements, the convergence rate is still tied to the least eigenvalue of the Gram matrix, leading to slow convergence. We then develop the positive definiteness of Gram matrices with general smooth activation functions and provide the convergence analysis of natural gradient descent (NGD) in training two-layer PINNs, demonstrating that the learning rate can be $\mathcal{O}(1)$ and at this rate, the convergence rate is independent of the Gram matrix. In particular, for smooth activation functions, the convergence rate of NGD is quadratic. Numerical experiments are conducted to verify our theoretical results.
Forward citations
Cited by 3 Pith papers
-
The Differential Neural Tangent Kernel and Its Positivity
The infinite-width Differential Neural Tangent Kernel is positive definite for shallow and deep networks under RePU or smooth non-polynomial activations and all linear differential operators.
-
Feature Learning for the High Dimensional Stationary Sch\"odinger Equation with Deep Ritz Method
Gradient descent on single-index and two-neuron models provably recovers feature directions of the Schrödinger equation source term in the deep Ritz framework.
-
A Sketch-and-Project Analysis of Subsampled Natural Gradient Algorithms
For linear least squares, SNGD and SPRING are proved equivalent to accelerated regularized Kaczmarz methods, yielding the first fast rates and first SPRING guarantee; the general quadratic analysis holds under strong ...
Discussion (0). Sign in to comment.