REVIEW 3 cited by
Optimal approximation of continuous functions by very deep ReLU networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We consider approximations of general continuous functions on finite-dimensional cubes by general deep ReLU neural networks and study the approximation rates with respect to the modulus of continuity of the function and the total number of weights $W$ in the network. We establish the complete phase diagram of feasible approximation rates and show that it includes two distinct phases. One phase corresponds to slower approximations that can be achieved with constant-depth networks and continuous weight assignments. The other phase provides faster approximations at the cost of depths necessarily growing as a power law $L\sim W^{\alpha}, 0<\alpha\le 1,$ and with necessarily discontinuous weight assignments. In particular, we prove that constant-width fully-connected networks of depth $L\sim W$ provide the fastest possible approximation rate $\|f-\widetilde f\|_\infty = O(\omega_f(O(W^{-2/\nu})))$ that cannot be achieved with less deep networks.
Forward citations
Cited by 3 Pith papers
-
On Universality of Deep Equivariant Networks
Deep equivariant networks are universal over the entry-wise separable regime once depth stabilizes separation or a convolutional readout is added, unifying prior architecture-specific results.
-
Dimension independent bounds for general shallow networks
An abstract theorem gives dimension-independent, smoothness-improving approximation rates for shallow kernel networks, covering ReLU networks, RBFs, manifold learning, and quasirandom integration.
-
Deep ReLU network approximation of functions on a manifold
Deep ReLU networks approximate beta-Hoelder functions on a d*-dimensional manifold with O(epsilon^{-d*/beta} log(1/epsilon)) nonzero parameters, and empirical risk minimization achieves risk n^{-2 beta/(2 beta + d*)} ...
Discussion (0). Continue with ORCID to comment.