Pith. sign in

REVIEW 1 cited by

Optimal Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.15994 v2 pith:QD6OMWBW submitted 2022-03-30 cs.LG cs.NAmath.NAstat.ML

classification cs.LGcs.NAmath.NAstat.ML
keywords dataoptimallearningnearproblemclasserrorfunction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper studies the problem of learning an unknown function $f$ from given data about $f$. The learning problem is to give an approximation $\hat f$ to $f$ that predicts the values of $f$ away from the data. There are numerous settings for this learning problem depending on (i) what additional information we have about $f$ (known as a model class assumption), (ii) how we measure the accuracy of how well $\hat f$ predicts $f$, (iii) what is known about the data and data sites, (iv) whether the data observations are polluted by noise. A mathematical description of the optimal performance possible (the smallest possible error of recovery) is known in the presence of a model class assumption. Under standard model class assumptions, it is shown in this paper that a near optimal $\hat f$ can be found by solving a certain discrete over-parameterized optimization problem with a penalty term. Here, near optimal means that the error is bounded by a fixed constant times the optimal error. This explains the advantage of over-parameterization which is commonly used in modern machine learning. The main results of this paper prove that over-parameterized learning with an appropriate loss function gives a near optimal approximation $\hat f$ of the function $f$ from which the data is collected. Quantitative bounds are given for how much over-parameterization needs to be employed and how the penalization needs to be scaled in order to guarantee a near optimal recovery of $f$. An extension of these results to the case where the data is polluted by additive deterministic noise is also given.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Elliptic Regularity Theory in Barron Spaces and Applications to the Deep Ritz Method

    math.AP 2026-07 accept novelty 7.0 of 10

    Harmonic functions with Barron Dirichlet data fail to be Lipschitz or H², yet admit Barron approximants of norm ~|log ε| with error ~ε on half-spaces and 2D rectangles, giving Deep Ritz a priori rates.

Pith tools