REVIEW 4 cited by
Early Stopping without a Validation Set
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Early stopping is a widely used technique to prevent poor generalization performance when training an over-expressive model by means of gradient-based optimization. To find a good point to halt the optimizer, a common practice is to split the dataset into a training and a smaller validation set to obtain an ongoing estimate of the generalization performance. We propose a novel early stopping criterion based on fast-to-compute local statistics of the computed gradients and entirely removes the need for a held-out validation set. Our experiments show that this is a viable approach in the setting of least-squares and logistic regression, as well as neural networks.
Forward citations
Cited by 4 Pith papers
-
ScoreStop: Gradient-based early stopping using functional score tests
ScoreStop introduces a functional score test for early stopping in gradient boosting, testing the null that the current predictor minimizes population risk with a scale-invariant statistic of known asymptotic distribution.
-
Early Stopping Against Label Noise Without Validation Data
The first local minimum of prediction changes on the noisy training set selects a near-optimal early stopping point without any validation data.
-
Towards Realistic Practices In Low-Resource Natural Language Processing: The Development Set
Using a development set for early stopping can change reported accuracy for low-resource NLP models, with per-language differences up to 18 percentage points compared to tuning training length on other languages.
-
AutoML: A Survey of the State-of-the-Art
A survey that organizes AutoML into a four-stage pipeline and reviews neural architecture search methods, their performance, and open problems.
Discussion (0). Continue with ORCID to comment.