Pith. sign in

REVIEW 1 cited by

Scalable Gradient-Based Tuning of Continuous Regularization Hyperparameters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1511.06727 v3 pith:KJDBXXD6 submitted 2015-11-20 cs.LG

classification cs.LG
keywords hyperparametersgradient-basedhyperparametermodelregularizationtrainingapproachcomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hyperparameter selection generally relies on running multiple full training trials, with selection based on validation set performance. We propose a gradient-based approach for locally adjusting hyperparameters during training of the model. Hyperparameters are adjusted so as to make the model parameter gradients, and hence updates, more advantageous for the validation cost. We explore the approach for tuning regularization hyperparameters and find that in experiments on MNIST, SVHN and CIFAR-10, the resulting regularization levels are within the optimal regions. The additional computational cost depends on how frequently the hyperparameters are trained, but the tested scheme adds only 30% computational overhead regardless of the model size. Since the method is significantly less computationally demanding compared to similar gradient-based approaches to hyperparameter optimization, and consistently finds good hyperparameter values, it can be a useful tool for training neural network models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sampling Imbalanced Data with Multi-objective Bilevel Optimization

    cs.LG 2025-06 reject novelty 6.0 of 10

    A heuristic bilevel optimization wrapper around SVM-SMOTE is claimed to improve minority-class F1 by selecting training samples that increase model-output variance and reduce overlap.

Pith tools