Pith. sign in

REVIEW 17 cited by

Hyper-Parameter Optimization: A Review of Algorithms and Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.05689 v1 pith:64XQQX76 submitted 2020-03-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords algorithmsdeeplearningoptimizationmajornetworkshyper-parametermodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since deep neural networks were developed, they have made huge contributions to everyday lives. Machine learning provides more rational advice than humans are capable of in almost every aspect of daily life. However, despite this achievement, the design and training of neural networks are still challenging and unpredictable procedures. To lower the technical thresholds for common users, automated hyper-parameter optimization (HPO) has become a popular topic in both academic and industrial areas. This paper provides a review of the most essential topics on HPO. The first section introduces the key hyper-parameters related to model training and structure, and discusses their importance and methods to define the value range. Then, the research focuses on major optimization algorithms and their applicability, covering their efficiency and accuracy especially for deep learning networks. This study next reviews major services and toolkits for HPO, comparing their support for state-of-the-art searching algorithms, feasibility with major deep learning frameworks, and extensibility for new modules designed by users. The paper concludes with problems that exist when HPO is applied to deep learning, a comparison between optimization algorithms, and prominent approaches for model evaluation with limited computational resources.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. End-to-end differentiable retrieval of molecular spectra using hydrodynamics, chemistry, and radiative transfer

    astro-ph.IM 2026-07 conditional novelty 6.0 of 10

    An end-to-end differentiable JAX pipeline couples 1D hydrodynamics, time-dependent chemistry, and radiative transfer, and recovers shock and rate parameters from synthetic HCO+ spectra.

  2. Exploiting Structural Properties for Efficient Constraint-Aware HNSW Hyperparameter Tuning

    cs.DB 2026-07 conditional novelty 6.0 of 10

    CHAT uses HNSW-specific monotonic and unimodal structure plus resource surrogates to tune M, efc, and efs under constraints, beating black-box tuners by up to 45% throughput or 11% recall and up to 44× faster convergence.

  3. Mock Deep Testing: Toward Separate Development of Data and Models for Deep Learning

    cs.SE 2025-02 conditional novelty 6.0 of 10

    KUnit's mock-based unit testing identified 63 issues in 50 DL programs and supported developers in resolving 63 issues in a user study with 36 participants.

  4. Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging

    cs.AI 2025-02 conditional novelty 6.0 of 10

    An automated multi-fidelity search framework discovers layer-wise and depth-wise model merging recipes that improve single- and multi-objective LLM reasoning performance without retraining.

  5. Mantis Shrimp: Exploring Photometric Band Utilization in Computer Vision Networks for Photometric Redshift Estimation

    astro-ph.IM 2025-01 conditional novelty 6.0 of 10

    A multi-survey CNN estimates photometric redshifts from GALEX, PanSTARRS, and UnWISE cutouts, with early and late image fusion performing comparably.

  6. Imaging Anisotropic Conductivity from Internal Measurements with Mixed Least-Squares Deep Neural Networks

    math.NA 2024-11 conditional novelty 6.0 of 10

    A mixed least-squares deep neural network recovers anisotropic conductivity tensors from internal gradient measurements, with error estimates and 2D/3D numerical demonstrations.

  7. ExpTest: Automating Learning Rate Searching and Tuning with Insights from Linearized Neural Networks

    cs.LG 2024-11 conditional novelty 6.0 of 10

    ExpTest auto-selects and tunes the learning rate by testing whether the training loss decays exponentially, without needing an initial learning rate choice.

  8. Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Adaptive re-sampling of RAHBO finds reliable RL hyperparameters more sample-efficiently than fixed-replication risk-averse or risk-neutral BO on offline multi-seed datasets.

  9. Fredholm Neural Networks for inverse problems in elliptic PDEs

    math.NA 2025-07 conditional novelty 5.0 of 10

    A boundary-integral based 'Fredholm neural network' converts fixed-point iterations into network layers and learns source terms for elliptic PDEs by backpropagating through the solver.

  10. Machine learning-based classification for Single Photon Space Debris Light Curves

    astro-ph.IM 2024-11 conditional novelty 5.0 of 10

    Machine learning, especially feature-based Random Forest and XGBoost, can classify single-photon space debris light curves with accuracies up to about 90.7 percent.

  11. Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A double ensemble that fuses features from pretrained CNNs and ViTs and ensembles tuned ML classifiers reaches 97.5% to 99.3% accuracy on three public brain MRI datasets, but the gains are not benchmarked against a he...

  12. Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification

    cs.CV 2025-06 reject novelty 4.0 of 10

    A ViT feature ensemble plus ML classifier voting pipeline is evaluated on two binary brain MRI datasets, reporting up to 99.8% accuracy without a same-dataset comparison against prior methods.

  13. Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A best-practices workflow for unsupervised scientific discovery, illustrated by a stability- and generalizability-driven clustering case study of Milky Way globular clusters using APOGEE data.

  14. Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification

    cs.CE 2025-07 conditional novelty 3.0 of 10

    DLRS adjusts the learning rate per epoch from the normalized first-to-last batch loss slope, and the authors report faster convergence on PINN and image classification benchmarks.

  15. A Unified Hyperparameter Optimization Pipeline for Transformer-Based Time Series Forecasting Models

    cs.LG 2025-01 conditional novelty 3.0 of 10

    A unified HPO pipeline built on Optuna and Ray Tune is applied to six time series forecasting models across three datasets, providing empirical guidance on hyperparameter choices.

  16. Crack Detection in Infrastructure Using Transfer Learning, Spatial Attention, and Genetic Algorithm Optimization

    cs.CV 2024-11 reject novelty 3.0 of 10

    The proposed Attention-ResNet50-GA pipeline reports 0.9967 precision and 0.9983 F1 for crack detection, but missing evaluation details prevent verification.

  17. Material synthesis through simulations guided by machine learning: a position paper

    cs.LG 2024-11 reject novelty 2.0 of 10

    The paper proposes ML-guided simulation for marble sludge mix design but only benchmarks porosity regression on an existing concrete dataset, leaving the central reuse claim untested.

Pith tools