REVIEW 17 cited by
Hyper-Parameter Optimization: A Review of Algorithms and Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Since deep neural networks were developed, they have made huge contributions to everyday lives. Machine learning provides more rational advice than humans are capable of in almost every aspect of daily life. However, despite this achievement, the design and training of neural networks are still challenging and unpredictable procedures. To lower the technical thresholds for common users, automated hyper-parameter optimization (HPO) has become a popular topic in both academic and industrial areas. This paper provides a review of the most essential topics on HPO. The first section introduces the key hyper-parameters related to model training and structure, and discusses their importance and methods to define the value range. Then, the research focuses on major optimization algorithms and their applicability, covering their efficiency and accuracy especially for deep learning networks. This study next reviews major services and toolkits for HPO, comparing their support for state-of-the-art searching algorithms, feasibility with major deep learning frameworks, and extensibility for new modules designed by users. The paper concludes with problems that exist when HPO is applied to deep learning, a comparison between optimization algorithms, and prominent approaches for model evaluation with limited computational resources.
Forward citations
Cited by 17 Pith papers
-
End-to-end differentiable retrieval of molecular spectra using hydrodynamics, chemistry, and radiative transfer
An end-to-end differentiable JAX pipeline couples 1D hydrodynamics, time-dependent chemistry, and radiative transfer, and recovers shock and rate parameters from synthetic HCO+ spectra.
-
Exploiting Structural Properties for Efficient Constraint-Aware HNSW Hyperparameter Tuning
CHAT uses HNSW-specific monotonic and unimodal structure plus resource surrogates to tune M, efc, and efs under constraints, beating black-box tuners by up to 45% throughput or 11% recall and up to 44× faster convergence.
-
Mock Deep Testing: Toward Separate Development of Data and Models for Deep Learning
KUnit's mock-based unit testing identified 63 issues in 50 DL programs and supported developers in resolving 63 issues in a user study with 36 participants.
-
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
An automated multi-fidelity search framework discovers layer-wise and depth-wise model merging recipes that improve single- and multi-objective LLM reasoning performance without retraining.
-
Mantis Shrimp: Exploring Photometric Band Utilization in Computer Vision Networks for Photometric Redshift Estimation
A multi-survey CNN estimates photometric redshifts from GALEX, PanSTARRS, and UnWISE cutouts, with early and late image fusion performing comparably.
-
Imaging Anisotropic Conductivity from Internal Measurements with Mixed Least-Squares Deep Neural Networks
A mixed least-squares deep neural network recovers anisotropic conductivity tensors from internal gradient measurements, with error estimates and 2D/3D numerical demonstrations.
-
ExpTest: Automating Learning Rate Searching and Tuning with Insights from Linearized Neural Networks
ExpTest auto-selects and tunes the learning rate by testing whether the training loss decays exponentially, without needing an initial learning rate choice.
-
Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL
Adaptive re-sampling of RAHBO finds reliable RL hyperparameters more sample-efficiently than fixed-replication risk-averse or risk-neutral BO on offline multi-seed datasets.
-
Fredholm Neural Networks for inverse problems in elliptic PDEs
A boundary-integral based 'Fredholm neural network' converts fixed-point iterations into network layers and learns source terms for elliptic PDEs by backpropagating through the solver.
-
Machine learning-based classification for Single Photon Space Debris Light Curves
Machine learning, especially feature-based Random Forest and XGBoost, can classify single-photon space debris light curves with accuracies up to about 90.7 percent.
-
Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification
A double ensemble that fuses features from pretrained CNNs and ViTs and ensembles tuned ML classifiers reaches 97.5% to 99.3% accuracy on three public brain MRI datasets, but the gains are not benchmarked against a he...
-
Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification
A ViT feature ensemble plus ML classifier voting pipeline is evaluated on two binary brain MRI datasets, reporting up to 99.8% accuracy without a same-dataset comparison against prior methods.
-
Unsupervised Machine Learning for Scientific Discovery: Workflow and Best Practices
A best-practices workflow for unsupervised scientific discovery, illustrated by a stability- and generalizability-driven clustering case study of Milky Way globular clusters using APOGEE data.
-
Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
DLRS adjusts the learning rate per epoch from the normalized first-to-last batch loss slope, and the authors report faster convergence on PINN and image classification benchmarks.
-
A Unified Hyperparameter Optimization Pipeline for Transformer-Based Time Series Forecasting Models
A unified HPO pipeline built on Optuna and Ray Tune is applied to six time series forecasting models across three datasets, providing empirical guidance on hyperparameter choices.
-
Crack Detection in Infrastructure Using Transfer Learning, Spatial Attention, and Genetic Algorithm Optimization
The proposed Attention-ResNet50-GA pipeline reports 0.9967 precision and 0.9983 F1 for crack detection, but missing evaluation details prevent verification.
-
Material synthesis through simulations guided by machine learning: a position paper
The paper proposes ML-guided simulation for marble sludge mix design but only benchmarks porosity regression on an existing concrete dataset, leaving the central reuse claim untested.
Discussion (0). Continue with ORCID to comment.