Pith. sign in

REVIEW 6 cited by

Optimizing Large-Scale Hyperparameters via Automated Learning Algorithm

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.09026 v1 pith:2MVB3RX5 submitted 2021-02-17 cs.LG

classification cs.LG
keywords optimizationhyperparameterhozoghyperparametersproblemalgorithmalgorithmsapproaches
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Modern machine learning algorithms usually involve tuning multiple (from one to thousands) hyperparameters which play a pivotal role in terms of model generalizability. Black-box optimization and gradient-based algorithms are two dominant approaches to hyperparameter optimization while they have totally distinct advantages. How to design a new hyperparameter optimization technique inheriting all benefits from both approaches is still an open problem. To address this challenging problem, in this paper, we propose a new hyperparameter optimization method with zeroth-order hyper-gradients (HOZOG). Specifically, we first exactly formulate hyperparameter optimization as an A-based constrained optimization problem, where A is a black-box optimization algorithm (such as deep neural network). Then, we use the average zeroth-order hyper-gradients to update hyperparameters. We provide the feasibility analysis of using HOZOG to achieve hyperparameter optimization. Finally, the experimental results on three representative hyperparameter (the size is from 1 to 1250) optimization tasks demonstrate the benefits of HOZOG in terms of simplicity, scalability, flexibility, effectiveness and efficiency compared with the state-of-the-art hyperparameter optimization methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CAZO uses curvature-informed, forward-only optimization to adapt models at test time, achieving strong accuracy with about 70% lower memory use than backprop methods.

  2. LiBOG: Lifelong Learning for Black-Box Optimizer Generation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LiBOG combines inter-task elastic weight consolidation and a new intra-task elite behavior consolidation, enabling a reinforcement-learning-based black-box optimizer generator to learn sequential task distributions wi...

  3. A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization

    math.OC 2025-05 conditional novelty 6.0 of 10

    A plug-and-play bilevel optimization framework attains O(√(n+m)/ε) sample complexity for PAGE, ZeroSARAH, and mixed stochastic estimators, matching lower bounds.

  4. qNBO: quasi-Newton Meets Bilevel Optimization

    cs.LG 2025-02 conditional novelty 6.0 of 10

    qNBO coordinates lower-level quasi-Newton iterations with inverse Hessian-vector products to approximate bilevel hypergradients, giving BFGS and SR1 instantiations and a non-asymptotic BFGS convergence rate.

  5. Predicting Barge Presence and Quantity on Inland Waterways using Vessel Tracking Data: A Machine Learning Approach

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Machine learning on AIS vessel tracking data can predict barge presence (F1 0.932) and quantity in six bins (F1 0.886), but the metrics are based on a small, post-hoc grouped sample.

  6. Zeroth-Order Optimization is Secretly Single-Step Policy Optimization

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Randomized finite-difference zeroth-order optimization is exactly single-step REINFORCE with baseline f(theta; xi), and adding an averaged baseline plus query reuse yields a faster ZOO algorithm (ZoAR).

Pith tools