Pith. sign in

REVIEW 1 cited by

Robust Bayesian optimization with reinforcement learned acquisition functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.00476 v1 pith:J3NY65XI submitted 2022-10-02 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords optimizationselectionbayesianacquisitionblack-boxefficientexpensiveexploitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Bayesian optimization (BO) for expensive black-box optimization tasks, acquisition function (AF) guides sequential sampling and plays a pivotal role for efficient convergence to better optima. Prevailing AFs usually rely on artificial experiences in terms of preferences for exploration or exploitation, which runs a risk of a computational waste or traps in local optima and resultant re-optimization. To address the crux, the idea of data-driven AF selection is proposed, and the sequential AF selection task is further formalized as a Markov decision process (MDP) and resort to powerful reinforcement learning (RL) technologies. Appropriate selection policy for AFs is learned from superior BO trajectories to balance between exploration and exploitation in real time, which is called reinforcement-learning-assisted Bayesian optimization (RLABO). Competitive and robust BO evaluations on five benchmark problems demonstrate RL's recognition of the implicit AF selection pattern and imply the proposal's potential practicality for intelligent AF selection as well as efficient optimization in expensive black-box problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Direct Regret Optimization in Bayesian Optimization

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A decision transformer, trained offline on ROI-filtered, early-stopped GP-ensemble rollouts and refined by sparse real evaluations, is proposed as a non-myopic policy for Bayesian optimization.

Pith tools