Pith. sign in

Robust Bayesian optimization with reinforcement learned acquisition functions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In Bayesian optimization (BO) for expensive black-box optimization tasks, acquisition function (AF) guides sequential sampling and plays a pivotal role for efficient convergence to better optima. Prevailing AFs usually rely on artificial experiences in terms of preferences for exploration or exploitation, which runs a risk of a computational waste or traps in local optima and resultant re-optimization. To address the crux, the idea of data-driven AF selection is proposed, and the sequential AF selection task is further formalized as a Markov decision process (MDP) and resort to powerful reinforcement learning (RL) technologies. Appropriate selection policy for AFs is learned from superior BO trajectories to balance between exploration and exploitation in real time, which is called reinforcement-learning-assisted Bayesian optimization (RLABO). Competitive and robust BO evaluations on five benchmark problems demonstrate RL's recognition of the implicit AF selection pattern and imply the proposal's potential practicality for intelligent AF selection as well as efficient optimization in expensive black-box problems.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Direct Regret Optimization in Bayesian Optimization

cs.LG · 2025-07-09 · conditional · novelty 6.0

A decision transformer, trained offline on ROI-filtered, early-stopped GP-ensemble rollouts and refined by sparse real evaluations, is proposed as a non-myopic policy for Bayesian optimization.

citing papers explorer

Showing 1 of 1 citing paper.

  • Direct Regret Optimization in Bayesian Optimization cs.LG · 2025-07-09 · conditional · none · ref 27 · internal anchor

    A decision transformer, trained offline on ROI-filtered, early-stopped GP-ensemble rollouts and refined by sparse real evaluations, is proposed as a non-myopic policy for Bayesian optimization.