Pith. sign in

REVIEW 1 cited by

Reinforced In-Context Black-Box Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17423 v3 pith:JYGICH64 submitted 2024-02-27 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords optimizationalgorithmhistoriesribboalgorithmsblack-boxdatain-context
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular components of BBO algorithms to speed up optimization and get rid of tedious hand-crafted heuristics. As an extension, learning the entire algorithm from data requires the least labor from experts and can provide the most flexibility. In this paper, we propose RIBBO, a method to reinforce-learn a BBO algorithm from offline data in an end-to-end fashion. RIBBO employs expressive sequence models to learn the optimization histories produced by multiple behavior algorithms and tasks, leveraging the in-context learning ability of large models to extract task information and make decisions accordingly. Central to our method is to augment the optimization histories with \textit{regret-to-go} tokens, which are designed to represent the performance of an algorithm based on cumulative regret over the future part of the histories. The integration of regret-to-go tokens enables RIBBO to automatically generate sequences of query points that satisfy the user-desired regret, which is verified by its universally good empirical performance on diverse problems, including BBO benchmark functions, hyper-parameter optimization and robot control problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ConfigX: Modular Configuration for Evolutionary Algorithms via Multitask Reinforcement Learning

    cs.LG 2024-12 reject novelty 6.0 of 10

    A unified RL policy can configure modular evolutionary algorithms within a family, but the claimed universal zero-shot generalization across algorithm families is not supported.

Pith tools