Pith. sign in

Explainable Benchmarking for Iterative Optimization Heuristics

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Benchmarking heuristic algorithms is vital to understand under which conditions and on what kind of problems certain algorithms perform well. In most current research into heuristic optimization algorithms, only a very limited number of scenarios, algorithm configurations and hyper-parameter settings are explored, leading to incomplete and often biased insights and results. This paper presents a novel approach we call explainable benchmarking. Introducing the IOH-Xplainer software framework, for analyzing and understanding the performance of various optimization algorithms and the impact of their different components and hyper-parameters. We showcase the framework in the context of two modular optimization frameworks. Through this framework, we examine the impact of different algorithmic components and configurations, offering insights into their performance across diverse scenarios. We provide a systematic method for evaluating and interpreting the behaviour and efficiency of iterative optimization heuristics in a more transparent and comprehensible manner, allowing for better benchmarking and algorithm design.

fields

cs.NE 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery

cs.NE · 2025-07-04 · conditional · novelty 5.0

Comparing six LLaMEA prompt and selection variants on 5D BBOB problems, the 1+1 elitist variant using both simplify and random-perturbation prompts produced the best anytime performance, and behaviour metrics link this to stronger exploitation and less stagnation.

citing papers explorer

Showing 1 of 1 citing paper.

  • Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery cs.NE · 2025-07-04 · conditional · none · ref 29 · internal anchor

    Comparing six LLaMEA prompt and selection variants on 5D BBOB problems, the 1+1 elitist variant using both simplify and random-perturbation prompts produced the best anytime performance, and behaviour metrics link this to stronger exploitation and less stagnation.