Pith. sign in

REVIEW 9 cited by

Benchmarking in Optimization: Best Practice and Open Issues

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.03488 v2 pith:IGHFFH2U submitted 2020-07-07 cs.NE cs.PFmath.OCstat.AP

Benchmarking in Optimization: Best Practice and Open Issues

classification cs.NE cs.PFmath.OCstat.AP
keywords benchmarkingbestdifferentgoaloptimizationpracticeactiveadequate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This survey compiles ideas and recommendations from more than a dozen researchers with different backgrounds and from different institutes around the world. Promoting best practice in benchmarking is its main goal. The article discusses eight essential topics in benchmarking: clearly stated goals, well-specified problems, suitable algorithms, adequate performance measures, thoughtful analysis, effective and efficient designs, comprehensible presentations, and guaranteed reproducibility. The final goal is to provide well-accepted guidelines (rules) that might be useful for authors and reviewers. As benchmarking in optimization is an active and evolving field of research this manuscript is meant to co-evolve over time by means of periodic updates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SPEC CPU: The Next Generation

    cs.PF 2026-05 unverdicted novelty 7.0

    SPEC CPU 2026 presents a new benchmark suite using open-source apps, expanded multithreading, and Rolling-Round-Robin Rate to address gaps in evaluating heterogeneous multiprogrammed CPU performance.

  2. Block-Bench: A Framework for Controllable and Transparent Discrete Optimization Benchmarking

    cs.NE 2026-04 unverdicted novelty 7.0

    Block-Bench constructs controllable discrete optimization benchmarks from block functions and dependency graphs to enable transparent analysis of algorithm behavior beyond the objective value.

  3. Learning to Assess the Reliability of Number-of-Runs Estimation in Stochastic Optimization

    cs.LG 2026-05 unverdicted novelty 6.0

    Supervised classifiers trained on 23 features from Nevergrad runs on COCO can detect unreliable run-number estimates with high minority-class recall in within-optimizer settings.

  4. Large-scale benchmarking of multi-objective soft-computing metaheuristics for redundancy allocation in repairable k-out-of-n systems

    cs.NE 2025-12 conditional novelty 6.0

    A 65-algorithm benchmark on a bi-objective repairable redundancy-allocation problem shows that algorithm rankings are budget-dependent and that Scaled Binomial Initialization changes relative performance.

  5. Standardizing case study descriptions for multi-energy systems and networks modeling

    eess.SY 2026-06 unverdicted novelty 4.0

    The authors adapt an existing standard into a unified description framework for multi-energy systems case studies, apply it to diverse cases, and develop a review checklist through cross-author evaluation.

  6. Improving Evaluation of Recombination-based Cartesian Genetic Programming

    cs.NE 2026-05 unverdicted novelty 4.0

    Hyperparameter optimization yields performance improvements for recombination-based Cartesian Genetic Programming on SRBench.

  7. From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors

    cs.LG 2026-03 conditional novelty 4.0

    Prompting LLMs with strong benchmark algorithm code, rather than relying on linguistic instructions, improves LLM-driven black-box optimization; the proposed BAG method outperforms five baselines on pbo and bbob.

  8. Time-Fair Benchmarking for Metaheuristics: A Restart-Fair Protocol for Fixed-Time Comparisons

    cs.NE 2025-09 conditional novelty 3.0

    A fixed-time, restart-fair benchmarking protocol for metaheuristics, using ERT and time-based performance profiles, is proposed without empirical validation.

  9. Asymmetry PRISM: A CPU/GPU Portfolio Optimization Engine for Deadline-Bounded Institutional Rebalancing

    q-fin.CP 2026-06 unverdicted novelty 2.0

    Asymmetry PRISM-CPU achieves 4.5x-24.1x speedups over reference solvers on N=100-2000 problems and GPU completes all 500 accounts in 109.5s where OSQP completes 4.