REVIEW 3 major objections 4 minor 47 references
How NOT to Fool the Masses When Giving Performance Results for Quantum Computers
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Quantum performance claims should meet the fair-test standards that reformed parallel-computing benchmarks.
desk verdict A sensible, well-written restatement of classical benchmarking ethics for quantum performance papers, weakened by an unproven prevalence claim and a soft normative premise, but worth publishing as a perspective piece. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs through three load-bearing tools. The first is the two-dimensional outcome space of heuristic solvers: every solver is instantiated with a work parameter $W$ and produces a pair ($S$, $T$) of solution quality and computation time, so comparative claims must be judged on the Pareto frontier of those pairs rather than on solution quality alone. The second is the no-free-lunch principle from optimization, which says that benchmark outcomes are governed mainly by the agreement between input structure and solver parameters, making tuning policy a component of the result rather than an external detail. The third is the time-to-solution (TTS) metric used in quantum speedup studies, combining a measured time per core operation with the number of core operations needed to succeed; the paper warns that TTS curves measured on small devices are commonly extrapolated to imaginary platforms, producing runtimes that cannot be observed.
What would settle it
A systematic survey of recent quantum performance papers that recorded whether runtimes, tuning times, and platform statuses were disclosed would settle the urgency claim; if omission was rare, or if re-analysis found that adding the missing data never changed any reported outperformance conclusion, the paper's central claim that these practices mislead the masses would be refuted.
Extended reading notes
Core claim
The central claim is that a performance claim of the form 'solver X outperforms solver Y' in quantum computing is an empirical statement whose validity depends on four pieces of disclosure. Because heuristic solvers trade solution quality against computation time through a user-set work parameter, a comparison that reports only success probability is incomplete: without runtimes, the reader cannot know whether X won because it is genuinely better or because it was allowed more time. Because benchmark outcomes are governed by the match between input structure and instantiated solver parameters, a tuned solver compared with untuned rivals is not a fair race, and the time spent tuning must be disclosed. Because the time-to-solution metric measured on small devices mixes real measurements with arithmetic adjustments, extrapolating those curves to larger or future platforms is not evidence of comparative performance. Because input sets in quantum studies are necessarily small, cherry-picking is a hazard that must be handled by justification and explicit qualification of the scope of conclusions. The paper applies these principles to anonymized examples from gate-model and quantum-annealing literature, and in one recalculated example the corrected sample sizes reverse the claimed success-probability comparison.
Load-bearing premise
The advice only binds if quantum performance comparisons aimed at the public should be judged by classical fair-test benchmarking rules, and it only matters if the four problematic practices are actually common; the paper asserts both rather than proving them.
Editorial extensions
If this is right
- A quantum paper that claims 'outperforms' but reports no runtimes is, on this view, unsupported; a reader can only treat the claim as an artifact of test design.
- Tuning effort is part of the performance result, so a QAOA comparison that omits parameter-search time cannot be compared fairly with a solver using defaults.
- Time-to-solution curves measured on small devices cannot be extrapolated to future hardware with any confidence; the paper cites cases where generational improvements changed speedups by four orders of magnitude.
- Cherry-picking is unavoidable, but it is acceptable only when the authors justify input and solver selection and qualify the scope of their conclusions.
- If the practices persist, the public audience may lose trust in quantum computing results, a dynamic the paper compares to the downsizing that followed parallel-computing benchmark scandals.
Reading between the lines
- Inference: the four principles could be turned into a one-page reporting checklist for quantum benchmarking papers, asking authors to state the work parameter setting, the tuning search cost, and the platform status for every solver.
- Inference: a systematic audit of recent gate-model and annealing papers would likely show that tuning time is the most commonly omitted quantity, since many physics-style studies treat parameter sweeps as preprocessing rather than part of the measurement.
- Inference: the paper's semantic-discord framing predicts that dispute resolution will come less from new benchmarks than from relabeling: physics-style studies should describe themselves as error-modeling investigations, reserving 'outperforms' for claims that include runtimes.
- Inference: extending the extrapolation warning, even well-calibrated quantum error correction and fault tolerance could change runtime ratios faster than any asymptotic model predicts, so performance predictions for post-fault-tolerance machines should carry explicit uncertainty bounds.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that quantum computing performance claims, especially comparative statements such as "X outperforms Y," should be held to the fair-test standards developed in classical benchmarking: report runtimes, disclose tuning time, avoid presenting extrapolations to nonexistent hardware as empirical results, and justify and qualify choices of test instances and comparison solvers. It illustrates each of these four principles with anonymized examples from the gate-model and quantum-annealing literature, discusses how disciplinary differences in the meaning of "performance" lead to miscommunication, and concludes with advice for researchers, journal editors, reviewers, and the public.
Significance. If read conditionally, the paper offers a useful and accessible checklist that could improve clarity, reproducibility, and trust in quantum performance comparisons. Its main strength is that it brings an established body of classical benchmarking methodology (Bailey, Johnson, and others) to a quantum audience that may be unfamiliar with it, and it does so with explicit caveats that physics-style studies are not being criticized per se. The anonymization of examples is a deliberate and mostly successful way to avoid ad hominem attacks, and the self-citations [20]–[22] are not load-bearing because the four principles are independently supported by other references. However, the paper's categorical framing overreaches its own caveats, and the prevalence claim in the abstract is not supported by the evidence presented, so the central claim needs revision rather than minor polishing.
major comments (3)
- [Four Principles (Principles 1–2) and Some Modest Suggestions] The four principles are stated as categorical 'Don'ts,' but the manuscript itself concedes that 'performance means different things to different groups' and that research papers about quantum performance are 'normally aimed at other researchers, not at the masses.' For a physics-style study of noise and error behavior, success probability may be the complete intended measure of output quality, with computation time deliberately excluded because the research question concerns fidelity and error rates, not speed. The paper never defends why the benchmarking/efficiency interpretation of 'outperforms' should take precedence over the authors' own stated definitions, nor why potential misreading by nonexpert audiences imposes an obligation on researchers writing for other researchers. Since the final section says 'If your paper is not intended to be read as a commercial benchmarking study, choose your words carefully,' the practical version of the principles is conditional, not categorical. The categorical version needs either a defended normative premise or an explicit reframing as audience-relative communication advice.
- [Abstract and Introduction] The abstract claims that 'the mistakes of three decades ago are being repeated by a new batch of researchers' and that a 'new generation' is committing these errors. This prevalence claim is not established by the four anonymized examples [R1]–[R4], which are presented as illustrations rather than as a systematic survey. The paper later qualifies itself: 'the discussion herein should not be interpreted as critiquing experimental methodology per se, but rather as illustrating how best practice in one field can be seen as biased and misleading practice in another.' These two statements are in tension. The abstract should either be softened to a claim about possible or illustrative miscommunication, or supported with systematic evidence about how common the practices are in the current literature.
- [Principle 3] The principle 'Don't claim faster runtimes for (or in comparison to) solvers running on imaginary platforms' is broadly sound as a warning about extrapolating TTS curves to nonexistent hardware. However, the section's first paragraph acknowledges that studying asymptotic performance on abstract models of computation is standard and valuable. As written, the principle could be read as condemning all extrapolation, including legitimate theoretical analysis. The paper should clarify that the prohibition concerns reporting estimated wall-clock times on devices that do not exist as if those times were empirical measurements, not asymptotic analysis on abstract models per se. This distinction is important because one of the paper's own examples [R3] involves literal extrapolated runtimes of 10^19 seconds, which is a different practice from theoretical scaling analysis.
minor comments (4)
- [Four Principles, Principle 1] In the Pareto-frontier discussion, the text says 'assuming lower is better in both dimensions,' but the same section uses success probability π, for which higher is better, and solution quality S, which is also normally higher-is-better. Please reconcile the notation, for example by defining S as an error rate or by explicitly stating 'lower is better for T and higher is better for S.'
- [Principle 4] The word 'inadvertantly' in the sentence about selecting inputs is a typo and should be 'inadvertently.'
- [References] Reference [41] lists 'Gerhard Reinhelt,' but the TSPLIB author is Gerhard Reinelt; please correct the spelling.
- [References] Several reference entries have truncated URLs in the text I reviewed (for example, reference [3] ends in '...computers-'); please ensure that complete URLs or DOIs are provided in the final version.
Circularity Check
No significant circularity: the four benchmarking principles are external norms applied to quantum reporting; the author's self-citations [20]-[22] are minor and not load-bearing.
full rationale
This is a position and guidance paper, not a derivation: the central claim is a set of four normative benchmarking principles, and there is no equation chain in which a predicted quantity reduces to an input quantity by construction. Walking the claimed chain: Principle 1 rests on the classical definition of heuristic performance as solution quality per unit computation time (the Pareto-frontier argument) and on external fair-test standards cited from the classical literature [4-9,15]; Principle 2 rests on the no-free-lunch principle [29,30] and fair-tuning guidelines [9,12,13,19]; Principle 3 rests on extrapolation warnings from Bailey [2] and Johnson [19]; Principle 4 rests on the NFL principle and scope-of-testing arguments from Johnson [19] and Hooker [18]. The only quantitative step is the illustrative expected time-to-solution example (pi_X=0.6 with T=1s vs. pi_Y=0.1 with T=1ms), which is the standard T/pi computation from the TTS literature, not a fitted parameter renamed as a prediction. The author's self-citations [20]-[22] appear in a survey sentence ('A large literature has also developed around guidelines for empirical evaluation of algorithms... [11] (see also [12-23])') and in the beginner's reading list ('papers marked with * in the reference list are a good place to start', which includes [22]); none of the four principles is justified by these self-citations, and deleting them would not change the argument, so they are minor and non-load-bearing. The paper also explicitly limits its own scope: 'the discussion herein should not be interpreted as critiquing experimental methodology per se,' it cautions 'Please remember, however, that correlation is not causation,' and its concluding recommendations are audience-conditional ('If your paper is not intended to be read as a commercial benchmarking study, choose your words carefully'). The skeptic's objection that the paper never defends the masses' interpretation of 'performance' over a physics-style success-probability definition is an unsettled normative premise, not a circular reduction, and belongs under correctness risk rather than circularity. Verdict: no significant circularity; score 1 reflects the minor non-load-bearing self-citations.
Assumptions & free parameters
assumptions (4)
- domain assumption Classical fair-test benchmarking standards should be the yardstick for quantum performance claims.
- domain assumption The no free lunch principle applies to instantiated heuristic solvers, so benchmark outcomes mainly reflect parameter and input agreement.
- ad hoc to paper The anonymized examples [R1]-[R4] are representative of current practices in quantum performance benchmarking.
- domain assumption Extrapolating runtime curves to untested sizes is unreliable for heuristic solvers.
Cite this review
Pith. "Pith review of How NOT to Fool the Masses When Giving Performance Results for Quantum Computers." pith.science (2026). https://pith.science/paper/XTUC6EJ5
@misc{pith2026241108860,
author = {Pith},
title = {Pith review of: How NOT to Fool the Masses When Giving Performance Results for Quantum Computers},
year = {2026},
howpublished = {\url{https://pith.science/paper/XTUC6EJ5}},
note = {Machine review of arXiv:2411.08860}
}
read the original abstract
In 1991, David Bailey wrote an article describing techniques for overstating the performance of massively parallel computers. Intended as a lighthearted protest against the practice of inflating benchmark results in order to ``fool the masses" and boost sales, the paper sparked development of procedural standards that help benchmarkers avoid methodological errors leading to unjustified and misleading conclusions. Now that quantum computers are starting to realize their potential as viable alternatives to classical computers, we can see the mistakes of three decades ago being repeated by a new batch of researchers who are unfamiliar with this history and these standards. Inspired by Bailey's model, this paper presents four suggestions for newcomers to quantum performance benchmarking, about how not to do it. They are: (1) Don't claim superior performance without mentioning runtimes; (2) Don't report optimized results without mentioning the tuning time needed to optimize those results; (3) Don't claim faster runtimes for (or in comparison to) solvers running on imaginary platforms; and (4) No cherry-picking (without justification and qualification). Suggestions for improving current practice appear in the last section.
Reference graph
Works this paper leans on
-
[20]
Catherine C. McGeoch. 1996. Feature article: Toward an expe rimental method for algorithm simulation, INFORMS Journal on Computing 8.1:1- 15. 15
work page 1996
-
[22]
*Catherine C. McGeoch. 2019. Principles and guidelines for quantum performance analysis, Quantum Technology and Optimization Problems (QTOP), LNCS 11413:36-48, Springer
work page 2019
-
[1]
David H. Bailey. 1991. Twelve ways to fool the masses when giving p erfor- mance results on parallel computers, Supercomputing Review 54-59. Re- trieved 2024 from davidhbailey.com/dhbpapers/twelve-ways.pdf
work page 1991
-
[2]
David H. Bailey. 1992. Misleading performance reporting in the sup ercom- puting field, Scientific Programming 1.2:141-151
work page 1992
-
[3]
John Markoff. 1991. Technology; Measuring How Fast Comput- ers Really Are, New York Times 3:14. Retrieved 2024 from www.nytimes.com/1991/09/22/business/technology-measuring-how-fast-computers-
work page 1991
-
[4]
*Jack Dongarra et al. 1987. Computer benchmarking: Paths and pit falls, IEEE Spectrum 24.7:38-43
work page 1987
-
[5]
Roger Hockney. 1996. The Science of Computer Benchmarking , So- ciety for Industrial and Applied Mathematics. Retrieved 2024 from epubs.siam.org
work page 1996
-
[6]
Torsten Hoefler and Roberto Belli. 2015. Scientific benchmarking of par- allel computing systems: Twelve ways to tell the masses when repor ting performance results, Proc. Int. Conf. High Performance Computing, Net- working, Storage and Analysis , 1-12. 14
work page 2015
Show all 47 references
-
[7]
Paul Fortier and Howard Mitchel. 2003. Computer Systems Performance Evaluation and Prediction , Elsevier
2003
-
[8]
*Lizy Kurian John and Lieven Eeckhout. 2017. Performance Evaluation and Benchmarking , CRC Press
2017
-
[9]
Mark Raasveldt et al. 2018. Fair benchmarking considered difficult : Com- mon pitfalls in database performance testing, Proc. of the Workshop on Testing Database Systems
2018
-
[10]
Benchmark (computing). (n.d.). Wikipedia, Wikimedia Foundation. Re- trieved 2024 from en.wikipedia.org/wiki/
2024
-
[11]
Crowder et al
Harlan P. Crowder et al. 1978. Reporting computational exper iments in mathematical programming, Mathematical Programming, 15.1:316-329
1978
-
[12]
*Richard S Barr et al. 1995. Designing and reporting on computationa l experiments with heuristic methods, Journal of Heuristics 1:9-32
1995
-
[13]
Thomas Bartz-Beielstein and Mark Preuss. 2010. Chapter 2: T he future of experimental research, in Thomas Bartz-Beielstein et al. (eds), Experimen- tal Methods for the Analysis of Optimization Algorithms . Berlin:Springer
2010
-
[14]
Thomas Bartz-Beielstein et al. 2020. Benchmarking in optimizatio n: Best practice and open issues. arXiv:2007.03488
2020 arXiv
-
[15]
Blackburn et al
*Stephen M. Blackburn et al. 2016. The truth, the whole truth, and noth- ing but the truth: A pragmatic guide to assessing empirical evaluatio ns, ACM Transactions on Programming Languages and Systems (TOP LAS) 38.4.15:1-20
2016
-
[16]
Paul R. Cohen. 1995. Empirical Methods for Artificial Intelligence . MIT Press
1995
-
[17]
Gent et al
Ian P. Gent et al. 1997. How Not To Do It , University of Leeds School of Computer Science Research Report Series, Report 97
1997
-
[18]
John Hooker. 1995. Testing Heuristics: We have it all wrong, Journal of Heuristics 1:33-42
1995
-
[19]
*David S. Johnson. 2001. A theoretician’s guide to the experimen- tal analysis of algorithms, in M. H. Goldwasser et al. (eds), Data Structures, Near Neighbor Searches, and Methodology: Fift h and Sixth DIMACS Implementation Challenges 59, AMS. Retrieved from dimacs.rutgers.ed...
2001
-
[21]
Catherine C. McGeoch. 2012. A Guide to Experimental Algorithmics , Cambridge Press
2012
-
[23]
Rardin and Reha Uzsoy
Ronald L. Rardin and Reha Uzsoy. 2001. Experimental evaluatio n of heuristic optimization algorithms: A tutorial, Journal of Heuristics 7:261- 304
2001
-
[24]
Hans Mittelmann. 2020. Benchmarking optimization software – a (hi)story, SN Operations Research Forum 1.2. See also the slide deck presented at EURO 2019 , retrieved 2024 from plato.asu.edu/talks/euro2019.pdf
2020
-
[25]
November 7, 2018
Announcement by Gurobi. November 7, 2018. Retrieved 2024 f rom plato.asu.edu/ftp/apology.pdf
2018
-
[26]
—. 2017. SPEC CPU ™ 2017 Run and Reporting Rules, SPEC™ Open Systems Group . Retrieved 2024 from spec.org/cpu2017/Docs/runrules.html
2017
-
[27]
Jack Dongarra et al. 2023. The LINPACK benchmark: Past, pr esent, and future, Concurrency and Computation 15.9:803-220
2023
-
[28]
Timothy Proctor et al. 2024. Benchmarking quantum computer s, arXiv:2407.088281
2024
-
[29]
Wolpert and William G
David H. Wolpert and William G. Macready. 1997. No free lunch theo rems for optimization, IEEE Transactions on evolutionary computation. 1.1:67- 68
1997
-
[30]
James McDermott. 2019. When and why metaheuristics resear chers can ignore NFL (Section 4). arXiv:1906.03280
2019 arXiv
-
[31]
Frank Hutter et al. 2006. Performance prediction and automa ted tuning of randomized and parametric algorithms: An initial investigation, Proc. 12th Int. Conf. on Principles and Practices of Constraint Pr ogramming (CP-06) 213-222
2006
-
[32]
Andreas Beham et al. 2018. Algorithm selection on generalized qu adratic assignment problem landscapes, Proc. of the Genetic and Evolutionary Computation Conference 253-260
2018
-
[33]
Lennart Bittel and Martin Kliesch. 2021. Training variational qu antum algorithms is NP-hard, Physical Review Letters 127.120502. 16
2021
-
[34]
Elijah Pelofske et al. 2024. Short-depth QAOA circuits and quan tum an- nealing on higher-order Ising models, npj Quantum Inf 10.30
2024
-
[35]
Roennow et al
Troels F. Roennow et al. 2014. Defining and detecting quantum s peedup, Science 345.6195:420-424
2014
-
[36]
Wolfgang Panny. 2010. Deletions in random binary search trees : A story of errors, Journal of Statistical Planning and Inference 140.8:2335-2345
2010
-
[37]
Derek Atkins et al. 1994. The magic words are squeamish ossifra ge, Proc. 4th Int. Conf. on Theory and Applications of Cryptology
1994
-
[38]
2024, MaxSAT Evaluation 2024
Jeremias Berg et al. 2024, MaxSAT Evaluation 2024 . Retrieved from maxsat-evaluations.github.io/2024
2024
-
[39]
Iain Dunning et al. 2018. What works best when? A systematic ev aluation of heuristics for Max-Cut and QUBO, INFORMS Journal on Computing ,
2018
-
[40]
Hans Mittelmann. (n.d.). Decision Tree for Optimization Software . Re- trieved 2024 from plato.asu.edu/guide.html
2024
-
[41]
Gerhard Reinhelt. (n.d.). TSPLIB, Ruprecht-Karls-Universit¨ at Heidel- berg. Retrieved 2024 from comopt.ifi.uni-heidelberg.de
2024
-
[42]
Hoos and Thomas St¨ utzle
Holger H. Hoos and Thomas St¨ utzle. 2000. SATLIB: An online re search resource for research on SAT, in I.P. Gent et al., editors, SAT 2000, IOS Press. SATLIB files are available from www.satlib.org
2000
-
[43]
George Santayana. 1917. The Life of Reason , Scribner’s Sons
1917
-
[44]
Stuart Rosenberg. 1967. Cool Hand Luke , Warner Bros/Seven Arts
1967
-
[45]
Hank Dietz. 1996. Blog post: Is parallel processing dead? Retr ieved 2024 from aggregate.ee.engr.uky.edu/Opinions/pardead.html
1996
-
[46]
B. Furht. 1994. Parallel computing: Glory and collapse, Computer 27.4:74- 75. 17
1994
-
[2018]
See also github.com/MQLIB
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.