Pith. sign in

REVIEW 3 major objections 4 minor 47 references

How NOT to Fool the Masses When Giving Performance Results for Quantum Computers

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Quantum performance claims should meet the fair-test standards that reformed parallel-computing benchmarks.

desk verdict A sensible, well-written restatement of classical benchmarking ethics for quantum performance papers, weakened by an unproven prevalence claim and a soft normative premise, but worth publishing as a perspective piece. read the letter →

arxiv 2411.08860 v1 pith:XTUC6EJ5 submitted 2024-11-13 quant-ph cs.ET

classification quant-phcs.ET
keywords quantumbenchmarkingperformanceclaimsruntimereportingtuningtimefairtestingcherry-pickingspeedupno-free-lunchprinciple
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that quantum-computer performance comparisons are repeating the benchmark-reporting mistakes that once undermined trust in parallel computers. It turns a famous 1991 satirical list of ways to overstate performance into four positive rules for quantum researchers: report runtimes alongside solution quality, report the time spent tuning parameters, do not compare against solvers running on imagined platforms, and do not cherry-pick inputs or competitors without justification and qualification. The claim is that a public 'X outperforms Y' statement is not self-justifying: without these disclosures, readers cannot tell whether the result is genuine or an artifact of the test design. The paper locates the root cause in a clash between physics-style studies of hardware error and the expectations of commercial benchmarking, and it offers separate advice for researchers, editors, and the public.

What carries the argument

The argument runs through three load-bearing tools. The first is the two-dimensional outcome space of heuristic solvers: every solver is instantiated with a work parameter $W$ and produces a pair ($S$, $T$) of solution quality and computation time, so comparative claims must be judged on the Pareto frontier of those pairs rather than on solution quality alone. The second is the no-free-lunch principle from optimization, which says that benchmark outcomes are governed mainly by the agreement between input structure and solver parameters, making tuning policy a component of the result rather than an external detail. The third is the time-to-solution (TTS) metric used in quantum speedup studies, combining a measured time per core operation with the number of core operations needed to succeed; the paper warns that TTS curves measured on small devices are commonly extrapolated to imaginary platforms, producing runtimes that cannot be observed.

What would settle it

A systematic survey of recent quantum performance papers that recorded whether runtimes, tuning times, and platform statuses were disclosed would settle the urgency claim; if omission was rare, or if re-analysis found that adding the missing data never changed any reported outperformance conclusion, the paper's central claim that these practices mislead the masses would be refuted.

Watch

Extended reading notes

Core claim

The central claim is that a performance claim of the form 'solver X outperforms solver Y' in quantum computing is an empirical statement whose validity depends on four pieces of disclosure. Because heuristic solvers trade solution quality against computation time through a user-set work parameter, a comparison that reports only success probability is incomplete: without runtimes, the reader cannot know whether X won because it is genuinely better or because it was allowed more time. Because benchmark outcomes are governed by the match between input structure and instantiated solver parameters, a tuned solver compared with untuned rivals is not a fair race, and the time spent tuning must be disclosed. Because the time-to-solution metric measured on small devices mixes real measurements with arithmetic adjustments, extrapolating those curves to larger or future platforms is not evidence of comparative performance. Because input sets in quantum studies are necessarily small, cherry-picking is a hazard that must be handled by justification and explicit qualification of the scope of conclusions. The paper applies these principles to anonymized examples from gate-model and quantum-annealing literature, and in one recalculated example the corrected sample sizes reverse the claimed success-probability comparison.

Load-bearing premise

The advice only binds if quantum performance comparisons aimed at the public should be judged by classical fair-test benchmarking rules, and it only matters if the four problematic practices are actually common; the paper asserts both rather than proving them.

Editorial extensions

If this is right

  • A quantum paper that claims 'outperforms' but reports no runtimes is, on this view, unsupported; a reader can only treat the claim as an artifact of test design.
  • Tuning effort is part of the performance result, so a QAOA comparison that omits parameter-search time cannot be compared fairly with a solver using defaults.
  • Time-to-solution curves measured on small devices cannot be extrapolated to future hardware with any confidence; the paper cites cases where generational improvements changed speedups by four orders of magnitude.
  • Cherry-picking is unavoidable, but it is acceptable only when the authors justify input and solver selection and qualify the scope of their conclusions.
  • If the practices persist, the public audience may lose trust in quantum computing results, a dynamic the paper compares to the downsizing that followed parallel-computing benchmark scandals.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the four principles could be turned into a one-page reporting checklist for quantum benchmarking papers, asking authors to state the work parameter setting, the tuning search cost, and the platform status for every solver.
  • Inference: a systematic audit of recent gate-model and annealing papers would likely show that tuning time is the most commonly omitted quantity, since many physics-style studies treat parameter sweeps as preprocessing rather than part of the measurement.
  • Inference: the paper's semantic-discord framing predicts that dispute resolution will come less from new benchmarks than from relabeling: physics-style studies should describe themselves as error-modeling investigations, reserving 'outperforms' for claims that include runtimes.
  • Inference: extending the extrapolation warning, even well-calibrated quantum error correction and fault tolerance could change runtime ratios faster than any asymptotic model predicts, so performance predictions for post-fault-tolerance machines should carry explicit uncertainty bounds.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that quantum computing performance claims, especially comparative statements such as "X outperforms Y," should be held to the fair-test standards developed in classical benchmarking: report runtimes, disclose tuning time, avoid presenting extrapolations to nonexistent hardware as empirical results, and justify and qualify choices of test instances and comparison solvers. It illustrates each of these four principles with anonymized examples from the gate-model and quantum-annealing literature, discusses how disciplinary differences in the meaning of "performance" lead to miscommunication, and concludes with advice for researchers, journal editors, reviewers, and the public.

Significance. If read conditionally, the paper offers a useful and accessible checklist that could improve clarity, reproducibility, and trust in quantum performance comparisons. Its main strength is that it brings an established body of classical benchmarking methodology (Bailey, Johnson, and others) to a quantum audience that may be unfamiliar with it, and it does so with explicit caveats that physics-style studies are not being criticized per se. The anonymization of examples is a deliberate and mostly successful way to avoid ad hominem attacks, and the self-citations [20]–[22] are not load-bearing because the four principles are independently supported by other references. However, the paper's categorical framing overreaches its own caveats, and the prevalence claim in the abstract is not supported by the evidence presented, so the central claim needs revision rather than minor polishing.

major comments (3)
  1. [Four Principles (Principles 1–2) and Some Modest Suggestions] The four principles are stated as categorical 'Don'ts,' but the manuscript itself concedes that 'performance means different things to different groups' and that research papers about quantum performance are 'normally aimed at other researchers, not at the masses.' For a physics-style study of noise and error behavior, success probability may be the complete intended measure of output quality, with computation time deliberately excluded because the research question concerns fidelity and error rates, not speed. The paper never defends why the benchmarking/efficiency interpretation of 'outperforms' should take precedence over the authors' own stated definitions, nor why potential misreading by nonexpert audiences imposes an obligation on researchers writing for other researchers. Since the final section says 'If your paper is not intended to be read as a commercial benchmarking study, choose your words carefully,' the practical version of the principles is conditional, not categorical. The categorical version needs either a defended normative premise or an explicit reframing as audience-relative communication advice.
  2. [Abstract and Introduction] The abstract claims that 'the mistakes of three decades ago are being repeated by a new batch of researchers' and that a 'new generation' is committing these errors. This prevalence claim is not established by the four anonymized examples [R1]–[R4], which are presented as illustrations rather than as a systematic survey. The paper later qualifies itself: 'the discussion herein should not be interpreted as critiquing experimental methodology per se, but rather as illustrating how best practice in one field can be seen as biased and misleading practice in another.' These two statements are in tension. The abstract should either be softened to a claim about possible or illustrative miscommunication, or supported with systematic evidence about how common the practices are in the current literature.
  3. [Principle 3] The principle 'Don't claim faster runtimes for (or in comparison to) solvers running on imaginary platforms' is broadly sound as a warning about extrapolating TTS curves to nonexistent hardware. However, the section's first paragraph acknowledges that studying asymptotic performance on abstract models of computation is standard and valuable. As written, the principle could be read as condemning all extrapolation, including legitimate theoretical analysis. The paper should clarify that the prohibition concerns reporting estimated wall-clock times on devices that do not exist as if those times were empirical measurements, not asymptotic analysis on abstract models per se. This distinction is important because one of the paper's own examples [R3] involves literal extrapolated runtimes of 10^19 seconds, which is a different practice from theoretical scaling analysis.
minor comments (4)
  1. [Four Principles, Principle 1] In the Pareto-frontier discussion, the text says 'assuming lower is better in both dimensions,' but the same section uses success probability π, for which higher is better, and solution quality S, which is also normally higher-is-better. Please reconcile the notation, for example by defining S as an error rate or by explicitly stating 'lower is better for T and higher is better for S.'
  2. [Principle 4] The word 'inadvertantly' in the sentence about selecting inputs is a typo and should be 'inadvertently.'
  3. [References] Reference [41] lists 'Gerhard Reinhelt,' but the TSPLIB author is Gerhard Reinelt; please correct the spelling.
  4. [References] Several reference entries have truncated URLs in the text I reviewed (for example, reference [3] ends in '...computers-'); please ensure that complete URLs or DOIs are provided in the final version.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the four benchmarking principles are external norms applied to quantum reporting; the author's self-citations [20]-[22] are minor and not load-bearing.

full rationale

This is a position and guidance paper, not a derivation: the central claim is a set of four normative benchmarking principles, and there is no equation chain in which a predicted quantity reduces to an input quantity by construction. Walking the claimed chain: Principle 1 rests on the classical definition of heuristic performance as solution quality per unit computation time (the Pareto-frontier argument) and on external fair-test standards cited from the classical literature [4-9,15]; Principle 2 rests on the no-free-lunch principle [29,30] and fair-tuning guidelines [9,12,13,19]; Principle 3 rests on extrapolation warnings from Bailey [2] and Johnson [19]; Principle 4 rests on the NFL principle and scope-of-testing arguments from Johnson [19] and Hooker [18]. The only quantitative step is the illustrative expected time-to-solution example (pi_X=0.6 with T=1s vs. pi_Y=0.1 with T=1ms), which is the standard T/pi computation from the TTS literature, not a fitted parameter renamed as a prediction. The author's self-citations [20]-[22] appear in a survey sentence ('A large literature has also developed around guidelines for empirical evaluation of algorithms... [11] (see also [12-23])') and in the beginner's reading list ('papers marked with * in the reference list are a good place to start', which includes [22]); none of the four principles is justified by these self-citations, and deleting them would not change the argument, so they are minor and non-load-bearing. The paper also explicitly limits its own scope: 'the discussion herein should not be interpreted as critiquing experimental methodology per se,' it cautions 'Please remember, however, that correlation is not causation,' and its concluding recommendations are audience-conditional ('If your paper is not intended to be read as a commercial benchmarking study, choose your words carefully'). The skeptic's objection that the paper never defends the masses' interpretation of 'performance' over a physics-style success-probability definition is an unsettled normative premise, not a circular reduction, and belongs under correctness risk rather than circularity. Verdict: no significant circularity; score 1 reflects the minor non-load-bearing self-citations.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central argument rests on normative premises about how performance claims should be evaluated, plus the NFL principle and an assumption that selected examples are representative. There are no fitted parameters or invented entities.

assumptions (4)
  • domain assumption Classical fair-test benchmarking standards should be the yardstick for quantum performance claims.
    This normative premise underpins all four principles; the author acknowledges that physics-style studies have different goals but asserts the covenant with users should apply when performance comparisons reach public audiences (Introduction, page 4).
  • domain assumption The no free lunch principle applies to instantiated heuristic solvers, so benchmark outcomes mainly reflect parameter and input agreement.
    Invoked in Principle 2 (page 6) to argue that tuning policies must be disclosed; based on Wolpert and Macready [29].
  • ad hoc to paper The anonymized examples [R1]-[R4] are representative of current practices in quantum performance benchmarking.
    The paper does not provide a systematic survey; it selects examples to illustrate each principle, which is acknowledged in the Introduction.
  • domain assumption Extrapolating runtime curves to untested sizes is unreliable for heuristic solvers.
    Used in Principle 3, supported by Johnson [19] and Bailey [2]; assumes no valid model of how work parameter W scales with problem size N.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How NOT to Fool the Masses When Giving Performance Results for Quantum Computers." pith.science (2026). https://pith.science/paper/XTUC6EJ5

@misc{pith2026241108860,
  author       = {Pith},
  title        = {Pith review of: How NOT to Fool the Masses When Giving Performance Results for Quantum Computers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XTUC6EJ5}},
  note         = {Machine review of arXiv:2411.08860}
}
read the original abstract

In 1991, David Bailey wrote an article describing techniques for overstating the performance of massively parallel computers. Intended as a lighthearted protest against the practice of inflating benchmark results in order to ``fool the masses" and boost sales, the paper sparked development of procedural standards that help benchmarkers avoid methodological errors leading to unjustified and misleading conclusions. Now that quantum computers are starting to realize their potential as viable alternatives to classical computers, we can see the mistakes of three decades ago being repeated by a new batch of researchers who are unfamiliar with this history and these standards. Inspired by Bailey's model, this paper presents four suggestions for newcomers to quantum performance benchmarking, about how not to do it. They are: (1) Don't claim superior performance without mentioning runtimes; (2) Don't report optimized results without mentioning the tuning time needed to optimize those results; (3) Don't claim faster runtimes for (or in comparison to) solvers running on imaginary platforms; and (4) No cherry-picking (without justification and qualification). Suggestions for improving current practice appear in the last section.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 45 canonical work pages

  1. [20]

    Catherine C. McGeoch. 1996. Feature article: Toward an expe rimental method for algorithm simulation, INFORMS Journal on Computing 8.1:1- 15. 15

  2. [22]

    *Catherine C. McGeoch. 2019. Principles and guidelines for quantum performance analysis, Quantum Technology and Optimization Problems (QTOP), LNCS 11413:36-48, Springer

  3. [1]

    David H. Bailey. 1991. Twelve ways to fool the masses when giving p erfor- mance results on parallel computers, Supercomputing Review 54-59. Re- trieved 2024 from davidhbailey.com/dhbpapers/twelve-ways.pdf

  4. [2]

    David H. Bailey. 1992. Misleading performance reporting in the sup ercom- puting field, Scientific Programming 1.2:141-151

  5. [3]

    John Markoff. 1991. Technology; Measuring How Fast Comput- ers Really Are, New York Times 3:14. Retrieved 2024 from www.nytimes.com/1991/09/22/business/technology-measuring-how-fast-computers-

  6. [4]

    *Jack Dongarra et al. 1987. Computer benchmarking: Paths and pit falls, IEEE Spectrum 24.7:38-43

  7. [5]

    Roger Hockney. 1996. The Science of Computer Benchmarking , So- ciety for Industrial and Applied Mathematics. Retrieved 2024 from epubs.siam.org

  8. [6]

    Torsten Hoefler and Roberto Belli. 2015. Scientific benchmarking of par- allel computing systems: Twelve ways to tell the masses when repor ting performance results, Proc. Int. Conf. High Performance Computing, Net- working, Storage and Analysis , 1-12. 14

Show all 47 references
  1. [7]

    Paul Fortier and Howard Mitchel. 2003. Computer Systems Performance Evaluation and Prediction , Elsevier

  2. [8]

    *Lizy Kurian John and Lieven Eeckhout. 2017. Performance Evaluation and Benchmarking , CRC Press

  3. [9]

    Mark Raasveldt et al. 2018. Fair benchmarking considered difficult : Com- mon pitfalls in database performance testing, Proc. of the Workshop on Testing Database Systems

  4. [10]

    Benchmark (computing). (n.d.). Wikipedia, Wikimedia Foundation. Re- trieved 2024 from en.wikipedia.org/wiki/

  5. [11]

    Crowder et al

    Harlan P. Crowder et al. 1978. Reporting computational exper iments in mathematical programming, Mathematical Programming, 15.1:316-329

  6. [12]

    *Richard S Barr et al. 1995. Designing and reporting on computationa l experiments with heuristic methods, Journal of Heuristics 1:9-32

  7. [13]

    Thomas Bartz-Beielstein and Mark Preuss. 2010. Chapter 2: T he future of experimental research, in Thomas Bartz-Beielstein et al. (eds), Experimen- tal Methods for the Analysis of Optimization Algorithms . Berlin:Springer

  8. [14]

    Thomas Bartz-Beielstein et al. 2020. Benchmarking in optimizatio n: Best practice and open issues. arXiv:2007.03488

  9. [15]

    Blackburn et al

    *Stephen M. Blackburn et al. 2016. The truth, the whole truth, and noth- ing but the truth: A pragmatic guide to assessing empirical evaluatio ns, ACM Transactions on Programming Languages and Systems (TOP LAS) 38.4.15:1-20

  10. [16]

    Paul R. Cohen. 1995. Empirical Methods for Artificial Intelligence . MIT Press

  11. [17]

    Gent et al

    Ian P. Gent et al. 1997. How Not To Do It , University of Leeds School of Computer Science Research Report Series, Report 97

  12. [18]

    John Hooker. 1995. Testing Heuristics: We have it all wrong, Journal of Heuristics 1:33-42

  13. [19]

    *David S. Johnson. 2001. A theoretician’s guide to the experimen- tal analysis of algorithms, in M. H. Goldwasser et al. (eds), Data Structures, Near Neighbor Searches, and Methodology: Fift h and Sixth DIMACS Implementation Challenges 59, AMS. Retrieved from dimacs.rutgers.ed...

  14. [21]

    Catherine C. McGeoch. 2012. A Guide to Experimental Algorithmics , Cambridge Press

  15. [23]

    Rardin and Reha Uzsoy

    Ronald L. Rardin and Reha Uzsoy. 2001. Experimental evaluatio n of heuristic optimization algorithms: A tutorial, Journal of Heuristics 7:261- 304

  16. [24]

    Hans Mittelmann. 2020. Benchmarking optimization software – a (hi)story, SN Operations Research Forum 1.2. See also the slide deck presented at EURO 2019 , retrieved 2024 from plato.asu.edu/talks/euro2019.pdf

  17. [25]

    November 7, 2018

    Announcement by Gurobi. November 7, 2018. Retrieved 2024 f rom plato.asu.edu/ftp/apology.pdf

  18. [26]

    —. 2017. SPEC CPU ™ 2017 Run and Reporting Rules, SPEC™ Open Systems Group . Retrieved 2024 from spec.org/cpu2017/Docs/runrules.html

  19. [27]

    Jack Dongarra et al. 2023. The LINPACK benchmark: Past, pr esent, and future, Concurrency and Computation 15.9:803-220

  20. [28]

    Timothy Proctor et al. 2024. Benchmarking quantum computer s, arXiv:2407.088281

  21. [29]

    Wolpert and William G

    David H. Wolpert and William G. Macready. 1997. No free lunch theo rems for optimization, IEEE Transactions on evolutionary computation. 1.1:67- 68

  22. [30]

    James McDermott. 2019. When and why metaheuristics resear chers can ignore NFL (Section 4). arXiv:1906.03280

  23. [31]

    Frank Hutter et al. 2006. Performance prediction and automa ted tuning of randomized and parametric algorithms: An initial investigation, Proc. 12th Int. Conf. on Principles and Practices of Constraint Pr ogramming (CP-06) 213-222

  24. [32]

    Andreas Beham et al. 2018. Algorithm selection on generalized qu adratic assignment problem landscapes, Proc. of the Genetic and Evolutionary Computation Conference 253-260

  25. [33]

    Lennart Bittel and Martin Kliesch. 2021. Training variational qu antum algorithms is NP-hard, Physical Review Letters 127.120502. 16

  26. [34]

    Elijah Pelofske et al. 2024. Short-depth QAOA circuits and quan tum an- nealing on higher-order Ising models, npj Quantum Inf 10.30

  27. [35]

    Roennow et al

    Troels F. Roennow et al. 2014. Defining and detecting quantum s peedup, Science 345.6195:420-424

  28. [36]

    Wolfgang Panny. 2010. Deletions in random binary search trees : A story of errors, Journal of Statistical Planning and Inference 140.8:2335-2345

  29. [37]

    Derek Atkins et al. 1994. The magic words are squeamish ossifra ge, Proc. 4th Int. Conf. on Theory and Applications of Cryptology

  30. [38]

    2024, MaxSAT Evaluation 2024

    Jeremias Berg et al. 2024, MaxSAT Evaluation 2024 . Retrieved from maxsat-evaluations.github.io/2024

  31. [39]

    Iain Dunning et al. 2018. What works best when? A systematic ev aluation of heuristics for Max-Cut and QUBO, INFORMS Journal on Computing ,

  32. [40]

    Hans Mittelmann. (n.d.). Decision Tree for Optimization Software . Re- trieved 2024 from plato.asu.edu/guide.html

  33. [41]

    Gerhard Reinhelt. (n.d.). TSPLIB, Ruprecht-Karls-Universit¨ at Heidel- berg. Retrieved 2024 from comopt.ifi.uni-heidelberg.de

  34. [42]

    Hoos and Thomas St¨ utzle

    Holger H. Hoos and Thomas St¨ utzle. 2000. SATLIB: An online re search resource for research on SAT, in I.P. Gent et al., editors, SAT 2000, IOS Press. SATLIB files are available from www.satlib.org

  35. [43]

    George Santayana. 1917. The Life of Reason , Scribner’s Sons

  36. [44]

    Stuart Rosenberg. 1967. Cool Hand Luke , Warner Bros/Seven Arts

  37. [45]

    Hank Dietz. 1996. Blog post: Is parallel processing dead? Retr ieved 2024 from aggregate.ee.engr.uky.edu/Opinions/pardead.html

  38. [46]

    B. Furht. 1994. Parallel computing: Glory and collapse, Computer 27.4:74- 75. 17

  39. [2018]

    See also github.com/MQLIB

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.