{"id":"7ac84e59-250f-4d4a-aafd-ef7eb3b460d4","arxiv_id":"2411.08860","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A former D-Wave scientist adapts David Bailey's 1991 benchmarking warnings to quantum computing and proposes four principles for honest performance reporting.","lead":"This paper lays out four rules for honestly reporting quantum computer performance results, drawn from lessons learned in classical high-performance computing benchmarking. It argues that many recent quantum papers omit runtimes, hide tuning time, extrapolate to imaginary hardware, or cherry-pick examples, and that these habits breed distrust.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never defends why the masses' interpretation of 'performance' should override the stated goals of physics-style studies; the four principles' normative force depends on that unsettled premise.","rationale":"The reader's weakest assumption was that the paper assumes classical fair-test benchmarking standards should govern public-facing quantum performance claims, without defending why that perspective should override physics-style authors' own definitions of performance. I agree this is the most load-bearing concern: the four principles are stated categorically, but their normative force depends entirely on fixing the intended audience and the meaning of 'performance.' If an author explicitly disclaims benchmarking intent and defines performance as success probability at a fixed or unspecified time, then omitting runtime is not a logical fallacy; it is a scope restriction. The paper itself acknowledges this in the concluding suggestions, advising authors who are not writing benchmarking reports to avoid giving that impression. This self-acknowledgment indicates the principles are conditional on audience and intent, yet the main text presents them as unconditional norms. The reader also noted the evidentiary issue of anonymized examples; that is secondary because even a single valid example of each pitfall illustrates the logical point, though prevalence claims would need a systematic survey. The paper is well-reasoned and offers sound practical advice, and the conditionality is easy to repair by adding an explicit scope condition to each principle. Thus the verdict of CONDITIONAL is appropriate, and my read does not change it. I would ask the author to state more directly that the principles apply to claims intended or likely to be read as efficiency comparisons, and to justify why the masses' interpretation should be accommodated in papers aimed at researchers.","tokens_in":10851,"tokens_out":8033,"duration_ms":80250,"concrete_test":"Construct a miniature 'physics-style' performance comparison in the style of [R1] that explicitly states: 'This is a study of algorithmic error rates, not a benchmarking study. All solvers are executed for equal wall-clock time; performance is defined as success probability.' Omit runtimes from the abstract, as [R1] does, and present the claim 'X outperforms Y.' Then run a survey of quantum computing researchers (including the author's target audience) asking whether the claim is misleading or methodologically unsound. If a substantial fraction accepts it as a legitimate, well-defined claim despite the omitted runtime, Principle 1 is overbroad and must be qualified to apply only to statements intended as efficiency comparisons.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that a comparative performance statement like 'X outperforms Y' is unsupported unless runtimes (and tuning times, and platform realism) are disclosed. That claim presupposes that 'performance' is, or should be, understood as an efficiency measure involving computation time. The author explicitly concedes that 'performance means different things to different stakeholders,' and that many quantum performance studies are physics-style investigations of noise and error behavior, not benchmarking reports aimed at end users. For such a study, success probability may be the complete, intended measure of output quality, with time deliberately excluded because the research question is about fidelity, not speed. The paper never supplies an argument for why the benchmark-oriented interpretation of 'outperforms' should take precedence over the authors' own stated definitions, nor why potential misreading by the masses imposes an obligation on researchers whose intended audience is other researchers. The result is that Principles 1 and 2 (and to some extent 3 and 4) are presented as categorical 'Don'ts,' but the manuscript's own concluding suggestions soften them to 'if your paper is not intended to be read as a commercial benchmarking study, choose your words carefully.' Without a defense of the normative premise, the principles reduce to conditional communication advice. This is load-bearing because if a physics-study author explicitly defines performance as success probability at fixed or unreported wall-clock time, omitting runtime does not logically corrupt the claim; it merely changes the meaning of 'outperforms.' The paper's evidence for the practices being widespread is also anecdotal and anonymized, but the normative gap is the deeper issue because it affects the generality of all four principles.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that quantum computing performance claims, especially comparative statements such as \"X outperforms Y,\" should be held to the fair-test standards developed in classical benchmarking: report runtimes, disclose tuning time, avoid presenting extrapolations to nonexistent hardware as empirical results, and justify and qualify choices of test instances and comparison solvers. It illustrates each of these four principles with anonymized examples from the gate-model and quantum-annealing literature, discusses how disciplinary differences in the meaning of \"performance\" lead to miscommunication, and concludes with advice for researchers, journal editors, reviewers, and the public.","tokens_in":11015,"tokens_out":6256,"duration_ms":62270,"significance":"If read conditionally, the paper offers a useful and accessible checklist that could improve clarity, reproducibility, and trust in quantum performance comparisons. Its main strength is that it brings an established body of classical benchmarking methodology (Bailey, Johnson, and others) to a quantum audience that may be unfamiliar with it, and it does so with explicit caveats that physics-style studies are not being criticized per se. The anonymization of examples is a deliberate and mostly successful way to avoid ad hominem attacks, and the self-citations [20]–[22] are not load-bearing because the four principles are independently supported by other references. However, the paper's categorical framing overreaches its own caveats, and the prevalence claim in the abstract is not supported by the evidence presented, so the central claim needs revision rather than minor polishing.","major_comments":[{"comment":"The four principles are stated as categorical 'Don'ts,' but the manuscript itself concedes that 'performance means different things to different groups' and that research papers about quantum performance are 'normally aimed at other researchers, not at the masses.' For a physics-style study of noise and error behavior, success probability may be the complete intended measure of output quality, with computation time deliberately excluded because the research question concerns fidelity and error rates, not speed. The paper never defends why the benchmarking/efficiency interpretation of 'outperforms' should take precedence over the authors' own stated definitions, nor why potential misreading by nonexpert audiences imposes an obligation on researchers writing for other researchers. Since the final section says 'If your paper is not intended to be read as a commercial benchmarking study, choose your words carefully,' the practical version of the principles is conditional, not categorical. The categorical version needs either a defended normative premise or an explicit reframing as audience-relative communication advice.","section":"Four Principles (Principles 1–2) and Some Modest Suggestions"},{"comment":"The abstract claims that 'the mistakes of three decades ago are being repeated by a new batch of researchers' and that a 'new generation' is committing these errors. This prevalence claim is not established by the four anonymized examples [R1]–[R4], which are presented as illustrations rather than as a systematic survey. The paper later qualifies itself: 'the discussion herein should not be interpreted as critiquing experimental methodology per se, but rather as illustrating how best practice in one field can be seen as biased and misleading practice in another.' These two statements are in tension. The abstract should either be softened to a claim about possible or illustrative miscommunication, or supported with systematic evidence about how common the practices are in the current literature.","section":"Abstract and Introduction"},{"comment":"The principle 'Don't claim faster runtimes for (or in comparison to) solvers running on imaginary platforms' is broadly sound as a warning about extrapolating TTS curves to nonexistent hardware. However, the section's first paragraph acknowledges that studying asymptotic performance on abstract models of computation is standard and valuable. As written, the principle could be read as condemning all extrapolation, including legitimate theoretical analysis. The paper should clarify that the prohibition concerns reporting estimated wall-clock times on devices that do not exist as if those times were empirical measurements, not asymptotic analysis on abstract models per se. This distinction is important because one of the paper's own examples [R3] involves literal extrapolated runtimes of 10^19 seconds, which is a different practice from theoretical scaling analysis.","section":"Principle 3"}],"minor_comments":[{"comment":"In the Pareto-frontier discussion, the text says 'assuming lower is better in both dimensions,' but the same section uses success probability π, for which higher is better, and solution quality S, which is also normally higher-is-better. Please reconcile the notation, for example by defining S as an error rate or by explicitly stating 'lower is better for T and higher is better for S.'","section":"Four Principles, Principle 1"},{"comment":"The word 'inadvertantly' in the sentence about selecting inputs is a typo and should be 'inadvertently.'","section":"Principle 4"},{"comment":"Reference [41] lists 'Gerhard Reinhelt,' but the TSPLIB author is Gerhard Reinelt; please correct the spelling.","section":"References"},{"comment":"Several reference entries have truncated URLs in the text I reviewed (for example, reference [3] ends in '...computers-'); please ensure that complete URLs or DOIs are provided in the final version.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a perspective/opinion piece rather than a technical research paper. It is likely to be of interest to a broad quantum computing audience, especially if the journal publishes commentary on research practice. The anonymized examples are probably identifiable to community members; the decision to anonymize is defensible but may make independent verification of the factual claims difficult. The main revisions needed are a conditional reframing of the four principles and a softening of the prevalence claim in the abstract; both are feasible within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a concise, well-written position paper that imports David Bailey's 1991 rules for honest benchmarking into quantum computing. The four principles (report runtimes, report tuning time, don't extrapolate to imaginary hardware, don't cherry-pick without justification) are not new, but the packaging with anonymized cautionary examples from gate-model and annealing literature is a useful service to the community. McGeoch writes clearly and is fair: she stresses that the problematic papers were probably not intentionally deceptive, and that the issue is cross-disciplinary semantic discord.\n\nThe references to classical benchmarking standards are solid and appropriately credited. The suggestion section is practical, especially the advice to authors to define 'performance' explicitly and to state when a paper is not a benchmarking study.\n\nThe main weakness is evidentiary. The claim that 'the mistakes of three decades ago are being repeated by a new generation' is supported only by a handful of anonymized cases chosen without any systematic survey. That makes the prevalence claim unverifiable. The stress-test worry about the normative premise is worth noting but not fatal: the paper does not fully defend why the masses' interpretation of 'outperforms' should override a physics-style study's own definition. However, the concluding suggestions soften this by framing the principles as conditional communication advice ('if your paper is not intended to be read as a commercial benchmarking study, choose your words carefully'). So the categorical 'Don'ts' are less strong than they appear, and the underlying recommendation is sensible: authors should anticipate how their claims will be read.\n\nWho is this for? Quantum performance researchers, especially those with industry affiliations, and editors/reviewers who handle benchmarking papers. It would make a good invited commentary or perspective piece. It deserves peer review—not because it breaks new technical ground, but because it addresses an important communication problem and does so honestly. I'd be happy to see it published with a more explicit acknowledgment that the examples are illustrative, not a representative sample.","headline":"A sensible, well-written restatement of classical benchmarking ethics for quantum performance papers, weakened by an unproven prevalence claim and a soft normative premise, but worth publishing as a perspective piece.","tokens_in":11602,"tokens_out":2760,"would_cite":true,"duration_ms":26948,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Quantum performance claims should meet the fair-test standards that reformed parallel-computing benchmarks.","keywords":["quantum benchmarking","performance claims","runtime reporting","tuning time","fair testing","cherry-picking","quantum speedup","no-free-lunch principle"],"falsifier":"A systematic survey of recent quantum performance papers that recorded whether runtimes, tuning times, and platform statuses were disclosed would settle the urgency claim; if omission was rare, or if re-analysis found that adding the missing data never changed any reported outperformance conclusion, the paper's central claim that these practices mislead the masses would be refuted.","tokens_in":10578,"feed_emoji":"⚖️","tokens_out":8844,"duration_ms":71125,"temperature":0.7,"pith_summary":"This paper argues that quantum-computer performance comparisons are repeating the benchmark-reporting mistakes that once undermined trust in parallel computers. It turns a famous 1991 satirical list of ways to overstate performance into four positive rules for quantum researchers: report runtimes alongside solution quality, report the time spent tuning parameters, do not compare against solvers running on imagined platforms, and do not cherry-pick inputs or competitors without justification and qualification. The claim is that a public 'X outperforms Y' statement is not self-justifying: without these disclosures, readers cannot tell whether the result is genuine or an artifact of the test design. The paper locates the root cause in a clash between physics-style studies of hardware error and the expectations of commercial benchmarking, and it offers separate advice for researchers, editors, and the public.","feed_headline":"Quantum performance claims need runtimes, tuning time, and caveats","feed_subtitle":"A benchmarking veteran adapts classic fair-test rules to keep quantum claims honest.","key_machinery":"The argument runs through three load-bearing tools. The first is the two-dimensional outcome space of heuristic solvers: every solver is instantiated with a work parameter $W$ and produces a pair ($S$, $T$) of solution quality and computation time, so comparative claims must be judged on the Pareto frontier of those pairs rather than on solution quality alone. The second is the no-free-lunch principle from optimization, which says that benchmark outcomes are governed mainly by the agreement between input structure and solver parameters, making tuning policy a component of the result rather than an external detail. The third is the time-to-solution (TTS) metric used in quantum speedup studies, combining a measured time per core operation with the number of core operations needed to succeed; the paper warns that TTS curves measured on small devices are commonly extrapolated to imaginary platforms, producing runtimes that cannot be observed.","core_discovery":"The central claim is that a performance claim of the form 'solver X outperforms solver Y' in quantum computing is an empirical statement whose validity depends on four pieces of disclosure. Because heuristic solvers trade solution quality against computation time through a user-set work parameter, a comparison that reports only success probability is incomplete: without runtimes, the reader cannot know whether X won because it is genuinely better or because it was allowed more time. Because benchmark outcomes are governed by the match between input structure and instantiated solver parameters, a tuned solver compared with untuned rivals is not a fair race, and the time spent tuning must be disclosed. Because the time-to-solution metric measured on small devices mixes real measurements with arithmetic adjustments, extrapolating those curves to larger or future platforms is not evidence of comparative performance. Because input sets in quantum studies are necessarily small, cherry-picking is a hazard that must be handled by justification and explicit qualification of the scope of conclusions. The paper applies these principles to anonymized examples from gate-model and quantum-annealing literature, and in one recalculated example the corrected sample sizes reverse the claimed success-probability comparison.","pith_inferences":["Inference: the four principles could be turned into a one-page reporting checklist for quantum benchmarking papers, asking authors to state the work parameter setting, the tuning search cost, and the platform status for every solver.","Inference: a systematic audit of recent gate-model and annealing papers would likely show that tuning time is the most commonly omitted quantity, since many physics-style studies treat parameter sweeps as preprocessing rather than part of the measurement.","Inference: the paper's semantic-discord framing predicts that dispute resolution will come less from new benchmarks than from relabeling: physics-style studies should describe themselves as error-modeling investigations, reserving 'outperforms' for claims that include runtimes.","Inference: extending the extrapolation warning, even well-calibrated quantum error correction and fault tolerance could change runtime ratios faster than any asymptotic model predicts, so performance predictions for post-fault-tolerance machines should carry explicit uncertainty bounds."],"forward_implications":["A quantum paper that claims 'outperforms' but reports no runtimes is, on this view, unsupported; a reader can only treat the claim as an artifact of test design.","Tuning effort is part of the performance result, so a QAOA comparison that omits parameter-search time cannot be compared fairly with a solver using defaults.","Time-to-solution curves measured on small devices cannot be extrapolated to future hardware with any confidence; the paper cites cases where generational improvements changed speedups by four orders of magnitude.","Cherry-picking is unavoidable, but it is acceptable only when the authors justify input and solver selection and qualify the scope of their conclusions.","If the practices persist, the public audience may lose trust in quantum computing results, a dynamic the paper compares to the downsizing that followed parallel-computing benchmark scandals."],"supporting_citations":[{"why":"supplies the original satirical twelve-ways article that the paper adapts into its four principles.","marker":"[1]"},{"why":"extends the original warning to extrapolation and misleading performance reporting, used in Principle 3.","marker":"[2]"},{"why":"supplies the guiding quotation that claims predicated on missing facts are inappropriate.","marker":"[15]"},{"why":"provides the authoritative fair-testing and extrapolation guidance that the four principles rest on.","marker":"[19]"},{"why":"gives the no-free-lunch theorem that explains why parameter tuning determines benchmark outcomes.","marker":"[29]"},{"why":"shows when no-free-lunch reasoning can be relaxed, supporting the demand for justified scope.","marker":"[30]"},{"why":"proves that optimizing QAOA parameters is NP-hard, explaining why tuning time can dominate reported results.","marker":"[33]"},{"why":"is the third-party QAOA dataset whose sample-size misreading reverses the comparison in the paper's first example.","marker":"[34]"},{"why":"defines the time-to-solution and quantum speedup framework that Principle 3 warns against extrapolating.","marker":"[35]"}],"fun_headline_variants":["Quantum speedup claims need runtimes and tuning truth","Fair quantum benchmarks: disclose runtimes, tuning, limits","Old benchmark lessons apply to quantum: no cherry-picking","Quantum performance honesty: four rules from 1991","Don't claim quantum superiority without these disclosures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The advice only binds if quantum performance comparisons aimed at the public should be judged by classical fair-test benchmarking rules, and it only matters if the four problematic practices are actually common; the paper asserts both rather than proving them.","fun_headline_variants_meta":{"raw":{"variants":["Quantum speedup claims need runtimes and tuning truth","Fair quantum benchmarks: disclose runtimes, tuning, limits","Old benchmark lessons apply to quantum: no cherry-picking","Quantum performance honesty: four rules from 1991","Don't claim quantum superiority without these disclosures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1300,"prompt_tokens":965,"completion_tokens":335,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":581,"tokens_out":335,"duration_ms":143193,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:14:13.472399+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic survey of recent quantum performance papers that recorded whether runtimes, tuning times, and platform statuses were disclosed would settle the urgency claim; if omission was rare, or if re-analysis found that adding the missing data never changed any reported outperformance conclusion, the paper's central claim that these practices mislead the masses would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the third-party QAOA dataset whose sample-size misreading reverses the comparison in the paper's first example."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the original satirical twelve-ways article that the paper adapts into its four principles."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"extends the original warning to extrapolation and misleading performance reporting, used in Principle 3."},{"cited_title":"Blackburn et al","cited_arxiv_id":null,"evidence_quote":"supplies the guiding quotation that claims predicated on missing facts are inappropriate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the authoritative fair-testing and extrapolation guidance that the four principles rest on."},{"cited_title":"Wolpert and William G","cited_arxiv_id":null,"evidence_quote":"gives the no-free-lunch theorem that explains why parameter tuning determines benchmark outcomes."},{"cited_title":"When and Why Metaheuristics Researchers Can Ignore \"No Free Lunch\" Theorems","cited_arxiv_id":"1906.03280","evidence_quote":"shows when no-free-lunch reasoning can be relaxed, supporting the demand for justified scope."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"proves that optimizing QAOA parameters is NP-hard, explaining why tuning time can dominate reported results."},{"cited_title":"Roennow et al","cited_arxiv_id":null,"evidence_quote":"defines the time-to-solution and quantum speedup framework that Principle 3 warns against extrapolating."}],"review_version":1}