{"id":"1722acf4-3656-4eff-875d-3872ffb6e541","arxiv_id":"2509.03078","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 329 quantum benchmarking studies yields a stack-aligned taxonomy and definitions for hardware-, software-, and application-focused benchmarks.","lead":"This paper is a systematic literature review that organizes quantum computer benchmarking methods into a three-level taxonomy (hardware, software, and application focus) with definitions and a map of how benchmark types depend on each other. It aims to give researchers and industry a common vocabulary and a structured way to choose benchmarks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy category assignments are not shown to be reproducible; without inter-rater reliability, the 'common language' claim is unsupported.","rationale":"The reader's weakest_assumption correctly identified that taxonomy assignments may be unstable, and the paper itself admits borderline subjectivity in Sec. 7. My concern is load-bearing because the central claim is not merely the collection of benchmarks but the provision of a structured, stakeholder-aligned framework. If category assignments are not reproducible, the framework's organizing power and 'common language' promise are weakened. I agree with the reader's CONDITIONAL verdict: the issue is addressable by adding an inter-rater reliability study or by explicitly revising the taxonomy to separate stakeholder from methodological dimensions. The reader also flagged the 'most comprehensive' claim and reproducibility artifacts; those are secondary to the taxonomy's validity. My proposed test directly settles whether the ambiguity is material.","tokens_in":48964,"tokens_out":2734,"duration_ms":35151,"concrete_test":"Sample 50 studies from the 329-study corpus. Give three expert raters (not authors) the taxonomy definitions from Sec. 5 and ask each to independently assign every study to one top-level focus and one subcategory. Compute Fleiss' kappa for top-level assignments and record how many studies are considered ambiguous (raters split across categories). If kappa < 0.6 or more than 10% of studies are ambiguous, the taxonomy lacks the reliability needed to support the 'common language' claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a unified, stack-aligned taxonomy that provides a common language for QC benchmarking. That claim depends on the three top-level categories (Hardware Focus, Software Focus, Application Focus) being assignable without material ambiguity. The paper's own definitions blur the boundaries: Sec. 5.2 states that Software-Focus benchmarks are 'also relevant for hardware developers seeking to assess... hardware as a holistic system,' and Sec. 5.3 places Q-score and VQE ground-state benchmarks in Application Focus even though they are algorithm-driven (QAOA/VQE) and would fit the Software-Focus definition of 'evaluating... algorithms and frameworks.' The paper acknowledges in Sec. 7 that 'the categorization and interpretation of borderline cases inherently involve a degree of subjectivity,' but provides no reproducibility check (e.g., inter-rater agreement) to show the assignments are stable across experts. If independent raters applying the paper's own definitions disagree materially, the taxonomy fails as a shared vocabulary and the 'common language' contribution is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a systematic literature review of benchmarking for gate-based quantum computers. The authors searched ACM Digital Library, Google Scholar, and ProQuest (412 initial records, 329 studies after screening and forward-backward search), used BERTopic clustering plus manual refinement to organize the literature, and propose a three-level taxonomy — Hardware Focus, Software Focus, Application Focus — with subcategories such as tomography, randomized benchmarks, volumetric benchmarks, compiler benchmarks, ground-state energy calculations, quantum machine learning, and quantum optimization. They define benchmark characteristics, map interdependencies across stack layers, and identify research gaps. The paper claims to be the most comprehensive systematic review of QC benchmarking to date and to provide a common language for the field.","tokens_in":49158,"tokens_out":7108,"duration_ms":82779,"significance":"If the taxonomy is stable and reproducible, this would be a valuable reference: the review is more systematic than most prior narrative reviews, the technical descriptions of randomized benchmarking, XEB, GST, quantum volume, Q-score, and VQE benchmarks are largely faithful to the underlying literature, and the stakeholder/stack alignment gives practitioners a useful map. The paper also provides a transparent search protocol with inclusion/exclusion counts, a forward-backward search, and NLP clustering details in Appendix A, which strengthens confidence in coverage. However, the value of the central contribution depends on the boundary assignments being trustworthy; that is currently not demonstrated.","major_comments":[{"comment":"The central claim is that the taxonomy provides a 'common language' for QC benchmarking, but category assignments are not shown to be reproducible. Sec. 5.2 defines Software Focus as 'system-wide performance' and says it is 'also relevant for hardware developers seeking to assess and compare their hardware as a holistic system'; Sec. 5.3 places Q-score and VQE ground-state benchmarks in Application Focus even though these are algorithm-driven and could satisfy the Software Focus definition of 'evaluating... algorithms and frameworks.' Sec. 7 concedes that 'the categorization and interpretation of borderline cases inherently involve a degree of subjectivity,' but no inter-rater reliability measure (e.g., Cohen's kappa) or explicit decision rule is reported. Because a shared vocabulary depends on stable assignment of benchmarks to categories, this is load-bearing. Please add a coding proto","section":"Secs. 5.2–5.3 and Sec. 7"},{"comment":"The benchmark definition adopted from Acuaviva et al. restricts benchmarks to tests that evaluate 'a quantum processor or hardware component for a certain task.' The review then includes compiler benchmarks (Sec. 5.2.6) and dequantization benchmarks (Sec. 5.2.4), which evaluate software and algorithmic components rather than a quantum processor or hardware component. This makes the foundational definition inconsistent with the taxonomy's Software and Application Focus categories. Please either broaden the benchmark definition, or explicitly show how compiler and dequantization evaluations fit the adopted definition; otherwise the proposed 'standard terminology' is not internally coherent.","section":"Sec. 2.1 vs. Secs. 5.2.4/5.2.6"}],"minor_comments":[{"comment":"The quantum volume equation is miswritten: it uses 'arg max' where the expression should be a maximum. The correct form is log2 VQ = max_m min(m, d(m)), and d(m) denotes the achievable circuit depth for width m, not 'the number of qubits in the largest square circuit' as stated in the text. Please correct the notation.","section":"Sec. 5.2.1, Eq. (11)"},{"comment":"The claim of 'the most comprehensive systematic literature review to date' is not substantiated by a quantitative or qualitative comparison with the cited prior reviews [12,37,38] on search coverage, time span, inclusion criteria, or number of included studies. Please provide such a comparison or soften the claim to avoid an unsupported superlative.","section":"Abstract and Sec. 3, final paragraph"},{"comment":"The final search string requires 'benchmark*' in the abstract and may miss relevant papers that use terms such as 'characterization,' 'evaluation,' or 'validation' without explicitly mentioning benchmarks. The forward-backward search mitigates this, but a sentence acknowledging the residual limitation would strengthen the methodology discussion.","section":"Sec. 4.2"},{"comment":"The phrase 'double-blind setting' is ambiguous: the text describes two reviewers independently applying inclusion/exclusion criteria, which is not the usual meaning of double-blind review. Consider using 'dual independent screening' to avoid confusion.","section":"Sec. 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the self-citation [197] is used only as an example, which is acceptable. The main risk is that the taxonomy's reproducibility, not its technical accuracy, determines whether the 'common language' contribution is convincing. The overclaim about comprehensiveness may draw reviewer criticism and should be toned down."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, useful systematic review of gate-based QC benchmarking, and the taxonomy it proposes is the main contribution. The paper is not a new measurement or algorithm; it is organizational work, and it is good organizational work. If you work in benchmarking, the three-level hardware/software/application split with subcategories like dequantization, error-correction, and compiler benchmarks is a genuinely useful way to navigate the literature. The technical descriptions of the benchmarks (T1/T2, RB, XEB, Q-score, tomography variants) are accurate in the places I checked. The search process is described in detail, and the authors are transparent about their screening decisions.\n\nThe soft spots are real but manageable. First, the 'most comprehensive systematic literature review to date' claim is not supported by evidence. The authors report 329 studies considered and 273 cited, but they do not provide the study list, the search protocol details, or the NLP clustering code and data. For a systematic review, that is a reproducibility gap. A reader cannot verify completeness or re-run the analysis. This should be fixed before publication.\n\nSecond, the taxonomy's boundary conditions are acknowledged to be subjective. The stress-test concern about inter-rater reliability is fair: the definitions of Software Focus and Application Focus overlap for algorithm-driven benchmarks like Q-score and VQE. That said, I don't think this is fatal. The top-level categories are broad, and the paper's own examples are defensible. Q-score is explicitly a full-stack application benchmark; VQE ground-state energy is an application problem. The subjectivity is in the borderline cases, and the paper admits it. A reliable taxonomy would benefit from an inter-rater agreement study, but its absence does not sink the contribution. It just means the 'common language' claim should be tempered to 'a proposed common language.'\n\nCitation pattern looks fine. The single self-citation [197] is used as an example of an optimization workflow, not as a load-bearing source.\n\nBottom line: this paper deserves a serious referee. It is a useful reference for anyone entering QC benchmarking or wanting a structured overview. I would accept it for peer review with requests for the study list, clustering artifacts, and a more modest claim about comprehensiveness. The core taxonomy is sound and likely to be adopted.","headline":"A well-executed systematic review whose taxonomy is the real contribution; fix the reproducibility gaps and temper the 'most comprehensive' claim.","tokens_in":49650,"tokens_out":2542,"would_cite":true,"duration_ms":27680,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review sorts the whole field of gate-based quantum-computing benchmarks—329 studies—into one stack-aligned taxonomy, and claims the result is a shared language for comparing quantum systems.","keywords":["quantum computing","benchmarking","systematic literature review","taxonomy","quantum stack","hardware focus","software focus","application focus"],"falsifier":"Have two independent panels of QC researchers apply the taxonomy's definitions to a random sample of, say, 100 benchmark papers from the 329 in the review and measure inter-rater agreement. If agreement on the top-level hardware/software/application assignment is at or below chance (e.g., Cohen's kappa below 0.6), the taxonomy fails to provide the unambiguous common language it claims. Alternatively, a single well-known gate-based benchmark that the taxonomy cannot place under any subcategory would falsify its completeness.","tokens_in":48872,"feed_emoji":"⚛️","tokens_out":3945,"duration_ms":46468,"temperature":0.7,"pith_summary":"The paper argues that quantum-computing benchmarking is fragmented: many protocols exist, few definitions agree, and no single benchmark covers everything. It claims to be the most comprehensive systematic literature review of gate-based QC benchmarking to date—329 studies—and to solve the fragmentation by introducing a three-part taxonomy: hardware focus, software focus, and application focus, aligned with the quantum stack and with the people who use benchmarks. If the taxonomy is right, benchmark designers, hardware vendors, and application users get a common vocabulary, a map of what exists, and a clear view of where the gaps are. The review also claims that benchmark categories are interdependent, so progress claims should be checked across the whole stack rather than on a single metric.","feed_headline":"One taxonomy maps all gate-based quantum benchmarks","feed_subtitle":"Hardware, software, and application categories give quantum teams a shared language for fair comparison.","key_machinery":"The taxonomy itself is the central object. It is built by combining BERTopic-based natural-language clustering of 329 papers with expert refinement, then aligning categories with the layers of the quantum stack (application, algorithm, programming language, compiler/runtime, instruction set, microarchitecture, quantum-classical interface, quantum chip). Each benchmark is assigned first by intended stakeholder and stack layer, then by methodological principle; the definitions and the mapping of interdependencies are what carry the argument. The adopted benchmark definition from ref. [12] sets the boundary: a benchmark is a test measuring performance of a quantum processor or hardware componen","core_discovery":"On its own terms, the paper's central claim is that the landscape of gate-based quantum-computer benchmarking can be organized into a single hierarchical taxonomy: three primary categories—hardware focus, software focus, application focus—that mirror the lower, middle, and upper layers of the quantum stack, each subdivided by methodological principle (component-level metrics, tomography, randomized benchmarks, volumetric benchmarks, algorithm-based benchmarks, error-correction benchmarks, ground-state energy calculations, quantum machine learning, quantum optimization). It claims these categories come with precise definitions, that interdependencies between them (gate fidelity to algorithm s","pith_inferences":["The taxonomy's usefulness depends on how reliably independent experts assign borderline benchmarks to the same category; an inter-rater reliability test on a sample of the 329 papers would make the claim to a common language checkable.","If the taxonomy were applied to non-gate-based platforms such as quantum annealers or analog simulators, the categories might need new top-level entries, suggesting the three-category structure is tied to the gate-based model.","The paper's own interdependence analysis implies that single-number vendor metrics like quantum volume or Q-score should be read as partial views; a device that scores high on one category may still fail on application benchmarks, and that is informative, not contradictory.","A natural extension is a living repository that continuously re-classifies new benchmarks; if new protocols repeatedly straddle the existing subcategories, the taxonomy would need revision, which the paper already anticipates."],"forward_implications":["A common vocabulary lets competing benchmark papers be compared on what they actually measure rather than on their names.","Hardware vendors and researchers can see which benchmark categories serve which decisions, and which categories are missing from current practice.","The interdependence map implies that improving one layer, such as gate fidelity, has predictable effects upward, so progress reports can be checked for stack-wide consistency.","Fairer evaluation becomes possible because categories and definitions expose hidden choices like compiler sensitivity, not just device quality.","Research gaps become explicit: composite cross-level benchmarks that link low-level metrics to application outcomes are identified as an open direction."],"supporting_citations":[{"why":"Supplies the BERTopic NLP clustering method used to create the initial benchmark groupings that the taxonomy refines.","marker":"[11]"},{"why":"Provides the benchmark definition the review adopts, and a recent framework that motivates the push for standardized terminology.","marker":"[12]"},{"why":"Prior three-class categorization (qubit, computer, fault-tolerant benchmarks) that the new taxonomy extends and contrasts with.","marker":"[9]"},{"why":"Supplies benchmark design principles and a categorization of benchmarks that this review builds on and systematizes.","marker":"[8]"},{"why":"Example of an existing benchmark suite with a criticized categorization, motivating the need for clearer definitions.","marker":"[16]"},{"why":"Comprehensive QCVV overview used as a comparison point and source of key benchmarking techniques.","marker":"[42]"},{"why":"Basis for the quantum stack layer model used to align the taxonomy with the stack.","marker":"[17]"}],"fun_headline_variants":["One taxonomy unifies all gate-based quantum benchmarks","Systematic review maps every quantum benchmark type","Hierarchical taxonomy covers all QC benchmarking","Quantum stack benchmarks now have one framework","Comprehensive taxonomy for quantum computer benchmarks"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The taxonomy assumes every gate-based QC benchmark can be placed in exactly one of the three focus categories using primary stakeholder and stack layer, with borderline assignments being rare enough not to undermine the classification; the paper itself concedes that these assignments carry a degree of subjectivity.","fun_headline_variants_meta":{"raw":{"variants":["One taxonomy unifies all gate-based quantum benchmarks","Systematic review maps every quantum benchmark type","Hierarchical taxonomy covers all QC benchmarking","Quantum stack benchmarks now have one framework","Comprehensive taxonomy for quantum computer benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1126,"prompt_tokens":671,"completion_tokens":455,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":415,"completion_tokens_details":{"reasoning_tokens":391}},"tokens_in":415,"tokens_out":455,"duration_ms":5193,"temperature":1.0,"reasoning_tokens":391,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:05:33.442495+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent panels of QC researchers apply the taxonomy's definitions to a random sample of, say, 100 benchmark papers from the 329 in the review and measure inter-rater agreement. If agreement on the top-level hardware/software/application assignment is at or below chance (e.g., Cohen's kappa below 0.6), the taxonomy fails to provide the unambiguous common language it claims. Alternatively, a single well-known gate-based benchmark that the taxonomy cannot place under any subcategory would falsify its completeness.","supporting_citations":[],"review_version":1}