Pith. sign in

REVIEW 2 major objections 4 minor 3 cited by

This review sorts the whole field of gate-based quantum-computing benchmarks—329 studies—into one stack-aligned taxonomy, and claims the result is a shared language for comparing quantum systems.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-05 11:05 UTC pith:OSKBGEI4

load-bearing objection A well-executed systematic review whose taxonomy is the real contribution; fix the reproducibility gaps and temper the 'most comprehensive' claim. the 2 major comments →

arxiv 2509.03078 v1 pith:OSKBGEI4 submitted 2025-09-03 quant-ph

Quantum Computer Benchmarking: An Explorative Systematic Literature Review

classification quant-ph
keywords quantum computingbenchmarkingsystematic literature reviewtaxonomyquantum stackhardware focussoftware focusapplication focus
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that quantum-computing benchmarking is fragmented: many protocols exist, few definitions agree, and no single benchmark covers everything. It claims to be the most comprehensive systematic literature review of gate-based QC benchmarking to date—329 studies—and to solve the fragmentation by introducing a three-part taxonomy: hardware focus, software focus, and application focus, aligned with the quantum stack and with the people who use benchmarks. If the taxonomy is right, benchmark designers, hardware vendors, and application users get a common vocabulary, a map of what exists, and a clear view of where the gaps are. The review also claims that benchmark categories are interdependent, so progress claims should be checked across the whole stack rather than on a single metric.

Core claim

On its own terms, the paper's central claim is that the landscape of gate-based quantum-computer benchmarking can be organized into a single hierarchical taxonomy: three primary categories—hardware focus, software focus, application focus—that mirror the lower, middle, and upper layers of the quantum stack, each subdivided by methodological principle (component-level metrics, tomography, randomized benchmarks, volumetric benchmarks, algorithm-based benchmarks, error-correction benchmarks, ground-state energy calculations, quantum machine learning, quantum optimization). It claims these categories come with precise definitions, that interdependencies between them (gate fidelity to algorithm s

What carries the argument

The taxonomy itself is the central object. It is built by combining BERTopic-based natural-language clustering of 329 papers with expert refinement, then aligning categories with the layers of the quantum stack (application, algorithm, programming language, compiler/runtime, instruction set, microarchitecture, quantum-classical interface, quantum chip). Each benchmark is assigned first by intended stakeholder and stack layer, then by methodological principle; the definitions and the mapping of interdependencies are what carry the argument. The adopted benchmark definition from ref. [12] sets the boundary: a benchmark is a test measuring performance of a quantum processor or hardware componen

Load-bearing premise

The taxonomy assumes every gate-based QC benchmark can be placed in exactly one of the three focus categories using primary stakeholder and stack layer, with borderline assignments being rare enough not to undermine the classification; the paper itself concedes that these assignments carry a degree of subjectivity.

What would settle it

Have two independent panels of QC researchers apply the taxonomy's definitions to a random sample of, say, 100 benchmark papers from the 329 in the review and measure inter-rater agreement. If agreement on the top-level hardware/software/application assignment is at or below chance (e.g., Cohen's kappa below 0.6), the taxonomy fails to provide the unambiguous common language it claims. Alternatively, a single well-known gate-based benchmark that the taxonomy cannot place under any subcategory would falsify its completeness.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • A common vocabulary lets competing benchmark papers be compared on what they actually measure rather than on their names.
  • Hardware vendors and researchers can see which benchmark categories serve which decisions, and which categories are missing from current practice.
  • The interdependence map implies that improving one layer, such as gate fidelity, has predictable effects upward, so progress reports can be checked for stack-wide consistency.
  • Fairer evaluation becomes possible because categories and definitions expose hidden choices like compiler sensitivity, not just device quality.
  • Research gaps become explicit: composite cross-level benchmarks that link low-level metrics to application outcomes are identified as an open direction.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy's usefulness depends on how reliably independent experts assign borderline benchmarks to the same category; an inter-rater reliability test on a sample of the 329 papers would make the claim to a common language checkable.
  • If the taxonomy were applied to non-gate-based platforms such as quantum annealers or analog simulators, the categories might need new top-level entries, suggesting the three-category structure is tied to the gate-based model.
  • The paper's own interdependence analysis implies that single-number vendor metrics like quantum volume or Q-score should be read as partial views; a device that scores high on one category may still fail on application benchmarks, and that is informative, not contradictory.
  • A natural extension is a living repository that continuously re-classifies new benchmarks; if new protocols repeatedly straddle the existing subcategories, the taxonomy would need revision, which the paper already anticipates.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper presents a systematic literature review of benchmarking for gate-based quantum computers. The authors searched ACM Digital Library, Google Scholar, and ProQuest (412 initial records, 329 studies after screening and forward-backward search), used BERTopic clustering plus manual refinement to organize the literature, and propose a three-level taxonomy — Hardware Focus, Software Focus, Application Focus — with subcategories such as tomography, randomized benchmarks, volumetric benchmarks, compiler benchmarks, ground-state energy calculations, quantum machine learning, and quantum optimization. They define benchmark characteristics, map interdependencies across stack layers, and identify research gaps. The paper claims to be the most comprehensive systematic review of QC benchmarking to date and to provide a common language for the field.

Significance. If the taxonomy is stable and reproducible, this would be a valuable reference: the review is more systematic than most prior narrative reviews, the technical descriptions of randomized benchmarking, XEB, GST, quantum volume, Q-score, and VQE benchmarks are largely faithful to the underlying literature, and the stakeholder/stack alignment gives practitioners a useful map. The paper also provides a transparent search protocol with inclusion/exclusion counts, a forward-backward search, and NLP clustering details in Appendix A, which strengthens confidence in coverage. However, the value of the central contribution depends on the boundary assignments being trustworthy; that is currently not demonstrated.

major comments (2)
  1. [Secs. 5.2–5.3 and Sec. 7] The central claim is that the taxonomy provides a 'common language' for QC benchmarking, but category assignments are not shown to be reproducible. Sec. 5.2 defines Software Focus as 'system-wide performance' and says it is 'also relevant for hardware developers seeking to assess and compare their hardware as a holistic system'; Sec. 5.3 places Q-score and VQE ground-state benchmarks in Application Focus even though these are algorithm-driven and could satisfy the Software Focus definition of 'evaluating... algorithms and frameworks.' Sec. 7 concedes that 'the categorization and interpretation of borderline cases inherently involve a degree of subjectivity,' but no inter-rater reliability measure (e.g., Cohen's kappa) or explicit decision rule is reported. Because a shared vocabulary depends on stable assignment of benchmarks to categories, this is load-bearing. Please add a coding proto
  2. [Sec. 2.1 vs. Secs. 5.2.4/5.2.6] The benchmark definition adopted from Acuaviva et al. restricts benchmarks to tests that evaluate 'a quantum processor or hardware component for a certain task.' The review then includes compiler benchmarks (Sec. 5.2.6) and dequantization benchmarks (Sec. 5.2.4), which evaluate software and algorithmic components rather than a quantum processor or hardware component. This makes the foundational definition inconsistent with the taxonomy's Software and Application Focus categories. Please either broaden the benchmark definition, or explicitly show how compiler and dequantization evaluations fit the adopted definition; otherwise the proposed 'standard terminology' is not internally coherent.
minor comments (4)
  1. [Sec. 5.2.1, Eq. (11)] The quantum volume equation is miswritten: it uses 'arg max' where the expression should be a maximum. The correct form is log2 VQ = max_m min(m, d(m)), and d(m) denotes the achievable circuit depth for width m, not 'the number of qubits in the largest square circuit' as stated in the text. Please correct the notation.
  2. [Abstract and Sec. 3, final paragraph] The claim of 'the most comprehensive systematic literature review to date' is not substantiated by a quantitative or qualitative comparison with the cited prior reviews [12,37,38] on search coverage, time span, inclusion criteria, or number of included studies. Please provide such a comparison or soften the claim to avoid an unsupported superlative.
  3. [Sec. 4.2] The final search string requires 'benchmark*' in the abstract and may miss relevant papers that use terms such as 'characterization,' 'evaluation,' or 'validation' without explicitly mentioning benchmarks. The forward-backward search mitigates this, but a sentence acknowledging the residual limitation would strengthen the methodology discussion.
  4. [Sec. 4.3] The phrase 'double-blind setting' is ambiguous: the text describes two reviewers independently applying inclusion/exclusion criteria, which is not the usual meaning of double-blind review. Consider using 'dual independent screening' to avoid confusion.

Circularity Check

0 steps flagged

No circularity: literature review taxonomy is a synthesis, not a derivation; the single self-citation [197] is not load-bearing.

full rationale

This is a systematic literature review, so there is no mathematical derivation chain whose output could be equivalent to its input by construction. The taxonomy is produced from the reviewed corpus (141 initial studies plus forward/backward search) through BERTopic clustering and manual expert refinement, then organized into Hardware/Software/Application Focus categories aligned with the QC stack. That is a synthesizing and organizational activity, not a prediction from fitted parameters. The only self-citation, [197], appears in Sec. 5.3 as an example of 'practical QC workflows'; it is not used to justify the taxonomy, to define categories, or to force any classification, so it is not load-bearing. The definitions and equations in the paper (e.g., T1/T2 decays, fidelity formulas, quantum volume, Q-score) are standard results cited from the literature and are presented as descriptions, not as predictions derived from the taxonomy. The paper explicitly acknowledges in Sec. 7 that 'the categorization and interpretation of borderline cases inherently involve a degree of subjectivity,' which is a transparency statement about expert judgment, not evidence of circularity. No self-definitional step, fitted-input-called-prediction step, uniqueness-imported-from-authors step, ansatz-smuggled-in-via-citation step, or renaming-known-result step is present. The central claim is a framework built from external sources plus declared expert interpretation, so the derivation chain is self-contained in the sense appropriate to a literature review.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The review introduces no new free parameters or invented entities. Its load-bearing assumptions are the adopted benchmark definition, the QC stack as the organizing axis, the gate-based scope restriction, and the validity of the NLP clustering as a starting point for the taxonomy.

axioms (4)
  • domain assumption A QC benchmark is 'a test, or set of tests, that aims to measure or evaluate the performance, efficiency, or other properties of a quantum processor or hardware component for a certain task' (Acuaviva et al., Sec. 2.1).
    Adopted from an external source as the scope-defining definition of what counts as a benchmark for inclusion in the review.
  • domain assumption The QC stack, from application down to the quantum chip, is the correct organizing backbone for categorizing benchmarks.
    Used throughout Sec. 5 to assign benchmarks to hardware, software, and application focus categories. If the stack is not the natural organizing axis, the taxonomy's partitions lose their grounding.
  • domain assumption Only universal, gate-based QC is in scope; quantum annealers and analog simulators are excluded.
    Stated as an inclusion/exclusion criterion in Sec. 4.3 and acknowledged as a limitation in Sec. 7, but it shapes the claimed completeness of the review.
  • domain assumption The NLP-based BERTopic clustering provides a meaningful initial structure that can be refined by manual expert analysis.
    The clustering step is described in Appendix A, but the validity of the resulting clusters is not quantitatively assessed, and the taxonomy relies on the authors' manual refinement to correct any cluster misassignments.

pith-pipeline@v1.4.0-alltime-deepseek-medium · 48723 in / 8007 out tokens · 85752 ms · 2026-08-05T11:05:33.442495+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Quantum Computer Benchmarking: An Explorative Systematic Literature Review." pith.science (2026). https://pith.science/paper/OSKBGEI4

@misc{pith2026250903078,
  author       = {Pith},
  title        = {Pith review of: Quantum Computer Benchmarking: An Explorative Systematic Literature Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OSKBGEI4}},
  note         = {Machine review of arXiv:2509.03078}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As quantum computing (QC) continues to evolve in hardware and software, measuring progress in this complex and diverse field remains a challenge. To track progress, uncover bottlenecks, and evaluate community efforts, benchmarks play a crucial role. But which benchmarking approach best addresses the diverse perspectives of QC stakeholders? We conducted the most comprehensive systematic literature review of this area to date, combining NLP-based clustering with expert analysis to develop a novel taxonomy and definitions for QC benchmarking, aligned with the quantum stack and its stakeholders. In addition to organizing benchmarks in distinct hardware, software, and application focused categories, our taxonomy hierarchically classifies benchmarking protocols in clearly defined subcategories. We develop standard terminology and map the interdependencies of benchmark categories to create a holistic, unified picture of the quantum benchmarking landscape. Our analysis reveals recurring design patterns, exposes research gaps, and clarifies how benchmarking methods serve different stakeholders. By structuring the field and providing a common language, our work offers a foundation for coherent benchmark development, fairer evaluation, and stronger cross-disciplinary collaboration in QC.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. StabilizerBench: A Benchmark for AI-Assisted Quantum Error Correction Circuit Synthesis

    quant-ph 2026-04 conditional novelty 8.0

    StabilizerBench is a new benchmark for evaluating AI agents on generating, optimizing, and making fault-tolerant stabilizer circuits for quantum error correction, with efficient verification and multi-tier scoring.

  2. Auditing Empirical Comparisons in Quantum Software

    cs.SE 2026-07 unverdicted novelty 7.0

    CLAIMSTAB-QC audits 455 comparative claims from 119 quantum-software papers and identifies a materialization gap where only 8 claims provide enough matched evidence for direct auditing, yielding 2 sustained, 4 unresol...

  3. Unitary Channel Testing Under a Depolarizing Noise Assumption

    quant-ph 2026-06 unverdicted novelty 7.0

    Optimal query algorithms for testing unitary channels under depolarizing noise yield Θ(1/ε) complexity with matching lower bounds even for adaptive ancilla-assisted protocols.

Reference graph

Works this paper leans on

269 extracted references · 12 canonical work pages · cited by 3 Pith papers · 1 internal anchor

  1. [1]

    Preskill, in Feynman Lectures on Computation (CRC Press, 2023), pp

    J. Preskill, in Feynman Lectures on Computation (CRC Press, 2023), pp. 193– 244

  2. [2]

    Shor, Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer

    P.W. Shor, Polynomial-time algorithms for prime factorization and discrete logarithms on a quantum computer. SIAM review 41(2), 303–332 (1999)

  3. [3]

    Schuld, N

    M. Schuld, N. Killoran, Is quantum advantage the right goal for quantum machine learning? Prx Quantum 3(3), 030101 (2022)

  4. [4]

    Hoefler, T

    T. Hoefler, T. H¨ aner, M. Troyer, Disentangling hype from practicality: On real- istically achieving quantum advantage. Communications of the ACM 66(5), 82–87 (2023)

  5. [5]

    Preskill, Quantum computing in the nisq era and beyond

    J. Preskill, Quantum computing in the nisq era and beyond. Quantum 2, 79 (2018)

  6. [6]

    Lubinski, S

    T. Lubinski, S. Johri, P. Varosy, J. Coleman, L. Zhao, J. Necaise, C.H. Bald- win, K. Mayer, T. Proctor, Application-oriented performance benchmarks for quantum computing. IEEE Transactions on Quantum Engineering 4, 1–32 (2023)

  7. [7]

    Murphy, K.R

    D.C. Murphy, K.R. Brown, Controlling error orientation to improve quantum algorithm success rates. Physical Review A 99(3), 032318 (2019)

  8. [8]

    Proctor, K

    T. Proctor, K. Young, A.D. Baczewski, R. Blume-Kohout, Benchmarking quantum computers. Nature Reviews Physics pp. 1–14 (2025)

  9. [9]

    Resch, U.R

    S. Resch, U.R. Karpuzcu, Benchmarking quantum computers and the impact of quantum noise. ACM Computing Surveys (CSUR) 54(7), 1–35 (2021)

  10. [10]

    Bowles, S

    J. Bowles, S. Ahmed, M. Schuld, Better than classical? the subtle art of bench- marking quantum machine learning models. arXiv preprint arXiv:2403.07059 (2024)

  11. [11]

    Grootendorst, Bertopic: Neural topic modeling with a class-based tf-idf procedure

    M. Grootendorst, Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794 (2022)

  12. [12]

    Acuaviva, D

    A. Acuaviva, D. Aguirre, R. Pe˜ na, M. Sanz, Benchmarking quantum com- puters: Towards a standard performance evaluation approach. arXiv preprint arXiv:2407.10941 (2024)

  13. [13]

    J. v. Kistowski, J.A. Arnold, K. Huppler, K.D. Lange, J.L. Henning, P. Cao, How to build a benchmark , in Proceedings of the 6th ACM/SPEC international conference on performance engineering (2015), pp. 333–336 54

  14. [14]

    W. Dai, D. Berleant, Benchmarking contemporary deep learning hardware and frameworks: A survey of qualitative metrics , in 2019 IEEE First International Conference on Cognitive Machine Intelligence (CogMI) (IEEE, 2019), pp. 148– 155

  15. [15]

    Donkers, K

    H. Donkers, K. Mesman, Z. Al-Ars, M. M¨ oller, Qpack scores: Quantitative per- formance metrics for application-oriented quantum computer benchmarking. arXiv preprint arXiv:2205.12142 (2022)

  16. [16]

    Tomesh, P

    T. Tomesh, P. Gokhale, V. Omole, G.S. Ravi, K.N. Smith, J. Viszlai, X.C. Wu, N. Hardavellas, M.R. Martonosi, F.T. Chong, Supermarq: A scalable quantum benchmark suite, in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) (IEEE, 2022), pp. 587–603

  17. [17]

    X. Fu, L. Riesebos, L. Lao, C.G. Almudever, F. Sebastiano, R. Versluis, E. Charbon, K. Bertels, A heterogeneous quantum computer architecture , in Proceedings of the ACM International Conference on Computing Frontiers (Association for Computing Machinery, New York, NY, USA, 2016), CF ’16, p. 323–330. https://doi.org/10.1145/2903150.2906827. URL https://do...

  18. [18]

    X. Fu, M.A. Rol, C.C. Bultink, J. Van Someren, N. Khammassi, I. Ashraf, R. Vermeulen, J. De Sterke, W. Vlothuizen, R. Schouten, et al.,An experimental microarchitecture for a superconducting quantum processor, in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture(2017), pp. 813–825

  19. [19]

    Jones, R

    N.C. Jones, R. Van Meter, A.G. Fowler, P.L. McMahon, J. Kim, T.D. Ladd, Y. Yamamoto, Layered architecture for quantum computing. Phys. Rev. X 2, 031007 (2012). https://doi.org/10.1103/PhysRevX.2.031007. URL https: //link.aps.org/doi/10.1103/PhysRevX.2.031007

  20. [20]

    Van Meter, D

    R. Van Meter, D. Horsman, A blueprint for building a quantum computer. Commun. ACM 56(10), 84–93 (2013). https://doi.org/10.1145/2494568. URL https://doi.org/10.1145/2494568

  21. [21]

    C´ orcoles, A

    A.D. C´ orcoles, A. Kandala, A. Javadi-Abhari, D.T. McClure, A.W. Cross, K. Temme, P.D. Nation, M. Steffen, J.M. Gambetta, Challenges and opportuni- ties of near-term quantum computing systems. Proceedings of the IEEE 108(8), 1338–1352 (2020). https://doi.org/10.1109/JPROC.2019.2954005

  22. [22]

    Grover, A fast quantum mechanical algorithm for database search , in Pro- ceedings of the twenty-eighth annual ACM symposium on Theory of computing (1996), pp

    L.K. Grover, A fast quantum mechanical algorithm for database search , in Pro- ceedings of the twenty-eighth annual ACM symposium on Theory of computing (1996), pp. 212–219

  23. [23]

    Da Rosa, R

    E.C.R. Da Rosa, R. De Santiago, Ket quantum programming. ACM Journal on Emerging Technologies in Computing Systems (JETC) 18(1), 1–25 (2021) 55

  24. [24]

    URL https://github.com/ microsoft/qsharp-language/tree/main/Specifications/Language#q-language

    Microsoft, Q# Language Specification (2020). URL https://github.com/ microsoft/qsharp-language/tree/main/Specifications/Language#q-language

  25. [25]

    Wecker, K.M

    D. Wecker, K.M. Svore. Liqui—¿: A software design architecture and domain- specific language for quantum computing (2014). URL https://arxiv.org/abs/ 1402.4467

  26. [26]

    Green, P.L

    A.S. Green, P.L. Lumsdaine, N.J. Ross, P. Selinger, B. Valiron, Quipper: a scal- able quantum programming language. ACM SIGPLAN Notices 48(6), 333–342 (2013). https://doi.org/10.1145/2499370.2462177. URL http://dx.doi.org/10. 1145/2499370.2462177

  27. [27]

    Javadi-Abhari, M

    A. Javadi-Abhari, M. Treinish, K. Krsulich, C.J. Wood, J. Lishman, J. Gacon, S. Martiel, P.D. Nation, L.S. Bishop, A.W. Cross, B.R. Johnson, J.M. Gambetta. Quantum computing with Qiskit (2024). https://doi.org/10.48550/arXiv.2405. 08810

  28. [28]

    Developers, Cirq (Zenodo, 2024)

    C. Developers, Cirq (Zenodo, 2024). https://doi.org/10.5281/ZENODO. 4062499. URL https://zenodo.org/doi/10.5281/zenodo.4062499

  29. [29]

    Bergholm, J

    V. Bergholm, J. Izaac, M. Schuld, C. Gogolin, S. Ahmed, V. Ajith, M.S. Alam, G. Alonso-Linaje, B. AkashNarayanan, A. Asadi, J.M. Arrazola, U. Azad, S. Banning, C. Blank, T.R. Bromley, B.A. Cordier, J. Ceroni, A. Delgado, O.D. Matteo, A. Dusko, T. Garg, D. Guala, A. Hayes, R. Hill, A. Ijaz, T. Isacsson, D. Ittah, S. Jahangiri, P. Jain, E. Jiang, A. Khandel...

  30. [30]

    AbuGhanem, Ibm quantum computers: Evolution, performance, and future directions

    M. AbuGhanem, Ibm quantum computers: Evolution, performance, and future directions. The Journal of Supercomputing 81(5), 687 (2025)

  31. [31]

    Chong, D

    F.T. Chong, D. Franklin, M. Martonosi, Programming languages and compiler design for realistic quantum hardware. Nature 549(7671), 180–187 (2017)

  32. [32]

    Thomsen, R

    M.K. Thomsen, R. Gl¨ uck, H.B. Axelsen, Reversible arithmetic logic unit for quantum arithmetic. Journal of Physics A: Mathematical and Theoretical 43(38), 382002 (2010)

  33. [33]

    Killoran, J

    N. Killoran, J. Izaac, N. Quesada, V. Bergholm, M. Amy, C. Weedbrook, Strawberry fields: A software platform for photonic quantum computing. Quan- tum 3, 129 (2019). https://doi.org/10.22331/q-2019-03-11-129. URL http: //dx.doi.org/10.22331/q-2019-03-11-129 56

  34. [34]

    Cross, L.S

    A.W. Cross, L.S. Bishop, J.A. Smolin, J.M. Gambetta. Open quantum assembly language (2017). URL https://arxiv.org/abs/1707.03429

  35. [35]

    Khammassi, G.G

    N. Khammassi, G.G. Guerreschi, I. Ashraf, J.W. Hogaboam, C.G. Almudever, K. Bertels. cqasm v1.0: Towards a common quantum assembly language (2018). URL https://arxiv.org/abs/1805.09607

  36. [36]

    Reilly, Engineering the quantum-classical interface of solid-state qubits

    D.J. Reilly, Engineering the quantum-classical interface of solid-state qubits. npj Quantum Information 1(1), 1–10 (2015)

  37. [37]

    J. Wang, G. Guo, Z. Shan, Sok: Benchmarking the performance of a quantum computer. Entropy 24(10), 1467 (2022)

  38. [38]

    Lorenz, T

    J.M. Lorenz, T. Monz, J. Eisert, D. Reitzner, F. Schopfer, F. Barbaresco, K. Kurowski, W. van der Schoot, T. Strohm, J. Senellart, et al., System- atic benchmarking of quantum computers: status and recommendations. arXiv preprint arXiv:2503.04905 (2025)

  39. [39]

    Eisert, D

    J. Eisert, D. Hangleiter, N. Walk, I. Roth, D. Markham, R. Parekh, U. Chabaud, E. Kashefi, Quantum certification and benchmarking. Nature Reviews Physics 2(7), 382–390 (2020)

  40. [40]

    Bharti, A

    K. Bharti, A. Cervera-Lierta, T.H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, H. Heimonen, J.S. Kottmann, T. Menke, et al., Noisy intermediate- scale quantum algorithms. Reviews of Modern Physics 94(1), 015004 (2022)

  41. [41]

    Huang, X.Y

    H.L. Huang, X.Y. Xu, C. Guo, G. Tian, S.J. Wei, X. Sun, W.S. Bao, G.L. Long, Near-term quantum computing techniques: Variational quantum algorithms, error mitigation, circuit compilation, benchmarking and classical simulation. Science China Physics, Mechanics & Astronomy 66(5), 250302 (2023)

  42. [42]

    Hashim, L.B

    A. Hashim, L.B. Nguyen, N. Goss, B. Marinelli, R.K. Naik, T. Chistolini, J. Hines, J.P. Marceaux, Y. Kim, P. Gokhale, T. Tomesh, S. Chen, L. Jiang, S. Ferracin, K. Rudinger, T. Proctor, K.C. Young, R. Blume-Kohout, I. Sid- diqi. A practical introduction to benchmarking and characterization of quantum computers (2024). URL https://arxiv.org/abs/2408.12064

  43. [43]

    Helsen, I

    J. Helsen, I. Roth, E. Onorati, A.H. Werner, J. Eisert, General framework for randomized benchmarking. PRX Quantum 3(2), 020357 (2022)

  44. [44]

    Y. Xiao, M. Watson, Guidance on conducting a systematic literature review. Journal of planning education and research 39(1), 93–112 (2019)

  45. [45]

    Biolchini, P.G

    J. Biolchini, P.G. Mian, A.C.C. Natali, G.H. Travassos, Systematic review in software engineering. System engineering and computer science department COPPE/UFRJ, Technical Report ES 679(05), 45 (2005) 57

  46. [46]

    Grant, A

    M.J. Grant, A. Booth, A typology of reviews: an analysis of 14 review types and associated methodologies. Health information & libraries journal 26(2), 91–108 (2009)

  47. [47]

    Brereton, B.A

    P. Brereton, B.A. Kitchenham, D. Budgen, M. Turner, M. Khalil, Lessons from applying the systematic literature review process within the software engineering domain. Journal of systems and software 80(4), 571–583 (2007)

  48. [48]

    Bishop, S

    L.S. Bishop, S. Bravyi, A. Cross, J.M. Gambetta, J. Smolin, Quantum volume. Quantum Volume. Technical Report (2017)

  49. [49]

    A. Wack, H. Paik, A. Javadi-Abhari, P. Jurcevic, I. Faro, J.M. Gambetta, B.R. Johnson. Quality, speed, and scale: three key attributes to measure the perfor- mance of near-term quantum computers (2021). URL https://arxiv.org/abs/ 2110.14108

  50. [50]

    Martiel, T

    S. Martiel, T. Ayral, C. Allouche, Benchmarking quantum coprocessors in an application-centric, hardware-agnostic, and scalable way. IEEE Transactions on Quantum Engineering 2, 1–11 (2021)

  51. [51]

    D. Lall, A. Agarwal, W. Zhang, L. Lindoy, T. Lindstr¨ om, S. Webster, S. Hall, N. Chancellor, P. Wallden, R. Garcia-Patron, et al., A review and collection of metrics and benchmarks for quantum computers: definitions, methodologies and software. arXiv preprint arXiv:2502.06717 (2025)

  52. [52]

    van der Schoot, R

    W. van der Schoot, R. Wezeman, P.T. Eendebak, N.M. Neumann, F. Phillip- son, Evaluating three levels of quantum metrics on quantum-inspire hardware. Quantum Information Processing 22(12), 451 (2023)

  53. [53]

    Gilchrist, N.K

    A. Gilchrist, N.K. Langford, M.A. Nielsen, Distance measures to compare real and ideal quantum processes. Physical Review A—Atomic, Molecular, and Optical Physics 71(6), 062310 (2005)

  54. [54]

    Linke, D

    N.M. Linke, D. Maslov, M. Roetteler, S. Debnath, C. Figgatt, K.A. Landsman, K. Wright, C. Monroe, Experimental comparison of two quantum computing architectures. Proceedings of the National Academy of Sciences 114(13), 3305– 3310 (2017)

  55. [55]

    Shi, Both toffoli and controlled-not need little help to do universal quantum computation

    Y. Shi, Both toffoli and controlled-not need little help to do universal quantum computation. arXiv preprint quant-ph/0205115 (2002)

  56. [56]

    Ramsey, A molecular beam resonance method with separated oscillating fields

    N.F. Ramsey, A molecular beam resonance method with separated oscillating fields. Phys. Rev. 78, 695–699 (1950). https://doi.org/10.1103/PhysRev.78.695. URL https://link.aps.org/doi/10.1103/PhysRev.78.695

  57. [57]

    Hahn, Spin echoes

    E.L. Hahn, Spin echoes. Phys. Rev. 80, 580–594 (1950). https://doi.org/10. 1103/PhysRev.80.580. URL https://link.aps.org/doi/10.1103/PhysRev.80.580 58

  58. [58]

    Korotkov, Error matrices in quantum process tomography

    A.N. Korotkov, Error matrices in quantum process tomography. arXiv preprint arXiv:1309.6405 (2013)

  59. [59]

    Flammia, Y.K

    S.T. Flammia, Y.K. Liu, Direct fidelity estimation from few pauli measurements. Physical review letters 106(23), 230501 (2011)

  60. [60]

    Zhang, M

    X. Zhang, M. Luo, Z. Wen, Q. Feng, S. Pang, W. Luo, X. Zhou, Direct fidelity estimation of quantum states using machine learning. Physical Review Letters 127(13), 130503 (2021)

  61. [61]

    S. Aaronson, Shadow tomography of quantum states , in Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing (Associa- tion for Computing Machinery, New York, NY, USA, 2018), STOC 2018, p. 325–338. https://doi.org/10.1145/3188745.3188802. URL https://doi.org/10. 1145/3188745.3188802

  62. [62]

    Schumacher, Sending entanglement through noisy quantum channels

    B. Schumacher, Sending entanglement through noisy quantum channels. Phys- ical Review A 54(4), 2614 (1996)

  63. [63]

    Nielsen, A simple formula for the average gate fidelity of a quantum dynamical operation

    M.A. Nielsen, A simple formula for the average gate fidelity of a quantum dynamical operation. Physics Letters A 303(4), 249–252 (2002)

  64. [64]

    Horodecki, P

    M. Horodecki, P. Horodecki, R. Horodecki, General teleportation channel, singlet fraction, and quasidistillation. Physical Review A 60(3), 1888 (1999)

  65. [65]

    Proctor, K

    T. Proctor, K. Rudinger, K. Young, M. Sarovar, R. Blume-Kohout, What ran- domized benchmarking actually measures. Physical review letters 119(13), 130502 (2017)

  66. [66]

    Aharonov, A

    D. Aharonov, A. Kitaev, N. Nisan, Quantum circuits with mixed states , in Pro- ceedings of the thirtieth annual ACM symposium on Theory of computing(1998), pp. 20–30

  67. [67]

    Benenti, G

    G. Benenti, G. Strini, Computing the distance between quantum channels: use- fulness of the fano representation. Journal of Physics B: Atomic, Molecular and Optical Physics 43(21), 215508 (2010)

  68. [68]

    H. Neven. Meet Willow, our state-of-the-art quantum chip — blog.google. https://blog.google/technology/research/google-willow-quantum-chip/ (2024). [Accessed 27-04-2025]

  69. [69]

    McKay, C.J

    D.C. McKay, C.J. Wood, S. Sheldon, J.M. Chow, J.M. Gambetta, Efficient z gates for quantum computing. Phys. Rev. A 96, 022330 (2017). https:// doi.org/10.1103/PhysRevA.96.022330. URL https://link.aps.org/doi/10.1103/ PhysRevA.96.022330 59

  70. [70]

    Arute, K

    F. Arute, K. Arya, R. Babbush, D. Bacon, J.C. Bardin, R. Barends, R. Biswas, S. Boixo, F.G. Brandao, D.A. Buell, et al., Quantum supremacy using a programmable superconducting processor. Nature 574(7779), 505–510 (2019)

  71. [71]

    N¨ unnerich, D

    M. N¨ unnerich, D. Cohen, P. Barthel, P.H. Huber, D. Niroomand, A. Ret- zker, C. Wunderlich, Fast, robust, and laser-free universal entangling gates for trapped-ion quantum computing. Physical Review X 15(2), 021079 (2025)

  72. [72]

    Torlai, G

    G. Torlai, G. Mazzola, J. Carrasquilla, M. Troyer, R. Melko, G. Carleo, Neural-network quantum state tomography. Nature Physics 14(5), 447–450 (2018). https://doi.org/10.1038/s41567-018-0048-5. URL https://doi.org/10. 1038/s41567-018-0048-5

  73. [73]

    Lange, M

    H. Lange, M. Kebric, M. Buser, U. Schollw¨ ock, F. Grusdt, A. Bohrdt, Adaptive quantum state tomography with active learning. Quantum 7, 1129 (2023)

  74. [74]

    Leonhardt, Discrete wigner function and quantum-state tomography

    U. Leonhardt, Discrete wigner function and quantum-state tomography. Physi- cal Review A 53(5), 2998 (1996)

  75. [75]

    K. He, M. Yuan, Y. Wong, S. Chakram, A. Seif, L. Jiang, D.I. Schuster, Efficient multimode wigner tomography. Nature communications 15(1), 4138 (2024)

  76. [76]

    Schmied, Quantum state tomography of a single qubit: comparison of meth- ods

    R. Schmied, Quantum state tomography of a single qubit: comparison of meth- ods. Journal of Modern Optics 63(18), 1744–1758 (2016). https://doi.org/10. 1080/09500340.2016.1142018. URL http://dx.doi.org/10.1080/09500340.2016. 1142018

  77. [77]

    Binosi, G

    D. Binosi, G. Garberoglio, D. Maragnano, M. Dapor, M. Liscidini, A tailor-made quantum state tomography approach. APL Quantum 1(3) (2024)

  78. [78]

    Y. Guo, S. Yang, Quantum state tomography with locally purified density operators and local measurements. Communications Physics 7(1), 322 (2024)

  79. [79]

    Wood, J.D

    C.J. Wood, J.D. Biamonte, D.G. Cory, Tensor networks and graphical calculus for open quantum systems. Quantum Info. Comput. 15(9–10), 759–811 (2015)

  80. [80]

    Blumoff, K

    J.Z. Blumoff, K. Chou, C. Shen, M. Reagor, C. Axline, R. Brierley, M. Silveri, C. Wang, B. Vlastakis, S.E. Nigg, et al., Implementing and characterizing precise multiqubit measurements. Physical Review X 6(3), 031041 (2016)

Showing first 80 references.