Pith. sign in

REVIEW 4 major objections 3 minor 1 cited by

The paper introduces Clifford Volume and Free Fermion Volume as scalable, classically verifiable benchmarks whose score is the largest qubit width at which random Clifford or free-fermion unitaries pass threshold-based verification, and it

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 14:42 UTC pith:BHOXT6E2

load-bearing objection Two genuinely new volumetric benchmarks with a real hardware demo, but the headline Clifford Volume of 34 rests on a factor-of-two error in the paper's own pass criterion. the 4 major comments →

arxiv 2512.19413 v2 pith:BHOXT6E2 submitted 2025-12-22 quant-ph

Clifford Volume and Free Fermion Volume: Complementary Scalable Benchmarks for Quantum Computers

classification quant-ph PACS 03.67.Lx
keywords Clifford groupfree fermionsquantum benchmarkingvolumetric benchmarkstabilizer statesclassical verificationnoise modeltrapped-ion device
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to give quantum computers a device-level score that stays meaningful as systems grow: run random circuits drawn from two groups—Clifford operations and free-fermion operations—that are individually easy for a classical computer to verify, and record the largest number of qubits n for which the device reliably passes threshold tests on measured expectation values. If this works, it gives a platform-independent, scalable alternative to Quantum Volume, whose output probabilities become classically intractable at scale. The two families together generate all quantum circuits, so the benchmarks are not sampling an unrepresentative corner of computation; they also serve as real algorithmic primitives in shadow tomography and quantum chemistry. The paper shows numerically that the scores respond smoothly to two-qubit gate and readout noise, and validates the Clifford benchmark end-to-end on a trapped-ion device, reporting a score of 34.

Core claim

The central claim is that a volumetric benchmark can be built from the Clifford group and the free-fermion group without sacrificing classical verifiability. For each width n, the protocol samples K=4 random n-qubit unitaries from the chosen group, compiles them to the device, measures four stabilizer and four destabilizer Pauli expectation values (or, for free fermions, selected Majorana-mode linear combinations), and declares the width passed only if every measured value clears the prescribed thresholds—stabilizers above 1/e and destabilizers below 1/(2e)—by at least two standard deviations, with averages clearing by five standard deviations. The benchmark score is the largest n for which

What carries the argument

The load-bearing objects are two group families: the n-qubit Clifford group, whose unitaries permute Pauli operators and whose stabilizer states have ideal expectation values of 1 on stabilizer generators and 0 on non-stabilizer Paulis, and the free-fermion group, whose evolutions are represented by orthogonal matrices O in SO(2n) acting on Majorana operators. For Clifford circuits, verification uses the stabilizer/destabilizer dichotomy with thresholds 1/e and 1/(2e). For free-fermion circuits, verification uses the orthogonality relation sum_k O_ki O_kj = delta_ij, reconstructed from measured Majorana expectation values, with the same thresholds. The mechanism that makes the protocol scala

Load-bearing premise

The entire score depends on the assumption that four random circuits, four measured observables per circuit, 512 shots, and the hand-picked thresholds 1/e and 1/(2e) are enough to tell a working device from a failing one—a statistical choice the paper does not justify with a power analysis.

What would settle it

Run the Clifford Volume protocol at 34 qubits on the same device type with the same compiled circuits but increase the shot count from 512 to 10,000 per observable and draw a fresh set of K=4 random Clifford unitaries; if any measured stabilizer expectation value minus 2sigma falls below 1/e, or any destabilizer absolute value plus 2sigma rises above 1/(2e), the claimed score of 34 is refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Clifford Volume and Free Fermion Volume can be evaluated at widths beyond what Quantum Volume allows, because the ideal output probabilities are classically computable for both groups; the simulations cover 5–100 qubits.
  • Because the two families together form a universal gate set, the benchmarks jointly cover the full space of quantum circuits while remaining individually verifiable; hybrid Clifford–free-fermion circuits could probe the boundary of classical simulability.
  • The free-fermion benchmark tracks progress that Clifford Volume cannot: its T-gate cost scales quadratically with logical width, similar to Quantum Volume, while Clifford Volume requires no T gates and will saturate in the early fault-tolerant regime.
  • The abstract, hardware-agnostic definitions mean that connectivity and native-gate differences enter only through compilation; the compiled circuit data shown for different topologies indicate that routing overhead, not the benchmark definition, dominates resource scaling.
  • The end-to-end trapped-ion demonstration at 34 qubits shows that the protocol is practically executable on current hardware and yields an interpretable, reproducible device-level score.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The reported score of 34 is a statement about this specific protocol instance, not a hardware invariant: with K=4 and 512 shots, a different random draw or a slightly noisier day could plausibly change the verdict by a qubit or two, since the measured stabilizer values at n=34 sit close to the 1/e boundary.
  • A natural extension the authors do not develop is a statistical framework that converts the measured expectation values and shot counts into a confidence interval for the Clifford Volume, rather than a point score.
  • The same verification trick—testing expectation values of operators whose ideal values are 0 and 1—could be applied to other classically simulable families (e.g., matchgates or Gaussian boson sampling) or to magic-state-injected circuits, as a scalable proxy for quantum advantage.
  • Because the benchmark deliberately captures readout errors alongside gate errors, scores on a given platform could improve simply by better measurement; this is a feature for holistic tracking but means the benchmark does not isolate gate fidelity the way randomized benchmarking does.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper introduces two volumetric quantum-computing benchmarks, Clifford Volume (CLV) and Free Fermion Volume (FFV), in which a device is tested at increasing qubit width n by executing K random n-qubit unitaries from the Clifford group or from the free-fermion (SO(2n)) group and verifying expectation values that are classically computable. The benchmark score is defined as the largest n for which the protocol's acceptance criteria hold. The authors provide numerical simulations under simple depolarizing and readout noise models, an open-source implementation, and an experimental demonstration on Quantinuum H2-1, reporting a Clifford Volume of 34. The paper argues that these benchmarks are scalable, classically verifiable, and platform-independent, and that Clifford and free-fermion operations are complementary because together they form a universal gate set.

Significance. If the protocol is made precise, this is a useful and timely contribution. The choice of classically simulable yet algorithmically relevant unitary families is well motivated, and the experimental demonstration is a genuine end-to-end test. The paper also ships numerical baselines and an open-source suite, which is valuable for adoption. I do not see a circularity problem: the H2-1 score is a measurement, not a fitted parameter, and the cited references supply standard background results. However, the central experimental claim and the FFV protocol definition contain statistical inconsistencies that must be resolved before the benchmark scores can be considered reliable. These issues are fixable, so my recommendation is major revision rather than rejection.

major comments (4)
  1. [Eq. (8) vs Appendix A; Table IV; Fig. 13] The published average-criterion formula is ambiguous. Eq. (8) defines sigma_{P^k} = sqrt( (1/4) sum_i sigma_{P_i^k}^2 ), while Appendix A, Performance Evaluation II, defines sigma_{P^k} = (1/4) sqrt( sum_i sigma_{P_i^k}^2 ). These differ by a factor of 2. For the C1 round at n=34 in Table IV, the stabilizer average is <S> = 0.486 and the RMS individual sigma is about 0.0388. Under Eq. (8), the condition <S> - 5 sigma >= 1/e gives 0.486 - 0.194 = 0.292, which is below 1/e, so the average stabilizer criterion fails. Under Appendix A, sigma is about 0.0194, giving 0.486 - 0.097 = 0.389, a pass. Thus the statement in Fig. 13 that all four rounds satisfy the benchmark criteria holds only with the Appendix A formula. Since the reported Clifford Volume of 34 is inferred from this pass, the central experimental claim is not well-defined as printed.
  2. [Sec. IV, Table IV, Appendix A step 4] The H2-1 result is statistically fragile. At n=34, C1, S1 has <S> - 2 sigma = 0.448 - 0.079 = 0.369, which exceeds the threshold 1/e=0.3679 by only about 0.001. With L=512 shots, a single bit-flip changes the estimate by 2/512 = 0.0039, enough to flip the round. No statistical power analysis is given for K=4 random unitaries, n_m=4 observables per category, or L=512 shots, and the thresholds are acknowledged to be chosen by hand. Consequently, the protocol does not quantify the probability that a device near the boundary passes or fails by chance. This is not a purely philosophical concern: the reported score 34 rests on a one-shot margin.
  3. [Eq. (27) with Eq. (29)] The FFV parallel uncertainty formula is incorrect as printed. For the reduced parallel combination Y_parallel = ( sum_{l in J} O_{li} <m_l> ) / ( sum_{l in J} O_{li}^2 ), propagation of shot noise gives a standard error of sqrt( sum_l O_{li}^2 sigma_l^2 ) / ( sum_l O_{li}^2 ). Equation (29) instead prints sqrt( sum_l O_{li}^2 (1-<m_l>^2)/L / sum_l O_{li}^2 ), which is missing one factor of the denominator. This changes the 2-sigma acceptance condition in Eq. (28) and makes the FFV pass/fail criterion depend on an uncorrected error-propagation expression. Please correct or clarify the intended formula.
  4. [Eq. (24) and Appendix B step 1(c)] The FFV measurement cut-off is inconsistently specified. The main text, Eq. (24), sets N(n) = 20 + floor(n/5) for n > 10, while Appendix B step 1(c)(ii) sets n_m = 20 + floor(n/20). Appendix D's r=10% analysis (Eqs. D12-D14) matches neither value: for n=100, Eq. (24) selects 40 of 200 entries (20%), Appendix B selects 25 (12.5%), and Appendix D uses 10%. Because the measured subset directly enters the renormalized combination in Eq. (27), the FFV protocol is not uniquely defined. Please specify the intended rule and make the appendices consistent.
minor comments (3)
  1. [Throughout] There are numerous typos, e.g., 'operetions' in Sec. I, 'te number' in Sec. III, 'the the expected value' and 'expected values of of' in Sec. II.C. A careful proofread is needed.
  2. [References] References [48] and [54] appear to be the same paper (Merkel et al., 'When Clifford benchmarks are sufficient...') with different arXiv identifiers; please consolidate.
  3. [Fig. 2 caption] The caption states subfigures (a)-(d) correspond to n=5,15,25,35, but the visible histograms are hard to distinguish; consider adding explicit n labels or a table of values.

Circularity Check

0 steps flagged

No circular derivation: benchmark scores are protocol outputs, not fitted predictions; self-citations are background.

full rationale

The paper's central derivation—the Clifford Volume and Free Fermion Volume acceptance criteria in Eqs. (4)-(8) and (27)-(29)—follows from the algebraic definitions of stabilizer states, Pauli expectation values, and free-fermion time evolution. No parameter appearing in the success criteria is fitted to the reported H2-1 Clifford Volume of 34. The thresholds tau_S=1/e and tau_D=1/(2e), K=4, n_m=4, and L>=512 are fixed protocol choices that the paper explicitly describes as somewhat arbitrary, and the headline score is an experimental measurement under those stated conditions rather than a fitted output. The self-citations (e.g., Refs. [49], [50], [60], and the software suite [52]) supply standard background results in group theory, fermionic simulation, and benchmarking; they are not used to infer the benchmark score. The internal inconsistency between the sigma formula in Eq. (8) and the sigma formula in Appendix A is a correctness and robustness concern about the printed pass/fail criterion, not a circularity: neither expression is obtained by fitting or by referencing the target result. Overall, the derivation chain is self-contained, and no prediction reduces to an input by construction.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 0 invented entities

The central claims rest on a handful of protocol parameters chosen by hand (thresholds, sample sizes, measurement reductions) rather than on fitted physical constants. The mathematical axioms are standard background results from the cited literature. No new physical entities are introduced.

free parameters (7)
  • Stabilizer threshold tau_S = 1/e ≈ 0.3679
    Chosen as an exponential-decay 'lifetime' under a depolarizing model; explicitly acknowledged as arbitrary. Every Clifford Volume and Free Fermion Volume score depends on it.
  • Destabilizer/orthogonal threshold tau_D = 1/(2e) ≈ 0.1839
    Chosen to be moderately close to the ideal value 0; arbitrary, and determines pass/fail for the non-stabilizer and orthogonal linear combinations.
  • Worst-case and average sigma margins = 2 sigma (worst-case), 5 sigma (average)
    Chosen by hand as a 'practical trade-off'; no calibration of the resulting false-pass or false-fail probabilities.
  • Number of random unitaries K = 4
    Very small sample size per width; no power analysis. The reported Clifford Volume of 34 is based on four random Clifford instances at each tested width.
  • Measured observables per Clifford instance n_m = 4 stabilizers + 4 destabilizers
    Choice reduces experimental cost; sensitivity of the benchmark depends on this number.
  • Measurement shots L = 512 (4096 in some simulations)
    Shots determine the statistical uncertainty; 512 was used on H2-1 and is the minimum allowed in the protocol.
  • FFV measurement subset N(n) = 2n for n ≤ 10; 20 + floor(n/5) or floor(n/20) for n > 10
    Formulas in Eq. (24) and Appendix B are inconsistent. The reduction is arbitrary and affects the sensitivity of the Free Fermion Volume score.
axioms (6)
  • standard math Uniform sampling from the n-qubit Clifford group is efficient and Clifford circuits are classically simulable (Gottesman-Knill).
    Cited from Refs. [40, 56, 57]; the basis for classical verifiability of the Clifford Volume benchmark.
  • standard math Free-fermion (matchgate) unitaries are classically simulable via SO(2n) orthogonal matrices and the Jordan-Wigner transformation.
    Cited from Refs. [41, 60]; the basis for classical verifiability of the Free Fermion Volume benchmark.
  • standard math Clifford operations together with free-fermion operations form a universal gate set.
    Cited from Ref. [49]; motivates the complementarity claim but is not proved in this paper.
  • domain assumption Depolarizing two-qubit errors plus independent bit-flip readout errors are representative of current hardware for benchmarking purposes.
    Used for all numerical simulations in Sec. III; only qualitatively connected to real hardware through the H2-1 emulator and experiment.
  • domain assumption Entries of random SO(2n) matrices can be approximated as independent Gaussian variables for the purpose of selecting a small measurement subset.
    Appendix D explicitly invokes this approximation to justify the top-10% selection and renormalization in the FFV protocol.
  • ad hoc to paper A device's capacity at width n is adequately assessed by K=4 random unitaries, 4 observables per unitary, and 512 shots.
    Load-bearing for the reported Clifford Volume of 34; no statistical power analysis is given for this sample size.

pith-pipeline@v1.3.0-alltime-deepseek · 31644 in / 16184 out tokens · 167521 ms · 2026-08-03T14:42:13.714604+00:00 · methodology

0 comments
read the original abstract

As quantum computing advances toward the late-NISQ and early fault-tolerant eras, scalable and platform-independent benchmarks are essential for quantifying computational capacity in a classically verifiable manner. We introduce two volumetric benchmarks, Clifford Volume and Free Fermion Volume, that assess quantum hardware by testing the execution of random Clifford and free fermion operations. These two groups of unitaries possess a combination of properties that make them ideal for benchmarking: (i) each is individually efficient to simulate classically, enabling verification at scale; (ii) together they form a universal gate set; (iii) they serve as essential algorithmic primitives in practical applications (including shadow tomography and quantum chemistry); and (iv) their definitions are formulated abstractly, without explicit reference to hardware-specific features such as qubit connectivity or native gate sets. This framework thus enables scalable and fair cross-platform comparisons and tracks meaningful computational advancement. We demonstrate the practical feasibility of these benchmarks through extensive numerical simulations across realistic noise parameters and through experimental validation on Quantinuum's H2-1 trapped-ion quantum computer, which achieves a Clifford Volume of 34.

Figures

Figures reproduced from arXiv: 2512.19413 by Attila Portik, Orsolya K\'alm\'an, Thomas Monz, Zolt\'an Zimbor\'as.

Figure 1
Figure 1. Figure 1: FIG. 1. Graphical representation of the [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: shows an example of how the average expec [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. The simulated minimum stabilizer expectation values [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Distribution of simulated expectation values for [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Schematic representation of the Free-Fermion Vol [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Distribution of measured expectation values of Ma [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. FFV benchmark scores for different pairs of error [PITH_FULL_IMAGE:figures/full_fig_p014_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8. The lowest simulated stabilizer expectation values ob [PITH_FULL_IMAGE:figures/full_fig_p015_8.png] view at source ↗
Figure 10
Figure 10. Figure 10: FIG. 10. Results collected during the initial scan of the CLV [PITH_FULL_IMAGE:figures/full_fig_p015_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: FIG. 11. CLV benchmark results for a single measurement [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figure 13
Figure 13. Figure 13: FIG. 13. CLV benchmark results for a qubit count of 34 on [PITH_FULL_IMAGE:figures/full_fig_p016_13.png] view at source ↗
Figure 12
Figure 12. Figure 12: FIG. 12. CLV benchmark results for a single measurement [PITH_FULL_IMAGE:figures/full_fig_p016_12.png] view at source ↗
Figure 14
Figure 14. Figure 14: FIG. 14. Characteristics of compiled circuit implementations [PITH_FULL_IMAGE:figures/full_fig_p022_14.png] view at source ↗
Figure 16
Figure 16. Figure 16: FIG. 16. Statistical justification of the reducing strategy used [PITH_FULL_IMAGE:figures/full_fig_p024_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Coherent-State Propagation: A Computational Framework for Simulating Bosonic Quantum Systems

    quant-ph 2026-04 unverdicted novelty 8.0

    Coherent-state propagation enables quasi-polynomial classical simulation of bosonic circuits with logarithmically many Kerr gates at exponentially small trace-distance error, with polynomial runtime in the weak-nonlin...

Reference graph

Works this paper leans on

67 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Eisert, D

    J. Eisert, D. Hangleiter, N. Walk, I. Roth, D. Markham, R. Parekh, U. Chabaud, and E. Kashefi, Quantum cer- tification and benchmarking, Nature Reviews Physics2, 382 (2020)

  2. [2]

    Proctor, K

    T. Proctor, K. Young, A. D. Baczewski, and R. Blume- Kohout, Benchmarking quantum computers, Nature Re- views Physics7, 105 (2025), arXiv:2407.08828 [quant- ph]

  3. [3]

    Hashim, L

    A. Hashim, L. B. Nguyen, N. Goss, B. Marinelli, R. K. Naik, T. Chistolini, J. Hines, J. P. Marceaux, Y. Kim, P. Gokhale, T. Tomesh, S. Chen, L. Jiang, S. Fer- racin, K. Rudinger, T. Proctor, K. C. Young, R. Blume- Kohout, and I. Siddiqi, A practical introduction to benchmarking and characterization of quantum comput- 18 ers, PRX Quantum6, 030202 (2025), a...

  4. [4]

    D. Lall, A. Agarwal, W. Zhang, L. Lindoy, T. Lind- str¨ om, S. Webster, S. Hall, N. Chancellor, P. Wallden, R. Garcia-Patron,et al., A review and collection of metrics and benchmarks for quantum computers: def- initions, methodologies and software, arXiv preprint arXiv:2502.06717 (2025)

  5. [5]

    J. M. Lorenz, T. Monz, J. Eisert, D. Reitzner, F. Schopfer, F. Barbaresco, K. Kurowski, W. van der Schoot, T. Strohm, J. Senellart,et al., Systematic bench- marking of quantum computers: status and recommen- dations, arXiv preprint arXiv:2503.04905 (2025)

  6. [6]

    Emerson, R

    J. Emerson, R. Alicki, and K. ˙Zyczkowski, Scalable noise estimation with random unitary operators, Journal of Optics B: Quantum and Semiclassical Optics7, S347 (2005), arXiv:quant-ph/0503243 [quant-ph]

  7. [7]

    Knill, D

    E. Knill, D. Leibfried, R. Reichle, J. Britton, R. B. Blakestad, J. D. Jost, C. Langer, R. Ozeri, S. Sei- delin, and D. J. Wineland, Randomized benchmarking of quantum gates, Physical Review A77, 012307 (2008), arXiv:0707.0963 [quant-ph]

  8. [8]

    Magesan, J

    E. Magesan, J. M. Gambetta, B. R. Johnson, C. A. Ryan, J. M. Chow, S. T. Merkel, M. P. da Silva, G. A. Keefe, M. B. Rothwell, T. A. Ohki, M. B. Ketchen, and M. Steffen, Efficient measurement of quantum gate error by interleaved randomized benchmarking, Physical Re- view Letters109, 080505 (2012), arXiv:1203.4550 [quant- ph]

  9. [9]

    J. J. Wallman and S. T. Flammia, Randomized bench- marking with confidence, New Journal of Physics16, 103032 (2014), arXiv:1404.6025 [quant-ph]

  10. [10]

    Blume-Kohout, J

    R. Blume-Kohout, J. K. Gamble, E. Nielsen, J. Mizrahi, J. D. Sterk, and P. Maunz, Robust, self-consistent, closed-form tomography of quantum logic gates on a trapped ion qubit (2013), arXiv:1310.4492

  11. [11]

    S. T. Merkel, J. M. Gambetta, J. A. Smolin, S. Poletto, A. D. C´ orcoles, B. R. Johnson, C. A. Ryan, and M. Stef- fen, Self-consistent quantum process tomography, Physi- cal Review A87, 062119 (2013), arXiv:1211.0322 [quant- ph]

  12. [12]

    Nielsen, J

    E. Nielsen, J. K. Gamble, K. Rudinger, T. Scholten, K. Young, and R. Blume-Kohout, Gate set tomography, Quantum5, 557 (2021), arXiv:2009.07301 [quant-ph]

  13. [13]

    N¨ oller, N

    J. N¨ oller, N. Miklin, M. Kliesch, and M. Gachechiladze, Classical certification of quantum gates under the dimen- sion assumption, Quantum9, 1825 (2025)

  14. [14]

    A. W. Cross, L. S. Bishop, S. Sheldon, P. D. Nation, and J. M. Gambetta, Validating quantum computers us- ing randomized model circuits, Physical Review A100, 032328 (2019), arXiv:1811.12926 [quant-ph]

  15. [15]

    Blume-Kohout and K

    R. Blume-Kohout and K. C. Young, A volumetric frame- work for quantum computer benchmarks, Quantum4, 362 (2020), arXiv:1904.05546 [quant-ph]

  16. [16]

    Boixo, S

    S. Boixo, S. V. Isakov, V. N. Smelyanskiy, R. Babbush, N. Ding, Z. Jiang, M. J. Bremner, J. M. Martinis, and H. Neven, Characterizing quantum supremacy in near- term devices, Nature Physics14, 595 (2018)

  17. [17]

    Erhard, J

    A. Erhard, J. J. Wallman, L. Postler, M. Meth, R. Stricker, E. A. Martinez, P. Schindler, T. Monz, J. Emerson, and R. Blatt, Characterizing large-scale quantum computers via cycle benchmarking, Nature Communications10, 5347 (2019), arXiv:1902.08543 [quant-ph]

  18. [18]

    Hines and T

    J. Hines and T. Proctor, Scalable full-stack benchmarks for quantum computers (2023), arXiv:2312.14107

  19. [19]

    Demarty, J

    M. Demarty, J. Mills, K. Hammam, and R. Garcia- Patron, Entropy density benchmarking of near-term quantum circuits, arXiv preprint arXiv:2412.18007 (2024)

  20. [20]

    Proctor, A

    T. Proctor, A. Tran, X. Liu, A. Dhumuntarao, S. Seritan, A. Green, and N. M. Linke, Featuremetric benchmarking: Quantum computer benchmarks based on circuit features (2025), arXiv:2504.12575

  21. [21]

    Zindorf, L

    B. Zindorf, L. Braccini, D. Das, and S. Bose, How ”quantum” is your quantum computer? macrorealism- based benchmarking via mid-circuit parity measure- ments, arXiv preprint arXiv:2511.15881 (2025)

  22. [22]

    Zimbor´ as, A

    Z. Zimbor´ as, A. Portik, D. Aguirre, R. Pe˜ na, D. Svastits, A. P´ alyi, A. M´ arton, J. K. Asb´ oth, A. Frisk Kockum, M. Sanz, O. K´ alm´ an, and T. Monz, The EU Quan- tum Flagship’s Key Performance Indicators for Quantum Computing (2025), Manuscript in preparation

  23. [23]

    Tomesh, P

    T. Tomesh, P. Gokhale, V. Omole, G. S. Ravi, K. N. Smith, J. Viszlai, X.-C. Wu, N. Hardavellas, M. R. Martonosi, and F. T. Chong, Supermarq: A scalable quantum benchmark suite, in2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)(IEEE, 2022) p. 587–603

  24. [24]

    K. Chen, Z. Liang, J. Liao,et al., Veriqbench: A bench- mark for multiple types of quantum circuits (2022), arXiv:2206.10880 [quant-ph]

  25. [25]

    Lubinski, J

    T. Lubinski, J. J. Goings, K. Mayer, S. Johri, N. Reddy, A. Mehta, N. Bhatia, S. Rappaport, D. Mills, C. H. Baldwin,et al., Quantum algorithm exploration us- ing application-oriented performance benchmarks, arXiv preprint arXiv:2402.08985 (2024)

  26. [26]

    Dong and L

    Y. Dong and L. Lin, Random circuit block-encoded ma- trix and a proposal of quantum linpack benchmark, Phys- ical Review A103, 062412 (2021)

  27. [27]

    Martiel, T

    S. Martiel, T. Ayral, and C. Allouche, Benchmark- ing quantum coprocessors in an application-centric, hardware-agnostic, and scalable way, IEEE Transactions on Quantum Engineering2, 1 (2021)

  28. [28]

    van der Schoot, R

    W. van der Schoot, R. Wezeman, N. Neumann, F. Phillipson, and R. Kooij, Extending the Q-score to an application-level quantum metric framework, in2024 IEEE International Conference on Quantum Computing and Engineering (QCE)(IEEE, 2024) p. 941–951

  29. [29]

    Montanez-Barrera, K

    J. Montanez-Barrera, K. Michielsen, and D. E. B. Neira, Evaluating the performance of quantum process- ing units at large width and depth, arXiv preprint arXiv:2502.06471 (2025)

  30. [30]

    Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

    J. Preskill, Quantum computing in the NISQ era and beyond, Quantum2, 79 (2018)

  31. [31]

    Zimbor´ as, B

    Z. Zimbor´ as, B. Koczor, Z. Holmes, E.-M. Borrelli, A. Gily´ en, H.-Y. Huang, Z. Cai, A. Ac ´ ın, L. Aolita, L. Banchi,et al., Myths around quantum computation before full fault tolerance: What no-go theorems rule out and what they don’t, arXiv preprint arXiv:2501.05694 (2025)

  32. [32]

    Eisert and J

    J. Eisert and J. Preskill, Mind the gaps: The fraught road to quantum advantage, arXiv preprint arXiv:2510.19928 (2025)

  33. [33]

    Katabarwa, K

    A. Katabarwa, K. Gratsea, A. Caesura, and P. D. Johnson, Early fault-tolerant quantum computing, PRX quantum5, 020101 (2024). 19

  34. [34]

    Haghshenas, E

    R. Haghshenas, E. Chertkov, M. Mills, W. Kadow, S.- H. Lin, Y.-H. Chen, C. Cade, I. Niesen, T. Beguˇ si´ c, M. S. Rudolph,et al., Digital quantum magnetism at the frontier of classical simulations, arXiv preprint arXiv:2503.20870 (2025)

  35. [35]

    Google Quantum AI, Observation of constructive inter- ference at the edge of quantum ergodicity, Nature646, 825 (2025)

  36. [36]

    F. Alam, J. L. Bosse, I. ˇCepait˙ e, A. Chapman, L. Clinton, M. Crichigno, E. Crosson, T. Cubitt, C. Derby, O. Dow- inton,et al., Fermionic dynamics on a trapped-ion quan- tum computer beyond exact classical simulation, arXiv preprint arXiv:2510.26300 (2025)

  37. [37]

    Quantum advantage tracker, https://quantum- advantage-tracker.github.io/ (2025)

  38. [38]

    Acuaviva, D

    A. Acuaviva, D. Aguirre, R. Pe˜ na, and M. Sanz, Benchmarking quantum computers: Towards a stan- dard performance evaluation approach, arXiv preprint arXiv:2407.10941 (2024)

  39. [39]

    S. K. Seritan, A. Dhumuntarao, A. Q. Wilber-Gauthier, K. M. Rudinger, A. E. Russo, R. Blume-Kohout, A. D. Baczewski, and T. Proctor, Benchmarking quantum com- puters with any quantum algorithm, arXiv preprint arXiv:2508.05754 (2025)

  40. [40]

    Aaronson and D

    S. Aaronson and D. Gottesman, Improved simulation of stabilizer circuits, Physical Review A70, 052328 (2004)

  41. [41]

    B. M. Terhal and D. P. DiVincenzo, Classical simulation of noninteracting-fermion quantum circuits, Physical Re- view A65, 10.1103/physreva.65.032325 (2002)

  42. [42]

    Huang, R

    H.-Y. Huang, R. Kueng, and J. Preskill, Predicting many properties of a quantum system from very few measure- ments, Nature Physics16, 1050 (2020)

  43. [43]

    K. Wan, W. J. Huggins, J. Lee, and R. Babbush, Match- gate shadows for fermionic quantum simulation, Commu- nications in Mathematical Physics404, 629 (2023)

  44. [44]

    N. C. Rubin, J. Lee, and R. Babbush, Compressing many-body fermion operators under unitary constraints, Journal of Chemical Theory and Computation18, 1480 (2022)

  45. [45]

    A. D. C´ orcoles, J. M. Gambetta, J. M. Chow, J. A. Smolin, M. Ware, J. Strand, B. L. Plourde, and M. Stef- fen, Process verification of two-qubit quantum gates by randomized benchmarking, Physical Review A87, 030301 (2013)

  46. [46]

    D. C. McKay, S. Sheldon, J. A. Smolin, J. M. Chow, and J. M. Gambetta, Three-qubit randomized benchmarking, Physical Review Letters122, 200502 (2019)

  47. [47]

    Proctor, K

    T. Proctor, K. Rudinger, K. Young, E. Nielsen, and R. Blume-Kohout, Measuring the capabilities of quan- tum computers, Nature Physics18, 75–79 (2021)

  48. [49]

    Oszmaniec and Z

    M. Oszmaniec and Z. Zimbor´ as, Universal extensions of restricted classes of quantum operations, Physical Review Letters119, 220502 (2017)

  49. [50]

    Oszmaniec, N

    M. Oszmaniec, N. Dangniam, M. E. Morales, and Z. Zim- bor´ as, Fermion sampling: A robust quantum computa- tional advantage scheme using fermionic linear optics and magic input states, PRX Quantum3, 020328 (2022)

  50. [51]

    Y. Ibe, Y. Hirano, Y. Ozu, T. Kawakubo, and K. Fu- jii, Measurement-based fault-tolerant quantum com- putation on high-connectivity devices: A resource- efficient approach toward early ftqc, arXiv preprint arXiv:2510.18652 (2025)

  51. [52]

    European quantum computing benchmarks, https://gitlab.com/qcpi/eqcb (2025)

  52. [53]

    J. Chen, D. Ding, C. Huang, and L. Kong, Linear cross- entropy benchmarking with Clifford circuits, Physical Review A108, 10.1103/physreva.108.052613 (2023)

  53. [54]

    S. A. Merkel, T. Proctor, S. Ferracin, J. Hines, S. Bar- ron, L. C. G. Govia, and D. McKay, When Clifford bench- marks are sufficient: Estimating application performance with scalable proxy circuits (2025), arXiv:2503.05943

  54. [55]

    S. T. Flammia and Y.-K. Liu, Direct fidelity estimation from few pauli measurements, Physical Review Letters 106, 230501 (2011), arXiv:1104.4695 [quant-ph]

  55. [56]

    E. v. d. Berg, A simple method for sampling random Clifford operators (2020)

  56. [57]

    Bravyi and D

    S. Bravyi and D. Maslov, Hadamard-free circuits expose the structure of the Clifford group, IEEE Transactions on Information Theory67, 4546 (2021)

  57. [58]

    Winderl, Q

    D. Winderl, Q. Huang, A. M. van de Griend, and R. Ye- ung, Architecture-aware synthesis of stabilizer circuits from Clifford tableaus (2024), arXiv:2309.08972 [quant- ph]

  58. [59]

    T. J. Proctor, S. Seritan, K. Rudinger, E. Nielsen, R. Blume-Kohout, and K. C. Young, Scalable random- ized benchmarking of quantum computers using mirror circuits, Physical Review Letters129, 150502 (2022), arXiv:2112.09853 [quant-ph]

  59. [60]

    Zimbor´ as, R

    Z. Zimbor´ as, R. Zeier, M. Keyl, and T. Schulte- Herbr¨ uggen, A dynamic systems approach to fermions and their relation to spins, EPJ Quantum Technology1, 11 (2014)

  60. [61]

    Jordan and E

    P. Jordan and E. Wigner, ¨Uber das paulische ¨ aquivalenzverbot, Zeitschrift f¨ ur Physik47, 631 (1928)

  61. [62]

    Miller, Z

    A. Miller, Z. Zimbor´ as, S. Knecht, S. Maniscalco, and G. Garc ´ ıa-P´ erez, Bonsai algorithm: Grow your own fermion-to-qubit mappings, PRX quantum4, 030314 (2023)

  62. [63]

    Gidney, Stim: a fast stabilizer circuit simulator, Quan- tum5, 497 (2021)

    C. Gidney, Stim: a fast stabilizer circuit simulator, Quan- tum5, 497 (2021)

  63. [64]

    Van den Nest, Classical simulation of quantum com- putation, the gottesman–knill theorem, and slightly be- yond, arXiv preprint arXiv:0811.0898 (2008)

    M. Van den Nest, Classical simulation of quantum com- putation, the gottesman–knill theorem, and slightly be- yond, arXiv preprint arXiv:0811.0898 (2008)

  64. [65]

    Gluza, M

    M. Gluza, M. Kliesch, J. Eisert, and L. Aolita, Fidelity witnesses for fermionic quantum simulations, Phys. Rev. Lett.120, 190501 (2018)

  65. [66]

    Bak´ o, Z

    B. Bak´ o, Z. Kolarovszki, and Z. Zimbor´ as, Fermionic born machines: Classical training of quantum genera- tive models based on fermion sampling, arXiv preprint arXiv:2511.13844 (2025)

  66. [67]

    Maslov and B

    D. Maslov and B. Zindorf, Depth optimization of CZ, CNOT, and Clifford circuits, IEEE Transactions on Quantum Engineering3, 1–8 (2022)

  67. [68]

    Maslov and W

    D. Maslov and W. Yang, CNOT circuits need little help to implement arbitrary Hadamard-free Clifford transfor- mations they generate, npj Quantum Information9, 96 (2023). 20 Appendix A: Step-by-step description of the CL V benchmark protocol The Clifford Volume benchmark protocol shall start by determining the candidate values (or range of values) of n, fo...