Pith. sign in

REVIEW 1 major objections 6 minor 60 references

Classical Simulation and Design Frontiers for IBM's Doped Clifford Sampling Experiment

T0 review · 1 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A deterministic tensor-network sweep computes exact amplitudes for IBM's 70-qubit doped Clifford sampling experiment, and the same width formula maps where this circuit family stays classically easy to simulate.

desk verdict Clean, self-contained width-35 contraction of IBM's 70x70 doped Clifford circuit, run in 37 minutes; the one unresolved risk is float32 arithmetic at d=70. read the letter →

arxiv 2608.13110 v1 pith:MEF3GRL3 submitted 2026-08-13 quant-ph

classification quant-ph
keywords randomcircuitsamplingdopedCliffordcircuitstensornetworkcontractionwidthcross-entropybenchmarkingquantumadvantageGPUsimulationbrickwork
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that IBM's headline 70-qubit, 70-layer 'doped Clifford' sampling experiment can be checked classically by a deterministic tensor-network contraction that exploits geometry rather than Clifford structure. The central claim is a width formula: for an open chain whose entangling gates have operator Schmidt rank 2, one exact amplitude (no bond truncation) costs contraction width $\lceil d/2\rceil$, independent of the values or placement of one-qubit gates and of the 468 $T$ gates. Applied to the IBM instance, the formula gives a width-35 path whose largest intermediate tensor holds $2^{35}$ complex64 entries (256 GiB), and the authors report evaluating all 2051 published-output amplitude batches in 37.3 minutes on 32 nodes with eight H100 GPUs each. The resulting log-XEB estimate is 0.35034 with a 95% interval [0.29763, 0.40305], which they present as numerically compatible with IBM's reported fidelity lower bound under the standard Porter\textendash Thomas and scrambled-noise assumptions. The same machinery doubles as a design map, quantifying how entangler bond dimension, boundary geometry, depth, and memory capacity move the classical-simulation cost by orders of magnitude.

What carries the argument

The load-bearing object is the temporal-boundary sweep over the PEPS-like tensor network obtained by replacing each CZ gate (operator Schmidt rank 2) with two rank-3 tensors joined by a 2-dimensional bond, then absorbing all one-qubit gates into neighboring tensors. The circuit becomes a rectangular planar network of temporal width $\lceil d/2\rceil$ and spatial length $n$. Algorithm 1 sweeps a 'boundary tensor' along the spatial direction, reversing time on alternate rows, and contracts pairs of tensors whenever a direct absorption would exceed rank $\lceil d/2\rceil$; Proposition 3.1 proves this keeps every intermediate at width $\lceil d/2\rceil$. A batched variant (Algorithm 2) leaves $k \leq \lfloor d/4\rfloor$ output indices open and returns $2^k$ amplitudes at the same width, which is what turns amplitudes into samples. Ratcatcher supplies the optimality witness on tested instances, and a hyper-optimized path search provides the comparison baseline and the slicing-overhead measurements.

What would settle it

Evaluate one or more of the 2051 amplitude batches at $d=70$ using complex128 arithmetic on enough nodes to hold the 512 GiB tensor, or using an independent contraction path, and compare with the published complex64 probabilities: the authors' own $d\leq 56$ validation gives a complex64-vs-complex128 total variation below $9\times10^{-7}$, so any larger discrepancy at depth 70 would falsify the numerical claim.

Watch

Extended reading notes

Core claim

The paper's discovery is that an open-boundary, one-dimensional brickwork circuit whose entangling gates have operator Schmidt rank 2 possesses an exact, deterministic contraction path with width $\lceil d/2\rceil$ for depth $d$, and that this path is width-minimal on every tested instance, as certified by the Ratcatcher carving-width algorithm. For the IBM circuit with $n=d=70$, this yields a width-35 schedule, a 256 GiB largest-intermediate tensor, and a measured 37.3-minute makespan on 32 eight-H100 nodes for all 2051 published amplitude batches. The authors use these exact probabilities to compute log-XEB = 0.35034 with 95% interval [0.29763, 0.40305], which they interpret, under the Porter\textendash Thomas and scrambled-noise assumptions, as numerically compatible with IBM's fidelity lower bound of 0.284. They further show that the favorable width is a geometric property: increasing the entangler bond dimension from 2 to 4, or closing the chain into a ring, roughly doubles the width, while slicing one bit below the natural width incurs a very large overhead. Consequently, they present the method both as an independent diagnostic of the experimental output and as a quantitative tool for designing future doped Clifford sampling experiments.

Load-bearing premise

The whole schedule depends on the IBM circuit being exactly an open one-dimensional chain of operator-Schmidt-rank-2 CZ gates, and on complex64 rounding at depth 70 staying as small as it was measured to be through depth 56; if any entangler had rank 4, the chain were closed, or rounding grew, the width or the amplitudes would no longer match the experiment.

Editorial extensions

If this is right

  • For any open-boundary, bond-dimension-2 brickwork circuit with $n\geq 4$ and $7\leq d\leq 2n$, one exact amplitude can be evaluated with contraction width $\lceil d/2\rceil$, time $O(nd2^{\lceil d/2\rceil})$, and space $O(2^{\lceil d/2\rceil})$; changing single-qubit gates or adding and moving $T$ gates does not change these counts.
  • The 2051 published output bitstrings of IBM's 70-qubit experiment are now individually calibrated: their exact ideal probabilities give log-XEB = 0.35034 with 95% interval [0.29763, 0.40305], an independent check of the reported fidelity scale.
  • For the IBM instance, the fidelity-weighted classical sampling workload (583 exact contractions for fidelity 0.284) projects to about 10.6 minutes on 32 eight-H100 nodes, while the measured full verification run took 37.3 minutes.
  • Memory, not arithmetic alone, is the practical bottleneck: slicing the width from 35 to 34 already creates many independent subtasks with very large cost overhead, and the aggregate-memory frontier for this circuit family sits near depth $d\approx 86$ with 1024 H100 GPUs.
  • Circuit design can tune hardness against this method: replacing CZ with a rank-4 entangler doubles the width, and closing the chain raises the width to $\min(n,d)$, though both changes carry physical-gate and spacetime-code costs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the width formula holds beyond the tested range, the cliff for this circuit family sits near depth 86 for the open $\chi=2$ geometry under 1024-H100 memory assumptions; any 'hard' instance in this family should be checked against that ceiling before invoking quantum advantage.
  • The independence of the schedule from $T$-gate count suggests that, for open one-dimensional chains, the practical hardness of doped Clifford circuits comes from geometry and entangler rank rather than from magic; closing the chain or using rank-4 entanglers is the natural hardening route, though it may also break the efficient spacetime-code verification the experiment relies on.
  • A direct testable extension is to run the same sweep on other rank-2 entanglers, such as CNOT or iSWAP-like gates, to see whether the width formula survives when the entangler is not diagonal; the paper's construction suggests it should as long as the network remains an open brickwork.
  • The projected 10.6-minute classical sampling time for the fidelity-weighted workload, compared with the experiment's 16.1-minute sampling run, suggests this specific 70-qubit instance sits close to parity on the sampling task itself with a 32-node classical cluster; whether that counts as 'advantage' depends on like-for-like resource and fidelity accounting.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 6 minor

Summary. Manabe, Gu, and Pan present a deterministic temporal-boundary tensor network contraction algorithm for open-boundary, one-dimensional brickwork circuits whose entangling gates have operator Schmidt rank 2. For an n-qubit circuit of depth d, the algorithm evaluates an exact amplitude with contraction width ceil(d/2) and cost O(nd 2^ceil(d/2)), and the authors use Ratcatcher to certify that no smaller width exists for the n=70 instances tested. A batched variant with k open output bits retains the same width for 1<=k<=floor(d/4). The paper applies the method to IBM's 70-qubit, 70-layer doped Clifford circuit with 468 T gates, computing all 2051 published-output amplitude batches in 37.3 minutes on 32 eight-H100 nodes. The resulting probabilities give log-XEB = 0.35034 with a 95% confidence interval [0.29763,0.40305], which the authors interpret as numerically compatible with IBM's fidelity lower bound under Porter-Thomas and scrambled-noise assumptions. The paper also analyzes how slicing, entangling-gate bond dimension, and periodic boundary conditions affect simulation cost, producing a classical-simulatability map for circuit design.

Significance. If the results hold, this is a significant advance in classical simulation of a specific doped Clifford RCS experiment. The width theorem is parameter-free and machine-verifiable via the Ratcatcher certification, and the independence of the dense contraction cost from T-gate count and placement is a clean structural insight. The reported execution is a concrete resource comparison against IBM's hardware experiment, and the public release of amplitudes and contraction paths supports reproducibility. The main unresolved risk is the use of complex64 arithmetic at the d=70 operating point, which is validated only up to d=56; a spot-check at d=70 would make the numerical claim fully secure.

major comments (1)
  1. [§5.1–5.2, Fig. 7, Table 1] The d=70 production calculation uses complex64 arithmetic, but the complex64-vs-complex128 validation in Fig. 7 (right panel) is only reported through d=56. The d=70 operating point has a 2^35-entry largest intermediate and roughly 4,900 contraction steps, and the paper does not report any complex128 comparison at depths above 56 nor a forward-error bound. Because the headline numerical results (log-XEB and its 95% interval) are computed from these d=70 amplitudes, the finite-precision error at the operating point is an unverified quantity that directly bears on the numerical claim. A complex128 spot-check for a subset of the 2051 batches at d=70 would settle this risk; the 512 GiB complex128 payload of the largest intermediate still fits on eight 80 GB H100 GPUs.
minor comments (6)
  1. [Abstract and §1] The abstract's phrase '256 times smaller than IBM's estimation' is ambiguous because IBM's memory-constrained cotengra estimate (2^30 scalars) is actually smaller than the paper's 2^35-entry intermediate; the qualifications in Section 1 about differing tensorization and memory-accounting conventions should be carried into the abstract.
  2. [§3.1, Proposition 3.1] The proof of Proposition 3.1 is quite terse; it analyzes two local configurations but does not state the full sweep order or an inductive invariant. Since the width bound is a central claim, an explicit inductive proof (or a more detailed pseudocode-level invariant) would make the argument easier to verify.
  3. [§3.1 and §4.3] The proofs of Proposition 3.1 and Proposition 4.1 use the word 'rank' to mean the number of tensor legs, not the multilinear rank; this is clear in context, but a one-line definition at first use would prevent confusion for readers who interpret 'rank' differently.
  4. [§5.2 and Table 1] The text reports a makespan of 37 min 16 s while the abstract and Table 1 round to 37.3 minutes; the rounding convention should be stated or the numbers made consistent.
  5. [Figure 7] The right panel of Figure 7 is labeled 'FP32 error,' while the text describes the total variation distance between complex64 and complex128 conditional distributions; relabeling the axis as 'TV distance (complex64 vs complex128)' would align the figure with the text.
  6. [Data availability] Releasing the contraction schedules and amplitudes on Zenodo is helpful; making the simulator code available as well would further strengthen reproducibility of the numerical results.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the width theorem and log-XEB diagnostic are self-contained, with only a non-load-bearing self-citation to the authors' distributed-contraction framework.

full rationale

The central derivation is self-contained. Proposition 3.1's width bound W(P_con)=ceil(d/2) follows from the open-boundary brickwork geometry, the operator-Schmidt-rank-2 CZ decomposition, and Algorithm 1's explicit sweep, not from any fitted parameter; the lower bound is certified by Ratcatcher, an external algorithm for carving width, with the constructive path used only as an upper-bound witness. Proposition 4.1 extends the same argument to batched amplitudes with k open output bits and again uses Ratcatcher/Ratcon only for independent width certification. The log-XEB estimate 0.35034 is computed from exact ideal probabilities evaluated for IBM's published bitstrings; IBM's data and fidelity lower bound are used as external benchmarks, not as inputs to the contraction or to any fitted quantity. The fidelity-weighted 583-contraction and 10.6-minute projection is explicitly labeled a resource projection, not a measured prediction. The only self-citation that plays any role in the implementation is Ref. [53], the authors' earlier multi-GPU tensor-contraction framework, and it is used as an engineering method rather than as justification of the mathematical or statistical claim; no uniqueness or optimality assertion is imported from it. The paper's own caveats, that 'exact' means no bond truncation and that complex64 finite-precision arithmetic was validated only through d=56 with TV below 9e-7, are acknowledged numerical limitations rather than circular reductions. No derivation step was found that equates a claimed prediction to its input by construction.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central result has no fitted parameters: the width bound is derived from geometry and gate topology. The axioms are standard graph-theoretic equivalences, the domain assumption that the IBM circuit is exactly the stated open-chain CZ brickwork, and the interpretive assumptions for log-XEB and finite-precision arithmetic. No new particles, forces, or mechanical entities are introduced.

free parameters (1)
  • batch size k = 8
    The IBM run fixes k=8 open output bits, giving 256 amplitudes per contraction. It is chosen by hand within the proved range k <= floor(d/4) and is not fitted to data; it affects throughput but not the width bound.
assumptions (4)
  • standard math Contraction trees of a tensor network correspond to carving decompositions of its weighted planar graph, so Ratcatcher's polynomial-time carving-width decision certifies minimum contraction width for these planar networks.
    Invoked in Section 3.2 via Refs [44,45] to translate Ratcatcher outputs into certified minimum widths W*. The equivalence is standard but not proved in this paper.
  • domain assumption Every entangling gate in the considered family is exactly a nearest-neighbor CZ gate on an open chain with the alternating brickwork pattern, giving operator Schmidt rank 2 and a planar strip topology.
    Definition of the circuit family in Section 2.1, taken from IBM's Ref [1]. The width formula and the execution both rely on this exact geometry.
  • domain assumption The 2051 experimental bitstrings can be treated as independent draws from a Porter-Thomas-like distribution under a scrambled-noise model, allowing log-XEB to be interpreted as a fidelity proxy.
    Used in Section 5.2 to compare log-XEB 0.35034 with IBM's 0.284 lower bound. The paper states the assumption explicitly; it is an interpretive assumption, not part of the amplitude calculation.
  • domain assumption Finite-precision complex64 arithmetic reproduces exact complex128 probabilities to within negligible total variation distance at d=70.
    The comparison in Section 5.1 shows TV distance below 9e-7 through d=56; the d=70 run uses complex64 without a direct same-size complex128 check. This is a practical numerical assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Classical Simulation and Design Frontiers for IBM's Doped Clifford Sampling Experiment." pith.science (2026). https://pith.science/paper/MEF3GRL3

@misc{pith2026260813110,
  author       = {Pith},
  title        = {Pith review of: Classical Simulation and Design Frontiers for IBM's Doped Clifford Sampling Experiment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MEF3GRL3}},
  note         = {Machine review of arXiv:2608.13110}
}
abstract

We classically simulate the IBM doped Clifford random circuit sampling experiment, comprising $70$ qubits, $70$ entangling layers, and $468$ inserted $T$ gates. A deterministic temporal-boundary tensor network contraction approach is specifically designed to tackle such open-boundary one-dimensional brickwork circuits with operator-Schmidt-rank-$2$ entangling gates. For an $n$-qubit circuit of depth $d$, the resulting unsliced path evaluates an exact amplitude with contraction width $\lceil d/2\rceil$; Ratcatcher calculations certify that no smaller width is possible for the tested instances. Because one-qubit gates are absorbed without changing the network topology, the width and dense scheduled contraction cost are independent of their values and of the number and placement of $T$ gates. For the IBM instance, its largest intermediate tensor contains $2^{35}$ complex64 entries (256 times smaller than IBM's estimation), corresponding to a tensor payload of $256$ GiB, and is distributed across eight GPUs within a node. Using 32 nodes, with eight NVIDIA H100 GPUs per node, we completed all 2051 amplitude batches corresponding to IBM's published output bitstrings in 37.3 minutes. The resulting probabilities yield a log-XEB estimate of $0.35034$ with a 95\% interval of $[0.29763,0.40305]$. Under the Porter--Thomas and scrambled-noise assumptions, this is numerically compatible with IBM's fidelity lower bound; separately, fidelity-weighted resource accounting projects a 583-contraction workload with a 10.6-minute makespan on the same 32 nodes. More broadly, the approach provides a practical diagnostic for experimental outputs and a quantitative tool for designing future doped Clifford sampling experiments.

Figures

Figures reproduced from arXiv: 2608.13110 by the authors.

Figure 1
Figure 1. Mapping an open-boundary one-dimensional brickwork circuit to its tensor network [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Deterministic contraction path for exact amplitude contraction. The numbers indicate [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The two possible paired steps. The purple region contains the tensors already absorbed [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Total contraction cost (left) and contraction width (right) for the [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Three-region decomposition for batched contraction with [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Resource scaling with the number of open output indices [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Single-GPU execution for the n = 70 and k = 8 contractions. The left and center panels show the contraction time and the peak device-memory allocation above the initialized baseline, respectively. The open diamond marks the analytically determined 256-GiB size of the l…
Figure 8
Figure 8. Figure 8: Log-XEB estimate for the 2051 experimental bitstrings. The blue line and shaded [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Slicing overhead for brickwork tensor networks with [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Effect of the entangling-bond dimension on the total contraction cost (left) and [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: Effect of closing the spatial boundary on the total contraction cost (left) and contraction [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Classical simulatability map for computing one exact amplitude of the one-dimensional [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 47 canonical work pages

  1. [1]

    Sampling hard circuits with verifiably high fidelity

    S. Martiel, J.-U. Chung, A. Seif, S. Ghosh, I. Hincks, A. Deshpande, B. Fefferman, J. M. Gambetta, and A. Javadi-Abhari, Sampling hard circuits with verifiably high fidelity, (2026), arXiv:2607.25941

  2. [2]

    P. W. Shor, Polynomial-Time Algorithms for Prime Factorization and Discrete Logarithms on a Quantum Computer, SIAM Journal on Computing26, 1484 (1997)

  3. [3]

    L. K. Grover, A fast quantum mechanical algorithm for database search, in Proceedings of the twenty-eighth annual ACM symposium on Theory of computing - STOC ’96 (1996), pp. 212–219

  4. [4]

    R. P. Feynman, Simulating physics with computers, International Journal of Theoretical Physics21, 467 (1982)

  5. [5]

    Farhi, J

    E. Farhi, J. Goldstone, and S. Gutmann, A Quantum Approximate Optimization Algorithm, (2014), arXiv:1411.4028

  6. [6]

    Gidney and M

    C. Gidney and M. Eker˚ a, How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits, Quantum5, 433 (2021)

  7. [7]

    M. E. Beverland, P. Murali, M. Troyer, K. M. Svore, T. Hoefler, V. Kliuchnikov, G. H. Low, M. Soeken, A. Sundaram, and A. Vaschillo, Assessing requirements to scale to practical quantum advantage, (2022), arXiv:2211.07629

  8. [8]

    A. W. Harrow and A. Montanaro, Quantum Computational Supremacy, Nature549, 203 (2017)

Show all 60 references
  1. [9]

    Aaronson and L

    S. Aaronson and L. Chen, Complexity-Theoretic Foundations of Quantum Supremacy Experiments, in 32nd Computational Complexity Conference (CCC 2017), Vol. 79, Leibniz International Proceedings in Informatics (LIPIcs) (2017), 22:1–22:67

  2. [10]

    Bouland, B

    A. Bouland, B. Fefferman, C. Nirkhe, and U. Vazirani, On the complexity and verification of quantum random circuit sampling, Nature Physics15, 159 (2019)

  3. [11]

    Hangleiter and J

    D. Hangleiter and J. Eisert, Computational Advantage of Quantum Random Sampling, Reviews of Modern Physics95, 035001 (2023)

  4. [12]

    Boixo, S

    S. Boixo, S. V. Isakov, V. N. Smelyanskiy, R. Babbush, N. Ding, Z. Jiang, M. J. Bremner, J. M. Martinis, and H. Neven, Characterizing Quantum Supremacy in Near-Term Devices, Nature Physics14, 595 (2018)

  5. [13]

    Arute, K

    F. Arute, K. Arya, R. Babbush, D. Bacon, J. C. Bardin, R. Barends, R. Biswas, S. Boixo, F. G. S. L. Brandao, D. A. Buell, et al., Quantum supremacy using a programmable superconducting processor, Nature574, 505 (2019). 21

  6. [14]

    Wu, W.-S

    Y. Wu, W.-S. Bao, S. Cao, F. Chen, M. -C. Chen, X. Chen, T. -H. Chung, H. Deng, Y. Du, D. Fan, et al., Strong quantum computational advantage using a superconducting quantum processor, Physical Review Letters127, 180501 (2021)

  7. [15]

    Q. Zhu, S. Cao, F. Chen, M. -C. Chen, X. Chen, T. -H. Chung, H. Deng, Y. Du, D. Fan, M. Gong, et al., Quantum computational advantage via 60-qubit 24-cycle random circuit sampling, Science Bulletin67, 240 (2022)

  8. [16]

    Morvan, B

    A. Morvan, B. Villalonga, X. Mi, S. Mandr` a, A. Bengtsson, P. V. Klimov, Z. Chen, S. Hong, C. Erickson, I. K. Drozdov, et al., Phase transitions in random circuit sampling, Nature 634, 328 (2024)

  9. [17]

    D. Gao, D. Fan, C. Zha, J. Bei, G. Cai, J. Cai, S. Cao, X. Zeng, F. Chen, J. Chen, et al., Establishing a New Benchmark in Quantum Computational Advantage with 105-qubit Zuchongzhi 3.0 Processor, Physical Review Letters134, 090601 (2025)

  10. [18]

    Zlokapa, B

    A. Zlokapa, B. Villalonga, S. Boixo, and D. A. Lidar, Boundaries of quantum supremacy via random circuit sampling, npj Quantum Information9, 36 (2023)

  11. [19]

    Huang, F

    C. Huang, F. Zhang, M. Newman, X. Ni, D. Ding, J. Cai, X. Gao, T. Wang, F. Wu, G. Zhang, et al., Efficient parallelization of tensor network contraction for simulating quantum computation, Nature Computational Science1, 578 (2021)

  12. [20]

    Villalonga, S

    B. Villalonga, S. Boixo, B. Nelson, C. Henze, E. Rieffel, R. Biswas, and S. Mandr` a, A flexible high-performance simulator for verifying and benchmarking quantum circuits implemented on real hardware, npj Quantum Information5, 86 (2019)

  13. [21]

    Villalonga, D

    B. Villalonga, D. Lyakh, S. Boixo, H. Neven, T. S. Humble, R. Biswas, E. G. Rieffel, A. Ho, and S. Mandr` a, Establishing the Quantum Supremacy Frontier with a 281 Pflop/s Simulation, Quantum Science and Technology5, 034003 (2020)

  14. [22]

    Y. Liu, Y. Chen, C. Guo, J. Song, X. Shi, L. Gan, W. Wu, W. Wu, H. Fu, X. Liu, et al., Verifying Quantum Advantage Experiments with Multiple Amplitude Tensor Network Contraction, Physical Review Letters132, 030601 (2024)

  15. [23]

    X. Gao, M. Kalinowski, C. -N. Chou, M. D. Lukin, B. Barak, and S. Choi, Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage, PRX Quantum5, 010334 (2024)

  16. [24]

    Delfosse and A

    N. Delfosse and A. Paetznick, Spacetime codes of Clifford circuits, (2023), arXiv: 2304. 05943

  17. [25]

    Martiel and A

    S. Martiel and A. Javadi-Abhari, Low-overhead error detection with spacetime codes, (2025), arXiv:2504.15725

  18. [26]

    De Raedt, K

    K. De Raedt, K. Michielsen, H. De Raedt, B. Trieu, G. Arnold, M. Richter, T. Lippert, H. Watanabe, and N. Ito, Massively parallel quantum computer simulator, Computer Physics Communications176, 121 (2007)

  19. [27]

    H¨ aner and D

    T. H¨ aner and D. S. Steiger, 0.5 Petabyte Simulation of a 45-Qubit Quantum Circuit, in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Nov. 2017), pp. 1–10

  20. [28]

    H. D. Raedt, F. Jin, D. Willsch, M. Nocon, N. Yoshioka, N. Ito, S. Yuan, and K. Michielsen, Massively parallel quantum computer simulator, eleven years later, Computer Physics Communications237, 47 (2019)

  21. [29]

    Y. Zhou, E. M. Stoudenmire, and X. Waintal, What Limits the Simulation of Quantum Computers?, Physical Review X10, 041038 (2020)

  22. [30]

    Bravyi and D

    S. Bravyi and D. Gosset, Improved classical simulation of quantum circuits dominated by Clifford gates, Physical Review Letters116, 250501 (2016). 22

  23. [31]

    Bravyi, D

    S. Bravyi, D. Browne, P. Calpin, E. Campbell, D. Gosset, and M. Howard, Simulation of quantum circuits by low-rank stabilizer decompositions, Quantum3, 181 (2019)

  24. [32]

    Kissinger and J

    A. Kissinger and J. van de Wetering, Simulating quantum circuits with ZX-calculus reduced stabiliser decompositions, Quantum Science and Technology7, 044001 (2022)

  25. [33]

    A. F. Mello, A. Santini, and M. Collura, Hybrid Stabilizer Matrix Product Operator, Physical Review Letters133, 150604 (2024)

  26. [34]

    Liu and B

    Z. Liu and B. K. Clark, Classical simulability of Clifford + T circuits with Clifford-augmented matrix product states, Physical Review Research8, 023116 (2026)

  27. [35]

    B. A. Chase and F. Labib, Clifft: Fast Exact Simulation of Near-Clifford Quantum Circuits, arXiv,10.48550/arXiv.2604.27058(2026), arXiv:2604.27058

  28. [36]

    K. Noh, L. Jiang, and B. Fefferman, Efficient classical simulation of noisy random quantum circuits in one dimension, Quantum4, 318 (2020)

  29. [37]

    Aharonov, X

    D. Aharonov, X. Gao, Z. Landau, Y. Liu, and U. Vazirani, A polynomial-time classical algorithm for noisy random circuit sampling, in Proceedings of the 55th Annual ACM Symposium on Theory of Computing (June 2023), pp. 945–957

  30. [38]

    I. L. Markov and Y. Shi, Simulating quantum computation by contracting tensor networks, SIAM Journal on Computing38, 963 (2008)

  31. [39]

    M. C. Ba˜ nuls, M. B. Hastings, F. Verstraete, and J. I. Cirac, Matrix Product States for dynamical simulation of infinite chains, Physical Review Letters102, 240603 (2009)

  32. [40]

    Lerose, M

    A. Lerose, M. Sonner, and D. A. Abanin, Influence Matrix Approach to Many-Body Floquet Dynamics, Physical Review X11, 021040 (2021)

  33. [41]

    Pan and P

    F. Pan and P. Zhang, Simulation of Quantum Circuits Using the Big-Batch Tensor Network Method, Physical Review Letters128, 030501 (2022)

  34. [42]

    F. Pan, K. Chen, and P. Zhang, Solving the Sampling Problem of the Sycamore Quantum Circuits, Physical Review Letters129, 090502 (2022)

  35. [43]

    P. D. Seymour and R. Thomas, Call Routing and the Ratcatcher, Combinatorica14, 217 (1994)

  36. [44]

    O’Gorman, Parameterization of Tensor Network Contraction, in 14th Conference on the Theory of Quantum Computation, Communication and Cryptography (TQC 2019), Vol

    B. O’Gorman, Parameterization of Tensor Network Contraction, in 14th Conference on the Theory of Quantum Computation, Communication and Cryptography (TQC 2019), Vol. 135, Leibniz International Proceedings in Informatics (LIPIcs) (2019), 10:1–10:19

  37. [45]

    Jakes-Schauer, D

    J. Jakes-Schauer, D. Anekstein, and P. Wocjan, Carving-width and contraction trees for tensor networks, (2019), arXiv:1908.11034

  38. [46]

    DeCross, R

    M. DeCross, R. Haghshenas, M. Liu, E. Rinaldi, J. Gray, Y. Alexeev, C. H. Baldwin, J. P. Bartolotta, M. Bohn, E. Chertkov, et al., Computational Power of Random Quantum Circuits in Arbitrary Geometries, Physical Review X15, 021052 (2025)

  39. [47]

    Verstraete, V

    F. Verstraete, V. Murg, and J. I. Cirac, Matrix Product States, Projected Entangled Pair States, and variational renormalization group methods for quantum spin systems, Advances in Physics57, 143 (2008)

  40. [48]

    Gray and S

    J. Gray and S. Kourtis, Hyper-optimized tensor network contraction, Quantum5, 410 (2021)

  41. [49]

    I. V. Hicks, Planar Branch Decompositions II: The Cycle Method, INFORMS Journal on Computing17, 413 (2005)

  42. [50]

    I. L. Markov, A. Fatima, S. V. Isakov, and S. Boixo, Massively Parallel Approximate Simulation of Hard Quantum Circuits, in 2020 57th ACM/IEEE Design Automation Conference (DAC) (July 2020), pp. 1–6. 23

  43. [51]

    Kalachev, P

    G. Kalachev, P. Panteleev, P. Zhou, and M. -H. Yung, Classical Sampling of Random Quantum Circuits with Bounded Fidelity, arXiv (2021), arXiv:2112.15083

  44. [52]

    Bayraktar, A

    H. Bayraktar, A. Charara, D. Clark, S. Cohen, T. Costa, Y. -L. L. Fang, Y. Gao, J. Guan, J. Gunnels, A. Haidar, et al., cuQuantum SDK: A High-Performance Library for Accelerating Quantum Science, in 2023 IEEE International Conference on Quantum Computing and Engineering (QCE),...

  45. [53]

    F. Pan, H. Gu, P. Springer, and X. Li, Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs, (2026), arXiv:2606.01852

  46. [54]

    NVIDIA,Multi-Process support - cuTENSORMp (Beta) — cuTENSOR, NVIDIA cuTEN- SOR Documentation, https://docs.nvidia.com/cuda/cutensor/latest/user_guide_ cutensorMp.html

  47. [55]

    R. Fu, Z. Su, H.-S. Zhong, X. Zhao, J. Zhang, F. Pan, P. Zhang, X. Zhao, M.-C. Chen, C.-Y. Lu, et al., Surpassing Sycamore: Achieving Energetic Superiority Through System-Level Circuit Simulation, in SC24: International Conference for High Performance Computing, Networking, St...

  48. [56]

    X.-H. Zhao, H. -S. Zhong, F. Pan, Z. -H. Chen, R. Fu, Z. Su, X. Xie, C. Zhao, P. Zhang, W. Ouyang, et al., Leapfrogging Sycamore: harnessing 1432 GPUs for 7×faster quantum random circuit sampling, National Science Review12, nwae317 (2025)

  49. [57]

    C. Oh, M. Liu, Y. Alexeev, B. Fefferman, and L. Jiang, Classical algorithm for simulating experimental Gaussian boson sampling, Nature Physics20, 1461 (2024)

  50. [58]

    Tindall, M

    J. Tindall, M. Fishman, E. M. Stoudenmire, and D. Sels, Efficient Tensor Network Simulation of IBM’s Eagle Kicked Ising Experiment, PRX Quantum5, 010308 (2024)

  51. [59]

    Tindall, A

    J. Tindall, A. Mello, M. Fishman, M. Stoudenmire, and D. Sels, Dynamics of disordered quantum systems with two- and three-dimensional tensor networks, Science392, 868 (2026)

  52. [60]

    Rausch, S

    R. Rausch, S. Singh, S. S. Jahromi, A. Kshetrimayum, and R. Orus, Pushing the Classical Frontier of 1D Fermi-Hubbard Quench Dynamics Beyond Current Quantum Simulations, (2026), arXiv:2606.04771. 24

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.