REVIEW 4 major objections 3 minor 2 cited by
Sampling hard circuits with verifiably high fidelity
T0 review · 4 major / 3 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read A 70-qubit doped Clifford circuit achieves a certified fidelity lower bound of 0.284 while remaining classically hard to simulate, by placing T gates only where they commute with the error-detection checks.
desk verdict A genuinely new error-detected sampling protocol with a real 97-qubit demonstration; the headline fidelity bound is conditional on a noise-model assumption that is narrower than the paper's stated generality. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the spacetime Pauli check: a Pauli operator supported on circuit wires whose back-propagation through the Clifford circuit is the identity (or a stabilizer of the input state), implemented via an ancilla measurement. The second ingredient is T-doping at check-commuting locations: each T gate is placed on a wire where a Z rotation commutes with all checks' back-cumulants, preserving the accepted fault set and syndrome distribution. The argument is carried by the identity F2 ≥ F1 − Pr(E∈H|E∈A), which turns the hard-to-measure doped fidelity into a difference of two efficiently accessible quantities, plus a counting/mixing argument showing that harmless faults are rar
What would settle it
Run the same 70-qubit DCS protocol while intentionally injecting a controlled Z-only error on a subset of T gates (or a correlated noise aligned with a measured check), and compare the measured fidelity of the doped state (via XEB at intermediate doping or DFE at low doping) against the certificate's predicted lower bound: if the fidelity drops by more than the certificate allows while the syndrome distribution is unchanged, the certificate is falsified.
Extended reading notes
Core claim
The central claim is that for a Clifford circuit C1 equipped with spacetime Pauli checks, and a circuit C2 obtained by adding T gates at wires that commute with all checks, the post-selected fidelity of the doped state satisfies F2 ≥ F1 − Pr(E∈H|E∈A), where F1 is the Clifford state's fidelity, A is the set of accepted (syndrome-passing) Pauli faults, and H is the subset of accepted faults that stabilize the Clifford output. The accepted fault set and post-selection probability are identical in the two circuits because the T gates commute with the checks. Since a harmless fault in a random linear-depth Clifford circuit is unlikely, the loss term Pr(E∈H|E∈A) is small; for the specific 70×70 in
Load-bearing premise
The certificate assumes that, after randomized compiling, the physical noise is exactly described by the Pauli-noise family swept in the Monte Carlo estimate of Pr(E∈H|E∈A); if the real noise contains components outside that family—such as globally correlated errors aligned with the spacetime stabilizers, or pure Z errors on the T gates—the syndromes can remain unchanged while the doped state's fidelity falls by more than the estimated 0.013.
Editorial extensions
If this is right
- Sampling experiments can now simultaneously be classically hard, error-suppressed through post-selection, and accompanied by a device-dependent fidelity certificate that needs only Pauli noise after randomized compiling.
- Because the certificate is instance-specific and does not require direct simulation of the doped circuit, it can be applied to larger or deeper circuits than currently simulable.
- The construction gives a systematic method for promoting a stabilizer state to a magic state while retaining the error-detection structure of the spacetime code, providing a bridge from near-term error detection to fault-tolerant quantum advantage.
- The 70×70, 468-T-gate instance is estimated to be infeasible for tensor-network and stabilizer-based classical simulations, and its effective CZ error rate after post-selection is 1.8×10^{-4}, making noisy-classical-simulation attacks impractical.
Reading between the lines
- If the harmless-fault probability estimate transfers to other Clifford skeletons, the certificate could become a generic template for verifying magic-state sampling on any device with good syndrome extraction; the key requirement is a random Clifford backbone with low-weight spacetime checks.
- The certificate as stated leaves a gap for adversarial noise that is invisible to the checks (e.g., Z-only errors on T gates or noise aligned with the code's stabilizers). A direct experimental test at 468 T gates, for instance by comparing measured syndrome-conditional fidelities at intermediate T counts, would determine whether the assumed noise family contains the real device noise.
- Because the certificate is numerical and noise-model-dependent, a fully device-independent verification remains out of reach; however, the structure suggests that planting circuit-specific secrets (peaked circuits) could be combined with DCS to give a verifiable advantage without any noise assumptions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a protocol, doped Clifford sampling (DCS), for quantum sampling experiments that are intended to be classically hard and at the same time verifiable: a random Clifford circuit is encoded in a spacetime code, post-selected on syndrome outcomes, and then doped with T gates that commute with the code checks. The central theoretical result is that the post-selected fidelity F2 of the doped state is lower-bounded by F2 ≥ F1 − Pr(E∈H|E∈A), where F1 is the fidelity of the Clifford state (measured by DFE) and the loss term is the conditional probability that an accepted Pauli fault is harmless in the Clifford circuit. The authors report a 70-qubit, depth-70 Clifford circuit with 468 T gates encoded into 97 physical qubits, F1 = 0.32(1), and a claimed 95% confidence lower bound F2 ≥ 0.284. They also provide asymptotic hardness arguments for the DCS ensemble and numerical studies of classical simulation difficulty at the experimental size.
Significance. If the numerical certificate were made rigorous, this would be a significant contribution: it combines a structured sampling proposal with code-based error detection and a fidelity estimate that does not rely on a classical simulation proxy for the hard state. The core inequality is clean, the DFE measurement of F1 is noise-model independent, and the use of virtual-Z T gates is a strong physical argument that doping itself introduces no additional gate error. The low- and medium-doping validation experiments (Fig. 4) provide meaningful empirical support. However, the headline 0.284 bound is conditional on a Monte Carlo estimate over a chosen Pauli-noise family, and the coherent-noise residual from using finitely many twirls is never bounded. These gaps are quantitative and affect the advertised 95% confidence statement.
major comments (4)
- [Main text Fig. 3c; Supplementary S2.3] The numerical loss Pr(E∈H|E∈A) ≤ 0.013(1) is obtained by Monte Carlo simulation 'sweeping various Pauli noise strengths and polarizations' (Fig. 3c). This is not a worst-case bound over the circuit-independent Pauli noise models allowed by Theorem S1. The asymptotic O(n^{-c}) argument in S2.3 is also not a finite-size certificate for the 70×70 instance. Since the experimental claim F2 ≥ 0.284 at 95% confidence relies directly on this number, the bound should either be proven as a true maximum over the claimed noise family, or be replaced by an interval computed from measured calibration data with rigorous uncertainty propagation. As written, the 95% confidence statement covers Monte Carlo sampling error within one parameterized family, not model uncertainty.
- [Supplementary S2.4.1; S6.4] Exact Pauli twirling removes the coherent interference term Γ_coh in Eq. (S23), but the experiment uses only 50 twirl instances (S6.4). The residual |Γ_coh| after finite twirling is never bounded. Consequently, the equality of post-selection rates and the reduction to the stochastic Pauli bound are only approximate. The paper needs a quantitative estimate of the finite-twirl residual—for example, a variance bound over the 50 randomizations actually used—and must include this residual in the claimed confidence interval if the certificate is to hold under coherent noise.
- [Supplementary S2.5] The adversarial scenarios listed in S2.5—Z-only errors on T gates and noise concentrated on the spacetime stabilizer group—are exactly cases where the syndrome distributions can remain unchanged while F2 drops by more than the estimated 0.013. The paper dismisses them as 'not physically realistic' but provides no quantitative argument or experimental test that excludes them. Given the stated goal of substantially weaker noise assumptions than XEB, these scenarios must be addressed explicitly, either by a measurable calibration check (e.g., injecting and amplifying such errors in a controlled way) or by a revised certificate that states the additional assumption and its experimental support.
- [Supplementary S8.3, Theorem S6] The average-case hardness theorem is proved for the interpolated distribution eD_{α*,K,Θ}, not for the uniform-rotation DCS ensemble that is actually sampled. The transition from exact computation of the truncated polynomial q_{β,K} on this interpolated distribution to approximate sampling from the original ensemble is not established; S8.3.2 explicitly notes that the required additive error 2^{-n/poly(n)} for 1/poly(n)-TV sampling remains open. The main-text claim that 'no classical sampler exists' under 'two appropriate conjectures' therefore overstates what Theorem S6 proves. The hardness statement should be qualified as conditional on additional interpolation/robustness conjectures, or Theorem S6 should be extended to the actual ensemble.
minor comments (3)
- [Main text, Fig. 4] The 'rescaled quantity' for XEB in Fig. 4a is described only in S4.3; the main-text caption should give a forward reference and define the rescaling explicitly, since it is essential for comparing XEB with DFE.
- [Supplementary S2.1, Eq. (S7)] The first-order expansion p2(ε)=p2(0)−2εk+O(ε²) should state precisely how k is defined for gates that are 'covered' by multiple checks, and how the O(ε²) term is controlled when using the measured post-selection rates to conclude that ε is small.
- [General presentation] There are minor typographical issues, including 'Tgates' for 'T gates' at several points, and inconsistent hyphenation of 'T-doped' versus 'T doping.' These do not affect the technical content.
Circularity Check
No significant circularity: the fidelity certificate is a derived inequality with independently measured and estimated inputs; acknowledged noise-model limitations are correctness risks, not circular reductions.
full rationale
The central bound F2 ≥ F1 − Pr(E∈H|E∈A) (Theorem S1, Eqs. S5/S9) is a non-circular mathematical consequence of the fault-path decomposition. F1=0.32(1) is measured by DFE, which the paper explicitly states makes no assumptions about hardware noise ([20,21], main text), and Pr(E∈H|E∈A) is estimated by Monte Carlo over a swept family of Pauli noise strengths and polarizations (Fig. 3c; S2.3), not fitted to the reported fidelity. The resulting 0.013(1) loss and the 95% lower bound 0.284 therefore do not reduce to the inputs by construction. The inference that T gates are noiseless from equal post-selection rates (S2.1) is an empirical model-dependent step; S2.5 explicitly concedes adversarial Z-only or stabilizer-aligned noise could leave syndromes unchanged while dropping fidelity, which is a limitation of the certificate's noise assumptions rather than a circular dependence of the bound on its conclusion. Self-citations [22] (spacetime-code formalism), [70] and [81] (open-source software) are backed by the paper's own S1.2 exposition and reproducible code, and are not used as uniqueness theorems; they are not load-bearing in a way that equates the derivation to its premises. The asymptotic hardness claim (Theorem S6) uses standard worst-to-average reductions and cited techniques, not self-referential inputs. No specific equation or fitted parameter is renamed as a prediction.
Assumptions & free parameters
free parameters (2)
- Harmless-fault Monte Carlo noise model family =
CZ error 0.1–0.2%, idle τ=100–200 µs, readout 0.3%, polarizations swept; max loss 0.013(1)
- Check-picking heuristic hyperparameters =
f=0.5, 1000 iterations, 5 random permutations, middle-out/binary ordering
assumptions (6)
- domain assumption After twirling, the device noise is a Pauli channel (stochastic, gate-level)
- domain assumption T gates are noiseless
- domain assumption Circuit-independent noise: the noise model depends only on circuit size, not on the presence of S/T gates within the family
- domain assumption Random Clifford circuits scramble Paulis so that harmless accepted faults have probability O(n^{-c}) or exponentially small at high weight
- domain assumption Complexity conjectures: polynomial hierarchy does not collapse; average-case #P-hardness of output probabilities up to additive error 2^{-n/poly(n)}
- standard math Spacetime Pauli algebra: valid checks correspond to back-propagators equal to identity (Eqs. S2–S4)
Cite this review
Pith. "Pith review of Sampling hard circuits with verifiably high fidelity." pith.science (2026). https://pith.science/paper/Z6OCGHVN
@misc{pith2026260725941,
author = {Pith},
title = {Pith review of: Sampling hard circuits with verifiably high fidelity},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z6OCGHVN}},
note = {Machine review of arXiv:2607.25941}
}
abstract
Sampling-based proposals are prominent candidates for demonstrating quantum computations beyond the reach of classical supercomputers. However, it has been difficult to combine their complexity-theoretic hardness with two capabilities needed for scalable quantum computing more generally: suppressing hardware errors, and verifying the quantum computation itself. Here we address both issues by introducing structured circuits, which, in addition to provable hardness guarantees, admit an encoding in a quantum code. This allows us to simultaneously reach high fidelities at high circuit depths, and to certify an experimental fidelity via the circuit structure and measurement of code syndromes. The resulting certificate is device dependent, but requires substantially weaker noise assumptions than existing fidelity proxy benchmarks. We demonstrate our proposal with a $64$-qubit, depth-$73$ Clifford circuit, doped with $314$ $T$ gates. We use a total of $76$ physical qubits to encode this computation in spacetime codes, effectively suppressing gate error rates by $10\times$ after syndrome post-selection, and yielding a state with a fidelity lower bound of $0.349$ with $95\%$ confidence. Our construction is a systematic method for promoting a stabilizer state to a magic state while keeping an error-detected fidelity certificate.
Forward citations
Cited by 2 Pith papers
-
Classical Verification of Quantum Advantage via Clifford Obfuscation
Clifford circuit obfuscation with overlapping local patch updates is proposed as a heuristic method to create quantum circuits that are hard to classically simulate yet efficiently classically verifiable via stabilize...
-
Classical Simulation and Design Frontiers for IBM's Doped Clifford Sampling Experiment
A deterministic tensor network contraction computes exact probabilities for all of IBM's doped Clifford sampling outputs in 37.3 minutes on 32 H100 nodes.
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.