{"id":"e541a6e7-5a7e-4609-8726-17f8997f2c81","arxiv_id":"2607.25941","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 97-qubit experiment certifies a 0.284 fidelity lower bound for a 468-T-gate sampling circuit by combining spacetime-code error detection with the measured fidelity of an undoped Clifford reference.","lead":"IBM researchers ran a 70-qubit, depth-70 circuit with 468 T gates on 97 physical qubits, using error-detecting spacetime codes and syndrome post-selection to certify the state’s fidelity without simulating the hard circuit. The result is a 95%-confidence lower bound of 0.284 on the fidelity of the hard “doped” state, a step toward quantum advantage experiments that can both suppress errors and verify their output.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.284 fidelity bound is conditional on an unproven worst-case estimate of Pr(harmless|accepted); the Monte Carlo sweep covers only a parameterized Pauli-noise family and finite-twirl coherent residuals are unquantified.","rationale":"The reader's CONDITIONAL verdict and weakest assumption match the load-bearing issue I find: the numerical fidelity-loss bound is not a proven worst case over the noise class the certificate claims, and the finite-twirl coherent contribution is unquantified. I agree with the reader's assessment rather than escalating to REJECT, because the protocol has a clean inequality proof, the DFE measurement of F1 is assumption-free, the virtual-Z implementation makes T-gate noise small, and the low/medium-doping validation experiments provide real though not decisive support. The missing piece is a finite-size worst-case bound on Pr(E∈H|E∈A) for the actual circuit and a bound on finite-twirl interference. These are specific, checkable gaps, consistent with a CONDITIONAL verdict. Since the reader already assigned CONDITIONAL, I recommend no change.","tokens_in":45036,"tokens_out":6725,"duration_ms":69023,"concrete_test":"For the exact 70×70 circuit and its 27 checks, enumerate the Pauli fault paths that contribute to A and H and solve for max_q Σ_{E∈H} q_E / Σ_{E∈A} q_E over the class Q of circuit-independent Pauli noise channels constrained only by measured per-gate error rates and the observed total acceptance rate, e.g., by linear programming over fault-path probabilities. If the resulting worst-case loss exceeds 0.013(1), the claimed F2 ≥ 0.284 is not established for the stated noise model; if it stays below, additionally simulate the actual 50-twirl distribution to bound the residual coherent-interference term Γ_coh.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim F2 ≥ 0.284 follows from F2 ≥ F1 − Pr(E∈H|E∈A), with F1 = 0.32(1) measured by DFE and Pr(E∈H|E∈A) replaced by a numerical estimate of 0.013(1). That estimate is obtained by Monte Carlo simulation 'sweeping various Pauli noise strengths and polarizations' (Fig. 3c), i.e., maximizing over a small parameterized family of local depolarizing/polarized Pauli channels. It is not a worst-case bound over the class of circuit-independent Pauli noise models that Theorem S1 claims to allow, and no finite-size certificate is given for the specific 70×70 circuit. The asymptotic O(n^{-c}) argument in S2.3 relies on random-circuit mixing of S-gate insertions and does not yield a rigorous finite-n bound for this instance. Separately, S2.4.1 shows that exact Pauli twirling removes coherent interference, but the experiment uses 50 twirls; the residual Γ_coh is never bounded, so the equality of post-selection rates is only approximate. The paper itself lists adversarial scenarios in S2.5—noise concentrated on spacetime stabilizers, Z-only errors on T gates—that leave syndromes unchanged while reducing the doped-state fidelity. These gaps are quantitative and addressable, not merely philosophical. The low- and medium-doping validations provide independent support, and the virtual-Z implementation makes pure T-gate Z-error physically unlikely, but the headline 95% confidence bound still depends on a noise-model assumption narrower than the paper's stated generality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a protocol, doped Clifford sampling (DCS), for quantum sampling experiments that are intended to be classically hard and at the same time verifiable: a random Clifford circuit is encoded in a spacetime code, post-selected on syndrome outcomes, and then doped with T gates that commute with the code checks. The central theoretical result is that the post-selected fidelity F2 of the doped state is lower-bounded by F2 ≥ F1 − Pr(E∈H|E∈A), where F1 is the fidelity of the Clifford state (measured by DFE) and the loss term is the conditional probability that an accepted Pauli fault is harmless in the Clifford circuit. The authors report a 70-qubit, depth-70 Clifford circuit with 468 T gates encoded into 97 physical qubits, F1 = 0.32(1), and a claimed 95% confidence lower bound F2 ≥ 0.284. They also provide asymptotic hardness arguments for the DCS ensemble and numerical studies of classical simulation difficulty at the experimental size.","tokens_in":45388,"tokens_out":6860,"duration_ms":72103,"significance":"If the numerical certificate were made rigorous, this would be a significant contribution: it combines a structured sampling proposal with code-based error detection and a fidelity estimate that does not rely on a classical simulation proxy for the hard state. The core inequality is clean, the DFE measurement of F1 is noise-model independent, and the use of virtual-Z T gates is a strong physical argument that doping itself introduces no additional gate error. The low- and medium-doping validation experiments (Fig. 4) provide meaningful empirical support. However, the headline 0.284 bound is conditional on a Monte Carlo estimate over a chosen Pauli-noise family, and the coherent-noise residual from using finitely many twirls is never bounded. These gaps are quantitative and affect the advertised 95% confidence statement.","major_comments":[{"comment":"The numerical loss Pr(E∈H|E∈A) ≤ 0.013(1) is obtained by Monte Carlo simulation 'sweeping various Pauli noise strengths and polarizations' (Fig. 3c). This is not a worst-case bound over the circuit-independent Pauli noise models allowed by Theorem S1. The asymptotic O(n^{-c}) argument in S2.3 is also not a finite-size certificate for the 70×70 instance. Since the experimental claim F2 ≥ 0.284 at 95% confidence relies directly on this number, the bound should either be proven as a true maximum over the claimed noise family, or be replaced by an interval computed from measured calibration data with rigorous uncertainty propagation. As written, the 95% confidence statement covers Monte Carlo sampling error within one parameterized family, not model uncertainty.","section":"Main text Fig. 3c; Supplementary S2.3"},{"comment":"Exact Pauli twirling removes the coherent interference term Γ_coh in Eq. (S23), but the experiment uses only 50 twirl instances (S6.4). The residual |Γ_coh| after finite twirling is never bounded. Consequently, the equality of post-selection rates and the reduction to the stochastic Pauli bound are only approximate. The paper needs a quantitative estimate of the finite-twirl residual—for example, a variance bound over the 50 randomizations actually used—and must include this residual in the claimed confidence interval if the certificate is to hold under coherent noise.","section":"Supplementary S2.4.1; S6.4"},{"comment":"The adversarial scenarios listed in S2.5—Z-only errors on T gates and noise concentrated on the spacetime stabilizer group—are exactly cases where the syndrome distributions can remain unchanged while F2 drops by more than the estimated 0.013. The paper dismisses them as 'not physically realistic' but provides no quantitative argument or experimental test that excludes them. Given the stated goal of substantially weaker noise assumptions than XEB, these scenarios must be addressed explicitly, either by a measurable calibration check (e.g., injecting and amplifying such errors in a controlled way) or by a revised certificate that states the additional assumption and its experimental support.","section":"Supplementary S2.5"},{"comment":"The average-case hardness theorem is proved for the interpolated distribution eD_{α*,K,Θ}, not for the uniform-rotation DCS ensemble that is actually sampled. The transition from exact computation of the truncated polynomial q_{β,K} on this interpolated distribution to approximate sampling from the original ensemble is not established; S8.3.2 explicitly notes that the required additive error 2^{-n/poly(n)} for 1/poly(n)-TV sampling remains open. The main-text claim that 'no classical sampler exists' under 'two appropriate conjectures' therefore overstates what Theorem S6 proves. The hardness statement should be qualified as conditional on additional interpolation/robustness conjectures, or Theorem S6 should be extended to the actual ensemble.","section":"Supplementary S8.3, Theorem S6"}],"minor_comments":[{"comment":"The 'rescaled quantity' for XEB in Fig. 4a is described only in S4.3; the main-text caption should give a forward reference and define the rescaling explicitly, since it is essential for comparing XEB with DFE.","section":"Main text, Fig. 4"},{"comment":"The first-order expansion p2(ε)=p2(0)−2εk+O(ε²) should state precisely how k is defined for gates that are 'covered' by multiple checks, and how the O(ε²) term is controlled when using the measured post-selection rates to conclude that ε is small.","section":"Supplementary S2.1, Eq. (S7)"},{"comment":"There are minor typographical issues, including 'Tgates' for 'T gates' at several points, and inconsistent hyphenation of 'T-doped' versus 'T doping.' These do not affect the technical content.","section":"General presentation"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The core idea is new and clever: encode a random Clifford circuit in a spacetime code, use the code to detect errors, then dope with T gates at locations that commute with all checks. The fidelity of the Clifford state is measurable by DFE, and the doped state's fidelity is bounded via F2 >= F1 - Pr(harmless|accepted). That inequality is clean and not circular, and the 97-qubit experiment with 468 T gates is a substantial step beyond prior sampling demonstrations. The validation at low and medium doping is genuinely independent and supportive, and the virtual-Z implementation makes pure Z-error on T gates physically unlikely. Credit also for shipping data and code.\n\nNow the soft spots, in proportion. The main one is exactly what the reader flagged: the numerical loss of 0.013(1) is a Monte Carlo estimate over a swept family of Pauli noise channels, not a rigorous finite-size worst-case bound for the 70x70 circuit. The theorem gives an asymptotic O(n^{-c}) argument, but that does not certify this instance. The paper itself lists adversarial scenarios in S2.5—noise concentrated on spacetime stabilizers, Z-only errors on T gates—that would leave syndromes unchanged while dropping the doped fidelity. These are acknowledged but not quantitatively bounded. The finite-twirl coherent residual is also not quantified; exact twirling is proved to recover the Pauli bound, but the experiment uses 50 twirls and the interference term Gamma_coh is never bounded. I do not think these are fatal. They are specific, addressable gaps, and the system-level evidence (validation experiments, syndrome distribution agreement) makes the bound plausible. But the headline \"95% confidence\" overstates the status: it is a confidence interval conditional on a noise-model family, not a distribution-free certificate.\n\nThe hardness analysis is honest—two standard conjectures, plus a careful numerical study of tensor network, MPS, stabilizer, and hybrid methods. The simulation claims look reasonable for this size, with the usual caveat that classical algorithms improve. The asymptotic complexity of the doping itself is only heuristic at finite size, but the paper says as much.\n\nBottom line: this is a serious paper with a real new method and a substantial experiment. The certificate is the load-bearing part, and it needs a rigorous finite-size worst-case bound or an explicit characterization of the noise family under which the 0.284 bound holds. I would send it to referees, and tell the authors to sharpen exactly that point. The reading group would enjoy it, and I would cite it for the protocol even while being careful about the numerical claim.","headline":"A genuinely new error-detected sampling protocol with a real 97-qubit demonstration; the headline fidelity bound is conditional on a noise-model assumption that is narrower than the paper's stated generality.","tokens_in":45923,"tokens_out":651,"would_cite":true,"duration_ms":10560,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Lx","03.67.Pp"],"model":"deepseek-v4-flash","headline":"A 70-qubit doped Clifford circuit achieves a certified fidelity lower bound of 0.284 while remaining classically hard to simulate, by placing T gates only where they commute with the error-detection checks.","keywords":["quantum advantage","sampling hardness","fidelity certification","spacetime codes","Clifford circuits","T-gate doping","error detection","quantum verification"],"falsifier":"Run the same 70-qubit DCS protocol while intentionally injecting a controlled Z-only error on a subset of T gates (or a correlated noise aligned with a measured check), and compare the measured fidelity of the doped state (via XEB at intermediate doping or DFE at low doping) against the certificate's predicted lower bound: if the fidelity drops by more than the certificate allows while the syndrome distribution is unchanged, the certificate is falsified.","tokens_in":44905,"feed_emoji":"⚛️","tokens_out":6578,"duration_ms":51734,"temperature":0.7,"pith_summary":"Sampling-based quantum advantage experiments have struggled to combine three things: provable classical hardness, suppression of hardware errors, and trustworthy verification of the output. This paper introduces doped Clifford sampling (DCS), which encodes a random Clifford circuit in a spacetime code, then injects T gates at locations that commute with the code's checks. Because the T gates are noiseless (implemented by frame tracking) and do not change the syndrome statistics, the measured fidelity of the Clifford state and the syndrome-passing rate can be used to lower-bound the fidelity of the hard doped state. The paper demonstrates this on a 70-qubit, depth-70 circuit with 468 T gates and 97 physical qubits, reporting a 10× effective suppression of gate error rates and a 95%-confidence fidelity lower bound of 0.284.","feed_headline":"Verified hard circuit: 70 qubits, 468 T gates, fidelity ≥ 0.284","feed_subtitle":"Spacetime codes suppress errors 10× while T-doping keeps the circuit hard and the syndrome statistics unchanged.","key_machinery":"The load-bearing object is the spacetime Pauli check: a Pauli operator supported on circuit wires whose back-propagation through the Clifford circuit is the identity (or a stabilizer of the input state), implemented via an ancilla measurement. The second ingredient is T-doping at check-commuting locations: each T gate is placed on a wire where a Z rotation commutes with all checks' back-cumulants, preserving the accepted fault set and syndrome distribution. The argument is carried by the identity F2 ≥ F1 − Pr(E∈H|E∈A), which turns the hard-to-measure doped fidelity into a difference of two efficiently accessible quantities, plus a counting/mixing argument showing that harmless faults are rar","core_discovery":"The central claim is that for a Clifford circuit C1 equipped with spacetime Pauli checks, and a circuit C2 obtained by adding T gates at wires that commute with all checks, the post-selected fidelity of the doped state satisfies F2 ≥ F1 − Pr(E∈H|E∈A), where F1 is the Clifford state's fidelity, A is the set of accepted (syndrome-passing) Pauli faults, and H is the subset of accepted faults that stabilize the Clifford output. The accepted fault set and post-selection probability are identical in the two circuits because the T gates commute with the checks. Since a harmless fault in a random linear-depth Clifford circuit is unlikely, the loss term Pr(E∈H|E∈A) is small; for the specific 70×70 in","pith_inferences":["If the harmless-fault probability estimate transfers to other Clifford skeletons, the certificate could become a generic template for verifying magic-state sampling on any device with good syndrome extraction; the key requirement is a random Clifford backbone with low-weight spacetime checks.","The certificate as stated leaves a gap for adversarial noise that is invisible to the checks (e.g., Z-only errors on T gates or noise aligned with the code's stabilizers). A direct experimental test at 468 T gates, for instance by comparing measured syndrome-conditional fidelities at intermediate T counts, would determine whether the assumed noise family contains the real device noise.","Because the certificate is numerical and noise-model-dependent, a fully device-independent verification remains out of reach; however, the structure suggests that planting circuit-specific secrets (peaked circuits) could be combined with DCS to give a verifiable advantage without any noise assumptions."],"forward_implications":["Sampling experiments can now simultaneously be classically hard, error-suppressed through post-selection, and accompanied by a device-dependent fidelity certificate that needs only Pauli noise after randomized compiling.","Because the certificate is instance-specific and does not require direct simulation of the doped circuit, it can be applied to larger or deeper circuits than currently simulable.","The construction gives a systematic method for promoting a stabilizer state to a magic state while retaining the error-detection structure of the spacetime code, providing a bridge from near-term error detection to fault-tolerant quantum advantage.","The 70×70, 468-T-gate instance is estimated to be infeasible for tensor-network and stabilizer-based classical simulations, and its effective CZ error rate after post-selection is 1.8×10^{-4}, making noisy-classical-simulation attacks impractical."],"fun_headline_variants":["Verifiable fidelity: 70 qubits, 468 T gates, 10x error suppression","Hard but verifiable: 70-qubit circuit with 468 T gates hits fidelity 0.284","Spacetime codes: 10x error suppression, verifiable 0.284 fidelity at depth 70","Certified hard sampling: 70-qubit circuit, 468 T gates, fidelity ≥0.284","Verifiable fidelity at scale: 70 qubits, 468 T gates, 10× suppression"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The certificate assumes that, after randomized compiling, the physical noise is exactly described by the Pauli-noise family swept in the Monte Carlo estimate of Pr(E∈H|E∈A); if the real noise contains components outside that family—such as globally correlated errors aligned with the spacetime stabilizers, or pure Z errors on the T gates—the syndromes can remain unchanged while the doped state's fidelity falls by more than the estimated 0.013.","fun_headline_variants_meta":{"raw":{"variants":["Verifiable fidelity: 70 qubits, 468 T gates, 10x error suppression","Hard but verifiable: 70-qubit circuit with 468 T gates hits fidelity 0.284","Spacetime codes: 10x error suppression, verifiable 0.284 fidelity at depth 70","Certified hard sampling: 70-qubit circuit, 468 T gates, fidelity ≥0.284","Verifiable fidelity at scale: 70 qubits, 468 T gates, 10× suppression"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000641,"raw_usage":{"total_tokens":2799,"prompt_tokens":769,"completion_tokens":2030,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":1901}},"tokens_in":513,"tokens_out":2030,"duration_ms":12052,"temperature":1.0,"reasoning_tokens":1901,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:02:38.511123+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 70-qubit DCS protocol while intentionally injecting a controlled Z-only error on a subset of T gates (or a correlated noise aligned with a measured check), and compare the measured fidelity of the doped state (via XEB at intermediate doping or DFE at low doping) against the certificate's predicted lower bound: if the fidelity drops by more than the certificate allows while the syndrome distribution is unchanged, the certificate is falsified.","supporting_citations":[],"review_version":1}