{"id":"812221c2-32ba-4000-be59-8f08a95ef82a","arxiv_id":"2607.08767","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"Plaquette compiles realistic quantum hardware noise models into multiple sampler representations, showing that Pauli-twirled approximations can misestimate logical error rates by an order of magnitude compared to leakage-aware and near-Clifford methods.","lead":"This paper presents Plaquette, a software framework that simulates how real quantum hardware errors (leakage, heating, coherent over-rotations) affect the logical performance of fault-tolerant quantum computers. It matters because hardware teams can use it to decide which imperfections to fix first and to get accurate overhead estimates without relying on oversimplified noise models.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"XPauli's incoherence assumption is validated only at d=3; the d=5–19 threshold results that anchor the paper's headline claims lack full-state verification.","rationale":"The reader correctly identified the XPauli incoherence assumption as the load-bearing concern. I agree that the lack of full-state validation beyond d=3 is the single most important gap. The paper is honest about this — it does not claim full-state verification at large distances — but the headline quantitative claims (20% threshold shift, order-of-magnitude discrepancy, 2.6× overestimate) are all drawn from the unvalidated regime. The concern is not that the incoherence assumption is necessarily wrong; for the specific physical models considered (Markovian Lindblad dynamics with well-separated energy levels), it may well be an excellent approximation. The concern is that the paper provides no bound on when it breaks down, and the physical regimes where it could fail — multiple leaked levels with comparable lifetimes, coherent Rabi oscillations between leaked states, or leakage events that create effective non-Markovian correlations across rounds — are not exotic. The closed-source software (limitation 2) and the Pauli-twirled decoder initialization (limitation 3) are real but secondary: the decoder issue is acknowledged and the software issue is a verification concern rather than a correctness concern. The d=5 full-state check I propose is computationally feasible (3^5 ≈ 243-dimensional state space per qubit group, manageable with the GPU backend described in Section IIIC3) and would directly settle whether the incoherence approximation is distance-stable. If it passes, the paper's claims are well-supported and the verdict could move toward ACCEPT. If it fails, the threshold estimates need revision. The CONDITIONAL verdict is appropriate as-is.","tokens_in":31735,"tokens_out":858,"duration_ms":231345,"concrete_test":"Run a full-state simulation of the superconducting leakage model (Eqs. 15–16) at d=5 with the same γ values used in Fig. 7, retaining the three-level transmon Hilbert space. Compare the full-state logical error rate to XPauli at 3–4 representative γ values spanning the threshold region. If the full-state and XPauli curves diverge by more than the statistical uncertainty at d=5, the incoherence approximation is distance-dependent and the d=5–19 threshold estimate is unreliable. If they agree within error bars, the d=3 validation likely extends and the headline claims are secure.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claims — 20% threshold shift (Fig. 7), order-of-magnitude logical error rate discrepancy at d=19, and the 2.6× scattering-axis overestimate (Fig. 8) — all come from XPauli simulations at distances where no full-state reference exists. The paper validates XPauli against full-state only at d=3 (Fig. 5b, Fig. 6a). At d=9, the paper explicitly drops full-state (Fig. 6b caption: 'no full-state leg') and reports a 25–56× gap between XPauli and Stim, but cannot confirm XPauli itself is tracking the true channel. The incoherence assumption (Section IIIC1, paragraph 2: 'coherence between different leaked levels, sectors, or environment labels can be discarded') could introduce distance-dependent systematic errors: as code distance grows, leakage events accumulate across more rounds and qubits, and coherent phases between leaked levels — if physically present — could interfere in ways that affect syndrome extraction differently at d=19 than at d=3. The generalized Pauli twirl that XPauli applies (Table III) discards off-diagonal structure in the leakage subspace, and the paper provides no bound on how the resulting approximation error scales with circuit depth or code distance. If the approximation degrades with distance, the threshold estimates themselves could be unreliable, not just different from the Pauli-twirled baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper presents Plaquette, a software framework for simulating fault-tolerant quantum error correction under realistic, non-Pauli hardware noise. The core technical contribution is the XPauli sampler, which extends stabilizer simulation by tracking leaked levels and environment sectors as classical labels while retaining coherence within the computational subspace. The framework also integrates near-Clifford samplers for coherent errors and full-state simulation for reference calculations. The paper validates XPauli and near-Clifford samplers against full-state simulation on small codes (distance-3 repetition code), then applies XPauli to three hardware-relevant scenarios: superconducting qubit leakage, neutral-atom intermediate-state scattering, and trapped-ion heating. The central quantitative finding is that Clifford-only Pauli-twirled simulations can underestimate logical error rates by over an order of magnitude and shift thresholds by approximately 20 percent relative to XPauli, which the authors argue is the more reliable approximation.","tokens_in":32478,"tokens_out":1428,"duration_ms":184912,"significance":"The paper addresses a genuine and important gap in FTQC design: the mismatch between the stochastic Pauli noise assumed by scalable stabilizer simulators and the richer noise structure of real hardware. The XPauli sampler is a well-motivated contribution that fills a practical niche between Clifford-only and full-state simulation. The channel-first design framework, which compiles a single physical error model into multiple sampler-specific representations, is a useful engineering contribution. The three hardware demonstrations are illustrative and span the major matter-qubit platforms. The key caveat, discussed below, is that the headline quantitative claims at large code distances rest on XPauli without independent full-state verification at those distances.","major_comments":[{"comment":"Section IIIC1, paragraph 2: The XPauli sampler's efficiency rests on the assumption that 'coherence between different leaked levels, sectors, or environment labels can be discarded.' This incoherence assumption is validated against full-state simulation only at distance 3 (Fig. 5b, Fig. 6a). At distance 9 (Fig. 6b), full-state is dropped and a 25–56× gap between XPauli and Stim is reported, but XPauli itself is not independently verified there. The threshold results at d=5–19 (Fig. 7), the 2.6× scattering-axis overestimate (Fig. 8), and the order-of-magnitude discrepancy at d=19 all depend on XPauli accuracy at distances where no full-state reference exists. The paper provides no bound on how the generalized Pauli twirl's approximation error scales with circuit depth or code distance. If leakage events accumulate coherently across rounds in a way the incoherence approximation misses, the","section":null},{"comment":"Section IIID, paragraph on the trapped-ion model: The sector-dependent depolarization model (Eq. 22) uses a linear relation p_depol(n) = p_0 + kappa(2n+1) with p_0 = 1e-4 and kappa = 5e-3, described as 'a modelling relation based on Ref. [84].' This relation is the sole coupling between the vibrational sector and the computational qubit, and thus the entire trapped-ion demonstration rests on it. The paper should clarify whether this is a physically motivated mapping or a purely illustrative choice, and ideally provide a sensitivity analysis showing how the threshold estimate changes with different sector-to-rate assignments. Without this, it is unclear whether the trapped-ion results generalize beyond the specific functional form chosen.","section":null}],"minor_comments":[{"comment":"Fig. 5: The caption states '5,000 full-state shots and 200,000 shots for the other samplers,' but the main text (Section IVA, paragraph preceding Fig. 5) says '5,000 shots with the full-state sampler and 200,000 shots with the others.' Fig. 6a caption says '10^5 shots for Stim and XPauli and 10^4 shots for full-state,' which is inconsistent with the 5,000 figure. Please reconcile.","section":null},{"comment":"Section IVB: The dimensionless parameters (omega=4.0, alpha=2.0, g=0.005, tau_CZ~444) are stated to be 'illustrative rather than tuned to a specific device.' It would help to note whether these are at least in a physically reasonable range for transmons, or whether they are purely abstract.","section":null},{"comment":"Section IVD, Eq. (22): The statement 'p_0 = 10^{-4} and kappa = 5 x 10^{-3}, so that the depolarizing probability grows from about 0.5% in sector sigma=0 to 4.5% in sector sigma=4' appears to contain an arithmetic inconsistency: p_0 + kappa*(2*0+1) = 10^{-4} + 5e-3 = 5.1e-3 ~ 0.5%, which is consistent, but p_0 + kappa*(2*4+1) = 10^{-4} + 5e-3*9 = 4.51e-2 ~ 4.5%, which is also consistent. The values are fine; please just verify the intermediate sectors are as intended.","section":null},{"comment":"Table II: The 'Near-Clifford sampling' entry lists four sub-methods but the distinction between 'general' and 'unitary' routes (channel-level vs. operator-level decomposition) could be clearer in the table. A brief note on when a user should choose one over the other would help.","section":null},{"comment":"Section IIC: The threshold surface construction via barycentric scan directions is described, but the number of directions used (15 in Fig. 8) is mentioned only in the figure caption. Stating this in the main text would clarify the resolution of the surface.","section":null},{"comment":"The paper uses 'Plaquette' both for the framework and the software suite. Occasional clarification that these refer to the same entity would help readers unfamiliar with the product.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a software/framework announcement with supporting technical validation. The XPauli sampler is a legitimate contribution, but the paper's structure blurs the line between a methods paper and a product demonstration. The key scientific question — whether XPauli's incoherence approximation remains valid at the distances where the headline claims are made — is not resolved. A single full-state data point at d=5 (even for one of the three hardware models) would substantially strengthen the paper. Without it, the threshold estimates are internally consistent but not externally validated at the distances that matter. The authors may also want to consider whether the near-Clifford samplers, which are exact, could serve as an intermediate-distance cross-check for the leakage models, since those samplers handle coherent structure that XPauli discards."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper introduces XPauli, a stabilizer-based sampler that tracks leakage levels and environmental sectors as classical labels while keeping full coherence within the computational subspace. It's a genuine extension of prior leakage-simulation work (Fowler, Suchara et al., the Pauli+ and Bosonic Pauli+ approaches), unified into a single framework that handles correlated Pauli-leakage errors, skipped gates, and sector-dependent noise. The channel-first design philosophy — specify a CPTP channel once, compile automatically to whatever sampler you need — is clean and practically useful. The near-Clifford samplers for coherent errors are also a real addition, and the validation against full-state at d=3 (Fig. 5) is solid: XPauli and the unitary near-Clifford sampler both land within confidence intervals of exact simulation while Pauli-twirled Stim gives noticeably wrong answers. That's a real result. The three hardware examples (transmon leakage, neutral-atom scattering, trapped-ion heating) are well-constructed and show the framework working end-to-end from Lindbladian physics to logical error rates. The 20% threshold shift in the superconducting example and the 2.6× scattering-axis overestimate in the neutral-atom example are the kind of quantitative findings that matter for hardware teams. The stress-test concern about the incoherence assumption is legitimate but probably overstated. The assumption — that coherence between leaked levels can be discarded — is physically well-motivated for the systems they model. Transmon leakage to |2⟩ decoheres fast, and the neutral-atom intermediate-state scattering they model is inherently incoherent (spontaneous emission). The real gap is that they validate against full-state only at d=3, then run d=5–19 without that check. That's a reasonable practical limitation — full-state with three-level qudits at d=19 is genuinely infeasible — but they should acknowledge more explicitly that the d=9 panel (Fig. 6b), where XPauli reports 25–56× above Stim, has no independent ground truth. The decoder issue is minor and they flag it themselves: using a Pauli-twirled DEM even when the sampler tracks richer noise is a known suboptimality, not a flaw in the results. The closed-source software is a real limitation for independent verification, though the physics and methods are described in enough detail to assess them. This paper is for people designing FTQC architectures who need to know whether Pauli approximations are good enough for their hardware. The answer, demonstrated clearly, is often no. It deserves a serious referee. The main thing I'd push for in revision is either a bound on how the incoherence approximation scales with distance, or at least a d=5 full-state spot check to bridge the gap between d=3 validation and d=19 claims.","headline":"New simulation method (XPauli) for leakage and environmental noise in QEC, validated at small scale but with unverified approximation at the distances where the headline claims live","tokens_in":32856,"tokens_out":664,"would_cite":true,"duration_ms":118946,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Pp","03.67.Lx","03.65.Yz"],"model":"glm-5.2","headline":"Pauli-noise shortcuts can misjudge quantum error correction by 10x","keywords":["quantum error correction","fault-tolerant quantum computing","stabilizer simulation","leakage noise","Pauli twirling","threshold estimation","surface code","open quantum systems"],"falsifier":"Construct a hardware noise model where leaked levels or environment sectors retain quantum coherence that affects logical error rates, and show that XPauli's logical error rate estimate diverges from full-state simulation beyond statistical uncertainty while a method that retains that coherence does not.","tokens_in":31717,"feed_emoji":"🔧","tokens_out":1927,"duration_ms":84360,"temperature":0.7,"pith_summary":"Fault-tolerant quantum computers must be simulated under the noise that real hardware actually produces — leakage, coherent over-rotation, heating, scattering — not the simplified stochastic Pauli noise that fast simulators handle. This paper introduces a framework, called Plaquette, that takes a hardware error model specified once as a quantum channel (from Kraus operators, Lindblad dynamics, or experiment) and automatically compiles it into the representation each of four simulation methods needs. The central technical contribution is the XPauli sampler, which extends efficient stabilizer simulation to qubits that leak into higher energy levels or interact with environmental modes: it tracks those non-computational degrees of freedom as classical labels while retaining full stabilizer coherence for qubits that remain in the computational subspace. This hybrid representation keeps simulation cost polynomial rather than exponential. The paper validates XPauli and a near-Clifford sampler for coherent errors against exact full-state simulation on small codes, showing agreement within statistical uncertainty, while the standard Pauli-twirled approximation can underestimate logical error rates by over an order of magnitude and shift threshold estimates by roughly twenty percent. Three realistic hardware demonstrations — transmon leakage, neutral-atom scattering, and trapped-ion heating — show that the magnitude of the discrepancy depends on the platform and noise process, making the choice of simulation method a material design decision rather than a technical detail.","feed_headline":"Pauli-noise shortcuts can misjudge quantum error correction by 10x","feed_subtitle":"A new simulator tracks leakage and heating that standard tools ignore, shifting threshold estimates by 20 percent or more","key_machinery":"The XPauli sampler's state representation: a stabilizer state on computational qubits, tensored with classical leakage-level labels and classical sector labels for environmental states. A generalized Pauli twirl converts hardware-derived Kraus channels into transition tables among these labels, with conditional Pauli errors on remaining computational qubits. The near-Clifford samplers decompose non-Clifford channels or unitaries as signed sums of Clifford elements, sampling Clifford circuits with quasiprobability weights that reconstruct the original channel in expectation.","core_discovery":"The XPauli sampler rests on a specific structural assumption about quantum noise: that coherence between different leaked levels, environment sectors, or between leaked and computational subspaces can be discarded without losing the information that matters for logical error correction. Under this assumption, the state of a multi-qubit system with leakage and environmental coupling factors into a stabilizer state on the computational qubits tensored with classical labels for everything else. A generalized Pauli twirl converts any CPTP channel into a transition table over these labels plus conditional Pauli errors on computational qubits. Sampling then proceeds in two stages — first a random-","pith_inferences":["If the incoherence assumption fails — for instance, if leaked levels retain phase coherence that interferes with subsequent gate operations or if leakage-mediated entanglement between qubits affects syndrome extraction — XPauli would introduce uncontrolled approximation errors whose magnitude is not bounded in the paper. The validity of the assumption likely depends on the specific hardware platfo","The observation that Pauli twirling can underestimate logical error rates by over an order of magnitude raises a concern for the broader QEC literature: many published threshold estimates based on Clifford-only simulation with Pauli-twirled noise may be systematically optimistic, and the direction of the bias (optimistic vs pessimistic) may depend on the specific noise structure in ways that are n","The framework's channel-first design could, in principle, be extended to simulate non-Markovian noise by embedding memory into explicit environment levels or sectors, but the paper notes this is limited to cases where the relevant memory can be represented within the circuit-level channel formalism — genuinely long-range temporal correlations may require fundamentally different simulation strategi"],"forward_implications":["Hardware teams evaluating whether their device is below threshold should not rely on Pauli-twirled approximations alone — the resulting threshold and logical error rate estimates can be off by an order of magnitude or more, leading to misallocated engineering effort.","The XPauli sampler's classical-label approach could be extended to other non-Pauli noise structures beyond leakage and heating, such as crosstalk with classical memory or time-correlated noise, as long as the incoherence assumption holds for the additional degrees of freedom.","The separation between sampler representation and decoder model means decoders initialized from Pauli-twirled detector error models may be suboptimal for leakage or coherent noise — richer decoder models (correlated, leakage-aware, trajectory-conditioned) could improve correction performance beyond what current decoders achieve.","Threshold surfaces in multi-parameter hardware noise spaces, as demonstrated for neutral atoms, could become a standard design tool: rather than reporting a single threshold number, hardware characterization would map the full tradeoff surface among competing imperfections."],"fun_headline_variants":["Pauli noise assumptions can misjudge logical error rates 10x","Coherent errors and leakage shift fault-tolerance thresholds beyond Pauli models","From open-system physics to logical thresholds: Plaquette tests fault tolerance","Twirling can miss leakage effects that shift error correction thresholds","New sampler matches full-state simulation where Pauli twirling falls short"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The XPauli sampler assumes that quantum coherence between different leaked levels, between leaked and computational subspaces, or between different environment sectors can be safely discarded — that these degrees of freedom behave classically. If leaked levels or environmental modes retain phase coherence that influences syndrome extraction, the generalized Pauli twirl underlying XPauli would introduce approximation errors whose magnitude the paper does not bound.","fun_headline_variants_meta":{"raw":{"variants":["Pauli noise assumptions can misjudge logical error rates 10x","Coherent errors and leakage shift fault-tolerance thresholds beyond Pauli models","From open-system physics to logical thresholds: Plaquette tests fault tolerance","Twirling can miss leakage effects that shift error correction thresholds","New sampler matches full-state simulation where Pauli twirling falls short"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":766,"prompt_tokens":677,"completion_tokens":89,"prompt_tokens_details":null},"tokens_in":677,"tokens_out":89,"duration_ms":20296,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T01:31:46.691168+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Construct a hardware noise model where leaked levels or environment sectors retain quantum coherence that affects logical error rates, and show that XPauli's logical error rate estimate diverges from full-state simulation beyond statistical uncertainty while a method that retains that coherence does not.","supporting_citations":[],"review_version":1}