{"id":"77a08a1f-cf23-43cc-a48b-c0baeef581c6","arxiv_id":"2608.05518","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Tuning the ratio of X-to-Z syndrome-extraction rounds in defect-adapted surface codes, based on noise bias, reduces logical error rate by up to 8.46x in simulations.","lead":"This paper shows that in a surface code adapted to broken qubits and couplers, the error-correcting rounds that check for X-type and Z-type errors do not need to be balanced: under noise that favors one type of error, tilting the round ratio toward the opposite check type can cut the logical error rate by up to 8.5x in simulations. The result gives hardware makers a calibration-based rule for choosing the ratio, without running a fresh simulation per device.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The p_Y=0 noise model is load-bearing: in the T1/T2-twirled channels used to motivate eta, p_Y=p_X, and Y errors produce X- and Z-syndromes at different rates under asymmetric schedules, so omitting them may shift the optimal ratio and inflate the reported improvements.","rationale":"I read the paper as establishing a previously unstudied scheduling knob for defect-adapted surface codes under biased noise. The simulation pipeline (Stim/PyMatching) is standard, the qualitative intuition is plausible, and the claim that the optimal ratio is approximately independent of code distance is backed by the 5%-tolerance bars in Figure 6. The single most fragile link is the Pauli channel: the paper explicitly sets p_Y=0 and justifies it with an unproved rescaling argument. Under alternating schedules, this argument is suspect because a Y error produces two syndrome events sampled at different rates in the X and Z subgraphs, so the decoder's matching problem changes with the schedule, not merely the overall error rate. Moreover, the T1/T2 relation used to connect eta to hardware calibration data has p_Y = p_X, so the simplified model does not match the intended hardware regime. The concrete test would settle whether R* and the improvement factors survive a realistic p_Y. I do not find other objections severe enough to move the verdict: the crosstalk section is separately modeled, and the lack of error bars, while worth reporting, is secondary to the noise-model question. I therefore keep the reader's CONDITIONAL verdict, with the condition that the p_Y=0 assumption be tested quantitatively.","tokens_in":17572,"tokens_out":6324,"duration_ms":61975,"concrete_test":"Rerun the Section IV sweep for d=13, dr=1% and 2%, eta=5, using the same 30 defect-adapted patches and shot counts, but replace the Pauli channel by p_Y = p_X, p_Z = 5 p_X, with p_X + p_Y + p_Z = 1e-3. Record R* and the LER improvement over R=1:1. If R* moves by more than one ratio step (e.g., 9:1 to 5:1) or the improvement changes by more than about 20%, the p_Y=0 simplification is load-bearing and the quantitative claims need qualification. A secondary check: repeat with p_X + p_Z fixed at 1e-3 and p_Y = p_X to separate the normalization effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III sets p_Y=0 and asserts that Y errors 'flip both syndromes, so including them rescales the LER by a constant factor at fixed bias, without changing the trends observed.' This assertion is not derived and is unlikely to hold for alternating schedules. A Y error is a correlated XZ event: in a schedule with R_X:R_Z != 1, its X-syndrome and Z-syndrome components are sampled at different rates and enter the MWPM decoding graph as separate detection events separated in time. Changing p_Y therefore changes the relative weight of the two decoding subgraphs, not just the overall error rate. The paper's own calibration relation eta = T1/T2 - 1/2 (Section IV-B and Fig. 11) comes from a T1/T2 Pauli-twirled channel in which p_Y = p_X, so the p_Y=0 model is inconsistent with the hardware regime used to choose eta. Since the headline numbers (4.25x, 8.46x) and the claim that R* is set only by eta depend on this choice, the p_Y=0 assumption is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the optimal ratio of X-type to Z-type syndrome-extraction rounds for a defect-adapted surface code (using the Snakes and Ladders framework) under biased noise. The authors simulate memory experiments with Stim and PyMatching, setting p_Y=0 and fixing p=p_X+p_Z, and sweep R_X:R_Z for various code distances, defect rates, and bias values eta=p_Z/p_X. They report that a non-uniform schedule reduces the average logical error rate by up to 4.25x at 1% defect rate and 8.46x at 2% defect rate for d=13 at eta=5, and that the optimal ratio R* is essentially independent of code distance and defect rate, depending mainly on eta. They also extend the idea to defect-free codes under a crosstalk model and report up to 4.5x improvement.","tokens_in":17853,"tokens_out":13054,"duration_ms":111836,"significance":"If the claims hold, the paper identifies a practical, calibration-driven tuning knob for surface codes with fabrication defects, complementing code-level bias tailoring. The use of standard tools (Stim, PyMatching) and the SnL framework is appropriate, and the empirical finding that R* depends primarily on noise bias is interesting and potentially useful. The paper is clearly written and the crosstalk section is presented as an illustrative extension. However, the two central claims rest on a restrictive p_Y=0 noise model and on simulation studies with no statistical uncertainty quantification, both of which need to be addressed before the quantitative results can be relied upon.","major_comments":[{"comment":"The simulation model sets p_Y=0 and asserts that Y errors 'flip both syndromes, so including them rescales the LER by a constant factor at fixed bias, without changing the trends observed.' This assertion is not derived and is inconsistent with Section IV-B, where the calibration relation eta = T1/T2 - 1/2 comes from a Pauli-twirled T1/T2 channel in which p_Y = p_X. Under a non-uniform schedule R_X:R_Z != 1, a Y error is a correlated XZ event whose X- and Z-syndrome components are sampled in alternating rounds at different rates; changing p_Y therefore alters the relative weight of the X and Z decoding subgraphs rather than only the overall error rate. The optimal ratio R* could shift with p_Y, so the headline improvements (4.25x at 1% defects, 8.46x at 2% defects) and the claim that R* depends only on eta are conditional on this unverified assumption. Please repeat the key sweeps (at least Figures 4 and 6) with nonzero p_Y, e.g., p_Y = p_X as implied by T1/T2 twirling, and report whether R* and the improvement factors change.","section":"Section III (Pauli channel setup)"},{"comment":"No shot counts, error bars, or confidence intervals are reported for any LER estimate, and each configuration uses only 10 to 30 stochastic defect patches (Figure 4 uses 10 patches; Figure 5 uses 30 patches). The definition of R* in Section IV-A as the smallest ratio whose average LER is within 1% of the observed minimum, and the 5% bars in Figure 6, therefore cannot be interpreted without knowing the sampling noise. The claim that R* is essentially independent of code distance and defect rate is a null result that needs uncertainty quantification; with a coarse ratio grid and small patch counts, the observed flatness of R* versus d could be an artifact of sampling variation. Please report the number of shots per patch, the number of patches per point, and standard errors or confidence intervals for the LER and for R*.","section":"Sections IV-A and IV-B (statistical reporting)"},{"comment":"The paper advertises that the optimal ratio can be selected directly from measured calibration data, without running device-specific simulations. However, the only quantitative output is Figure 6, which shows R* for eta in {1,2,4,6,8,10} and for discrete ratios up to 49:1, with no fitted curve, interpolation rule, or statement of how sensitive the LER is to a small error in R*. Since the transferability claim relies on the relationship R*(eta) being robust, the manuscript should provide either a table of recommended ratios, a fit R*(eta), or an explicit error tolerance (e.g., the range of R where LER is within 1% of the minimum) as a function of eta.","section":"Section IV-B (calibration-based selection)"}],"minor_comments":[{"comment":"The text 'performing O(d) rounds of syndrome extraction' is underspecified; state the exact total number of rounds used for each d and confirm that the total number of rounds is held fixed across all schedules so that LER comparisons are not confounded by different memory times.","section":"Section III (simulation setup)"},{"comment":"The improvement at d=13, dr=1%, eta=5 is reported as 4.46x in Figure 4 and 4.25x in Figure 8, while the abstract gives 4.25x; please reconcile these numbers and state which patch set and ratio grid each value refers to.","section":"Figures 4 and 8"},{"comment":"The 1% tolerance used to define R* is not justified; please report how the choice of tolerance affects the identified R* (e.g., a sensitivity sweep) or provide shot-noise-based error bars that motivate the tolerance.","section":"Section IV-A (optimality tolerance)"},{"comment":"The relation eta = T1/T2 - 1/2 is stated without derivation; consider adding the Pauli-twirling expressions for p_X, p_Y, p_Z or a citation with the explicit formula to make the connection to Section III's p_Y=0 model transparent.","section":"Section IV-B (calibration relation)"},{"comment":"The crosstalk model sets the simultaneous-to-isolated gate error ratio alpha in [1,2] but cites only general randomized-benchmarking references; state which measured values or specific experimental results motivate this range.","section":"Section V (crosstalk model)"},{"comment":"There are minor typographical issues: 'X-ZRound' appears in the title line, 'eta=p z/px' should use proper subscripts, and the notation R_X:R_Z should be defined explicitly in the abstract or introduction.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of QCE and the core question is timely. The main technical concern (the p_Y=0 assumption) is addressable with additional simulations, and the statistical reporting can be improved. I would encourage the authors to release the Stim scripts and data, as the small patch counts make reproducibility checking valuable. The discrepancy between Figure 4 and Figure 8 should be resolved before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's core idea—tuning the X:Z round scheduling ratio in defect-adapted surface codes under biased noise—is new and probably right in direction, but the headline quantitative claims rest on a noise model that sets p_Y=0, and that assumption is not as harmless as the authors claim. I read the stress-test note and it holds up.\n\nWhat is actually new: prior defect-adaptation frameworks (SnL, LUCI, ACID, MIA) are benchmarked under symmetric or depolarizing noise, while bias-tailored codes like XZZX are studied on defect-free lattices. The joint setting, and the observation that the optimal ratio R* is set primarily by bias and is roughly independent of distance and defect rate, is a genuinely useful empirical result if it survives. The crosstalk extension is a nice bonus.\n\nThe paper does several things well. The simulations use Stim and PyMatching, which are standard. The qualitative trend—that under Z-biased noise you want more X rounds—has a clear mechanism and is consistent across distances and defect rates. The related-work survey is fair and situates the contribution correctly. They also provide a calibration recipe (choose R* from T1/T2 data) that would be valuable if validated.\n\nThe soft spots are the ones you'd expect. First and most serious: p_Y=0 is load-bearing. The authors assert that Y errors flip both X and Z syndromes, so including them would only rescale the LER. That is not generally true for asymmetric schedules. A Y error is an XZ correlation; under R_X:R_Z != 1, its X and Z syndrome components are sampled at different rates and appear in the decoding graph at different times. Changing p_Y changes the relative weight of the two subgraphs, not just the overall rate. There is also an internal inconsistency: the paper motivates eta via T1/T2 with Pauli twirling, and in that channel p_Y = p_X. So the noise model used for the main results is not the same one used to justify the bias range. This needs to be fixed before the 4.25x and 8.46x numbers can be taken seriously.\n\nSecond, there are no error bars or shot counts. Results are based on 10–30 stochastic defect patches per configuration, and the 1% tolerance used to define R* is arbitrary. The qualitative trend is probably robust, but the claimed improvement factors are unquantified. Third, I see no mention of code or data release; that should be part of the package.\n\nOverall, this is a solid engineering paper with a correct central direction but an overly strong conclusion. The p_Y=0 issue is fixable—run the same sweeps with p_Y > 0 and report whether R* shifts. If it does, the headline numbers change. If it doesn't, the qualitative claim stands.\n\nI would send this to peer review rather than desk reject. A competent referee should push for Y-error simulations and statistical reporting. If those come back clean, this is a citable contribution. As it stands, I would not cite the quantitative results.","headline":"New and plausible scheduling knob for defect-adapted surface codes, but the p_Y=0 model makes the headline LER reductions unreliable.","tokens_in":18361,"tokens_out":3471,"would_cite":false,"duration_ms":29955,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under biased noise, reallocating X- and Z-check rounds in defect-adapted surface codes lowers logical error rates by up to 8.46x.","keywords":["surface code","quantum error correction","biased noise","defect adaptation","round scheduling","syndrome extraction","super-stabilizers","logical error rate"],"falsifier":"Run the same scheduling sweep at $\\eta=5$, $d=13$, and 1-2% defect rates with a biased Pauli channel that includes nonzero $p_Y$ (for example $p_Y = \\sqrt{p_X p_Z}$), and check whether the optimal ratio $R_X:R_Z$ shifts from the $p_Y=0$ prediction and whether the logical error rate improvement over 1:1 shrinks.","tokens_in":17355,"feed_emoji":"⚛️","tokens_out":7172,"duration_ms":56872,"temperature":0.7,"pith_summary":"This paper argues that when a surface code is adapted to fabrication defects, the alternating X/Z syndrome-check schedule imposed by the adaptation is not a fixed cost but a tunable design parameter. Under biased noise, where one Pauli error type dominates, measuring the syndrome basis that detects the dominant error more often than the other lowers the logical error rate. In circuit-level simulations with bias $\\eta = p_Z/p_X$ up to 10, the optimal ratio $R_X:R_Z$ moves away from 1:1 and improves the average logical error rate by 1.12x to 8.46x. The optimal ratio is set mainly by the noise bias and is nearly independent of code distance and defect rate, so it can be chosen from device calibration data instead of per-device simulation. The same scheduling idea also helps defect-free codes under CNOT crosstalk, where separating X and Z rounds improves the logical error rate by up to 4.5x.","feed_headline":"Rebalancing syndrome-check rounds cuts logical errors up to 8.46x","feed_subtitle":"A simple scheduling change lets defect-adapted surface codes match noise bias using only calibration data.","key_machinery":"The central object is the scheduling ratio $R = R_X:R_Z$, the relative number of consecutive syndrome-extraction rounds devoted to X-type versus Z-type gauge checks in a defect-adapted surface code. Defect adaptation converts affected stabilizers into lower-weight gauge checks that anti-commute across bases, so X and Z super-stabilizers must be measured in alternating rounds; the paper treats the relative frequency of those rounds as a free parameter. The mechanism has two effects: more frequent rounds in one basis sample the corresponding syndrome sooner, suppressing the logical error rate in that basis, and repeating the same gauge-check type across consecutive rounds makes individual gauge outcomes deterministic, letting the decoder discount measurement errors. The optimal ratio balances these effects and turns out to be governed by the noise bias $\\eta$ rather than by code size or defect density.","core_discovery":"The paper's central claim is that under biased noise, the uniform $R=1:1$ round schedule used by defect-adapted surface codes is suboptimal, and reallocating rounds toward the basis that detects the dominant error type lowers the average logical error rate. Concretely, at $\\eta=5$ and distance 13, the optimal schedule improves the logical error rate by 4.25x at a 1% defect rate and 8.46x at a 2% defect rate, relative to 1:1. The optimal ratio $R^*$ is, to a good approximation, determined only by the noise bias $\\eta$: it stays constant across code distances $d=5$ through 13 and across defect rates 0.5% through 2%, up to sampling variation. This claim extends beyond defective hardware: in defect-free patches where parallel CNOTs suffer crosstalk, measuring X and Z stabilizers in separate rounds reduces the parallel-CNOT count per layer and yields up to a 4.5x logical error rate improvement at $d=13$, $\\alpha=2$, and $\\eta=5$.","pith_inferences":["The same round-rebalancing logic should apply to other defect-adaptation schemes that produce anti-commuting gauge checks, so the ratio could become a universal compilation knob rather than a setting specific to the defect-adaptation framework used here.","The $p_Y=0$ assumption is the main point to test on hardware: because Y errors flip both X and Z syndromes, a realistic nonzero $p_Y$ may rescale the optimal ratio or dampen the reported improvements, so a sweep with nonzero Y error would bound the effect.","Spatially heterogeneous bias, like the qubit-to-qubit variation shown in the calibration data, suggests a future extension: a position-dependent scheduling ratio that dedicates more rounds to locally dominant error regions rather than a single global optimum.","The crosstalk extension implies a trade-off between parallelism and error sampling: separating X and Z rounds halves per-layer CNOT concurrency but doubles the number of rounds, so the 4.5x gain is tied to the simulated crosstalk penalty and could weaken for smaller $\\alpha$."],"forward_implications":["A manufacturer can measure the bias $\\eta$ from calibration data and set the scheduling ratio once, reusing it across all code distances without running per-device simulation sweeps.","Larger code distances amplify the benefit: at $\\eta=5$, the improvement over 1:1 grows from 2.07x at $d=9$ to 4.46x at $d=13$.","Higher defect rates increase the benefit: at $d=13$ and $\\eta=5$, the gain rises from 4.25x at 1% defects to 8.46x at 2% defects, even though absolute logical error rates worsen.","In defect-free architectures with CNOT crosstalk, scheduling X-only and Z-only rounds separately can reduce the logical error rate by up to 4.5x at $d=13$ with a 2x crosstalk penalty."],"supporting_citations":[{"why":"Supplies the defect-adaptation construction whose anti-commuting gauge checks create the alternating X/Z schedule that this paper tunes.","marker":"[9]"},{"why":"Stim is the circuit-level stabilizer simulator running all memory experiments.","marker":"[33]"},{"why":"Sparse Blossom minimum-weight matching provides the decoder used to extract logical error rates.","marker":"[3]"},{"why":"Introduces the leading bias-tailored surface-code variant against which this paper positions round scheduling as complementary.","marker":"[14]"},{"why":"Provides the simultaneous-CNOT crosstalk noise model used in the non-defective architecture extension.","marker":"[46]"},{"why":"Supplies the simultaneous-randomized-benchmarking ratio that quantifies crosstalk strength as $\\alpha$.","marker":"[47]"}],"fun_headline_variants":["Round scheduling boosts surface-code error correction up to 8.46x","Optimal X/Z round ratio from noise bias alone improves logical error rate 8.46x","Defect-tolerant surface codes: rebalancing rounds yields up to 8.46x error reduction","Surface code with defects: schedule X/Z rounds by noise bias for 8.46x gain","Crosstalk-aware scheduling: separate X/Z rounds cut errors 4.5x in biased noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulation fixes $p_Y=0$ and assumes Y errors flip both X and Z syndromes, so including them would only rescale the logical error rate at fixed bias rather than change the optimal scheduling ratio; if Y errors interact differently with the alternating gauge-check schedule, the optimal ratio and the reported improvements could shift.","fun_headline_variants_meta":{"raw":{"variants":["Round scheduling boosts surface-code error correction up to 8.46x","Optimal X/Z round ratio from noise bias alone improves logical error rate 8.46x","Defect-tolerant surface codes: rebalancing rounds yields up to 8.46x error reduction","Surface code with defects: schedule X/Z rounds by noise bias for 8.46x gain","Crosstalk-aware scheduling: separate X/Z rounds cut errors 4.5x in biased noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001194,"raw_usage":{"total_tokens":4987,"prompt_tokens":1070,"completion_tokens":3917,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":3799}},"tokens_in":686,"tokens_out":3917,"duration_ms":24756,"temperature":1.0,"reasoning_tokens":3799,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:38:45.159820+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same scheduling sweep at $\\eta=5$, $d=13$, and 1-2% defect rates with a biased Pauli channel that includes nonzero $p_Y$ (for example $p_Y = \\sqrt{p_X p_Z}$), and check whether the optimal ratio $R_X:R_Z$ shifts from the $p_Y=0$ prediction and whether the logical error rate improvement over 1:1 shrinks.","supporting_citations":[{"cited_title":"Snakes and ladders: Adapting the surface code to defects,","cited_arxiv_id":null,"evidence_quote":"Supplies the defect-adaptation construction whose anti-commuting gauge checks create the alternating X/Z schedule that this paper tunes."},{"cited_title":"Stim: a fast stabilizer circuit simulator,","cited_arxiv_id":null,"evidence_quote":"Stim is the circuit-level stabilizer simulator running all memory experiments."},{"cited_title":"Sparse Blossom: correcting a million errors per core second with minimum-weight matching,","cited_arxiv_id":null,"evidence_quote":"Sparse Blossom minimum-weight matching provides the decoder used to extract logical error rates."},{"cited_title":"The xzzx surface code,","cited_arxiv_id":null,"evidence_quote":"Introduces the leading bias-tailored surface-code variant against which this paper positions round scheduling as complementary."},{"cited_title":"Software mitigation of crosstalk on noisy intermediate-scale quantum computers,","cited_arxiv_id":null,"evidence_quote":"Provides the simultaneous-CNOT crosstalk noise model used in the non-defective architecture extension."},{"cited_title":"Optimized noise suppression for quantum circuits,","cited_arxiv_id":null,"evidence_quote":"Supplies the simultaneous-randomized-benchmarking ratio that quantifies crosstalk strength as $\\alpha$."}],"review_version":1}