{"id":"5a765736-df90-49e3-95d7-3d94fa5aeb6a","arxiv_id":"2506.17190","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For distance-3 spin-qubit codes in silicon, a hybrid Loss-DiVincenzo/singlet-triplet encoding outperforms an all-Loss-DiVincenzo encoding, and the Bacon-Shor code gives much lower logical state-preparation error than the surface code, with gate errors dominating the error budget.","lead":"This paper simulates two quantum error-correcting codes, the surface code and the Bacon-Shor code, on silicon spin qubits using either all single-electron qubits or a hybrid of single-electron and singlet-triplet qubits. It finds the hybrid design performs better under current experimental parameters and that gate errors, not memory decoherence, set the logical error rate floor for these distance-3 codes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gate-error dominance may be inflated by treating LD and ST single-qubit gate errors as a single uniform parameter in the error-subset sampler.","rationale":"The reader's verdict is CONDITIONAL, with the weakest assumption identified as the transferred GaAs ST two-qubit gate infidelity and the ST T2* scaling. My stress-test found a different, more internal concern: the error-subset sampler as described uses one probability per error category, but Table I assigns different error rates to LD and ST qubits. If a single (presumably the larger ST) value is used for all single-qubit gates, the LD data-qubit Ry gates, which Sec. V.B identifies as the dominant single-qubit gate error source, are modeled with a 10x too-high infidelity. This would make gate errors appear more dominant over memory errors than they actually are, potentially undermining the central claim that coherence-time improvements will not help at current parameters. The paper does not provide code or a detailed mapping of Table I onto the sampler parameters, so the ambiguity cannot be resolved from the manuscript alone. The concern is load-bearing because it targets the key numerical conclusion, but it is not an outright refutation; a revised simulation with separate parameters could either confirm or dissolve it. Therefore the existing CONDITIONAL verdict is appropriate, and I recommend no change to the reader's verdict. I do not identify a basis for REJECT, as the qualitative ranking (hybrid vs all-LD, surface vs BS) may survive even if the quantitative gate-dominance claim shifts. My agreement with the reader is partial because the reader's weakest assumption concerns the value of the ST gate error (from a GaAs experiment), whereas my concern is about the internal treatment of the two distinct gate-error rates in the sampler; both relate to gate-error parameterization but are not the same issue.","tokens_in":26290,"tokens_out":21388,"duration_ms":205359,"concrete_test":"Re-run the surface-17 hybrid QEC simulation at the state-of-the-art integration time (2.0 us) with the LD and ST gate infidelities kept separate as in Table I, tracking LD and ST single-qubit and two-qubit errors as distinct subset parameters. Then check whether the T2*=21 us logical error rate still saturates the T2*->infinity bound, and repeat the Fig. 9(d) sensitivity sweep by reducing only the ST single-qubit gate infidelity from 4e-3 to 4e-4 while holding LD at 4e-4. If the logical error rate decreases materially or the saturation disappears, the gate-error-dominance claim is an artifact of the uniform-parameter approximation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the logical error rate is fundamentally limited by gate errors and not memory errors (Sec. V.B, Fig. 11) rests on the relative size of gate errors versus idle dephasing. The error-subset sampler described in Sec. IV.B uses m=8 independent parameters with a single probability per category (e.g., 'after 1-qubit gates', 'after 2-qubit gates'), and the subset-weight formula A_w is a product of binomial factors with one p_i per category. However, Table I assigns distinctly different infidelities to LD and ST qubits: single-qubit gate infidelity is 4e-4 for LD but 4e-3 for ST, and two-qubit gate infidelity is 2e-3 for LD but 4e-3 for ST. The text in Sec. II explicitly says these probabilities 'will be different depending on the type of qubit.' If the simulator instead used a single value (for example, the larger ST value 4e-3) for all single-qubit gates, then the 24 LD Ry gates per round on data qubits would carry a 10x inflated error rate. Section V.B attributes the largest sensitivity to p1q precisely to these Ry gates on data qubits. With the correct LD value of 4e-4, the two-error contribution from data-qubit Ry gates would drop by roughly two orders of magnitude, potentially lowering the gate-error floor below the memory-error contribution at T2*=21 us. The saturation of the T2*=21 us curve to the T2*->infinity bound in Fig. 11, which is the direct evidence for the headline claim, could then disappear. The manuscript does not explain how the two distinct gate infidelities in Table I are mapped onto the 8 sampler parameters, nor whether the sensitivity sweeps vary LD and ST separately. This is an internal consistency concern about the simulation's error model, independent of whether the chosen ST values are realistic.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies distance-3 surface and Bacon-Shor codes implemented on silicon spin qubits, comparing an all-Loss-DiVincenzo (LD) encoding with a hybrid LD-data/ST-ancilla encoding. Using stabilizer simulations with a circuit-level Pauli noise model parameterized by experimental values from Table I, the authors compute logical error rates for one fault-tolerant QEC step and for logical |0> and |+> state preparation. They report that the hybrid scheme outperforms the all-LD scheme because of the much shorter ST readout, that surface-17 slightly beats BS-17 in the QEC step but BS-17 is about two orders of magnitude better in logical state preparation, and that the logical error rate is dominated by 1- and 2-qubit gate errors rather than memory errors.","tokens_in":26593,"tokens_out":7923,"duration_ms":79434,"significance":"If the results hold, the paper provides a practically useful comparison for near-term silicon spin-qubit QEC experiments. Its strengths include the explicit, table-driven physical parameter set; the use of lower and upper bounds on logical error rates via an error-subset sampler; and the observation that the surface and BS circuits can be identical, differing only in classical decoding. The claim that gate errors, not memory errors, set the logical error floor for distance-3 hybrid schemes is a concrete, falsifiable prediction that would usefully guide hardware improvements. The simulation methodology is standard and carefully described, and the work is non-circular: no parameter was fitted to the target logical error rates.","major_comments":[{"comment":"The error-subset sampler in Sec. IV.B is described with m=8 independent parameters, each with a single probability p_i in Eq. (1), yet Sec. II and Table I assign different error probabilities to LD and ST qubits for single-qubit gates (4e-4 vs 4e-3) and two-qubit gates (2e-3 vs 4e-3). The manuscript never explains how these distinct values are mapped onto the sampler's categories, for example whether the 'after 1-qubit gates' category uses one value or separate values per qubit type. If the implementation uses a single value, such as the larger ST value, for all single-qubit gates, then the RY(±π/2) gates on the LD data qubits would carry a tenfold-inflated error rate; since Sec. V.B identifies these data-qubit RY gates as the dominant source of the p1q sensitivity, the gate-error floor in Fig. 11 could be substantially overestimated, and the saturation of the T2*=21 µs curve at the T2*→∞ bound, which is the direct evidence for the headline claim, might disappear when the correct LD value of 4e-4 is used. This ambiguity must be resolved and, if necessary, the simulations rerun with per-type gate-error parameters.","section":"IV.B and Table I"},{"comment":"The quantitative conclusions depend on two extrapolated ST-qubit parameters: the two-qubit gate infidelity of 4e-3 is taken from a GaAs LD-ST experiment [48] and assumed for silicon, and the ST coherence time is set to T2*,ST = T2*,LD/√2. The authors mention the expected improvement in silicon, but the gate-error-dominance claim in Sec. V.B is stated for 'current parameters' and would shift if these extrapolations are off by even a factor of a few. A sensitivity analysis over a plausible range of these ST parameters, for example p2q,ST from 1e-3 to 1e-2 and T2*,ST from 5 µs to 30 µs, would make the robustness of the hybrid advantage and of the gate-error floor quantitative rather than assumed.","section":"Table I and Sec. V.B"}],"minor_comments":[{"comment":"The phrase 'have been showed recently' should be 'have been shown recently'.","section":"Introduction"},{"comment":"The affiliation contains a typo: 'Univeristy' should be 'University'.","section":"Author affiliations"},{"comment":"The sentence 'The decoding (inferring the error that occurred based on the syndromes).' is an incomplete sentence; it should be rephrased.","section":"Sec. III"},{"comment":"The caption contains a typo: 'inital state preparation error probability' should be 'initial state preparation error probability'.","section":"Fig. 9 caption"},{"comment":"The caption contains a typo: 'doted line' should be 'dotted line'.","section":"Fig. 11 caption"},{"comment":"The caption contains a duplicated article: 'For the the surface-17 code' should be 'For the surface-17 code'.","section":"Fig. 12 caption"},{"comment":"The coherence-time labels '21 s' and '210 s' should be '21 µs' and '210 µs' to match the units used in the text.","section":"Fig. 7 caption"},{"comment":"The word 'Thisconstitutes' is missing a space and should read 'This constitutes'.","section":"Sec. V.C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope and the citation pattern appears appropriate. The main risk is the unresolved mapping of the per-type LD/ST error parameters onto the m=8 sampler categories; this directly affects the central gate-error-dominance claim, so it should be clarified and, if needed, the simulations redone before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a stabilizer-simulation comparison of distance-3 surface and Bacon-Shor codes on silicon spin qubits, with an all-Loss-DiVincenzo (LD) versus hybrid LD/singlet-triplet (ST) encoding. The two findings worth taking seriously are the roughly two-order-of-magnitude advantage of BS-17 over surface-17 for fault-tolerant logical state preparation (thanks to the coherent, GHZ-based protocol), and the demonstration that the hybrid scheme's fast ST readout is what makes QEC viable at current parameters. The simulation method is standard, the error-subset sampler with lower/upper bounds is sensible, and the authors are upfront about the main limitations: uncorrelated Pauli noise, GaAs-sourced ST two-qubit gate error, and the absence of noise correlations.\n\nThe real soft spot is internal consistency of the noise model as reported. Table I gives single-qubit gate infidelity 4e-4 for LD and 4e-3 for ST, and two-qubit gate infidelity 2e-3 for LD and 4e-3 for ST. The sampler, however, is described with m=8 parameters, one per error category, with a single probability per category. The text in Sec. II explicitly says error probabilities will differ by qubit type, but Sec. IV.B never explains how the distinct LD/ST values are mapped onto eight parameters. Since the hybrid circuit applies Y rotations to both LD data and ST ancilla qubits, and all two-qubit gates are between LD and ST, this is not a trivial detail. The sensitivity sweeps in Fig. 9(d) use a single p1q whose baseline is 4e-3, the ST value, which suggests they may have conservatively used the larger infidelity for all single-qubit gates. The authors attribute the strong p1q sensitivity to RY gates on the LD data qubits; if those actually have infidelity 4e-4, the gate-error floor in Fig. 11 could drop noticeably, and the load-bearing claim that logical error is fundamentally limited by gate errors rather than memory errors might weaken. I cannot tell from the text whether this is a real error or just an under-specified description, but it is a serious reproducibility gap, compounded by the absence of code or data.\n\nMinor issues: the abstract's \"consistently outperforms\" overstates Fig. 7, where at very short readout times the all-LD scheme does better due to better LD gate fidelities, and the ST T2* value is an unannotated estimate (LD T2*/sqrt(2)).\n\nWho benefits? Experimentalists planning proof-of-concept distance-3 QEC on silicon spin qubits. The practical guidance — focus on reducing 1- and 2-qubit gate errors, use hybrid encoding, consider BS for state preparation — is actionable. I would send this to peer review: the BS state-prep result and the architecture comparison are worth publishing once the authors clarify the noise-model parameter mapping and ideally release the simulation wrapper. Accept conditionally, with that clarification required.","headline":"Useful spin-qubit QEC architecture comparison with a genuinely valuable BS state-prep result, but the central gate-error-dominance claim is undercut by an unexplained mapping of distinct LD/ST gate infidelities to a single sampler parameter.","tokens_in":27217,"tokens_out":6131,"would_cite":false,"duration_ms":58376,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Pp","03.67.Lx"],"model":"deepseek-v4-flash","headline":"Simulations of distance-3 surface and Bacon-Shor codes on silicon spin qubits find that a hybrid encoding—single-electron Loss–DiVincenzo data qubits paired with fast-readout singlet–triplet ancillas—outperforms an all-Loss–DiVincenzo…","keywords":["silicon spin qubits","quantum error correction","surface code","Bacon-Shor code","Loss-DiVincenzo qubit","singlet-triplet qubit","hybrid encoding","logical error rate"],"falsifier":"Measure the logical error rate of the hybrid surface-17 code at the optimal integration time in two otherwise identical silicon devices whose ST two-qubit gate infidelities differ by tenfold while T2* is held fixed; the paper predicts a clear drop in logical error, whereas a memory-limited scenario predicts none. A cheaper check is to re-run the simulation with T2* set to infinity: the paper predicts the logical floor stays near 1e-2, limited by gates.","tokens_in":26064,"feed_emoji":"⚛️","tokens_out":9181,"duration_ms":83516,"temperature":0.7,"pith_summary":"This paper simulates two quantum error-correcting codes—the surface code and the Bacon-Shor code—on silicon spin qubits, using realistic experimental error rates. It compares a layout where every qubit is a single-electron Loss–DiVincenzo spin with a hybrid layout where data qubits are Loss–DiVincenzo and ancilla qubits are faster-read singlet–triplet qubits. The central finding is that the hybrid layout wins by more than an order of magnitude, because the short singlet–triplet readout prevents data qubits from dephasing during syndrome measurement. The paper also finds that at current parameters the logical error rate is limited by one- and two-qubit gate errors, not by memory or coherence errors, and that the Bacon-Shor code's fault-tolerant logical state preparation is about two orders of magnitude better than the surface code's because it can be done coherently without projective measurements.","feed_headline":"Gate errors, not memory, limit spin-qubit error correction","feed_subtitle":"A hybrid silicon-qubit design keeps logical error rates low; cutting one- and two-qubit gate errors matters most.","key_machinery":"The load-bearing object is the hybrid qubit layout: nine Loss–DiVincenzo data qubits (single-electron spins in single dots) and eight singlet–triplet ancilla qubits (two-electron states in double dots), coupled by exchange-mediated CZ gates, with the ST ancillas reading out roughly ten times faster than LD qubits. Two codes share the same stabilizer-measurement circuit and differ only in the classical decoder: the surface code uses all weight-2 and weight-4 outcomes directly, while the Bacon-Shor code discards some information and multiplies outcomes to get weight-6 parity checks. The performance evaluation is carried by a stabilizer simulator with an error-subset importance sampler that returns lower and upper bounds on logical error rates for eight independent Pauli noise parameters. For the Bacon-Shor advantage, the key mechanism is a shallow GHZ-state preparation circuit that is naturally fault-tolerant at distance 3, requiring row connectivity between data qubits.","core_discovery":"On the paper's own terms, the central discovery is a quantitative ranking of two architectures for a distance-3 logical qubit in silicon. Using state-of-the-art Si/SiGe parameters, the authors simulate one round of fault-tolerant quantum error correction and fault-tolerant logical state preparation for the surface-17 and Bacon-Shor-17 codes. They find that the hybrid LD-ST scheme reduces the logical error rate by more than an order of magnitude compared with the all-LD scheme, that the surface code is slightly better than Bacon-Shor in the QEC cycle, and that Bacon-Shor is about two orders of magnitude better at logical state preparation because its |0⟩ and |+⟩ states can be prepared by a unitary GHZ circuit rather than projective stabilizer measurements. The paper further claims that, for these fast distance-3 protocols, the logical error floor is fundamentally limited by gate errors—chiefly one- and two-qubit gate infidelities—and not by memory (dephasing) errors.","pith_inferences":["The gate-error-dominance result, if it holds beyond the simulated circuits, implies that near-term silicon spin-qubit roadmaps should spend their fidelity budget on faster and more accurate gates rather than on prolonging T2*; coherence gains only pay when code distance grows.","The Bacon-Shor coherent state-preparation advantage suggests that early fault-tolerant demonstrations might be better served by choosing a code with a shallow unitary logical-preparation circuit than by choosing the code with the best QEC cycle, since state injection is often the bottleneck.","A clean experimental discriminator would be to measure the surface-17 hybrid logical error floor at the optimal integration time while varying only the ST two-qubit gate error; the paper's model predicts a roughly 1.75-fold drop when the CZ infidelity goes from 4e-3 to 4e-4.","Because the paper uses uncorrelated Pauli noise, spatial noise correlations seen in silicon spin-qubit pairs could alter the ranking; if correlated errors mimic higher-weight errors, distance-3 codes would fail faster than these simulations predict."],"forward_implications":["A proof-of-concept distance-3 experiment on silicon should prioritize lowering one- and two-qubit gate infidelities: reducing them together by tenfold lowers the surface-17 hybrid logical error floor from about 1.7e-2 to below 1e-3, while cutting CZ duration or state-preparation error alone has almost no effect.","The hybrid logical qubit at state-of-the-art readout times makes a slightly better quantum memory than an unprotected physical LD qubit, but a simple spin-echo pulse on a physical qubit outperforms it; the logical qubit is not yet a memory win.","For the Bacon-Shor-17 code, fault-tolerant preparation of logical |0⟩ and |+⟩ at about 4.05e-4 beats physical state preparation by roughly an order of magnitude and the surface-17 projective preparations by two orders of magnitude, provided direct entangling gates exist between data qubits along rows or columns.","The optimal ST readout integration time for the logical error rate is about 0.88 microseconds, shorter than the roughly 1.4 microseconds that minimizes readout infidelity, so choosing the integration time for QEC is a separate optimization from choosing it for readout alone.","The conclusion that gate errors dominate memory errors is confined to fast distance-3 protocols; the paper expects memory errors to become significant for higher-distance codes that require more syndrome-measurement rounds."],"supporting_citations":[{"why":"Supplies the hybrid LD-ST two-qubit gate infidelity of 4e-3 and demonstrates the exchange-mediated interface between the two encodings.","marker":"[48]"},{"why":"Supplies the ST readout time of 2.4 microseconds, readout infidelity of 4e-4, and the measured readout-infidelity-versus-integration-time curve used in Appendix B.","marker":"[46]"},{"why":"Supplies the LD two-qubit CZ gate infidelity of 2e-3 and gate time of 40 ns used in the noise model.","marker":"[17]"},{"why":"Supplies the ST single-qubit gate infidelity of 4e-3 from resonantly driven singlet-triplet qubits.","marker":"[55]"},{"why":"Supplies the LD dephasing time T2* = 21 microseconds used for the data qubits.","marker":"[103]"},{"why":"Supplies LD single-qubit gate, readout, and state-preparation infidelities included in Table I.","marker":"[105]"},{"why":"Supplies the earlier observation that the optimal readout integration time for QEC differs from the integration time that minimizes readout infidelity.","marker":"[49]"},{"why":"Supplies the experimental precedent for fault-tolerant Bacon-Shor logical state preparation via GHZ states in trapped ions, which motivated the coherent preparation circuit.","marker":"[77]"}],"fun_headline_variants":["Hybrid spin-qubit encoding cuts logical error rate tenfold","Spin-qubit QEC: gate errors, not memory, are the limit","Silicon spin qubits: surface vs Bacon-Shor, hybrid wins","For spin qubits, gate errors dominate over memory errors","Bacon-Shor excels at state prep, hybrid qubits overall better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the faster-readout singlet-triplet ancilla qubits will have roughly the same two-qubit gate quality in silicon as in the gallium-arsenide experiments their parameters are borrowed from, and that noise on different qubits is uncorrelated; if either assumption fails, the quantitative sizes of the hybrid advantage and the gate-vs-memory conclusion change.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid spin-qubit encoding cuts logical error rate tenfold","Spin-qubit QEC: gate errors, not memory, are the limit","Silicon spin qubits: surface vs Bacon-Shor, hybrid wins","For spin qubits, gate errors dominate over memory errors","Bacon-Shor excels at state prep, hybrid qubits overall better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3256,"prompt_tokens":925,"completion_tokens":2331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":2239}},"tokens_in":541,"tokens_out":2331,"duration_ms":15579,"temperature":1.0,"reasoning_tokens":2239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:10:20.334438+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the logical error rate of the hybrid surface-17 code at the optimal integration time in two otherwise identical silicon devices whose ST two-qubit gate infidelities differ by tenfold while T2* is held fixed; the paper predicts a clear drop in logical error, whereas a memory-limited scenario predicts none. A cheaper check is to re-run the simulation with T2* set to infinity: the paper predicts the logical floor stays near 1e-2, limited by gates.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies LD single-qubit gate, readout, and state-preparation infidelities included in Table I."},{"cited_title":"Levy, Universal quantum computation with spin-1/2 pairs and heisenberg exchange, Phys","cited_arxiv_id":null,"evidence_quote":"Supplies the experimental precedent for fault-tolerant Bacon-Shor logical state preparation via GHZ states in trapped ions, which motivated the coherent preparation circuit."}],"review_version":2}