{"id":"24e212e4-caa8-4dd1-b3ac-53f9bdb28c30","arxiv_id":"2608.11072","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A silicon spin qubit achieves π/2 gate fidelity of 99.99920(2)% by reducing reservoir-induced noise and using Gaussian pulses to avoid off-resonant driving artifacts.","lead":"Silicon spin qubits reach single-qubit gate fidelity above 99.999 percent after removing nearby charge reservoirs and using smooth microwave pulses. The paper also reveals a measurement artifact that makes qubits appear worse than they are, giving practical conditions for correct fidelity benchmarking in spin-qubit arrays.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RB-derived per-π/2 fidelity assumes identical gate-independent, Markovian errors; if the error rate is non-uniform across Xπ/2/Yπ/2 or the decay is non-exponential, the quoted 99.99920(2)% is a model-dependent average.","rationale":"The paper's central claim is a specific per-π/2 gate fidelity, and the only route from the data to that number is the RB conversion that the reader identified. I read the manuscript in good faith: the authors explicitly identify off-resonant driving and heating artifacts, use Gaussian pulses to suppress them, and perform a purity-benchmarking decomposition to separate coherent and incoherent errors. These are real, independent supporting checks. Nevertheless, the conversion from per-Clifford RB decay to per-π/2 fidelity is load-bearing, and the manuscript's own observations of heating-induced frequency shifts and reference-qubit artifacts make the Markovian, gate-independent assumption the least secure part of the chain. The proposed interleaved-RB test directly probes this assumption and is feasible in the same setup. I do not see grounds to reject the paper; the appropriate posture remains conditional on data release and on such a direct per-primitive check. My verdict therefore remains unchanged from the reader's CONDITIONAL assessment.","tokens_in":14374,"tokens_out":8694,"duration_ms":81710,"concrete_test":"Perform interleaved randomized benchmarking on the two primitive rotations Xπ/2 and Yπ/2 actually used in the Clifford decomposition, with the same Gaussian pulse shape, reservoir-isolated configuration, and ZZ readout. Interleaved RB yields direct per-gate depolarizing parameters without invoking the 3.25-average divisor. If the interleaved fidelities do not both agree with 99.99920(2)% within the reported uncertainty, or if the per-sequence-length RB decay shows curvature inconsistent with a single exponential, the headline value is an artifact of the averaging and Markovianity assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline 99.99920(2)% is extracted from a randomized-benchmarking fit by dividing the per-Clifford depolarizing parameter by 3.25, the average number of π/2 rotations per Clifford. That conversion is exact only if every primitive π/2 gate has the same error rate and if the errors are gate-independent and Markovian. The manuscript itself provides reasons to question this: Fig. 3 shows microwave-induced frequency shifts that are time- and power-dependent, a history-dependent, non-Markovian error source, and Fig. 4 shows that off-resonant driving of the reference qubit can masquerade as extra decay in parity readout. The methods section also concedes that GST and PB, when performed in the single-qubit space, fail to account for off-resonant rotations of the reference qubit. Since the final RB data are taken with ZZ parity readout, the reference qubit is always present; any residual off-resonant rotation, even if small for Gaussian pulses, adds a contribution to the observed decay that is not part of Q2's true per-gate error. A further symptom is that the abstract value, infidelity 8.0×10^{-6}, differs from the error-budget run's ε_RB = 8.6(3)×10^{-6}; this may be run-to-run variation, but it illustrates how sensitive the reported precision is to extraction assumptions. If the Xπ/2 and Yπ/2 primitives have unequal errors, or if the decay is not a single exponential, the 3.25-average conversion does not yield a well-defined per-gate fidelity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports single-qubit gate fidelities above 99.999% in a 28Si/SiGe spin qubit, with a headline π/2 rotation fidelity of 99.99920(2)% extracted from randomized benchmarking (RB). The authors identify two main sources of error: (i) proximity to a grounded reservoir, which degrades the driven coherence time T1ρ, and (ii) off-resonant driving of the reference qubit during parity-readout RB, which creates a benchmarking artifact that can be mistaken for qubit error. They show that removing a proximal reservoir, operating at 100 mK to suppress microwave-induced frequency shifts, and using Gaussian-shaped pulses rather than rectangular pulses mitigate these effects. Purity benchmarking is used to separate incoherent and coherent contributions, and the authors conclude that the residual error is predominantly incoherent, with a coherent error per π/2 rotation of about 1.5(4)×10^-6.","tokens_in":14736,"tokens_out":7065,"duration_ms":70683,"significance":"If the reported fidelity is taken at face value, this is a state-of-the-art result for single-qubit control in silicon spin qubits: an infidelity near 8×10^-6 exceeds previous demonstrations by roughly an order of magnitude. The paper is also valuable for identifying and isolating a subtle benchmarking artifact (off-resonant driving of a parity-readout ancilla) that could otherwise corrupt high-fidelity RB claims, and for connecting driven coherence (T1ρ) to device geometry through the reservoir-proximity study. A clear strength is that the central fidelity is a direct experimental measurement with stated statistical error bars, not a derived quantity from a fitted model, and the RB-versus-purity comparison provides evidence for the incoherent-limited claim. The supporting capacitance model and the off-resonant-driving simulation are explicitly presented as fitted or approximate, and they do not feed back into the headline fidelity extraction.","major_comments":[{"comment":"The headline value 99.99920(2)% is obtained by fitting the per-Clifford RB decay and converting to a per-π/2 error by dividing by 3.25, the average number of π/2 rotations per Clifford gate. This conversion is exact only if every primitive π/2 gate (Xπ/2, Yπ/2, and the four-pulse identity) has the same error rate and if the noise is Markovian and gate-independent. The manuscript does not provide an experimental test of these assumptions at the 10^-6 level, and the microwave-induced frequency shifts described in the section on heating are a known time-dependent, potentially non-Markovian error source. I ask the authors to quantify the systematic uncertainty in the per-π/2 fidelity from gate-dependent errors and non-Markovianity, for example by measuring primitive gate fidelities directly or by interleaved RB, or by demonstrating that the RB decay is single-exponential over the full range of m with residuals consistent with the model.","section":"Methods, Randomized benchmarking; Fig. 5a"},{"comment":"The final RB and purity-benchmarking data are taken with ZZ parity readout, so the reference qubit Q1 is always present during the sequence. The methods section explicitly concedes that single-qubit-space GST and PB 'fail to account for off-resonant rotations of the reference qubit,' and the justification for using Gaussian pulses under ZZ readout rests on a simulation (Supplementary Fig. 5) rather than an experimental comparison. If the simulated off-resonant depolarization q is an underestimate, the observed RB decay includes a contribution that is not part of Q2's per-gate error, and the RB-versus-purity decomposition could also be affected. I request an experimental bound on q at the Gaussian-pulse operating point (for example, repeating the IZ-versus-ZZ comparison of Fig. 4d with Gaussian pulses) or an explicit sensitivity analysis of the simulation to its assumptions (noise amplitude A, waiting time, pulse truncation).","section":"Methods, Error budgeting (purity benchmarking); Supplementary Fig. 5"}],"minor_comments":[{"comment":"The conversion from per-Clifford decay to per-π/2 fidelity is described only in prose; please state the explicit relation used (e.g., ε_π/2 = ε_Clifford / 3.25 or p_π/2 = p_Clifford^(1/3.25)) so that the reader can reproduce the number.","section":"Methods, Randomized benchmarking"},{"comment":"The Data availability section contains the placeholder '{URL}' for the Zenodo record; a working DOI should be provided before publication.","section":"Data availability"},{"comment":"The sentence 'Modifying the 500 nm offset in the reservoir geometry changes these constants but does not affect the overall result' is vague; please state explicitly which aspect is robust (the functional form T1ρ ∝ (η^2 + B)^-1, the qualitative distance scaling, or the value of the exponent).","section":"Supplementary Fig. 9"},{"comment":"Several occurrences of 'π/2' appear as 'π22' in the abstract and main text; the final version should use the correct mathematical notation.","section":"Abstract and main text"},{"comment":"The caveat that single-qubit-space GST and PB fail to account for off-resonant rotations of the reference qubit is important and should also appear in the main-text discussion of the error budget, not only in the Methods.","section":"Methods, Error budgeting"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the 3.25 conversion is substantive and lands: the central claim is a per-π/2 fidelity quoted to four significant figures, so the authors should either validate the gate-independence assumption experimentally or bound the systematic error. The off-resonant-artifact issue for Gaussian pulses is also worth an experimental check, although a possible positive bias from the artifact would only make the headline fidelity conservative. With those additions, I would likely support acceptance; the paper is otherwise carefully executed and the central result is likely to be important for the silicon spin-qubit community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a strong experimental paper. It reports a genuine state-of-the-art single-qubit gate fidelity in a 28Si/SiGe spin qubit—99.99920(2)% for pi/2 gates—and, more usefully, identifies two concrete mechanisms that limit both the fidelity and its measurement. The reservoir-proximity effect on T1rho is clearly demonstrated with a control experiment, and the off-resonant-driving artifact in parity-readout RB is a new benchmarking caveat that people working with multi-spin devices should know about. The paper earns its claims.\n\nWhat's new: the fidelity number beats the group's previous 99.99% result by an order of magnitude in infidelity, and the claim that residual error is mostly incoherent is supported by a same-run purity-benchmarking comparison that separates coherent from incoherent contributions. The pulse-shaping fix (Gaussian pulses) and the IZ readout control are well-designed experiments that isolate the artifact. The simulation of off-resonant driving matches the observed oscillation period in gate time, which is a nice quantitative check.\n\nSoft spots, in proportion: the headline fidelity is extracted from standard RB by dividing the per-Clifford error by 3.25, the average number of pi/2 gates per Clifford. That's the usual convention, but it yields an average per-rotation fidelity and assumes the two primitive rotations have comparable errors. The paper doesn't test that assumption directly. Given the error is already at 8e-6, any imbalance between Xpi/2 and Ypi/2 would change the interpretation slightly but not the practical conclusion. The stress-test concern about non-Markovianity is real in spirit—microwave-induced frequency shifts are shown to exist—but the authors mitigate it with 100 mK operation, automated recalibration, and pulse shaping. The remaining artifact from the reference qubit in parity readout is shown to be negligible for Gaussian pulses via simulation. So I don't think the central claim is fragile.\n\nOther caveats: the reservoir-distance scaling is fitted with only three configurations and a capacitance model with adjustable constants; that part is suggestive, not definitive. Data and code are promised but not yet available; the Zenodo link is a placeholder. That's worth a referee noting, but not a blocker.\n\nWho it's for: experimentalists in semiconductor spin qubits and anyone benchmarking high-fidelity single-qubit gates. It deserves serious peer review; I'd send it to a strong experimental journal. I'd cite it for the benchmarking artifact alone.","headline":"A strong experimental paper reporting a state-of-the-art silicon spin-qubit gate fidelity and two practically important benchmarking artifacts; the central claim holds up, with minor caveats about RB extraction assumptions and missing data.","tokens_in":15282,"tokens_out":2627,"would_cite":true,"duration_ms":23512,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By removing a nearby reservoir, operating at 100 mK, and shaping microwave pulses, a driven $^{28}$Si/SiGe spin qubit reaches a $\\pi/2$ gate fidelity of 99.99920(2)\\%, with the remaining error dominated by incoherent noise.","keywords":["silicon spin qubits","single-qubit gate fidelity","randomized benchmarking","purity benchmarking","spin-locking coherence","off-resonant driving","pulse shaping","28Si/SiGe quantum dots"],"falsifier":"Run interleaved randomized benchmarking experiments that isolate $X_{\\pi/2}$ and $Y_{\\pi/2}$ rotations, or run randomized benchmarking with only $X_{\\pi/2}$ gates; if the inferred per-rotation error differs from $8.6(3)\\times10^{-6}$ by more than the error bars, the equal-splitting, gate-independent assumption fails and the headline fidelity is not the average. A second check is to look for curvature or non-exponential tails in the randomized-benchmarking decay, which would signal noise that remembers previous pulses.","tokens_in":14189,"feed_emoji":"⚛️","tokens_out":15168,"duration_ms":125987,"temperature":0.7,"pith_summary":"This paper is trying to establish that a driven silicon spin qubit can be pushed to a $\\pi/2$ gate fidelity above 99.999%, and that the remaining error is governed by decoherence rather than by poor control. The authors report a fidelity of 99.99920(2)% in a $^{28}$Si/SiGe spin qubit after fixing three problems: a nearby reservoir shortens the spin-locking coherence time $T_{1\\rho}$, microwave heating shifts the qubit frequency, and rectangular pulses off-resonantly excite the neighboring qubit used in parity readout, which distorts randomized benchmarking. Removing the reservoir, operating at 100 mK, and using Gaussian pulse shapes suppress all three effects. If correct, the work turns a record number into a reproducible recipe for high-fidelity single-qubit gates in larger silicon arrays.","feed_headline":"Silicon spin qubit reaches 99.999% gate fidelity","feed_subtitle":"With nearby reservoirs removed and pulses shaped, decoherence alone sets the per-gate error at 8 in a million.","key_machinery":"The load-bearing object is the spin-locking coherence time $T_{1\\rho}$, the decay time of a qubit under continuous microwave irradiation, connected to noise by $T_{1\\rho}^{-1} = \\frac12 T_1^{-1} + \\frac{(2\\pi)^2}{2} S_\\parallel(f_{\\rm Rabi})$, where $S_\\parallel$ is the longitudinal noise spectral density at the Rabi frequency. Reservoir proximity enters through a capacitive coupling factor $\\eta(d)^2$ that scales the induced voltage noise. A second mechanism is off-resonant driving in randomized benchmarking: with parity readout, the reference qubit's extra depolarization $q$ adds to the target qubit's depolarizing parameter $p$, so the joint decay looks worse than the true fidelity; the synchronization condition $t_g(n)=\\sqrt{4n^2-1}/(4|\\Delta f|)$ reveals the 7.9 ns periodicity, and Gaussian pulses suppress the artifact. Purity benchmarking closes the argument by comparing the decay of $\\mathrm{Tr}(\\rho^2)$ with the RB decay, separating coherent from incoherent error.","core_discovery":"The central claim is that, in a $^{28}$Si/SiGe single-spin qubit, a $\\pi/2$ rotation fidelity of 99.99920(2)% is achievable, and the fidelity is limited by incoherent noise. The paper identifies the mechanisms that previously stood in the way: proximal reservoirs degrade the driven coherence time $T_{1\\rho}$ via capacitive coupling and longitudinal noise at the Rabi frequency; microwave drive heats the device and shifts the qubit resonance; and rectangular pulses create spectral sidelobes that off-resonantly depolarize the parity-readout partner, adding a spurious decay to the benchmarking signal. With the reservoir depleted, operation at 100 mK, and Gaussian pulses, the measured RB decay rate is $\\varepsilon_{\\rm RB}=8.6(3)\\times10^{-6}$ per $\\pi/2$ gate, the purity decay rate is $\\varepsilon_{\\rm PB}=7.1(2)\\times10^{-6}$, giving a coherent contribution of only $1.5(4)\\times10^{-6}$ and confirming decoherence as the dominant error.","pith_inferences":["The off-resonant artifact is unlikely to be confined to this device: any parity-readout architecture with a shared drive line and qubits separated by tens of MHz will see the same inflated randomized-benchmarking decay, so comparing IZ and ZZ readouts is a cheap diagnostic for it.","The reservoir-proximity dependence of $T_{1\\rho}$ implies that qubits sitting next to charge sensors in dense arrays will have uneven driven coherence; sparse layouts with shuttling, or reservoirs far from qubits, should produce more homogeneous fidelities.","Purity benchmarking's low cost makes it a natural tool for two-qubit error budgets, separating coherent cross-talk from decoherence in CNOT or CZ gates rather than only single-qubit rotations.","If heating scales with total microwave power, scaling to many simultaneously driven qubits will require baseband or low-frequency control; a direct test is to measure the qubit frequency shift while increasing the number of concurrently driven qubits."],"forward_implications":["With the reservoir depleted and Gaussian pulses, single-qubit randomized benchmarking at $f_{\\rm Rabi}=3$ MHz should consistently return fidelities above 99.999%, rather than the 99.9% level commonly reported.","Operating at 100 mK reduces the transient microwave-induced frequency shift to below 50 kHz, so long sequences of tens of thousands of $\\pi/2$ pulses no longer accumulate a coherent detuning error.","Rectangular-pulse randomized benchmarking combined with parity readout systematically under-reports fidelity for any qubit pair separated by tens of MHz; pulse shaping is therefore a prerequisite for trustworthy multi-qubit benchmarks in this architecture.","Because the purity decay rate is close to the randomized-benchmarking decay rate, further improvement in this device will come mainly from lengthening $T_{1\\rho}$ and reducing incoherent noise, not from better pulse calibration."],"supporting_citations":[{"why":"Supplies the five-qubit device, the Clifford-to-pi/2 gate decomposition, and the >99.99% single-qubit fidelity baseline that this work deepens.","marker":"13"},{"why":"Documents the microwave-heating-induced qubit frequency shift and the higher-temperature mitigation used to suppress it.","marker":"15"},{"why":"Provides the pulse-shaping framework behind the spectrally narrow Gaussian pulses that remove off-resonant driving artifacts.","marker":"16"},{"why":"Gives the relaxation relation connecting spin-locking time to noise spectral densities, used to interpret reservoir proximity data.","marker":"20"},{"why":"Earlier measurement of induced frequency shifts in silicon spin qubits; the present device's smaller shift is compared against this baseline.","marker":"21"},{"why":"Introduces the purity-versus-randomized-benchmarking comparison that separates coherent from incoherent error.","marker":"22"},{"why":"Demonstrates purity benchmarking on silicon qubits approaching incoherent noise limits, the direct precedent for the error budget here.","marker":"23"},{"why":"Describes the rapid parity spin readout whose joint-spin sensitivity makes off-resonant excitation of the reference qubit visible in randomized benchmarking.","marker":"14"}],"fun_headline_variants":["Silicon spin qubit gate fidelity passes 99.999%","Spin qubit: 99.999% fidelity, decoherence-limited","Reservoir removal and pulse shaping yield 99.999% qubit gates","Driven silicon spin qubit reaches 99.9992% gate fidelity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline fidelity assumes the noise that makes randomized benchmarking decay is the same from pulse to pulse and has no memory, and that the error per Clifford gate divides evenly among the average 3.25 $\\pi/2$ rotations; if the noise remembers previous pulses or favors certain rotations, the quoted 99.99920(2)% may not equal the true average gate fidelity.","fun_headline_variants_meta":{"raw":{"variants":["Silicon spin qubit gate fidelity passes 99.999%","Spin qubit: 99.999% fidelity, decoherence-limited","Reservoir removal and pulse shaping yield 99.999% qubit gates","Driven silicon spin qubit reaches 99.9992% gate fidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000337,"raw_usage":{"total_tokens":1887,"prompt_tokens":994,"completion_tokens":893,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":811}},"tokens_in":610,"tokens_out":893,"duration_ms":7911,"temperature":1.0,"reasoning_tokens":811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:49:42.218374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run interleaved randomized benchmarking experiments that isolate $X_{\\pi/2}$ and $Y_{\\pi/2}$ rotations, or run randomized benchmarking with only $X_{\\pi/2}$ gates; if the inferred per-rotation error differs from $8.6(3)\\times10^{-6}$ by more than the error bars, the equal-splitting, gate-independent assumption fails and the headline fidelity is not the average. A second check is to look for curvature or non-exponential tails in the randomized-benchmarking decay, which would signal noise that remembers previous pulses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the microwave-heating-induced qubit frequency shift and the higher-temperature mitigation used to suppress it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pulse-shaping framework behind the spectrally narrow Gaussian pulses that remove off-resonant driving artifacts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the relaxation relation connecting spin-locking time to noise spectral densities, used to interpret reservoir proximity data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier measurement of induced frequency shifts in silicon spin qubits; the present device's smaller shift is compared against this baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the purity-versus-randomized-benchmarking comparison that separates coherent from incoherent error."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the rapid parity spin readout whose joint-spin sensitivity makes off-resonant excitation of the reference qubit visible in randomized benchmarking."}],"review_version":1}