{"id":"081a8e9a-fd5f-4a22-963b-6d00ae4c17d5","arxiv_id":"2607.07586","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":8,"one_line_summary":"Numerical leakage-aware randomized benchmarking shows that operating a Si:P donor spin system as a native ququart (C4 Clifford group) yields 40-50% lower error rates than encoded two-qubit operation (C2^⊗2) under charge noise, due to reduced circuit complexity.","lead":"This paper numerically simulates a silicon-phosphorus donor spin system operated as a native four-level qudit (ququart) versus an encoded two-qubit system, finding ~40-50% lower error rates for the qudit under charge noise. A generalist might read it to understand whether quantum computers should use higher-dimensional units of information instead of standard two-level qubits.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The 40-50% quantitative advantage is robust to noise model changes in direction but not necessarily in magnitude; the structural argument (fewer operations per Clifford) is the load-bearing claim and is largely noise-independent.","rationale":"The reader correctly identified the noise model as the weakest assumption. My independent reading confirms this is the most load-bearing concern. The paper's structural argument—C4 requires fewer generators, ESR pulses, EDSR pulses, and ramps per Clifford element (Fig. 3)—is a group-theoretic property independent of noise and provides strong support for the qualitative claim. The quantitative 40-50% figure, however, is computed under a single specific noise configuration (one symmetric TLF, quasi-static regime) and its robustness to noise model variations is not tested. The paper is transparent about the quasi-static nature of the noise and frames itself as a 'case study,' which is appropriate. The driving parameters are conservative but applied identically to both groups, so the relative comparison is fair. The use of leakage-aware RB (Ref. 37) is appropriate given the reversible leakage in this system. No code is shipped, which is a reproducibility gap for a numerical study but does not affect soundness. The verdict of CONDITIONAL with MODERATE confidence is appropriate: the structural advantage is well-supported, but the quantitative magnitude needs sensitivity analysis to be fully established. The paper would be strengthened by the concrete test proposed above, but the current evidence is sufficient to support the central qualitative claim.","tokens_in":14693,"tokens_out":3641,"duration_ms":121982,"concrete_test":"Rerun the leakage-aware RB simulation (same gate sets, same driving parameters, same ramp type) with f_RTN swept from 1 Hz to 1 MHz, and additionally with 2-3 independent TLFs with amplitudes drawn from a distribution. If the ratio ε^LB_PT(C4)/ε^LB_PT(C2^⊗2) stays within 0.5-0.65 (i.e., the 40-50% reduction holds) across these conditions, the quantitative claim is robust. If the ratio crosses 0.7 for any parameter set, the 40-50% figure should be reported as regime-specific rather than general.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader correctly identifies the noise model as the weakest assumption. The paper uses a single symmetric TLF with A_TLF = 10 V/m and f_RTN = 10 Hz (Sec. II.A, Eq. 5), which is deep in the quasi-static regime: the mean dwell time (100 ms) vastly exceeds gate durations (~microseconds), so the TLF acts as a random static detuning offset per RB shot. The authors acknowledge this explicitly (Sec. V.A: 'charge noise is quasistatic with very rare switching events... the TLF therefore acts primarily as a detuning offset during a given RB shot'). In this regime, the per-gate error is approximately proportional to gate duration × noise sensitivity at the operating point, and total sequence error scales with total operation count. Since Fig. 3 shows C4 requires fewer ESR pulses, fewer EDSR pulses, and fewer ramps per Clifford element, the structural advantage is expected to hold under any quasi-static noise model. However, the quantitative 40-50% figure is specific to the chosen A_TLF and the single-fluctuator assumption. Under faster switching (f_RTN comparable to inverse gate durations), motional narrowing effects could change the per-operation error rates differently for ESR (interface, low sensitivity) vs EDSR (ionization point, high sensitivity) operations, potentially altering the ratio. Similarly, multiple fluctuators with distributed parameters would broaden the noise distribution and could shift the magnitude. The concern lands on the quantitative claim, not the qualitative one. The paper's own framing ('structural rather than an artifact of driving conditions') is supported by the resource-count analysis independent of noise, but the specific 40-50% number is not established as robust to noise model variations.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This manuscript presents a numerical fidelity analysis comparing native ququart Clifford group C4 operation versus encoded two-qubit Clifford group C2^⊗2 operation on a Si:P donor spin system. The Hamiltonian (Eqs. 1-3) follows Tosi et al. [26], with control via ESR pulses near the interface and EDSR pulses at the ionization point, connected by adiabatic displacement ramps. Three ramp types (linear, raised cosine, K-adiabatic) are compared. The authors employ leakage-aware randomized benchmarking following Chen & Baldwin [37] to extract population-transfer decay parameters r_PT under a single symmetric two-level fluctuator (TLF) noise model (Eq. 5, A_TLF = 10 V/m, f_RTN = 10 Hz). The central finding is that C4 consistently achieves approximately 40-50% lower lower-bound error rates ε^LB_PT compared to C2^⊗2, attributed to reduced circuit complexity (fewer ESR pulses, EDSR pulses, and ramps per Clifford element, as shown in Fig. 3). The authors argue this advantage is structural rather than an artifact of driving conditions.","tokens_in":15117,"tokens_out":1899,"duration_ms":298549,"significance":"The manuscript addresses a well-motivated question: whether native qudit operation provides fidelity advantages over encoded qubit operation within the same Hilbert space, specifically for donor spin systems. The use of leakage-aware RB (following Ref. [37]) is appropriate for the system under study, where leakage outside the computational subspace is non-negligible. The comparison protocol is fair in that the same random seeds and noise realizations are used across gate sets. The resource-count analysis (Fig. 3) provides a transparent, group-theoretic explanation for the observed fidelity difference. The K-adiabatic ramp construction (Eqs. 9-11) is a reasonable adaptation of prior work [26] to the multi-ramp context. The structural argument—that C4 requires fewer operations per Clifford element (768 vs 11,520 elements)—is largely noise-independent and is the strongest aspect of the paper. However, the quantitative 40-50% figure is specific to the chosen noise model parameters.","major_comments":[{"comment":"Sec. II.A, Eq. (5) and Sec. V.A: The noise model uses a single symmetric TLF with A_TLF = 10 V/m and f_RTN = 10 Hz, placing it deep in the quasi-static regime (mean dwell time ~100 ms >> gate durations ~microseconds). The authors acknowledge this (Sec. V.A: 'the TLF therefore acts primarily as a detuning offset during a given RB shot'). In this regime, per-gate error scales approximately with gate duration times noise sensitivity, and total sequence error scales with total operation count. This means the 40-50% quantitative advantage is largely a consequence of the resource counts shown in Fig. 3, and the specific A_TLF value mainly sets the overall error scale. The manuscript should state more explicitly that the quantitative figure is specific to this noise model and that the structural advantage (fewer operations) is the noise-independent claim. A brief sensitivity analysis varying A_","section":null},{"comment":"Sec. V.C, Fig. 4d: The lower-bound error ε^LB_PT = 1-(1+3r_PT)/4 is plotted for both groups as a function of ramp duration τ. The text states 'there is a consistent 40-50% reduction of the lower bound error for C4 compared to C2^⊗2.' However, the error bars (99% confidence intervals, shown as shaded regions in panels a-c but not explicitly in panel d) are not reported for panel d. Given that the absolute error rates are on the order of 3-8% (from Table I), the statistical significance of the 40-50% reduction should be verified. Please add confidence intervals or error bars to Fig. 4d and confirm that the reduction is statistically significant across the full range of τ.","section":null},{"comment":"Sec. IV, Fig. 3: The resource-count comparison (panels a-d) shows distributions over all Clifford elements, but the connection between resource counts and the RB error rates is only qualitatively stated. The text says 'C2^⊗2 requires not only more ESR pulses, but also more EDSR pulses and more ramps' (Sec. V.C). Since ESR and EDSR pulses have different noise sensitivities (ESR at interface vs EDSR at ionization point), and ramps contribute nonadiabatic leakage, a more quantitative decomposition of the total error into contributions from each operation type would strengthen the causal claim. At minimum, the mean operation counts (ESR, EDSR, ramps) per Clifford for each group should be reported numerically, not only visually in the violin plots.","section":null}],"minor_comments":[{"comment":"Sec. II.A, Eq. (5): The notation switches between A_TLF (in Eq. 5 and Sec. V.A) and A_RTN (in Sec. V.A text: 'A_RTn = 10 V/m'). Please use consistent notation.","section":null},{"comment":"Sec. V.B, Fig. 4 caption: Panel (a) is labeled 'S_PT,4(m)' but the y-axis label reads 'S_PT, 4 (m) (%)'. The subscript formatting is inconsistent across the figure.","section":null},{"comment":"Sec. III.B, Eq. (6): The function s(α) is introduced, but in Eq. (10) the notation switches to s as the integration variable for ∆E_z(s). This is potentially confusing since s(t) in Eq. (5) denotes the TLF state. Consider using a different symbol for one of these.","section":null},{"comment":"Table I: The column header 'S(m) / S_PT(m)' is ambiguous—presumably S(m) is the standard survival and S_PT(m) is the leakage-aware version, but this should be stated explicitly. Also, the F_RB column appears to report standard RB fidelity while F_PT reports the leakage-aware bounds; the relationship to Eqs. (27) and (32)-(33) should be clarified.","section":null},{"comment":"Sec. V.B: The text mentions 'n_seeds = 10 and n_trials = 150' but the caption of Fig. 4 says 'n_seeds = 10, n_trials = 150.' In Sec. V.C, the text says '10 RB seeds of 150 trials each.' These are consistent but the total sample size of 1500 per τ should be stated once clearly.","section":null},{"comment":"Sec. IV, Eq. (14): The gate decomposition U_Gate(ϕ1,...,ϕ4; θ1,...,θ6) uses subscripts on Y_i that refer to the {|i-1⟩,|i⟩} subspace, but the relationship between the index i (1-3) and the six θ parameters is not immediately clear. A brief clarification would help readers unfamiliar with Ref. [21].","section":null},{"comment":"Sec. II: The detuning field is defined as ∆E_z ≡ E_z - E^0_z, but the orbital Hamiltonian (Eq. 2) uses (E_z - E^0_z)/h. The factor of h (Planck's constant) in Eq. (2) versus its absence in the definition of ∆E_z should be clarified—presumably ∆E_z is in units where h is absorbed, but this is not stated.","section":null},{"comment":"Fig. 2b: The y-axis label 'P_initial' is unclear. It should be 'Survival probability' or 'P_survival' for consistency with the text.","section":null},{"comment":"Sec. V.A: The phrase 'and, which ideally leads to the final state being |ψ_init⟩ = |0⟩' contains a grammatical error (stray comma after 'and').","section":null}],"recommendation":"minor_revision","confidential_remarks":"The reader's assessment of the noise model as the weakest assumption is correct, but on reading the paper the concern is somewhat mitigated by the authors' own acknowledgment (Sec. V.A) that the TLF acts as a quasi-static detuning offset. In this regime, the structural advantage from fewer operations (Fig. 3) is indeed largely noise-independent in direction, and the quantitative magnitude is primarily determined by operation counts rather than noise spectral shape. The paper would benefit from making this argument more explicitly rather than leaving it implicit. The paper is a solid numerical case study; the main question for the editor is whether the journal's scope includes numerical simulation studies of this type, as there are no new experimental results."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee correctly identifies the structural resource-count argument as the strongest, noise-independent aspect of the paper and raises three points: (1) the quantitative 40-50% figure is specific to the chosen noise parameters and should be stated more explicitly, with a sensitivity analysis; (2) confidence intervals are missing from Fig. 4d and statistical significance should be verified; (3) mean operation counts should be reported numerically and a quantitative error decomposition by operation type would strengthen the causal claim. We agree with all three points and will revise accordingly.","responses":[{"response":"The referee's analysis is correct. In the quasi-static regime (mean dwell time ~100 ms >> gate durations ~microseconds), per-gate error scales approximately with gate duration times noise sensitivity, and total sequence error scales with total operation count. The 40-50% figure is therefore largely a consequence of the resource counts shown in Fig. 3, with A_TLF setting the overall error scale rather than the ratio between the two gate sets. We will revise Sec. V.A and the conclusion to state explicitly that (i) the quantitative 40-50% reduction is specific to the chosen noise model parameters (A_TLF = 10 V/m, f_RTN = 10 Hz), and (ii) the structural advantage—fewer operations per Clifford element for C4 versus C2^⊗2—is the noise-independent claim. We will also add a brief sensitivity analysis varying A_TLF over a range of values (e.g., 1-50 V/m) to demonstrate that the ratio of error rates between the two groups remains approximately constant while the absolute error scale changes, confirming that A_TLF sets the overall scale but not the relative advantage.","revision_made":"yes","referee_comment":"Sec. II.A, Eq. (5) and Sec. V.A: The noise model uses a single symmetric TLF with A_TLF = 10 V/m and f_RTN = 10 Hz, placing it deep in the quasi-static regime. The 40-50% quantitative advantage is largely a consequence of resource counts, and the specific A_TLF value mainly sets the overall error scale. The manuscript should state more explicitly that the quantitative figure is specific to this noise model and that the structural advantage is the noise-independent claim. A brief sensitivity analysis varying A_TLF is requested."},{"response":"We agree that confidence intervals should be shown in Fig. 4d. The 99% confidence intervals are available from the fitting procedure used for panels a-c but were omitted from panel d for visual clarity. We will add shaded confidence regions to Fig. 4d. Based on our data (10 seeds × 150 trials = 1500 samples per τ value), the confidence intervals for C4 and C2^⊗2 do not overlap across the full range of τ shown, confirming that the 40-50% reduction is statistically significant. We will state this explicitly in the revised text.","revision_made":"yes","referee_comment":"Sec. V.C, Fig. 4d: The error bars (99% confidence intervals) are not reported for panel d. Given that the absolute error rates are on the order of 3-8%, the statistical significance of the 40-50% reduction should be verified. Please add confidence intervals or error bars to Fig. 4d and confirm that the reduction is statistically significant across the full range of τ."},{"response":"We agree that the mean operation counts should be reported numerically. We will add a table reporting the mean ESR pulses, EDSR pulses, and ramps per Clifford element for both C4 and C2^⊗2 alongside the existing Fig. 3. Regarding the quantitative error decomposition by operation type: a full decomposition would require separate RB experiments isolating each operation type (ESR-only, EDSR-only, ramp-only sequences), which is not straightforward within the Clifford RB framework since each Clifford element contains a mixture of operation types. However, we can provide a partial decomposition by noting that ESR pulses operate near the interface (low charge-noise sensitivity), EDSR pulses operate at the ionization point (high sensitivity), and ramps contribute primarily nonadiabatic leakage. We will add a paragraph in Sec. V.C that qualitatively discusses these different noise sensitivities and their relative contributions to the total error, and we will note that a full quantitative decomposition is a direction for future work. We believe the numerical mean counts plus this discussion sufficiently strengthen the causal claim for the present manuscript.","revision_made":"partial","referee_comment":"Sec. IV, Fig. 3: The connection between resource counts and RB error rates is only qualitatively stated. A more quantitative decomposition of the total error into contributions from each operation type would strengthen the causal claim. At minimum, the mean operation counts (ESR, EDSR, ramps) per Clifford for each group should be reported numerically, not only visually in the violin plots."}],"tokens_in":14810,"tokens_out":1056,"duration_ms":153096,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper applies the C4-vs-C2^⊗2 Clifford comparison (previously done for transmons by Seifert et al.) to Si:P donor spins with ESR/EDSR control and adiabatic displacement ramps, and finds ~40-50% lower error rates for the native ququart encoding. The structural argument — fewer pulses and ramps per Clifford element for C4 — is sound and essentially noise-independent. The specific 40-50% number is tied to their single-fluctuator noise model and should be read as illustrative, not predictive of real devices. What's genuinely new is the platform-specific treatment: the adiabatic ramp design (linear, raised cosine, K-adiabatic), the operating-point strategy (parking at the interface for ESR, ionization point only for EDSR), and the application of leakage-aware RB to this system. The ramp-shape comparison in Fig. 2b is a nice concrete result — the K-adiabatic ramp clearly suppresses survival-probability oscillations. The resource-count analysis (Fig. 3) is straightforward group theory but does the job of making the structural advantage visible. The RB protocol follows Chen & Baldwin's leakage-aware framework correctly, and the comparison is fair: same seeds, same noise realizations across gate sets. The soft spot is the noise model, and the reader and stress-test are right to flag it. A single symmetric TLF with A_TLF = 10 V/m and f_RTN = 10 Hz is deep in the quasi-static regime — the TLF is effectively a random static detuning offset per shot. The authors acknowledge this explicitly. In this regime, total error scales roughly with total operation count × per-operation noise sensitivity, so the C4 advantage is expected to survive any quasi-static noise model. What could change the magnitude: faster switching (where motional narrowing affects ESR and EDSR differently), multiple fluctuators with distributed parameters, or 1/f noise. None of these would reverse the sign of the effect, but the 40-50% figure could shrink or grow. A sensitivity sweep over f_RTN and A_TLF, or a quick multi-fluctuator check, would have strengthened the paper considerably. No code is shipped, which is a minor issue for a simulation study but would help reproducibility. This is a well-constructed numerical study with an honest framing of its limitations. The central claim holds. It deserves a serious referee who should push for at least one robustness check on the noise model parameters before acceptance.","headline":"Solid comparative simulation; structural advantage is real, quantitative magnitude is noise-model-dependent","tokens_in":15545,"tokens_out":1119,"would_cite":false,"duration_ms":60053,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Native ququart gates cut donor-spin errors 40–50% vs encoded qubits","keywords":["donor spin","qudit","ququart","randomized benchmarking","silicon phosphorus","Clifford group","leakage","charge noise"],"falsifier":"If the number of ESR pulses, EDSR pulses, and displacement ramps per Clifford element were found to be comparable between C4 and C2⊗2 under a different decomposition scheme, or if the noise model were changed such that the additional operations required by C2⊗2 did not meaningfully increase error exposure, the 40–50% advantage would not hold.","tokens_in":14905,"feed_emoji":"🎚️","tokens_out":923,"duration_ms":159970,"temperature":0.7,"pith_summary":"This paper argues that operating a phosphorus-donor spin system in silicon as a genuine four-level qudit (a ququart) is structurally more efficient than forcing the same four levels to behave as two encoded qubits. The central object is the Si:P donor spin system, whose electron and nuclear spin define four computational basis states. The authors compare two ways to run Clifford gates on this system: the native ququart Clifford group C4, which treats all four levels as one unit, and the encoded two-qubit Clifford group C2⊗2, which partitions the four levels into two logical qubits. Using numerical leakage-aware randomized benchmarking under a charge-noise model (a single two-level fluctuator producing random telegraph noise), the authors find that C4 consistently achieves 40–50% lower lower-bound error rates than C2⊗2. The mechanism is circuit complexity: the native ququart gate set requires fewer ESR pulses, fewer EDSR pulses, and fewer displacement ramps per Clifford element, so the system spends less time exposed to noise and leakage. The authors also show that adiabatic displacement ramps, which shuttle the electron between the ionization point (for EDSR) and the interface (for ESR), are essential for suppressing leakage outside the computational subspace, with a K-adiabatic ramp informed by avoided crossings performing best. The advantage is attributed to the structure of the gate sets themselves rather than to any particular choice of driving parameters.","feed_headline":"Native ququart gates cut donor-spin errors 40–50% vs encoded qubits","feed_subtitle":"Si:P spin system needs fewer pulses as a genuine four-level qudit than as two forced qubits, cutting noise exposure.","key_machinery":"The comparison hinges on two gate sets defined on the same four-level Hilbert space: C4 (768 elements, generated by the qudit Fourier gate F4, phase gate S4, and Z4) versus C2⊗2 (11,520 elements, generated by Hadamard, S, and CNOT gates on two encoded qubits). Both are decomposed into Givens rotations (two-level subspace rotations) implemented by ESR and EDSR pulses, with virtual Z-gates. Leakage-aware randomized benchmarking extracts a population-transfer decay parameter r_PT that bounds the average gate fidelity from below, accounting for reversible leakage into excited orbital states. Adiabatic displacement ramps (linear, raised cosine, and K-adiabatic profiles) shuttle the electron to a ","core_discovery":"The native ququart Clifford group C4 requires fewer physical operations (ESR pulses, EDSR pulses, and displacement ramps) per Clifford element than the encoded two-qubit Clifford group C2⊗2, and this structural economy translates into a consistent 40–50% reduction in lower-bound error rates under charge noise. The comparison is made through leakage-aware randomized benchmarking on a Si:P donor spin system, where adiabatic ramps suppress leakage by positioning the electron at the ionization point only during EDSR control and near the interface during ESR control.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Native ququart operation cuts donor spin errors up to 50% vs qubits","Genuine qudit control reduces Si:P donor spin errors 40-50%","Fewer pulses make Si:P ququarts less error-prone than encoded qubits","Native C4 gates yield lower error rates than encoded two-qubit systems","Donor spin qudits achieve 40-50% lower error than encoded qubit pairs"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The noise model uses a single symmetric two-level fluctuator with a fixed amplitude and switching frequency, acting as a quasistatic detuning offset per experimental shot. Real Si:P devices exhibit 1/f charge noise with multiple fluctuators and distributed parameters, so the quantitative magnitude of the 40–50% advantage could shift under different noise spectra, though the structural advantage from fewer operations would likely persist.","fun_headline_variants_meta":{"raw":{"variants":["Native ququart operation cuts donor spin errors up to 50% vs qubits","Genuine qudit control reduces Si:P donor spin errors 40-50%","Fewer pulses make Si:P ququarts less error-prone than encoded qubits","Native C4 gates yield lower error rates than encoded two-qubit systems","Donor spin qudits achieve 40-50% lower error than encoded qubit pairs"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1508,"prompt_tokens":590,"completion_tokens":918,"prompt_tokens_details":null},"tokens_in":590,"tokens_out":918,"duration_ms":56471,"temperature":1.0,"reasoning_tokens":740,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T05:58:05.899880+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the number of ESR pulses, EDSR pulses, and displacement ramps per Clifford element were found to be comparable between C4 and C2⊗2 under a different decomposition scheme, or if the noise model were changed such that the additional operations required by C2⊗2 did not meaningfully increase error exposure, the 40–50% advantage would not hold.","supporting_citations":[],"review_version":1}