{"id":"8c58a852-8845-4a33-82ab-db6db449c104","arxiv_id":"2506.09131","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Qutrit-enabled protocols reduce the gauge ambiguity in SPAM and gate Pauli noise characterization by using extra energy levels to tighten positivity constraints, as shown theoretically and experimentally.","lead":"This paper proposes using higher energy levels of a quantum device, such as qutrits, to reduce the ambiguity in characterizing state preparation and measurement (SPAM) noise, and demonstrates the approach on a superconducting processor. The technique could make quantum noise characterization and error mitigation more precise on current hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The qutrit-enhancement claim rests on noiseless single-qutrit control (SM Eq. S6); with realistic qutrit gate error the gauge transformation no longer preserves the noise model, so the claimed reduction in gauge ambiguity lacks quantitative support.","rationale":"The reader's weakest_assumption identifies exactly the assumption I find most load-bearing: noiseless (or gate-independent and absorbable) single-qudit control. The paper's Proposition 1 and the experimental ambiguity reductions are both consequences of this assumption, and the Discussion explicitly acknowledges the limitation. My concern does not move the verdict because the paper already receives a CONDITIONAL verdict based on this issue; the condition should be that the authors quantitatively validate qutrit gate fidelity or show that the enhancement survives realistic qutrit gate noise. I agree with the reader that the core idea is coherent and the experimental demonstration is suggestive, but the central quantitative claim is not yet supported with rigorous treatment of the dominant systematic. I would not escalate to REJECT: the theoretical framework is self-contained, the first-order gauge analysis is explicit, and the limitation is honestly stated. The missing piece is a concrete error-budget analysis for the qutrit gates, which is addressable and would settle the concern.","tokens_in":12600,"tokens_out":6284,"duration_ms":72774,"concrete_test":"On the same device used for Figs. 1–3, measure the per-gate error rates of the qutrit gates X_{0,2}, X_{1,2}, and any other generalized X gates used in the protocol, e.g., via qutrit randomized benchmarking. Then simulate the qutrit-enhanced SPAM protocol with a noisy-gate model (each ideal gate followed by a depolarizing or amplitude-damping/leakage channel with the measured rate r), and recompute the positivity-constrained feasible interval for the qubit-subspace SPAM parameters. If the interval width is not smaller than the qubit-only protocol for the measured r, the central claim fails; at minimum, report the crossover gate-error rate r* at which the qutrit-enhanced interval equals the qubit-only interval and add a systematic error bar to Figs. 1–3 that includes this term.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that qutrit-enhanced protocols reduce SPAM and gate-noise gauge ambiguity. The whole gauge-counting argument and the positivity bounds that produce the ambiguity intervals rely on SM Eq. (S6), where the gauge transformation ρ0 → Λ_{Ω,p}(ρ0), E_l → Λ^{-1}_{Ω,p}(E_l) leaves all observables unchanged because every single-qudit unitary commutes with the depolarizing channel (SM Eq. S7). If the generalized X gates used in the protocol are noisy, this conjugation no longer maps the gate set to itself: a noisy X gate is not transformed into the same noiseless unitary, the equivalence class used to count 2^n−1 gauge DOFs breaks, and the first-order equations (SM Eq. S15) acquire extra contributions from qutrit gate errors. The paper's own Discussion concedes this: 'single-qudit gate might be more noisy due to their complexity,' and leakage/seepage is listed as a possible source of the experimental discrepancy. However, no qutrit gate fidelity is reported for the device used in Figs. 1–3, and no sensitivity analysis shows how the reported ambiguity reduction degrades as a function of qutrit gate error. Because the claimed enhancement is quantitatively the difference between qubit-only and qutrit-enhanced intervals, a per-gate qutrit error comparable to that difference could erase or invert the claimed benefit. This is not a stylistic simplification: qutrit gates are typically less coherent than qubit gates and often leak, so the noiseless-control assumption is a genuine correctness risk.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using higher energy levels (qudits) to reduce gauge ambiguities in qubit-subspace SPAM and Pauli-gate noise characterization. The authors develop a first-order gauge theory for n-qudit SPAM noise under the assumption of noiseless single-qudit control, proving that the gauge freedoms are exactly the 2^n−1 subsystem depolarizing transformations and giving an explicit protocol with at most 2d^n circuits that determines all learnable parameters. They then use positivity constraints to bound the residual gauge ambiguity, arguing that qutrit levels with smaller populations tighten the bound. Experimental results on one- and two-superconducting-qutrit devices compare qubit-only and qutrit-enhanced protocols, showing smaller ambiguity intervals for the latter, and a gate-noise application bounds non-identifiable Pauli fidelities for CZ more tightly using qutrit-enhanced SPAM data.","tokens_in":12892,"tokens_out":8253,"duration_ms":81247,"significance":"If the central claim holds, this is a practical and elegant contribution: it extends the self-consistent Pauli-noise-learning framework to qudit systems, gives a proof of the gauge-DOF count, and demonstrates a concrete use of higher levels already present in superconducting devices. The positivity-based ambiguity bound (Eq. (9)) and the explicit circuit constructions in the SM are useful. The paper also provides an experimental demonstration, which strengthens the case. The main caveat is that the entire gauge analysis assumes noiseless single-qudit gates; the authors acknowledge this but do not quantify how the benefit degrades with qutrit gate error. The proof of Proposition 1 is explicit and the protocol is concrete, which is a strength, as is the inclusion of experimental data rather than only numerics.","major_comments":[{"comment":"The noiseless single-qudit control assumption is load-bearing for the central claim. The gauge transformation in Eq. (S6) preserves the gate set only because Λ_p commutes with every single-qudit unitary (S7). When the generalized X gates used in the protocol are noisy, conjugation by the gate no longer maps the noise model to itself, and the first-order equations (S15) acquire additional contributions from qutrit gate errors. The Discussion explicitly states that single-qudit gates 'might be more noisy due to their complexity' and lists leakage as a possible source of the experimental discrepancy. However, the paper does not report qutrit gate fidelities for the device used in Figs. 1–3 and does not provide a sensitivity analysis showing how the reported reduction in gauge ambiguity depends on single-qutrit gate error. Since the claimed enhancement is the difference between qubit-only and qutrit-enhanced intervals, a per-gate qutrit error comparable to that difference could erase the benefit. This needs to be quantified, e.g., by reporting gate fidelities and by numerical simulation of the protocol with finite single-qutrit gate error.","section":"SM Eq. (S6)–(S7); 'Single-qudit SPAM characterization'"},{"comment":"The experimental demonstration does not provide statistical uncertainty on the ambiguity reduction itself. The light bars are described as one standard error for deciding the region, but the ambiguity width is a nonlinear function of the estimated error probabilities, involving minima (Eq. (9)); the plug-in estimate of a minimum is biased, and the reported reduction could be within noise. The authors should provide confidence intervals or bootstrap/tolerance intervals for the ambiguity intervals, or a statistical test comparing the qubit-only and qutrit-enhanced widths.","section":"Figs. 1–3 and Eq. (9)"},{"comment":"The quantitative claims are first-order in ε, but the manuscript does not bound the size of the neglected O(ε^2) terms for the experimentally measured error rates (up to about 3.5%). Figure 4's blue shaded region shows a visible spread attributed to the first-order approximation, indicating that these corrections are not always negligible relative to the reported ambiguity differences. The authors should provide an estimate or upper bound on second-order contributions for the parameters in Figs. 1–3, or validate the first-order bounds against a higher-order or numerical calculation for those parameters.","section":"Eqs. (5), (8), (9); Figs. 3–4"}],"minor_comments":[{"comment":"The sentence 'We can all the independent noise parameters in a vector v' contains a typo and should read 'We can call' or 'We can collect'.","section":"SM text before Eq. (S8)"},{"comment":"The phrase 'the experiment of X_{π*(k)(0)}' appears to be a typo; the probability expression in Eq. (S20) uses the gate X_{π*(k)}, not X_{π*(k)(0)}. Please clarify.","section":"SM paragraph after Eq. (S20)"},{"comment":"The notation '(0,h_i)' in the definition of π*(k) is not defined in that context; please specify the transposition or cycle notation used for the permutation.","section":"SM Eq. (S19)"},{"comment":"The caption says the blue shaded region is 'due to the first-order approximation'; since the max/min curves also differ because of the remaining gauge choices, the caption should distinguish the two sources of spread.","section":"Fig. 4 caption"},{"comment":"The manuscript does not report standard experimental details such as the number of shots, single-qutrit gate decompositions, device calibration parameters, or the exact procedure used to assign the light error bars in Figs. 1–3; providing these details in the Supplemental Materials would improve reproducibility.","section":"Experimental sections"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a quantum information journal. The referee's main concern is that the central experimental claim rests on an unquantified assumption (noiseless qutrit gates) and lacks uncertainty quantification on the ambiguity reduction. If the authors can supply the requested gate-fidelity data, sensitivity analysis, and error bars on the ambiguity intervals, the paper would be a solid contribution. The Note added correctly discloses concurrent work; the editor may wish to verify the status of Ref. [28]."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The qutrit-enhancement idea is real, and the theory is clean. The new piece is Proposition 1: for n qudits with perfect single-qudit control, the SPAM gauge DOFs are exactly the 2^n-1 subsystem depolarizing maps, and O(2 d^n) circuits suffice to learn everything else. That proof is self-contained in the SM. The application idea is simple and sensible: higher levels have smaller populations and smaller readout confusion, so positivity bounds pin the gauge harder. The experimental data in Figs. 1-3 show the expected effect: qutrit-enhanced protocols give narrower ambiguity intervals than qubit-only ones, both for SPAM and for non-identifiable Pauli fidelities of CZ. I believe the qualitative result.\n\nThe soft spots are real. The gauge-counting argument assumes noiseless single-qudit control, and that is the load-bearing assumption. If the generalized X gates on the qutrit are noisy, the conjugation in Eq. (S7) does not preserve the gate set, the gauge equivalence class breaks, and the derived first-order equations pick up extra terms from gate errors. The authors know this; the Discussion says 'single-qudit gate might be more noisy due to their complexity.' But they do not report qutrit gate fidelities for the device used in the experiments, nor do they provide a sensitivity analysis showing how the ambiguity reduction degrades with gate error. For a claim that is quantitatively about interval widths, that is a real gap. I also agree with the reader that the ambiguity intervals in the figures have no rigorous error bars on the ambiguity itself; the light bars are standard errors on parameter estimates, not on the gauge range.\n\nNone of this kills the paper. The theory is solid under the stated assumption, the assumption is common in this subfield, and the experimental demonstration is a proof-of-principle. But the central quantitative claim needs a direct check: measure qutrit gate errors, or at least bound them, and show the reported effect survives. That is exactly what a referee should ask for.\n\nThis paper is for anyone working on Pauli noise learning, gate set tomography, or error mitigation. It is a useful extension of the self-consistent framework. I recommend sending it to peer review, and asking the authors to add the missing sensitivity analysis and error bars before publication.","headline":"A clean theory of qudit SPAM gauge ambiguities, with a plausible experimental demonstration; the main risk is the noiseless-control assumption, which the authors flag but do not quantify.","tokens_in":13435,"tokens_out":2049,"would_cite":true,"duration_ms":20300,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68"],"pacs":["03.67.-a"],"model":"deepseek-v4-flash","headline":"Qutrit-enhanced protocols shrink the gauge ambiguity in SPAM and gate Pauli noise characterization, with experiments on superconducting devices.","keywords":["quantum noise characterization","SPAM errors","gauge freedom","qutrit","qudit","Pauli noise","superconducting qubits","noise identifiability"],"falsifier":"Run the qutrit-enhanced SPAM protocol on a device while deliberately tuning the fidelity of the |1>-|2> gate; if the estimated gauge ambiguity interval A does not grow as that gate fidelity drops, or if the inferred ε^S_j and ε^M_{l,k} drift, the noiseless-control premise is falsified. A simpler check: compare qutrit-enhanced estimates of qubit-subspace SPAM with and without randomized compiling on the qutrit gates; they should agree if the gate noise is gate-independent.","tokens_in":12407,"feed_emoji":"⚛️","tokens_out":5835,"duration_ms":54881,"temperature":0.7,"pith_summary":"This paper aims to show that the extra energy levels already present in superconducting devices—the |2>, |3>, ... levels above the qubit subspace—can be used to sharpen quantum noise characterization. The authors develop a theory of state-preparation-and-measurement (SPAM) noise for n qudits under perfect single-qudit control, proving that the unlearnable gauge degrees of freedom are exactly the subsystem depolarizing gauges, 2^n - 1 in total, and that all learnable parameters can be recovered with at most 2d^n circuits. They then use qutrit information to constrain the gauge through positivity of error probabilities, shrinking the ambiguity interval for qubit-subspace SPAM. Experiments on superconducting qutrits show that the qutrit-enhanced protocol yields smaller gauge ambiguity than qubit-only characterization, and that this improvement carries over to bounds on non-identifiable Pauli fidelities of CZ gates. If the central claim is right, noise characterization on existing hardware can be improved without cooling or special entangling gates, simply by using the levels already available.","feed_headline":"Qutrit levels shrink gauge ambiguity in quantum noise tests","feed_subtitle":"Using the extra |2> level of a transmon cuts noise-parameter ambiguity in both SPAM and gate characterization, with theory and hardware…","key_machinery":"The load-bearing object is the subsystem depolarizing gauge: for any subset Ω of the n qudits, the map D^Ω_p(·)=(1-p)(·)+p Tr_Ω(·)⊗ I_Ω/$d^{{|Ω|}}$ commutes with every parallel single-qudit unitary, so transforming ρ0 with D^Ω_p and each measurement operator with its inverse leaves all outcome distributions unchanged. The paper linearizes these transformations to first order and shows the 2^n - 1 non-empty subsets give linearly independent gauge directions; it then constructs a protocol using generalized Pauli-X gates (permutation gates X_π) that determines every parameter orthogonal to all gauge directions, with no more than 2d^n circuits. The positivity constraints—every ε^S_j and ε^M_{l,k} is a probability—turn the gauge directions into a bounded interval, and it is the size of that interval that qutrit levels shrink.","core_discovery":"For an n-qudit system whose SPAM noise is incoherent and whose single-qudit unitary control is perfect, the paper proves (Proposition 1) that the noise model has exactly 2^n - 1 gauge degrees of freedom, generated by subsystem depolarizing maps D^Ω_p; no experiment within the gate set can distinguish the transformation ρ0 → D^Ω_p(ρ0), E_l → (D^Ω_p)^{-1}(E_l). Up to first order, each gauge parameter is confined by positivity: the initialization and measurement error probabilities must stay non-negative. For a single qudit the resulting ambiguity is A = min_{j≠0} ε^S_j + min_{k≠l} ε^M_{l,k}; since the population of the |2> level of a superconducting transmon is much smaller than that of |1>, adding qutrit information tightens A. The paper experimentally demonstrates this on one- and two-qubit/qutrit superconducting devices, showing that qutrit-enhanced SPAM characterization reduces the gauge ambiguity compared to qubit-only protocols both with and without heralding. It further feeds the tighter SPAM bounds into intercept cycle benchmarking of a two-qubit CZ gate, producing smaller intervals for non-identifiable Pauli fidelities such as λ_XI.","pith_inferences":["If single-qutrit gate noise is gate-independent, it can likely be absorbed into the gauge framework just as single-qubit gate noise is absorbed in Pauli noise learning; testing this would require interleaving qutrit gate calibrations into the protocol.","The positivity-constraint mechanism might apply to leakage detection: population in |2> is itself a leakage signature, so the same extra-level data could jointly bound leakage and SPAM gauge in transmons.","A natural stress test is to run the protocol on a device with controllable |1>-|2> gate error; the predicted gauge interval should widen monotonically as that gate's error increases.","The gauge DOF count being independent of d suggests that using qutrits rather than qubits is always at least as informative even if higher levels are slightly noisier, because the added levels give more positivity constraints without adding new gauge directions."],"forward_implications":["The gauge ambiguity for a single qudit shrinks from roughly ε^S_1 + min{ε^M_{1,0}, ε^M_{0,1}} to the qutrit-constrained minimum over j≠0 and k≠l, which is much smaller on thermal transmon devices.","Tighter SPAM bounds directly tighten the estimate of Pauli fidelities λ_a for gates whose Clifford action changes the support of P_a, where λ_a is not individually identifiable.","All learnable SPAM parameters of an n-qudit system can be learned with at most 2d^n circuits, within a factor of two of the information-theoretic lower bound of d^n + 1.","Qutrit-enhanced characterization works alongside heralding; the two methods combine, since heralding reduces ε^S_1 and qutrit constraints further restrict the gauge.","The same framework applies to correlated multi-qubit SPAM noise: for two qutrits the three gauge DOFs are the two single-qudit depolarizing gauges plus the joint {1,2} depolarizing gauge."],"supporting_citations":[{"why":"Establishes which Pauli eigenvalues are identifiable and introduces intercept cycle benchmarking, which the present protocol builds on for gate noise characterization.","marker":"[3]"},{"why":"Provides the gauge-transformation framework for Pauli noise learning that the paper extends to qudit SPAM noise.","marker":"[10]"},{"why":"Prior algorithmic-cooling method for suppressing SP noise, used as a baseline and as the heralding procedure in the experiments.","marker":"[5]"},{"why":"Shows how noiseless entangling gates can resolve SPAM gauge ambiguity; the present work instead uses only single-qudit control.","marker":"[6]"},{"why":"Demonstrates high-fidelity single-qutrit control, the key assumption that makes the proposed protocol experimentally justified.","marker":"[11]"},{"why":"Randomized compiling justifies twirling generic noise into the incoherent and Pauli noise models used throughout.","marker":"[17]"},{"why":"Validates randomized compiling on a superconducting processor and supports the assumption that single-qubit control noise is much weaker than SPAM or multi-qubit gate noise.","marker":"[7]"},{"why":"Cycle error reconstruction (CER) provides the sanity-check comparison for identifiable Pauli eigenvalue combinations.","marker":"[18]"},{"why":"Supplies the thermal-population ordering ε^S_2 << ε^S_1 that underlies the qutrit benefit in superconducting transmons.","marker":"[14]"}],"fun_headline_variants":["Extra energy levels tighten quantum noise calibration","Qutrit levels cut ambiguity in quantum noise tests","Extra qudit levels narrow noise gauge freedom","Qutrit-enhanced noise calibration reduces ambiguity","Qutrits help pin down quantum noise parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that single-qudit control is noiseless (or that its noise is gate-independent and absorbable into the noise model), so if the extra-level gates themselves have appreciable noise, the claimed gauge reduction may not survive.","fun_headline_variants_meta":{"raw":{"variants":["Extra energy levels tighten quantum noise calibration","Qutrit levels cut ambiguity in quantum noise tests","Extra qudit levels narrow noise gauge freedom","Qutrit-enhanced noise calibration reduces ambiguity","Qutrits help pin down quantum noise parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000496,"raw_usage":{"total_tokens":2434,"prompt_tokens":951,"completion_tokens":1483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":1429}},"tokens_in":567,"tokens_out":1483,"duration_ms":12783,"temperature":1.0,"reasoning_tokens":1429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:56:05.041051+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the qutrit-enhanced SPAM protocol on a device while deliberately tuning the fidelity of the |1>-|2> gate; if the estimated gauge ambiguity interval A does not grow as that gate fidelity drops, or if the inferred ε^S_j and ε^M_{l,k} drift, the noiseless-control premise is falsified. A simpler check: compare qutrit-enhanced estimates of qubit-subspace SPAM with and without randomized compiling on the qutrit gates; they should agree if the gate noise is gate-independent.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes which Pauli eigenvalues are identifiable and introduces intercept cycle benchmarking, which the present protocol builds on for gate noise characterization."},{"cited_title":"Laflamme, J","cited_arxiv_id":null,"evidence_quote":"Prior algorithmic-cooling method for suppressing SP noise, used as a baseline and as the heralding procedure in the experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows how noiseless entangling gates can resolve SPAM gauge ambiguity; the present work instead uses only single-qudit control."},{"cited_title":"Morvan, V","cited_arxiv_id":null,"evidence_quote":"Demonstrates high-fidelity single-qutrit control, the key assumption that makes the proposed protocol experimentally justified."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates randomized compiling on a superconducting processor and supports the assumption that single-qubit control noise is much weaker than SPAM or multi-qubit gate noise."}],"review_version":1}