REVIEW 3 major objections 5 minor 1 cited by
Benchmarking Quantum Instruments
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A single long run of randomly compiled measurements directly exposes a quantum instrument's error rate through the decay of its all-success probability.
desk verdict A solid theory paper proving exponential decay for randomized compiled instruments, with a real ideal-gate caveat that the abstract overstates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is randomized compiling for instruments: before each measurement a random Weyl operator is applied, the measurement is performed, and a compensating random Weyl operator is applied to the post-measurement state, with the reported outcome derandomized so the ideal action is unchanged. Averaging over these random operations maps any noisy instrument $\mathcal{M}$ to a stochastic instrument $\hat{\mathcal{M}}$ whose components are a probability distribution $\nu_{a,b}$ over register shifts, and the generalized Pauli fidelities $\tilde{\nu}_{s,t}$—overlaps of the instrument's outcome maps with the Weyl basis—are invariant under this averaging. The key quantitative tool is Theorem 3's exponential-decay law: the all-zero survival probability is $\langle\!\langle I| \hat{\mathcal{M}}_0^m |\rho\rangle\!\rangle$, and a nonnegative-matrix eigenvalue bound (proved in the appendix via disk and oval eigenvalue estimates) guarantees a unique dominant eigenvalue $\mu = \nu_{0,0} + O(\epsilon^2)$ with all other eigenvalues bounded by $\epsilon$. This separation of eigenvalues is what turns a generic survival probability into a clean exponential whose base is the error rate.
What would settle it
Run Algorithm 1 on a device whose instrument error rate is independently known from process tomography: if the fitted decay base disagrees with $1 - \epsilon$ beyond the paper's $O(\epsilon^2)$ bound (with $\epsilon < 1/3$), or if the all-zero probability is not exponential, the central theorem is false. A second check is to deliberately add gate-dependent noise to the randomizing gates and observe that the decay base shifts, since the theorem requires ideal gates.
Extended reading notes
Core claim
The central claim is that the error rate of a noisy quantum instrument can be measured directly from the probability that $m$ successive randomly compiled measurements all return the ideal outcome. Under ideal randomizing gates and for error rate $\epsilon < 1/3$, the authors prove $\Pr(\vec 0) = A\mu^m + O(\binom{m}{\xi}\epsilon^{m-\xi})$ with $\mu = 1 - \epsilon + O(\epsilon^2)$, where $\xi + 1$ is the largest Jordan-block dimension of the relevant transfer matrix. The proof works by randomized compiling: averaging over random Weyl operators converts the noisy instrument into a stochastic channel whose error structure is a probability distribution $\nu_{a,b}$ over register shifts, and the diamond-distance error rate is exactly $\epsilon = 1 - \nu_{0,0}$. A matrix eigenvalue bound ensures the all-zero branch of the compiled channel has a unique dominant eigenvalue $\mu$, so fitting the all-success probability to an exponential estimates $\epsilon$. The paper further shows that generalized Pauli fidelities are invariant under randomized compiling and can be recovered from the same data; for single-qubit instruments the characterization is complete up to an unavoidable gauge transformation that swaps pre-measurement and post-measurement error probabilities.
Load-bearing premise
The exponential-decay guarantee assumes the randomizing gates used for compiling are themselves ideal or have errors that fold into the instrument, so if those gates carry significant gate-dependent noise, the fit slope no longer provably equals the measurement's error rate.
Editorial extensions
If this is right
- A single sequence of $m$ measurements yields the whole decay curve by marginalizing later outcomes, so no separate experiments per sequence length are needed.
- The fitted rate equals the diamond distance between the randomly compiled instrument and the ideal instrument, giving a state-independent figure of merit for syndrome measurements.
- With phase-based post-processing, the same dataset also estimates the generalized Pauli fidelities, giving a partial stochastic error model without full process or gate-set tomography.
- For single-qubit instruments, the total error probabilities $\nu_{0,0}$ and $\nu_{1,1}$ are unambiguous, while only the pre-measurement/post-measurement split of the remaining errors is gauge-ambiguous.
- The protocol needs no unmeasured idling qudits and no entangling gates, reducing the experimental overhead of earlier mid-circuit measurement benchmarks.
Reading between the lines
- The single-run property suggests a natural calibration mode: run one sufficiently long sequence on an idle qubit and extract the measurement error rate in real time, an operational use the paper does not develop.
- The gauge ambiguity is argued to be fundamental for a single qubit; a concrete multi-qubit extension would be to check whether inserting entangling gates between measurements resolves the $\nu_{0,1}/\nu_{1,0}$ split, as the paper hints in its discussion.
- A testable diagnostic suggested by the contrast between compiled and uncompiled runs is to compare the fitted decay base with and without randomized compiling: if they differ, the error rate is state-dependent, quantifying exactly what randomized compiling protects against.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a randomized-compiling protocol for characterizing noisy quantum instruments (nondestructive measurements with feed-forward). The main theoretical result is that, under ideal randomizing gates, the probability that m successive randomly compiled measurements all return the ideal outcome decays exponentially in m with base μ = 1 − ε + O(ε²), where ε = 1 − ν_{0,0} is the error probability of the randomly compiled instrument. The authors prove that generalized Pauli fidelities are invariant under randomized compiling and that they determine the compiled instrument, and they show how post-processing can estimate these fidelities up to a sign/gauge ambiguity. An appendix proves an eigenvalue bound for nonnegative matrices that is used to control the sub-dominant eigenvalues in the decay theorem.
Significance. If the ideal-gate limitation can be overcome, the protocol would be a significantly simpler alternative to existing mid-circuit measurement benchmarking: it requires no idling qudits, no entangling circuits, and no separate experiments for each sequence length, since data for different lengths are obtained by marginalizing a single long sequence. The proof of Theorem 8 is a clean and potentially reusable eigenvalue bound, and the invariance theorem plus the explicit gauge ambiguity are useful conceptual contributions. The central proofs are coherent given their stated assumptions, and the paper is honest about the ideal-gate caveat, but the practical claim in the abstract is broader than what is proven, and the absence of finite-sample analysis limits the protocol's immediate applicability.
major comments (3)
- [Section II, before Theorem 1; Algorithm 1] The only treatment of noise on the randomizing gates is the sentence 'gate-dependent errors require a perturbative treatment' immediately before Theorem 1, but no perturbative theorem, lemma, or bound is supplied. The proofs of Theorems 1-3 rely on exact Pauli twirling identities such as the second equality in eq. (18), which require the inserted gates X^x, Z^a, Z^b to be noiseless. In Algorithm 1, step 3b inserts the physical gate Z^{β_i} X^{α_{i-1}-α_i}; if that gate carries a gate-dependent error, the averaged channel is not of the form in eq. (20), the matrix N in Theorem 3 is not guaranteed to be a nonnegative matrix with entries summing to 1, and the fitted base μ can absorb errors from the randomizing gates as well as from the instrument. Since the abstract claims that the error rate 'of such a measurement' is directly estimated, the practical claim is broader than the proven statement. Please either add a quantitative perturbative analysis (e.g., a first-order bound with explicit constants) or explicitly restrict the stated scope to ideal or gate-independent randomizing gates.
- [Section III, Theorem 3 and Fig. 2] The protocol estimates Pr(0) from empirical frequencies and fits the exponential decay of Theorem 3, but the paper contains no finite-sample analysis: there are no confidence intervals, no sample-complexity bounds, no discussion of bias in the fitted μ, and no guidance on choosing m. Equations (25)-(27) are exact expectation statements, and Theorem 3 is an asymptotic statement about the true probability; converting a finite number of observed all-zero counts into an estimate of ε requires statistical guarantees. The numerical demonstration in Fig. 2 reports a single point estimate (ν_{0,0} ≈ 0.952) with no error bars, no noise model, and no fitting details, so it does not fill this gap. For a benchmarking protocol whose stated advantage is efficiency, this missing statistical analysis is load-bearing.
- [Section III, eqs. (40)-(43)] The gauge transformation B defined in eq. (43) is not completely positive for generic values of the fidelity ratio: for a single qubit, a Pauli-diagonal superoperator with eigenvalues (1, 1, 1, r) is CPTP only for r = 1. Therefore B M_k B^{-1} in the transformed instrument may not be a valid quantum instrument, and the two roots of eq. (40) may not both correspond to physical noise models. The claim that the sign ambiguity is 'fundamental' needs qualification: if it is meant in the gate-set-tomography gauge sense, that should be stated explicitly; if it is meant as an ambiguity among physical instruments, the authors should provide a proof or example that both roots satisfy complete positivity. Without this, the stated ambiguity between errors before or after the measurement may be resolvable by physicality constraints.
minor comments (5)
- [Abstract and Section III] The phrase 'all the data can be obtained using a single sufficiently large number of successive measurements' is potentially misleading: estimating a probability still requires many repetitions of the long sequence; what is avoided is the need for separate experiments at each sequence length. Please rephrase to clarify this.
- [Proof of Theorem 3, around eq. (28)] The bound on the entries of J^m does not by itself bound the term ⃗I† S J^m S^{-1} ⃗ρ; one needs to multiply by the condition number of the eigenvector matrix S. The O() statement remains correct with a constant that depends on N, but the proof should state this explicitly.
- [Fig. 2] The simulation figure should include error bars, a description of the simulated noise model (e.g., which ν_{a,b} are nonzero), and the fitting procedure used to obtain ν_{0,0} ≈ 0.952.
- [Section IV] The sentence 'This optimization is not possible in the protocols outlined in from [16–18]' contains a typo ('in from'); it should read 'outlined in [16–18]'.
- [Theorem 8] The phrase 'unique maximal eigenvalue' is imprecise; the theorem shows uniqueness in the leading Gershgorin disc and dominance in modulus for ε < 1/3. Please reword to 'unique eigenvalue in the leading Gershgorin disc, which is dominant in modulus under the stated condition'.
Circularity Check
No significant circularity: the exponential-decay estimate of the instrument error rate is derived, not fitted as an input.
full rationale
The paper's central claim is Theorem 3, which derives Pr(0) = A mu^m + O(...) with mu = 1 - epsilon + O(epsilon^2). This is a genuine derivation: Theorem 2 expresses the randomly compiled instrument in terms of the generalized Pauli fidelities via character orthogonality, and Theorem 8 (proved in the appendix using Gershgorin discs and Brauer's eigenvalue bound) provides the dominant-eigenvalue result needed for the exponential decay. The quantity nu_{0,0} is an output of the model, and the error rate epsilon is defined through Eq. (24) as 1 - nu_{0,0}, so fitting the decay to estimate nu_{0,0} is not an input/output reversal. The paper does rely on prior work [13,15] for the interpretation of the nu_{a,b} as a probability distribution and for the diamond-distance identification, and these references share authors with the present paper. However, these are supporting structural facts, not the target exponential-decay result, and the central derivation is self-contained. The ideal-gate assumption is a stated validity limitation, not a circular step. No specific reduction of a predicted quantity to a fitted input or to a self-citation chain is present.
Assumptions & free parameters
assumptions (8)
- domain assumption Randomizing Pauli gates are ideal or have gate-independent errors foldable into the instrument.
- domain assumption All qudits are measured, with no idling qudits.
- domain assumption The instrument is trace-preserving, so ν0,0 = 1 − ε is the error probability.
- domain assumption The error rate satisfies ε < 1/3.
- ad hoc to paper The generalized Pauli fidelities are real and positive enough for the log and quadratic-root post-processing to be defined.
- standard math Schur orthogonality of Weyl characters.
- standard math Gershgorin disc theorem and Brauer's eigenvalue bound.
- standard math Jordan decomposition of matrices.
Cite this review
Pith. "Pith review of Benchmarking Quantum Instruments." pith.science (2026). https://pith.science/paper/RKIONZAP
@misc{pith2026250200179,
author = {Pith},
title = {Pith review of: Benchmarking Quantum Instruments},
year = {2026},
howpublished = {\url{https://pith.science/paper/RKIONZAP}},
note = {Machine review of arXiv:2502.00179}
}
read the original abstract
Quantum measurements with feed-forward are crucial components of fault-tolerant quantum computers. We show how the error rate of such a measurement can be directly estimated by fitting the probability that successive randomly compiled measurements all return the ideal outcome. Unlike conventional randomized benchmarking experiments and alternative measurement characterization protocols, all the data can be obtained using a single sufficiently large number of successive measurements. We also prove that generalized Pauli fidelities are invariant under randomized compiling and can be combined with the error rate to characterize the underlying errors up to a gauge transformation that introduces an ambiguity between errors happening before or after measurements.
Figures
Forward citations
Cited by 1 Pith paper
-
Readout sweet spots for spin qubits with strong spin-orbit interaction
Readout back-action in spin qubits from g-tensor modulation is minimized when the magnetic field is oriented so the static Zeeman field is parallel to the sensor-induced Zeeman fluctuation (gB parallel to g'B), a cond...
Reference graph
Works this paper leans on
- [18]
-
[1]
E. B. Davies and J. T. Lewis, An operational approach to quantum probability, Communications in Mathematical Physics 17, 239 (1970)
work page 1970
- [2]
-
[3]
I. L. Chuang and M. A. Nielsen, Prescription for ex- perimental determination of the dynamics of a quantum black box, Journal of Modern Optics 44, 2455 (1997)
1997
-
[4]
S. T. Merkel, J. M. Gambetta, J. A. Smolin, S. Poletto, A. D. C´ orcoles, B. R. Johnson, C. A. Ryan, and M. Stef- fen, Self-consistent quantum process tomography, Phys. Rev. A 87, 062119 (2013)
2013
-
[5]
R. Stricker, D. Vodola, A. Erhard, L. Postler, M. Meth, M. Ringbauer, P. Schindler, R. Blatt, M. M¨ uller, and T. Monz, Characterizing quantum instruments: From nondemolition measurements to quantum error correc- tion, PRX Quantum 3, 030318 (2022)
work page 2022
-
[6]
K. Rudinger, G. J. Ribeill, L. C. Govia, M. Ware, E. Nielsen, K. Young, T. A. Ohki, R. Blume-Kohout, and T. Proctor, Characterizing midcircuit measurements on a superconducting qubit using gate set tomography, Phys. Rev. Appl. 17, 014014 (2022)
work page 2022
-
[7]
L. Pereira, J. J. Garc ´ ıa-Ripoll, and T. Ramos, Com- plete physical characterization of quantum nondemoli- tion measurements via tomography, Phys. Rev. Lett. 129, 010402 (2022)
work page 2022
Show all 22 references
-
[8]
Pereira, J
L. Pereira, J. J. Garc ´ ıa-Ripoll, and T. Ramos, Parallel tomography of quantum non-demolition measurements in multi-qubit devices, npj Quantum Information 9, 22 (2023)
2023
-
[9]
Magesan, J
E. Magesan, J. M. Gambetta, and J. Emerson, Scalable and robust randomized benchmarking of quantum pro- cesses, Physical review letters 106, 180504 (2011)
2011
-
[10]
Dankert, R
C. Dankert, R. Cleve, J. Emerson, and E. Livine, Exact and approximate unitary 2-designs and their application to fidelity estimation, Phys. Rev. A 80, 012304 (2009). 7
2009
-
[11]
Erhard, J
A. Erhard, J. J. Wallman, L. Postler, M. Meth, R. Stricker, E. A. Martinez, P. Schindler, T. Monz, J. Emerson, and R. Blatt, Characterizing large-scale quantum computers via cycle benchmarking, Nature communications 10, 5347 (2019)
2019
-
[12]
J. J. Wallman and J. Emerson, Noise tailoring for scalable quantum computation via randomized compiling, Phys. Rev. A 94, 052325 (2016)
2016
-
[13]
McLaren, M
D. McLaren, M. A. Graydon, and J. J. Wallman, Stochastic errors in quantum instruments, arXiv preprint arXiv:2306.07418 (2023)
2023 arXiv
-
[14]
S. T. Flammia and J. J. Wallman, Efficient estimation of pauli channels, ACM Transactions on Quantum Com- puting 1, 10.1145/3408039 (2020)
2020 doi
-
[15]
S. J. Beale and J. J. Wallman, Randomized compiling for subsystem measurements (2023), arXiv:2304.06599 [quant-ph]
2023 arXiv
-
[16]
Zhang, S
Z. Zhang, S. Chen, Y. Liu, and L. Jiang, A generalized cycle benchmarking algorithm for characterizing mid- circuit measurements (2024), arXiv:2406.02669 [quant- ph]
2024 arXiv
-
[17]
Hines and T
J. Hines and T. Proctor, Pauli noise learning for mid- circuit measurements (2024), arXiv:2406.09299 [quant- ph]
2024 arXiv
-
[19]
J. Lin, J. J. Wallman, I. Hincks, and R. Laflamme, In- dependent state and measurement characterization for quantum computers, Phys. Rev. Res. 3, 033285 (2021)
2021
-
[20]
S. A. Gershgorin, ¨Uber die abgrenzung der eigenwerte einer matrix, Bulletin of the Academy of Sciences of the Union of Soviet Socialist Republics. VII series. Mathe- matics and natural sciences class 6, 749 (1931)
1931
-
[21]
R. A. Horn and C. R. Johnson, Matrix analysis (Cam- bridge university press, 2012)
2012
-
[22]
Brauer, Limits for the characteristic roots of a matrix
A. Brauer, Limits for the characteristic roots of a matrix. II., Duke Mathematical Journal 14, 21 (1947). Appendix A: Eigenvalue bound for nonnegative matrices We prove an eigenvalue bound for nonnegative matri- ces that may be of independent interest. To prove these bounds, w...
1947
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.