REVIEW 2 major objections 5 minor 1 cited by
Zero-noise extrapolation can silently collapse into a fixed rescaling of a single noisy measurement, producing apparent improvements that are artefacts once noise destroys the amplified signal.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
2026-08-04 04:22 UTC pith:2E6QRUS5
load-bearing objection Eq. (2) is a genuine and useful catch; the 'widespread threat' claim is partly under-evidenced but the paper deserves refereeing. the 2 major comments →
Benchmarking Error Mitigation: Artefactual Improvements in Zero-Noise Extrapolation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that once the expectation values E(λ) at all amplified noise levels λ>1 reach the observable floor f, the Richardson estimate collapses to Ê(0)=c₁E(λ₁)+(1−c₁)f: a fixed affine rescaling of the single noisy value E(λ₁), independent of how the noise was amplified. Because c₁>1 for the standard scale set {1,3,5}, this rescaling always points upward, manufacturing an apparent improvement that can pass through the true ideal value or overshoot it. On the IQM Euro-Q-Exa hardware, a 4-qubit Trotter circuit folded at depths where the amplified parity values hit the floor produces estimates that track this rescaling and overshoot the ideal by 21%. The authors call this a 'horosc
What carries the argument
Richardson extrapolation: a linear combination Ê(0)=Σ c_k E(λ_k) with Σ c_k=1, where the coefficients c_k depend only on the scale factors; for {1,3,5}, c₁=15/8. The collapse identity (Eq. 2) is the key object: when E(λ>1)=f, the estimate becomes one-point rescaling. The matched-cost garbage-folding control (folds that do not reduce to identity but have the same gate count and error rate) and the negative-probability weight W_neg (sum of negative per-basis-state extrapolated probabilities) are the practical diagnostics. They carry the argument by distinguishing genuine improvement from arithmetic artefact.
Load-bearing premise
The conclusion that this artefact is a widespread threat to current benchmarks rests on the premise that the signal-destroyed regime—where amplified expectation values hit the floor—is reached quickly on real hardware for non-trivial circuits, a premise supported by one device and ideal-depolarising-noise simulations.
What would settle it
Run a non-trivial circuit (e.g., a 4-qubit Trotter circuit at depth d≥5) on a different hardware device or qubit chain, measure E(λ) at scale factors {1,3,5}, and check whether E(λ3) and E(λ5) stay clearly above the floor while the Richardson estimate does not match c₁E(λ1)+(1−c₁)f. If the amplified signal is retained and Richardson recovers the ideal value, the collapse has not occurred; finding a case where E(λ>1)>floor yet the estimate still equals the rescaling would falsify the sufficiency of the floor condition.
If this is right
- If collapse is reached, any linear extrapolator with c₁>1 yields the same fixed rescaling; the reported improvement is invariant to the noise-amplification method and can be manufactured by signal-destroying folds.
- Garbage-folding reports a larger apparent improvement than genuine folding in the destroyed regime, so the magnitude of a claimed improvement is not evidence of its correctness.
- The zero-cost negative-probability weight W_neg separates valid from artefactual runs with no overlap across 405 hardware runs (AUC=1.0), and requires only per-λ counts a benchmark already holds.
- Benchmarks that report only the scalar estimate Ê(0) and its distance from the raw value can misreport artefacts; the proposed three-item checklist (signal retention, W_neg, overshoot flag) can be adopted at no additional computational cost.
Where Pith is reading between the lines
- The collapse identity (Eq. 2) depends only on the linear coefficient sum and the floor, so non-linear extrapolators that saturate at a floor—exponential or multi-exponential fits—may exhibit an analogous artefact; the paper does not analyse those.
- A matched-cost negative control like garbage folding could be adopted as a routine sanity check across QEM benchmarking suites, not only ZNE, since it tests whether a claimed gain is specific to the amplification mechanism.
- Because the onset of the artefact depends on device noise and circuit depth, the per-λ decay curves used for the diagnostic could be summarised into a single 'signal-retention depth' figure of merit for device benchmarking.
- If the empirical premise generalises beyond the one hardware device studied, previously published ZNE improvements may need re-examination; applying the zero-cost W_neg check to existing per-λ data would settle the question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper identifies and characterises a failure mode in Richardson zero-noise extrapolation (ZNE). The central result is Eq. (2): if, for all amplified noise levels k>1, the measured expectation values collapse to the observable floor f, then the Richardson estimate reduces to a deterministic rescaling \hat{E}(0)=c_1 E(λ1)+(1-c_1)f; for {1,3,5} this is (15/8)E(λ1)-(7/8)f. The authors demonstrate this regime on IQM Euro-Q-Exa with a 4-qubit quantum Trotter circuit using genuine {1,3,5} folding, reporting an overshoot of up to 21% at depth d=3. They further introduce a signal-retention taxonomy, a matched-cost garbage-folding negative control, a zero-cost negative-probability diagnostic W_neg, and a reporting checklist for ZNE benchmarks.
Significance. The mathematical content is elementary but the paper's contribution is the identification and empirical demonstration of an underappreciated benchmarking artefact. Eq. (2) is derived cleanly from the Lagrange-coefficient property and the stated floor condition, and the hardware experiment appears carefully controlled: matched gate counts, an identity-fold control, blocked/interleaved acquisition, and 15 repetitions with reported error bars. The proposed W_neg diagnostic is useful because it does not require knowledge of the ideal value. The availability of code and data in a reproduction package is a notable strength. If the artefact is common in practice, the paper has real value for the QEM benchmarking community.
major comments (2)
- [Abstract and Section VII] The abstract states that the signal-destroyed regime is 'quickly reached on current hardware for non-trivial circuits' and that the artefact poses a 'rarely considered threat to the validity of many empirical evaluations in quantum computing.' This generality is not supported by the evidence: Section VII correctly concedes that 'hardware evidence is limited to one device and circuit family,' and the simulations assume ideal depolarising noise. Figure 2 also shows circuits (QFT mirror, QTC at d=1) in which the signal is retained and the artefact does not occur. Because the paper's stated significance rests on this empirical premise, the authors should either (a) provide additional hardware evidence across devices and circuit families, or (b) explicitly reframe the claims as showing that the artefact 'can occur' and is a risk to be checked, rather than a demonstrated widespread failure. Th
- [Section III-A and Figure 1] The empirical confirmation of the collapse regime is described informally: the text says the amplified values 'have reached the floor' and that \hat{E}(0) tracks the rescaling 'to within 5%,' but no quantitative collapse criterion or uncertainty intervals are given. The 15 repetitions are mentioned later in Section V-D, yet Figure 1 does not show error bars. Please specify a precise operational threshold for 'reached the floor' (for example, |E(λ_k)-f| relative to the statistical error), report the per-λ values and uncertainties for d=3 and d=5, and add error bars to Figure 1. This would make the hardware demonstration more robust and would prevent a reader from mistaking a qualitative visual trend for a quantitative confirmation.
minor comments (5)
- [Section IV / Figure 2] The figure legend and axis label '0(extrap.)' are slightly confusing: the extrapolated value at λ=0 should be labelled separately from the scale-factor axis. Please clarify the plotting convention in the caption.
- [Section V-B] The garbage-folding results for QFT and QTC are given only in the text as ρ=16 and ρ=125. These are striking values; consider showing them in a table or in Figure 2 to make them reproducible and easy to locate.
- [Section II-a] The statement that the tightly spaced set yields a variance bound '194× larger' than {1,3,5} is correct only because 681/3.5 ≈ 194. Please state this arithmetic explicitly to avoid apparent arbitrariness.
- [Section V-D] The phrase 'matched in gate count and residual error rate' for garbage folding is not fully defined. Please specify how 'residual error rate' is estimated or matched; if it is inferred from gate counts alone, say so.
- [Section VI] The negative-probability diagnostic W_neg is reported with AUC=1.0 on the hardware data. Since the regimes were defined using the same data, please state whether this separation is in-sample or whether a cross-validated or out-of-sample assessment was performed.
Circularity Check
No significant circularity; Eq. (2) is a derived consequence and hardware confirmation uses fresh measurements.
full rationale
The central derivation is self-contained. Eq. (2) follows algebraically from Eq. (1) and the stated floor hypothesis E(lambda_k) ≈ f for k>1, using only the Lagrange-coefficient sum property Σc_k = 1; no fitted parameter is renamed as a prediction. The Euro-Q-Exa confirmation is a fresh measurement of a 4-qubit QTC whose amplified parity values happen to collapse to the floor, not a fit of Eq. (2) to data. The garbage-folding negative control is deliberately constructed to maximize signal destruction, the precondition of Eq. (2), but the resulting overshoot and W_neg values are computed from the measurement statistics, not imposed by the construction. The self-citation [11] appears only as contextual positioning ('extending our broader study...') and is not load-bearing for any theorem or numerical result. The paper's own Section VII caveat that 'hardware evidence is limited to one device and circuit family' is a generality limitation, not a circular step. Thus no step reduces, by the paper's own equations or by self-citation, to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (1)
- p_2q (two-qubit depolarising error rate in simulation) =
1e-3 (Section IV); 6.4e-4 (Section VI)
axioms (4)
- standard math Lagrange interpolation coefficients satisfy Σ c_k = 1
- domain assumption Ideal depolarising noise model for simulation experiments
- domain assumption The observable floor f is known and stationary (f = 1/2^N for probabilities, f = 0 for parity)
- domain assumption The hardware calibration and qubit chain selection (qubits 8-11) remain stable over the experiment
read the original abstract
Reliable benchmarking of Quantum Error Mitigation (QEM) requires distinguishing genuine improvements from artefacts of the post-processing arithmetic. In this paper, we expose a failure mode in Richardson Zero-Noise Extrapolation (ZNE), a widely used technique routinely (and often implicitly) relied upon in benchmarks and experiments. When noise amplification operates beyond usable signals - a regime that is quickly reached on current hardware for non-trivial circuits - we show that the extrapolation no longer reflects the underlying physics, but collapses into a fixed rescaling of a single noisy measurement, producing a bogus apparent improvement that is independent of noise amplification. This poses a rarely considered threat to the validity of many empirical evaluations in quantum computing. Measurements on real hardware (IQM Euro-Q-Exa) confirm this collapse with ordinary folding alone: as circuit depth erodes the signal, the reported estimate decouples from the truth and overshoots the ideal by up to 21%. We further introduce a matched-cost "garbage-folding" negative control that carries no usable signal yet reports a larger apparent improvement than genuine folding - showing that the magnitude of an improvement is not evidence of its correctness - alongside a zero-cost check flagging the artefact from data a benchmark already holds. We distil both into a short reporting checklist for ZNE benchmarks.
Figures
Forward citations
Cited by 1 Pith paper
-
Classically Augmented Zero-Noise Extrapolation
Classically Augmented Zero-Noise Extrapolation replaces high-noise Richardson nodes with classically simulated estimates, yielding exponential sampling-variance reduction for linear node spacings at fixed cutoff.
Reference graph
Works this paper leans on
-
[1]
Greiwe, T
F. Greiwe, T. Krüger, and W. Mauerer, Effects of imperfections on quantum algorithms: A software engineering perspective,QSW, 2023
2023
-
[2]
Thelen, H
S. Thelen, H. Safi, and W. Mauerer, Approximating under the influence of quantum noise and compute power,QCE, 2024
2024
-
[3]
Carbonelli et al., Challenges for quantum software engineering: An industrial use case perspective,Quantum Software: Aspects of Theory and System Design, Springer-Nature, 2024
C. Carbonelli et al., Challenges for quantum software engineering: An industrial use case perspective,Quantum Software: Aspects of Theory and System Design, Springer-Nature, 2024
2024
-
[4]
Temme, S
K. Temme, S. Bravyi, and J. M. Gambetta, Error mitigation for short- depth quantum circuits,PRL, 2017
2017
-
[5]
Li and S
Y . Li and S. C. Benjamin, Efficient variational quantum simulator incorporating active error minimization,PRX, 2017
2017
-
[6]
S. Endo, S. C. Benjamin, and Y . Li, Practical quantum error mitigation for near-future applications,PRX, 2018
2018
-
[7]
Cai et al., Quantum error mitigation,Rev
Z. Cai et al., Quantum error mitigation,Rev. Mod. Phys., 4 2023
2023
-
[8]
Giurgica-Tiron et al., Digital zero noise extrapolation for quantum error mitigation,QCE, 2020
T. Giurgica-Tiron et al., Digital zero noise extrapolation for quantum error mitigation,QCE, 2020
2020
-
[9]
Krebsbach, B
M. Krebsbach, B. Trauzettel, and A. Calzona, Optimization of richard- son extrapolation for quantum error mitigation,Phys. Rev. A, 6 2022
2022
-
[10]
Govia et al., Bounding the systematic error in quantum error mitigation due to model violation,PRX Quantum, 1 2025
L. Govia et al., Bounding the systematic error in quantum error mitigation due to model violation,PRX Quantum, 1 2025
2025
-
[11]
Köster and W
D. Köster and W. Mauerer, Claim against measurement: Statistical artefacts in quantum error mitigation benchmarks, 2026
2026
-
[12]
Russo et al., Testing Platform-Independent Quantum Error Mitiga- tion on Noisy Quantum Computers,TQE, 2023
V . Russo et al., Testing Platform-Independent Quantum Error Mitiga- tion on Noisy Quantum Computers,TQE, 2023
2023
-
[13]
Mauerer and S
W. Mauerer and S. Scherzinger, 1-2-3 reproducibility for quantum software experiments,SANER, 2022
2022
-
[14]
Majumdar et al., Best practices for quantum error mitigation with digital zero-noise extrapolation,QCE, 2023
R. Majumdar et al., Best practices for quantum error mitigation with digital zero-noise extrapolation,QCE, 2023
2023
-
[15]
Kandala et al., Error mitigation extends the computational reach of a noisy quantum processor,Nature, 2019
A. Kandala et al., Error mitigation extends the computational reach of a noisy quantum processor,Nature, 2019
2019
-
[16]
S. R. Maschek et al., Make some noise! measuring noise model quality in real-world quantum software,QSW, 2025
2025
-
[17]
Mohammadipour and X
P. Mohammadipour and X. Li, Direct analysis of zero-noise extrapo- lation: Polynomial methods, error bounds, and simultaneous physical- algorithmic error mitigation,Quantum, 2025
2025
-
[18]
V . S. Alfaro, The finite-shot help-harm boundary of zero-noise ex- trapolation, 2026
2026
-
[19]
Miranskyy et al., Improving zero-noise extrapolation via physically bounded models, 2026
A. Miranskyy et al., Improving zero-noise extrapolation via physically bounded models, 2026
2026
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.