{"id":"035ed8eb-0dc9-45ed-a4f9-b364f134c823","arxiv_id":"2607.02888","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Expectation-value accuracy is not decision-complete for shift-invariant QEM workflows; residual gap geometry (the decision kernel) governs argmin/ranking risk and can invert MSE rankings.","lead":"Quantum error mitigation is usually judged by how accurately it estimates expectation values, but many near-term algorithms only use those values to pick a best option or ranking. This paper shows that accuracy gains can leave those choices unchanged or worse, and that the right object to compare methods is the residual gap law induced by device noise through the mitigation map.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Operational risk and ordering claims rest on a Gaussian/LAN gap surrogate whose non-asymptotic fidelity to multinomial shot records remains open.","rationale":"The reader correctly isolates the Gaussian/fixed-allocation surrogate as the weakest load-bearing modeling choice for the operational half of the strongest claim while recognizing that the structural theorems stand independently. The paper is unusually careful: it pre-registers endpoints, reports that practical and dynamic targets are missed, and withholds device-level claims when the hardware micro-cell fails its stricter criterion. No stronger internal inconsistency or hidden assumption was found; the pullback geometry and constructible accuracy–decision separations (CDR flat, PEC coexistence) are supported inside the declared regimes. Therefore the CONDITIONAL verdict—accept theory and constructible separation, withhold aggregate practical benefit and device claims—remains the right one; the concrete multinomial-versus-Gaussian rate check would only tighten or loosen the already-stated caveat.","tokens_in":43231,"tokens_out":626,"duration_ms":25916,"concrete_test":"On the same QAOA-MaxCut instance pool, recompute empirical decision-failure rates under pure multinomial sampling at budgets B=2^12…2^16 (no Gaussian plug-in) and extract the observed large-shot rates −(1/B)log Rm; if these rates deviate from the Gaussian Im of Eq. (106) by more than the o(1) term permitted by the paper’s own CLT/Berry–Esseen discussion (Prop 7.1), or if method orderings reverse relative to GRI, the surrogate-based operational claims weaken.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The structural core (quotient factorization Thm 3.3, gap-law minimality Thm 4.2/4.5, marginal no-go Thm 5.2, pullback Thm 6.1) is decision-complete without Gaussianity. The operational content of the strongest claim—Gaussian risk Rm(B)=1-Φd(√B μm; Σm) (Prop 7.2), diagonal exponent Im=min μi^{2}/(2Σii) (Thm 7.4/7.5), and the fixed-allocation two-point converse (Thm 9.1/Cor 9.2)—treats residual gaps as a local Gaussian/LAN experiment with fixed non-adaptive allocations and (for the converse) equal-covariance alternatives. Section 7 and Limitations explicitly mark a full multinomial-to-Gaussian large-deviation bridge and a measurement-independent quantum-Fisher decision converse as open. Consequently, finite-shot method orderings and the claim that accuracy-oriented QEM can leave decisions unchanged or worse are rigorously established only inside the surrogate plus the declared Aer noise models; they are not yet controlled for the primitive shot measure at the rare-event scale that sets large-B ranking.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that expectation-value accuracy (MSE) and reliability of shift-invariant downstream decisions (argmin, ranking, top-k, etc.) are controlled by different objects. For such decisions the minimal decision-complete object is the residual gap law L(LE_m); in Gaussian finite-shot regimes this reduces to the effective margin–kernel pair (μ_m, Σ_m)=(Δ+La_m, L K_m L^⊤), where Σ_m is the pullback of shared physical device noise through the mitigation map. The authors prove quotient factorization, gap-law sufficiency (including a Blackwell form), a marginal no-go for MSE/pointwise benchmarks, a QEM pullback restriction, Gaussian risk and large-deviation formulas, and a fixed-allocation shot-level converse. Finite-shot Qiskit Aer QAOA-MaxCut simulations show constructible accuracy–decision separation (positive-affine CDR is decision-flat while improving MSE; PEC can improve MSE while worsening decision risk via overhead). Decision-aware selection modestly reduces static held-out failure relative to MSE selection, often by retaining Raw, but does not meet the pre-set practical threshold or the dynamic success target. Pre-registered device-noise and hardware micro-cell stress tests are reported with carefully limited claims.","tokens_in":43565,"tokens_out":1530,"duration_ms":18540,"significance":"If the structural results hold—as they appear to under standard measure-theoretic and Gaussian LDP tools—the paper cleanly separates ambient accuracy benchmarks from decision reliability for a large class of near-term quantum workflows. The QEM-specific content is the restricted pullback geometry of the decision kernel, which is absent from generic ranking-and-selection models and explains why MSE rankings and decision rankings can invert. Strengths include: explicit theorems with proofs or deferred proofs; a linear-pullback realizable no-go witness; pre-registered endpoints and falsification criteria; public hash-locked reproduction package; and honest reporting of non-passes (static reduction 0.1146 < 0.15; dynamic target not reached; hardware decision-direction criterion not met). The operational message—compare methods in residual gap geometry, and sometimes keep Raw—is actionable and proportionate to the evaluated regimes. The main limitation is that finite-shot risk formulas and method orderings are controlled inside a Gaussian/LAN surrogate plus declared Aer noise models, with a full multinomial-to-Gaussian rare-event bridge left open.","major_comments":[{"comment":"Section 7 (Prop. 7.1–7.2, Thm. 7.4) and Limitations: the operational risk Rm(B)=1−Φ_d(√B μ_m; Σ_m), the diagonal exponent I_m=min_i μ_i²/(2Σ_ii), and the method-ordering claims used in §12 rest on a local Gaussian/LAN gap surrogate with fixed non-adaptive allocation. The manuscript correctly marks a full multinomial-to-Gaussian large-deviation bridge as open. For the operational implication in the abstract and conclusion (“select QEM methods through residual gap geometry”), please fence more sharply which claims are theorem-level (gap-law completeness, pullback identity, marginal no-go) versus surrogate-level (finite-shot risk ranking and Algorithm 1), and state explicitly that large-B ordering is controlled only inside the surrogate plus the declared Aer models unless the rare-event bridge is supplied.","section":null},{"comment":"Section 12, Table 5 and the static/dynamic endpoints: P4 fails the pre-set practical threshold (relative reduction 0.1146 < 0.15) and P5 fails the aggregate success target 0.80; the decision-aware selector often retains Raw. The abstract and §14 still recommend selecting through residual gap geometry “in the evaluated regimes.” That recommendation is defensible as a diagnostic practice, but the manuscript should avoid any residual implication of an established aggregate decision benefit. Please align the abstract’s operational sentence and the conclusion with Table 5’s not-pass rows so that the positive claim is limited to constructible accuracy–decision separation, pullback ordering, and modest/sub-threshold static improvement under ranking-preserving noise.","section":null},{"comment":"Theorem 9.1 / Corollary 9.2: the shot-level converse is fixed-allocation, classical-information, and (in the Gaussian corollary) equal-covariance local-minimax. That is appropriate and carefully scoped, but §1 and the contribution list present it alongside the structural core as one of the “seven main results.” Please ensure the introduction and result hierarchy do not over-weight this converse relative to the decision-representation theorems (3.3, 4.2, 4.5, 5.2, 6.1), and restate in the main text (not only Limitations) that it is not a measurement-independent quantum-Fisher or adaptive fixed-confidence bound.","section":null}],"minor_comments":[{"comment":"Figure 1 and the pipeline notation (E_m → LE_m → (μ_m,Σ_m) → R_m(D)) are clear; consider adding a one-line caption note that MSE stops at the ambient residual field so readers scanning figures alone see the mismatch.","section":null},{"comment":"Notation for per-unit vs finite-budget kernels (K_m vs K_m^{(B)}, Σ_m vs Σ_m^{(B)}) is careful but dense; a short notation table early in §3 would reduce cognitive load.","section":null},{"comment":"Related work: Demarty et al. [10] is well positioned; a sentence on how virtual distillation [7] and symmetry verification [45] fit (or do not fit) the pullback picture would help completeness without changing the claim.","section":null},{"comment":"Algorithm 1 uses plug-in Gaussian risk; cross-reference Prop. 7.2 and the critical-band selection warning (selection bias) in the algorithm caption so implementers do not estimate C and Σ on the same data.","section":null},{"comment":"Typos/style: “Vicenzo” vs common “Vincenzo” is author choice; ensure consistent hyphenation of “decision-aware” and “finite-shot” throughout; check “Cramér–Wold” accent consistency.","section":null},{"comment":"Appendix D hardware micro-cell: the re-selection amendment is properly sealed; a one-sentence pointer in the main §12 hardware paragraph to the held-out generalization error (0.095) would help readers who skip the appendix.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is unusually careful about pre-registration, non-passes, and not claiming hardware decision advantage. Fit for a quant-ph theory-plus-simulation venue is good. The Gaussian rare-event bridge is the main scientific loose end; I would not block on it if the operational claims are fenced as requested. No novelty or citation-pattern concerns beyond the usual independent-researcher literature breadth."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple and right: for shift-invariant QEM workflows, MSE is the wrong object. The minimal decision-complete thing is the residual gap law, and in second-order regimes the pullback kernel Σ_m = L K_m L^⊤ induced by device noise through the mitigation map. That joint identification is what is new; “gaps matter” is classical, but confining the kernel to the physical pullback family and proving the marginal no-go for MSE/CI selection is not.\n\nThe math is clean. Quotient factorization, gap-law sufficiency/Blackwell form, the no-go, and the pullback identities are standard tools applied carefully; the fixed-allocation shot-level converse is a local two-point barrier, not a universal quantum-Fisher claim, and the paper says so. Worked ZNE/CDR/PEC signatures make the geometry operational. The simulations are unusually honest: CDR is decision-flat while crushing MSE, PEC can improve accuracy while worsening risk via overhead, static selection beats MSE selection but misses the pre-set 0.15 practical threshold, dynamic success never hits 0.80, and the hardware micro-cell supports covariance match but fails the stricter decision-direction criterion so they make no device claim. Pre-registration and a public reproduction package help.\n\nSoft spots are real but proportional. Operational risk formulas and large-shot exponents live inside a Gaussian/LAN surrogate with fixed non-adaptive allocation; the multinomial-to-Gaussian rare-event bridge and measurement-independent converse are marked open. So finite-shot method orderings are rigorous inside the surrogate and the declared Aer models, not yet controlled at the primitive shot measure for large-B ranking. That is a scope limit, not a hole in the structural core. Free parameters (0.15, τ_R, 0.80) are declared thresholds, not fitted conclusions.\n\nThis is for people who actually select QEM for VQE/QAOA ranking, top-k, or step acceptance. Theory readers get a clean decision object; practitioners get a reason to keep Raw when mitigation spends shots on the wrong directions. I would send it to peer review. Engage with the theory and the constructible separation; treat aggregate practical benefit and device-level claims as not established.","headline":"Solid structural paper: residual gap law + physical pullback kernel is the real contribution; Gaussian operational claims are scoped honestly and the experiments refuse to overclaim.","tokens_in":44222,"tokens_out":549,"would_cite":true,"duration_ms":6264,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Accuracy gains from quantum error mitigation need not improve the downstream decisions those estimates are used to make.","keywords":["quantum error mitigation","decision kernel","residual gap law","finite-shot estimation","quotient space","Clifford data regression","probabilistic error cancellation","QAOA"],"falsifier":"Find a finite-shot QEM regime in which a method that strictly improves residual gap margins and decision-kernel geometry still loses on the downstream decision to a method that only improves ambient mean-squared error, under the same total shot budget and the same shift-invariant decision rule.","tokens_in":44074,"feed_emoji":"⚛️","tokens_out":975,"duration_ms":7500,"temperature":0.7,"pith_summary":"Quantum error mitigation is almost always scored by how close estimated expectation values are to the truth. Many near-term workflows, however, only use those values to choose something: the best parameter, a top-k set, an optimizer step, or a phase label. Those choices depend on gaps between values, not on absolute levels, so they live in a quotient space that ignores global shifts. The paper builds a finite-shot theory around that fact. The minimal object that fully determines every shift-invariant decision is the residual gap law; under Gaussian shot noise it is summarized by effective margins and a decision kernel that is the pullback of shared device noise through the mitigation map. Because that kernel is physically constrained, a method can cut mean-squared error while leaving the decision unchanged or worse. Simulations of Clifford-data regression and probabilistic error cancellation show the split, and decision-aware selection often keeps the raw estimator. The practical reading is to compare mitigation methods by residual gap geometry, not by accuracy alone.","feed_headline":"Better estimates need not make better quantum decisions","feed_subtitle":"Gap geometry, not mean-squared error, controls argmin and ranking after error mitigation","key_machinery":"The decision kernel Σ_m = L K_m L^⊤, the gap-space pullback of residual covariance through a contrast map L whose kernel is the all-ones vector. It, together with the effective margins, is the second-order Gaussian object that determines argmin, ranking, top-k, and related finite gap decisions, and it is confined to the physical family generated by device noise and the mitigation map.","core_discovery":"Estimator accuracy and decision reliability are controlled by different objects. For shift-invariant downstream tasks the minimal decision-complete object is the residual gap law of the mitigated landscape; in Gaussian finite-shot regimes that law is summarized by the effective margin–kernel pair induced by the gap map applied to residual bias and covariance. The decision kernel is not free: it is the pullback of shared physical device noise through the mitigation map. Accuracy-oriented mitigation can therefore reduce ambient mean-squared error while leaving decisions flat or worse, so methods should be selected through residual gap geometry rather than expectation-value accuracy alone.","pith_inferences":["Any variational or combinatorial workflow that ends in an argmin, ranking, or top-k filter inherits the same ambient-versus-gap mismatch, not only quantum error mitigation.","Shot budgets and mitigation maps could be co-designed to minimize gap-kernel risk rather than ambient MSE, treating the decision kernel as the design objective.","Readout and device-calibration pipelines that inject common-mode or positive-affine residuals will look accurate under MSE while leaving decisions untouched.","A natural next test is whether the same gap-kernel selector improves decisions under adaptive allocation or fixed-confidence stopping, which the paper leaves open."],"forward_implications":["QEM method rankings should be produced in gap space; accuracy rankings can invert decision rankings.","Positive-affine Clifford-data regression can improve MSE while remaining decision-flat, so retaining the raw estimator can be optimal.","Probabilistic error cancellation can lower MSE yet raise decision risk through sampling overhead.","Critical-band estimation of residual margins and the decision kernel is enough for operational method selection.","Under ranking-preserving noise the decision-aware choice is often to decline mitigation."],"fun_headline_variants":["Accuracy gains need not improve quantum downstream decisions","Gap geometry not MSE controls QEM argmin and ranking","Mitigation can cut error while leaving decisions flat or worse","Residual gap kernels dictate when QEM helps real choices","Select error mitigation by decision margins not accuracy alone"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The operational risk formulas and the shot-level converse rest on a local Gaussian approximation to residual gap noise under fixed non-adaptive shot allocation; a full non-asymptotic bridge from multinomial counts and a measurement-independent quantum bound are left open.","fun_headline_variants_meta":{"raw":{"variants":["Accuracy gains need not improve quantum downstream decisions","Gap geometry not MSE controls QEM argmin and ranking","Mitigation can cut error while leaving decisions flat or worse","Residual gap kernels dictate when QEM helps real choices","Select error mitigation by decision margins not accuracy alone"]},"model":"grok-4.5","effort":"low","cost_usd":0.00511,"raw_usage":{"total_tokens":1477,"prompt_tokens":890,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":51100000,"prompt_tokens_details":{"text_tokens":890,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":510,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":890,"tokens_out":77,"duration_ms":4927,"temperature":1.0,"reasoning_tokens":510,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T06:21:28.041355+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a finite-shot QEM regime in which a method that strictly improves residual gap margins and decision-kernel geometry still loses on the downstream decision to a method that only improves ambient mean-squared error, under the same total shot budget and the same shift-invariant decision rule.","supporting_citations":[],"review_version":1}