{"id":"458ecad0-693a-4f4f-a941-181910eb120a","arxiv_id":"2412.16690","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A random-basis, minimum-of-means stabilizer certification protocol accepts good states and rejects bad states with error probabilities exponentially small in qubit number under a wide fidelity gap.","lead":"The paper proposes a protocol for checking whether a quantum computer prepared a desired stabilizer state, using n measurement settings with many repeated shots. It matters because it offers a mathematically analyzed certification method for NISQ hardware, though only under a narrow set of practical conditions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"False-positive guarantee rests on Propositions 13/16 in a separate full paper; the arXiv artifact does not contain their proof.","rationale":"The reader's conditional verdict identifies the same load-bearing gap: the false-positive guarantee is asserted in the extended abstract but not proven. The deterministic no-false-negative part is solid: any state with fidelity ≥1−δ satisfies tr(sρ)≥1−2δ for every non-identity stabilizer, so ω=2 removes intrinsic false negatives. The unresolved point is the combinatorial bound on the measure of bases whose elements all have expectation above 1−αε, for an arbitrary bad state. My own quick derivation, using the sum identity over stabilizer expectations and the projection bound P_−≤I−|ψ⟩⟨ψ|, yields |G|≤2^n/(2−α) and hence an intrinsic false-positive probability at most (1/(2−α))^n / c, which is even stronger than the claimed O(((α+1)/2)^n). This suggests the concern is likely resolvable, but the submitted artifact does not contain this proof, so a conditional verdict is appropriate until the full paper or an independent check confirms it. The secondary concern about wall-clock speedup is explicitly acknowledged by the author in §1.3 and does not affect the correctness theorem. Therefore I recommend no change to the reader's verdict.","tokens_in":12113,"tokens_out":44194,"duration_ms":372174,"concrete_test":"Fetch the full paper at dojt.srht.site/storage/2412.16690 and independently verify Propositions 13 and 16. Alternatively, for small n (e.g., n=4,5,6) and α=1/2, perform an exhaustive numerical search over discretized syndrome distributions p on the 2^n eigenstates to maximize Pr_basis[∀i, Σ_{e: b_i·e=1} p_e < αε/2], and check that the maximum never exceeds ((1+α)/2)^n; a single violation would invalidate the central false-positive claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that BMoM rejects every bad state with fidelity ≤1−ε with overwhelming probability depends on the intrinsic false-positive bound: for every such state, the probability over a uniformly random stabilizer basis that every basis element has expectation >1−αε is at most η_n(α)=O(((α+1)/2)^n). This is stated in §2.2.b.i as fact (b), with proofs deferred to Propositions 13 and 16 of the full paper, which is not part of the arXiv artifact. The deterministic no-false-negative part (ω=2) is elementary and correct, but the false-positive part is the load-bearing step: if the combinatorial estimate were wrong, BMoM could accept bad states with non-negligible probability. A short independent argument (using the sum identity ∑_{s≠I}(1−tr(sρ))=2^n(1−F) and the bound tr(sρ)≥2F−1) gives |{s: tr(sρ)>1−αε}| ≤ 2^n/(2−α), suggesting the lemma is true, but this proof is not in the submitted text. The artifact is therefore an extended abstract whose main theorem is not self-contained.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes BMoM (Basis-Min-of-Means), a protocol for certifying stabilizer states. The protocol samples a random basis of the stabilizer group, estimates the expectation value of each basis element by m identical shots, and accepts if the minimum of these estimates is at least 1−β. The paper claims deterministic absence of intrinsic false negatives for good states (fidelity ≥ 1−δ) when the gap parameter ω=2, and bounds on the intrinsic false-positive probability for bad states (fidelity ≤ 1−ε): an upper bound O(((α+1)/2)^n) and a lower bound 2^{−Ω(n/α)}, where α is a parameter in Eq. (2) relating δ and ε. The statistical-noise analysis is sketched through Chernoff bounds and a partition of the fidelity gap. The manuscript is explicitly an extended abstract, with the proofs of the main false-positive bounds deferred to Propositions 13 and 16 of an external full paper.","tokens_in":12308,"tokens_out":6920,"duration_ms":57951,"significance":"If the deferred combinatorial bounds are correct, the paper offers a conceptually clean result: randomizing over stabilizer bases and taking the minimum of the empirical means amplifies the fidelity gap, reducing the worst-case factor from O(n) in Somma et al.'s bound to O(1). The deterministic no-false-negative part (ω=2) is elementary and correct, and the author is unusually candid about the method's practical limitations. However, the central false-positive guarantee is not proved in the submitted artifact; the main theorem is therefore conditional on an external proof. The paper also contains no formal statement of the total error probability combining intrinsic and statistical errors. The significance is real but presently unverified from the manuscript alone.","major_comments":[{"comment":"The central false-positive guarantee is stated as fact (b) in §2.2.b.i: for every bad state, the probability over a uniformly random stabilizer basis that all basis elements have expectation above 1−αε is at most η_n(α)=O(((α+1)/2)^n). This bound is load-bearing: it is the only quantitative support for the claim that BMoM rejects bad states with overwhelming probability. The proof is deferred to Propositions 13 and 16 of an external full paper, which is not part of the arXiv artifact, and the submitted text does not even sketch the argument. Because the abstract promises 'mathematically rigorous analysis' of the false-positive rate, the proof (or at least a complete proof sketch) must be included in the submitted manuscript; otherwise the main theorem remains unverifiable.","section":"§2.2.b.i and Introduction (p.2)"},{"comment":"The paper never states a formal theorem that combines the intrinsic-error bounds (facts (a) and (b)) with the statistical-noise analysis into a total error probability for the full protocol. The abstract claims that the protocol 'with overwhelming probability, accepts any good state ... and rejects any bad state,' but the only explicit quantitative statements in the artifact are the intrinsic bounds; the choice of m, β, γ_good, γ_bad to achieve a target total error probability p is described only informally ('one uses Chernoff's theorem' in §2.1). A precise statement of the achieved false-negative and false-positive probabilities as functions of all parameters is needed to substantiate the abstract's claim.","section":"§2.2.b and §2.2.b.i"}],"minor_comments":[{"comment":"The passage 'No clear instances of outright grammatical errors or significant misuses of language appear in this text...' appears to be a leftover note to the author or editor and should be removed; it is not part of a scientific manuscript.","section":"§2.2.c"},{"comment":"The example 'α = 1/4 if n ≥ 47' seems inconsistent with the typical choice 'ε = 8δ' stated in §1.1: Eq. (2) requires 2δ < αε, so for α=1/4 one needs ε > 8δ, leaving no room for γ_good and γ_bad when ε = 8δ. Please clarify the required margins or correct the numeric examples.","section":"§2.2.b and §1.1"},{"comment":"The caption references Propositions 13 and 16 of the full paper, which are not in the submitted artifact; since the figure is meant to display the paper's bounds, the caption should either describe the plotted functions explicitly or the relevant definitions should be included in the artifact.","section":"Figure 2 caption"},{"comment":"The footnote 'A useless observation in view of our stated target of n ≪ 1000...' is informal and distracts from the technical presentation; consider rewriting it as a neutral remark about the asymptotic regime.","section":"Footnote 4 (p.2)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an extended abstract whose main technical results are proved in an external document not included in the arXiv artifact. If the journal does not accept submissions with proofs external to the manuscript, this would alone justify rejection. Even if such a format is permitted, the central false-positive bound needs to be made available in the submitted version. The author's own critical evaluation is candid and helpful, but it also highlights that the practical motivation is not currently supported by the theory. I recommend that the editor require the author to include the missing proofs or clearly mark the unproved statements as conjectures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a genuinely new protocol idea, clearly presented, but the central false-positive proof is not in the submitted artifact — it lives in a separate full paper. Worth refereeing, with the full paper attached.\n\nWhat is actually new: Somma et al. fix a basis and use an affine combination of expectation values; BMoM randomizes the basis and uses the minimum of per-observable means. That change is simple but not in the prior literature, and the paper gives rigorous worst-case bounds for both error directions, including an exponential-in-n intrinsic false-positive probability. The deterministic no-false-negative fact (ω=2) is elementary and correct. The paper also has a refreshingly honest critical evaluation section: it states plainly that BMoM works only for n≪1000 and ε>2δ, that the lower/upper bounds are far apart, that the lower bound lacks a closed form, and that the practical speed-up depends on shots being much cheaper than basis switching. Credit where due: the author identifies the narrow regime of applicability better than most.\n\nThe soft spot: fact (b) in §2.2.b.i — the upper bound on the intrinsic false-positive probability — is asserted, with proofs deferred to Propositions 13 and 16 of a full paper that is not part of the arXiv artifact. This is the load-bearing step. If that bound were wrong, BMoM could accept bad states with non-negligible probability. I agree with the stress-test note that this is the real issue. That said, the note's own short counting argument (using ∑_{s≠I}(1−tr(sρ)) = 2^n(1−F)) suggests the lemma is true; I'm not accusing the author of a false claim, just noting the artifact is not self-contained. No code or simulations accompany the text, and the non-linear fault-catching motivation is explicitly a hypothesis. These are addressable weaknesses, not fatal ones.\n\nWho should read this: anyone working on stabilizer-state verification or NISQ testing. The protocol is simple enough to implement, and the analysis framework (intrinsic vs statistical mistakes) is useful. For peer review: yes, send it out, but require the full paper's proof of Propositions 13 and 16 to be included or at least stated completely. The extended abstract alone doesn't establish the main theorem.","headline":"Worth a referee's time, provided the full paper with Propositions 13 and 16 is attached; the idea is genuinely new and the self-criticism is unusually honest.","tokens_in":12883,"tokens_out":3842,"would_cite":true,"duration_ms":33894,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Randomly sampling a stabilizer basis and taking the minimum of per-observable estimates certifies stabilizer states with n settings and provable error bounds.","keywords":["stabilizer states","quantum state certification","direct fidelity estimation","minimum of means","false-positive rate","false-negative rate","NISQ testing","random stabilizer basis"],"falsifier":"Evaluate the probability numerically for a concrete bad state, e.g., a pure state orthogonal to the target, at n around 20 with α=1/4: if the fraction of random stabilizer bases whose minimum expectation exceeds 1−αε is not below the claimed O(((α+1)/2)^n) bound, the false-positive analysis is wrong.","tokens_in":11840,"feed_emoji":"⚛️","tokens_out":7104,"duration_ms":57974,"temperature":0.7,"pith_summary":"The paper proposes a quantum-state certification protocol, Basis-Min-of-Means (BMoM), that checks whether an experimentally prepared state is close to a known stabilizer target state. Instead of averaging over many random stabilizer measurements, it samples one random basis of the stabilizer group, estimates each basis element's expectation value with many identical shots, and accepts only if the minimum of those estimates is high. The author proves that this protocol needs only n measurement settings rather than O(1/epsilon), and that it comes with rigorous worst-case guarantees: good states (fidelity at least 1−δ) are never intrinsically rejected, and bad states (fidelity at most 1−ε) are rejected except with probability exponentially small in n for fixed gap parameters. The motivation is in-situ testing of near-term quantum computers, where switching measurement bases is expensive while repeating identical shots is cheap.","feed_headline":"Certify stabilizer states with just n measurement settings","feed_subtitle":"Tuned for NISQ testing: identical shots are cheap, so n settings with many shots beats many settings.","key_machinery":"The protocol's load-bearing object is a uniformly random basis of the stabilizer group of the target state, together with a certificate function that takes the minimum of the n estimated expectation values. The key mechanism is amplification through randomization: for a fixed basis the minimum expectation of a bad state can be as low as 1−O(ε/n), but when the basis is drawn uniformly at random, the probability that every basis element's expectation exceeds 1−αε becomes exponentially small in n (for α<1 fixed). This randomization converts a weak fixed-basis bound (inherited from Somma et al.) into a strong worst-case guarantee.","core_discovery":"The central claim is that BMoM certifies a stabilizer state with overwhelming probability: any state with fidelity at least 1−δ is accepted and any state with fidelity at most 1−ε is rejected. The false-negative part is deterministic with ω=2 — for every good state, every element of every basis of the stabilizer group has expectation at least 1−2δ — and the intrinsic false-positive probability is bounded above by O(((α+1)/2)^n) and below by $2^{{−Ω(n/α)}}$, where the parameters are tied by ωδ(1+γ_good) = β = (1−γ_bad)αε. This is achieved with exactly n measurement settings, each repeated m times, in contrast to DFE-C which requires O(1/ε) settings. The paper also claims, as a hypothesis rather than a proven result, that the non-linear minimum-of-means certificate can catch some faults that linear averages miss.","pith_inferences":["If the omitted combinatorial bound holds, the same random-basis-plus-minimum amplification could be applied to other groups of observables, such as generators of graph states or hypergraph states, yielding verification protocols with few settings for those families too.","The paper leaves open replacing the deterministic factor ω=2 by a probabilistic ω<2; if that succeeds, the effective gap requirement shrinks and BMoM becomes viable for ε closer to 2δ, where DFE-C operates.","The fault-detection advantage of the non-linear certificate is untested; a simulation study of specific fault models, e.g., a two-qubit gate that intermittently misfires, could confirm or refute it before any hardware trial.","Should future NISQ hardware make basis switching as cheap as identical shots, the latency-based motivation disappears and BMoM's remaining value would rest on the still-unproven fault-detection hypothesis."],"forward_implications":["A stabilizer state can be certified with exactly n measurement settings, each repeated m times, with worst-case bounds on both accepting bad states and rejecting good states.","When switching measurement bases is the dominant cost, BMoM can certify a state with fewer wall-clock seconds than DFE-C, even though its total shot count is larger.","The protocol requires a wide fidelity gap: the analysis needs ε > 2δ, and suggests ε ≈ 8δ for small qubit counts, to make intrinsic false positives negligible.","Using the minimum instead of the mean gives the certificate a non-linear component, which in principle can distinguish some faults whose effects cancel in a linear average, though the paper presents this as a hypothesis.","The intrinsic false-positive probability decreases exponentially in n/α, so larger systems and wider gaps automatically become safer for fixed α."],"supporting_citations":[{"why":"Supplies the Direct Fidelity Estimation framework that BMoM modifies, including the measurement of random stabilizer observables.","marker":"[6]"},{"why":"Gives the fixed-basis infidelity bound ε ≤ (n/2)(1−μ) that BMoM replaces by a minimum and amplifies by randomization.","marker":"[14]"},{"why":"Defines the quantum state certification task with fidelity thresholds δ and ε that BMoM is designed to solve.","marker":"[8]"}],"fun_headline_variants":["Certify stabilizer states with n bases and many shots","Stabilizer check: n measurement settings, m shots each","Few observables, many shots: rigorous stabilization verification","BMoM: n settings for stabilizer certification","Quantum verification: n bases, many shots per basis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The false-positive guarantee depends on a combinatorial probability estimate — that for every bad state a uniformly random stabilizer basis makes every basis element's expectation exceed 1−αε with probability at most η_n(α) — which is only stated, not proved, in this extended abstract.","fun_headline_variants_meta":{"raw":{"variants":["Certify stabilizer states with n bases and many shots","Stabilizer check: n measurement settings, m shots each","Few observables, many shots: rigorous stabilization verification","BMoM: n settings for stabilizer certification","Quantum verification: n bases, many shots per basis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000331,"raw_usage":{"total_tokens":1783,"prompt_tokens":823,"completion_tokens":960,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":881}},"tokens_in":439,"tokens_out":960,"duration_ms":8723,"temperature":1.0,"reasoning_tokens":881,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:21:55.820045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the probability numerically for a concrete bad state, e.g., a pure state orthogonal to the target, at n around 20 with α=1/4: if the fraction of random stabilizer bases whose minimum expectation exceeds 1−αε is not below the claimed O(((α+1)/2)^n) bound, the false-positive analysis is wrong.","supporting_citations":[{"cited_title":"Lower bounds for the fidelity of entangled state preparation","cited_arxiv_id":"quant-ph/0606023","evidence_quote":"Gives the fixed-basis infidelity bound ε ≤ (n/2)(1−μ) that BMoM replaces by a minimum and amplifies by randomization."}],"review_version":1}