{"id":"2e22fcbb-fd05-4a2a-9691-f8ab0d490181","arxiv_id":"2607.14212","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Under local depolarizing noise, the fidelity of a quantum circuit on classically simulable stabilizer-scar states approximates its fidelity on classically hard typical states, enabling scalable verification of non-equilibrium quantum simulation.","lead":"This paper proposes a method for checking a quantum simulator's accuracy on hard problems by measuring its accuracy on easy \"stabilizer scar\" states run through the same circuit. If correct, this gives a practical benchmark for when noisy quantum devices truly beat classical computers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fidelity equivalence proven only for Haar-averaged inputs; transfer to fixed circuit-evolved benchmark states rests solely on n=10 numerics.","rationale":"The reader identifies the transfer from Haar-averaged inputs to fixed circuit-evolved states as the weakest assumption. This is indeed the most load-bearing gap: the paper's analytical proof of fidelity equivalence (Eqs. 8–15) applies to averages over input states to the error channel, whereas the proposed benchmark measures fidelity for specific initial bitstrings evolved through the circuit. The numerical evidence is at n=10, which is far from the 'at scale' regime, uses random subspace-preserving circuits rather than the Trotterized evolution of the physical model, and does not provide a quantitative bound on the deviation for individual states. The Lévy concentration argument does not bridge the gap because circuit-evolved states form a highly structured subset of Hilbert space. The paper itself acknowledges the limitation, so a conditional verdict is appropriate. No new independent concern outweighs this one; the MCMC sampler inefficiency is a practical issue that could be resolved with a better sampler, while the average-to-individual transfer is a core proof gap. Therefore, the reader's CONDITIONAL verdict should stand unchanged.","tokens_in":26260,"tokens_out":8423,"duration_ms":89417,"concrete_test":"Using exact state-vector simulation, run the noisy Trotterized circuit for the model of Eq. (11) with local depolarizing noise p=0.01 and p=0.1 for system sizes n=12,14,16,18,20 (2L qubits, odd L). For each n, choose 10 scar bitstrings and 10 generic bitstrings that produce non-scar states via the state-preparation circuit of Fig. S1(c). Compute the exact output fidelities F(ρ_S,σ_S) and F(ρ_H,σ_H) and record the maximum and mean absolute difference over the sampled inputs. If the mean difference does not decrease with n (beyond the trivial saturation at 1/d when p_eff→1), the equivalence for fixed circuit-evolved states is not supported. Additionally, compute the spread over the 10 inputs to check concentration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (8) defines the average fidelity F(μ)=E_{|ψ⟩∼μ}⟨ψ|E(|ψ⟩⟨ψ|)|ψ⟩, and Eqs. (9)–(15) show F(μ_H)≈F(μ_S) for a single local depolarizing channel acting on Haar-random input states. The actual benchmark, however, uses a fixed computational-basis bitstring fed through a state-preparation circuit and a deep evolution circuit with depolarizing noise after every gate. The state entering each noise channel is a deterministic function of the input, not a Haar-random state. The Discussion explicitly concedes: '(iii) has only been proven as an average over input states to the error channel.' The numerical support (Fig. 4) is limited to n=10 qubits, random subspace-preserving circuits, and one effective error rate p_eff; it does not test the actual Trotterized evolution of Eq. (11) at large n, nor does it bound the deviation for fixed states. The Lévy concentration argument in SM S3c concerns Haar-random states and need not apply to the structured set of states produced by the circuit. Since the headline claim is that the measured scar-state fidelity 'bounds the fidelity of classically intractable simulations' for a specific experiment, the missing step is precisely the transfer from average to individual circuit-evolved states.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a benchmarking protocol for non-equilibrium quantum simulation based on stabilizer scars. For a Z2-lattice-gauge-theory-dual Ising model with a stabilizer-scar subspace, the authors argue (i) scar dynamics are classically simulable; (ii) direct fidelity estimation (DFE) is efficient for scar states, using bounds on the stabilizer Rényi entropy and a Markov-chain sampler; and (iii) under local depolarizing noise, the fidelity of scar states approximately equals the fidelity of typical, classically intractable non-scar states, so scar fidelity can certify device performance on hard dynamics. The analytic core, Eqs. (8)--(15), computes the average channel fidelity for Haar-random inputs in the full Hilbert space and in the scar subspace and bounds the difference. Numerical demonstrations include MCMC-based DFE for up to n=142 qubits and noisy qiskit simulations for n=10 qubits. The authors explicitly concede in the Discussion that claim (iii) is proven only as an average over input states to the error channel.","tokens_in":26610,"tokens_out":12437,"duration_ms":123162,"significance":"If the central equivalence were established for the actual circuit-evolved states used in a benchmark experiment, the protocol would be a genuinely useful tool: it would convert a classically simulable stabilizer-scar subspace into a quantitative, efficiently measurable proxy for device fidelity on thermalizing, classically hard dynamics. The analytic treatment is parameter-free and clearly presented, the bounds in Eqs. (9)--(15) are explicit, and the numerical MCMC demonstrations with both exact and shot-limited data are concrete and encouraging. The paper is also refreshingly candid about the main limitation. However, the missing transfer from Haar-averaged single-channel fidelity to fixed, circuit-evolved states is the load-bearing step for the headline claim, so the result as stated is not yet sufficiently supported.","major_comments":[{"comment":"The central claim (iii) is proven only for a single depolarizing channel acting on Haar-random input states (Eq. (8) and SM S3). The benchmark protocol, by contrast, initializes a fixed computational-basis bitstring, applies a deterministic circuit, and inserts depolarizing noise after every gate; the state entering each noise channel is a deterministic circuit output, not a Haar-random state. Haar-measure concentration (SM S3c) does not apply to this structured set. The numerical evidence in Fig. 4 is n=10 with random subspace-preserving circuits (Eq. (16)) and a single effective error rate; it does not test the Trotterized evolution of Eq. (11) at large n, nor does it provide a bound for fixed states. Since the abstract claims that scar fidelity 'bounds the fidelity of classically intractable simulations,' this average-to-individual transfer must be proved, or the claim must be restric","section":"Fidelity estimation for classically non-simulable states; Discussion"},{"comment":"The scalable-verification claim requires an efficient classical procedure to sample from the Pauli distribution P_rho of the target scar state. The MCMC algorithm used here proposes uniformly from P_S, whose size is d(n^2-n+2) -- exponential in n -- and no mixing-time bound is supplied. The Discussion explicitly states that the MCMC algorithm 'is not efficient.' Without a polynomial-time sampler or a mixing-time guarantee, the total classical overhead of the DFE protocol is not established, and the protocol cannot yet be called scalable at large n. The n=142 numerical chains are useful empirical evidence but do not constitute a complexity guarantee.","section":"Numerical Simulations; SM S4; Discussion"}],"minor_comments":[{"comment":"The tail bound in Eq. (6) does not by itself justify the claim that samples with Pauli expectation significantly below the mean are 'exponentially unlikely.' Hoeffding's inequality with range [0,1] gives the exponent 2 gamma^2; to rule out X < mean/2 one needs gamma ~ 1/(2 d_s^4), for which the bound 2 exp(-2 gamma^2) is close to 2 and vacuous. The rigorous efficiency guarantee comes from the epsilon-truncation argument in SM S1 (Eq. (S12)) rather than from Eq. (6). Please revise the text so that the truncation argument is the primary support, and avoid overstating the concentration result.","section":"Eq. (6) and SM S1"},{"comment":"The symbol 's' is used both for the input bitstring and for the number of MCMC samples. This creates confusion in Fig. 4(a) and the surrounding text; please use distinct notation.","section":"Fig. 4 caption and main text"},{"comment":"There is a typo: 'A next step is is the implementation' should read 'A next step is the implementation.'","section":"Discussion, final paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper's own Discussion concedes the central gap (iii) is unproven for fixed circuit-evolved states; this is not a hidden flaw but a stated limitation. The revision should either prove a typicality or worst-case statement for the specific circuit family or explicitly reframe the protocol as an average-fidelity benchmark. The MCMC-efficiency issue is a second substantial gap that should be addressed. The analytic bounds and numerical demonstrations are solid enough that a major revision, rather than rejection, seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a genuinely new and clean analytic result — under local depolarizing noise the average fidelity over a stabilizer-scar subspace matches the average fidelity over the full Hilbert space, asymptotically in n and for small p. The derivation in Eqs. (9)–(15) is parameter-free and the trace algebra is clean. The tighter bounds in SM S3 are a real technical contribution. The DFE analysis for scar subspaces is also careful: the stabilizer-Renyi-entropy bound gives a polynomial shot cost, and the MCMC demonstration at up to 142 qubits shows the estimator behaves as advertised.\n\nThe soft spot is exactly where you expect it. The headline claim — that scar-state fidelity bounds the fidelity of classically intractable simulations — requires transferring the average-fidelity equivalence to fixed, circuit-evolved non-scar states. What is proven is only an average over Haar-random input states to the error channel. The Discussion concedes this, and the numerical support is n=10 qiskit simulations with random subspace-preserving circuits, not the actual Trotterized evolution at scale. That is a load-bearing gap. It may be addressable — a stronger statement about local depolarizing noise being subspace-agnostic, or a proof for structured states, or at least larger numerics on the actual circuit — but as written the transfer is a conjecture.\n\nTwo smaller caveats: the MCMC sampler they use is not efficient, and the paper says so; the scalable DFE claim therefore depends on a future proposal scheme. And the local-depolarizing-noise assumption is strong. The authors acknowledge both, so it is not a hidden flaw, but it does limit how ready the protocol is for direct experimental use.\n\nWho is this for: people working on quantum simulation verification and benchmarking, particularly those interested in using many-body scars or stabilizer structure to certify advantage. The paper deserves a serious referee: the analytic core is solid, the limitations are explicit, and referee time could push the authors to close the gap or reframe the claim. I would read it and cite the average-fidelity equivalence, but I would not yet use the protocol as a black-box benchmark in an experiment.\n\nRecommendation: send to peer review, with a request to either prove or substantially strengthen (iii), and to soften the abstract to match what is proven.","headline":"Genuinely new analytic result with an honestly admitted gap: the average-fidelity equivalence is proven, but the transfer to fixed circuit-evolved states rests on n=10 numerics; worth refereeing, not worth taking as a proven benchmark.","tokens_in":27036,"tokens_out":2443,"would_cite":true,"duration_ms":25060,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes that measuring fidelity on classically simulable stabilizer-scar states yields a benchmark for the fidelity of classically hard quantum simulations, under local depolarizing noise.","keywords":["quantum advantage","benchmarking quantum devices","stabilizer scars","direct fidelity estimation","quantum many-body scars","local depolarizing noise","non-equilibrium dynamics","lattice gauge theory"],"falsifier":"A concrete test would be to classically emulate a circuit with a specific non-scar bitstring input at gradually larger qubit counts (e.g., n=12, 14, 16) under local depolarizing noise and compute the exact fidelity of the output; if the non-scar fidelity deviates from the scar-subspace bound by more than the predicted subleading correction as n grows, the transfer from average to individual states fails. A second test would repeat the same comparison under a different error model, such as amplitude damping or cross-talk noise, to see whether the equivalence persists.","tokens_in":26193,"feed_emoji":"⚛️","tokens_out":4114,"duration_ms":43719,"temperature":0.7,"pith_summary":"The paper proposes a protocol for verifying quantum simulations of non-equilibrium dynamics at scales where classical simulation is intractable. It argues that, if errors are dominated by local depolarizing noise, the fidelity of states in a special classically simulable 'stabilizer scar' subspace equals, to within asymptotically vanishing corrections, the fidelity of typical classically hard states evolved under the same circuit. Because scar-state fidelities can be measured efficiently with direct fidelity estimation, one can certify a device's performance on hard states by checking easy states. This turns quantum many-body scars from a physical curiosity into a practical benchmark for quantum advantage.","feed_headline":"Easy states can certify hard quantum simulations","feed_subtitle":"Measure fidelity on stabilizer-scar states; under local noise it matches the fidelity of classically hard states.","key_machinery":"The method combines three ingredients: (i) an exactly solvable stabilizer-scar subspace of dimension polynomial in qubit number, spanned by stabilizer states with closed dynamics; (ii) direct fidelity estimation (DFE), which estimates the fidelity of a scar target state from a polynomial number of Pauli samples and measurement shots; (iii) an average-fidelity equivalence between scar and Haar-random input states under local depolarizing noise, proved for the channel and supported by a Markov Chain Monte Carlo sampler for the Pauli distribution.","core_discovery":"The central claim is that stabilizer scars—a polynomially large subspace spanned by stabilizer states that remains invariant under a non-integrable Hamiltonian—provide a scalable verification handle for quantum simulation. The authors show that, under a physically motivated local depolarizing error model, the average fidelity of states drawn uniformly from the scar subspace is asymptotically equal to the average fidelity of Haar-random states from the full Hilbert space. Since scar states are classically simulable and their fidelity can be estimated efficiently by sampling Pauli observables, the measured scar fidelity serves as a bound/proxy for the fidelity of classically intractable non-sc","pith_inferences":["If the average-to-specific transfer holds on real hardware, the protocol could be used as a routine pre-flight diagnostic: measure scar fidelity before running a hard simulation to predict the fidelity of the hard output.","The underlying principle—a classically simulable subspace with efficient DFE can certify fidelity of hard dynamics—may extend beyond stabilizer scars to other structured subspaces embedded in thermalizing spectra.","The benchmark's reliability hinges on the noise being local and unstructured; testing with correlated noise, amplitude damping, or non-Markovian error models would reveal how far the equivalence extends beyond depolarizing channels."],"forward_implications":["A device's performance on classically hard non-equilibrium states can be certified by measuring only easy scar-subspace states, without classically solving the hard problem.","The measurement overhead scales polynomially with system size, enabling benchmarks at scales where classical simulation is impossible.","The protocol gives a concrete route to quantum-advantage experiments in non-equilibrium dynamics, particularly for lattice gauge theory models that host such scars.","Concentration of measure implies the fidelity equivalence holds not just on average but for the overwhelming majority of input states.","The method provides a uniform benchmark across scar and non-scar dynamics, with corrections that vanish asymptotically in system size or noise strength."],"fun_headline_variants":["Stabilizer scars certify quantum advantage","Scar states verify hard quantum simulations","Easy-state benchmark for quantum sims","Quantum sim checks via scar fidelity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The fidelity equivalence is proven as an average over Haar-random input states to the noise channel, but the benchmark relies on the untested assumption that this same equivalence holds for the specific fixed non-scar bitstring input states that an actual benchmark circuit would use; numerical evidence is only provided for a small system size.","fun_headline_variants_meta":{"raw":{"variants":["Stabilizer scars certify quantum advantage","Scar states verify hard quantum simulations","Easy-state benchmark for quantum sims","Quantum sim checks via scar fidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1117,"prompt_tokens":579,"completion_tokens":538,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":323,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":323,"tokens_out":538,"duration_ms":5705,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:46:40.541887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to classically emulate a circuit with a specific non-scar bitstring input at gradually larger qubit counts (e.g., n=12, 14, 16) under local depolarizing noise and compute the exact fidelity of the output; if the non-scar fidelity deviates from the scar-subspace bound by more than the predicted subleading correction as n grows, the transfer from average to individual states fails. A second test would repeat the same comparison under a different error model, such as amplitude damping or cross-talk noise, to see whether the equivalence persists.","supporting_citations":[],"review_version":1}