{"id":"4647db74-f5ad-48d6-8554-0ecf1a7051e0","arxiv_id":"2502.09015","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Random circuits satisfy an ergodicity condition for positive-coefficient polynomials, and its deviation can benchmark quantum chip fidelity, recovering and generalizing linear cross-entropy benchmarking.","lead":"This paper introduces 'ergodicity' to random circuit sampling, showing that averages of certain functions of output probabilities over bitstrings converge to averages over random circuits. It uses deviations from this ergodicity to build cross-entropy benchmarks that estimate circuit fidelity, generalizing Google's linear XEB.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 2.1's weak-noise fidelity estimate rests on an invalid Lévy-lemma concentration step: the stated Lipschitz bound is a lower bound and yields a trivial probability bound, so the central XEB-recovery claim is unproven.","rationale":"The reader's verdict CONDITIONAL is appropriate. Theorem 1 and the global depolarizing case study are supported by explicit Haar-integral calculations and are not challenged here. The load-bearing gap is Proposition 2.1: the weak-noise fidelity formula depends on a concentration statement that is not proved. The concrete defects are (i) the weak-correlation assumption, as written for f(p) = N^2 p^2, is used in Eq. (36) for the linear function N^2 P, so the mean E[C_f] ≈ 1 is not derived; and (ii) the Lévy-lemma step uses a lower bound on the Lipschitz constant where an upper bound is required, and the resulting exponent is O(1/N^2), so the claimed high-probability bound is absent. Because the paper's abstract and introduction advertise fidelity estimation for weakly correlated noise and recovery of Google's linear XEB, this unproven proposition is central, not peripheral. I agree with the reader that the ergodicity theorems stand and that the weak-noise claim needs a corrected proof or an explicit concentration bound; hence the verdict remains CONDITIONAL. I mark agreement as partial because the reader focused on the Lévy bound but did not flag the mismatch between the assumed weak-correlation function and the one used in Eq. (36).","tokens_in":22034,"tokens_out":13151,"duration_ms":116677,"concrete_test":"Recompute the Lévy-lemma probability in Proposition 2.1: with ε = 1/√N, η = N√2, and dimension factor 2N+1, the RHS of Lemma 2 is 2 exp[-C(2N+1)/(2N^3)], which is ≈2 for large N, whereas the paper claims 2 exp[-C(2N+1)N]. This single arithmetic check settles that the concentration step is invalid. As a second check, compute the variance of C_f(P_U, χ_U) over Haar U for the noise model χ_U = |0⟩⟨0| (so C_f = N P_U(0)); the standard deviation is Θ(1), demonstrating that Lipschitz constants of order N do not yield O(1/√N) concentration. If such a noise model also satisfies (32), the proposition is false; if not, the paper must prove that (32) rules it out, which it does not.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline benchmarking claim for weakly correlated noise is Proposition 2.1, which asserts F = 1 - DE_f for f(p) = N^2 p^2 and thereby recovers Google's linear XEB. The proof requires C_f(P_U, χ_U) = N Σ_x P_U(x) χ_U(x) to concentrate around 1 with fluctuations O(1/√N). This is not established. First, the weak-correlation assumption (32) is stated for f(p) = N^2 p^2: E[χ_U f(P_U)] ≈ E[χ_U] E[f(P_U)]. But Eq. (36) applies it to the linear function N^2 P_U, i.e., it needs E[χ_U P_U] ≈ E[χ_U] E[P_U]. The quadratic assumption does not imply the linear one, so the computation of E[C_f] = 1 + O(1/(N√N)) is unsupported. Second, even conditional on that mean, the concentration argument via Lévy's lemma is invalid: the proof only gives a lower bound η ≥ N√2 on the Lipschitz constant, whereas Lévy's lemma requires an upper bound. Worse, inserting η = N√2 and ε = 1/√N into Lemma 2 yields a probability bound 2 exp[-C(2N+1)/(2N^3)] ≈ 2, not the claimed 2 exp[-C(2N+1)N]. The bound is trivial, so one cannot conclude C_f = E[C_f] ± O(1/√N). Indeed, for noise χ_U peaked on a single computational basis state, C_f = N P_U(0) has standard deviation Θ(1), not O(1/√N), absent further restrictions. Consequently, the fidelity estimation F = 1 - DE_f for weakly correlated noise is not a theorem; the main advertised application to linear XEB is left as an assumption. The global depolarizing noise analysis (Section III.B.a) is unaffected and appears correct.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a notion of ergodicity for random circuit ensembles: for functions f of the output probabilities, the average of f over output bitstrings of a typical circuit is close to the ensemble average. Theorem 1 proves this for positive-coefficient polynomials of degree t when the circuit ensemble is a unitary 2t-design, via explicit Haar-integral computations showing negative covariances between distinct bitstrings. Theorem 2 extends the result to f(p)=p ln p for Haar-random unitaries. The authors then define a benchmarking quantity, the deviation of ergodicity (DE_f), and analyze it for global depolarizing noise and for a weak-correlation noise model. For the quadratic choice f(p)=N^2 p^2, they claim to recover linear cross-entropy benchmarking (XEB), i.e., that DE_f estimates the circuit fidelity F. They also analyze Google's Sycamore data for several polynomial scheme functions.","tokens_in":22385,"tokens_out":5503,"duration_ms":54202,"significance":"The ergodicity theorems (Theorem 1 and Theorem 2) are a genuine mathematical contribution: they establish a clean concentration baseline for random circuit output statistics, with proofs based on explicit Haar integrals and negative covariance bounds. The global depolarizing noise analysis (Section III.B.a) is a correct and useful sanity check, showing that DE_f is a linear function of fidelity in that model. However, the paper's headline applied claim, the recovery of linear XEB for weakly correlated noise (Proposition 2.1), is not proven and, as stated, is false. Since the abstract and introduction advertise this recovery as a central result, the paper's main benchmarking claim fails. The experimental section inherits this problem because its interpretation of the Sycamore data relies on the weak-noise relation that is not established.","major_comments":[{"comment":"The weak-correlation assumption (32) is stated only for f(p)=N^2 p^2, i.e., E[χ_U f(P_U)] ≈ E[χ_U] E[f(P_U)]. However, Eq. (36) applies the same factorization to the linear function N^2 P_U, requiring E[χ_U P_U] ≈ E[χ_U] E[P_U]. The quadratic assumption does not imply the linear one, so the computation of E[C_f(P_U,χ_U)] = 1 + O(1/(N√N)) is unsupported by the stated hypothesis.","section":"Section III.B.b, Eq. (36)"},{"comment":"The proof states that C_f(P_U,χ_U) is Lipschitz with constant η ≥ N√2 and then invokes Lévy's lemma (Lemma 2, Appendix F). Lévy's lemma requires an upper bound on the Lipschitz constant; a lower bound is useless for concentration. Moreover, inserting η = N√2 and ε = 1/√N into Lemma 2 gives the trivial bound 2 exp[-C(2N+1)/(2N^3)] ≈ 2, not the claimed 2 exp[-C(2N+1)N]. The concentration claim is therefore not derived.","section":"Section III.B.b, Eq. (39)"},{"comment":"Even if the mean were correctly computed, the fluctuation of C_f(P_U,χ_U) around its mean need not be O(1/√N). Consider noise where χ_U is a random computational basis state |π(U)⟩⟨π(U)|, with π(U) uniformly random and independent of U. This satisfies E[χ_U(x0)] = 1/N and also satisfies Eq. (32) exactly, because E[χ_U(x0) f(P_U(x0))] = (1/N) E[f(P_U(x0))] = E[χ_U(x0)] E[f(P_U(x0))]. Yet C_f(P_U,χ_U) = N P_U(π(U)), whose variance over U is (N-1)/(N+1) ≈ 1. Thus C_f fluctuates by order 1, not O(1/√N), and Eq. (40) fails. Consequently Proposition 2.1 is false as stated, and the claimed recovery of linear XEB, Eq. (45), does not hold under the stated assumptions.","section":"Section III.B.b, Eqs. (37)-(40)"}],"minor_comments":[{"comment":"The word \"ergdicity\" should be \"ergodicity\".","section":"Fig. 2 caption"},{"comment":"The sentence \"our analysis in Section III B dose not contradict\" contains a typo: \"dose\" should be \"does\".","section":"Section III.D"},{"comment":"There are several OCR-style spacing artifacts in the text, such as \"o ffer\" and \"su fficient\"; these should be cleaned up.","section":"Section I"},{"comment":"The statement that the correction term o(p) can be expressed as O(1) independent of N is imprecise; the limit should be justified more carefully, though this does not affect the main results.","section":"Appendix A, Eq. (A9)"}],"recommendation":"reject","confidential_remarks":"The ergodicity theorems and the global depolarizing analysis are sound and could form the basis of a useful paper. However, Proposition 2.1, which is the main advertised application, is not merely unproven but false under the stated weak-correlation assumption, as shown by the independent-random-permutation counterexample. Substantial reworking of the claims and assumptions would be needed, and the current manuscript should not be accepted. The editor may wish to encourage the authors to resubmit a version that either drops the weak-noise XEB claim or replaces it with a properly justified sufficient condition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read. Two things to know: the ergodicity theorem is real and clean; the advertised weak-noise recovery of linear XEB is not proved.\n\nWhat is new: Theorem 1 shows that for a unitary 2t-design, the bitstring average of a positive-coefficient polynomial in output probabilities concentrates around the Haar ensemble average, with Chebyshev and explicit negative covariances. The covariance integral in Appendix D is correct. Theorem 2 adds p ln p for Haar. The global depolarizing noise case study is straightforward and correct, and it genuinely generalizes the estimator to f=N^i p^i. The experimental section is minor—reanalysis of Google's data—but it does not overclaim.\n\nThe soft spot is Proposition 2.1, and it is load-bearing. The weak-correlation assumption (32) is about the quadratic f(p)=N^2p^2, i.e. a weighted second moment of P, but equation (36) uses it for the linear function N^2P, a first moment. The quadratic assumption does not imply the linear one, so the computation of E[C_f(P,χ)]=1 is unsupported. The concentration step is also invalid: the proof gives η≥N√2, a lower bound on the Lipschitz constant, while Lévy's lemma requires an upper bound; plugging the lower bound into the lemma gives a probability bound of order 2e^{-C/N^2}, trivially close to 2. The paper's own Appendix F makes exactly this point about bitstring averages, and Proposition 2.1 repeats the mistake. The flaw is not a technicality: for a noise term peaked on one computational basis state, C_f can fluctuate by order one, not 1/√N. So the claim that the framework recovers linear XEB under weak correlated noise is currently unproved. The global depolarizing analysis is not affected.\n\nWho gets value? People working on quantum certification will want Theorem 1 and the depolarizing-noise estimators. The weak-noise claim needs either a corrected proof or an explicit downgrade to an assumption. I would send this to peer review, because the solid part is substantial and the broken part is isolated enough that a referee can demand a fix.","headline":"Solid ergodicity theorem and clean global-depolarizing benchmark; the advertised weak-noise recovery of linear XEB rests on an invalid Lévy step and an unmatched weak-correlation assumption.","tokens_in":22992,"tokens_out":4601,"would_cite":true,"duration_ms":45523,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","81P45","60B20"],"pacs":["03.67.-a","03.67.Lx"],"model":"deepseek-v4-flash","headline":"Random circuit ensembles satisfy an ergodicity condition—bitstring averages concentrate around ensemble averages—and the deviation from it estimates a chip's fidelity, generalizing linear cross-entropy benchmarking.","keywords":["ergodicity","random circuit sampling","cross-entropy benchmarking","unitary t-design","circuit fidelity","depolarizing noise","concentration of measure","quantum certification"],"falsifier":"Take a concrete weakly correlated noise model, such as single-qubit depolarizing noise applied after a few gates, and for small $N$ (e.g., $N=8$ or $16$) compute $C_f(P_U,\\chi_U)$ exactly for many Haar-random unitaries $U$. If the empirical probability that $|C_f-\\mathbb{E}[C_f]|\\ge 1/\\sqrt{N}$ decays far more slowly than the claimed bound $2\\exp[-C(2N+1)N]$, or does not decay at all, then Proposition 2.1's concentration step fails and the fidelity formula $F=1-\\widehat{DE}_f$ is unsupported.","tokens_in":21741,"feed_emoji":"🎲","tokens_out":12464,"duration_ms":115131,"temperature":0.7,"pith_summary":"This paper establishes an ergodicity property of random circuit sampling: for a circuit drawn from an ensemble that forms a unitary $2t$-design, the average of any degree-$t$ polynomial with non-negative coefficients over the output bitstrings of a single circuit is close—with probability at least $1-1/\\alpha^2$—to the same polynomial averaged over the circuit ensemble, up to a gap of order $\\sigma_f/\\sqrt{N}$ with $N=2^n$. The paper then turns the size of this gap into a benchmarking signal, defining the deviation of ergodicity and showing that it estimates circuit fidelity under global depolarizing noise and under weakly correlated noise. For the quadratic scheme function $f(p)=N^2p^2$ it recovers linear cross-entropy benchmarking as a special case, and for $f(p)=p\\ln p$ it gives a logarithmic variant. The practical point is that the benchmark needs only $T=O(\\mathrm{poly}(\\log N))$ samples, far fewer than the Hilbert space dimension, making it usable in the supremacy regime.","feed_headline":"Random circuits pass an ergodicity test that benchmarks chips","feed_subtitle":"Bitstring averages concentrate around Haar averages; the gap estimates fidelity, generalizing linear XEB.","key_machinery":"The engine is the unitary $2t$-design—a finite set of unitaries that matches the Haar measure through the $2t$-th moment—together with Lemma 1's covariance identity, which shows that the output probabilities at two distinct bitstrings are negatively correlated over the unitary ensemble. This negative covariance is what lets Chebyshev's inequality produce the $1/\\alpha^2$ concentration bound, and it is also the reason the argument needs a $2t$-design rather than a $t$-design. For the benchmarking scheme, the operative object is the deviation of ergodicity $DE_f$, estimated from $T$ experimental samples via $\\widehat{C}_f=\\frac{1}{NT}\\sum_i f(P_U(x_i))/P_U(x_i)$, which is unbiased and needs only $T=O(\\mathrm{poly}(\\log N))$ samples. A second mechanism, Levy's lemma, is invoked to concentrate the noise correlation $C_f(P_U,\\chi_U)$ in the weakly correlated noise analysis.","core_discovery":"The paper's central claim is Theorem 1: a unitary $2t$-design is ergodic relative to any polynomial $f(p)=\\sum_{i=1}^{t}a_ip^i+b$ with $a_i\\ge 0$, meaning that for every $\\alpha>0$, the probability that the bitstring average $\\frac{1}{N}\\sum_x f(P_U(x))$ differs from the ensemble average $\\mathbb{E}_U[f(P_U)]$ by at least $\\alpha\\sigma_f/\\sqrt{N}$ is at most $1/\\alpha^2$. The proof runs through Chebyshev's inequality and the negative-covariance identity $\\mathrm{Cov}(P_U^{q_1}(x),P_U^{q_2}(y))<0$ for distinct bitstrings $x,y$, which makes the variance of the bitstring average smaller than $\\sigma_f^2/N$. The paper defines the deviation of ergodicity $DE_f=|\\mathbb{E}_U[f(P_U)]-C_f(P_U,Q_U)|$ with $C_f$ the correlation between ideal and experimental output distributions, proves the deviation estimates fidelity for global depolarizing noise, and gives a sufficient condition under which it does so for weakly correlated noise, recovering the linear cross-entropy benchmark for $f(p)=N^2p^2$.","pith_inferences":["The benchmark's sensitivity is not fixed by the theorem: replacing a single polynomial degree by several degrees $i=2,3,4$ gives independent estimates of the same fidelity, and discrepancies among them would flag correlated noise that the weakly correlated model cannot handle; this is a natural extension the paper does not develop.","The positive-coefficient condition in Theorem 1 suggests a sharp boundary for ergodicity: testing functions with mixed signs, such as $f(p)=p-p^2$, could reveal whether the concentration relies specifically on monotone post-processing and might lead to benchmark functions tailored to particular noise channels.","Since the deviation of ergodicity is computed from ideal probabilities of sampled bitstrings, it is exposed to the same classical spoofing threats as linear cross-entropy benchmarking; a classical sampler that mimics the ideal distribution's statistics could keep $\\widehat{DE}_f$ small without high fidelity, so the scheme's certification power depends on the hardness of such spoofing.","The paper's Levy-lemma route to concentrating $C_f(P_U,\\chi_U)$ appears to give a trivial bound; if a different argument can close that gap, the weakly correlated noise result would become a theorem, and if not, the scheme's validity in practical regimes rests on empirical noise assumptions."],"forward_implications":["For a noiseless device the estimator $\\widehat{DE}_f$ is $O(\\sigma_f/\\sqrt{N})$ with high probability, so a statistically significant large deviation directly signals noise.","Under global depolarizing noise with fidelity $F$, the deviation satisfies $\\widehat{DE}_f=(1-F)(i-1)!(i-1)\\pm O(1/\\sqrt{T})$ for the polynomial $f(p)=N^ip^i$, giving an explicit way to extract $F$ from samples.","Choosing $f(p)=N^2p^2$ reproduces the linear cross-entropy benchmarking estimator $F_{\\mathrm{XEB}}=F\\pm O(1/\\sqrt{T})$ under the paper's weak-correlation condition, so the new framework contains the standard XEB protocol as one instance.","The ergodicity result for $f(p)=p\\ln p$ provides a logarithmic cross-entropy variant with the same sample-count scaling.","Because the estimate requires only $T=O(\\mathrm{poly}(\\log N))$ samples, the scheme remains practical when the number of qubits is large enough that most bitstrings are never observed twice."],"supporting_citations":[{"why":"It introduces the Porter-Thomas distribution and linear cross-entropy benchmarking, the baseline the generalized scheme extends and reproduces.","marker":"[16]"},{"why":"It supplies the 53-qubit random-circuit sampling dataset reanalyzed in the numerical section and the linear XEB fidelity result that the framework recovers.","marker":"[17]"},{"why":"It shows weak unital local noise scrambles into an effective global depolarizing channel, which is the justification for the depolarizing-noise fidelity analysis.","marker":"[57]"},{"why":"It provides the evidence that adversarially correlated noise can separate benchmarking metrics from fidelity, bounding the regime where the deviation-of-ergodicity estimate is valid.","marker":"[59]"},{"why":"It defines unitary designs, the object appearing in the hypothesis of Theorem 1.","marker":"[60]"},{"why":"It supplies structural results on unitary designs used in stating the $2t$-design relaxation of the Haar ensemble.","marker":"[61]"},{"why":"It states Levy's lemma, which Proposition 2.1 invokes to concentrate the noise correlation term.","marker":"[66]"},{"why":"It gives the Haar-integration moment machinery used to prove the negative-covariance identity in Lemma 1.","marker":"[68]"},{"why":"It provides the closed-form Haar moment formula for $|\\langle i|U|0\\rangle|^{2q_i}$ used in the covariance proof.","marker":"[71]"}],"fun_headline_variants":["Ergodicity test benchmarks quantum chips from random circuits","Random circuit benchmarking via ergodicity, not just XEB","Bitstring averages yield new quantum fidelity estimator","Ergodicity links random circuits and noise, chalks fidelity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the correlation between ideal output probabilities and the noise part stays close to its average for every circuit; if it does not, the deviation of ergodicity is not a reliable measure of fidelity under weakly correlated noise, and the paper's argument for that concentration is incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Ergodicity test benchmarks quantum chips from random circuits","Random circuit benchmarking via ergodicity, not just XEB","Bitstring averages yield new quantum fidelity estimator","Ergodicity links random circuits and noise, chalks fidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000122,"raw_usage":{"total_tokens":1146,"prompt_tokens":1043,"completion_tokens":103,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":38}},"tokens_in":659,"tokens_out":103,"duration_ms":1746,"temperature":1.0,"reasoning_tokens":38,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T22:55:02.513430+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a concrete weakly correlated noise model, such as single-qubit depolarizing noise applied after a few gates, and for small $N$ (e.g., $N=8$ or $16$) compute $C_f(P_U,\\chi_U)$ exactly for many Haar-random unitaries $U$. If the empirical probability that $|C_f-\\mathbb{E}[C_f]|\\ge 1/\\sqrt{N}$ decays far more slowly than the claimed bound $2\\exp[-C(2N+1)N]$, or does not decay at all, then Proposition 2.1's concentration step fails and the fidelity formula $F=1-\\widehat{DE}_f$ is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It introduces the Porter-Thomas distribution and linear cross-entropy benchmarking, the baseline the generalized scheme extends and reproduces."},{"cited_title":"Tight bounds on the convergence of noisy random circuits to the uniform distribution","cited_arxiv_id":"2112.00716","evidence_quote":"It provides the evidence that adversarially correlated noise can separate benchmarking metrics from fidelity, bounding the regime where the deviation-of-ergodicity estimate is valid."},{"cited_title":"Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage","cited_arxiv_id":"2112.01657","evidence_quote":"It defines unitary designs, the object appearing in the hypothesis of Theorem 1."},{"cited_title":"Ledoux, The concentration of measure phenomenon, nachdr","cited_arxiv_id":null,"evidence_quote":"It states Levy's lemma, which Proposition 2.1 invokes to concentrate the noise correlation term."},{"cited_title":"Ledoux, The concentration of measure phenomenon , 89 (American Mathematical Soc., 2001)","cited_arxiv_id":null,"evidence_quote":"It gives the Haar-integration moment machinery used to prove the negative-covariance identity in Lemma 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the closed-form Haar moment formula for $|\\langle i|U|0\\rangle|^{2q_i}$ used in the covariance proof."}],"review_version":1}