{"id":"e146b5a5-0ab4-483d-8a0f-a28e16e5d21c","arxiv_id":"1909.01295","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Analogue randomized benchmarking (ARB) measures the average error rate per time evolution for a family of Hamiltonians on an analogue quantum simulator, demonstrated in classical simulations of XY spin chains.","lead":"This paper adapts randomized benchmarking, the standard error-measurement tool for digital quantum computers, to analogue quantum simulators that evolve continuously under a Hamiltonian. In simulations of six-spin models with several noise sources, the adapted protocol returns stable average error rates and clearly fails for one unitary set that does not scramble enough.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unproven 2-design assumption is load-bearing; the reported f and r values are not protected from an alternative averaging explanation.","rationale":"The paper's stated goal is to provide a scalable, SPAM-independent benchmarking method for analogue simulators, and its central quantity is the average error rate r. Every step that converts the measured survival probability into r goes through the assumption that the unitary set is an epsilon-approximate 2-design (Sec III A; Thm III.1; Thm III.2; App F). The authors themselves write in the conclusion that 'we have not formally proven that our unitary sets {Uk} are epsilon-approximate 2-designs'. The numerical support is indirect: a good fit to an exponential is compatible with the presence of a 2-design, but also with the averaging over 1000 unitaries and the specific noise models producing an effectively depolarizing ensemble. Section IV A 1 explicitly acknowledges this confound. Because no frame potential or second-moment distance is computed, the precondition is not demonstrated. The reader's conditional verdict is therefore appropriate. My read agrees with the reader's weakest-assumption identification; I do not see a different, more basic defect, though the unphysical perfect-inversion assumption and the absence of SPAM simulations are secondary limitations that reinforce the conditional status.","tokens_in":28434,"tokens_out":15514,"duration_ms":161279,"concrete_test":"For each of the four sets from Figs. 1 and 2 (nearest-neighbour/all-to-all × global/local, N=6, K=1000, dt=0.005), Monte Carlo estimate the unitary frame potential F_m = E_{U,V}|tr(U^dagger V)|^4 for random products of m unitaries drawn from the set, for m = 1, 10, 100, 1000 and for m up to the longest used sequence length T_max/dt. Compare F_m to the Haar value 2 (Def. E.2). If F_m does not approach 2 within the uncertainty of the ARB fits, the 2-design assumption fails and the decay interpretation is unsupported; if F_m is close to 2, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ARB extracts an average error rate r from the decay P_T = A + B f^T requires the sampled unitaries {U_k = exp(-iH_k dt)} to form an epsilon-approximate 2-design (Sec. III A, Thm. III.1 and III.2). The authors explicitly concede in Sec. V that this has not been formally proven, and the only supporting evidence is that the simulated survival data fit the exponential curve. That evidence is confounded: a finite set of K=1000 unitaries plus averaging over many sequences and the simple noise models could produce an approximately exponential decay without true 2-design twirling. The paper itself notes this possibility in Sec. IV A 1 ('it is possible that the averaging during ARB rather than twirling over an approximate 2-design is what causes the errors to behave like a depolarising channel'). No frame potential or other 2-design distance (Defs. E.2/E.3) is reported for the actual sets, so the premise that locks r to the average infidelity is unsupported. If the sets are only weak approximate designs, the l·epsilon bound in Thm III.1 is too loose for the long sequences used, and the fitted f cannot be interpreted as an error rate per evolution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes analogue randomized benchmarking (ARB), an adaptation of digital randomized benchmarking to programmable analogue quantum simulators. The protocol replaces discrete gates with unitaries U_k = exp(-i H_k dt) generated from a native Hamiltonian with added disorder, uses systematic time-inversion of each preceding unitary, and asserts that the average survival probability follows P_T = A + B f^T, from which an average error rate r = (d-1)(1-f)/d per unit time is extracted. The authors present classical simulations for nearest-neighbour and all-to-all XY models with several noise models (static fluctuations, weakly time-dependent errors, spontaneous emission, and noisy inversion), reporting fitted decay curves and 95% confidence intervals for r. They explicitly acknowledge that the unitary sets have not been proven to form epsilon-approximate 2-designs and that this is an assumption underlying the protocol.","tokens_in":28674,"tokens_out":6623,"duration_ms":69867,"significance":"If the central premise is established, ARB would fill a real gap: a scalable, SPAM-robust benchmarking method for analogue simulators that uses native operations rather than compiled gate sets. The paper is commendably transparent about its main assumption, and the numerical study is useful as a proof-of-principle for the protocol's curve-fitting machinery. However, the current evidence does not yet secure the interpretation of the fitted f as an average error rate: the 2-design property is unverified, the associated epsilon is never quantified, and the claimed SPAM robustness is not actually simulated. The contribution is therefore a promising protocol proposal with honest numerical illustrations, rather than a fully supported benchmarking method.","major_comments":[{"comment":"The load-bearing assumption that the sets {U_k = exp(-i H_k dt)} form an epsilon-approximate unitary 2-design is unproven, and the evidence offered for it is partly circular. The paper states in Sec. V that this has not been formally proven, and in Sec. IV A 1 it notes that the averaging during ARB, rather than twirling over an approximate 2-design, could be what makes the errors behave like a depolarising channel. The fact that simulated survival data fit an exponential decay does not distinguish these mechanisms. No frame potential or second-moment-operator distance (Defs. E.2 and E.3) is reported for the actual K=1000 sets, so the premise that connects the fitted f to average infidelity remains unsupported. I would ask the authors to compute and report a direct numerical estimate of epsilon for the unitary sets used, or otherwise provide a non-circular validation of the 2-design property.","section":"Sec. III A, Assumption III.1, Sec. V"},{"comment":"The stated bounds do not protect the reported r values. Theorem III.1 gives |P_alpha_l - P_mu_l| <= l*epsilon, which grows linearly with sequence length; for the long sequences actually fitted (e.g., up to T J = 150 with dt = 0.005, corresponding to l = 30000), this bound is vacuous unless epsilon is extraordinarily small. Theorem III.2 and Lemma F.5 give r' within epsilon of r only at l=1 and ps=1, but the protocol extracts f from a fit over all l, and Lemma F.3 shows the error in f grows with l. Since epsilon is never estimated, Eq. (16) is invoked without content, and the 95% confidence intervals reported in Eqs. (14)-(15) and (19)-(20) are purely statistical. The systematic error from the approximate design is therefore unquantified, and the claim that ARB produces a meaningful average error rate is not yet established.","section":"Thms. III.1-III.2, Eq. (16), Sec. IV A 1"},{"comment":"The abstract claims that ARB incorporates SPAM errors, but the numerical simulations are run with no state-preparation or measurement errors: Sec. IV A states 'we run the protocol with no errors in state preparation or measurement', and Eq. (10) fixes A = 1/d, B = (d-1)/d, which assumes no SPAM. The robustness of the A + B f^T form to SPAM is never tested, so the central advertised advantage over fidelity-estimation methods is not demonstrated. I would ask for at least one simulation with explicit SPAM errors (e.g., imperfect initial-state preparation and measurement misclassification) showing that the decay curve retains the A + B f^T form and that r is unaffected.","section":"Sec. IV A, Eq. (10), Abstract"}],"minor_comments":[{"comment":"In the comparison with directly computed average infidelities, the confidence interval for random product states is reported as 0.00150 (0.00145, 0.0000155); the upper endpoint 0.0000155 is almost certainly a typo and should likely be 0.00155.","section":"Sec. IV A 2"},{"comment":"The exponent in P_T = A + B' f^{2(T-1*dt)} is notationally confusing: as written it mixes time T and time-step dt in a way that is dimensionally inconsistent. It should be expressed in terms of the integer sequence length l (e.g., f^{2(l-1)}) or defined explicitly.","section":"Eq. (21), Sec. IV B 3"},{"comment":"The text describing Fig. 8 uses 'M = 6' where the system size is elsewhere denoted N; please use consistent notation.","section":"Appendix D"},{"comment":"References [19] and [20] spell the author name as 'Mageson'; the correct spelling is 'Magesan'.","section":"References"},{"comment":"The statement that 'with the standard error on our result we bound the unknown parameter epsilon' is not justified: standard errors quantify statistical fluctuations, not the systematic error epsilon entering through the approximate 2-design property. This sentence should be revised to reflect that epsilon remains unquantified.","section":"Sec. V"}],"recommendation":"major_revision","confidential_remarks":"The core idea is worth publishing if the central assumption is validated. I would regard a direct numerical estimate of epsilon (e.g., via the frame potential or moment operators of the actual unitary sets) and at least one SPAM-including simulation as necessary before the paper can be accepted; without these, the abstract's claims are stronger than the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the translation of randomized benchmarking from digital gate sets to continuous Hamiltonian evolutions. The authors construct unitary sets from disordered Hamiltonians, benchmark an error rate per unit time, and replace the single inversion gate with systematic Loschmidt-echo-style inversion. I have not seen that combination in the cited digital, approximate, or direct RB literature, so the novelty is genuine, not just a relabeling. The paper is also honest about the practical obstacles, especially time-reversal, which is the main experimental bottleneck.\n\nWhat the paper does well: the theory under the stated assumptions is standard, and the numerical case studies cover a reasonable range—nearest-neighbour and all-to-all XY models, global and local disorder, and several physically motivated noise models including spontaneous emission and weakly time-dependent errors. The fits to the predicted decay curve are shown with confidence intervals, and the comparison of r against directly computed average infidelities for random states is a sensible sanity check. The authors also flag their own confound: averaging during ARB, rather than true twirling over a 2-design, could produce the observed exponential decay.\n\nThe soft spots are real but not disqualifying. The load-bearing assumption that {U_k = exp(-i H_k dt)} forms an epsilon-approximate 2-design is not proven, and no frame potential or second-moment distance is reported for the actual sets. The fits to the decay curve are indirect evidence and are partly circular, since the same exponential form is assumed in the analysis. Second, SPAM robustness is claimed but never simulated—the case studies explicitly run with no SPAM errors, so the A+B f^T curve does not demonstrate the SPAM-absorption claim. Third, no code or data are provided to reproduce the figures. These are addressable gaps: compute (or bound) the 2-design distance for the actual unitaries, and rerun the protocol with modeled SPAM errors.\n\nThis paper is for the analogue-simulator and trapped-ion community, and for people working on randomized benchmarking beyond the Clifford group. It opens a direction rather than closing it, and the central conjecture—that disordered Hamiltonians generate approximate 2-designs—is a fair open problem. I would send it to serious peer review and ask for the 2-design evidence and SPAM simulation before acceptance. A desk reject would be a mistake.","headline":"A genuine adaptation of RB to analogue simulators with an honest but unproven 2-design assumption; worth refereeing for the analogue community.","tokens_in":29198,"tokens_out":1719,"would_cite":true,"duration_ms":17917,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Randomized benchmarking can be extended to analogue quantum simulators using disordered Hamiltonian evolutions, yielding a scalable average error rate per unit time that is independent of state-preparation and measurement errors.","keywords":["randomized benchmarking","analogue quantum simulation","unitary 2-design","depolarising channel","SPAM errors","disordered Hamiltonians","noise characterisation","Loschmidt echo"],"falsifier":"Compute the frame potential $F(\\{U_k\\}) = \\frac{1}{K^2}\\sum_{k,k'} |\\mathrm{tr}(U_k^\\dagger U_{k'})|^4$ for the specific sets used in the fits; if it is far from the Haar value 2 over the sequence lengths where the decay is fit, the sets are not approximate 2-designs and the fitted $r$ is not the average error rate. Alternatively, run ARB with a deliberately correlated, non-depolarising noise model on a small system: if the survival probabilities do not follow $P_T=A+Bf^T$ for long $T$, the twirling mechanism is not doing the work.","tokens_in":28215,"feed_emoji":"⚛️","tokens_out":6936,"duration_ms":66034,"temperature":0.7,"pith_summary":"This paper sets out to bring randomized benchmarking, the standard digital tool for measuring average gate error, to analogue quantum simulators. It claims that by replacing gates with native Hamiltonian evolutions $U_k=e^{-iH_k dt}$, adding disorder to generate a large family of unitaries, and inverting each step, one can measure the average error rate per unit time for that family. The metric is read from the decay of the survival probability, $P_T=A+Bf^T$, with SPAM errors absorbed into $A$ and $B$. If the claim holds, analogue simulators can be tested scalably, independently of state-preparation and measurement errors, and across devices running the same family of Hamiltonians.","feed_headline":"Single decay curve benchmarks analogue quantum simulators","feed_subtitle":"Disordered evolutions stand in for random gates, yielding the average error per unit time from a single decay curve.","key_machinery":"The load-bearing mechanism is the unitary 2-design twirl. When a noisy channel is conjugated and averaged over an exact 2-design, a finite set whose averages match the Haar measure for polynomials up to degree two, it becomes a depolarising channel, so the entire noise reduces to one number. ARB uses an $\\epsilon$-approximate 2-design in the diamond-norm sense, samples long sequences of disordered evolutions, and inverts each step; the paper proves $|P^\\alpha_l - P^\\mu_l| \\le l\\epsilon$ and, with no SPAM errors and $l=1$, $r-\\epsilon \\le r' \\le r+\\epsilon$. The candidate 2-designs are generated by adding symmetry-breaking disorder terms to $H_s$, and convergence is diagnosed by how well survival data follow the exponential decay.","core_discovery":"Analogue randomized benchmarking (ARB) claims that an analogue quantum simulator can be characterized by a single average error rate per unit time for a family of Hamiltonian evolutions, with state-preparation and measurement errors absorbed into fit parameters. The protocol replaces digital gates with time-evolution unitaries $U_k = e^{-iH_k dt}$ built from a base Hamiltonian $H_s$ plus disorder terms, and replaces the single inversion gate of standard RB by systematic inversion of each unitary in an echo-style sequence. Under the assumption that the set $\\{U_k\\}$ forms an $\\epsilon$-approximate 2-design, the average survival probability obeys $P_T = A + B f^T$, with $f$ related to the average error rate by $r = (d-1)(1-f)/d$. Simulations on a six-spin XY model with nearest-neighbour and all-to-all couplings show fits to this curve for several noise models, with the locally disordered all-to-all set fitting best; the globally disordered all-to-all set does not fit, which the authors interpret as failure to converge to a 2-design.","pith_inferences":["Because the bound grows linearly with sequence length, the protocol's reliability may be improved by estimating $\\epsilon$ directly, for example through frame-potential or second-moment comparisons, rather than only through fit quality.","If the inversion step can be implemented in Trotterised digital form on hybrid trapped-ion devices, ARB could benchmark the same hardware in both digital and analogue modes, giving a per-time error rather than per-gate error; the paper mentions this direction but does not develop it.","The observed failure signature for global all-to-all disorder suggests a practical protocol-development heuristic: local, site-dependent disorder produces richer scrambling and should be preferred when designing benchmark sets for long-range Hamiltonians."],"forward_implications":["ARB turns a noisy analogue simulator into a single number $r$ per unit time for a family of Hamiltonians, so device quality can be compared without knowing SPAM errors.","Because it uses only native time evolutions, it avoids the compilation overhead that limits digital RB on hardware without native Clifford gates.","The exponential decay form gives a built-in diagnostic: when a disordered set fails to scramble enough to approximate a 2-design, the data visibly depart from $A+Bf^T$, as seen for the globally disordered all-to-all set.","The bound $r-\\epsilon \\le r' \\le r+\\epsilon$ lets experimenters quote the average error rate with an explicit uncertainty coming from how close the set is to a 2-design."],"supporting_citations":[{"why":"Supplies the motion-reversal RB structure of running imperfect unitaries followed by their inverses, which ARB adapts to analogue evolution.","marker":"[13]"},{"why":"Establishes that RB with a distribution close to a 2-design yields meaningful average error rates, the central justification for using approximate designs.","marker":"[22]"},{"why":"Provides theoretical evidence that locally disordered Hamiltonians approximate unitary 2-designs, motivating the disorder construction.","marker":"[29]"},{"why":"Supplies the normal-distributed disorder-potential construction used to generate the unitary sets in ARB.","marker":"[30]"},{"why":"Defines the diamond norm used in the paper's definition of an $\\epsilon$-approximate 2-design and in the error bounds.","marker":"[42]"},{"why":"Defines unitary 2-designs and $\\epsilon$-approximate 2-designs, the property the disordered unitary sets are assumed to approach.","marker":"[43]"}],"fun_headline_variants":["ARB: single decay curve benchmarks analogue quantum simulators","Analogue randomized benchmarking: one error rate per time step","Single average error rate for analogue Hamiltonian families","Randomized benchmarking goes analogue with a single curve","Benchmarking analogue simulators: one curve for all errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole interpretation rests on the disordered set $\\{U_k\\}$ being an $\\epsilon$-approximate 2-design over the sequence lengths used; if the set does not scramble enough, the noise is not depolarised and the fitted decay curve cannot be read as an average error rate, and the paper itself notes that this has not been formally proven.","fun_headline_variants_meta":{"raw":{"variants":["ARB: single decay curve benchmarks analogue quantum simulators","Analogue randomized benchmarking: one error rate per time step","Single average error rate for analogue Hamiltonian families","Randomized benchmarking goes analogue with a single curve","Benchmarking analogue simulators: one curve for all errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1684,"prompt_tokens":1022,"completion_tokens":662,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":585}},"tokens_in":638,"tokens_out":662,"duration_ms":6146,"temperature":1.0,"reasoning_tokens":585,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:21:57.514388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the frame potential $F(\\{U_k\\}) = \\frac{1}{K^2}\\sum_{k,k'} |\\mathrm{tr}(U_k^\\dagger U_{k'})|^4$ for the specific sets used in the fits; if it is far from the Haar value 2 over the sequence lengths where the decay is fit, the sets are not approximate 2-designs and the fitted $r$ is not the average error rate. Alternatively, run ARB with a deliberately correlated, non-depolarising noise model on a small system: if the survival probabilities do not follow $P_T=A+Bf^T$ for long $T$, the twirling mechanism is not doing the work.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that RB with a distribution close to a 2-design yields meaningful average error rates, the central justification for using approximate designs."},{"cited_title":"Approximate Randomized Benchmarking for Finite Groups","cited_arxiv_id":"1803.03621","evidence_quote":"Provides theoretical evidence that locally disordered Hamiltonians approximate unitary 2-designs, motivating the disorder construction."},{"cited_title":"Direct randomized benchmarking for multi-qubit devices","cited_arxiv_id":"1807.07975","evidence_quote":"Supplies the normal-distributed disorder-potential construction used to generate the unitary sets in ARB."}],"review_version":1}