{"id":"1b1c043a-5cd6-426b-b84a-482b6b3178c1","arxiv_id":"2607.14441","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Scheduling-induced timing jitter makes dynamical decoupling pulses counterproductive above about 2 Hz in a trapped-ion QCCD; a real-time pulse-tracking protocol restores most of the benefit.","lead":"This paper studies how timing jitter in trapped-ion quantum computers degrades a standard error-fighting technique called dynamical decoupling, and proposes a smarter way to schedule the correction pulses. The authors show that too many correction pulses can make memory worse, and that tracking the jitter lets you cancel most of the damage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The predicted 2 Hz DD optimum rests on untested i.i.d. scheduling errors: correlated pulse-timing errors would change the 4Nσ² DC shelf in Eq. (33), potentially eliminating the central claim. Fig. 2 measures only single-pulse offsets.","rationale":"The i.i.d. scheduling-error model is the linchpin of the paper's headline quantitative claim. The analytic derivations are internally consistent, and the Monte Carlo and experimental results are consistent with the model, but the model's crucial assumption—independence across pulses—is supported only by a single-pulse histogram. The derivation in Sec. III.D, Eq. (33), directly shows that the linear-in-N DC shelf is a sum over pulse pairs weighted by covariances; any correlation structure changes the scaling. Positive correlations, which are plausible in a QCCD scheduler (a delayed pulse can shift subsequent transport and gating), would reduce the shelf and could make N=4 or higher preferable to N=2. This is an empirical question, not a mathematical flaw. I considered the apparent impossibility of implementing the 'compile-time' robust DD of Sec. V.A (since it requires knowing ε_{j-1} at runtime), but that concern affects a secondary contribution and not the central 2 Hz claim. The reader's weakest_assumption identifies the same issue, and I agree. The verdict should remain CONDITIONAL: the paper is promising but should not be relied upon for the quantitative pulse-rate recommendation until joint pulse-timing statistics are measured on H2-1 and the infidelity optimum is recomputed with the observed covariance. No code or data release further hinders independent verification, but that is an addressable issue, not a fatal one.","tokens_in":23453,"tokens_out":12721,"duration_ms":115711,"concrete_test":"On Quantinuum H2-1, run a compile-time CPMG sequence (e.g., N=4 and N=8, T=1 s) recording the actual timestamps of all N pulses across many repetitions. Compute the empirical covariance matrix C_{jk}=cov(ε_j,ε_k) and the effective DC shelf Δ_N = 4Σ_{j,k=1}^N (-1)^{j+k} C_{jk}. Compare Δ_N with 4Nσ². If Δ_N/(4Nσ²) is significantly below 1, re-simulate the infidelity of Fig. 5 using the measured covariance instead of the i.i.d. model; if the optimal N shifts above 2, the central claim fails. A simpler preliminary check: compute adjacent-pulse correlation ρ=C_{12}/σ²; |ρ|>0.2 would already cast doubt on the i.i.d. assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that scheduling errors create a low-frequency shelf growing linearly with pulse count, making N>2 counterproductive—derives from the i.i.d. assumption on pulse-timing errors in Sec. III.A. Equation (33) gives the DC filter-function component as F_ideal(0,T)+...+4Nσ². Equivalently, expanding F(ω,T) to lowest order in ω, the DC shelf is 4·var(Σ_j (-1)^j ε_j); with i.i.d. errors this becomes 4Nσ². If errors are correlated, the shelf is instead 4Σ_{j,k} (-1)^{j+k} cov(ε_j, ε_k). Positive correlations between pulses of opposite sign (e.g., adjacent CPMG pulses) reduce this alternating sum: for perfect positive correlation in N=2, the shelf vanishes entirely, eliminating the predicted optimum. The only empirical support, Fig. 2, is a histogram of the first-pulse offset measured on a 28-qubit H2 machine—not a joint distribution of multiple pulses on the H2-1 hardware used in the experiments. Thus the scaling law and the quantitative 2 Hz optimum are not yet validated for the actual system; a multi-pulse measurement could either confirm the i.i.d. model or show that correlations substantially weaken the central conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes how scheduling-induced pulse-timing errors degrade dynamical decoupling (DD) in trapped-ion QCCD quantum computers. Using the filter-function formalism and a cumulant expansion, it models scheduling errors as independent Gaussian timing jitter, derives the scheduling-error-averaged filter function, and shows that the DC component acquires a term 4Nσ² that grows with the number of pulses. For a 1/f² memory-error spectrum, this predicts an optimum at a low DD pulse rate (about 2 Hz for the parameters used), so that adding more pulses becomes counterproductive. The authors also introduce real-time DD (RTDD) and a scheduling-error-robust protocol that uses the measured displacement of earlier pulses to adjust later pulse timings. Analytic predictions are compared with Monte Carlo simulations, and compile-time CPMG sequences are tested on Quantinuum H2-1 in Ramsey-type experiments.","tokens_in":23773,"tokens_out":5978,"duration_ms":60497,"significance":"If the central claims hold, the paper provides a practically important and somewhat counterintuitive design rule for trapped-ion QCCD memories: for low-frequency dephasing noise, scheduling errors can make a small number of DD pulses preferable to a larger number. The analytic averaged-filter-function expressions are a useful contribution, and I verified that the main derivations are internally consistent: Eq. (32) follows from the stated model, Eq. (33) gives the 4Nσ² DC shelf for ideal sequences, and the Monte Carlo results in Figs. 3, 4, and 6 match the analytic formulas. The proposed scheduling-error-robust protocol is elegant and, if implementable, would remove the linear-N penalty. The experimental data are from a real 56-qubit QCCD and include a useful null check of XY4 versus CPMG, supporting the longitudinal-noise assumption. The principal weaknesses are that the key i.i.d. assumption on scheduling errors is not directly validated, the robust protocol is not demonstrated experimentally, and the experimental results do not actually show that larger N is counterproductive, only that N=4 does not improve over N=2.","major_comments":[{"comment":"The central prediction that the DC filter-function shelf grows as 4Nσ² relies on the assumption that scheduling errors are i.i.d. Gaussian. For an ideal DD sequence, the DC shelf is more generally 4 Σ_{j,k} (-1)^{j+k} cov(ε_j, ε_k); it reduces to 4Nσ² only when cov(ε_j, ε_k)=σ²δ_{j,k}. Positive correlations between pulses of opposite sign can substantially reduce or even cancel this term. Fig. 2 provides evidence only for the marginal distribution of the first pulse, measured on the 28-qubit H2, not on the 56-qubit H2-1 used in the experiments, and it does not test independence or correlations across pulses in a sequence. Because the optimal-N conclusion and the benefit of the robust protocol both depend on this structure, please provide multi-pulse joint timing measurements on the relevant hardware, or alternatively present a sensitivity analysis showing that the qualitative conclusions","section":"Sec. III.A, Eq. (33)"},{"comment":"The 'robust compile-time DD' protocol defines target times u_j = t_j + ε_{j-1} for j>1, where ε_{j-1} is the realized scheduling error of the previous pulse. At compile time, ε_{j-1} is not known; it can only be known after the previous pulse has actually been scheduled/executed. If the intended implementation is real-time feedback, the text should say so explicitly and discuss the feedback latency and how it is incorporated into the error budget. As written, the phrase 'implemented at compile time' in the opening of Sec. V is inconsistent with the equations. This is not merely a wording issue: it determines whether the proposed cancellation is actually realizable for the claimed compile-time scenario.","section":"Sec. V.A, Eq. (44)"},{"comment":"The experimental data do not support the abstract's statement that 'increasing the DD pulse frequency beyond 2 Hz is generally counterproductive.' The measurements compare N=0, 2, and 4 and find that N=4 does not significantly improve over N=2. This is consistent with the model, but it is also consistent with saturation or with a very weak dependence; it does not show that N=4 is worse. To substantiate the counterproductivity claim, the authors need either data at larger N (e.g., N=6 or 8) showing a statistically significant decrease in survival probability, or a softened claim such as 'no further improvement' rather than 'counterproductive.'","section":"Sec. VI, Fig. 8"},{"comment":"The abstract says 'We demonstrate our methods on Quantinuum H2-1,' but the experimental section only demonstrates compile-time CPMG (N=0,2,4). The RTDD and scheduling-error-robust protocols are not tested on hardware; their support is purely analytic and numerical. Either clarify that only the compile-time standard method was experimentally demonstrated, or add experimental data for RTDD/robust DD. As it stands, the claim of experimental demonstration exceeds what Sec. VI reports.","section":"Sec. VI, Abstract"}],"minor_comments":[{"comment":"There are several OCR/typographical issues: 'IMP ACT OF SCHEDULING ERRORS' in the section header, 'to to' in Sec. III.C, and garbled characters in Fig. 8 (e.g., '□0:2'). These should be cleaned before publication.","section":"Throughout"},{"comment":"The caption should state explicitly that only the first-pulse offset ε₁ is histogrammed and that no multi-pulse joint statistics are shown. This is important because the i.i.d. assumption of Sec. III.A is a modeling input rather than an empirical finding.","section":"Fig. 2 caption"},{"comment":"The notation ⟨⟨·⟩_mem⟩_sch and the subsequent approximation signs are a little compressed. It would help readers to spell out that C^(1)(T) and C^(2)(T) are functions of the random pulse times and that the second line is a cumulant expansion in those random variables.","section":"Eq. (16)"},{"comment":"Please clarify the noise PSD convention: whether S(ω) is one-sided or two-sided and how the 'Detuning noise (Hz²/Hz)' vertical axis relates to the S(2πf) used in Eq. (15). This will prevent unit confusion in the overlap integral.","section":"Fig. 4"},{"comment":"The estimation of the cumulants via four readout phases is described briefly. It would be helpful to state explicitly that the four phases are measured on the same physical qubits in separate experimental runs and how statistical uncertainties are propagated to the error bars in Fig. 8.","section":"Sec. VI, Eq. (57)-(60)"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and mostly self-consistent paper, with correct analytic derivations and a genuinely useful practical message. The main risk is that the i.i.d. scheduling-error model is load-bearing and currently rests on a single-pulse histogram from a different machine; a correlated-error distribution could change the scaling law and the quantitative optimum. The authors should be asked to either measure joint pulse-timing statistics or provide a robustness analysis. I would also insist that the experimental claims be aligned with the data: the abstract currently overstates both the demonstration of the robust/RTDD protocols and the evidence for 'counterproductive' higher pulse numbers. These are fixable with additional measurements and/or careful rewording, so I do not recommend rejection, but the current version is not yet suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The paper's main result — random scheduling jitter puts a low-frequency shelf in the DD filter function that grows with pulse count — is real. Eq. (33)'s DC term 4Nσ² follows from the model, and the Monte Carlo curves match the analytic expressions. The scheduling-error-robust protocol (shift the next pulse by the previous pulse's measured displacement) is genuinely neat and gives dephasing independent of pulse number; that's a practical contribution for QCCD compilers.\n\nThe experiment on H2-1 is well done as far as it goes: Ramsey with N=0,2,4 CPMG, N=4 not beating N=2, and XY4 matching CPMG, consistent with the central claim. Credit where due: the paper does not oversell the parameter-free part; the analytic formulas are derived rather than fitted.\n\nSoft spots, in order. First, the abstract says the methods are 'demonstrated on H2-1,' but RTDD and robust RTDD are only simulated; the experiments cover compile-time CPMG/XY4 only. That overclaim needs to be fixed before publication. Second, and more load-bearing: the 4Nσ² scaling and the predicted 2 Hz optimum depend on i.i.d. Gaussian pulse-timing errors. The only empirical support is Fig. 2, which is a single-pulse offset histogram measured on the 28-qubit H2, not on H2-1, and says nothing about correlations across pulses. If scheduling errors are positively correlated across adjacent CPMG pulses, the alternating sum can shrink or vanish, and the quantitative optimum shifts. The stress-test note is right about this. It is a genuine limitation, not a manufactured one. Third, no code or data are released, so the figures can't be independently reproduced.\n\nNone of this kills the paper. The robust protocol remains valuable even if correlations are present, because its cancellation is per-realization and does not rely on i.i.d. structure. And the qualitative prediction — more DD pulses eventually hurt in low-frequency-dominated memories — is very likely right. But the specific '2 Hz is optimal' number should be presented as system-dependent until a multi-pulse joint-timing characterization is done on H2-1.\n\nWho this is for: anyone working on DD or error mitigation for trapped-ion QCCDs. It deserves a serious referee. I'd send it to review with a request to revise the abstract, temper the quantitative claim, and ideally add a joint-distribution measurement.","headline":"Worth engaging: the analytic filter-function result for timing jitter is solid and the robust protocol is useful, but the quantitative 2 Hz optimum rests on an untested i.i.d. scheduling-error assumption.","tokens_in":24329,"tokens_out":3852,"would_cite":true,"duration_ms":39892,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scheduling-induced timing jitter degrades dynamical decoupling by adding a low-frequency noise shelf that grows with each added pulse; on a trapped-ion QCCD the optimum is a modest ~2 Hz pulse rate, and a real-time protocol can cancel the a","keywords":["dynamical decoupling","scheduling error","trapped-ion QCCD","filter function","memory error","timing jitter","dephasing","pulse-sequence optimization"],"falsifier":"Measure the timing of every pulse in a multi-pulse CPMG sequence on the same 56-qubit machine used for the experiments and compute the joint distribution of the errors; significant correlation between adjacent pulses would break the 4Nσ² scaling. A simpler experiment: Ramsey memory test at T=1 s with CPMG N=8; the paper's claim predicts N=8 is worse than N=2, so a result where higher N improves survival would falsify it.","tokens_in":23287,"feed_emoji":"⚛️","tokens_out":6238,"duration_ms":57452,"temperature":0.7,"pith_summary":"The paper claims that the timing jitter from transporting trapped ions and sharing gate zones makes each dynamical-decoupling (DD) pulse add a small amount of dephasing, and this effect accumulates with the number of pulses. In the filter-function picture, each pulse contributes a positive term 4σ² to the zero-frequency response, so a sequence with more pulses becomes more sensitive to the very low-frequency magnetic-field noise that DD is meant to cancel. As a result, increasing the pulse rate eventually makes memory errors worse, and for typical trapped-ion memory noise the best average rate is about 2 Hz. The paper introduces a scheduling-error-robust DD protocol that updates each pulse's target time based on the displacement of the previous pulse, cancelling the accumulation so the dephasing rate no longer grows with pulse count. Ramsey experiments on a 56-qubit trapped-ion QCCD show that N=4 pulses do not outperform N=2, consistent with the model.","feed_headline":"More decoupling pulses can worsen trapped-ion memory","feed_subtitle":"Timing jitter adds a noise shelf that grows with pulse count, so fewer pulses win.","key_machinery":"The filter-function formalism, which reduces the performance of a DD sequence to a frequency-domain overlap between the noise power spectral density and the control filter function F(ω,T). The paper augments this with a second cumulant expansion that averages F over the Gaussian scheduling-error distribution, giving the DC component ⟨F(0,T)⟩_sch = F_ideal(0,T)+...+4Nσ². The companion mechanism is the scheduling-error-robust update rule, in which the target time of pulse j is shifted by the measured displacement of pulse j−1, so that the only remaining timing error is that of the final pulse. This turns the N-scaling of the dephasing rate from 4Nσ²m² into a constant 4σ²m².","core_discovery":"The central result is an analytic expression for the filter function averaged over scheduling errors: even for an ideal DD sequence, the DC (zero-frequency) component contains an additive 4Nσ² term, where σ is the standard deviation of the timing jitter and N is the pulse count. Because memory-error noise in trapped-ion QCCDs is concentrated at low frequencies, the overlap between this noise and the filter function is dominated by the DC shelf, so the dephasing rate grows linearly with N. For the measured jitter σ≈5 ms and a 1/f² noise spectrum, the numerical optimum is N=2 for a 1 s idle time (about 2 Hz), and experiments on a 56-qubit device confirm that increasing to N=4 gives no statisti","pith_inferences":["The 4Nσ² scaling is derived from an i.i.d. Gaussian model; measuring the joint distribution of timing errors across a real pulse sequence would reveal whether the next-order corrections (from correlations or non-Gaussian tails) matter in practice.","The same error-compensation rule—shift the next pulse by the previous pulse's realized error—could be applied to any control sequence with measurable timing jitter, including other qubit platforms, as long as displacements can be tracked in real time.","The model suggests a hardware guideline: reducing scheduling jitter σ has a direct, linear payoff in allowing higher DD pulse rates, which sets a concrete engineering target for transport and compilation latency.","For noise spectra that are not strictly 1/f², one can use the same machinery to locate the optimum: the best N is where the added 4Nσ² DC shelf exactly balances the suppression of the remaining spectrum; this could be tested on hardware with different magnetic-field shielding."],"forward_implications":["Compile-time DD schedules on QCCD machines should use a modest pulse rate (≈2 Hz for current hardware) rather than the maximum possible, because extra pulses add more low-frequency sensitivity than they remove.","Real-time DD has a similar tradeoff: the threshold time Δt must balance remainder-time memory error (growing as Δt²) against scheduling error (growing as 1/Δt); the measured optimum lies near 0.1 s.","Scheduling-error-robust DD removes the N-dependence of the scheduling-error contribution, so more pulses can be added without the low-frequency penalty, making higher-order sequences viable if high-frequency noise matters.","The Ramsey experiments confirm that the noise is dominantly longitudinal dephasing, since CPMG and XY4 give the same survival probabilities.","The analytic filter-function model provides a design rule: given the jitter σ and the noise PSD, one can compute the optimal pulse number and threshold for any idle time."],"fun_headline_variants":["Timing jitter makes extra DD pulses backfire in ion traps","For trapped-ion memory, fewer decoupling pulses can be better","Scheduling jitter flips DD benefit: optimal pulse count is low","Ion-trap memory: extra decoupling pulses add noise, not benefit","Trapped-ion QCCD: pulse jitter turns DD gain into loss"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that the timing error of every pulse is an independent, identically distributed Gaussian variable with the same variance, and that these errors are uncorrelated with the magnetic-field noise; this is supported only by a single-pulse histogram from a 28-qubit device, so correlated or state-dependent timing errors would change the 4Nσ² scaling and the robust protocol's benefit.","fun_headline_variants_meta":{"raw":{"variants":["Timing jitter makes extra DD pulses backfire in ion traps","For trapped-ion memory, fewer decoupling pulses can be better","Scheduling jitter flips DD benefit: optimal pulse count is low","Ion-trap memory: extra decoupling pulses add noise, not benefit","Trapped-ion QCCD: pulse jitter turns DD gain into loss"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1017,"prompt_tokens":679,"completion_tokens":338,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":423,"completion_tokens_details":{"reasoning_tokens":243}},"tokens_in":423,"tokens_out":338,"duration_ms":3906,"temperature":1.0,"reasoning_tokens":243,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:04:44.299711+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the timing of every pulse in a multi-pulse CPMG sequence on the same 56-qubit machine used for the experiments and compute the joint distribution of the errors; significant correlation between adjacent pulses would break the 4Nσ² scaling. A simpler experiment: Ramsey memory test at T=1 s with CPMG N=8; the paper's claim predicts N=8 is worse than N=2, so a result where higher N improves survival would falsify it.","supporting_citations":[],"review_version":1}