{"id":"a9f73b9f-51cd-4a21-984b-ab881d06b308","arxiv_id":"2505.23622","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new hierarchical segmentation algorithm disentangles concurrent discrete frequency fluctuations in transmons, tracking them at tens of milliseconds resolution and attributing them to charge parity switching and a charge-dipole two-level system.","lead":"Researchers developed a method to track a superconducting qubit's frequency fluctuations in real time, resolving jumps as short as tens of milliseconds over many hours, and automatically separating overlapping fluctuation sources. The approach identifies specific physical causes, such as charge parity switching and two-level system defects, which could improve qubit calibration and error mitigation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HDFA's disentanglement and physical-origin claims lack synthetic ground-truth validation; fΔmax and δng statistics may be segmentation artifacts.","rationale":"The reader's weakest assumption about the Markovian single-frequency model inside each averaging window is legitimate and feeds directly into δf(t). However, the more load-bearing unverified assumption is that HDFA, which is the central new method, correctly recovers a known hierarchy of overlapping fluctuations. The paper provides no synthetic validation of HDFA itself, and the physical-origin conclusions (fΔmax matching ΔCP and |δng| accompanying S2 jumps) are extracted from HDFA outputs in ways that could be biased by segmentation artifacts. The proposed emulator with and without a charge-offset jump would settle this directly. The reader's CONDITIONAL verdict already captures the need for additional evidence, so no verdict change is needed.","tokens_in":39330,"tokens_out":8509,"duration_ms":97319,"concrete_test":"Run a full emulator matching the experimental protocol (same Nτ, τmax, Ns, WG, T1, and switching rates) with known ground-truth δf(t) generated from a known hierarchy: a fast RTN whose amplitude follows |cos(2π ng(t))|, plus a slow RTN in fc with a prescribed simultaneous ng jump, exactly as in the charge-dipole TLS model. Generate binary outcomes from Eq. (3), then run the complete pipeline: Gaussian averaging, Markov fit, HDFA with automatic hyperparameter selection, and the Appendix G2 ng extraction. Compare recovered fΔ distributions, fΔmax, S2 rates, and |δng| against the true values. Repeat with a null model where the slow fc jumps are accompanied by no ng change. If HDFA recovers the ground truth in the first case and gives zero |δng| in the null case, the disentanglement and origin claims are supported; if not, they are not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that HDFA faithfully recovers the hierarchy of fluctuations from the measured δf(t) series. The paper validates the Gaussian averaging and Markovian noise fit with emulations (Appendices B and D), but it never validates HDFA itself on data with known multi-level fluctuation structure. HDFA has two free hyperparameters (λll and Lmin) chosen by elbow/RMSE heuristics, and it is designed to segment any noisy time series into piecewise-constant RTN segments; agreement with experimental data therefore does not by itself demonstrate correctness. The physical attributions depend directly on HDFA outputs: fΔmax in Table I is the maximum observed S1 amplitude, and because the maximum of a noisy estimate is upward biased, the comparison with ΔCP is not a valid quantitative test unless this bias is calibrated. Similarly, the charge-offset jumps |δng| in Appendix G2 are derived from fΔ(t) via Eq. (G2); if HDFA misassigns points near fΔ minima or close to f_c jumps, the apparent ng jump could be an artifact of the segmentation rather than evidence for a charge-dipole TLS. Without a synthetic benchmark with known ground truth, the disentanglement and the CP/TLS origin assignments remain conditional.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript develops and demonstrates a two-stage framework for characterizing time-dependent qubit frequency fluctuations. Stage one uses repeated idle-circuit tomography with Gaussian-weighted few-repetition averaging to estimate time-dependent parameters δf(t), Γ1(t), Γϕ(t) by fitting a Markovian Lindblad model (Eq. (3)). Stage two, HDFA, recursively segments the δf(t) series into a hierarchy of random-telegraph-noise components, each described by a state, center, and magnitude, with the hierarchy relation f^{(n+1)} = f_c^{(n)}; switching rates and magnitudes are then extracted. On three transmon qubits of the IBM Lima device, the authors report fast RTN fluctuations at rates of 3–5 s^{-1} whose maximum amplitudes roughly match independently computed charge dispersion, and slower asymmetric RTN fluctuations attributed to charge-dipole TLSs, with correlated charge-offset jumps |δng| of order 0.001–0.011. The paper claims temporal resolution of tens of milliseconds over hours and positions the framework as a tool for calibration, error mitigation, and physical diagnostics.","tokens_in":39596,"tokens_out":11199,"duration_ms":114517,"significance":"If the physical attributions hold, the framework is a useful addition to qubit-noise characterisation: the minimal-data Gaussian averaging is tested on simulated data (Appendices B and D), the bootstrap uncertainty propagation in Appendix C is carefully specified, and the comparison of fΔmax with independent charge-dispersion calculations in Table I is a falsifiable check rather than a circular fit. The authors also acknowledge several limitations in the text, including HDFA misclassification when fΔ is small and Γϕ artifacts near jumps. However, the central novelty is the hierarchical auto-segmentation, and the manuscript does not yet provide a synthetic benchmark with known multi-level fluctuation structure; the physical attributions are consistency arguments rather than unique identifications. The value of the paper therefore depends on closing that validation gap.","major_comments":[{"comment":"The HDFA disentanglement is not tested on synthetic data with known hierarchical fluctuation structure. Appendix B validates only a single-level RTN with constant f_c and f_Δ, and Appendix D tests the noise-model fitting chain rather than the multi-level segmentation. The hyperparameters λll and Lmin are chosen by elbow and RMSE heuristics on the experimental data (Figs. 10 and 11), and because any noisy time series can be segmented into piecewise-constant RTN pieces, the visual agreement in Fig. 4 does not by itself demonstrate that recovered f_c^{(n)}, f_Δ^{(n)}, s^{(n)}, and rates are the true generators. Please add emulations with known two-level (and, if possible, continuously drifting f_c or f_Δ) hierarchies, and report state-assignment error, rate bias, and bias in the f_Δ distribution as functions of λll and Lmin. This is load-bearing because Table I and Appendix G2 derive physical attributions directly from HDFA outputs.","section":"Secs. II.D and III.D; Appendices B and F"},{"comment":"The Markovian-constant-within-window assumption used to fit Eq. (3) is violated whenever a frequency jump falls inside the Gaussian averaging window, and the resulting artifact is explicitly visible as sharp Γϕ peaks in Fig. 2(c). These jump-adjacent δf(t) estimates are fed into HDFA without exclusion or refitting with the two-frequency model of Ref. [25]. The statement in Sec. II.C that 'for the majority of times' a Markovian model is adequate is not quantified on the experimental data. Please estimate the fraction of time steps whose windows contain a jump (from the fitted s^{(1)}(t) and WG), and show that removing or refitting those points does not materially alter f_c^{(1)}, f_Δ^{(1)}, the level-1 rates, or the level-2 outputs. Without such a check, the bias near jumps is an uncontrolled systematic in the input to the central algorithm.","section":"Sec. II.C and Fig. 2(c)"},{"comment":"Using the maximum observed HDFA amplitude fΔmax as the estimate of the charge dispersion ΔCP is subject to unquantified sampling and noise biases: maximization over noisy segment estimates tends to be upward-biased, while finite observation time may also under-sample ng values near the dispersion maximum. The quoted uncertainties reflect segmentation and fitting noise only, not the bias of the maximum statistic, so the agreement in Table I is not yet a quantitative consistency test. The same fΔmax enters Eq. (G2) to convert fΔ(t) into ng(t), so any bias propagates directly into the charge-offset jumps |δng| of Table II. Please calibrate this using synthetic data with known ΔCP and finite observation length, or replace fΔmax by a robust estimator with known sampling distribution, and propagate that uncertainty through the |δng| extraction.","section":"Table I and Eq. (G2)"},{"comment":"The charge-dipole TLS attribution rests on the claim that the distribution of |δng| across level-2 switching events is peaked away from zero. The statistical significance of this peak is not quantified, and for qubit 0 the counts in Fig. 12 are very small. In addition, the calculation uses ng values obtained from HDFA-segmented fΔ(t) via Eq. (G2), so segmentation errors near fΔ minima or near f_c^{(2)} jumps could produce an apparent δng correlation. Please provide a null distribution from randomized jump positions (or from emulations with known TLS-free noise), state the number of events and significance for both qubits, and either add a direct experimental check of the TLS (e.g., avoided crossings or frequency-tuned spectroscopy) or soften the abstract and Sec. III.E from 'identify the origins' to 'consistent with' a charge-dipole TLS origin.","section":"Sec. III.E and Appendix G2, Fig. 12"}],"minor_comments":[{"comment":"The statement that crosstalk across the three parallel qubits was verified to be negligible is asserted without data, a metric, or a description of the verification; please add this evidence or an explicit caveat.","section":"Sec. III.A"},{"comment":"The manuscript does not provide raw data or code; given the algorithmic nature of HDFA and the hyperparameter choices, a data/code availability statement or repository would materially aid reproducibility.","section":"General"},{"comment":"There are several typographical errors that should be corrected: 'shaded greed region' should be 'shaded green region' in the Fig. 12 caption, Appendix G3 contains an unfinished sentence beginning 'The TLS is Since', and Sec. III.D contains 'is applied to to the results'.","section":"Fig. 12 caption and Appendix G3"},{"comment":"The missed-jump correction in Eq. (F4) assumes that jumps shorter than τmin are completely missed while all longer jumps are detected, but Appendix B shows partial detection and temporal smearing for finite WG; please test this rate correction on simulated hierarchical data or state its assumptions more carefully.","section":"Eq. (F4)"}],"recommendation":"major_revision","confidential_remarks":"The main risk is not circularity but under-validation of the new segmentation tool: the Appendices validate the measurement chain and single-level HDFA, while the strongest claims (disentanglement and TLS attribution) depend on multi-level outputs that have no synthetic ground truth. I would encourage the editor to request the HDFA benchmark and the bias-calibrated comparison with ΔCP before publication. I saw no concerns about citation or novelty disclosure beyond the points in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the HDFA algorithm is a real addition to the qubit-noise toolbox, and the experiment is real work, but the paper's central disentanglement and physical-origin claims are only partially validated. The method itself is not tested on data with known multi-level fluctuation structure, so treat the CP/TLS attributions as conditional.\n\nWhat is genuinely new: HDFA recursively segments a time series into a hierarchy of two-state fluctuations with time-dependent centers and amplitudes. That combination is not in the cited literature. The authors also do several things well. The few-repetition Gaussian averaging with bootstrap uncertainties is tested on emulated data in Appendices B and D, and the emulations reproduce the main experimental features, including the Gamma-phi peaks at frequency jumps. The hyperparameter selection is described, and the uncertainty propagation from segmentation is more careful than most papers in this area. They also flag their own weak points: the Markovian-within-window assumption is violated around jumps, small-amplitude segments can be misclassified, and the crosstalk check is asserted rather than shown.\n\nThe soft spots are real, and the main stress-test concern lands. Every emulation in the paper validates the noise extraction step, not HDFA itself. A synthetic time series with known multiple overlapping RTNs and drifts would establish that the segmentation recovers the hierarchy rather than inventing it. Without that, the fDelta,max comparison to charge dispersion is weakened because a maximum over a noisy estimate is upward biased, and the charge-offset jumps in Appendix G2 inherit whatever artifacts HDFA introduces. The TLS parameter ranges in Table II are consistent with literature, but they come from a minimal model with assumed junction thickness and TLS temperature, so they are plausibility arguments rather than measurements. No raw data or code is provided, which limits reproducibility.\n\nAll that said, this paper deserves a serious referee. The method is novel, the experiments cover hours at tens-of-millisecond resolution on three qubits, and the internal consistency checks are good. The target audience is experimentalists and error-mitigation people; they will get the most value. I would not cite it in my own work until the HDFA synthetic benchmark and data availability are added. If I were the editor, I would send it to review and ask for those additions.","headline":"A genuinely new segmentation method for qubit frequency noise with a credible but incompletely validated experimental demonstration; the CP/TLS attributions are plausible, not proven.","tokens_in":40118,"tokens_out":3177,"would_cite":false,"duration_ms":32409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A measurement-plus-segmentation framework tracks qubit frequency fluctuations at tens-of-millisecond resolution over hours and attributes them to overlapping charge-parity switching and two-level-system defects.","keywords":["qubit noise characterisation","transmon qubits","charge parity switching","two-level systems","random telegraph noise","hierarchical discrete fluctuation auto-segmentation","Gaussian moving average","fast frequency tracking"],"falsifier":"A supervised emulation with known fast and slow random-telegraph ground truth, run at jump durations near the detection limit, would falsify the disentanglement claim if HDFA fails to recover the known rates and magnitudes within the quoted uncertainties.","tokens_in":39113,"feed_emoji":"⚛️","tokens_out":11276,"duration_ms":89101,"temperature":0.7,"pith_summary":"This paper builds a two-stage method for following the fast, overlapping frequency fluctuations that limit superconducting qubits. The measurement stage runs idle-qubit tomography circuits with very few repetitions and averages the binary outcomes with a Gaussian window, so the qubit frequency can be estimated at intervals of a few tens of milliseconds. The analysis stage, called hierarchical discrete fluctuation auto-segmentation (HDFA), automatically decomposes the resulting frequency time series into nested two-state jump processes ordered by timescale. Applied to three transmon qubits over hours, the method resolves fast jumps with switching rates of about 3–5 per second plus slower jumps on timescales of seconds to minutes, and disentangles them from each other. The paper concludes that the fast fluctuations are charge-parity switching and the slower ones are an off-resonant charge-dipole two-level system, based on matching predicted charge dispersions and observed charge-offset jumps.","feed_headline":"Fast qubit noise tracked at 30 ms and traced to two physical sources","feed_subtitle":"Tens-of-millisecond resolution over hours reveals which physics drives qubit frequency noise.","key_machinery":"The load-bearing object is the hierarchical discrete fluctuation auto-segmentation (HDFA) algorithm, built on hidden Markov models and time-series segmentation. It models each fluctuation level as a random telegraph $f^{(n)}(t) = f_c^{(n)}(t) + s^{(n)}(t) f_\\Delta^{(n)}(t)/2$ and then recursively strips off the fastest level: after fitting an HMM with Baum-Welch and decoding states with Viterbi, it segments the series by growing each segment until a mean-log-likelihood threshold flags a change in the centre or magnitude, then feeds the centre $f_c^{(n)}$ into the next level. The measurement side is the few-repetition Gaussian averaging protocol, which weights binary measurement outcomes by $w(i,r) = \\exp[-(i-r)^2/(2 W_G^2)]$ with $W_G = 2$ to $4$, giving tens-of-millisecond resolution while still allowing a Markovian single-frequency noise model to be fitted at each time step. Together these parts turn raw single-shot outcomes into a time-ordered stack of disentangled fluctuation processes with rates, magnitudes, and uncertainties.","core_discovery":"The paper claims that qubit frequency noise, which usually appears as a messy mixture of overlapping stochastic processes, can be resolved into a clean hierarchy of independent random-telegraph fluctuations if one combines few-repetition Gaussian averaging with recursive auto-segmentation. At each hierarchy level the frequency obeys $f^{(n)}(t) = f_c^{(n)}(t) + s^{(n)}(t) f_\\Delta^{(n)}(t)/2$, with the centre of the fast random telegraph becoming the input to the next slower level, $f^{(n+1)} = f_c^{(n)}$. On the transmon data, the fastest level switches at a few times per second with roughly symmetric rates and a fluctuation magnitude whose maximum matches the device's predicted charge dispersion, identifying it as charge-parity switching; the next level switches much more slowly and asymmetrically, and each switch is accompanied by a small measurable jump in the qubit charge offset, matching an off-resonant charge-dipole two-level-system model. The paper presents this as a demonstration that concurrent fluctuations can be disentangled automatically and attributed to physical origins without assuming that one source dominates.","pith_inferences":["A direct hardware test of the charge-parity assignment would be simultaneous parity-sensitive readout on the same qubit: the fastest HDFA state switches should coincide with independently observed parity flips, which would also reveal how often the assignment misses very short parity dwells.","Because the averaging window is optimized for frequency, the same protocol could be re-optimized for relaxation or pure-dephasing fluctuations; the paper's emulations suggest fast relaxation fluctuations are mostly statistical, but a window optimized for that parameter would settle whether true fast relaxation fluctuations exist.","The HDFA hierarchy treats each level as a two-state random telegraph, so continuous drifts are approximated by nesting many small segments; a testable extension is to compare HDFA output against a multi-state or continuous-state model on simulated drifts to quantify the residual approximation error.","The observed correlation between two-level-system frequency shifts and charge-offset jumps suggests that monitoring the charge offset alone could act as a cheap early-warning signal for impending TLS switching, potentially enabling pre-emptive recalibration."],"forward_implications":["Qubit frequency noise can be tracked at tens-of-millisecond resolution for hours, revealing concurrent random-telegraph fluctuations whose rates span more than three orders of magnitude.","The fastest fluctuations were attributed to charge-parity switching because their symmetric rates (about 3–5 s$^{-1}$) and maximum magnitudes match the device charge dispersion; if this attribution is right, charge-parity noise is present and dominant at sub-second timescales on these qubits.","The slower fluctuations were attributed to an off-resonant charge-dipole two-level system because each slow frequency jump coincides with a charge-offset jump, and the extracted two-level-system parameters fall in ranges reported for such defects.","The disentangled fluctuation information can be used to update qubit calibration on the fly, mitigate errors, and inform error-correction scheduling, since the discrete noise state at the time of an algorithm run becomes knowable.","The framework generalizes to other qubit platforms and to fluctuations in parameters other than frequency, provided the noise can be approximated as piecewise-constant within the averaging window."],"supporting_citations":[{"why":"Supplies the idle-qubit noise model (detuning, relaxation, pure dephasing) that this paper simplifies to a Markovian single-frequency fit for each averaging window.","marker":"[25]"},{"why":"Establishes charge-parity switching as a source of transmon frequency random-telegraph noise and provides the charge-parity frequency relation used to compare observed fluctuation magnitudes with predicted charge dispersion.","marker":"[50]"},{"why":"Gives the transmon charge dispersion formula and design parameters used to predict the maximum frequency swing expected from charge-parity switching.","marker":"[52]"},{"why":"Prior observation of concurrent charge-parity and two-level-system fluctuations in a superconducting qubit, which this work generalizes to comparable-magnitude overlapping fluctuations.","marker":"[59]"},{"why":"Documents low-frequency telegraphic noise from single fluctuators (two-level systems) in transmon qubits and supports the assignment of slow fluctuations to two-level systems.","marker":"[42]"},{"why":"Provides the hidden Markov model machinery (Baum-Welch and Viterbi algorithms) that the HDFA segmentation routine runs within each segment.","marker":"[77]"},{"why":"Supplies the time-series segmentation and change-point detection perspective that HDFA adapts for discrete fluctuations.","marker":"[80]"},{"why":"Underpins the Gaussian moving-average weighting used to estimate probabilities from few repetitions while keeping temporal resolution.","marker":"[88]"}],"fun_headline_variants":["Millisecond qubit noise resolved into charge parity and TLS switching","Automatic splitting of qubit noise reveals two physical culprits","Qubit noise disentangled: overlapping jumps and drifts separated at last","Hierarchical segmentation catches qubit frequency flips at 30 ms","Fast qubit noise tracking at 30 ms identifies charge parity and TLS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that at most time points the qubit's noise parameters are constant inside the short Gaussian averaging window, so each fitted frequency value is the true instantaneous frequency; a frequency jump inside the window breaks this assumption and shows up as a pure-dephasing spike that the analysis excludes rather than models.","fun_headline_variants_meta":{"raw":{"variants":["Millisecond qubit noise resolved into charge parity and TLS switching","Automatic splitting of qubit noise reveals two physical culprits","Qubit noise disentangled: overlapping jumps and drifts separated at last","Hierarchical segmentation catches qubit frequency flips at 30 ms","Fast qubit noise tracking at 30 ms identifies charge parity and TLS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000264,"raw_usage":{"total_tokens":1605,"prompt_tokens":949,"completion_tokens":656,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":564}},"tokens_in":565,"tokens_out":656,"duration_ms":7469,"temperature":1.0,"reasoning_tokens":564,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:41:51.416022+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A supervised emulation with known fast and slow random-telegraph ground truth, run at jump durations near the detection limit, would falsify the disentanglement claim if HDFA fails to recover the known rates and magnitudes within the quoted uncertainties.","supporting_citations":[{"cited_title":"Truong, L","cited_arxiv_id":null,"evidence_quote":"Underpins the Gaussian moving-average weighting used to estimate probabilities from few repetitions while keeping temporal resolution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes charge-parity switching as a source of transmon frequency random-telegraph noise and provides the charge-parity frequency relation used to compare observed fluctuation magnitudes with predicted charge dispersion."},{"cited_title":"Measuring NISQ Gate-Based Qubit Stability Using a 1+1 Field Theory and Cycle Benchmarking","cited_arxiv_id":"2201.02899","evidence_quote":"Documents low-frequency telegraphic noise from single fluctuators (two-level systems) in transmon qubits and supports the assignment of slow fluctuations to two-level systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the hidden Markov model machinery (Baum-Welch and Viterbi algorithms) that the HDFA segmentation routine runs within each segment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the time-series segmentation and change-point detection perspective that HDFA adapts for discrete fluctuations."}],"review_version":1}