{"id":"5d91c8ac-e088-40b2-af09-5ace3faed9ac","arxiv_id":"2506.15420","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A dual-rail logical qubit built from the two modes of a fixed-frequency multimode transmon shows error-detected bit-flip and phase-flip lifetimes 48x and 11x longer than the underlying physical modes.","lead":"Researchers encoded a quantum bit in the two microwave modes of a single fixed-frequency superconducting device and showed that the most common type of error can be turned into a detectable warning signal. If it works at scale, this erasure approach could lower the hardware overhead needed for quantum error correction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unquantified |00>-to-logical readout confusion can account for the apparent 48x/11x coherence improvements; a confusion-matrix check is required.","rationale":"The reader's CONDITIONAL verdict is already driven by the absence of readout characterization, and the present stress-test sharpens that concern into a quantitative failure mode. The central claim is that error-detected logical bit-flip and phase-flip rates are more than an order of magnitude below the physical rates. That claim depends on the fidelity with which the GMM classifier separates |00> from |01> and |10>, because the erasure population grows monotonically with delay and any |00>-to-logical misclassification is inherited by the postselected logical population in Eq. (3). A confusion probability of only ~2% between |00> and the wrong logical state is sufficient to generate the reported 48x bit-flip improvement from the erasure rate alone, and ~9% confusion is sufficient for the 11x phase-flip improvement. Since the paper does not report a confusion matrix, assignment fidelities, or independent test-set statistics for the GMM, the observed ratios are not yet distinguishable from readout contamination. The concrete test is a single held-out confusion-matrix analysis with the contamination subtracted; it directly settles whether the reported coherence ratios survive. This does not change the reader's CONDITIONAL verdict, but it does make the required condition more specific and more clearly necessary.","tokens_in":16981,"tokens_out":12893,"duration_ms":148841,"concrete_test":"Re-analyze the raw single-shot IQ data (and ideally a re-run of the 50-hour sequence) to build a held-out confusion matrix for the EOL GMM classifier after preparing |00>, |01>, |10> (and, if possible, |11>, |20>, |02>). Estimate ε_00→01 and ε_00→10. Then subtract the contamination term ε_i [1-exp(-t/T_erasure)] from the postselected logical populations in Fig. 2(d,e) before re-fitting the 30 µs slopes. If the readout-corrected median T_L1 and T_L2E no longer exceed ~10x the physical medians, the central claim fails; if they remain above 10x with ε < 1%, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing point is not simply that readout fidelities are missing; it is quantitative. The postselection in Eq. (3) renormalizes over the states classified as |01> and |10>. If a fraction ε of the true |00> erasure population is classified as one of the logical states, those shots enter the bit-flip numerator. Since the erasure population grows as 1-exp(-Γ_erasure t) ≈ Γ_erasure t on the 30 µs fit window, this adds an apparent logical bit-flip rate δΓ_L ≈ ε Γ_erasure. With physical Γ_erasure ≈ 1/(60–80 µs) ≈ 13–17 kHz and the reported median T_L1 = 3.12 ms (Γ_L1 ≈ 0.32 kHz), any ε ≳ 2×10^-2 produces the whole '48x improvement' even if the true logical bit-flip rate is zero. For T_L2E = 0.76 ms, the threshold is ε ≳ 0.09. The paper reports no confusion matrix for the GMM classifier; Fig. 1(c) shows partially overlapping IQ distributions, so ε of order a few percent is not ruled out. Until this contamination is bounded, the headline coherence ratios are not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a dual-rail logical qubit encoded in the single-excitation subspace of a fixed-frequency two-mode multimode transmon ('dimon'), with logical states |01> and |10> and an end-of-line readout that classifies the |00> state as an erasure. By postselecting on non-erasure shots, the authors extract effective logical bit-flip and phase-flip rates and report improvements over the physical mode coherence times: median T_L1 = 3.12 ms (48x), T_L2E = 0.76 ms (11x), and T_L2R = 0.07 ms (4x), based on roughly 50-hour interleaved measurements on three devices. Section II describes the device, Hamiltonian, and three-state GMM readout; Section III defines the coherence metrics and postselection; Section IV presents noise analysis via Allan deviation and spectral density, attributing the observed logical coherence improvement to suppression of common-mode dephasing sources while retaining sensitivity to differential noise.","tokens_in":17278,"tokens_out":8441,"duration_ms":87605,"significance":"If the reported rates are correct, this is a significant step for erasure-conversion hardware: a flux-free, fixed-frequency, three-island transmon provides all-microwave dual-rail encoding with detectable amplitude-damping errors in a compact coaxial architecture, with repeatability across three devices. The long time series, interleaved measurements, and bootstrap bounds are notable strengths, as is the explicit stability analysis over 50 hours. However, the headline improvements depend on the three-state GMM readout being nearly confusion-free and on the definition of the short-time logical error rates; without a confusion matrix and a clearer fitting model, the central claim is not yet fully established. The work is of high interest to the superconducting-qubit and quantum-error-correction communities, but requires additional readout and fit characterization.","major_comments":[{"comment":"The postselection in Eq. (3) is valid only if the GMM classifier separates |00> from the logical states with negligible error, but the manuscript reports no confusion matrix or readout fidelity for the three-state classifier. Because the erasure population grows approximately as Gamma_erasure * t with Gamma_erasure ~ 13-17 kHz, a fraction epsilon of |00> shots misclassified as |01> or |10> contributes an apparent bit-flip rate of order epsilon * Gamma_erasure. For the reported median T_L1 = 3.12 ms (Gamma_L1 ~ 0.32 kHz), epsilon of a few percent is sufficient to produce the claimed 48x bit-flip improvement even if the true logical bit-flip rate were zero, and the overlapping distributions in Fig. 1(c) make such confusion non-negligible. The authors should report the full confusion matrix for the EOL readout, including |00>->|01>, |00>->|10>, and |01><->|10> misclassification probabilities, and show that the extracted logical rates are stable under an upper bound on these probabilities.","section":"Section III, Eq. (3) and Fig. 1(c)"},{"comment":"The logical T_L1 and T_L2E values are obtained by fitting a linear slope with constant offset to data over the first 30 microseconds, whereas the physical coherence times are extracted from exponential fits. The manuscript does not report fit residuals, uncertainties, or a model justifying the linear approximation over this window; if the true short-time logical bit-flip probability is quadratic in time (as expected when a bit flip requires relaxation followed by re-excitation), the fitted slope depends on the chosen window length. This makes the reported 'logical error rate' not directly comparable to the physical decay rates. Please provide the fits with residuals, a stability test of the extracted rates under varying the fit window (e.g., 10, 20, and 40 microseconds), and, if possible, fits to the expected short-time functional form.","section":"Section III and Fig. 2(d)-(f)"},{"comment":"The statement that there is no direct single-photon decay channel from |01> to |10> due to the different symmetries of the modes is load-bearing for the logical bit-flip model, since any direct |01><->|10> matrix element would contribute a first-order bit-flip rate. The manuscript gives no derivation or reference for the vanishing of this matrix element. Please provide a quantitative symmetry argument (or a citation) showing that the relaxation coupling operator has zero matrix element between |01> and |10>, or bound the residual direct transfer rate experimentally, for example by preparing |01> and measuring the |10> population in the absence of re-excitation.","section":"Section II, 'no direct single photon decay channel'"}],"minor_comments":[{"comment":"There is a typographical error in the denominators of Eq. (3): 'P (|01))' should read 'P (|01>).'","section":"Eq. (3)"},{"comment":"The sentence 'The erasure rate is a important metric' should read 'an important metric.'","section":"Section III, paragraph on erasure rate"},{"comment":"The caption says 'Y-axis in (d) showing difference'; this should be 'shows the difference' for grammatical completeness.","section":"Fig. 2 caption"},{"comment":"Equation (7) is introduced without derivation or a reference; please provide a derivation sketch or citation so that the dephasing ratio can be checked.","section":"Section IV, Eq. (7)"},{"comment":"The units of the noise amplitudes A and B should be stated explicitly when the fit model sigma_delta_f(tau) = (B/2)^{1/2} tau^{-1/2} + (2 ln 2 A)^{1/2} is introduced.","section":"Section IV, text near Fig. 4(i)"}],"recommendation":"major_revision","confidential_remarks":"The key technical issue is the missing confusion matrix for the three-state GMM readout. The quantitative argument in the report shows that a few percent |00>->logical misclassification could account for the claimed bit-flip improvement, so the authors must supply this characterization or a reanalysis with bounded confusion. The 50-hour multi-device dataset and the noise-analysis framework are valuable and the paper is well within the scope of the journal; with the requested additions it could become acceptable. I do not see a fundamental flaw requiring rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a real experimental step: dual-rail encoding in a fixed-frequency multimode transmon (dimon), all-microwave control, three devices, and 50 hours of data. The physical mode characterization is careful, the time series are extensive, and the noise analysis (Allan deviation, spectral density, TLS-like jumps) is thoughtful. The three-device reproducibility and bootstrap bounds are genuine strengths. The paper is also honest about what it has not identified, e.g., the precise noise source limiting the logical Ramsey coherence.\n\nWhat is genuinely new is applying dual-rail erasure conversion to this specific device class in a coaxial architecture, with mode detunings (0.76–1.05 GHz) larger than prior tunable-transmon dual-rail work. That is a useful hardware-oriented contribution, not a conceptual breakthrough.\n\nThe load-bearing soft spot is readout. The postselection in Eq. (3) renormalizes over shots classified as |01> or |10>. If even a couple percent of true |00> erasure shots are misclassified as one of the logical states, those shots add an apparent logical bit-flip rate of roughly ε·Γ_erasure. With Γ_erasure ≈ 13–17 kHz and the claimed Γ_L1 ≈ 0.32 kHz, ε ≈ 2% is enough to produce the entire 48x improvement even if the true logical bit-flip rate is zero. The paper reports no numeric readout assignment fidelity or confusion matrix; Fig. 1(c) shows overlapping IQ distributions and Fig. 1(d) is described only qualitatively. This is not a minor omission—it bears directly on the central quantitative claim. A referee needs to see the confusion matrix for the GMM classifier, ideally for the same integration time and drive settings used in the coherence measurements.\n\nOther soft spots are secondary: the 30 µs linear fit is short but defensible for short circuit depths, Eq. (7) is stated without derivation, the Q3 Ramsey data are absent, and raw data/code are not provided. None of these individually sinks the paper, but together they lower confidence that the reported ratios are robust.\n\nThe paper deserves a serious referee. The device work is real and the erasure-conversion angle is timely. My recommendation: send it to peer review with a request for major revision, and make the readout-fidelity quantification a condition for acceptance. If the confusion matrix shows ε well below the 2% threshold for all three devices, the central claim will stand; if not, the headline numbers need to be revised.","headline":"A plausible and well-executed fixed-frequency dual-rail erasure demo, but missing readout-confusion numbers leave the headline 48x/11x coherence improvements unproven.","tokens_in":17787,"tokens_out":3543,"would_cite":false,"duration_ms":36308,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fixed-frequency multimode superconducting qubit encodes a dual-rail logical qubit whose detected bit-flip and phase-flip rates are 48x and 11x below the physical modes.","keywords":["dual-rail encoding","erasure conversion","multimode transmon","dimon","superconducting qubit","amplitude damping","error detection","logical coherence"],"falsifier":"Retrain the GMM classifier on a different readout drive frequency or with a stricter $|00\\rangle$ rejection threshold and remeasure the logical bit-flip decay: if $T_1^L$ shifts, readout misclassification is biasing the reported suppression. Independently, search spectroscopically for any transition between $|01\\rangle$ and $|10\\rangle$; observing direct coupling between the logical states would break the symmetry argument and invalidate the erasure interpretation.","tokens_in":16785,"feed_emoji":"⚛️","tokens_out":11034,"duration_ms":105869,"temperature":0.7,"pith_summary":"The paper sets out to establish that a single fixed-frequency multimode transmon—a three-island, two-junction device supporting two transmonlike microwave modes—can host a dual-rail logical qubit in which amplitude damping, normally the dominant error, becomes a detectable erasure. It encodes $|0\\rangle_L=|10\\rangle$ and $|1\\rangle_L=|01\\rangle$, so relaxation of either logical state leaks into $|00\\rangle$; a three-state Gaussian-mixture readout flags that leakage and postselection (Eq. 3) removes it. On three devices measured over more than 50 hours, the error-detected logical coherence times are $T_1^L=3.12$ ms and $T_{2E}^L=0.76$ ms median, corresponding to 48x and 11x reductions in bit-flip and phase-flip rates relative to the physical modes. A sympathetic reader would care because erasure conversion is one of the more hardware-efficient routes to fault tolerance, and this implementation requires no flux tuning, no galvanic coupling, and no larger footprint than a conventional coaxial transmon.","feed_headline":"Dual-rail encoding in one transmon cuts bit-flip rates 48x","feed_subtitle":"End-of-line error detection converts damping into detectable leakage and holds stable across three devices for 50 hours.","key_machinery":"The central mechanism is a dual-rail encoding inside a dimon, a three-island, two-junction superconducting device with two transmonlike modes, $D$ and $Q$. The logical states $|0\\rangle_L = |10\\rangle$ and $|1\\rangle_L = |01\\rangle$ have different electric-field symmetries, so there is no direct single-photon transition between them; amplitude damping therefore exits the logical subspace to the ground state $|00\\rangle$. A readout tone placed between the two dispersive resonances produces three separable Gaussian clouds, and a Gaussian mixture model assigns each shot to $|00\\rangle$, $|01\\rangle$, or $|10\\rangle$. Eq. (3) renormalizes the surviving logical populations, converting a raw decay measurement into an error-detected coherence measurement. This combination—symmetry-forbidden logical leakage, three-state readout, and postselection—is what carries the reported $T_1^L$ and $T_{2E}^L$ improvements.","core_discovery":"The central discovery is that the two single-excitation states of the dimon form a dual-rail code with no direct single-photon decay channel between them, so energy relaxation from either logical state lands in $|00\\rangle$ and is heralded by the end-of-line readout. After discarding those heralded events, the logical subspace exhibits median $T_1^L = 3.12$ ms and $T_{2E}^L = 0.76$ ms, i.e. bit-flip and phase-flip rates suppressed by factors of 48 and 11 relative to the constituent $D$ and $Q$ modes, with Ramsey coherence improved by about 4x. The erasure rate closely tracks the physical $T_1$ fluctuations, confirming that the detected leakage is ordinary amplitude damping. The same numbers reproduce on three devices with mode detunings from 0.76–1.05 GHz, showing the improvement is architectural rather than sample-specific.","pith_inferences":["The paper measures only end-of-line detection; a natural extension would be a mid-circuit erasure check, which would turn the postselected memory into a repeat-until-success primitive whose overhead is set by the measured erasure rate.","The observed Ramsey improvement on Q2 (~4.5x) exceeds the photon-shot-noise-only estimate from Eq. (7) (~1–2x), suggesting that other common-mode noise is dominant; sweeping the readout photon number while tracking $T_{2E}^L$ would test this separation directly.","If junction-asymmetry noise is the limiting differential source, then fabricating devices with $r = E_{J1}/E_{J2}$ closer to unity, or post-processing the junctions, should raise $T_{2E}^L$; the paper's tabulated $r$ values make this a quantitative prediction across Q1–Q3."],"forward_implications":["A fixed-frequency multimode transmon can act as an erasure-converted logical qubit with all-microwave control and no flux or galvanic-coupling overhead.","The error-detected suppression is stable over 50-hour runs on three separate devices, so it is not a one-sample coincidence.","Because the logical subspace cancels common-mode frequency noise, comparing logical and physical Ramsey or Hahn-echo decays gives a per-device estimate of the common versus differential noise budget.","The erasure rate follows the physical $T_1$ fluctuations, so raising physical mode relaxation times directly lowers the postselection overhead of this logical protocol."],"supporting_citations":[{"why":"It introduces erasure conversion as a route to fault tolerance, providing the theoretical motivation for converting amplitude damping into detectable leakage.","marker":"[17]"},{"why":"It shows erasure qubits can overcome the $T_1$ limit in superconducting circuits, supplying the payoff the DDQ aims for.","marker":"[18]"},{"why":"It demonstrates dual-rail encoding in superconducting cavities, a prior platform for the same erasure-conversion idea.","marker":"[23]"},{"why":"It demonstrates a superconducting dual-rail cavity qubit with erasure-detected logical measurements, setting the experimental pattern for end-of-line detection.","marker":"[24]"},{"why":"It supplies the bit-flip measurement methodology and the tunable-transmon dual-rail baseline the DDQ results are compared against.","marker":"[27]"},{"why":"It establishes the multimode transmon's two-mode structure and charge-sensitivity properties used for the dimon device.","marker":"[29]"},{"why":"It defines the dimon device concept whose two-mode Hamiltonian underlies the logical encoding.","marker":"[41]"},{"why":"It provides the decoherence benchmarking and Allan-deviation analysis used to separate white and $1/f$ noise in the logical Ramsey data.","marker":"[54]"}],"fun_headline_variants":["Erasure conversion in one transmon cuts bit-flips 48x","Dual-rail code inside a single fixed-frequency transmon","Multimode transmon heralds decay to shield logical bit-flips","48x bit-flip suppression via dual-rail in a single qubit","Single device dual-rail: bit-flips 48x lower, phase-flips 11x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The protocol assumes that relaxation from either logical state always goes to the empty $|00\\rangle$ state, never directly between $|01\\rangle$ and $|10\\rangle$, and that the three-state readout separates $|00\\rangle$ from the logical states well enough that postselection removes essentially all amplitude-damping events.","fun_headline_variants_meta":{"raw":{"variants":["Erasure conversion in one transmon cuts bit-flips 48x","Dual-rail code inside a single fixed-frequency transmon","Multimode transmon heralds decay to shield logical bit-flips","48x bit-flip suppression via dual-rail in a single qubit","Single device dual-rail: bit-flips 48x lower, phase-flips 11x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000551,"raw_usage":{"total_tokens":2618,"prompt_tokens":922,"completion_tokens":1696,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":1594}},"tokens_in":538,"tokens_out":1696,"duration_ms":13168,"temperature":1.0,"reasoning_tokens":1594,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:35:14.629409+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the GMM classifier on a different readout drive frequency or with a stricter $|00\\rangle$ rejection threshold and remeasure the logical bit-flip decay: if $T_1^L$ shifts, readout misclassification is biasing the reported suppression. Independently, search spectroscopically for any transition between $|01\\rangle$ and $|10\\rangle$; observing direct coupling between the logical states would break the symmetry argument and invalidate the erasure interpretation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It introduces erasure conversion as a route to fault tolerance, providing the theoretical motivation for converting amplitude damping into detectable leakage."},{"cited_title":"Kubica, A","cited_arxiv_id":null,"evidence_quote":"It shows erasure qubits can overcome the $T_1$ limit in superconducting circuits, supplying the payoff the DDQ aims for."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It demonstrates dual-rail encoding in superconducting cavities, a prior platform for the same erasure-conversion idea."},{"cited_title":"Levine, A","cited_arxiv_id":null,"evidence_quote":"It supplies the bit-flip measurement methodology and the tunable-transmon dual-rail baseline the DDQ results are compared against."},{"cited_title":"Wills, G","cited_arxiv_id":null,"evidence_quote":"It establishes the multimode transmon's two-mode structure and charge-sensitivity properties used for the dimon device."},{"cited_title":"Hazra, K","cited_arxiv_id":null,"evidence_quote":"It defines the dimon device concept whose two-mode Hamiltonian underlies the logical encoding."}],"review_version":2}