{"id":"875610e5-2284-4bb9-8ba1-fd2eb5d79641","arxiv_id":"1908.01869","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Multilevel encoding and repeated quantum nondemolition readouts suppress relaxation-induced measurement errors in superconducting qubits, achieving logical assignment infidelities down to 5.8x10^-5.","lead":"By encoding qubits in multiple energy levels of superconducting circuits, the authors show that measurement errors caused by relaxation can be suppressed by orders of magnitude, reaching assignment infidelities as low as 5.8x10^-5 for bosonic Fock states. The technique also allows direct detection of transmon gate errors at the one-part-per-thousand level, which is normally obscured by measurement noise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline Fock-code infidelity of 5.8e-5 may be dominated by the 0.2% 'stuck' ancilla-reset events whose mechanism the authors explicitly leave unexplained.","rationale":"The paper's core idea is well supported by the direct transmon shelving measurements and by the systematic improvement of Fock-code assignment with code distance in Fig. 5; I do not see a reason to reject or downgrade the work. However, the most load-bearing number in the abstract, the 5.8e-5 Fock-code infidelity, is at risk from an acknowledged but uncharacterized rare process. The reader's weakest assumption emphasized QND-ness and the non-independence of the theory-model comparison; those are reasonable secondary concerns, but the more specific and more consequential gap is the lack of any conditional analysis linking the 0.2% stuck-ancilla fraction to the final misassignment counts. Because the stuck fraction exceeds the headline error by more than an order of magnitude, a small conditional error rate among stuck runs could completely explain the reported infidelity without any contribution from the intended multilevel-encoding physics. The manuscript already has the raw data needed to resolve this, so the appropriate verdict remains conditional: accept the scientific contribution, but require the stuck-run breakdown (or release of raw sequences) before the specific headline number is taken as a clean demonstration. This matches the reader's CONDITIONAL verdict, so no change is needed.","tokens_in":17397,"tokens_out":32645,"duration_ms":344895,"concrete_test":"Re-analyze the raw readout sequences underlying Table I. Flag every experiment whose reset protocol required five or more conditional feedforward iterations (the paper's own 'stuck' definition), then recompute the MLE assignment infidelity for the Fock |0> vs |5> code three ways: (i) including all runs, (ii) excluding stuck runs, and (iii) restricted to stuck runs only. Also report the fraction of all misassignments that occur in stuck runs. If the non-stuck infidelity remains within error bars of 5.8e-5, the headline is robust; if it drops sharply (e.g., below 3e-5), the reported value is dominated by the unexplained reset anomaly and must be quoted separately from the multilevel-encoding result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is the 5.8e-5 logical assignment infidelity for the Fock |0> vs |5> encoding in Table I. The paper states (Supplemental Fig. S3) that 0.2% of 1.8 million repeated-measurement experiments contain a 'stuck' ancilla reset, defined as requiring five or more feedforward reset iterations, and that the detailed mechanism of these high-transmon-level excitations is 'not apparent from our measurements.' The authors postselect away these events only for the theory comparison in Fig. 5, and explicitly state that Fig. 6 and Table I are not postselected. This creates a serious quantitative gap: the stuck-event rate (2e-3) is about 34 times larger than the headline infidelity (5.8e-5). If a stuck event causes a wrong MLE assignment with even a few percent probability, those rare events would account for essentially the entire reported error budget. The hidden Markov model uses actual reset times in the transition matrix, but it does not model the high-ancilla-level 'stuck' state; its emission matrix is calibrated for g/e/f/h readouts, so a run with a stuck ancilla is outside the model used to compute the reported assignment. Without a conditional breakdown of misassignments by reset history, the headline number cannot be attributed to the multilevel-encoding suppression of relaxation errors; it may instead be an artifact of an unmodeled, rare, correlated error process.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper demonstrates that encoding a qubit in multilevel states suppresses relaxation-induced measurement errors in superconducting circuits. Two experimental settings are studied: (i) a transmon whose higher excited states |f> and |h> are used to improve readout contrast, including a shelving technique that directly resolves g–e gate errors at the level of about 1.4e-3; and (ii) a bosonic storage mode measured through a transmon ancilla with repeated map–measure–reset cycles. For the bosonic encodings, the authors report logical assignment infidelities of 5.8e-5 for the Fock encoding |0> vs |5> and 4.2e-3 for the S=2,N=1 binomial code, obtained with a hidden-Markov-model/MLE classifier. The paper also proposes Eq. (4) to describe the tradeoff between repeated readouts and state relaxation and compares Fig. 5 data to this model. The central claim is that multilevel encoding plus repeated readout yields record-level measurement fidelities in circuit QED.","tokens_in":17712,"tokens_out":4147,"duration_ms":43410,"significance":"If the reported numbers are robust, this is an important experimental advance: the 5.8e-5 assignment infidelity and the 1.4e-3 direct resolution of gate errors are substantially better than standard transmon readout, and the multilevel-encoding principle is general and applicable to other platforms. The strength of the paper is that the headline infidelities are direct experimental assignments with error bars, not theory extrapolations; the QND-ness of the readout is separately calibrated (PD = 0.02%); and the shelving demonstration cleanly separates measurement error from gate error in a way that is useful for SPAM characterization. The main caveats concern the treatment of rare reset failures and the circularity of the model comparison, as detailed below.","major_comments":[{"comment":"The paper explicitly states that 0.2% of 1.8 million repeated-measurement experiments contain a 'stuck' ancilla reset requiring five or more feedforward iterations, and that the mechanism of these high-level excitations is not apparent from the measurements. These events are not modeled in the hidden-Markov-model transition matrix, whose emission matrix is calibrated only for g/e/f/h readouts. Since the results in Fig. 6 and Table I are not postselected, the reported Fock-code infidelity of 5.8e-5 is about 34 times smaller than the 2e-3 stuck-event rate. If even a few percent of stuck events corrupt the MLE assignment, these rare correlated events would account for essentially the entire reported error budget. The authors should provide a conditional breakdown of misassignments by reset history (e.g., with and without stuck events), or otherwise demonstrate that the headline infidelities are not dominated by this unmodeled process. As written, the central quantitative claim is not yet supported.","section":"Supplemental Material, 'Reset errors'; main text §IV, Fig. 6 and Table I"},{"comment":"The theory curves in Fig. 5 are not an independent test of the photon-loss model: delta0 and delta1 are taken as the first points of the same curves being compared, and kappa_up tau is fitted to the 1-photon curve. The abstract's statement that the tradeoff is 'shown to be consistent with the photon-loss model' should therefore be qualified. The measured infidelities themselves are direct assignments and are not invalidated by this circularity, but the model agreement in Fig. 5 does not add independent confirmation. The authors should either refit the model to independent calibration measurements or explicitly state which parameters are fixed a priori.","section":"§IV, Eq. (4); Supplemental Material, 'System parameters'"},{"comment":"The argument that preparation errors can be neglected relies on the assumption that the only significant residual error is photon loss during the final check measurement, which the protocol is robust to. This is plausible for the Fock and binomial codes studied, where a single loss event from |L> or |2>/|3> does not cross the S/ar S boundary. However, the argument is not fully quantitative: the paper does not provide an upper bound on other preparation errors (e.g., residual thermal population or mapping-pulse miscalibration) that could enter the reported infidelity. Adding an explicit error budget for state preparation would strengthen the claim that the quoted infidelities are measurement infidelities.","section":"Supplemental Material, 'State preparation'"}],"minor_comments":[{"comment":"Typo: 'beloning' should be 'belonging'.","section":"Fig. 2 caption"},{"comment":"Typo: 'simulataneous' should be 'simultaneous'.","section":"Supplemental Material, 'Simultaneous number-selective pulses'"},{"comment":"The caption mentions dashed and dash-dotted lines but does not define which is which in terms of the two contributions (majority-vote vs. state-transition). Please add a legend or textual description.","section":"§IV, Fig. 5"},{"comment":"The notation 'S=2,N=1 binomial code' is used without a definition in the main text; a brief sentence citing Ref. [25] and explaining S and N would help readers.","section":"Table I"},{"comment":"The phrase 'the measurement is only corrupted when multiple errors occur' is slightly imprecise for the transmon case, where a single relaxation event from |f> still leaves the state in the excited manifold; consider rewording to 'a single relaxation event does not corrupt the encoded bit' for clarity.","section":"§III"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically strong and the central idea is valuable, but the stuck-reset contamination issue is load-bearing because the headline infidelity is smaller than the rate of these unmodeled events. I would ask for a conditional analysis (with/without stuck resets) and a revision of the model-comparison claims. If the authors can show that the reported infidelities are stable under postselection on reset success, the paper would be a strong candidate for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi —\n\nQuick take on the Elder et al. paper: it does what it says. The multilevel encoding idea (which is not new; Hann et al. proposed it, and latching readout has been around) is actually implemented and pushed to the point where single-shot readout of transmon |g> vs |e> can resolve ~1e-3 gate errors, and the Fock |0>-vs-|5> assignment infidelity is 5.8e-5. Those numbers are direct experimental assignments, not fits. That's a genuine advance for cQED readout and gate calibration.\n\nThe paper is transparent in the right ways: it calibrates the QND-ness of the ancilla readout (PD = 0.02%), it uses a hidden Markov model with actual reset times in the transition matrix, and it clearly states that the Fig. 5 theory comparison is postselected on successful resets while Fig. 6 and Table I are not. The systematic improvement with code distance in Fig. 6 is visible by eye and doesn't depend on the model parameters.\n\nThe soft spots are real but not fatal. First, the 'consistent with the photon-loss model' claim in the abstract is weaker than it sounds. In Fig. 5, delta0 and delta1 are taken from the first points of the same curves, and kappa_up*tau is fitted to one curve, so the agreement isn't an independent test. The direct assignment data are the strong evidence; the model fit is illustrative. Second, and more importantly, the 0.2% stuck-ancilla reset events are a concern. The stuck rate is 2e-3, about 34x the headline 5.8e-5. If even a few percent of stuck events flip the MLE assignment, they would account for essentially all of the reported infidelity. The HMM doesn't model the high-level stuck state, and the paper doesn't give a conditional breakdown of misassignments by reset history. The authors are honest that the mechanism is unclear, but for a claim of 6e-5 infidelity, the reader needs to know whether the residual error is relaxation-limited or dominated by these rare technical events. This is the question I'd want answered before quoting the headline number. That said, it doesn't undermine the central point that multilevel encodings help; it just means the limiting error source for the best Fock code is not fully identified.\n\nWho is this for? Anyone working on superconducting readout, bosonic codes, or gate calibration. It's a solid experimental paper that deserves a serious referee. I'd send it to review with a request for the reset-history breakdown and a clarification of the model fitting. I'd also suggest the authors release the raw sequences/code, since that would resolve most of the circularity worries.\n\nNet: worth engaging with, worth citing once the stuck-event question is addressed.\n\nBest,","headline":"A real experimental advance in multilevel readout, but the headline 5.8e-5 infidelity may be limited by unmodeled 0.2% stuck reset events rather than relaxation.","tokens_in":18318,"tokens_out":6882,"would_cite":true,"duration_ms":67087,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Encoding a qubit across widely spaced multilevel states suppresses relaxation-induced measurement errors, giving logical assignment infidelity as low as $5.8\\times10^{-5}$ and resolving transmon gate errors near $1.4\\times10^{-3}$.","keywords":["multilevel encoding","state-preserving readout","superconducting qubits","bosonic codes","binomial codes","shelving readout","gate characterization","assignment infidelity"],"falsifier":"Prepare the same storage codewords with an independent tomographic verification rather than the heralded number-selective checks, then measure the logical assignment infidelity while scanning the storage loss rate $\\kappa_\\downarrow$ or the readout drive power; if the error does not follow the predicted $(\\kappa_\\downarrow\\tau)^L$ scaling, or if the inferred demolition probability rises above $0.02\\%$ with drive power, the central claim is false.","tokens_in":17204,"feed_emoji":"⚛️","tokens_out":12490,"duration_ms":113006,"temperature":0.7,"pith_summary":"Superconducting qubit readout is usually capped by a tradeoff: collect signal longer to beat noise, but risk the qubit relaxing during the measurement and being assigned the wrong bit. This paper tries to break that tradeoff by encoding the bit in states separated by several energy levels, so one relaxation event no longer flips the answer. It demonstrates the idea on a transmon with a shelving pulse, and on a bosonic storage cavity with repeated state-preserving readouts of Fock and binomial encodings. The result is a logical assignment infidelity of $5.8\\times10^{-5}$ when distinguishing $\\lvert 0\\rangle$ from $\\lvert 5\\rangle$, and $4.2\\times10^{-3}$ for a binomial error-correcting code, with transmon gate errors resolved directly at the $10^{-3}$ level and the Fock result surpassing previous circuit-QED readout fidelities.","feed_headline":"Multilevel codes cut qubit readout error to 5.8e-5","feed_subtitle":"Spacing qubit codewords apart defeats relaxation during readout and reveals gate errors at the 1e-3 level.","key_machinery":"The carrying object is the code distance $L$ between the logical codewords with respect to the photon-loss channel: for Fock codes the codewords are $\\lvert 0\\rangle$ and $\\lvert L\\rangle$, and the measured subspace $S$ is chosen so that losing or gaining a small number of photons does not move a state across the $\\{S,\\bar S\\}$ partition. The readout is formalized as a projective measurement of membership in $S$, with fidelity $F=1-P(S\\mid \\lvert 1_L\\rangle)-P(\\bar S\\mid \\lvert 0_L\\rangle)$, and the repeated-readout protocol combines a dispersive storage–ancilla map that flips the ancilla only when the storage state lies in $S$, a real-time ancilla reset that makes the readout effectively repeatable, and a hidden Markov maximum-likelihood classifier that weighs all votes against the growing chance of relaxation. The theoretical infidelity, Eq. (4), splits into majority-vote terms that fall with the number of readouts and transition terms proportional to $(\\kappa_\\downarrow\\tau)^{L-1}$ and $(\\kappa_\\uparrow\\tau)^2$ that rise as the code distance grows.","core_discovery":"The central discovery is that the leading relaxation error is promoted to a higher order when codewords are separated by $L$ photon-number steps: the probability that relaxation corrupts the measurement scales like $(\\kappa_\\downarrow \\tau)^L$ rather than $\\kappa_\\downarrow \\tau$, so a code of distance $L$ is robust until $L$ errors accumulate. On the transmon, this means reading out the $\\lvert g\\rangle$–$\\lvert h\\rangle$ manifold instead of $\\lvert g\\rangle$–$\\lvert e\\rangle$; applying an $e$–$f$ shelving pulse turns the usual qubit measurement into a higher-distance measurement and exposes gate errors of order $10^{-3}$ that ordinary readout hides. In the storage cavity, a number-selective pulse maps the encoded bit onto a transmon ancilla, the ancilla is read out and reset by real-time feedback, and the sequence of outcomes is classified by a maximum-likelihood hidden Markov model. The reported best logical assignment infidelities are $5.8\\times10^{-5}$ for the Fock encoding $\\lvert 0\\rangle$ versus $\\lvert 5\\rangle$ and $4.2\\times10^{-3}$ for the $S=2,N=1$ binomial code.","pith_inferences":["If the per-readout demolition probability could be pushed below the calibrated $0.02\\%$, the same protocol should reach assignment errors set by preparation and reset rather than by relaxation; this can be tested by measuring the $\\lvert 0\\rangle$–$\\lvert L\\rangle$ infidelity at still larger $L$.","The gap between the Fock result ($5.8\\times10^{-5}$) and the binomial result ($4.2\\times10^{-3}$) suggests that photon gain and single-round readout errors, not code distance, are the next limit; reducing the storage thermal excitation should narrow that gap.","A direct test of the model would vary $\\kappa_\\downarrow\\tau$ (by changing readout duration or cavity lifetime) and check that the optimal number of votes and the infidelity scale as Eq. (4) predicts.","Codes whose logical states share photon-number support would require a phase-respecting map, because the photon-number-selective flip used here only measures membership in a number subspace; extending the scheme to such codes is a nontrivial next step."],"forward_implications":["If the central claim is right, relaxation no longer sets the practical floor for superconducting-qubit readout; the floor moves to state preparation, ancilla reset, and the state-preserving quality of each readout.","Gate and state-preparation characterization can be made an order of magnitude more precise, because shelved readout resolves transmon gate errors directly at the $1.4\\times10^{-3}$ level instead of hiding them behind larger measurement errors.","For bosonic encodings, logical assignment infidelity should continue to improve roughly exponentially with $L$, so larger-distance Fock codes or higher-order binomial codes should push below $10^{-5}$ on the same hardware.","Repeated state-preserving readout with real-time reset is directly applicable to error-syndrome extraction, where the same ancilla must be reused many times without destroying the logical information.","Because the mechanism is generic to multilevel systems with a one-photon loss channel, the same encoding-and-voting strategy should transfer to other qubit platforms with relaxation-limited readout."],"supporting_citations":[{"why":"supplies the repeated-readout protocol, the hidden Markov model classifier, and the infidelity formula Eq. (4) that the experiment implements","marker":"[26]"},{"why":"provides the binomial code states and code distance that define the QEC-encoded logical qubit measurements","marker":"[25]"},{"why":"is the prior circuit QED single-shot readout result whose fidelity the multilevel measurements are compared against and surpass","marker":"[22]"},{"why":"defines the measurement fidelity metric and the relaxation-limited readout problem that the paper addresses","marker":"[23]"},{"why":"is the earlier latching readout idea extended by the transmon shelving measurement","marker":"[40]"},{"why":"supplies the sideband conversion method used to create higher Fock states in the storage cavity from ancilla excitations","marker":"[44]"}],"fun_headline_variants":["Multilevel encoding cuts qubit readout error to 5.8e-5","Spacing codewords defeats relaxation in superconducting readout","Distance-L codes slash readout infidelity to 5.8e-5","Readout reveals gate errors at 1e-3 via relaxation suppression","Higher-distance encoding yields 5.8e-5 qubit readout infidelity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire result rests on the repeated readout being essentially state-preserving and on independent photon loss and gain being the only significant error processes; if a readout or reset event occasionally shuffles the storage state in a way not captured by Eq. (4), then the quoted infidelities no longer describe errors about the true initial state.","fun_headline_variants_meta":{"raw":{"variants":["Multilevel encoding cuts qubit readout error to 5.8e-5","Spacing codewords defeats relaxation in superconducting readout","Distance-L codes slash readout infidelity to 5.8e-5","Readout reveals gate errors at 1e-3 via relaxation suppression","Higher-distance encoding yields 5.8e-5 qubit readout infidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00096,"raw_usage":{"total_tokens":4176,"prompt_tokens":1116,"completion_tokens":3060,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":732,"completion_tokens_details":{"reasoning_tokens":2959}},"tokens_in":732,"tokens_out":3060,"duration_ms":21071,"temperature":1.0,"reasoning_tokens":2959,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:01:07.673421+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Prepare the same storage codewords with an independent tomographic verification rather than the heralded number-selective checks, then measure the logical assignment infidelity while scanning the storage loss rate $\\kappa_\\downarrow$ or the readout drive power; if the error does not follow the predicted $(\\kappa_\\downarrow\\tau)^L$ scaling, or if the inferred demolition probability rises above $0.02\\%$ with drive power, the central claim is false.","supporting_citations":[{"cited_title":"Hann, Salvatore S","cited_arxiv_id":null,"evidence_quote":"supplies the repeated-readout protocol, the hidden Markov model classifier, and the infidelity formula Eq. (4) that the experiment implements"},{"cited_title":"Michael, Matti Silveri, R","cited_arxiv_id":null,"evidence_quote":"provides the binomial code states and code distance that define the QEC-encoded logical qubit measurements"},{"cited_title":"Walter, P","cited_arxiv_id":null,"evidence_quote":"is the prior circuit QED single-shot readout result whose fidelity the multilevel measurements are compared against and surpass"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the measurement fidelity metric and the relaxation-limited readout problem that the paper addresses"},{"cited_title":"Ong, Agustin Palacios-Laloy, Fran¸ cois Nguyen, Patrice Bertet, Denis Vion, and Daniel Esteve","cited_arxiv_id":null,"evidence_quote":"is the earlier latching readout idea extended by the transmon shelving measurement"},{"cited_title":"Pechal, L","cited_arxiv_id":null,"evidence_quote":"supplies the sideband conversion method used to create higher Fock states in the storage cavity from ancilla excitations"}],"review_version":1}