{"id":"ecd4002b-2a4c-47df-b2c8-fd77d99b24f3","arxiv_id":"2504.17082","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A compilation of experiments showing that small surface codes on superconducting transmons can realize logical initialization, measurement, and gates, with fault-tolerant variants outperforming non-fault-tolerant ones and logical error rates below physical error rates.","lead":"This PhD thesis reports a series of experiments on small superconducting quantum processors that implement surface-code error detection and correction building blocks, including high-fidelity gates, automatic calibration, and leakage reduction. It demonstrates logical operations on a distance-2 logical qubit and shows that fault-tolerant variants outperform non-fault-tolerant ones.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Logical-vs-physical error comparison in Supp 3.6.5 is internally inconsistent: the j+_L logical error rate exceeds the best physical T2 error rate.","rationale":"The paper's core contribution - a distance-2 surface code with a complete set of logical operations and fault-tolerant variants outperforming non-fault-tolerant variants - is supported by peer-reviewed publication (Nature Physics 18, 80 (2022)) and detailed characterization. The reader's identified assumption in Supp 3.6.11 (pre-rotation single-qubit gate errors negligible) is likely safe: Table 3.2 lists single-qubit gate fidelities 99.83-99.98%, while multiplexed readout fidelities are 94.2-98.9%, so single-qubit gate errors are at least an order of magnitude smaller. The more significant issue is the supplementary comparison of logical vs physical error rates, which is internally inconsistent for the j+_L state given the paper's own quoted numbers. This does not overturn the main logical-operations claim, but it means one clause of the reader's strongest_claim should be qualified. Because the unpublished chapters 4, 6, and 7 remain unavailable for full assessment and the leakage free parameter in Model 5 is fitted rather than directly measured, the CONDITIONAL verdict remains appropriate.","tokens_in":57136,"tokens_out":19817,"duration_ms":188609,"concrete_test":"Recompute Fig. 3.9 per-cycle logical error rates for each of the four logical states from the archived data (github.com/DiCarloLab-Delft/Logical_Qubit_Operations_Data) using the pipelined cycle time of 840 ns. For each state, compare against the corresponding physical decoherence error per cycle: |1_L> vs best T1, |0_L> vs ground-state error (residual excitation plus dephasing), and |+_L>/|-_L> vs best T2 (or T2*) using Eq. 3.23. If the fitted j+_L rate remains above the physical rate, the 'all states' claim in Supp 3.6.5 is contradicted and must be qualified to 'some states' or to an average over symmetric state pairs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Supplement 3.6.5 (Eqs. 3.21-3.23, Fig. 3.9) claims that logical error rates for all states are lower than the corresponding best physical error rates. The quoted per-cycle logical error rates for the pipelined scheme (840 ns cycle) are 0.43% (j0_L), 0.67% (j1_L), 0.49% (j+_L), and 0.15% (j-_L). Using the best echo T2 = 117 us (D4, Table 3.2), Eq. 3.23 gives a physical dephasing error per cycle of 1 - exp(-840 ns / (2 x 117 us)) ~ 0.36%. Thus the j+_L rate of 0.49% is about 36% larger than the best physical T2 error, directly contradicting the 'all states' claim. The comparison is also asymmetric for j0_L: T1 decay does not affect the physical |0> state, so comparing j0_L to the |1> T1 error is not the corresponding physical error, and a fair state-matched comparison would make j0_L look worse. The 'below best physical qubit error rates' clause in the reader's central claim is therefore not supported by the paper's own numbers.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This PhD thesis compiles experimental work on flux-tunable transmon surface-code processors. Chapter 2 presents the SNZ controlled-Z gate, reporting up to 99.93% gate fidelity and 0.10% leakage. Chapter 3 demonstrates a distance-2 Surface-7 logical qubit stabilized by repeated error detection, including logical initialization to arbitrary states, logical measurement in the ZL/XL/YL bases, and a universal single-qubit logical gate set; a central claim is that fault-tolerant variants outperform non-fault-tolerant variants. Chapter 4 describes automatic calibration and benchmarking of a 17-transmon Surface-17 device, including techniques for coping with two-level-system defects and frequency-targeting errors. Chapters 5-7 cover leakage reduction units, calibration of QEC cycles, and soft-information decoding for a bit-flip code. The central narrative is that small-scale surface-code error detection can realize a complete logical-qubit operation set and that the building blocks can be calibrated and optimized for larger codes.","tokens_in":57383,"tokens_out":11238,"duration_ms":101038,"significance":"If the results hold, the thesis provides valuable experimental milestones for surface-code quantum error correction with superconducting qubits: a complete logical-operation suite on a distance-2 surface code, a high-fidelity and easily tuned two-qubit gate, automatic calibration at the 17-qubit scale, and a demonstrated benefit of soft-information decoding. Strengths include the availability of processed data for Chapters 2 and 3, the use of standard randomized benchmarking and tomography with quoted error bars, the detailed comparison of pipelined versus parallel stabilizer readout, and the transparent treatment of leakage as a free parameter in simulations. The SNZ gate argument is concise and analytical in its ideal limit. However, one supporting comparison in the Chapter 3 supplementary material is internally inconsistent and needs correction before the manuscript can be accepted as a coherent record of logical performance.","major_comments":[{"comment":"The statement that \"the logical error rates for all states ... are lower than the corresponding best physical error rates\" is not supported by the paper's own numbers. For the |+_L> state, the fitted per-cycle logical error rate is 0.49%, while Eq. (3.23) with the best measured echo T2 = 117 us (D4, Table 3.2) and the 840 ns pipelined cycle gives a physical dephasing error of 1 - exp(-840 ns / (2 x 117 us)) ~ 0.36%; the logical rate is therefore about 36% larger, contradicting the \"all states\" claim. In addition, the comparison for |0_L> uses the T1 error of the physical |1> state (Eq. 3.22), which is not the corresponding error for a physical qubit prepared in |0>; a state-matched comparison would make the |0_L> rate of 0.43% per cycle look worse, not better. Please correct or remove this comparison, or present a state-matched and cycle-matched physical baseline.","section":"Supplement 3.6.5, Eq. (3.23), Fig. 3.9"}],"minor_comments":[{"comment":"The cross-reference to \"Table 5.1\" in these sections should point to Table 3.2, the device-characteristics table in Chapter 3; the current numbering is confusing because Chapter 5 has its own tables.","section":"Section 3.6.11 and Section 3.2.1"},{"comment":"The text contains an unresolved citation \"[77, 128?]\" with a literal question mark; this reference should be completed or removed.","section":"Section 3.4"},{"comment":"The notation 1 - e^{-t/T2/2} is ambiguous; please write the intended expression explicitly, e.g., 1 - exp(-t/(2T2)) or 1 - exp(-2t/T2), and ensure it matches the physical dephasing model used in the comparison.","section":"Eq. (3.23)"},{"comment":"There are several typographical errors, including \"benchmakring\" in the Chapter 4 title and headings, \"envolved\" in Section 3.4, and \"uncertasinties\" in the Fig. 4.3 caption; these should be corrected in a final proofreading pass.","section":"Chapter 4, title and figure captions"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a PhD thesis that reuses several already-published papers, in particular Chapters 2, 3, 5, and 7. If the venue expects original journal content, the editor should assess whether the new material in Chapters 4, 6, and 8 provides sufficient incremental contribution. The main technical issue identified in the report is localized to the supplementary logical-vs-physical comparison and is correctable without affecting the core logical-operation results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent experimental thesis built on two strong published results, and the engineering chapters are worth reading. But the supplementary claim that every logical error rate is below the best physical error rate does not survive the paper's own numbers. For the pipelined 840 ns cycle, the j+_L logical error rate is 0.49% while the best T2 dephasing error is about 0.36%; so the \"all states\" sentence in Supp. 3.6.5 is wrong. That is a real flaw, but it is a comparison claim in the supplement, not one of the load-bearing experimental results.\n\nWhat is genuinely new: the SNZ CZ gate and the Surface-7 logical operations were published in PRL and Nature Physics, and they are strong experiments with standard benchmarking, error bars, and linked processed data. The thesis adds useful unpublished material: Graph-Based Tuneup for a 17-transmon device, engineering around TLS defects and frequency trimming, leakage-reduction units, and soft-information decoding for a distance-3 bit-flip code. I could not fully check Chapters 4 and 6 because the text was truncated and no data/code are linked for those chapters; the engineering content looks plausible and is clearly presented.\n\nSoft spots, in order of importance. First, the leakage per CZ gate L1 in the Chapter 3 simulation is explicitly a fit parameter, and the authors admit their histogram-based estimates are not accurate enough to pin it down. That weakens the quantitative statement about leakage being the dominant missing error source. A direct leakage measurement, even on one gate, would firm this up. Second, the stress-test issue above: the logical-vs-physical comparison should be corrected or removed. Third, the logical readout fidelities rely on the assumption that single-qubit gate errors are negligible compared to readout errors; that is probably safe given Table 5.1, but it is an assumption and should be stated as such. The self-citation pattern is not a problem here; the earlier papers are independent benchmarks, and this is a thesis.\n\nWho should read it: experimentalists working on small surface codes and calibration automation, and theorists who want concrete numbers for error models. It deserves peer review if submitted as a review or monograph; a referee should ask for the leakage free parameter to be addressed and the unsupported comparison fixed. It is not a desk reject.","headline":"Solid thesis compilation with two strong published experiments and useful engineering content, but the supplementary claim that all logical error rates beat the best physical error rates is contradicted by the paper's own numbers.","tokens_in":57904,"tokens_out":4174,"would_cite":false,"duration_ms":40768,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Run under repeated error detection, a seven-transmon distance-2 surface code can initialize, measure, and gate a logical qubit through a universal single-qubit set, with fault-tolerant variants outperforming non-fault-tolerant ones.","keywords":["surface code","quantum error detection","logical qubit","superconducting transmon","fault-tolerant operations","logical Pauli transfer matrix","leakage reduction","soft-information decoding"],"falsifier":"Run the logical $X_L$ measurement with a deliberately inserted single-qubit pre-rotation error (for example, a 1% under-rotation on one data qubit) while keeping data-qubit readout unchanged, and extract $F_R^L$ with the paper's assignment-probability-matrix method; if the extracted value moves materially away from 98.7%, the negligible-single-qubit-error premise is false and the quoted logical readout fidelities need recalibration.","tokens_in":56895,"feed_emoji":"⚛️","tokens_out":9728,"duration_ms":79642,"temperature":0.7,"pith_summary":"This thesis argues that the surface-code architecture can be carried through the entire experimental stack on flux-tunable superconducting transmons, from calibrated two-qubit gates to logical-level operations. Its central experimental claim is that a distance-2 surface code using seven transmons supports a complete logical-qubit operation set—initialization anywhere on the logical Bloch sphere, measurement in all three cardinal bases, and universal single-qubit logical gates—while repeated stabilizer measurements detect errors and post-selection discards runs in which an error is seen. For every operation type, the fault-tolerant implementation outperforms the non-fault-tolerant one (logical readout fidelities of 99.8% for $Z_L$, 98.7% for $X_L$, 91.4% for $Y_L$), and the logical error rate per error-detection round is below the best physical-qubit error rate. Around this result the thesis develops the supporting building blocks: a sudden net-zero controlled-Z gate reaching 99.93% fidelity, automated calibration of a 17-transmon surface code, leakage-reduction units, and soft-information decoding. If the central claim holds, small surface codes can function as testbeds for fault-tolerant protocols rather than as merely stabilized memories.","feed_headline":"Seven-qubit surface code runs a full logical-qubit toolkit","feed_subtitle":"Fault-tolerant operations win, and logical errors fall below the best physical-qubit rates.","key_machinery":"The load-bearing object is the distance-2 surface-code logical qubit: seven flux-tunable transmons, four data and three ancilla, arranged so that the codespace is the even-parity subspace of the stabilizer set $S = \\{Z_{D1}Z_{D3},\\; X_{D1}X_{D2}X_{D3}X_{D4},\\; Z_{D2}Z_{D4}\\}$. The stabilizer measurements project the data into the codespace and flag errors by returning $-1$; because this code cannot uniquely identify every error, the protocol post-selects runs with no detected error. Fault tolerance is achieved by designing each logical operation so that any single fault either yields a non-trivial syndrome or leaves the desired logical state with a detectable error; the order of CZ gates in the weight-4 stabilizer circuit is what prevents ancilla faults from becoming logical errors. The second carrying mechanism is the logical Pauli transfer matrix, a $4\\times 4$ real matrix mapping input logical Pauli expectation values to outputs, which lets logical gates be benchmarked with a single average fidelity. The third is the sudden net-zero (SNZ) controlled-Z pulse, whose two square half-pulses realize an exact Mach-Zehnder interferometer for the $|11\\rangle$-$|02\\rangle$ transition, giving 99.93% gate fidelity and a regular leakage landscape that makes tuneup simple.","core_discovery":"The core discovery is that a distance-2 surface code (Surface-7), four data transmons and three ancillas, can be operated as a logical qubit under repeated error detection, not just as a stabilized quantum memory. The work demonstrates initialization of arbitrary logical states, fault-tolerant logical measurement in the $Z_L$ and $X_L$ bases, a non-fault-tolerant $Y_L$ measurement, and a universal single-qubit logical gate set that includes transversal $R_X^{\\pi}$ and $R_Z^{\\pi}$ gates and gate-by-measurement $R_X^{\\pi/2}$ and $T_L$ gates. It characterizes the gates with a logical Pauli transfer matrix, reporting average logical gate fidelities of 97.9%, 98.1%, 95.6%, and 97.3% for those four gates. For each operation type the fault-tolerant variant beats the non-fault-tolerant variant, and logical error rates per round (0.15% to 0.67%) lie below the best physical-qubit $T_1$ and $T_2$ error rates. The work further compares two scalable stabilizer-measurement schemes, pipelined and parallel, finding the pipelined scheme slightly better, and identifies leakage to higher transmon states during CZ gates as the dominant error source not captured by standard decoherence and readout models.","pith_inferences":["If the negligible-single-qubit-error premise in the logical-readout analysis is relaxed, the exact quoted $F_R^L$ values (99.8/98.7/91.4%) would need recomputation; the fault-tolerant ordering might survive, but the margins could narrow.","The same assignment-probability-matrix method used to extract logical readout fidelity without initialization corruption could be applied directly to the distance-3 code, giving a readout benchmark that is independent of state-preparation errors.","A natural next experiment is to run the Chapter 3 protocol on a distance-3 device with the automated calibration, leakage reduction, and soft-information decoding of Chapters 4, 5, and 7; if the per-round logical error rate drops when distance increases, the exponential-suppression claim would be tested directly."],"forward_implications":["A complete logical operation set, including a non-Clifford $T_L$ gate, can be executed on a distance-2 code with repeated error detection, providing a template for the logical-level primitives needed in higher-distance surface codes.","Fault-tolerant circuit design pays off even at the smallest code distance: every operation type shows a measurable improvement of the fault-tolerant variant over the non-fault-tolerant variant.","Logical error rates per round below the best physical qubit error rates imply that error detection already confers a logical benefit, and that the path to fault tolerance is through code distance rather than through better physical qubits alone.","A scalable stabilizer-measurement scheme with distance-independent cycle time, such as the pipelined scheme, can be used without sacrificing logical performance; its per-cycle error rate is about 3% lower than the parallel scheme.","Leakage out of the computational subspace, chiefly from CZ gates, must be actively managed in repeated stabilizer experiments; post-selection and leakage-reduction units are the near-term tools, and the thesis shows both working."],"supporting_citations":[{"why":"Supplies the prior distance-2 surface-code experiment that prepared logical cardinal states and measured in two bases; the present work extends it to a full logical operation suite.","marker":"[74]"},{"why":"Introduces the SNZ controlled-Z gate used throughout, with up to 99.93% fidelity and 0.10% leakage.","marker":"[69]"},{"why":"Provides the Net Zero bipolar flux-pulse method and the error-budget simulation approach that SNZ builds on.","marker":"[94]"},{"why":"Proposes the scalable quantum-hardware architecture and the pipelined/parallel stabilizer-measurement schemes compared in Chapter 3.","marker":"[98]"},{"why":"Gives the fault-tolerance criterion requiring single faults to produce non-trivial syndromes, used to justify the CZ gate ordering.","marker":"[121]"},{"why":"Supplies the leakage model and leakage-conditional phase formalism used to estimate CZ leakage and simulate its impact on error detection.","marker":"[77]"}],"fun_headline_variants":["Full logical qubit on 7-qubit surface code","Fault-tolerant gate set beats physical error rates","Logical errors fall below physical on Surface-7","Universal logical gates on 7-qubit code"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quoted logical readout fidelities rest on the assumption that single-qubit gate errors during the measurement pre-rotations are negligible compared with data-qubit readout errors; if that premise fails, the fault-tolerant versus non-fault-tolerant readout comparison shifts.","fun_headline_variants_meta":{"raw":{"variants":["Full logical qubit on 7-qubit surface code","Fault-tolerant gate set beats physical error rates","Logical errors fall below physical on Surface-7","Universal logical gates on 7-qubit code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1565,"prompt_tokens":1092,"completion_tokens":473,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":412}},"tokens_in":708,"tokens_out":473,"duration_ms":4874,"temperature":1.0,"reasoning_tokens":412,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:49:52.247928+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the logical $X_L$ measurement with a deliberately inserted single-qubit pre-rotation error (for example, a 1% under-rotation on one data qubit) while keeping data-qubit readout unchanged, and extract $F_R^L$ with the paper's assignment-probability-matrix method; if the extracted value moves materially away from 98.7%, the negligible-single-qubit-error premise is false and the quoted logical readout fidelities need recalibration.","supporting_citations":[],"review_version":1}