{"id":"738065d4-5b8e-4538-9a30-0dc2c0dd8ac3","arxiv_id":"2504.16303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Flexion selectively encodes only the qubits involved in two-qubit gates, cutting the overhead of full error correction for early fault-tolerant variational algorithms on trapped ions.","lead":"This paper proposes Flexion, a trapped-ion architecture that runs single-qubit gates on unprotected qubits and two-qubit gates on error-corrected logical qubits, converting between the two in place as needed. The approach could lower the resource barrier for early fault-tolerant quantum computing, a key near-term milestone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 15.5x-vs-FTQC headline is unsupported: the large-scale benchmarks are Clifford-only, yet the MSD-Logical baseline is charged a full magic-state factory, artificially reducing its code distance.","rationale":"The reader's weakest assumption focuses on unmodeled shuttling and LogicMove costs, which is a legitimate gap in the evaluation. I agree that the noise model omits movement and that this could erode the claimed gains. However, the more directly load-bearing issue, and the one I would prioritize, is the baseline mismatch in the Flexion-vs-MSD-Logical comparison. The paper explicitly restricts large-scale ansatze to Clifford circuits, yet charges MSD-Logical for a 2,594-qubit magic-state factory that is unnecessary for Clifford computation. This artificially reduces the baseline's code distance from 9 to 6 and is the primary driver of the reported 7.16x improvement on Heisenberg n=30 and the 15.5x headline claim. That is a concrete, textually verifiable flaw: it does not depend on speculative hardware parameters, and it can be settled by re-running the comparison without the unneeded factory. My recommendation remains CONDITIONAL because the core hybrid bare/logical architecture is coherent and the 8.0x-over-NISQ claim may survive a corrected comparison, but the second headline number should not be stated without a fair baseline. The reader's verdict of CONDITIONAL is therefore unchanged, though the reason differs: the shuttling gap is one concern, and the MSD baseline mismatch is another, more immediately quantifiable concern.","tokens_in":22974,"tokens_out":18259,"duration_ms":191323,"concrete_test":"Re-run the Flexion-vs-MSD-Logical comparison on the large-scale (n=20, 30) Clifford benchmarks with a corrected baseline: MSD-Logical is allowed to use the entire 4,860-qubit budget for distance-9 surface-code logical qubits and no magic-state factory, because the circuits are Clifford. If the VQA energy gap narrows or reverses, the 15.5x headline is an artifact of charging the baseline for an unused T-factory. A second check: run the same comparison on the small-scale non-Clifford benchmarks (e.g., BeH2 and H2O at n=10) and report Flexion's improvement over MSD-Logical without a factory, to see whether any advantage survives when both sides pay only for operations actually executed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section VII-A states that for circuits above 10 qubits 'we restrict the ansatz to Clifford circuits' so that Stim can be used. For a Clifford circuit there are no non-Clifford RZ(theta) gates and hence no T gates, so a standard FTQC baseline needs no magic-state factory. Nevertheless, the Flexion-vs-MSD-Logical comparison in Sec. VII-B gives MSD-Logical a T-factory of 2,594 physical qubits, leaving only enough qubits for distance-6 encoding of the 30 program qubits, while Flexion effectively uses distance-9 logical patches (30 x 2 x 9^2 = 4,860). The reported 7.16x energy improvement for Heisenberg n=30 and the aggregate '15.5x over standard FTQC' claim are therefore largely an artifact of reserving factory qubits for a workload that, as restricted, contains no T gates. Had MSD-Logical been allowed to use all 4,860 qubits for distance-9 logical encoding, the logical-error-rate gap that drives the comparison would be largely eliminated. This is not a noise-model calibration issue; it is a mismatch between the workload class used in the large-scale simulations and the overhead charged to the baseline. Since the abstract's central claim includes the 15.5x improvement over standard FTQC, the comparison as presented does not support that number.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Flexion, a hybrid encoding architecture for trapped-ion QCCD systems in which qubits remain bare for single-qubit operations and are dynamically encoded into surface-code patches only for two-qubit gates. It contributes an in-situ bare-to-logical conversion protocol inspired by gauge fixing, a hybrid ISA with bare, logical, and boundary regions plus LogicMove/Encode_Boundary instructions, and a compiler that schedules encoding conversions and routing. Evaluation uses VQA/UCCSD benchmarks with density-matrix simulation up to 10 qubits and Clifford-restricted Stim simulation for 20-30 qubits, together with Stim+Pymatching surface-code LER estimation. The paper reports an 8.0x improvement over NISQ-Bare execution and a 15.5x improvement over a fully logical MSD-Logical baseline under the same qubit budget.","tokens_in":23252,"tokens_out":7692,"duration_ms":72375,"significance":"The central idea of selective spatial and temporal QEC, with bare qubits for high-fidelity 1Q gates and encoded qubits only for noisy 2Q gates, is timely and potentially valuable for early fault tolerance on trapped-ion hardware. The paper contains several strong components: an in-situ conversion protocol that avoids post-selection, a concrete hybrid ISA, a compiler formulation as register allocation, and an end-to-end simulation pipeline using Stim and Pymatching. If the quantitative claims held, the paper would make a significant systems contribution. However, the headline comparison against standard FTQC is undermined by a baseline mismatch: the large-scale benchmarks are Clifford-only, yet the MSD-Logical baseline is charged a full magic-state factory. Omissions of shuttling and patch-movement errors, an optimistic SPAM rate, and an under-supported distance-independence claim for conversion error further reduce confidence in the reported numbers. The core idea remains defensible, but the evidence as presented does not yet support the central claims.","major_comments":[{"comment":"The headline comparison against MSD-Logical is not supported by the reported experiments. Section VII-A states that for circuits above 10 qubits the ansatz is restricted to Clifford circuits so that Stim can be used. A Clifford circuit contains no non-Clifford RZ(theta) rotations and hence no T gates, so the MSD-Logical baseline needs no magic-state factory. Nevertheless, Section VII-B computes the qubit budget as 30 x 2d^2 with d=9, subtracts a 2,594-qubit factory, and then gives MSD-Logical only distance-6 encoding of the 30 program qubits, while Flexion is effectively charged distance-9 patches. The reported 7.16x energy improvement for Heisenberg n=30 and the abstract's 15.5x improvement over standard FTQC are therefore largely artifacts of reserving factory qubits for a workload that, as restricted, contains no T gates. Figure 10(c)'s depth-overhead metric from Clifford+T decomposition is likewise inapplicable if the large-scale workloads are Clifford-only. Please rerun the comparison with MSD-Logical allowed to use all 4,860 qubits for distance-9 logical encoding (or an equivalently fairly resourced baseline), and report the resulting ratios and the aggregation rule behind the 15.5x number.","section":"Sec. VII-A and Sec. VII-B"},{"comment":"The end-to-end evaluation omits the very operations that make the architecture distinctive. The ISA relies on LogicMove_Vertical and LogicMove_Horizontal instructions that move entire distance-9 surface-code patches across a junction grid, and transversal CNOTs require aligning patches in shared traps. Section II-C asserts that ion shuttling has negligible decoherence, but Section VII-A's noise model contains only 1Q, 2Q, and measurement/reset error terms; no shuttling time, junction-crossing error, or parallel-gate crosstalk is simulated. If moving a distance-9 patch through a 4-way junction is not effectively noiseless, then the conversion counts, routing schedules, and the Flexion-vs-NISQ comparison all change. Please add a sensitivity analysis with nonzero shuttling/junction errors, or justify quantitatively why these terms can be dropped at the reported accuracy.","section":"Sec. II-C, Sec. V-B, Sec. VII-A"},{"comment":"The SPAM assumption in the noise model is inconsistent with the hardware numbers cited in the paper. Section II-C reports SPAM fidelities exceeding 99.99%, i.e., an error rate around 1e-4, while Section VII-A sets measurement and reset error rates at order 1e-6. The encoding-shrink step in Section IV-A consists of single-qubit measurements, so the SPAM rate directly enters the conversion error pc used throughout the analysis. Please either use SPAM = 1e-4, or explicitly justify the 1e-6 choice and report how the results shift under 1e-4 SPAM.","section":"Sec. II-C and Sec. VII-A"},{"comment":"The claim that conversion-induced logical error is independent of code distance is load-bearing and currently under-supported. Section IV-B states that only a constant number of critical qubit locations cause logical failure, so pc does not decrease with d, and Section IV-C uses a single measured ratio pc/p2 approximately 4 in the analytic comparison. The comparison is not circular, because the threshold pc/p2 < n2/nc is derived and then confirmed by compiler counts, but the numerical value of pc requires broader validation. If distance-independence fails, larger patches reduce conversion error and the trade-off between Flexion and full encoding shifts. Please provide the syndrome-level argument for distance independence in full, or a quantitative decoder simulation across d = 3,5,7,9 and at several physical error rates, rather than relying on Fig. 11 as the sole evidence.","section":"Sec. IV-B and Sec. IV-C"}],"minor_comments":[{"comment":"The legend lists 'iSwitch' alongside 'NISQ-Bare' and 'Ideal', but 'iSwitch' is never defined in the text; please clarify whether it denotes Flexion and distinguish it from the NISQ-Bare curve.","section":"Fig. 9"},{"comment":"The text describes Swap_Intra and Swap_Inter instructions, but Table I and the surrounding text define BareMove_Vertical and BareMove_Horizontal; please reconcile the naming.","section":"Sec. VI-B"},{"comment":"Reference [81] is cited twice in the same sentence, and the Fig. 7 caption says '(1) Mapping surface codes...' while the subfigures are labeled (a)-(d); please fix the numbering.","section":"Sec. II-C"},{"comment":"The ratio pc/p2 approximately 4 is stated without reporting the underlying measured values or simulation conditions; please add the numerical pc and p2 used.","section":"Sec. IV-C"},{"comment":"The text says the ansatz is restricted to Clifford circuits above 10 qubits, yet the benchmarks are described as approximating ground-state energies of non-Clifford Hamiltonians; please state explicitly that the reported energies are for the Clifford-restricted ansatz and may not correspond to the true ground state.","section":"Sec. VII-A"},{"comment":"The derivation of the aggregate 15.5x improvement over standard FTQC is not shown; please define how the per-benchmark ratios are aggregated into a single number.","section":"Sec. VII-B"}],"recommendation":"major_revision","confidential_remarks":"The Clifford-only/MSD-factory mismatch is, in my view, the decisive issue: the headline 15.5x claim should not appear until the baseline is fairly resourced. If the authors can rerun the comparison with an MSD-Logical baseline allowed to use all available qubits for distance-9 encoding, and add a shuttling-error sensitivity study, the paper could become a solid systems contribution. I would not accept it in its current form, but the central idea is worth a major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is real: run 1Q gates on bare ions and selectively encode qubits only for 2Q gates, using an in-situ gauge-fixing conversion protocol and a compiler that schedules conversions. That combination is new, and the runtime encoding design is the strongest part of the paper. The center-injection protocol, the deterministic-stabilizer analysis, and the comparison against state injection and corner enlargement are all reasonable and suggest the authors understand the underlying code-deformation physics. The compiler side is competent too: modeling logical patches as a register file and bare qubits as memory is a sensible abstraction, and the conversion-count results in Fig. 11(c) behave as expected.\n\nWhere the paper falls down is the headline comparison against standard FTQC. Section VII-A says the large-scale benchmarks are restricted to Clifford circuits because they are simulated with Stim. Clifford circuits contain no T gates and no non-Clifford rotations, so a standard FTQC baseline needs no magic-state factory. Yet the MSD-Logical baseline in Sec. VII-B is given a 2,594-qubit T-factory, which forces distance-6 encoding of program qubits while Flexion uses distance-9 patches. That mismatch, not an intrinsic advantage of Flexion, is what produces the 7.16x Heisenberg result and the aggregated 15.5x claim. This is a load-bearing flaw, and it needs to be fixed before the paper can claim superiority over full FTQC.\n\nThe other issues are smaller but real. The evaluation omits shuttling and patch-movement errors entirely, even though LogicMove instructions are central to the architecture; the text cites negligible shuttling decoherence, but moving whole distance-9 patches through junctions is not established to be noiseless at the level required. The SPAM rate in the noise model is 1e-6, which is optimistic relative to the 99.99% (1e-4) experimental numbers the paper itself cites. The NISQ-Bare baseline has no error mitigation, so the 8x-vs-NISQ numbers probably overstate practical gains against a reasonable NISQ competitor. And the 'iSwitch' curve in Fig. 9 is never defined; that is a concrete editorial error that should be fixed. None of these are fatal by themselves, but together they mean the reported magnitudes are upper bounds, not reliable predictions.\n\nWho benefits from reading this? Systems researchers working on early fault tolerance and trapped-ion architectures. The selective-encoding concept and the runtime conversion protocol are worth engaging with, even if the end-to-end evaluation is not yet trustworthy. It deserves a serious referee, but the authors should be told to rerun the large-scale comparison without a magic-state factory for Clifford workloads, add a shuttling error model, and tone down the headline claims to match what the evidence actually supports.","headline":"Flexion has a genuinely interesting selective-QEC idea and a coherent compiler/protocol stack, but the headline 15.5x-vs-FTQC number does not survive scrutiny because the large-scale baseline is charged a magic-state factory it does not need.","tokens_in":23823,"tokens_out":1791,"would_cite":false,"duration_ms":20563,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Selective encoding—bare qubits for 1Q gates, surface-code patches for 2Q gates—cuts early fault-tolerance overhead on trapped-ion systems.","keywords":["selective quantum error correction","hybrid bare-logical encoding","surface codes","trapped-ion QCCD architecture","gauge fixing","early fault tolerance","variational quantum algorithms","magic state distillation avoidance"],"falsifier":"Run the same VQA benchmarks with a noise model that includes per-junction ion-shuttling error and parallel-gate crosstalk, and compare final VQA energies; if moving a distance-9 logical patch through a 4-way junction adds error comparable to a 2Q gate or time comparable to a QEC cycle, the reported 8.0x and 15.5x gaps shrink or disappear.","tokens_in":22744,"feed_emoji":"⚛️","tokens_out":7196,"duration_ms":63883,"temperature":0.7,"pith_summary":"The paper argues that early fault-tolerant quantum computing does not need to encode every qubit all the time. It proposes Flexion, a hybrid scheme for trapped-ion QCCD systems where single-qubit gates run directly on bare physical qubits—whose $\\sim 10^{-6}$ error already meets the EFT target—and two-qubit CNOTs are executed on surface-code logical qubits, whose error is suppressed below $10^{-6}$. A runtime protocol converts any bare qubit state into a logical patch and back in place, using gauge fixing and center injection with conversion error independent of code distance. A hybrid instruction set and compiler decide when to encode and how to route patches. On VQA benchmarks the paper reports 8.0x improvement over NISQ execution and 15.5x improvement over standard full fault tolerance under the same qubit budget, because non-Clifford gates avoid gate synthesis, teleportation, and magic state distillation.","feed_headline":"Ion-trap quantum computing: encode only for two-qubit gates","feed_subtitle":"A hybrid bare/logical scheme claims 8.0x gains over noisy hardware and 15.5x over full error correction for early fault-tolerance workloads.","key_machinery":"The load-bearing object is the in-situ encoding-switch protocol built on gauge fixing: it initializes ancillas around a central bare qubit so that the largest possible number of stabilizers are deterministic (+1), measures the remaining random stabilizers and tracks them as gauge, then runs $d$ rounds of QEC to produce a distance-$d$ logical patch; shrinking measures the ancillas along $X_L$ and $Z_L$ and applies a single-qubit correction. This conversion makes the hybrid ISA possible, with Encode_Boundary and Shrink_Boundary restricted to a boundary region, LogicMove instructions shuttling entire patches across the 2D grid, and transversal physical CNOTs between co-located patches. The compiler treats logical patches as a register file and bare qubits as memory, using a greedy linear-scan allocation for encoding conversions and a SABRE-style routing pass for patch movement, minimizing conversion count because each conversion costs roughly $4\\times$ a 2Q gate error.","core_discovery":"Flexion's central claim is that selective QEC—bare qubits for 1Q gates, logical patches for 2Q gates—is enough to meet early fault tolerance LER targets while eliminating the three dominant overheads of standard FTQC: full encoding, T-gate synthesis, and magic state distillation. The paper's runtime encoding switch treats a bare qubit as a degenerate encoding and grows a full surface-code patch around it by initializing ancillas in a gauge-aware pattern, measuring stabilizers, and running $d$ QEC cycles; the reverse shrinks the patch by measuring ancillas along the logical operators. The protocol is designed so that conversion-induced logical error is independent of code distance, set by a few critical qubit locations, and the paper estimates conversion error $p_c \\approx 4 \\, p_2$ (about $4.3 \\times 10^{-3}$ at a 2Q error of $10^{-3}$), giving a net gain whenever the CNOT count exceeds the conversion count by more than about 4. Flexion then claims 8.0x energy-gap improvement over bare NISQ execution and 15.5x improvement over an MSD-based fully logical baseline under equal qubit budgets.","pith_inferences":["If shuttling noise is added to the evaluation model, the optimal encoding ratio could drop and LogicMove cost would become a first-class objective; a natural test is to recompute the 8.0x and 15.5x numbers with per-junction error rates from ion-transport measurements.","The same bare/logical split should transfer to other platforms with very high 1Q fidelity and modular 2D connectivity, such as neutral-atom arrays, as long as whole logical patches can be moved and aligned; the conversion protocol itself is platform-neutral.","The gauge-fixing switch could be generalized from bare-to-logical to distance $d_1$-to-$d_2$ switching, letting a program raise protection only for critical subcircuits—an adaptive-QEC knob the paper does not explore.","Because only about four qubit locations are critical during a switch, choosing which qubit carries the bare state could be folded into scheduling as a first-class fidelity decision, reducing conversion error further than the compiler currently does."],"forward_implications":["On circuits whose two-qubit gate count exceeds the conversion count by roughly $4\\times$ or more, Flexion's selective encoding beats fully bare NISQ execution; VQE and fermionic simulation circuits satisfy this condition.","Because non-Clifford 1Q rotations execute directly on bare qubits, Flexion avoids the Clifford+T depth blowup (measured average 14.9x) and the idling and memory error of waiting for distilled magic states.","Under an equal physical qubit budget, a full-FT baseline must reserve thousands of qubits for T factories, forcing lower-distance encoding for program qubits; Flexion instead uses the budget for program qubits, yielding lower logical error and better VQA energy.","Conversion error is independent of surface-code distance, so raising $d$ to lower the logical error rate does not make switches more expensive, only the QEC rounds after conversion.","The scheme targets the megaquop regime directly: roughly $20$–$50$ logical qubits and $10^4$–$10^6$ gates, where full FTQC overhead is prohibitive but selective QEC fits within thousands of physical qubits."],"supporting_citations":[{"why":"Supplies the gauge-fixing theory that the runtime encoding switch builds on for in-situ bare-to-logical conversion.","marker":"[54]"},{"why":"Defines standard full-FT surface-code overheads and the magic state distillation cost baseline used for comparison.","marker":"[25]"},{"why":"Supplies the (15-to-1)13,5,5 T-factory size and latency used in the MSD-Logical baseline and resource comparison.","marker":"[77]"},{"why":"Supplies the Clifford+T decomposition cost for RZ rotations that Flexion avoids by executing 1Q gates on bare qubits.","marker":"[40]"},{"why":"Provides the ion-shuttling QCCD architecture and all-to-all connectivity that make transversal logical CNOTs possible.","marker":"[49]"},{"why":"Provides the approximately $10^{-6}$ single-qubit gate fidelity that justifies executing 1Q gates on bare qubits.","marker":"[48]"},{"why":"Supplies the routing strategy adapted by the Flexion compiler for co-locating logical patches.","marker":"[97]"},{"why":"Supplies the stabilizer simulator used to estimate logical error rates in the evaluation.","marker":"[104]"},{"why":"Supplies the minimum-weight perfect matching decoder used for syndrome decoding in the evaluation.","marker":"[72]"}],"fun_headline_variants":["Selective QEC: encode only for 2-qubit gates on ion traps","Flexion: adaptive encoding cuts quantum error correction overhead","Ion-trap QEC made practical: encode only when needed","Hybrid bare/logical qubits for early fault-tolerant ion traps","Flexion: on-demand encoding slashes FTQC overhead"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The architecture assumes whole logical surface-code patches can be shuttled across the QCCD junction grid and aligned for transversal CNOTs with negligible additional error and time, while the evaluation noise model omits shuttling, movement, and parallel-gate crosstalk.","fun_headline_variants_meta":{"raw":{"variants":["Selective QEC: encode only for 2-qubit gates on ion traps","Flexion: adaptive encoding cuts quantum error correction overhead","Ion-trap QEC made practical: encode only when needed","Hybrid bare/logical qubits for early fault-tolerant ion traps","Flexion: on-demand encoding slashes FTQC overhead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001571,"raw_usage":{"total_tokens":6339,"prompt_tokens":1080,"completion_tokens":5259,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":696,"completion_tokens_details":{"reasoning_tokens":5170}},"tokens_in":696,"tokens_out":5259,"duration_ms":36384,"temperature":1.0,"reasoning_tokens":5170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:07:40.339784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same VQA benchmarks with a noise model that includes per-junction ion-shuttling error and parallel-gate crosstalk, and compare final VQA energies; if moving a distance-9 logical patch through a 4-way junction adds error comparable to a 2Q gate or time comparable to a QEC cycle, the reported 8.0x and 15.5x gaps shrink or disappear.","supporting_citations":[{"cited_title":"Code deformation and lattice surgery are gauge fixing","cited_arxiv_id":null,"evidence_quote":"Supplies the gauge-fixing theory that the runtime encoding switch builds on for in-situ bare-to-logical conversion."},{"cited_title":"High-fidelity preparation, gates, memory, and readout of a trapped-ion quantum bit","cited_arxiv_id":null,"evidence_quote":"Provides the approximately $10^{-6}$ single-qubit gate fidelity that justifies executing 1Q gates on bare qubits."},{"cited_title":"Tackling the qubit mapping prob- lem for nisq-era quantum devices","cited_arxiv_id":null,"evidence_quote":"Supplies the routing strategy adapted by the Flexion compiler for co-locating logical patches."}],"review_version":1}