{"id":"c3f19ff2-aa1f-4ade-b795-0a2e4bf1e135","arxiv_id":"2412.12529","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A linear-time CCZ gate can be implemented on a looped pipeline 2D architecture, but resource estimates show it is currently about two times slower and larger than magic state distillation.","lead":"This paper maps a known fault-tolerant non-Clifford gate for surface codes onto a shuttling-based 2D chip design, and works out its resource cost. It finds the gate currently loses to magic state distillation, mostly because of a bottleneck decoder, and identifies that decoder as the key thing to improve.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The factor-of-two overhead vs distillation rests on d_ccz ~ 100, extrapolated from five phenomenological single-code points on a different slice geometry; the full three-code circuit-level distance is unverified.","rationale":"The paper's novel contribution is the pipelined schedule and the in-place code transformations; those are described in enough detail that no obvious internal inconsistency emerges from the text. The strongest claim as framed by the reader, however, is quantitative: linear-time CCZ costs about twice as much as magic state distillation at P_ccz ~ 3e-10. That number is proportional to d_ccz, and the only numerical input for d_ccz is a single-code phenomenological extrapolation that the authors themselves call unreliable. Circuit-level noise and three-code error-spreading from the physical CCZ can only be assessed by simulating the actual protocol; without that, the exact factor is not established. This does not undermine the architectural feasibility claim, so the verdict should remain conditional rather than rejected or accepted. The concrete test above would either validate the factor or force a re-statement as a bound.","tokens_in":14274,"tokens_out":15423,"duration_ms":140325,"concrete_test":"Run a circuit-level simulation of the exact three-code, two-layer CCZ protocol on the proposed shuttling schedule for small distances (L=3,5,7), using the JIT decoder on the implemented syndrome sequence, including physical CCZ compiled as in Sec. III C. Compare the logical CCZ error rate and threshold to the phenomenological single-code curve in Fig. 12. If the circuit-level failure rate at L=5,7 is not within, say, an order of magnitude of the phenomenological prediction, re-fit d_ccz; if the extrapolated distance differs by more than 20%, replace the 'about twice' claim with an inequality and state the direction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative comparison (Sec. IV) is governed by d_ccz ~ 100, taken from Fig. 12: five Monte Carlo points for L in (30,40] for a single red code under phenomenological noise, with a linear fit and no reported error bars. The authors explicitly call this estimate unreliable and note that the simulated slices in Ref. [31] are not the two-layer slices used in their construction (Sec. IV B). The full CCZ involves three codes and a physical CCZ that maps single-qubit X errors to correlated XZZ errors across codes; no simulation accounts for this spread or for circuit-level noise from shuttling, ancilla preparation, and measurement in the proposed schedule. The paper also only 'expects' the protocol to be fault-tolerant from the JIT decoder proof in Ref. [6] (Sec. III B). If the true distance is 150 rather than 100, the time overhead T'_lin = 7 d_ccz rises from 700 to 1050 cycles and distillation is ~3x faster, not 2x; if the distance were much smaller, the claimed 2x disadvantage could disappear. Either way, the headline factor is not settled by the data presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a concrete implementation of Brown's linear-time CCZ gate (Ref. [6]) on a looped pipeline architecture, giving a detailed shuttling schedule, a mapping of the three surface-code layers and slices onto shuttling loops, and code-deformation procedures for converting standard rotated surface-code patches into the required lattices. It then compares the space-time cost of this in-place CCZ gate with the cost of magic state distillation for a large-scale factoring scenario (2048-bit RSA, N=6000 logical qubits), concluding that distillation is currently faster by roughly a factor of two and that the gap is dominated by the performance of the just-in-time decoder. The paper is explicitly framed as a resource study rather than a claim of practical advantage for the non-distillation approach.","tokens_in":14485,"tokens_out":5435,"duration_ms":48825,"significance":"If the construction is correct, the manuscript provides a concrete architectural blueprint for fault-tolerant non-Clifford gates on a 2D device without magic state distillation, using only short-range shuttling, local gates, and measurements. The detailed mapping of the two-layer slices and the three-step stabiliser measurement schedule is a useful technical contribution, and the paper is commendably transparent about the limitations of its numerical extrapolation. The data underlying Fig. 12 are made available, and the simulations use publicly available code. The negative resource comparison is itself useful as a benchmark and identifies the JIT decoder as the main bottleneck. However, the quantitative factor-of-two speed advantage claimed for distillation is load-bearing and rests on an extrapolation that the authors themselves describe as unreliable, which limits the strength of the headline comparison.","major_comments":[{"comment":"The quantitative resource comparison (T'_lin = 7 d_ccz = 700 code cycles vs T_msd ≈ 330 code cycles, giving a factor-of-two speed advantage for magic state distillation) depends critically on the estimate d_ccz ~ 100. This value is obtained by a linear fit to five Monte Carlo points (L in the range 30 to 40) for a single red code under phenomenological noise, not for the full three-code CCZ gate or for the proposed shuttling schedule. The authors themselves state that this estimate is 'unreliable' and that the simulated slices in Ref. [31] are not the two-layer slices used in the construction. No circuit-level simulation of the full protocol is provided. If the true distance is larger, the distillation advantage grows; if it is smaller (e.g., d_ccz = 50), T'_lin becomes comparable to T_msd and the claimed factor-of-two disadvantage disappears. Since this factor is a central result of the paper, the estimate needs either a substantial strengthening (e.g., circuit-level or at least three-code simulations) or the quantitative claim should be removed and the comparison stated only qualitatively as 'currently in favour of distillation'.","section":"Sec. IV B, Fig. 12"},{"comment":"The paper asserts in the introduction and Sec. III B that 'we thus expect the entire protocol to be fault-tolerant' based on the fault-tolerance proof for the JIT decoder in Ref. [6]. However, the specific implementation here differs from Ref. [6] in several ways: it uses two-layer slices of a different geometry, a three-step stabiliser measurement with a half-cycle advancement of inter-layer ancilla qubits, a compiled physical CCZ gate (six CNOTs), and the looped-pipeline shuttling schedule itself. No argument is given that the JIT decoding proof extends to these modifications, and no simulation covers the correlated XZZ errors that arise when an X error on one code passes through the physical CCZ to the other codes. The resource estimate implicitly assumes that the logical error rate of the full three-code gate is the same as that of the single-code phenomenological simulations. Please either provide a fault-tolerance argument for the exact schedule or explicitly state this as an additional unverified assumption and discuss how it would affect the resource comparison.","section":"Sec. III B and Sec. III C"}],"minor_comments":[{"comment":"There is a repeated typo 'looped pipleline' in the introduction and Sec. II B; this should read 'looped pipeline'.","section":"Introduction and Sec. II B"},{"comment":"The expression T_msd = 230 + 2d + 3d with d = 21 gives 335 code cycles, not 330; please correct the arithmetic or explicitly round to 335.","section":"Sec. IV A"},{"comment":"In the example with target Pccz = 10^-7, the value d_ccz = 50 is introduced without explaining how it is obtained from Fig. 12; please either show the extrapolation or label this as an assumption.","section":"Sec. V"},{"comment":"In the caption of Fig. 2, 'pipeline /f_low direction' appears to be a LaTeX error and should read 'pipeline/flow direction'. Also, Ref. [32] has a doubled 'https://' in its URL.","section":"Fig. 2 caption and Ref. [32]"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal and the architectural construction is valuable and clearly presented. The main issue is that the quantitative factor-of-two comparison to magic state distillation relies on the extrapolated d_ccz ~ 100, which the authors themselves call unreliable. I would encourage the editor to request either additional simulations (even at the single-code circuit level, but ideally for the full three-code protocol) or a rewording of the abstract and Sec. IV to avoid presenting the factor of two as a solid quantitative result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper is genuinely useful: it shows how Brown's linear-time CCZ can be implemented on a looped pipeline with a shuttling schedule only modestly more complex than the standard surface code, and it specifies the code deformations and the half-cycle ancilla advancement trick in detail. That part is new and looks careful. The resource comparison is also honest in an unusual way: the authors make deliberately favorable assumptions for the linear CCZ — ideal physical CCZ error rate, ignoring routing, and so on — and still find magic state distillation about twice as fast in time at the target error rate. They also identify the JIT decoder as the bottleneck and share their data.\n\nThe soft spot is exactly where the reader's report puts it. The d_ccz ~ 100 used for the headline factor is an extrapolation from five Monte Carlo points for a single red code under phenomenological noise, on slices that differ from the ones used here. The authors call it unreliable themselves and treat it as a lower bound. So the factor-of-two is not a settled number. If the true distance is larger, distillation wins by more; if smaller, the gap could shrink or disappear. The qualitative conclusion — that distillation beats this approach in the practical regime — is probably robust, because the favorable assumptions carry a lot of weight. Separately, the fault-tolerance of the full protocol is expected from Brown's proof rather than proven for this specific schedule; a referee should ask for that to be tightened.\n\nNone of this undermines the core contribution. The mapping stands on its own as a feasibility result, and the negative comparison is a valuable data point. It deserves a serious referee. My recommendation: send it to peer review, with the request that the authors either strengthen the distance estimate or clearly label the quantitative comparison as provisional.","headline":"A careful and honest architecture mapping for Brown's linear-time CCZ on a looped pipeline, with a resource comparison that favors distillation but rests on an extrapolated distance estimate.","tokens_in":15040,"tokens_out":2382,"would_cite":true,"duration_ms":22814,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that a looped pipeline architecture can implement a fault-tolerant non-Clifford CCZ gate on a 2D device, at a current time cost about twice that of magic-state distillation.","keywords":["looped pipeline architecture","surface codes","non-Clifford gate","CCZ gate","just-in-time decoding","magic state distillation","quantum error correction","qubit shuttling"],"falsifier":"A full circuit-level Monte Carlo simulation of the complete three-slice linear-time CCZ protocol at distances 50 through 100, measuring the logical CCZ error rate at $p = 5 \\times 10^{-4}$, would settle whether the distance needed for $P_{\\mathrm{ccz}} \\sim 3 \\times 10^{-10}$ is near 100 or far below it.","tokens_in":14041,"feed_emoji":"🔄","tokens_out":7360,"duration_ms":63824,"temperature":0.7,"pith_summary":"This paper establishes that a looped pipeline architecture—physical qubits shuttled in synchronised loops, with several qubits per loop forming stacked arrays—can implement the fault-tolerant, linear-time CCZ gate between three 2D surface codes. The shuttling schedule needed is only slightly more involved than the one already used for a standard surface code in the same architecture. The paper then compares this no-distillation route to magic-state distillation for the same architecture and finds that, at current decoder performance, distillation is about twice as fast in time at the target error rate $P_{\\mathrm{ccz}} \\sim 3 \\times 10^{-10}$, with comparable or better space use. The bottleneck is the just-in-time decoder, which forces very large code distances; the paper argues that this decoder has been little studied and could improve substantially.","feed_headline":"A 2D shuttling chip can run a no-distillation quantum gate","feed_subtitle":"The gate runs with short-range shuttling alone, yet magic-state factories still win by about 2x in time.","key_machinery":"The load-bearing object is the looped pipeline architecture: qubits circulate clockwise around fixed shuttling loops in synchrony, and placing multiple qubits per loop stacks independent layers of a code without extra hardware. Inter-loop interactions realise intra-layer stabiliser checks, while intra-loop interactions realise transversal gates between layers, giving effective 3D connectivity on a 2D chip. The second ingredient is the linear-time CCZ gate, which reorders the transversal CCZ of three 3D surface codes so each qubit experiences the same sequence—initialise, measure $Z$ stabilisers, apply physical CCZ, measure out—but slices of the 3D codes are processed one by one in time. Fault tolerance is supplied by just-in-time decoding, which guesses the syndrome of the corresponding 3D code at each timestep; the paper's architecture adds a three-step slice cycle and code-deformation stages for in-place operation.","core_discovery":"The central claim is that a strictly 2D device with short-range shuttling loops can run the linear-time CCZ gate of Ref. [6] using the same physical operations as multiple planar surface codes: short-range shuttling on fixed paths, single- and two-qubit gates, and single-qubit Pauli measurements. The implementation maps three different 2D layer codes (a standard square-lattice surface code and two mirrored kagome-lattice codes) and their two-layer 3D slices onto loops of a single grid, with intra-loop interactions supplying transversal operations between layers and inter-loop interactions supplying stabiliser measurements within layers. A three-step slice cycle (two full loop cycles plus a half-cycle ancilla shift) measures all needed $Z$ stabilisers. Resource estimates for a 2048-bit factoring workload put the in-place gate at roughly 700 code cycles and a $5d \\times 10d$ footprint, versus about 330 cycles for a CCZ distillation factory placed in the same corridor, so magic-state distillation is about twice as fast at $P_{\\mathrm{ccz}} \\sim 3 \\times 10^{-10}$; the gap narrows to about 1.5x at $P_{\\mathrm{ccz}} \\sim 10^{-7}$. The paper attributes the overhead almost entirely to the low threshold and poor sub-threshold scaling of the just-in-time decoder, not to the gate itself, and argues that improved decoders—or hybrid post-selection schemes—could change this balance.","pith_inferences":["A circuit-level simulation of the full three-slice protocol would be the natural next test: if an optimised just-in-time decoder approaches standard surface-code thresholds, the factor-of-two gap likely closes or reverses.","The same intra-loop transversal interaction mechanism could be used for other transversal operations among stacked codes, potentially reducing the cost of Clifford factories or enabling new code-deformation routines beyond CCZ.","The architecture's value may lie less in replacing distillation today than in providing a fallback universal gate set for platforms where shuttling is cheap and magic-state routing is expensive; the paper's own numbers quantify that trade-off."],"forward_implications":["A looped pipeline device can host a fault-tolerant non-Clifford gate without changing the physical operation set of the surface code, only the shuttling schedule and some local constant-depth circuits.","At the target error rate for 2048-bit factoring, the in-place linear-time CCZ costs about 700 code cycles and a $5d \\times 10d$ patch, while a CCZ distillation factory in the same space produces and teleports a gate in about 330 cycles.","Using linear-time CCZ to build factories instead of applying it in place gives a space overhead about 1.4 times larger and a time overhead about twice as large as distillation factories.","If the target CCZ error rate is relaxed to $10^{-7}$, the time gap shrinks to about 1.5x, and the space freed by the gate no longer accommodates a factory per three qubits.","Over asymptotically small error rates, the linear-time gate's exponential suppression for linear distance growth beats the exponential trial overhead of post-selection-based distillation, though not in the practical regime considered here."],"supporting_citations":[{"why":"Supplies the linear-time CCZ gate construction and its fault-tolerance proof, which the paper implements.","marker":"[6]"},{"why":"Supplies the looped pipeline architecture and its shuttling schedule for the standard surface code.","marker":"[11]"},{"why":"Supplies the numerical just-in-time decoding data and the code used for the distance extrapolation.","marker":"[12]"},{"why":"Provides the simulation code used to generate the new Monte Carlo points for the distance estimate.","marker":"[31]"},{"why":"Supplies the CCZ distillation factory design and its footprint and timing used as the comparison baseline.","marker":"[30]"},{"why":"Supplies the 2048-bit factoring resource targets that set the target logical error rates used in the comparison.","marker":"[28]"},{"why":"Supplies the formula used to choose the surface-code distance for the target logical cycle error rate.","marker":"[29]"},{"why":"Supplies the 3D surface-code slices and transversal CCZ structure that the linear-time gate reorders.","marker":"[7]"},{"why":"Introduces just-in-time decoding, the strategy whose performance drives the resource comparison.","marker":"[8]"}],"fun_headline_variants":["2D shuttling chip runs CCZ gate without distillation","No magic-state factory: 2D shuttling does CCZ","Distillation still faster, but 2D shuttling can do CCZ","CCZ gate on 2D shuttling chip: no distillation needed","2D shuttling enables no-distillation CCZ, but distillation still wins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The resource comparison depends on the extrapolated estimate that the three-code gate needs a code distance near 100; if circuit-level noise demands much less (or much more), the factor-of-two time gap shrinks or grows accordingly.","fun_headline_variants_meta":{"raw":{"variants":["2D shuttling chip runs CCZ gate without distillation","No magic-state factory: 2D shuttling does CCZ","Distillation still faster, but 2D shuttling can do CCZ","CCZ gate on 2D shuttling chip: no distillation needed","2D shuttling enables no-distillation CCZ, but distillation still wins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001625,"raw_usage":{"total_tokens":6513,"prompt_tokens":1045,"completion_tokens":5468,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":5369}},"tokens_in":661,"tokens_out":5468,"duration_ms":33336,"temperature":1.0,"reasoning_tokens":5369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:58:47.938647+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A full circuit-level Monte Carlo simulation of the complete three-slice linear-time CCZ protocol at distances 50 through 100, measuring the logical CCZ error rate at $p = 5 \\times 10^{-4}$, would settle whether the distance needed for $P_{\\mathrm{ccz}} \\sim 3 \\times 10^{-10}$ is near 100 or far below it.","supporting_citations":[{"cited_title":"The full CCZ requires the codes on the left and right to switch places, so each code needs to travel (2 dccz − 1) + 1 = 2dccz data qubit loops to the left or right","cited_arxiv_id":null,"evidence_quote":"Supplies the linear-time CCZ gate construction and its fault-tolerance proof, which the paper implements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the looped pipeline architecture and its shuttling schedule for the standard surface code."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the numerical just-in-time decoding data and the code used for the distance extrapolation."},{"cited_title":"Goel and J","cited_arxiv_id":null,"evidence_quote":"Supplies the CCZ distillation factory design and its footprint and timing used as the comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 2048-bit factoring resource targets that set the target logical error rates used in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the formula used to choose the surface-code distance for the target logical cycle error rate."},{"cited_title":"Gottesman, The heisenberg representation of quan- tum computers, arXiv (1998)","cited_arxiv_id":null,"evidence_quote":"Introduces just-in-time decoding, the strategy whose performance drives the resource comparison."}],"review_version":1}