{"id":"a95edc9e-700e-4655-bdd5-c45e0b651bc7","arxiv_id":"2411.17519","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A sparse second qubit layer (Bypass) shortens lattice surgery paths, reducing decoding bottlenecks and enabling a 2.5D FTQC architecture that is faster and uses fewer resources in simulation.","lead":"This paper proposes a 2.5D qubit layout, called the Bypass architecture, that shortens the paths used by lattice surgery operations in fault-tolerant quantum computers. In simulations of quantum phase estimation programs, it reports up to 1.73x speedup and 17% fewer hardware resources than a conventional 2D layout.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No single decisive flaw, but the d=23 resource claim rests on an unchecked proportionality between total program LER and total effective path length; run the stated FP-based test.","rationale":"I read the paper in good faith and agree with the reader that the Bypass architecture is a plausible, well-articulated proposal: the geometric path-length argument (Sec 5.2.2) is sound, the CBPI stack is a useful bottleneck-analysis methodology, and the qualitative claim that shorter effective paths reduce decoding and path penalties is supported by the simulation evidence. I also agree that the two fragile assumptions are the d=23 projection and the hardware feasibility of inter-layer CNOTs. However, my stress-test role asks for the single most load-bearing concern. I judge the d=23 scaling step to be the most load-bearing, because the headline 17% resource reduction is directly derived from it, while the inter-layer CNOT feasibility concern is partially mitigated by the paper's explicit sensitivity analysis (p' = 1p, 5p, 10p in Sec 6.1) and its honest caveat that performance is only claimed for the optimistic/practical range. I do not see an internal inconsistency in the path-length argument; the issue is a gap between the aggregate L' statistics and the assumed program-level LER scaling. The reader's weakest_assumption identifies the same LER-scaling concern (partial agreement), but my emphasis differs: I would not reject the d=23 number outright, only require a demonstration that the L' -> p_L -> d chain holds for a full program, including factory decoding. Because the paper is a simulation-based architecture proposal and the concern is a validation gap rather than a demonstrated error, the appropriate verdict remains CONDITIONAL (no change from the reader's verdict).","tokens_in":22265,"tokens_out":1689,"duration_ms":15989,"concrete_test":"Run the FTQC program end-to-end at both d=25 and d=23 for the chosen FH(200) benchmark, using the paper's cycle-accurate simulator augmented with a first-order LER estimate: track the per-instruction LER (weighted by measured path and duration) and compute the program-level logical failure probability p_L, then check whether scaling d from 25 to 23 changes p_L by roughly the factor predicted by the Sec 7.3 formula and whether the claimed CBPI values are unchanged. If p_L is not reduced by the predicted factor, or if CBPI shifts, the 17% resource reduction and 1.73x speedup are not supported by the stated model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (1.73x speedup, 17% resource reduction) requires the step in Sec 7.3 where total program LER is assumed proportional to total effective path length L', so a 1/X reduction in L' reduces required code distance by 2. This is the weakest load-bearing link: the paper presents aggregate L' histograms (Fig 13) but never verifies that aggregate-L' proportionality holds for the actual program, with its mixture of operation types, magic-state-dependent gates, and decoding tasks. The formula p_L ~ const x (p/p_th)^((d-1)/2) is for a single logical operation of a given size; the conversion from 'total path length in arbitrary cell units' to a required d for the whole program is asserted, not derived. The excluded factory decoding (Sec 7.2) and the assumption that CBPI remains unchanged when d is reduced from 25 to 23 (Fig 17, purple stars) mean the 17% resource reduction and 1.73x speedup are only as solid as this scaling extrapolation. The second fragile premise is Bypass-cell decoding weight 1/d (Sec 7.2), used for all decoding-penalty conclusions; it is a modeling choice, not a measured decoder cost.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the CBPI stack, a bottleneck-analysis methodology for lattice-surgery-based fault-tolerant quantum computing, and proposes the Bypass architecture, a 2.5D layout with a dense Logic layer and a sparse Bypass layer. The authors define an effective path length L', argue that the Bypass layer reduces L' by a factor of the code distance d, and evaluate the architecture with circuit-level Stim/PyMatching simulations for MEAS_ZZ operations and a cycle-accurate LS simulator on quantum phase estimation benchmarks. The central quantitative claims are a 1.73x speedup and a 17% reduction in quantum/classical hardware resources over a conventional 2D layout in the moderate-resources case, obtained after reducing the code distance from 25 to 23 on the basis of the reduced total effective path length.","tokens_in":22502,"tokens_out":6726,"duration_ms":64216,"significance":"If the central claims hold, the Bypass architecture is a plausible route to higher-density and lower-overhead lattice-surgery computation, and the CBPI stack is a useful organizing framework for architecture-level FTQC performance analysis. The paper has several concrete strengths: circuit-level LER simulations with Stim and PyMatching, systematic sweeps over nF, TPdec, Rdata, and problem size, 1000 random qubit assignments with reported standard deviations, and a clean proof in Appendix A that 50% is the maximum data-cell ratio for immediate-operation-capable arrangements. The central resource-reduction claim, however, rests on an extrapolation from total effective path length to program-wide logical error rate that is asserted rather than demonstrated; this is the main load-bearing point that needs additional validation.","major_comments":[{"comment":"The reduction from d=25 to d=23 is not supported by a calculated or simulated program-wide logical error rate. The text cites p_L ≈ const × (p/p_th)^((d-1)/2) for a single logical operation and then asserts that the LER of the entire FTQC program is proportional to the total effective path length of all LS operations; no derivation is given, and the aggregate L' histograms in Fig. 13 are not a program-LER computation. This assumption must be checked by simulating the full QPE program (or a representative subset, including magic-state teleportation and the factory decoding that is currently excluded in Sec. 7.2) at d=25 and d=23, or by deriving a bound from the per-operation LER curves in Fig. 10. Until this check is done, the 17% resource reduction and 1.73x speedup in Sec. 7.5 rest on an unverified scaling hypothesis.","section":"Sec. 7.3 (\"LS path length and program fidelity\")"},{"comment":"The conclusion that the Bypass architecture eliminates or reduces the decoding penalty depends on the assumption that \"the decoding task difficulty for a single cell in the Bypass layer is 1/d of that for a single cell in the Logic layer.\" This is a modeling choice, not a measured or simulated decoder cost. The inter-layer stabilizers have 6-degree connectivity and a different syndrome-graph geometry, so the 1/d factor needs justification with decoder-level data (e.g., matching-graph size or PyMatching runtime) or at least a sensitivity analysis. Since the CBPI results in Figs. 14-17 are directly affected by the decoding-penalty term, this assumption is load-bearing for the reported speedups.","section":"Sec. 7.2 (decoding model)"},{"comment":"The LER advantage of the Bypass layer is shown only for MEAS_ZZ operations above a crossover path length; for the pessimistic p'=10p case, Fig. 10(d) places the crossover at L≥31. The paper does not report the distribution of actual LS path lengths L for the benchmark programs, only the effective path length L' in Fig. 13, so it is not established that a sufficiently large fraction of operations in FH(200) lie above this crossover for the aggregate fidelity improvement to hold. Please report the program-level distribution of L and compute the total program LER under p'=5p and p'=10p, rather than relying on the MEAS_ZZ-only comparison.","section":"Sec. 6.2 and Sec. 5.2.2"},{"comment":"The purple-star curves for d=23 appear to be obtained by rescaling the resource axis of the d=25 cycle-accurate simulations, but the simulation model couples code distance to decoder throughput (TPdec per cell), decoding-task difficulty, and the Bypass-layer 1/d factor, so it is not obvious that CBPI remains unchanged when d changes from 25 to 23. The authors should rerun the LS simulation at d=23 for the reported configurations, or explicitly argue why CBPI is invariant under this change; otherwise the resource/performance trade-off curves for d=23 are extrapolations.","section":"Sec. 7.5 (Fig. 17)"}],"minor_comments":[{"comment":"The definition of L' as \"the number of data qubits involved in a given LS instruction divided by d^2\" would benefit from a sentence explaining why this dimensionless quantity is called an effective path length; as written it can be mistaken for an area or volume measure.","section":"Sec. 5.2.2"},{"comment":"The legend values \"Acc.\" and \"Ave.\" should be defined in the caption; \"Acc.\" appears to denote the total accumulated L' over the program, but this is not stated.","section":"Fig. 13 caption and legend"},{"comment":"The main text says factories generate magic states at \"regular intervals of several code beats,\" while Sec. 7.2 specifies 15 code beats per MSD circuit; please make the main-text description consistent with the simulation parameter.","section":"Sec. 3.2.2 and Sec. 7.2"},{"comment":"The relationship between TPdec and per-cell decoding capacity is described verbally as \"0.5 (1.0) ... half (all) of the cells\"; stating this as an explicit equation would improve reproducibility of the cycle-accurate simulator.","section":"Sec. 7.2"},{"comment":"No artifact availability statement is given for the Stim circuits or the cycle-accurate LS simulator; releasing code and data would substantially strengthen verification of the d=23 extrapolation and the CBPI results.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a quantum-architecture or FTQC-systems venue, and the architecture idea is timely. The main risk is not the architecture itself but the strength of the headline resource numbers: the d=23 claim depends on a program-LER scaling assumption that is plausible but unchecked. I would recommend asking the authors to provide the missing full-program LER check or to rephrase the headline claims as projections under an explicitly stated scaling assumption. The Bypass-layer decoding-cost assumption (1/d) is a second point where a small amount of additional data would materially increase confidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look. The CBPI stack is a practical decomposition tool for FTQC bottleneck analysis, and the Bypass layer is a real geometric idea: using a sparse logic layer to shorten effective lattice-surgery path length by O(1/d) and to let conflicting paths cross. The new 50% immediate-operation-capable arrangement with the Appendix A density proof is also a clean, small contribution. Credit where due: they simulate circuit-level LERs with Stim and PyMatching, they sweep key parameters (p', TPdec, nF, Rdata), and they report standard deviations. They also stay honest about hardware-readiness, citing flip-chip references rather than claiming a fabricated device.\n\nThe soft spots are real but not fatal. The main load-bearing issue is Section 7.3: the 17% resource reduction and the 1.73x speedup at d=23 assume that total program LER is proportional to total effective path length L'. They present histograms of L' and assert the proportionality, but they never test it against the actual program with its mixed operation types, magic-state-dependent branches, and decoding tasks. The p_L ~ (p/p_th)^((d-1)/2) formula is for a single logical operation; converting from aggregated cell-level path length to a program-wide required distance is a jump. This can be fixed: run the stated fault-path test or a direct simulation with d=25 and d=23 on the same program. The second fragile spot is the Bypass-cell decoding weight of 1/d (Section 7.2), a modeling choice that underlies all decoding-penalty conclusions. Minor but worth noting: no code or data is released, and factory decoding is excluded from the model (the paper says so explicitly).\n\nI do not think the stress-test note overstates the concern. On reading, the proportionality in Sec 7.3 is asserted, not derived, and the appendix proof does not rescue it. But the qualitative claims are solid: shorter paths reduce path and decoding penalties, and the Bypass layer makes those penalties vanish in simulation. The central architecture idea is sound.\n\nThis paper is for the FTQC architecture and compilation community, plus experimentalists thinking about 3D integration. It deserves a serious referee. I would recommend conditional acceptance: ask the authors to verify the aggregate-L' scaling assumption or retract the d=23 quantitative claims, and to release simulator code and data. A clear yes on engagement.","headline":"A genuinely useful architecture paper with a clever sparse Bypass layer, but the headline resource numbers rest on an unverified LER-proportionality step; conditional accept.","tokens_in":23130,"tokens_out":1399,"would_cite":true,"duration_ms":15208,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Lx"],"model":"deepseek-v4-flash","headline":"A sparse 2.5D 'Bypass' layer shortens the effective lattice-surgery path length by a factor of $d$, removing the path-conflict and decoding bottlenecks that limit dense 2D fault-tolerant quantum computers.","keywords":["fault-tolerant quantum computing","lattice surgery","surface code","2.5D architecture","logical error rate","CBPI stack","quantum phase estimation","magic state factory"],"falsifier":"Measure, by circuit-level simulation of a full QPE program at a fixed code distance, whether the total program logical error rate is indeed proportional to the total effective path length $L'$; if cutting $L'$ by a factor of about 2 does not cut the logical error rate correspondingly, the code-distance reduction from 25 to 23 is unsupported. Alternatively, fabricate or measure an inter-layer CNOT in a flip-chip device: if its error rate exceeds about $10p$, the Bypass advantage for short lattice-surgery paths disappears, and the crossover length for benefit rises above the $L=31$ threshold reported in the paper.","tokens_in":22026,"feed_emoji":"🔀","tokens_out":8428,"duration_ms":75144,"temperature":0.7,"pith_summary":"This paper sets out to show that the main performance limiters in lattice-surgery-based fault-tolerant quantum computing—conflicting paths between operations and the decoding load of long paths—can be removed by a 2.5-dimensional layout it calls the Bypass architecture. The Bypass layout adds a sparse qubit layer beneath the dense logic layer, so distant logical qubits can be merged along short, programmably chosen routes instead of roundabout 2D paths. Because the effective path length $L'$ of a lattice-surgery operation drops by a factor of the code distance $d$, the logical error rate of the whole program falls and the code distance can be reduced. In simulations of practical quantum phase estimation programs, the proposed layout achieves a 1.73x speedup and a 17% reduction in quantum and classical hardware compared with a conventional 2D layout in the moderate-resource case. The paper also introduces the CBPI stack, a hazard-decomposition analysis that lets architects see which penalty—magic-state supply, path conflicts, or decoding—dominates execution time.","feed_headline":"A sparse second layer cuts quantum surgery paths by factor d","feed_subtitle":"The 2.5D 'Bypass' layout also removes path conflicts, giving a 1.73x speedup with 17% fewer qubits.","key_machinery":"The carrying mechanism is the effective path length $L'$, the number of data qubits participating in a lattice-surgery operation divided by $d^2$. In a conventional 2D merge, $L'$ grows linearly with the path length $L$ because every intermediate cell contributes $d^2$ data qubits; in the Bypass layer, a merge between distant cells uses only the $d$-wide fragments in the sparse layer, so $L'$ scales as $O(L/d)$. This reduction is what lowers both the syndrome-graph size handed to the decoder and the per-operation logical error rate. The complementary mechanism is the CBPI stack, a decomposition of average code beats per instruction into Base, Magic, Path, and Decoding components, used to identify which hazard dominates; the paper also proves that any 'immediate-operation' data-cell arrangement has a data-cell ratio at most 50%, and exhibits a 50% arrangement the Bypass layout exploits.","core_discovery":"The central claim is that a dedicated sparse qubit layer can act as a programmable network for lattice surgery. In the Bypass architecture, each cell of the logic layer has a thin 'fragment' of qubits in the Bypass layer, and inter-layer CNOTs build wide-rectangle stabilizers that let a merge travel between distant cells using only $O(d^2 + Ld)$ data qubits instead of $O(Ld^2)$. Defining the effective path length as the number of data qubits involved divided by $d^2$, this turns an $O(L)$ path into an $O(L/d)$ path. Shorter paths mean lighter decoding tasks and lower logical error rates per operation; together with multiple path options that resolve conflicts, the architecture removes the path and decoding penalties that appear at dense data-cell arrangements. The paper argues, using the proportionality between program logical error rate and total effective path length, that the code distance $d$ can be lowered by 2 under practical assumptions, which is what converts improved fidelity into a 1.73x speedup and a 17% hardware reduction.","pith_inferences":["Because the advantage scales with the number of data cells and with the scarcity of decoding resources, the Bypass gain is likely to widen for larger FTQC programs or for cryogenic decoder budgets, an extrapolation beyond the simulated benchmarks.","The same sparse-layer idea could be applied to other topological codes or to neutral-atom or photon-based hardware, but only if inter-layer two-qubit gates can be kept close to intra-layer error rates; the paper's assumptions suggest an experimental requirement of physical error rate $p' \\lesssim 10p$.","The CBPI stack could be used as an iterative design loop: measure which hazard dominates, change the architecture, and re-measure; extending it to include factory decoding or non-greedy schedulers would be a natural next step not carried out here."],"forward_implications":["Long-range lattice-surgery merges, which dominate large QPE programs, become roughly $d$ times cheaper in the number of data qubits involved, so the advantage grows with problem size.","Path conflicts on a single plane disappear when one of two intersecting operations can be routed through the Bypass layer, allowing denser data-cell arrangements without path penalties.","Total effective path length for the Fermi-Hubbard benchmark falls to about 30% of the conventional 2D layout's value (and about half of a two-logic-layer 3D layout), which at fixed code distance lowers the program's logical error rate by the same factor.","With the assumed noise model and a code distance reduced from 25 to 23, the Bypass layout gives a 1.73x speedup and 17% fewer hardware resources in the moderate-resource case; in the limited-resource case the speedup is 2.35x with 13% savings.","The Bypass layout also reduces the variance in execution time over random logical-qubit placements, meaning less placement optimization is needed during compilation."],"supporting_citations":[{"why":"Defines the lattice-surgery merge and split operations that the Bypass layer routes.","marker":"[26]"},{"why":"Establishes the surface-code decoding-graph picture connecting syndrome-graph size to logical error rate.","marker":"[14]"},{"why":"Provides the circuit-level stabilizer simulator used to compute MEAS_ZZ logical error rates with and without the Bypass layer.","marker":"[15]"},{"why":"Supplies the minimum-weight matching decoder used in the error-rate simulations.","marker":"[23]"},{"why":"Gives the two-level 15-to-1 magic-state factory, its 15-code-beat generation rate, and the cell/factory layout assumptions used in the system simulation.","marker":"[33]"},{"why":"Provides the Fermi-Hubbard and Jellium QPE benchmarks, the SELECT-circuit structure, problem sizes, the $d=25$ reference distance, and the quantum-advantage size threshold.","marker":"[49]"},{"why":"Supplies the measured inter-chip to intra-chip error-rate ratio (about 5x) used for the practical $p'=5p$ scenario.","marker":"[19]"},{"why":"Supports the flip-chip/chiplet implementation assumptions for stacking the Logic and Bypass layers.","marker":"[40]"}],"fun_headline_variants":["2.5D bypass lattice surgery: 1.73x faster, 17% fewer qubits","Bypass architecture: d-fold shorter paths, 1.73x speedup","Sparse bypass layer cuts surgery path length by factor d","Lattice surgery gets a bypass: 1.73x faster, 17% less hardware"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premises are that the full program's logical error rate scales in direct proportion to the total effective lattice-surgery path length, and that inter-layer CNOTs can be fabricated with physical error rates no larger than about 5–10 times the intra-layer rate; if either fails, the claimed distance reduction, speedup, and resource savings do not follow.","fun_headline_variants_meta":{"raw":{"variants":["2.5D bypass lattice surgery: 1.73x faster, 17% fewer qubits","Bypass architecture: d-fold shorter paths, 1.73x speedup","Sparse bypass layer cuts surgery path length by factor d","Lattice surgery gets a bypass: 1.73x faster, 17% less hardware"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001149,"raw_usage":{"total_tokens":4802,"prompt_tokens":1019,"completion_tokens":3783,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":3692}},"tokens_in":635,"tokens_out":3783,"duration_ms":33025,"temperature":1.0,"reasoning_tokens":3692,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:01:38.053192+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, by circuit-level simulation of a full QPE program at a fixed code distance, whether the total program logical error rate is indeed proportional to the total effective path length $L'$; if cutting $L'$ by a factor of about 2 does not cut the logical error rate correspondingly, the code-distance reduction from 25 to 23 is unsupported. Alternatively, fabricate or measure an inter-layer CNOT in a flip-chip device: if its error rate exceeds about $10p$, the Bypass advantage for short lattice-surgery paths disappears, and the crossover length for benefit rises above the $L=31$ threshold reported in the paper.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the lattice-surgery merge and split operations that the Bypass layer routes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Fermi-Hubbard and Jellium QPE benchmarks, the SELECT-circuit structure, problem sizes, the $d=25$ reference distance, and the quantum-advantage size threshold."},{"cited_title":"Reagor, M","cited_arxiv_id":null,"evidence_quote":"Supplies the measured inter-chip to intra-chip error-rate ratio (about 5x) used for the practical $p'=5p$ scenario."}],"review_version":1}