{"id":"dfc26cd3-df95-4fb4-a431-8e777f8d0166","arxiv_id":"2608.11719","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A decoder-aware scheduler packs transversal CNOT gates into surface-code quantum programs as densely as decoder capacity allows, using hybrid decoder selection, template-based DEM stitching, and sub-window parallel decoding.","lead":"PACE is a scheduling framework for surface-code quantum computers with transversal CNOT gates; it balances quantum speedup against the classical decoder's latency, memory, and just-in-time error-model compilation costs. Its three techniques split, stitch, and merge decoding work so that transversal CNOT gates can be packed as densely as the decoder resources allow.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central speedup and feasibility claims rest on a volume-only decoder model fitted to single-patch memory tasks, but multi-patch TCNOT DEMs contain hyperedges whose density and topology are not modeled; this needs direct measurement before the headline numbers are accepted.","rationale":"The paper makes a genuine systems contribution: it identifies DEM preparation, decoder memory, and decoding latency as first-class scheduling resources for transversal-CNOT FTQC, and the three mitigation techniques are individually plausible with supporting experiments. The reader's conditional verdict is appropriate because the evaluation's central quantitative claims are computed from a capacity model whose exponents are fitted to single-patch memory windows and then extrapolated to multi-patch hyperedge-rich TCNOT windows. My own reading confirms that this is the most load-bearing assumption: almost every headline number, including the 7.86x speedup and the infeasibility of current unmatchable decoders even for All-d, is derived through Eq. (4), and the paper's own Sec. 7.6 explicitly disclaims the missing hyperedge dependence. The concern is not that the authors are careless; it is that the chosen abstraction, V=q w/W, has not been shown to be a sufficient statistic for the decoder cost of TCNOT workloads. The proposed concrete test would settle whether the extrapolation holds. I do not see a stronger attack on the central claim: the window-wise correction is verified against full-circuit decoding in Fig. 11, DEM Stitch has direct timing measurements in Fig. 12, and the scheduling algorithm is well specified. The absence of a released artifact and the lack of head-to-head comparison with Decoder Switching and GreenPeas are secondary and addressable, and they do not undermine the internal logic of the framework. Therefore the verdict should remain conditional: accept the framework as a valid proposal, but require validation of the decoder-cost model on multi-patch TCNOT windows before relying on the quantitative speedup and feasibility claims.","tokens_in":22410,"tokens_out":3752,"duration_ms":40406,"concrete_test":"Generate Stim DEMs for multi-patch TCNOT decoding windows with matched decoding volume V=q w/W but different hyperedge densities and topologies: e.g., two-patch, four-patch, and eight-patch windows containing a single correlated TCNOT versus several pairwise TCNOTs, keeping V constant. Measure BP-LSD and Relay-BP per-window latency and peak working memory over these windows at d=15,19,23, and fit the gamma_dec and gamma_mem exponents separately. If these multi-patch fits deviate materially from the single-patch memory fits (e.g., by more than 20% in the exponent, or by a large constant factor at fixed V), re-run the Sec. 7.4 capacity sweep with the corrected model and compare the speedups and the All-d feasibility conclusion.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing premise is the volume-only power-law capacity model in Table 1 and Eq. (4), with exponents fitted in Sec. 7.6 to one-patch memory windows (q=1, d<=w<=3d, p=1e-3) and then applied to multi-patch TCNOT windows whose DEMs contain hyperedges. The scaling of BP-LSD and Relay-BP latency and memory with hyperedge density and topology is not captured by V=q w/W: a single fault can flip detectors across k patches and appear as a k-way hyperedge, and unmatchable decoding cost can depend on hyperedge count and structure, not just on the detector count. Because Sec. 7.4's speedup results (up to 7.86x at V_dec=64) and Sec. 7.6's infeasibility conclusion are outputs of this model, an unmodeled hyperedge dependence could make Eq. (4) understate decoder load: schedules selected as feasible could miss real-time deadlines, and the All-d infeasibility claim could shift. The paper itself concedes in Sec. 7.6 that the model 'does not model additional dependence on hyperedge density or topology.' This is not an internal inconsistency, but it is a missing validation of the central quantitative claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes PACE, a decoder-aware scheduling framework for fault-tolerant quantum computation that uses transversal CNOT (TCNOT) gates on surface-code patches. PACE combines three decoder-side techniques: hybrid window decoding, which routes graphlike windows to fast matchable decoders and non-graphlike windows to unmatchable decoders; DEM Stitch, a template-based just-in-time method for constructing schedule-specific detector error models; and sub-window parallel decoding (SPD), which decomposes large decoding windows into smaller sub-windows via graph coloring. The scheduler then searches a bounded family of spacing policies to find the densest TCNOT schedule that satisfies power-law feasibility constraints on DEM-preparation latency, decoder memory, and decoding latency (Eqs. (4)-(5)). The evaluation shows that hybrid decoding preserves logical error rates, DEM Stitch reduces DEM compilation time by 108.4x for a d=23 memory window, SPD reduces peak decoding volume, and PACE achieves speedups over the All-d baseline up to 7.86x at normalized capacity V_dec=64. A case study concludes that current unmatchable-decoder configurations do not meet even the minimum capacities required for the All-d schedule.","tokens_in":22671,"tokens_out":7277,"duration_ms":77173,"significance":"If the central capacity-model extrapolation is validated, PACE would fill a real gap by making decoder capacity an explicit, schedulable resource for TCNOT-based FTQC. The three component techniques are individually well motivated and each has a concrete experiment supporting its component claim: hybrid decoding matches BP-LSD error rates, DEM Stitch shows a large compilation speedup, and SPD reduces decoding volume. The paper is also honest about many of its assumptions, and the sensitivity analysis with respect to gamma_dec is a useful addition. The concern that the scheduler is circular because it optimizes against the same model used for feasibility is not, in my reading, a defect: this is standard co-design, not hidden circularity. The real risk is external validity: the feasibility model in Sec. 5.1 and the headline speedups in Sec. 7.4 are built on a volume-only power-law scaling of decoder cost, fitted in Sec. 7.6 to single-patch memory windows and then applied to multi-patch TCNOT windows whose DEMs contain hyperedges. The paper itself concedes in Sec. 7.6 that this approximation does not model hyperedge density or topology.","major_comments":[{"comment":"The load-bearing feasibility model defines decoding volume as V=qw/W and models DEM-preparation latency, decoder working memory, and decoding latency as power-law functions of V. The exponents and constants are fitted in Sec. 7.6 to one-patch memory windows (q=1, d<=w<=3d) and then applied to multi-patch TCNOT windows whose DEMs contain hyperedges. The paper itself states in Sec. 7.6 that this volume-only approximation 'does not model additional dependence on hyperedge density or topology.' For unmatchable decoders such as BP-LSD and Relay-BP, the number and structure of hyperedges can significantly affect both runtime and memory usage; a single fault that propagates across k patches can appear as a k-way hyperedge, and decoder cost need not scale with detector count alone. If multi-patch TCNOT tasks scale worse than the fitted volume-only power law, Eq. (4) understates decoder load, the schedules selected in Sec. 7.4 could miss real-time deadlines, and the infeasibility conclusion for current decoders in Sec. 7.6 could shift. I ask the authors to provide direct measurements or a validated model of decoding latency and memory on multi-patch TCNOT windows (e.g., two- and four-patch windows spanning the relevant code distances) and to re-run the Sec. 7.4 and Sec. 7.6 analyses with that data. Until then, the quantitative speedup and infeasibility claims are contingent on an untested extrapolation.","section":"Sec. 5.1 / Table 1 / Eq. (4) / Sec. 7.6"},{"comment":"The correctness of the JIT DEM pipeline rests on the claim that DEM Stitch produces a window-level DEM equivalent to full-circuit construction. The evaluation in Sec. 7.2 measures only compilation latency and template-cache footprint; it does not verify that the stitched DEM and the full Stim DEM have identical detector check matrix H, logical action matrix A, and fault probabilities p for any TCNOT window. If stitching introduces errors in detector indexing, logical-observable propagation, or temporal-port resolution, then schedules deemed feasible by Eq. (4) could be decoded incorrectly even when latency and memory constraints are met. I ask the authors to add an equivalence test that compares syndrome and logical-observable predictions of stitched versus full-circuit DEMs across randomized schedules, and to report the result for at least one multi-patch TCNOT window.","section":"Sec. 4.2 / Algorithm 1 / Sec. 7.2"},{"comment":"The speedups reported in Fig. 14 (up to 7.86x at V_dec=64) are computed under normalized decoder capacities that, according to the paper's own Sec. 7.6 case study, are not met by any of the evaluated current unmatchable-decoder configurations: none reaches even the minimum capacities required for the All-d schedule (V_dec>=2 and V_unit>=2). Consequently, the headline speedup numbers do not represent what PACE would deliver with a decoder characterized in this paper; at feasible current capacities the selected schedules are close to All-d (1.06x at V_dec=4). The manuscript should state this distinction prominently in the abstract and in Sec. 7.4, and should frame the current-hardware result as 'PACE identifies infeasibility and required capacity improvements' rather than as a realized speedup. This does not invalidate the framework, but the current presentation overstates the practical benefit that the paper can support.","section":"Sec. 7.4 / Sec. 7.6"}],"minor_comments":[{"comment":"The paper states that PACE 'schedules TCNOTs as densely as the available decoder resources permit,' but the scheduler restricts the search to a bounded family of period-p spacing policies and selects the shortest candidate in that family. The manuscript should phrase the objective as finding the best schedule within the evaluated policy family, not as unconstrained densest scheduling.","section":"Sec. 5.3 / Algorithm 2"},{"comment":"Figure 12(b) reports preparation time versus number of patches, but the text does not specify whether all data points share the same code distance and window height; please state these parameters in the caption and the text.","section":"Sec. 7.2 / Fig. 12"},{"comment":"The Relay-BP FPGA projection assumes linear scaling of decoding latency with decoding volume based on a quoted 24 ns per iteration. Since this linearity is not measured for the multi-patch structures considered, please explicitly label it as an assumption and discuss how a superlinear scaling would affect the conclusions.","section":"Sec. 7.6 / Table 3"},{"comment":"The finite-decoder-pool discussion is useful but should be cross-referenced from Sec. 5.2, where the feasibility constraints assume sufficient concurrent DEM preparation and decoder instances; otherwise readers may mistake this assumption for a demonstrated architectural requirement.","section":"Sec. 8.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the quantum architecture/decoding community and makes a useful contribution if the capacity-model extrapolation is validated. My main concern is external validity of the volume-only power-law model for multi-patch TCNOT tasks, which the authors themselves flag in Sec. 7.6; I would not reject on that basis alone, but the abstract and Sec. 7.4 should be reframed to avoid implying realized speedups on current hardware. I would also require an explicit equivalence check for DEM Stitch, since the feasibility checks depend on its correctness. The paper should be encouraged to add the multi-patch validation rather than rely on memory-window fits."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"PACE is a real systems contribution you should know about: it's the first TCNOT scheduling work I've seen that treats the classical decoder and JIT DEM compilation as schedulable resources rather than fixed constraints. The three techniques are individually sensible, and the paper does honest work on each. The headline speedup numbers, though, come from a volume-only power-law model that the authors themselves flag as an approximation, so treat 7.86x as conditional on that model.\n\nWhat's genuinely new: the framing and the integration. Prior work (Cain, Sahay, Zhou) optimizes the quantum side of TCNOT speedup, but PACE makes the decoder capacity explicit in the schedule selection. The three components are not radically novel—per-window decoder switching, parallel window decoding, and JIT DEM compilation all have recent precursors—but the TCNOT-aware sub-window decomposition with graph coloring is a clean twist. DEM Stitch is the most concrete contribution: 108x faster than full Stim construction, with a modest template cache. The case study showing that current unmatchable decoders don't even meet All-d minimum capacities is a useful reality check.\n\nThe soft spot is the one the stress test hit: the feasibility model in Eq (4) and Table 1 treats decoding volume V=q w/W as the only scaling variable. The exponents are fitted to one-patch memory windows and then extrapolated to multi-patch TCNOT windows whose DEMs have hyperedges. Hyperedge density and topology could change the scaling of BP-LSD and Relay-BP, and that could shift both the selected schedules and the infeasibility claims. The paper says as much in Sec 7.6. So the speedup numbers are model-based estimates, not measurements, and the paper should be read that way. This is a limitation, not a fatal flaw, because the framework's value doesn't collapse if the exponents shift; the scheduling logic still stands, but the specific quantitative claims need validation on real multi-patch DEMs.\n\nAlso worth noting: DEM Stitch's equivalence to full Stim is asserted, not verified—they never show the stitched DEM produces the same error rates. No artifact is released. No head-to-head with Decoder Switching or GreenPeas. All of these are addressable.\n\nWho's this for: people working on FTQC system scheduling, real-time decoding, and neutral-atom architectures. It deserves a serious referee; I'd send it to peer review with the expectation of major revision focused on validating the capacity model on multi-patch DEMs and verifying DEM Stitch.","headline":"PACE is a legitimate systems contribution, but its headline speedup claims rest on a decoder model that the authors themselves flag as an approximation, so take them as conditional.","tokens_in":23267,"tokens_out":2145,"would_cite":true,"duration_ms":21183,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PACE claims that scheduling transversal CNOTs around classical decoder capacity speeds quantum execution up to 7.86x over the All-d baseline while keeping real-time decoding feasible.","keywords":["transversal CNOT gates","surface codes","real-time decoding","decoder-aware scheduling","window decoding","just-in-time detector error models","fault-tolerant quantum computing","neutral-atom architectures"],"falsifier":"Hold decoding volume fixed on multi-patch TCNOT windows while varying only hyperedge density and topology, measuring the unmatchable decoders' latency and memory; if these grow faster than the fitted $V^\\gamma$ power laws, then Eq. (4) is optimistic and PACE-selected schedules could miss their real-time deadlines.","tokens_in":22133,"feed_emoji":"⚛️","tokens_out":12143,"duration_ms":113793,"temperature":0.7,"pith_summary":"Transversal CNOT gates can cut the syndrome-extraction rounds between logical operations from $O(d)$ to $O(1)$, but this paper argues that the real bottleneck is the classical decoder: correlated errors across patches inflate decoding volume, schedule-dependent detector error models can no longer all be precomputed, and dense schedules outgrow decoder memory and latency. PACE is a scheduling framework that attacks these three costs with per-window choice between matchable and unmatchable decoders, just-in-time assembly of window detector error models from reusable templates, and sub-window parallel decoding that decomposes large windows. Its scheduler then packs TCNOTs as densely as the modeled decoder resources permit, converting spare decoding capacity into quantum execution speedup. On ten benchmark circuits, the selected schedules run up to 7.86x faster than the conservative All-d baseline at the largest evaluated decoder capacity, while currently configured unmatchable decoders would not reach even the All-d schedule's minimum capacity. The paper's contribution is to reframe the promise of transversal gates from a hardware question into a scheduling and decoder-resource question.","feed_headline":"Decoder-aware CNOT scheduling yields up to 7.86x speedups","feed_subtitle":"Transversal gates cut syndrome rounds; PACE packs them as densely as the decoder allows.","key_machinery":"The carrying object is the decoding volume $V=qw/W$, a scalar for a decoding task covering $q$ logical patches over a window length $w$ normalized by the standard window height $W$. PACE models each decoder-side cost, DEM preparation latency, working memory, and decoding latency, as a power law $\\alpha V^\\gamma$ of this volume, and the feasibility constraints in Eq. (4) bound both the largest per-sub-window volume and the accumulated latency of the colored decoding layers. The three techniques act on this machinery directly: Hybrid Window Decoding shrinks the set of tasks needing unmatchable decoders, DEM Stitch cuts the prefactor of DEM-preparation cost through reusable templates, and Sub-window Parallel Decoding replaces one large window with colored sub-windows so the volume constraints are checked at sub-window granularity.","core_discovery":"On its own terms, PACE establishes that the decoder-side costs of TCNOTs can be controlled by three mechanisms and that, once controlled, TCNOT density becomes a schedulable quantity. Hybrid Window Decoding routes each decoding window to a cheap matchable decoder when its detector error model is graphlike and to an unmatchable decoder otherwise, with a window-wise correction step that lets different decoder types coexist in one pipeline. DEM Stitch precompiles reusable local fault templates and stitches them into schedule-specific window detector error models just in time, avoiding exponential precompilation. Sub-window Parallel Decoding forms sub-windows from a TCNOT-event graph and colors the resulting conflict graph so that non-conflicting sub-windows decode in parallel. The scheduler searches a bounded family of spacing policies and selects the shortest schedule whose windows satisfy the capacity constraints in Eq. (4), yielding geometric-mean speedups over All-d from 1.06x at $V_{dec}=4$ to 7.86x at $V_{dec}=64$, with current decoder systems still unable to reach even the All-d feasibility point.","pith_inferences":["The volume-only power-law model is the weakest point to test first: a direct measurement on multi-patch TCNOT windows with constant volume but varying hyperedge density would show whether Eq. (4) needs a topology-dependent term before these speedups can be trusted on hardware.","The same capacity-versus-schedule tradeoff could be applied to other classical resources, such as magic-state distillation throughput or network bandwidth, by expressing them as similar per-window constraints.","Since PACE still produces beneficial schedules when non-graphlike tasks are converted to graphlike decoding, decoder-aware scheduling and fast correlated decoding are likely complementary and could be combined into one pipeline.","A concrete extension would be to replace $V=qw/W$ with an effective volume that weights hyperedges and then rerun the capacity sweeps; the gap between the two results would quantify how much the hyperedge approximation matters."],"forward_implications":["Under PACE's model, decoder capacity becomes a first-class scheduling resource: raising $V_{dec}$ or $V_{unit}$ shortens the selected TCNOT schedule without any change to the quantum hardware.","The exponential diversity of window-level detector error models means online DEM compilation is unavoidable for aggressive TCNOT scheduling, so template-based just-in-time assembly is a necessary system component.","Because current unmatchable-decoder configurations do not reach the minimum capacities of even the All-d schedule at $d=23$, TCNOT-based FTQC needs faster unmatchable decoding, larger decoder memory, or a longer syndrome-extraction period.","Even at the highest evaluated decoder capacity, more than half of the decoding volume on average remains graphlike, so hybrid routing to matchable decoders continues to matter.","Workloads with higher data-loading parallelism place more correlated TCNOTs in each decoding window and therefore need larger decoder capacity before dense schedules become feasible."],"supporting_citations":[{"why":"Supplies the All-d spacing baseline that PACE compares against and the prior strategy of bounding TCNOT concurrency with d-round separation.","marker":"[44]"},{"why":"Supplies the O(1)-round TCNOT timing model and the 1 ms syndrome-extraction period used in the evaluation.","marker":"[54]"},{"why":"Establishes the low-overhead transversal architecture that motivates scheduling TCNOTs at all.","marker":"[55]"},{"why":"Provides the full-circuit DEM construction baseline that DEM Stitch replaces and outperforms in compile-latency measurements.","marker":"[23]"},{"why":"Supplies the matchable decoder used in Hybrid Window Decoding and as the graphlike-capacity reference.","marker":"[31]"},{"why":"Supplies the unmatchable decoder whose measured latency and memory scaling define the decoder-capacity case study.","marker":"[32]"},{"why":"Supplies the parallel window decoding principle that Sub-window Parallel Decoding extends into sub-windows.","marker":"[45]"},{"why":"Supplies the existing just-in-time DEM compilation baseline for schedule-specific window DEMs.","marker":"[56]"},{"why":"Describes the fast correlated decoding approach that the paper identifies as a complementary backend for non-graphlike tasks.","marker":"[12]"}],"fun_headline_variants":["PACE scheduler speeds up CNOTs without overwhelming decoder","Schedule TCNOTs denser, decode faster with PACE","PACE avoids decoder overload for fast transversal CNOTs","TCNOT scheduling that keeps decoders from choking: PACE","PACE: decoder-aware scheduling for up to 7.86x speedup"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a single number, the decoding volume $V=qw/W$, predicts DEM preparation time, decoder memory, and decoding latency through fixed power laws even for multi-patch TCNOT windows, and the paper itself notes that this ignores hyperedge density and topology.","fun_headline_variants_meta":{"raw":{"variants":["PACE scheduler speeds up CNOTs without overwhelming decoder","Schedule TCNOTs denser, decode faster with PACE","PACE avoids decoder overload for fast transversal CNOTs","TCNOT scheduling that keeps decoders from choking: PACE","PACE: decoder-aware scheduling for up to 7.86x speedup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1837,"prompt_tokens":1101,"completion_tokens":736,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":647}},"tokens_in":717,"tokens_out":736,"duration_ms":7327,"temperature":1.0,"reasoning_tokens":647,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:30:10.143466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold decoding volume fixed on multi-patch TCNOT windows while varying only hyperedge density and topology, measuring the unmatchable decoders' latency and memory; if these grow faster than the fitted $V^\\gamma$ power laws, then Eq. (4) is optimistic and PACE-selected schedules could miss their real-time deadlines.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the full-circuit DEM construction baseline that DEM Stitch replaces and outperforms in compile-latency measurements."}],"review_version":1}