{"id":"f25e51c8-fb20-4fed-a72c-cc1a4d9eb3a4","arxiv_id":"2608.10576","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A streaming surface-code decoder called Snowflake runs on commercial FPGAs with sub-microsecond simulated throughput up to distance 21, validated in hardware up to distance 9 at room and cryogenic temperatures.","lead":"The paper builds a decoder for quantum error correction on FPGAs, tests it at room and cryogenic temperatures, and adds low-overhead confidence scores. It also sketches a 2D hardware layout that could fit bigger codes on smaller chips.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulated d=11–21 sub-microsecond claim hinges on an unvalidated 200 MHz assumption; the d=9 hardware point that had to be area-optimized to 60 MHz and gave 1.06 µs/round contradicts it.","rationale":"The paper's main contribution is a genuine hardware implementation of Snowflake for d=3..9, with cryogenic operation and DCS logic. That part is plausible and credit-worthy. The load-bearing issue is the headline throughput claim: 'sub-microsecond for d up to 21.' The data behind d=11..21 are simulation points computed at a target 200 MHz clock, but the only hardware point at which the design was forced to fit an actual device (d=9) required an area-optimized rebuild that ran at 60 MHz and produced 1.06 µs/round—already over budget. Resource scaling from Table 1 (roughly 290 LUTs/node at d=9 and growing) implies d≥13 would need millions of LUTs, which the paper itself admits is 'beyond today's technology.' The simulated curve therefore describes an algorithmic cycle count, not an achievable physical decoder. This is a correctable framing issue: the abstract already says 'extrapolated' and Section 5 proposes a 2D architecture that might realize it, but Section 4.2's categorical sentence overstates. The reader's weakest_assumption targeted exactly this point, and I agree. Other weaknesses (DCS overhead not quantified against baseline, DCS accuracy comparison without data) are real but minor relative to the central throughput claim. A timing-closed implementation at d=11 or a conservative projection using the d=9 achieved 60 MHz would settle the matter.","tokens_in":18994,"tokens_out":9535,"duration_ms":83310,"concrete_test":"Run full synthesis, place-and-route, and timing closure for the d=11 Snowflake configuration on the XCKU5P (or the largest available UltraScale+ device), and measure mean per-round decoding time over 1024 simulated rounds at the achieved Fmax. If the design fails to fit, fails timing at 200 MHz, or yields mean decoding time ≥1 µs, then Figure 6's d≥11 sub-microsecond points are not supported and the claim should be restated as a projection conditional on a 2D/time-multiplexed architecture or a non-existent FPGA.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim in §4.2 that Snowflake 'achieves a sub-microsecond average decoding time ... for code distances up to d=21' is not supported by the hardware evidence. The only distance at which the fully parallel PE design was forced onto a real FPGA is d=9, and there the design exceeded the XCKU5P's LUT capacity (Table 1: 237,929 LUTs), requiring an area-optimized rebalance that ran at 60 MHz and produced 1.06 µs/round (§4.3), already above the 1 µs budget. For d=11,13,...21, LUT counts scale roughly as d^2(d+1) per instance (from Table 1, ~290 LUTs per node at d=9 and growing), implying multi-million-LUT designs for d≥13 that do not fit on any current FPGA; the paper itself concedes that 'implementing a large-scale system would require either a large FPGA beyond today's technology or clusters of FPGAs.' Yet Figure 6's simulated d=11–21 points are presented at a target clock of 200 MHz with no timing closure, no placed-and-routed results, and no evidence that the PE logic achieves 200 MHz even at d=9 (the area-optimized d=9 achieves only 60 MHz). The claim therefore conflates an algorithmic cycle-count simulation with an achievable hardware throughput. A sub-microsecond result at d≥11 may be realizable with a different architecture (e.g., the proposed 2D time-multiplexed variant), but that architecture is only proposed, not implemented or timed. Consequently, the paper's strongest quantitative claim is an extrapolation, not a demonstrated result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an FPGA implementation of the Snowflake streaming Union–Find decoder for the unrotated surface code, augmented with three decoder confidence scores, and evaluates it at both room temperature and inside a cryostat. The authors report resource utilisation for code distances d=3–9, measured decoding throughput, and cryogenic power and temperature data. For d=11–21 they simulate decoding times at a target clock frequency of 200 MHz and claim a sub-microsecond average decoding time per stabiliser measurement round, which they argue satisfies the 1 µs real-time budget for superconducting quantum devices. They also propose a 2D time-multiplexed architecture intended to reduce resource usage, supported by an analysis of 'active depth' and by pseudocode.","tokens_in":19349,"tokens_out":5657,"duration_ms":49168,"significance":"If the extrapolation were backed by timing closure, the sub-microsecond streaming decoder claim would be a significant result for real-time quantum error correction. The measured hardware results for d=3–9, the cryogenic power characterisation, and the confidence-score implementation with negligible latency overhead are useful engineering contributions, and the external benchmarking against Helios, AQ2-RT, and QUEKUF is valuable. The active-depth analysis and the proposed 2D architecture are also interesting. However, the headline claim as stated in Section 4.2 is not supported by the presented evidence: the d=11–21 points are simulated at an assumed clock frequency, while the only d=9 hardware implementation runs at 60 MHz and already exceeds the 1 µs budget. The manuscript should either provide timing-closure evidence for the larger designs or explicitly reframe the claim as an extrapolation.","major_comments":[{"comment":"The statement that 'Snowflake achieves a sub-microsecond average decoding time for code distances up to d=21' is not supported by the hardware evidence. The d=11–21 curve is a cycle-count simulation evaluated at an assumed 200 MHz clock; no placed-and-routed results or timing closure are shown for these distances. The only d=9 hardware result, from the area-optimised design, runs at 60 MHz and yields 1.06 µs per round (Section 4.3), already above the 1 µs budget. The claim should be reworded to distinguish demonstrated hardware performance (d=3–9, with d=9 above the budget) from extrapolated simulation, or the extrapolation should be removed from the abstract and conclusions.","section":"Section 4.2, Figure 6"},{"comment":"The resource-scaling data directly conflict with the plausibility of the 200 MHz target at large distances. The d=9 design already exceeds the XCKU5P LUT capacity (237,929 LUTs), and the area-optimised variant that fits still runs at only 60 MHz. Since the number of processing elements scales as d^2(d+1), the d=11–21 designs would require many millions of LUTs, beyond any current FPGA; the paper itself concedes this in the abstract and in Section 5. The simulation should therefore be presented as an algorithmic extrapolation under an explicit assumption that the clock target is met, not as an achieved hardware result.","section":"Section 4.1, Table 1"},{"comment":"The estimate that QUEKUF crosses the 1 µs threshold at d≥11 assumes a constant clock frequency of 247.5 MHz, equal to the maximum reported at d=10. This optimistic assumption is stated in the text, but the comparison is used to position Snowflake's extrapolated performance; the same standard of evidence should be applied to both decoders. Please add a sentence noting that the QUEKUF figures are estimates under a favourable clock assumption, and consider showing sensitivity to the clock frequency.","section":"Section 4.2, QUEKUF comparison"}],"minor_comments":[{"comment":"The axis labels in Figure 6 appear garbled in the manuscript; please ensure that the fonts and symbols render correctly in the final version.","section":"Figure 6"},{"comment":"The sentence 'the architecture leaves ample room for further timing optimisation' is an unsupported assertion; consider replacing it with a discussion of the specific critical-path bottlenecks that limited the area-optimised d=9 design to 60 MHz.","section":"Section 4.2"},{"comment":"The claim that the confidence-score logic has 'minimal impact on the resource utilisation' is not quantified; state explicitly whether Table 1 includes the DCS logic and report the incremental LUT and FF counts attributable to the three DCS implementations.","section":"Section 3.2"},{"comment":"The abstract carefully says the results 'when extrapolated' remain within acceptable limits, but Section 4.2 states the sub-microsecond figure as an achieved value; the body should consistently use the same qualifier for all distances above the hardware-validated range.","section":"Abstract and Section 4.2"},{"comment":"The sentence 'This amounts to a ∼47% reduction in power draw at cryogenic temperatures compared to the XCKU5P FPGA' is ambiguous: the comparison appears to be between the Artix-7 d=3 implementation and the XCKU5P d=3 implementation, and this should be stated explicitly.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The hardware work is generally well executed, but the central extrapolation is currently presented too strongly, and the resource-scaling discussion undermines the assumed 200 MHz clock for large distances. The paper would be acceptable after the claims are reframed and the QUEKUF comparison is caveated. Please also check the rendering of Figure 6 and the competing-interest declaration, which is present."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: this is a genuine hardware implementation paper. They put Snowflake on two commercial FPGAs, ran it at room temperature and around 4 K, and measured resource use, throughput for d=3 to 9, and power/temperature. That is real engineering work, and the cryogenic measurements with calibrated die temperature are a useful addition to the FPGA-decoder literature. The active-depth analysis is a nice new idea, and the proposed 2D time-multiplexed architecture is sensible, though it is only pseudocode.\n\nThe soft spot is the headline. Section 4.2 states Snowflake \"achieves a sub-microsecond average decoding time for code distances up to d=21.\" That is not what the hardware shows. The d=9 implementation could only close at 60 MHz after an area-optimized rebalance, giving 1.06 us per round. The d=11 to 21 points are a simulation at a target 200 MHz with no timing closure, and the LUT counts for d>=13 would not fit on a current FPGA, which the paper itself concedes. So the claim conflates cycle-count simulation with hardware throughput. The abstract is more careful (\"when extrapolated\"), but the body overstates it. That should be fixed before publication.\n\nOther soft spots are minor. The DCS accuracy comparison is stated without data; they say cluster-size 1-norm is best but give no Monte Carlo numbers. The \"negligible overhead\" claim for DCS logic lacks a baseline without DCS to compare against. No code or artifacts are provided. All correctable.\n\nThe stress-test note is mostly right, but I would not call the d=9 point a contradiction of the 200 MHz assumption; the area-optimized design is a different design, so it does not prove 200 MHz is impossible, just unvalidated. The right fix is to present the large-distance numbers as extrapolated and, ideally, provide a timing-closed point at a moderate distance.\n\nThe comparison with AQ2-RT and QUEKUF is fair, and they acknowledge the error-rate difference. The citation pattern looks fine; Snowflake's own software analogue and localuf are legitimately used for validation.\n\nWho should read it: people building real-time decoding hardware for superconducting or spin systems, and anyone picking an FPGA decoder. It deserves a serious referee. I would send it to review with a request to temper the d<=21 claim and add DCS data. Verdict conditional, not reject.","headline":"Real cryogenic FPGA decoder work with honest limitations, but the sub-microsecond d=21 claim is extrapolation, not measurement.","tokens_in":19872,"tokens_out":2402,"would_cite":true,"duration_ms":21741,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"A streaming Union-Find decoder on an FPGA decodes a surface-code syndrome round in under a microsecond for code distances up to 21.","keywords":["quantum error correction","surface code","Union-Find decoder","streaming decoding","FPGA","cryogenic electronics","decoder confidence score","real-time decoding"],"falsifier":"Run the Verilog design for $d=11$ through place-and-route on the same FPGA family at a 200 MHz constraint; if timing closure fails or the measured average decode time per round exceeds 1 microsecond, the central extrapolation is refuted.","tokens_in":18783,"feed_emoji":"❄️","tokens_out":10379,"duration_ms":92844,"temperature":0.7,"pith_summary":"The paper aims to establish that Snowflake, a round-wise streaming Union-Find decoder for the surface code, can run on commercial FPGAs fast enough for real-time fault-tolerant quantum computing. The authors report an average per-round decoding time below one microsecond for code distances up to $d=21$ in simulation, satisfying the real-time budget for superconducting qubits, and validate the same design on hardware for distances 3 to 9 at both room and cryogenic temperatures. They also add a decoder confidence score, the Cluster Size 1-Norm Fraction, with no added latency and negligible resource cost, and propose a flattened 2D architecture that reduces the logic footprint by a factor of about the code distance while implementing the same algorithm exactly. If correct, this makes a single small FPGA a viable tier-1 decoder with soft output for a superconducting quantum processor.","feed_headline":"FPGA decoder clears the 1-microsecond budget up to distance 21","feed_subtitle":"Cryogenic FPGA tests validate the decoder at distances 3 to 9; simulation clears the 1-microsecond real-time budget up to 21.","key_machinery":"The load-bearing object is Snowflake's decoding window, a fixed-size connected subgraph of the surface-code decoding graph containing $d^2(d+1)$ nodes; the window slides upward by one layer per measurement round (a “drop”), and clusters grow and merge within it according to a Union-Find schedule. The hardware maps each node to a small finite-state-machine processing element and each edge to a communication link, with edge variables (growth and correction) kept in a global register array; a central controller issues the drop-grow-merge commands. The confidence machinery is the Cluster Size 1-Norm Fraction, obtained by summing growth variables in the edge table, which quantifies the total cluster volume in the window and is computed in the idle time between drops. The proposed 2D architecture replaces the full $d^2(d+1)$-node network with a $d\\times d$ grid of column processors plus memory, processing one horizontal sheet of the window at a time; active-depth measurements show most nodes are idle most of the time, so early stopping at the active depth is exact, not approximate.","core_discovery":"On the paper's own terms, the central claim is that Snowflake's inherently parallel decoding window maps naturally onto an FPGA fabric: every node of the fixed-size decoding window becomes a processing element and every edge a communication link, so the drop-grow-merge cycle runs as a distributed synchronous process at 200 MHz. With this implementation, simulated code distances up to $d=21$ decode each stabiliser-measurement round in sub-microsecond average time, meeting the 1 microsecond inverse-throughput budget for superconducting devices. Physical hardware tests for $d=3$ to 9 at room and cryogenic temperatures confirm the design operates correctly, though the $d=9$ build exceeded the FPGA's LUT capacity and required an area-optimised variant running at 60 MHz. The same circuit computes the Cluster Size 1-Norm Fraction confidence score between drop cycles at no timing overhead, and the authors further propose a 2D slice-processing architecture that would cut logic usage by a factor of $d$ and enable offload of distant decoding-window data to high-speed memory.","pith_inferences":["The active-depth result suggests the decoding window height could be made adaptive per noise realisation, shrinking the window when defects are shallow; the paper fixes the window size and only stops early.","The confidence score's negligible overhead generalises naturally: any local decoding graph whose edge variables are already in registers could emit a cluster-volume score, so the idea may transfer to other hardware decoders.","A memory-bound 2D design raises a new bottleneck the paper does not quantify: if the high-speed bus or DRAM bandwidth cannot feed sheets fast enough at 200 MHz, the factor-of-$d$ saving in logic will be traded for a memory-bandwidth ceiling.","Clusters of small FPGAs at the 50 K stage might be tested end-to-end with a realistic syndrome stream; the paper identifies the direction but stops short of a benchmark."],"forward_implications":["A single commercial FPGA running Snowflake clears the 1 microsecond per-round budget for distances up to $d=21$, so streaming decoding can coexist with fast superconducting readout.","Decoder confidence scores come essentially for free in this architecture, enabling decoder switching or postselection without sacrificing throughput.","Cryogenic operation works, but heat from a highly utilised FPGA at the 4 K stage is substantial, so large-scale deployments would move to higher-temperature stages or clusters.","The 2D slice architecture reduces logic resource use by roughly a factor of $d$, making Snowflake practical on smaller FPGAs or as the basis for a custom cryo-CMOS ASIC.","Early-stop logic based on active depth is exact—skipping deeper sheets never changes the decoder output—so the 2D design gains speed without losing accuracy."],"supporting_citations":[{"why":"Defines the Snowflake streaming Union-Find algorithm and the 1:1 and 2:1 cluster growth schedules that this hardware implements.","marker":"[1]"},{"why":"Provides the neural decoder baselines, including the real-time variant whose sub-microsecond throughput this work compares against.","marker":"[15]"},{"why":"Establishes the Cluster Size 1-Norm Fraction as a useful decoder confidence score for postselection.","marker":"[23]"},{"why":"Supplies the overlapping-window FPGA Union-Find decoder whose reported throughput is a primary comparison.","marker":"[43]"},{"why":"Provides the FPGA Union-Find decoder on the toric code used to compare extrapolated throughput at larger distances.","marker":"[45]"},{"why":"Documents cryogenic operation of the same Kintex UltraScale+ FPGA family, motivating the device choice and expected power savings.","marker":"[50]"},{"why":"Provides the cryogenic low-dropout regulator design adapted for stable FPGA power delivery inside the cryostat.","marker":"[53]"},{"why":"Supplies the circuit-level noise model used to generate the test syndromes for both simulation and hardware runs.","marker":"[54]"}],"fun_headline_variants":["Streaming FPGA decoder clears 1-microsecond budget to d=21","Cryo-tested FPGA decoder streams with confidence, hits sub-microsecond","Sub-microsecond stream decode on FPGA, with confidence scores at no cost","FPGA stream decoder: d=21 simulated, d=9 cryo-tested, confidence added"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sub-microsecond decoding claim for distances 11 to 21 rests on a simulated 200 MHz design, even though the largest hardware build (distance 9) already overflowed the FPGA's logic and had to run at 60 MHz.","fun_headline_variants_meta":{"raw":{"variants":["Streaming FPGA decoder clears 1-microsecond budget to d=21","Cryo-tested FPGA decoder streams with confidence, hits sub-microsecond","Sub-microsecond stream decode on FPGA, with confidence scores at no cost","FPGA stream decoder: d=21 simulated, d=9 cryo-tested, confidence added"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001797,"raw_usage":{"total_tokens":7065,"prompt_tokens":917,"completion_tokens":6148,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":6062}},"tokens_in":533,"tokens_out":6148,"duration_ms":38854,"temperature":1.0,"reasoning_tokens":6062,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:30:46.070106+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Verilog design for $d=11$ through place-and-route on the same FPGA family at a 200 MHz constraint; if timing closure fails or the measured average decode time per round exceeds 1 microsecond, the central extrapolation is refuted.","supporting_citations":[{"cited_title":"Snowflake: A distributed streaming decoder","cited_arxiv_id":null,"evidence_quote":"Defines the Snowflake streaming Union-Find algorithm and the 1:1 and 2:1 cluster growth schedules that this hardware implements."},{"cited_title":"Efficient post-selection for general quantum LDPC codes","cited_arxiv_id":null,"evidence_quote":"Establishes the Cluster Size 1-Norm Fraction as a useful decoder confidence score for postselection."},{"cited_title":"FPGA-based distributed union-find decoder for surface codes","cited_arxiv_id":null,"evidence_quote":"Supplies the overlapping-window FPGA Union-Find decoder whose reported throughput is a primary comparison."},{"cited_title":"QUEKUF: An FPGA union find de- coder for quantum error correction on the toric code","cited_arxiv_id":null,"evidence_quote":"Provides the FPGA Union-Find decoder on the toric code used to compare extrapolated throughput at larger distances."},{"cited_title":"Cryogenic characterization of commercial devices for application of quantum 11 computing electronics","cited_arxiv_id":null,"evidence_quote":"Documents cryogenic operation of the same Kintex UltraScale+ FPGA family, motivating the device choice and expected power savings."},{"cited_title":"Cryo- genic low-dropout voltage regulators for stable low-temperature electronics","cited_arxiv_id":null,"evidence_quote":"Provides the cryogenic low-dropout regulator design adapted for stable FPGA power delivery inside the cryostat."},{"cited_title":"Actis: A strictly local Union–Find decoder","cited_arxiv_id":null,"evidence_quote":"Supplies the circuit-level noise model used to generate the test syndromes for both simulation and hardware runs."}],"review_version":1}