{"id":"d9e905bd-530c-4d9a-90f2-f01a9bad6e6b","arxiv_id":"2607.10058","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"AI pre-decoders plus Chromobius cut color-code logical failure rates by up to 347x and runtime by 7.33x at d=31, p=0.3%, improving with distance.","lead":"AI pre-decoders for triangular color codes cut logical errors by hundreds of times and speed decoding as distance grows, working with parallel block decoding. This could make color codes competitive with surface codes for large fault-tolerant quantum computers.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Abstract-only review cannot verify the 347x LER claim; the load-bearing gap is whether training-data simplification and local pre-decoder accuracy hold under the (unstated) noise model and parallel block-wise evaluation at d=31.","rationale":"The Reader's verdict is already CONDITIONAL with LOW confidence precisely because the abstract alone cannot substantiate the training-data simplification, architecture, noise model, or validation protocol that underwrite the 347x figure. That is the correct posture for an abstract-only review: the direction is plausible and the numbers, if real, would be significant, but they remain unverified. No stronger internal inconsistency can be diagnosed without the full text; manufacturing one would violate the good-faith rule. Therefore the stress-test leaves the verdict unchanged and records agreement with the Reader's identification of the weakest assumption. The concrete test above is the minimal check that would settle whether that assumption lands once the paper becomes available.","tokens_in":2153,"tokens_out":493,"duration_ms":4170,"concrete_test":"Once the full paper (or code) is available: retrain the pre-decoder on the unsimplified feedforward circuit data (or on an independent circuit-level noise model matching the evaluation conditions) and re-evaluate LER at d=31, p=0.3% under the same parallel block-wise schedule used for the Chromobius baseline. If the LER gain falls below ~10x or the distance-scaling trend reverses, the headline claim does not hold under realistic training.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the assertion that a novel NN architecture, trained on simplified data from color-code syndrome-extraction circuits that include feedforward, produces local spacelike and timelike corrections that remain accurate and compatible with parallel block-wise decoding when composed with Chromobius. Because only the abstract is available, the noise model, training distribution, validation protocol, error bars, and any ablation of the simplification step are all unstated. Without those, it is impossible to confirm that the reported 347x LER improvement (and the claimed improvement-with-distance trend) is not an artifact of training/evaluation mismatch or of an overly optimistic noise model. The Reader correctly flags this as the weakest assumption; it is also the single most load-bearing concern for the headline numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes AI-based pre-decoders for triangular color codes, motivated by the need for local corrections that are compatible with parallel block-wise decoding in space and time and with lattice-surgery protocols. It introduces a novel neural-network architecture and methods to simplify training data arising from color-code syndrome-extraction circuits that include feedforward operations. The central empirical claim is that a pre-decoder + Chromobius pipeline improves both logical failure rate and runtime relative to raw Chromobius, with the gap widening as code distance grows; the headline numbers are a 347× LER improvement and a 7.33× runtime reduction at d=31 and p=0.3%.","tokens_in":2351,"tokens_out":1004,"duration_ms":16396,"significance":"If the quantitative claims hold under a clearly specified noise model and evaluation protocol, the work would be a meaningful step toward making color codes competitive with surface codes for large-scale FTQC, leveraging color codes’ advantages in lattice surgery and transversal Cliffords. The framing of pre-decoders as local spacelike/timelike correctors that compose with an existing decoder (Chromobius) and with parallel block-wise schemes is a useful architectural contribution. Credit is due for targeting the missing framework for AI-based decoding under parallel space–time blocking, and for reporting simultaneous LER and runtime gains that improve with distance—an unusual and potentially high-impact combination if verified.","major_comments":[{"comment":"The headline claim (abstract: 347× LER improvement and 7.33× runtime reduction at d=31, p=0.3% vs raw Chromobius) is load-bearing for the paper’s central result, but only the abstract is available for review. Without the full methods, noise model, baseline configuration of Chromobius, error bars/confidence intervals, number of Monte Carlo shots, and training/validation/test splits, these factors cannot be verified and must be treated as unverified. A complete evaluation section with tables and uncertainty quantification is required before the claim can be assessed.","section":null},{"comment":"The abstract states that methods were developed to simplify complex training data from feedforward syndrome-extraction circuits, and that the resulting pre-decoders remain accurate for local spacelike and timelike corrections under parallel block-wise decoding. This simplification-and-generalization step is the weakest load-bearing assumption: if training and evaluation distributions differ (or if feedforward simplification removes error mechanisms present at d=31), the reported LER gains could be artifacts. The manuscript must specify the noise model, the exact simplification procedure, and an ablation or hold-out validation showing that local corrections remain accurate when composed with Chromobius at the reported distances.","section":null},{"comment":"The abstract asserts that both LERs and runtimes improve relative to raw Chromobius as code distance increases. That trend is central to the claim that AI pre-decoding narrows the color-code vs surface-code gap at scale. Supporting distance-scaling plots (LER and wall-clock or cycle time vs d), with the same decoder settings and noise model across distances, are needed; without them the “improves with distance” claim cannot be checked and the d=31 point remains an isolated number.","section":null}],"minor_comments":[{"comment":"Abstract only: define or briefly name the noise model (e.g., circuit-level depolarizing with measurement errors) and whether thresholds or only fixed-p LER comparisons are reported, so readers can place the 0.3% operating point.","section":null},{"comment":"Abstract only: clarify what “runtime” measures (decoder wall-clock per shot, amortized over blocks, including or excluding NN inference) and on what hardware, so the 7.33× factor is interpretable.","section":null},{"comment":"Abstract only: a one-line description of the novel NN architecture (e.g., input features, locality of receptive field, separate spacelike vs timelike heads) would help readers assess compatibility with block-wise parallel decoding without waiting for the full methods section.","section":null}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was available; the full text was not provided. An abstract-only review cannot responsibly accept or reject the quantitative claims. I recommend the editor obtain the full manuscript (methods, noise model, training protocol, ablations, and distance-scaling data) and re-review. If the full paper does not supply those elements, major_revision or reject would be appropriate; if it does and the numbers hold, the contribution looks journal-worthy. Fit for quant-ph / QEC venues is good if the evaluation is solid."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: the abstract says an AI pre-decoder plus Chromobius gives 347x better logical failure rate and 7.33x faster runtime than raw Chromobius at d=31, p=0.3%, with both metrics improving as distance grows. That would matter a lot for color codes if it holds. We only have the abstract, so treat the numbers as unverified claims for now.\n\nWhat looks new and useful is the framing itself. They treat the AI piece as a local pre-decoder that does spacelike corrections on qubits and timelike corrections on stabilizers, so it slots into parallel block-wise decoding and lattice surgery without fighting the global decoder. They also claim a novel NN architecture plus a concrete way to simplify the messy training data that comes from color-code syndrome circuits with feedforward. That combination is a real engineering target: color codes already have transversal Cliffords and simpler surgery, but decoding has been the drag. If the pre-decoder actually leaves Chromobius an easier residual problem and the advantage grows with d, that narrows a long-standing practical gap with surface codes.\n\nSoft spots are exactly the ones you expect from abstract-only. Noise model, training distribution, validation splits, error bars, ablations on the data-simplification step, and any code are all missing. The load-bearing assumption is that the local corrections stay accurate and do not inject new correlations that Chromobius then mishandles at large distance. The improvement-with-distance trend is the most interesting claim and also the easiest place for a training/evaluation mismatch to hide. Circularity risk looks mild rather than fatal, but it is there until someone checks hold-out behavior.\n\nThis is for people building FTQC stacks who still care about color codes, and for decoder engineers who need something that composes with existing tools like Chromobius. It is not a theory paper; it is an empirical systems claim. It deserves a serious referee once the full manuscript appears with the missing pieces. I would not cite the headline numbers from the abstract alone, but I would read the paper carefully and send it to review rather than desk-reject. The problem is important enough that the claims should be stress-tested.","headline":"Abstract-only: big claimed LER/runtime wins for color-code AI pre-decoders that improve with distance, but the 347x figure and training simplifications are unverifiable without methods.","tokens_in":2976,"tokens_out":555,"would_cite":false,"duration_ms":11653,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"AI pre-decoders for triangular color codes cut logical failure by 347x and runtime by 7.33x versus raw Chromobius at distance 31, with gains growing as codes get larger.","keywords":["color codes","AI pre-decoders","neural-network decoding","fault-tolerant quantum computing","quantum error correction","Chromobius","lattice surgery","logical error rates"],"falsifier":"Reproduce the d=31, p=0.3% experiment under the paper's noise model and training protocol: if the pre-decoder-plus-Chromobius pipeline fails to improve logical failure rate by roughly two orders of magnitude and runtime by several times relative to raw Chromobius, the central claim is false.","tokens_in":3000,"feed_emoji":"⚛️","tokens_out":873,"duration_ms":13634,"temperature":0.7,"pith_summary":"Color codes offer simpler lattice surgery and transversal Clifford gates than surface codes, yet they have lagged because decoding is slower and logical failure rates are worse. This paper claims that local AI-based pre-decoders can close that gap while remaining compatible with the parallel space-and-time block decoding required for large-scale fault-tolerant computation. The authors introduce a neural-network architecture that applies spacelike corrections on physical qubits and timelike corrections on stabilizer measurements, together with methods that simplify the otherwise complex training data generated by color-code circuits containing feedforward operations. When the resulting pre-decoder is pipelined with Chromobius, both logical error rates and wall-clock runtimes improve relative to raw Chromobius, and the advantage grows with code distance. At distance 31 and physical error rate 0.3 percent the pipeline improves logical failure by a factor of 347 while reducing runtime by a factor of 7.33, bringing color codes closer to practical use.","feed_headline":"AI pre-decoders cut color-code failures 347x at d=31","feed_subtitle":"Local neural corrections plus Chromobius also run 7x faster as distance grows, closing the gap to surface codes","key_machinery":"The AI pre-decoder itself: a neural network that outputs local spacelike corrections on data qubits and timelike corrections on stabilizer outcomes; its locality makes it natively compatible with parallel block-wise decoding and lattice-surgery protocols, while the authors' data-simplification methods render the feedforward training circuits tractable.","core_discovery":"A novel neural-network pre-decoder for triangular color codes, trained on simplified data from feedforward syndrome-extraction circuits, produces local spacelike and timelike corrections that, when followed by Chromobius, simultaneously lower logical failure rates by hundreds of times and accelerate decoding, with both metrics improving as distance increases.","pith_inferences":["The same local-pre-decoder-plus-classical-decoder pipeline may transfer to other stabilizer codes whose syndrome-extraction circuits contain feedforward.","The training-data simplification technique could be reused for any quantum circuit whose measurement record is entangled by classical feedforward.","If the observed improvement continues past d=31, color codes could become preferable to surface codes for architectures that prioritize transversal Cliffords or reduced lattice-surgery overhead."],"forward_implications":["Color-code logical error rates become low enough that their transversal Clifford gates and simpler lattice-surgery protocols can be used at scale.","Decoding wall-clock time improves with distance rather than worsening, enabling larger codes under fixed latency budgets.","Parallel space-and-time block decoding schemes required for lattice surgery become practical for color codes.","The historical performance gap between color codes and surface codes narrows substantially at high distance."],"fun_headline_variants":["AI pre-decoders cut color-code LERs 347x at d=31 while speeding up","Neural pre-decoding yields 347x lower color-code failures at d=31","Local AI corrections slash color-code LER 347x and runtime 7x","Color-code pre-decoders improve LER 347x as distance grows","AI pre-decoders plus Chromobius: 347x better LER at d=31"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The simplified training data and novel network architecture produce pre-decoders whose local corrections remain accurate and compatible with parallel decoding under the noise model used for the reported distance-31 benchmarks.","fun_headline_variants_meta":{"raw":{"variants":["AI pre-decoders cut color-code LERs 347x at d=31 while speeding up","Neural pre-decoding yields 347x lower color-code failures at d=31","Local AI corrections slash color-code LER 347x and runtime 7x","Color-code pre-decoders improve LER 347x as distance grows","AI pre-decoders plus Chromobius: 347x better LER at d=31"]},"model":"grok-4.5","effort":"low","cost_usd":0.00425,"raw_usage":{"total_tokens":1345,"prompt_tokens":859,"num_sources_used":0,"completion_tokens":119,"cost_in_usd_ticks":42500000,"prompt_tokens_details":{"text_tokens":859,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":367,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":859,"tokens_out":119,"duration_ms":3147,"temperature":1.0,"reasoning_tokens":367,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T00:41:44.503207+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Reproduce the d=31, p=0.3% experiment under the paper's noise model and training protocol: if the pre-decoder-plus-Chromobius pipeline fails to improve logical failure rate by roughly two orders of magnitude and runtime by several times relative to raw Chromobius, the central claim is false.","supporting_citations":[],"review_version":1}