{"id":"384d139f-a85d-42ba-bc3a-a007616b9c38","arxiv_id":"2603.06730","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Under a matched LiDMaS+ protocol, matching-style decoding yields higher Pauli-reference thresholds than Union-Find, while hybrid CV-discrete crossings remain grid- and estimator-sensitive and require fallback diagnostics.","lead":"Surface-code threshold numbers change with the decoder and the analysis method used to extract them, under both standard Pauli noise and a digitized hybrid continuous-variable noise model. The work matters because quantum-computing roadmaps treat those threshold numbers as design inputs, so decoder- and estimator-dependence is an auditability issue.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the caveats already front-and-center in the reader verdict.","rationale":"The reader correctly identifies the paper as a careful methods/simulation study whose strongest claim is that threshold estimates depend on the inference pipeline and should be audited with fallback diagnostics. The Pauli-reference matching-style results (pc≈0.053, collapse consistent) and the hybrid interior crossings with explicit estimator caveats are presented as protocol-conditional, not as material constants. The high d=9 fallback rate and the low-LER (d=5,7) sensitivity are already reported and used to qualify interpretation. No internal contradiction, circularity, or unacknowledged assumption threatens the central methodological conclusion. Therefore the CONDITIONAL verdict (accept the methods point once caveats stay front-and-center) needs no adjustment. The concrete test above would only further insulate the claim from the already-flagged backend-proxy concern.","tokens_in":15767,"tokens_out":485,"duration_ms":4598,"concrete_test":"Re-run the dense hybrid transition-window (d=3,5,7, σ∈[0.30,0.50], step 0.01, 3000 trials) with a production Blossom/MWPM backend (zero greedy fallback) under the same seeds and digitization map; if decoder ordering and the existence of interior crossings after plateau exclusion are preserved, the methodological claim is confirmed independent of the fallback policy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s central claim is methodological: threshold summaries are pipeline outputs and must be reported with decoder, estimator, grid resolution, and fallback diagnostics. That claim is supported by the matched LiDMaS+ experiments and is stated with the right qualifications (interior vs boundary crossings, low-LER estimator sensitivity, d=9 fallback rates up to 0.747). The reader’s weakest_assumption correctly flags that the matching-style backend is not production Blossom and that the hybrid map is digitized without soft information; the manuscript itself already labels these limits (Table I, §II matching description, Table VII, Discussion). Within the stated scope there is no hidden inconsistency that would overturn the methodological point. Hybrid σc numbers are not offered as decoder-independent constants, so the proxy concern does not undercut the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript argues that surface-code threshold estimates are outputs of an inference pipeline and therefore depend on decoder backend, estimator, sweep resolution, and statistical budget. Within a single LiDMaS+ harness with matched seeds, grids, and trial counts, the authors compare a matching-style backend (exact small-instance solver with greedy fallback), Union-Find, and a lightweight neural-guided reweighter under a Pauli-reference channel and a digitized hybrid continuous-variable/discrete channel that maps Gaussian displacements to parity-mapped Pauli faults without soft information. In the Pauli-reference mode they report a matching-style crossing median pc=0.0531 [0.0415,0.0572] and collapse fit pc=0.052 (ν=1.35), while Union-Find yields no stable crossing. In a dense hybrid transition window they report interior matching-style crossings σc=0.4707 (d=3,5) and σc=0.3275 (d=5,7), the latter flagged as low-LER and estimator-sensitive; a d=9 extension shows high matching-fallback rates (up to 0.747) and larger Union-Find LER; a d=5 guidance sweep shows a modest mean-LER reduction under full reweighting. The methodological conclusion is that estimator resolution and backend fallback diagnostics belong in an auditable decoder comparison.","tokens_in":16033,"tokens_out":1595,"duration_ms":43086,"significance":"If the results hold under the stated protocol, the paper makes a useful methodological contribution: it treats threshold estimation as a controlled, reproducible pipeline comparison rather than as a decoder-free material constant, and it documents fallback and decoder-failure diagnostics alongside LER. The matched-seed design, Wilson intervals, explicit exclusion of exact-zero plateaus before crossing localization, and joint reporting of crossing versus collapse summaries for the Pauli control are strengths that support auditability. The absolute Pauli pc≈0.05 is well below standard MWPM bit-flip thresholds, and the hybrid model is deliberately digitized without soft information, so the work does not claim production decoder performance or a full GKP-style analog decoder study. Within that scope, the controlled comparison and the insistence on reporting estimator and fallback metadata are of practical value to groups that publish finite-distance threshold summaries under nonstandard noise maps.","major_comments":[{"comment":"The hybrid noise map is central to half the study (Table I; hybrid mode throughout §§III–IV) but is specified only as “Gaussian→parity-mapped Pauli” with no soft stream. An explicit, self-contained digitization rule (e.g., the quadrature-to-Pauli binning or parity map, including any GKP-style modular reduction and the precise relation between σ and the effective Pauli rate) is needed so that the dense-window σc values and multi-distance reversals can be reproduced without the full LiDMaS codebase. Without that equation, the hybrid crossings remain protocol-internal numbers rather than independently checkable results.","section":"Table I / §II hybrid mode"},{"comment":"The matching-style backend is exact only below a configured small-instance limit and uses a greedy fallback above it; the paper correctly reports a maximum fallback rate of 0.747 at d=9 (Table VII) and states that results are not production Blossom benchmarks. Fallback rates are not tabulated for the dense d=3,5,7 transition window or the coarse multi-distance hybrid sweeps that supply the main decoder-ordering and σc claims (Fig. 3, Table VI, Table IV). Because high-σ, larger-d points are exactly where defect sets grow, those LER curves and the (d=5,7) crossing may mix exact matching with the greedy policy. Reporting fallback rate versus (d,σ) for every hybrid curve used in crossing localization is load-bearing for interpreting “matching-style vs Union-Find” as a decoder-family comparison rather than a fallback-policy comparison.","section":"§II matching description; Table VI; Table VII"},{"comment":"The Pauli-reference matching-style crossing median pc=0.0531 and collapse pc=0.052 (Table V, Fig. 5) are substantially below the ~10% independent bit-flip MWPM thresholds standard in the literature for the surface code. The manuscript notes that the backend is not production Blossom, but it does not discuss this absolute discrepancy as a control-check on the protocol. A short quantitative discussion—attributing the reduction to the greedy fallback, the restricted noise branch (X→Z only), finite-distance effects, or another identified cause—would strengthen the claim that the Pauli mode is a “stable comparison point” rather than an anomalous internal control.","section":"Table V / §IV.C Pauli threshold"}],"minor_comments":[{"comment":"Several figure panels in the compiled text appear with corrupted axis labels or placeholder glyphs (e.g., Figs. 2–5 in the source dump). Ensure final production figures have legible axis labels, legends, and Wilson-band captions.","section":"Figs. 2–5"},{"comment":"Table IX compares heterogeneous published thresholds and metrics; the caption already warns against one-to-one benchmarking, but a column for noise model and decoder family would make the non-comparability more transparent.","section":"Table IX"},{"comment":"The neural-guided model is described as a linear reweighter trained simulator-in-the-loop (Alg. 4; Eqs. 7–8 in the appendix), but training set size, regularization, and whether the sensitivity sweep uses a held-out σ grid are not stated. A brief note would clarify that the 0.1773→0.1663 mean-LER change is an internal sensitivity result, not a learned-decoder benchmark.","section":"§II neural-guided matching; Table VIII"},{"comment":"Eq. (12) linearizes the crossing between adjacent grid points after plateau exclusion. State explicitly whether the reported σc values use the midpoints of the 0.01 grid or the continuous interpolant, and give the two bracketing (σ, ΔLER) pairs for the (d=5,7) estimate so readers can judge low-LER sensitivity.","section":"Eq. (12); Table VI"},{"comment":"Minor typography: “EXPERIMENT AL DESIGN”, “RESUL TS”, and “V alidity” spacing artifacts in headings should be cleaned for the camera-ready version.","section":"§III–V headings"}],"recommendation":"minor_revision","confidential_remarks":"The central methodological claim is sound and carefully qualified; I do not see a hidden inconsistency that would justify rejection. Novelty is moderate—decoder dependence of thresholds is known—but the auditability framing (fallback rates, plateau exclusion, matched harness) is a legitimate contribution for a methods-oriented quant-ph venue. Fit is better for a specialized QEC/methods journal than for a broad high-impact letter. Self-citation to LiDMaS is appropriate because it is the experimental harness. If the journal expects production-decoder or soft-information GKP results, the scope mismatch should be weighed by the editor; within the manuscript’s stated scope, minor revision is sufficient."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a methods paper with a clean punchline: surface-code threshold numbers are pipeline outputs, and if you do not report decoder backend, estimator, grid resolution, and fallback rates, the comparison is not auditable. That claim is supported by the experiments they actually ran.\n\nWhat is new is not “MWPM beats Union-Find” — that is already in the literature they cite. What is new is a single LiDMaS+ harness with matched seeds, grids, and trial budgets under both a Pauli reference and a digitized Gaussian→parity-mapped-Pauli hybrid map, plus explicit exclusion of exact-zero plateaus before crossing localization and explicit matching-fallback / decoder-failure diagnostics. The Pauli-reference side is the solid control: matching-style gives crossing median pc=0.0531 [0.0415,0.0572] and collapse pc=0.052 (ν=1.35), mutually consistent at the reported resolution; Union-Find does not yield a stable scalar on the same grid. The hybrid dense window then produces interior crossings (σc≈0.47 for d=3,5; ≈0.33 for d=5,7) that the authors themselves flag as low-LER and estimator-sensitive, and the d=9 extension shows fallback rates up to 0.747. That honesty is the paper’s main virtue.\n\nSoft spots are real but already on the page. The matching backend is exact only for small defect sets and greedy above a cutoff, so high-σ / high-d LER is partly fallback policy. The hybrid mode throws away continuous soft information. Distances stop at 9. No code is shipped, so reproducibility is reimplementation-level. None of that overturns the methodological claim; it just means the hybrid σc numbers are not material constants and should not be quoted as such.\n\nMath and estimators look standard (Wilson intervals, linearized crossings, collapse ansatz). Citations are appropriate. Self-citation to LiDMaS is the harness, not a circular prior threshold.\n\nWho it is for: people who report or compare surface-code thresholds, especially under hybrid/GKP-style digitization. A serious referee should see it. I would engage with the reporting standard; I would not treat the hybrid crossings as decoder-family truth without a production matching backend and soft-info follow-up.","headline":"Careful matched-decoder study that makes a real methodological point about threshold auditing under digitized hybrid noise, with hybrid numbers that stay estimator- and fallback-conditional.","tokens_in":16666,"tokens_out":568,"would_cite":true,"duration_ms":6239,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Surface-code threshold numbers are outputs of a decoder-and-estimator pipeline, not fixed properties of the code and noise alone.","keywords":["surface code","threshold estimation","decoder dependence","matching decoder","Union-Find","hybrid continuous-variable noise","logical error rate","finite-size scaling"],"falsifier":"Rerun the same hybrid dense window and distance-9 points with a production minimum-weight matching backend that never falls back to the greedy path, and check whether the (d=5,7) interior crossing stays near 0.33, moves toward the (d=3,5) value, or disappears once fallback rates are near zero.","tokens_in":16600,"feed_emoji":"⚛️","tokens_out":1009,"duration_ms":16080,"temperature":0.7,"pith_summary":"This paper argues that when you estimate a surface-code threshold, the number you get depends on which decoder you use and how you locate the crossing in finite data. The authors run matching-style decoding, Union-Find, and a light neural reweighting of matching inside one shared workflow, first under ordinary Pauli noise and then under a hybrid model that turns continuous Gaussian displacements into discrete faults. Matching consistently beats Union-Find under Pauli noise and yields a crossing near physical error rate 0.05; under hybrid noise the same backends give interior crossings that split by distance pair and stay sensitive to the grid and estimator. At larger distance the matching path falls back to a greedy rule for many shots, so the reported logical-error curves mix exact and approximate behavior. The practical claim is that auditable threshold work must publish decoder choice, sweep resolution, and backend fallback rates alongside the scalar number.","feed_headline":"Surface-code thresholds shift with the decoder you pick","feed_subtitle":"Matched Pauli and hybrid sweeps show matching beats Union-Find and hybrid crossings stay grid-sensitive.","key_machinery":"A controlled LiDMaS+ comparison protocol: identical seeds, distance sets, and noise grids across backends, with hybrid Gaussian noise first mapped to parity-encoded Pauli faults (no continuous soft stream), and with hybrid crossing localization that drops the initial exact-zero logical-error plateau before interpolating sign changes or reporting minimum-separation proxies.","core_discovery":"Under a single matched workflow, decoder and estimator choice change surface-code threshold summaries for both Pauli-reference noise and digitized hybrid continuous-variable noise. Matching-style decoding outperforms Union-Find on Pauli sweeps and returns a crossing median pc of 0.0531 with a consistent collapse fit near 0.052; hybrid dense-window interior crossings for matching are about 0.47 for distances 3 and 5 and about 0.33 for distances 5 and 7, with the latter low-error estimate remaining estimator-sensitive, while high fallback rates at distance 9 limit how far the hybrid numbers can be pushed.","pith_inferences":["If hybrid architectures keep digitizing continuous noise before decoding, threshold claims will need separate analog-aware and digitized pipelines rather than one shared discrete decoder stack.","Adaptive trial budgets concentrated where distance curves nearly touch would likely shrink the gap between coarse-grid proxies and interior hybrid crossings more than simply raising total shots.","Once production matching removes greedy fallback, remaining decoder gaps between matching and Union-Find under hybrid noise would cleanly isolate approximation quality from implementation artifacts."],"forward_implications":["Published surface-code thresholds should name the decoder, estimator, grid resolution, and fallback or decoder-failure rates, not only a scalar pc or σc.","Union-Find can preserve qualitative distance reversal under hybrid noise while still inflating logical error at larger distance and moderate-to-high noise relative to matching.","Hybrid crossings extracted after dropping exact-zero plateaus remain pair-dependent and low-error estimates stay grid-sensitive, so they should not be treated as a single converged critical point.","Lightweight learned reweighting of matching edges can modestly lower sampled mean logical error without becoming an end-to-end neural decoder.","High matching-fallback rates at larger distance make those logical-error curves partly measures of the fallback policy, not of the exact matching objective alone."],"fun_headline_variants":["Decoder choice shifts surface-code thresholds in matched sweeps","Matching beats Union-Find on Pauli surface-code pc estimates","Hybrid crossings stay estimator-sensitive at low LER","Matching backend yields pc=0.0531 over Union-Find on Pauli noise","Decoder and estimator alter hybrid sigma_c under dense sweeps"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That a matching backend which solves only small defect sets exactly and falls back to a greedy rule, plus a hybrid model that feeds only digitized Pauli events without soft information, is still a fair enough stand-in for production matching and for GKP-style continuous noise that decoder-family effects can be read off the reported curves.","fun_headline_variants_meta":{"raw":{"variants":["Decoder choice shifts surface-code thresholds in matched sweeps","Matching beats Union-Find on Pauli surface-code pc estimates","Hybrid crossings stay estimator-sensitive at low LER","Matching backend yields pc=0.0531 over Union-Find on Pauli noise","Decoder and estimator alter hybrid sigma_c under dense sweeps"]},"model":"grok-4.5","effort":"low","cost_usd":0.00427,"raw_usage":{"total_tokens":1361,"prompt_tokens":926,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":42700000,"prompt_tokens_details":{"text_tokens":926,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":367,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":926,"tokens_out":68,"duration_ms":3632,"temperature":1.0,"reasoning_tokens":367,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T14:10:44.421333+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the same hybrid dense window and distance-9 points with a production minimum-weight matching backend that never falls back to the greedy path, and check whether the (d=5,7) interior crossing stays near 0.33, moves toward the (d=3,5) value, or disappears once fallback rates are near zero.","supporting_citations":[],"review_version":1}