{"id":"2c950006-4416-4bc1-9fc8-157020c2efdd","arxiv_id":"2607.27153","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Pruning high-weight failures to their malignant cores plus subregion MCMC (resampling ~1/w_min of locations) estimates low-p logical error rates 2–10× faster than Bravyi–Vargo MCMC.","lead":"The paper gives two practical tools for checking quantum error-correcting codes at the tiny error rates real machines need: a pruning routine that strips away harmless noise to expose the true failure core, and a faster MCMC sampler that resamples chunks of the error pattern. Together they make low-error performance checks feasible on ordinary hardware instead of huge clusters.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Core-resampling speedup is demonstrated only when w_min is already known and errors are not sparse; the debugging use-case where w_min is unknown is not closed.","rationale":"The reader correctly isolated the fluff/core heuristic and p_r≈1/w_min as the weakest assumption and already assigned CONDITIONAL with moderate confidence. My concern sharpens the same point: the speedup numbers that constitute the strongest claim were generated under conditions where w_min is known and the error density is still above t, whereas the paper’s stated purpose includes discovering unknown/low w_min and operating at utility-scale p. That is a genuine gap in the evidence for the advertised use-case, but it is not an internal contradiction of the reported experiments, nor does it overturn the pruning results or the Monte-Carlo-matched logical-error curves of Fig. 5. No code release and the use of R̂ as error proxy remain secondary engineering limitations already noted by the reader. Consequently the verdict stays CONDITIONAL; I do not push toward REJECT or ACCEPT. A single targeted re-run on the planted non-FT instance would settle whether the gap is cosmetic or material.","tokens_in":17290,"tokens_out":743,"duration_ms":48887,"concrete_test":"Re-run the Fig. 7 protocol on the intentionally non-FT d=11 surface code of Fig. 1a (missing hook). First obtain w_est via Algorithm 1 on a few hundred high-p failures; then compare decode counts to R̂≤1.05 for (i) p_r=1/6, (ii) p_r=1/w_est, and (iii) BV single-location, at the same p ladder. If (ii) loses the reported 2–10× advantage over BV, or if (i) is no better than BV, the core-resampling claim does not transfer to the unknown-w_min setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central speed claim (Fig. 7: 2–10× fewer decodes than BV to reach R̂≤1.05) is measured in the “Core Resampling” regime (p_r,p_f)=(1/w_min,p_j) of §3.2. That choice presupposes knowledge of w_min. Yet the paper’s own motivation (§1–2, Fig. 1) is that w_min is often unknown a priori and can be strictly less than t+1 precisely when the implementation is non-fault-tolerant. Fig. 7 reports surface-code and Bacon-Shor runs under bit-flip noise at p values where the expected error count still exceeds t, using the design distance (or the known decoder-limited weight) to set p_r. The paper never shows the combined pipeline—high-p Monte Carlo → Algorithm 1 prune → feed the discovered |E| into p_r → subregion MCMC at lower p—on an instance whose true w_min is initially unknown or <t+1 (e.g., the planted-hook decoder of Fig. 1a). If p_r is misspecified relative to the actual malignant core, the heuristic that “the region hits one core location on average” fails and the mixing advantage is unproven. The sparse-error regime where MCMC is most needed is explicitly deferred (§3.3). Thus the load-bearing empirical claim is only partially supported for the use-case the introduction advertises.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper addresses estimation of quantum error-correcting code performance at low physical error rates, where direct Monte Carlo is infeasible and asymptotic scaling P(p)=α p^{w_min} depends on a minimum uncorrectable weight that is often implementation-dependent and unknown. Motivated by the observation that typical failures near threshold consist of a low-weight malignant core plus many easily correctable “fluff” errors, the authors introduce (i) a pruning algorithm (Alg. 1) that iteratively removes subsets of errors from high-p Monte Carlo failures to recover candidate cores and thereby diagnose w_min or decoder/circuit bugs, and (ii) subregion MCMC, a Metropolis–Hastings family that resamples a random fraction p_r of circuit locations at a flip rate p_f, interpolating between single-site Bravyi–Vargo (BV) MCMC and full Monte Carlo. With the heuristic core-resampling choice (p_r,p_f)=(1/w_min,p_j), they report logical rates matching Monte Carlo when Gelman–Rubin R̂≤1.05 (Fig. 5) and roughly 2–10× fewer circuit simulations to convergence than BV across surface-code distances and concatenated Bacon–Shor codes (Fig. 7), under full circuit-level noise in QVM.","tokens_in":17653,"tokens_out":1519,"duration_ms":43058,"significance":"If the empirical speedups and the pruning diagnostic hold more broadly, the work supplies practical engineering tools for verifying fault tolerance and for pushing logical-error estimates into the 10^{-10} regime without heroic Monte Carlo budgets. The fluff/core decomposition is a useful organizing idea; the parameterized proposal family is a clean, reusable contribution within the Metropolis–Hastings framework; and the methods are validated against independent Monte Carlo baselines and a planted non-fault-tolerant hook bug (Figs. 1–3). These are concrete, reproducible advances over BV MCMC for circuit-level QEC simulation.","major_comments":[{"comment":"§3.2 “Core Resampling” and Fig. 7: the reported 2–10× decode reduction versus BV is measured with (p_r,p_f)=(1/w_min,p_j), which presupposes knowledge of w_min (or the design distance). The introduction and §2 motivate the work precisely by the fact that w_min is often unknown a priori and can be <t+1 for non-fault-tolerant implementations (Fig. 1a). The manuscript never demonstrates the combined pipeline—high-p Monte Carlo → Alg. 1 prune → feed discovered |E| into p_r → subregion MCMC at lower p—on an instance whose true w_min is initially unknown or strictly less than t+1. Without that, or without a clear sensitivity study for misspecified p_r, the load-bearing speed claim is only partially supported for the use-case advertised in §§1–2.","section":"§3.2, Fig. 7"},{"comment":"§3.3 explicitly defers the sparse-error regime (expected physical errors ≲ t), which is exactly where MCMC is most needed relative to Monte Carlo and where the fluff/core heuristic and the p_r≈1/w_min mixing argument are least justified. The conjecture that larger subregions cut decorrelation time is stated without quantitative mixing-time or effective-sample-size comparisons beyond R̂ as a proxy. A minimal addition—either data at lower p or an explicit scoping statement that the speedup is demonstrated only when the expected error count still exceeds t—would keep the central claim from overreaching.","section":"§3.3"},{"comment":"§3.1: statistical error on the splitting ratios is deferred; R̂ is used both as convergence diagnostic and as a proxy for uncertainty, with a citation to the R̂–ESS relation. Because temporal correlations and possible false convergence (multiple disconnected components, incomplete burn-in; cf. Fig. 6) can bias the Bennett-style ratios, the logical-error points in Fig. 5 lack rigorous error bars. At minimum the text should state what can and cannot be claimed about uncertainty from R̂ alone, or report a simple ESS/MCSE estimate on the chains already generated.","section":"§3.1"}],"minor_comments":[{"comment":"Fig. 7 switches to a bit-flip error model “since BV MCMC was originally described in terms of bit or phase flips,” while the rest of the paper (and Fig. 5) uses depolarizing noise. A sentence clarifying that the relative speedup is expected to carry over, or a single depolarizing comparison point, would remove ambiguity.","section":"§3.3, Fig. 7"},{"comment":"Alg. 1 leaves the removal strategy (subset size distribution, K, M) as free parameters with only empirical guidance (“<5”, “1000–10000 rounds”, “a few hundred patterns”). A short sensitivity paragraph or default settings used for Figs. 2–3 would aid reproducibility.","section":"§2, Algorithm 1"},{"comment":"Eq. (12)–(14): the cancellation that makes A independent of p_r is neat but compressed; a one-line remark that locations with e_ℓ=e′_ℓ contribute factors of 1 would help readers verify the algebra.","section":"§3.2"},{"comment":"Table 1 caption and body: “Prob.<w_min” at p=10^{-3} is useful; stating the Poisson or binomial model used for the tail probabilities would make the table self-contained.","section":"Table 1"},{"comment":"Minor typography: “simualtedrounds”, “thathave”, “inthe”, missing spaces after periods in several places (e.g., near Eq. (1) and the start of §3), and inconsistent italicization of p_r, p_f, w_min.","section":null},{"comment":"References [4] and [16] are concurrent arXiv preprints on related rare-event QEC simulation; a sentence distinguishing subregion MCMC from those proposal adaptations would help place the contribution.","section":"§1, References"}],"recommendation":"major_revision","confidential_remarks":"The skeptic’s w_min-circularity point is real and should be fixed before acceptance, but it is a validation/scoping gap rather than an internal contradiction: pruning is offered precisely to discover w_min, and the authors simply never close the loop in the experiments. I would not reject on that basis. Fit for a methods-focused quant-ph venue is good; novelty relative to Bravyi–Vargo plus the two concurrent rare-event papers should be stated more sharply in revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: two usable tools for circuit-level QEC engineering. Pruning turns high-p Monte Carlo failures into low-weight cores that actually expose bugs (the planted hook is clean). Subregion MCMC is a genuine new proposal family that interpolates Monte Carlo and single-flip BV, and with p_r ≈ 1/w_min they show 2–10× fewer decodes to the same R̂ on surface codes through d=17 and on concatenated Bacon-Shor, while tracking Monte Carlo logical rates where both apply.\n\nWhat is new is concrete. The fluff-versus-core observation is not deep theory, but it motivates both algorithms and is illustrated well. The acceptance-rate cancellation when p_f = p_j is a nice, clean calculation. Algorithms are stated clearly; validation against independent Monte Carlo and a planted non-FT decoder is the right kind of evidence for a methods paper. Citations to Bravyi–Vargo, Bennett, and the recent circuit-level MCMC work look appropriate; they are not reinventing the splitting method, they are improving the proposal.\n\nSoft spots, in proportion. The stress-test is partly right: Fig. 7 sets p_r from a known w_min (design or decoder-limited), while the introduction sells the case where w_min is unknown. They never close the loop on the planted-hook instance—prune first, feed |E| into p_r, then run subregion at low p. That is an incomplete demonstration of the advertised pipeline, not a contradiction of the reported speedups. Sparse-error regime and proper ESS/statistical-error bars are deferred honestly; R̂ is used as a proxy. No code release, so independent reimplementation is needed. None of that sinks the central empirical claims on the codes and regimes they actually ran.\n\nWho it is for: people sizing codes, debugging syndrome circuits and decoders, or running rare-event QEC sims. Not a new code or threshold theorem. I would bring it to reading group, cite it if I am doing circuit-level logical-rate work, and send it to referees. Accept for peer review; ask for the combined prune→subregion pipeline on a non-FT example and a clearer statement of when the 1/w_min heuristic fails.","headline":"Practical QEC methods paper: pruning finds low-weight cores and subregion MCMC is a real 2–10× speedup over BV when the core heuristic fits, with one incomplete loop on the unknown-w_min use case.","tokens_in":18338,"tokens_out":584,"would_cite":true,"duration_ms":20544,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Typical quantum error-correction failures hide a small malignant core under correctable fluff; pruning and subregion MCMC exploit that structure to certify fault tolerance and speed low-error simulations.","keywords":["quantum error correction","Markov chain Monte Carlo","fault tolerance","surface code","logical error rate","minimum-weight failures","Metropolis-Hastings","circuit-level noise"],"falsifier":"On a code family whose failures are known not to be localized cores (or at physical rates where the expected number of errors falls below t), measure the number of circuit simulations needed for Gelman–Rubin R̂ ≤ 1.05; if subregion MCMC with p_r = 1/w_min no longer requires fewer simulations than single-location Metropolis, or if pruning never recovers weight ≤ t when non-fault-tolerant hooks are deliberately inserted, the central claim fails.","tokens_in":18093,"feed_emoji":"⚛️","tokens_out":1041,"duration_ms":26156,"temperature":0.7,"pith_summary":"Utility-scale quantum algorithms need logical error rates so low that direct Monte Carlo cannot reach them, so engineers extrapolate from higher rates using the minimum weight of uncorrectable errors. That weight is often unknown once the decoder and syndrome circuit are fixed, and ordinary Monte Carlo almost never sees the rare low-weight failures. The paper shows that, in the error-rate window near threshold, a typical failure is a small malignant core surrounded by many easily correctable “fluff” errors. A simple pruning procedure strips the fluff so the core can be inspected and the true minimum weight recovered; a new family of Metropolis–Hastings moves, subregion MCMC, then resamples a controlled fraction of the lattice at each step and converges to the failure distribution far faster than single-location Metropolis. Together the two tools let implementers both debug subtle non-fault-tolerance and obtain reliable logical-error curves at the rates needed for large algorithms.","feed_headline":"Core-plus-fluff failures speed quantum-code checks 2–10×","feed_subtitle":"Pruning and subregion MCMC turn rare logical errors into everyday desktop simulations","key_machinery":"Subregion MCMC: a Metropolis–Hastings proposal that selects each circuit location independently with probability p_r and resamples the chosen subregion from the noise model at flip rate p_f. With the core-resampling choice (p_r, p_f) ≈ (1/w_min, p_j) the acceptance probability is unity for uncorrectable patterns and the chain explores distinct logical cores efficiently; the method continuously interpolates between ordinary Monte Carlo (p_r = 1) and single-location Metropolis (p_r ∼ 1/N).","core_discovery":"At physical error rates not far below threshold, uncorrectable error patterns consist of a low-weight malignant core coexisting with a large number of easily correctable fluff errors. Removing random small subsets of errors repeatedly (pruning) isolates that core and thereby reveals whether an implementation achieves the expected minimum failure weight. The same structure motivates subregion MCMC: by resampling a random fraction of locations (ideally about one core error per step) one obtains a Metropolis–Hastings chain that mixes between distinct logical failures orders of magnitude faster than single-site flips while still sampling the correct conditional distribution of failures.","pith_inferences":["Adaptive schedules that raise p_r when the expected error count drops below w_min could close the sparse-error regime the authors defer.","The same core-extraction idea may let decoders themselves reject fluff on the fly, turning a diagnostic into an online decoder improvement.","If the mixing-time advantage scales with distance as suggested by the surface-code data, subregion MCMC becomes the practical default for any large-distance threshold study."],"forward_implications":["A few hundred pruned Monte Carlo failures suffice to certify that an implementation meets its design distance or to expose the exact low-weight bug.","Logical-error curves at 10^{-10} and below become obtainable on a desktop for surface and concatenated Bacon–Shor codes under full circuit noise.","The continuum of proposal distributions between Monte Carlo and single-site MCMC can be tuned by two scalar parameters without redesigning the sampler.","Any code whose failure geometry is core-plus-fluff inherits the same 2–10× reduction in simulation count relative to Bravyi–Vargo MCMC."],"fun_headline_variants":["Pruning isolates malignant cores in quantum error patterns","Subregion MCMC resamples fractions for faster failure mixing","Core-plus-fluff structure accelerates low-error QEC estimates","MCMC pruning reveals true minimum weights in code implementations","Resampling error subregions beats single-flip Metropolis for QEC"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim that failures are spatially localized cores plus fluff, and that resampling roughly one core location per step remains optimal, must hold for the codes, decoders, and sparse-error regimes of interest.","fun_headline_variants_meta":{"raw":{"variants":["Pruning isolates malignant cores in quantum error patterns","Subregion MCMC resamples fractions for faster failure mixing","Core-plus-fluff structure accelerates low-error QEC estimates","MCMC pruning reveals true minimum weights in code implementations","Resampling error subregions beats single-flip Metropolis for QEC"]},"model":"grok-4.5","effort":"low","cost_usd":0.004042,"raw_usage":{"total_tokens":1304,"prompt_tokens":891,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":40424000,"prompt_tokens_details":{"text_tokens":891,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":343,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":891,"tokens_out":70,"duration_ms":7184,"temperature":1.0,"reasoning_tokens":343,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T10:51:47.556981+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a code family whose failures are known not to be localized cores (or at physical rates where the expected number of errors falls below t), measure the number of circuit simulations needed for Gelman–Rubin R̂ ≤ 1.05; if subregion MCMC with p_r = 1/w_min no longer requires fewer simulations than single-location Metropolis, or if pruning never recovers weight ≤ t when non-fault-tolerant hooks are deliberately inserted, the central claim fails.","supporting_citations":[],"review_version":1}