{"id":"eea2e2da-11c7-4916-ba1f-fa4321b6dc28","arxiv_id":"2608.11516","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Trapping-set analysis on detector error models predicts the error floor of iterative quantum decoders, with RelayBP matching simulation closely and two other decoders within one order of magnitude.","lead":"The paper applies classical trapping-set enumeration directly to circuit-level detector error models of a quantum LDPC code, and predicts the low-error-rate error floors of three iterative decoders from a fixed structural catalog. For the RelayBP decoder the prediction matches Monte Carlo simulation; for two others it is within an order of magnitude.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The LETS-completeness assumption is load-bearing: the lower-bound estimate omits non-leafless, non-elementary, size>5, and weight-five supports, and the paper's own ImpulseBP/ELMS gaps show this omission is real; only the single RelayBP match supports the general claim.","rationale":"The reader's weakest assumption already identifies the LETS-completeness restriction, and I agree with that choice. The proposed membership audit is concrete and implementable with the public catalog plus the decoder code; it cleanly separates statistical fluctuation from structural incompleteness and directly tests whether the lower bound is tight because the enumerated universe contains the dominant failing supports. The alternative concern about missing Monte Carlo error bars is real but secondary: even perfect Monte Carlo statistics would not establish the framework if the enumerated universe excludes the dominant failure supports for the general decoder class. The paper's own limitation statements—the 'not clear a priori' caveat in the introduction and the explicit acknowledgment that ImpulseBP failures may require a=6, ETS with leaves, or non-elementary TS—support this reading. The verdict should remain CONDITIONAL because the framework is demonstrated for RelayBP on one code, but the general error-floor prediction claim needs either the completeness audit or an explicit scope reduction.","tokens_in":27247,"tokens_out":9198,"duration_ms":113611,"concrete_test":"Use the released enumeration catalog to audit Monte Carlo failure supports: record the exact DEM fault support for every simulated logical failure of ImpulseBP and ELMS at p=1e-4 and p=1e-3, and test membership of each support in the union U of enumerated LETS supports (e.g., by hashing the 92M-instance catalog). Compute the fraction of Monte Carlo logical-failure probability attributable to supports outside U. If this fraction is non-negligible, the lower-bound gap is structural incompleteness; if it is negligible, the gap is instead due to un-injected weight-five or higher-weight patterns, and the paper should re-state its conclusions accordingly. For RelayBP, the same audit would confirm whether the match is exact because all simulated failures fall inside U.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is completeness of the tested support universe. Equation (3) is exact only as a lower bound; it becomes a prediction of the error floor only if the omitted mass—non-LETS supports, LETS with a>5 or b>5, and weight-five fault patterns that are never injected—is negligible. The paper itself flags this assumption as 'not clear a priori' for QLDPC DEMs, and its own results show it is violated for two of the three decoders: Figures 2 and 3 lie systematically above the LETS estimates, and the text attributes the gap to failures not captured by the current LETS analysis, possibly requiring a=6, ETS with leaves, or non-elementary TS. Thus the only clean validation of completeness is the RelayBP match on a single [[144,12,12]] code with priors fixed at p0=1e-3. That validates the framework for one decoder/code configuration, but not the general claim that LETS 'capture a substantial part' of the error floor or that the shortfall remains within one order of magnitude across decoders and codes. The enumeration is also truncated at a<=5, b<=5 for memory reasons (92,088,583 instances), so the central claim rests on an untested truncation for other codes and on excluding weight-five fault patterns even though the code distance would, in principle, permit correction up to weight five.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a framework for analyzing low-error-rate failures of iterative decoders under circuit-level noise by applying leafless elementary trapping set (LETS) enumeration directly to the detector error model (DEM) of a [[144,12,12]] bivariate bicycle code. The authors enumerate 92,088,583 LETSs with a_max = b_max = 5, exhaustively inject weight-three and weight-four fault patterns supported on these structures, and compute a lower bound on the logical error rate from the noise model without fitting any parameter to simulation. The resulting estimate is compared with Monte Carlo simulations of three decoders: RelayBP, ImpulseBP, and a new ensemble layered min-sum decoder (ELMS). The estimate closely matches the RelayBP curve, lies within one order of magnitude for ImpulseBP and ELMS, and the failure statistics are strongly concentrated in a small number of LETS topology classes that are partly shared across decoders.","tokens_in":27539,"tokens_out":10988,"duration_ms":127323,"significance":"If the central claim is sustained, the framework offers a decoder-agnostic structural route to error-floor estimation that avoids the prohibitive cost of direct Monte Carlo simulation in the low-error regime, with no free parameters fitted to the simulated curves. The paper ships its enumeration code and dataset, inherits a proof of exhaustive enumeration of LETSs from the classical LDPC literature by Hashemi and Banihashemi, and provides a new simple decoder (ELMS) as a benchmark. The topology-level analysis, showing that only a small fraction of enumerated LETS structures are harmful and that certain topologies are harmful across decoders with very different heuristics, is a concrete and useful structural insight for decoder design. The main limitation is that the predictive claim rests on a single code and on a completeness assumption that the paper itself shows is violated for two of the three decoders; the strength of the claims in the abstract and introduction should be aligned with the actual evidence.","major_comments":[{"comment":"The paper's abstract and introduction assert that trapping-set analysis 'predicts' the error floor, but the evidence supports this only for RelayBP on the [[144,12,12]] code. For ImpulseBP and ELMS, Figures 2 and 3 show the LETS estimate systematically below the Monte Carlo curve, and the text attributes the gap to structures outside the enumerated universe (a=6, ETS with leaves, or non-elementary TS). Since Eq. (3) is only a lower bound unless the omitted support mass is negligible, the completeness of the LETS universe is exactly the load-bearing assumption that converts the bound into a prediction. Please either temper the abstract and Section I claims to 'lower-bound prediction' and 'partial structural characterization', or provide quantitative evidence on the missing mass, for example by a partial enumeration at a=6 or by estimating the contribution of non-LETS supports. Without this, the 'practical framework for predicting error floors' claim is not supported by the reported data.","section":"Section I.A and Section V.A; Eq. (3)"},{"comment":"No error bars or confidence intervals are reported for either the Monte Carlo data or the LETS estimates. For the randomized decoders (RelayBP and ELMS), Eq. (5) is a single Bernoulli draw per fault support; Eq. (6) gives a variance bound but is never evaluated, and the figures show a single realization of the estimator. The statement that the RelayBP estimate 'closely agrees' with simulation is therefore not a statistical statement. Please add (a) the physical error-rate range and the number of trials/failures for each Monte Carlo point, (b) confidence bands or standard errors for the estimates, and (c) a statement of how many independent realizations of the randomized estimator are shown. This is necessary for the reader to judge whether the discrepancy between the estimate and simulation is significant.","section":"Section IV.B and Figures 1-3"},{"comment":"The figures display per-round logical error rates down to 10^-15, yet the Monte Carlo protocol described in Section V uses at least 10,000 trials per point, which cannot produce 100 failures at such low rates. It is unclear over which physical error rates the Monte Carlo curves were actually simulated and which parts, if any, are extrapolations; if extrapolated, the fitting procedure should be stated. In addition, Eq. (7) converts the experiment-level logical failure probability over r=12 rounds into a per-round rate under a memoryless assumption, while the DEM contains faults spanning multiple rounds and hence correlated failures. Please clarify whether this conversion is used only for display and whether it materially affects the comparison, or present the experiment-level probabilities directly.","section":"Section V and Section IV.C; Eq. (7)"},{"comment":"ELMS is a new decoder introduced in this work and is one of the three anchors of the empirical claims, but its specification is incomplete. The layer partition of the check nodes is not defined, the 'minimum-weight estimate among converged' selection rule is ambiguous in the presence of degenerate syndrome-valid estimates, and the choices of 20 constituent decoders, 100 iterations, and damping alpha=0.95 are given without motivation or a sensitivity check. Please provide a complete pseudocode, or a reference to a public implementation, so that the ELMS results are fully reproducible.","section":"Section II.E"}],"minor_comments":[{"comment":"There is an internal inconsistency in the failure-weight statements: Section IV says RelayBP and ImpulseBP fail at weight four and ELMS also at weight three, while the Conclusions say 'all the analyzed decoders fail on some weight-three and weight-four fault patterns'. Please reword the conclusion to match the reported weight-specific observations.","section":"Section IV and Conclusions"},{"comment":"The decoder priors are fixed at p0=10^-3 for every physical error rate p. Since the LETS estimate and the Monte Carlo simulation both use the same fixed priors, the comparison is internally consistent, but the choice of p0 is arbitrary. A sentence reporting the sensitivity of the estimate (and of the MC comparison) to p0 would strengthen confidence in the results.","section":"Section IV.A and Figures 1-3"},{"comment":"Algorithm numbering is confusing: the qualitative sketch in Section III.C is named Algorithm 1, the full dpl-search in Appendix B is Algorithm 2, and the expansion-table generation in Ref. [13] is also Algorithm 1. Use distinct labels, for example 'Algorithm 1 (sketch)' and 'Algorithm 2 (full search)'.","section":"Appendix B"},{"comment":"The statement that the same three topologies form the top-three set for RelayBP and ImpulseBP is based on unweighted instance counts N_fail(tau). The text already notes this, but the caption of Table V should state explicitly that these are instance-level structural statistics and are not weighted by the physical probabilities of the injected fault patterns.","section":"Table V and Table VII"},{"comment":"The paper would benefit from a short note connecting the failure-spectrum approach of Ref. [28] with the LETS lower bound, for example whether the weight-four failures identified here are also found by the 'fail fast' search of Ref. [28]. This would help the reader place the two methods relative to each other.","section":"Section III.C"},{"comment":"In Eq. (1), the product over j not in S includes the probability that all other DEM fault variables are absent; for a DEM with many fault variables this factor is close to unity in the low-error regime, but for very large p it can be non-negligible. A brief remark that the estimator is dominated by the first product in the error-floor regime would help.","section":"Equation (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper's technical core—the lower-bound estimator, the exhaustive enumeration pipeline, and the topology-level analysis—appears sound and reproducible. The main risk is that the abstract and introduction make a predictive claim that the body of the paper only partially supports: the only clean validation of the completeness assumption is the RelayBP match on a single code, and the paper itself shows that the assumption is violated for ImpulseBP and ELMS. If the authors reframe their claims as lower-bound and partial structural characterization, or add evidence about the missing support mass, the paper would be a solid methods contribution. The statistical presentation (error bars, MC ranges, estimator variance) also needs attention before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, it is the first time anyone has run an exhaustive trapping-set search directly on a circuit-level detector error model, which is a real step past the code-capacity and phenomenological studies that dominate the QLDPC trapping-set literature. Second, the headline result is narrower than the abstract suggests: the LETS-based estimate reproduces the Monte Carlo error floor almost perfectly for RelayBP on the [[144,12,12]] BB code, but for ImpulseBP and ELMS the estimate sits below the simulated curves, and the paper itself attributes that gap to structures outside the enumerated universe. That is not a hidden flaw; it is stated clearly, but it means the \"substantial part\" claim is directly established for one decoder and one code.\n\nWhat the paper does well: the adaptation of dpl-search to the irregular DEM Tanner graph is nontrivial and the appendix gives enough detail to be checked. The lower-bound argument in Eq. (3) is correct, and the estimator for randomized decoders is unbiased, which is worth saying because a lot of error-floor work quietly fits constants instead of computing lower bounds. The public 92-million-instance catalog is a real asset. The topology-level analysis is also the most useful part for decoder design: the fact that RelayBP and ImpulseBP share the same top three harmful topologies, with ELMS not far behind, is a concrete, non-obvious finding.\n\nSoft spots, in proportion. The completeness assumption is load-bearing, and the paper knows it. The search is limited to LETS of size at most five, and the injection to fault patterns of weight at most four. Weight-five patterns are excluded on asymptotic grounds, which is fine for a lower bound, but the \"within one order of magnitude\" claim for the other two decoders is tied to weight-three and weight-four only. The Monte Carlo plots have no error bars, even though the text reports at least 100 failures per point; that is a minor fix but should be made. Enumeration completeness is inherited from Hashemi and Banihashemi without a re-proof for this graph family, though the explicit structural checks in the appendix partially mitigate that. And there is only one code. The authors note the [[288,12,18]] BB code as an open question, but that leaves the framework's generality untested.\n\nThe math and the data look solid to me; nothing here looks fitted or circular. This is a serious paper for anyone working on QLDPC iterative decoding or error-floor prediction. It deserves peer review, not desk rejection. I would ask the authors for error bars, a softening of the abstract's \"establish\" language to something like \"demonstrate for the studied code,\" and, if feasible, one more code point to show the framework transfers. I would also encourage them to keep the honest framing of the lower-bound gap; that honesty is a strength, not a liability.","headline":"First exhaustive LETS enumeration on a circuit-level DEM, with a genuine RelayBP error-floor match; the general claim rests on one code and one clean match, but the paper is honest about that.","tokens_in":28061,"tokens_out":2219,"would_cite":true,"duration_ms":28006,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P70","94B35"],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"A finite search over small leafless trapping sets in a circuit-level detector model predicts the error floor of iterative quantum decoders.","keywords":["trapping sets","detector error model","error floor","quantum LDPC codes","iterative decoding","circuit-level noise","bivariate bicycle code","message-passing decoders"],"falsifier":"Run the same three decoders at a physical error rate near p = $10^{-4}$, collect every Monte Carlo failure, and check whether each failing fault pattern is supported on an enumerated LETS with a_max=5; discovering any failing weight-four pattern whose support lies only in a non-elementary, leaf-containing, or size-six-or-larger structure would falsify the completeness claim for that decoder.","tokens_in":27024,"feed_emoji":"⚛️","tokens_out":9628,"duration_ms":102241,"temperature":0.7,"pith_summary":"This paper tries to establish that the rare failures of iterative quantum decoders under circuit-level noise can be predicted without Monte Carlo simulation, by enumerating the small graph structures that support them. The authors apply an exhaustive leafless elementary trapping-set search directly to the detector error model of a [[144,12,12]] bivariate bicycle code, then exhaustively inject every weight-three and weight-four fault pattern supported on those structures and sum the probabilities of the patterns that defeat each decoder. The resulting lower-bound estimate reproduces the simulated error floor for RelayBP and stays within one order of magnitude for ImpulseBP and for the paper's own ELMS decoder. If correct, this converts extremely rare logical failure analysis into a finite, reusable structural catalog and quantifies a gap between circuit-level distance and practical decoding.","feed_headline":"Trapping-set search predicts quantum decoder error floors","feed_subtitle":"For three circuit-level decoders, the structural estimate matches Monte Carlo or stays within an order of magnitude.","key_machinery":"The central object is the leafless elementary trapping set (LETS): a set of fault-variable nodes in the detector error model whose induced subgraph has every check of degree one or two, and whose normal-graph representation has no leaf. The carrying mechanism is the dpl-search algorithm, which grows LETS instances from short cycles using dot, path, and lollipop expansions controlled by a completeness table, enumerating all instances up to a_max=5 and b_max=5. The second half of the machinery is exhaustive injection: every weight-one through weight-four fault pattern supported on the union of LETS supports is decoded once, and the per-pattern probabilities of the failures are summed through Eq. (1), giving the lower-bound estimate of the logical error rate per round.","core_discovery":"The central discovery is that the low-weight error floor of circuit-level iterative decoders is governed by a finite set of leafless elementary trapping sets in the detector error model, and that exhaustive enumeration followed by exhaustive fault-pattern injection turns the floor into a weighted sum. For the [[144,12,12]] code the search found 92,088,583 LETS instances through (5,5); from their supports the authors built all distinct weight-one through weight-four fault supports and decoded each once. No decoder failed on weight-one or weight-two patterns, RelayBP and ImpulseBP failed on weight-four patterns, and ELMS also failed on weight-three patterns. The summed probabilities of the failing supports give a lower bound on the round logical error rate; for RelayBP it tracks the Monte Carlo curve, and for ImpulseBP and ELMS it lies within one order of magnitude. The paper also shows harmfulness is highly concentrated in a few topology classes, with the same three topologies most harmful for RelayBP and ImpulseBP and still highly ranked for ELMS.","pith_inferences":["If the completeness of LETS enumeration carries over to other QLDPC code families, the same pipeline could rank decoders by their structural weaknesses before expensive low-rate Monte Carlo runs; the growing number of LETS instances would likely force pruning by topology class rather than full enumeration.","The priors in the estimate are fixed at the p0 = 10^-3 operating point, so at substantially different physical error rates the relative importance of the identified supports could change and the lower bound would need rechecking.","One testable extension is to modify a decoder to penalize or escape the identified weight-three and weight-four failure supports; if the error floor drops accordingly, the causal role of the LETSs is confirmed.","The cross-decoder overlap of harmful topologies hints at intrinsic DEM subgraphs that are hard for message passing regardless of heuristic; characterizing those subgraphs algebraically could lead to code-design rules that avoid creating them."],"forward_implications":["The same LETS catalog can be reused to estimate the floor of a decoder at any physical error rate by reweighting fault-support probabilities, so extremely low-rate regimes become accessible without direct simulation.","All three decoders fail on weight-three or weight-four patterns even though the circuit-level distance is 12, so the gap between code distance and iterative-decoding performance is not an artifact of one heuristic.","Harmful failures concentrate in a small set of LETS topology classes; the top three are identical for RelayBP and ImpulseBP and remain highly ranked for ELMS, giving concrete targets for decoder redesign.","The estimate is a lower bound, so the shortfall seen for ImpulseBP and ELMS points to failures supported on structures outside the enumerated LETS universe, such as larger or non-elementary trapping sets."],"supporting_citations":[{"why":"Supplies the dpl-search algorithm and its expansion-table completeness argument, which the paper adapts to detector error models.","marker":"[13]"},{"why":"Supplies the LETS characterization and exhaustive-search framework that is the enumeration backbone.","marker":"[14]"},{"why":"Defines the bivariate bicycle code and syndrome-extraction circuit whose detector error model is the object of study.","marker":"[29]"},{"why":"Defines the RelayBP decoder whose Monte Carlo curve the LETS estimate is designed to match.","marker":"[4]"},{"why":"Defines the ImpulseBP decoder used as a second, deterministic ensemble-based test case.","marker":"[8]"},{"why":"Introduces quantum trapping sets, the conceptual basis for why zero-syndrome structures degrade message-passing decoders.","marker":"[9]"},{"why":"Provides the alternative failure-spectrum rare-event technique that the paper positions its trapping-set lower bound against.","marker":"[28]"},{"why":"Provides the implementation and the enumerated LETS dataset used in the paper's experiments.","marker":"[30]"}],"fun_headline_variants":["Trapping sets pinpoint quantum decoder error floors","Predicting error floors without Monte Carlo simulation","Finite structural search exposes decoder weaknesses","Low-weight traps forecast iterative decoder failures","Trapping-set analysis quantifies quantum error floors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every dominant low-weight failure in the low-error regime is supported on one of the enumerated leafless elementary trapping sets of size at most five; if an important fault pattern lives on a larger, leaf-containing, or non-elementary configuration, the lower-bound estimate misses it and the predicted floor is too low.","fun_headline_variants_meta":{"raw":{"variants":["Trapping sets pinpoint quantum decoder error floors","Predicting error floors without Monte Carlo simulation","Finite structural search exposes decoder weaknesses","Low-weight traps forecast iterative decoder failures","Trapping-set analysis quantifies quantum error floors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000585,"raw_usage":{"total_tokens":2773,"prompt_tokens":991,"completion_tokens":1782,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1717}},"tokens_in":607,"tokens_out":1782,"duration_ms":14561,"temperature":1.0,"reasoning_tokens":1717,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:36:58.446534+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three decoders at a physical error rate near p = $10^{-4}$, collect every Monte Carlo failure, and check whether each failing fault pattern is supported on an enumerated LETS with a_max=5; discovering any failing weight-four pattern whose support lies only in a non-elementary, leaf-containing, or size-six-or-larger structure would falsify the completeness claim for that decoder.","supporting_citations":[{"cited_title":"Hashemi and A","cited_arxiv_id":null,"evidence_quote":"Supplies the dpl-search algorithm and its expansion-table completeness argument, which the paper adapts to detector error models."},{"cited_title":"Hashemi and A","cited_arxiv_id":null,"evidence_quote":"Supplies the LETS characterization and exhaustive-search framework that is the enumeration backbone."},{"cited_title":"Bravyi, A","cited_arxiv_id":null,"evidence_quote":"Defines the bivariate bicycle code and syndrome-extraction circuit whose detector error model is the object of study."},{"cited_title":"Raveendran and B","cited_arxiv_id":null,"evidence_quote":"Introduces quantum trapping sets, the conceptual basis for why zero-syndrome structures degrade message-passing decoders."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the implementation and the enumerated LETS dataset used in the paper's experiments."}],"review_version":1}