{"id":"6036b9b7-9f06-4056-95e9-059199e82ff0","arxiv_id":"2608.08996","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A multi-agent search over balanced-product quantum LDPC codes finds new finite-length instances, including [[288,16,18]], [[288,18,18]], and [[234,28,18]], with leading rate-distance scores under fixed weight constraints.","lead":"This paper uses a multi-agent AI system to search for compact quantum error-correcting codes, reporting new instances with record or competitive parameters up to 400 physical qubits. It also finds structurally distinct codes from non-normal group constructions and shows several decode well under a standard error model.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 'best-known' claims rest on an unreleased 1,209-instance comparison benchmark; without it the central comparative results cannot be audited from the preprint.","rationale":"The reader's weakest assumption is the completeness of the curated comparison set, and my stress-test agrees. The paper is methodologically careful in distinguishing exact distances from upper bounds, and the SI gives explicit group/protograph data for each code, so the codes are in principle constructible. The residual risk is concentrated in the comparative and certification claims, both of which depend on artifacts not included in the preprint. The MILP certification procedure is described plausibly (a symplectic basis with 2k subproblems covering all logical classes), and the overall-weight recomputation is stated, so there is no internal inconsistency I can identify. Therefore the appropriate verdict remains CONDITIONAL: accept if the released benchmark confirms the best-known statements and the certificates replay. I do not see grounds to move the verdict.","tokens_in":20531,"tokens_out":7979,"duration_ms":74914,"concrete_test":"Release the full comparison set with per-instance recomputed overall weight w (Eq. 1) and distance-evidence labels; independently recompute Q=k d^2/n for every instance in each weight class and verify no omitted or misclassified instance increases the maximum Q in w=6..10. Separately, replay the MILP optimality certificates for the three flagship codes (e.g., by independent binary linear algebra or an exact distance solver such as GAP/QDistRnd) to confirm d=18,18,18.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Methods V.G describes a curated comparison set of 1,209 instances from 91 sources, but the Data availability statement says all comparison data 'will be made publicly available upon publication.' The central claims that [[288,16,18]], [[288,18,18]], and [[234,28,18]] are the strongest known exact-distance codes in weight classes w=7,9,10 are comparative statements. If the benchmark omits a stronger published code, or if any included code's overall weight (Eq. 1) was miscomputed from partial construction data, the 'best-known' assertions fail even though the codes themselves remain valid CSS codes with the stated certified parameters. Additionally, exact distance claims rely on Gurobi MILP certificates (Methods V.F) that are not included in the manuscript; the SI provides only construction data. Thus the two load-bearing supports for 'leading performance' — the benchmark's completeness and the distance certifications — are not replayable from the submitted materials. This is a verification gap rather than an identified error; the internal construction logic appears coherent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a multi-agent AI search framework for discovering finite-length quantum LDPC codes. The framework combines researcher and curator agents that propose hypotheses and maintain persistent lessons, worker agents that evolve executable code-family generators, and a fixed deterministic evaluator that constructs coset-orbit balanced-product CSS codes, enforces CSS orthogonality, block length n<=400 and overall weight w<=10 (Eq. 1), canonically labels Tanner graphs to remove isomorphic copies, and scores candidates by Qproxy = k d~^2/n using a capped QDistEvol distance upper bound (Eq. 2). The paper reports 20 highlighted codes in Table I and 7 structurally distinct codes in Table II, including claimed exact-distance leaders [[288,16,18]] at w=7, [[288,18,18]] at w=9, and [[234,28,18]] at w=10, as well as genuine balanced-product constructions with non-normal subgroup actions. The paper also reports BP-OSD code-capacity simulations showing low per-logical error rates and pseudo-thresholds around 5--9%.","tokens_in":20709,"tokens_out":7145,"duration_ms":67761,"significance":"If the claimed parameters and benchmark comparisons hold, the paper contributes practically relevant finite-length qLDPC codes under hardware-motivated weight constraints, and the multi-agent methodology with persistent memory and deterministic evaluation is a useful template for structured search in quantum error correction. Strengths of the manuscript include the explicit separation of exact distances (MILP) from upper bounds (QDistEvol), the deterministic evaluator with independent F2 arithmetic, the careful marking of upper-bound entries in Tables I--III, and the inclusion of full construction data for the highlighted codes in the Supplementary Information. The main weakness is not internal correctness but audibility: the comparative 'best-known' claims and the exact-distance certificates are not replayable from the submitted materials. These are verification gaps rather than identified errors.","major_comments":[{"comment":"The central comparative claims in Section III.A, for example that [[288,16,18]], [[288,18,18]], and [[234,28,18]] are the strongest known codes in their respective weight classes, rest entirely on the 1,209-instance literature benchmark described in Methods V.G, but that benchmark is not included in the preprint and is promised only upon publication. Since the overall weight under Eq. (1) is recomputed from defining matrices or construction data, the reader cannot check that every compared instance satisfies the intended weight class or that no stronger published code is omitted. Please provide the complete comparison set, including instances, sources, matrices or construction data, distance-evidence labels, and recomputed w and Q values, as a supplementary table or public repository at submission, or clearly restrict the claims to 'best among the compared set'.","section":"Methods V.G; Data availability"},{"comment":"The exact distance claims in Table I, such as [[288,16,18]], [[288,18,18]], and [[234,28,18]], depend on Gurobi 12.0.3 MILP certificates and explicit witnesses described in Methods V.F, but the evidence bundles containing matrix hashes, solver statuses, and witnesses are not included in the manuscript or SI, and the code is not available until publication. The SI provides only the balanced-product construction data, which is insufficient to replay the distance certification. Please include or release the MILP certificates, witnesses, and an independent replay script, or at minimum the final binary check matrices with verified parameters, so that the exact-distance assertions are checkable before acceptance.","section":"Methods V.F; Data availability"},{"comment":"Several headline 'leading' entries are distance upper bounds rather than exact distances, for example [[390,32,<=32]] and [[384,18,<=28]], and the abstract's phrase 'leading or competitive rate--distance performance in every weight class' covers both cases. The text is generally careful to label these entries, but some sentences in Section III.A, such as the statement that an upper-bound endpoint 'exceeds' a previous upper-bound endpoint, could be read as ordering actual code performance. Please add an explicit sentence stating that comparisons involving upper-bound entries are comparisons of upper endpoints, not of certified code distances, and ensure the abstract reflects this distinction.","section":"Section III.A; Table I"}],"minor_comments":[{"comment":"The word 'recoreds' should be 'records'.","section":"Data availability"},{"comment":"The term 'non-BB codes' is used without definition; please define it as 'non-bivariate-bicycle codes' at first use.","section":"Methods V.H"},{"comment":"The description of confidence intervals for adaptively stopped simulation runs would benefit from a concrete statement of how the negative-binomial interval is computed and how the stopping rule affects the reported shot counts.","section":"Methods V.H"},{"comment":"The proxy score Qproxy is a heuristic because the cap min(d_ub,1.3*sqrt(n)) can either overestimate or underestimate the true Q; a sentence stating explicitly that Qproxy is neither an upper nor a lower bound on Q would prevent misinterpretation.","section":"Section II, Eq. (2)"},{"comment":"The phrase 'the pseudo-threshold decreases monotonically' is based on the ten selected codes in Table III; please state that this is an observation on the selected set rather than a proven general trend.","section":"Section III.C and Table III"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the internal construction logic appears coherent, but the central comparative and exact-distance claims cannot currently be evaluated from the submitted materials. I would make release of the comparison benchmark and the MILP certificates/witnesses conditions of acceptance. The relation to Ref. [32] (OmniQEC) is addressed only in a brief footnote; the editors may wish to verify that the two contributions are sufficiently distinct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious look. The paper reports concrete new qLDPC instances — including [[288,16,18]] at w=7, [[288,18,18]] at w=9, and [[234,28,18]] at w=10 — and a multi-agent search framework that extends earlier structured concept evolution with an explicit subgroup-action level. The specific codes and the genuine balanced-product constructions in Table II are new as far as I can tell. The authors are careful with labeling: they distinguish exact distances from upper bounds, describe the MILP certification methodology honestly, and don't overclaim the BP–OSD results. The deterministic evaluator and the explicit weight definition (Eq. 1) are sensible, and the framework description is detailed enough to reproduce once code is released.\n\nThe main soft spot is a verification gap rather than an identified error. The \"strongest known\" claims for the weight classes rest on a curated comparison set of 1,209 instances from 91 sources, but that set is not included and is promised only upon publication. Without it, the comparative statements can't be audited from the preprint. Similarly, the exact distances rely on Gurobi solver attestations; the evidence bundles aren't in the SI. These are load-bearing for the headline results. A referee should ask for the benchmark and the MILP witnesses before accepting the \"best-known\" language.\n\nOne minor concern: the overall-weight definition sums X and Z qubit degrees, which is more restrictive than some literature conventions. The paper acknowledges one case ([[300,60,14]]) but it remains a judgment call that affects the comparison. It's not fatal, but readers should keep it in mind. The conjecture that lifted products are intrinsically favored in this regime is reasonable but unproven; fine as a conjecture.\n\nFor someone working on finite-length qLDPC or AI-guided code discovery, this paper has value: the instances are concrete, the search design is thoughtful, and the honest labeling sets a good tone. It deserves a serious referee. My recommendation: send it to peer review, conditional on the authors releasing the comparison set and the distance-certification evidence before final acceptance.","headline":"Plausible new finite-length qLDPC instances with a careful search pipeline, but the 'best-known' claims hinge on an unreleased comparison benchmark and MILP certificates.","tokens_in":21204,"tokens_out":1615,"would_cite":true,"duration_ms":17947,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P70","94B05"],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"A multi-agent search framework discovers finite-length quantum LDPC codes with leading or competitive rate–distance performance in every weight class from 6 to 10, including certified record-setters [[288,16,18]], [[288,18,18]], and…","keywords":["quantum LDPC codes","balanced product codes","multi-agent search","error correction","finite-length codes","distance certification","BP-OSD decoding","AI for scientific discovery"],"falsifier":"One concrete check is to independently compile all published binary CSS codes with n≤400 and overall weight w≤10 under the paper's weight convention, compute Q = k $d^{2}$/n for every code with exact distance, and see whether any code in the w=7, w=9, or w=10 class exceeds Q=18.00, Q=20.25, or Q=38.77 respectively; finding one would refute the record claims. A second check is to re-certify the distance of [[288,16,18]] by an independent MILP or exhaustive logical-operator search: if the true distance is below 18, the headline w=7 record fails.","tokens_in":20314,"feed_emoji":"⚛️","tokens_out":10874,"duration_ms":89143,"temperature":0.7,"pith_summary":"This paper claims that a multi-agent search framework can discover practical quantum LDPC codes that are competitive with or better than previously known finite-length codes under realistic hardware constraints. The framework evolves executable programs that generate coset-orbit balanced-product code families, and it couples program evolution with a scientific-reasoning loop of researcher agents, a curator that maintains lessons, and deterministic evaluation. Under the constraints of block length n ≤ 400 and overall weight w ≤ 10, it reports codes with the largest parameter scores Q = k $d^{2}$/n in the w = 7, 9, and 10 weight classes among codes with rigorously certified distances: [[288,16,18]], [[288,18,18]], and [[234,28,18]]. These matter because finite-length qLDPC codes are prime candidates for low-overhead quantum error correction in near-term hardware, and certified distance records provide concrete targets for experiments. The paper also shows the search reaches structurally new constructions, including genuine balanced-product codes with non-normal subgroup actions, and that selected discoveries have low logical failure rates under a standard decoder.","feed_headline":"Multi-agent search finds best-known quantum LDPC codes","feed_subtitle":"Record rate-distance scores at weights 7, 9, and 10, and competitive codes elsewhere, all with certified distances.","key_machinery":"The central object is the coset-orbit balanced-product code construction, which builds a CSS code from a finite host group G, subgroup assignments K_i and K_j, and two protograph matrices A(t), B(t) whose entries are F2 sums of double-coset orbits K_i g K_j; this space includes lifted products when subgroups are trivial or normal and extends to genuinely new codes for non-normal actions. The search machinery is a four-level executable representation—local terms, protograph shape, subgroup action, and host-group family—that a worker agent mutates under a MAP-Elites archive. Candidates are scored by the proxy Q_proxy = k $d_ub^{2}$/n using a distance upper bound from an evolutionary routine (QDistEvol), capped at 1.3√n to avoid rewarding loose bounds, and the most promising candidates are escalated to more expensive verification tiers. After the search, exact distances are certified by mixed-integer linear programming, with each claim backed by an explicit logical-operator witness and independent linear-algebra verification.","core_discovery":"The central claim is that a closed-loop search coupling scientific reasoning with executable-program evolution finds finite-length binary CSS codes that set or approach the best-known rate–distance trade-off under practical sparsity constraints. The strongest reported results are certified exact-distance codes [[288,16,18]] at w=7 (Q=18.00), [[288,18,18]] at w=9 (Q=20.25), and [[234,28,18]] at w=10 (Q=38.77), each stated to exceed the best previously known exact code in its weight class; in addition, upper-bound findings such as [[390,32,≤32]] (Q≤84.02 at w=10) indicate further headroom. The search also produced structurally distinct constructions, including a [[336,12,≤24]] candidate over PSL(2,11) and an exact [[368,18,16]] code over A6×Z2, both realized as genuine balanced products with non-normal subgroup actions, demonstrating that the framework explores beyond the lifted-product family. Under code-capacity depolarizing noise with a common BP-OSD decoder, all ten parameter champions have pseudo-thresholds between 5.4% and 9.3% and per-logical error rates below 1.7×$10^{-5}$ at p=0.03.","pith_inferences":["The paper leaves open whether the scientific-reasoning loop (researcher council and curator) is essential; a natural ablation would compare this framework against a version with only program evolution and fitness feedback, to quantify the contribution of persistent lessons.","The record claims are contingent on the comparison dataset, which is stated to be released upon publication; until then, an independent re-computation of Q for all published codes in the same regime is the only way to verify the 'best-known' statements.","The paper's conjecture that genuine balanced products may win at larger n or higher w could be tested by extending the search window to n≈1000, where balanced-product orbit geometry might overcome lifted products.","The proxy score Q=k d_ub^2/n with a cap rewards high-distance, high-rate codes but ignores decoding performance; a direct extension would be multi-objective search that adds pseudo-threshold or logical error rate as an archive dimension."],"forward_implications":["The certified codes, particularly [[288,16,18]], [[288,18,18]], and [[234,28,18]], provide concrete finite-length qLDPC instances that experimental groups can target directly, with rigorous distance proofs already in hand.","If the literature comparison is complete, these results shift the known finite-length rate–distance frontier in weight classes 7, 9, and 10, giving new reference points for future code design.","The framework's architecture is construction-agnostic: any qLDPC family with a deterministic constructor and a compact program representation can be plugged into the same search loop, so the method should carry over to other construction families.","The systematic decay of pseudo-threshold with increasing weight suggests that high-weight codes need better decoders or decoder-aware search; the paper proposes incorporating decoder feedback directly into the loop as a next step.","The discovery of genuine balanced-product codes with non-normal subgroup actions shows the search reaches structurally new territory, but the paper conjectures that lifted products may still dominate at n≤400, w≤10."],"supporting_citations":[{"why":"Defines the coset-orbit balanced-product construction that forms the search space.","marker":"[14]"},{"why":"Provides the previous best exact-distance w=6 code [[340,16,18]] and the w=6 bivariate-bicycle upper-bound benchmark [[360,12,≤24]].","marker":"[16]"},{"why":"Supplies the previous w=7 upper-bound Q≤17.78 and the w=8 benchmark [[306,22,≤24]] used to establish improvement.","marker":"[18]"},{"why":"Provides the previous best exact-distance w=10 code that [[234,28,18]] surpasses.","marker":"[20]"},{"why":"Supplies the executable-program evolution approach and the proxy-score cap that Eq. (2) inherits.","marker":"[26]"},{"why":"Supplies the three-level structured representation that the four-level hierarchy extends.","marker":"[27]"},{"why":"Provides QDistEvol, the evolutionary routine that yields distance upper bounds used in the proxy score.","marker":"[33]"},{"why":"Supplies the graph-canonicalization method used to remove isomorphic codes before scoring.","marker":"[35]"},{"why":"Defines the MAP-Elites archive that structures the search niches.","marker":"[38]"},{"why":"Supplies the MILP solver used to certify exact distances post-search.","marker":"[41]"}],"fun_headline_variants":["Multi-agent search finds record quantum LDPC codes","Agent-driven discovery nets practical quantum LDPC records","Quantum LDPC codes set new rate-distance records","Multi-agent framework finds best practical quantum LDPC codes","Agentic search achieves record quantum LDPC performance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the discovered codes are the strongest known in their weight classes rests on the completeness of a curated comparison set of 1,209 published qLDPC instances from 91 sources; if a published code with a larger Q was missed, the 'best-known' statements for those weight classes would fail, although the discovered codes would remain valid with their certified parameters.","fun_headline_variants_meta":{"raw":{"variants":["Multi-agent search finds record quantum LDPC codes","Agent-driven discovery nets practical quantum LDPC records","Quantum LDPC codes set new rate-distance records","Multi-agent framework finds best practical quantum LDPC codes","Agentic search achieves record quantum LDPC performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000597,"raw_usage":{"total_tokens":2888,"prompt_tokens":1137,"completion_tokens":1751,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":753,"completion_tokens_details":{"reasoning_tokens":1674}},"tokens_in":753,"tokens_out":1751,"duration_ms":14670,"temperature":1.0,"reasoning_tokens":1674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:18:10.300649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete check is to independently compile all published binary CSS codes with n≤400 and overall weight w≤10 under the paper's weight convention, compute Q = k $d^{2}$/n for every code with exact distance, and see whether any code in the w=7, w=9, or w=10 class exceeds Q=18.00, Q=20.25, or Q=38.77 respectively; finding one would refute the record claims. A second check is to re-certify the distance of [[288,16,18]] by an independent MILP or exhaustive logical-operator search: if the true distance is below 18, the headline w=7 record fails.","supporting_citations":[{"cited_title":"Breuckmann and Jens N","cited_arxiv_id":null,"evidence_quote":"Defines the coset-orbit balanced-product construction that forms the search space."},{"cited_title":"Generalized toric codes on twisted tori for quantum error correction.PRX Quantum, 6:020357, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the previous best exact-distance w=6 code [[340,16,18]] and the w=6 bivariate-bicycle upper-bound benchmark [[360,12,≤24]]."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the previous w=7 upper-bound Q≤17.78 and the w=8 benchmark [[306,22,≤24]] used to establish improvement."},{"cited_title":"Bayesianoptimization for quantum error-correcting code discovery.arXiv preprint arXiv:2601.18562, 2026","cited_arxiv_id":null,"evidence_quote":"Supplies the executable-program evolution approach and the proxy-score cap that Eq. (2) inherits."},{"cited_title":"OmniQEC: discovering practical quantum error-correcting codes by an AI scientist","cited_arxiv_id":"2607.25865","evidence_quote":"Provides QDistEvol, the evolutionary routine that yields distance upper bounds used in the proxy score."},{"cited_title":"Kovalev and Leonid P","cited_arxiv_id":null,"evidence_quote":"Defines the MAP-Elites archive that structures the search niches."}],"review_version":1}