{"id":"69c7a9a0-b01d-4ccd-8bed-0af2fc1d06dd","arxiv_id":"2504.20212","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Syndrome statistics alone can reconstruct the detector error models of repetition, surface, and color code memories, including hyperedge probabilities, assuming independent Pauli noise with known error-event structure.","lead":"This paper shows how to estimate the error rates that feed quantum error correction decoders directly from the syndrome measurements of a QEC experiment, including for codes that require hypergraph error models. A decoder calibrated this way lowers the logical error rate by up to about 18 percent in simulations when actual qubit error rates vary across the device.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hypergraph support is taken from Stim's exact DEM, so the color-code method estimates rates on a known hypergraph rather than reconstructing the DEM from syndrome statistics alone (Sec. III D).","rationale":"After checking the algebra, the graph-case estimator (Eqs. 3-4) is exact for any simple decoding graph under independent edge noise: with a_i=1-2<v_i> and a_ij=1-2<v_i>-2<v_j>+4<v_i v_j>, one obtains (a_i a_j)/a_ij=(1-2p_ij)^2, so p_ij is recovered independently of other incident edges. Thus the graph portion and the shot-count scaling are credible, and the Stim comparisons are a fair rate-calibration test. The weakness is concentrated in the hypergraph extension, the paper's main claim. There the estimator receives the hyperedge support from Stim's exact DEM (Sec. III D) and only fits rates, which contradicts the 'syndrome measurements alone' framing and leaves open whether structure can be identified from data. This is not fatal: for a known circuit, support can in principle be derived without knowing rates, and the numerical benchmarks are careful. But the support dependence should be stated as a condition, and a support-identifiability or recovery test is needed. This matches the Reader's secondary concern more than the independent-Pauli-noise limitation, which is explicitly scoped. I therefore keep the CONDITIONAL verdict.","tokens_in":21042,"tokens_out":13477,"duration_ms":136645,"concrete_test":"Run the color-code estimation at the Sec. III D operating point (d=5, r=2, p=2e-3, N=5e7) twice: once with the true Stim hyperedge support, and once with a candidate support that adds spurious detector triples and omits one true hyperedge while using the same syndrome statistics. Then compare (i) the least-squares rates assigned to spurious triples against a zero baseline under the observed shot noise, and (ii) the logical error rate of a decoder using the resulting weights against the true-DEM decoder. If spurious triples receive nonzero rates or the logical error rate degrades materially, syndrome statistics alone do not identify hyperedge support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The novel hypergraph results are presented as reconstructing the detector error model 'from the syndrome measurements alone without relying on prior information' (Sec. II). However, Sec. III D states that the estimator solves for 'any possible triplet that exists in the actual Z- or X-DEMs obtained by Stim,' and the d=3 color-code procedure is described as collecting 'all possible error mechanisms that exist in the DEM.' The support of the hypergraph, i.e., which detector triples are genuine error events, is therefore supplied by the exact DEM rather than learned from syndrome statistics; the statistics are used only to fit rates on that oracle-supplied support. This matters most for the hypergraph case, the main novelty. For graph-like DEMs the support is fixed by the known syndrome-extraction circuit, so Eqs. (3)-(4) are a legitimate rate-calibration procedure. But for hypergraphs the paper never shows that the estimator can distinguish a true hyperedge from coincidental combinations of lower-order errors. If the true DEM contained an unsuspected hyperedge, e.g., from crosstalk or a misidentified hook error, the fixed-support pipeline would silently miss it and the decoder would be miscalibrated. The claim to 'learn exactly all the information' in the color-code DEM is thus only conditional on knowing the support, which is not acknowledged in the abstract or Sec. II.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a method for reconstructing the detector error model (DEM) of memory quantum error correction experiments from syndrome statistics, avoiding full state or process tomography. For graph-like DEMs, two-point detector correlations are used with closed-form expressions from Ref. [31] to estimate edge probabilities. The method is extended to hypergraph DEMs by solving systems of configuration-probability equations (up to O(p^7)) via least squares, applied to the color code with bare-ancilla extraction and to the repetition code with Steane-style extraction. The reconstructed DEMs are decoded with MWPM and compared with Stim's exact DEMs for distances d=3..9 and physical error rates spanning below and above threshold; the logical error rates agree well. The paper also demonstrates that calibrating the decoder to spatially fluctuating error rates can improve the logical error rate relative to a decoder using the mean rates. Simulation code is publicly available.","tokens_in":21288,"tokens_out":20653,"duration_ms":189214,"significance":"The graph-like reconstruction is well validated and essentially reproduces and extends Ref. [31] to surface codes with circuit-level noise; this part is a solid contribution. The hypergraph extension is the main novelty, and the reported relative errors (below 1% for hyperedges, below 4% for edges) and logical-rate matches are encouraging. The release of the code and the comparison against Stim's independent DEM are strengths. However, the hypergraph extension as presented calibrates rates on a support that is supplied by the exact circuit-level DEM, so the advertised claim of reconstructing the DEM from syndrome statistics alone is not fully achieved for the hypergraph case. This caveat is load-bearing for the paper's central novelty and should be addressed before the manuscript can be accepted.","major_comments":[{"comment":"The hypergraph results do not support the claim that the decoder graph is generated 'from the syndrome measurements alone without relying on prior information' (Sec. II). In Sec. III D, the estimator solves for 'any possible triplet that exists in the actual Z- or X-DEMs obtained by Stim,' and the d=3 color-code procedure collects 'all possible error mechanisms that exist in the DEM.' The support of the hypergraph—the set of detector subsets that can be flipped together—is therefore supplied by the exact circuit-level DEM, and the syndrome statistics are used only to fit rates on that known support. The manuscript never demonstrates that a hyperedge absent from the assumed support (for example, arising from crosstalk or an unknown hook error) would be detected, nor that a genuine hyperedge is distinguished from coincidental combinations of lower-order events. Thus the claim to 'learn exactly all the information contained in' the color-code DEM is conditional on oracle knowledge of the support. The abstract and Sec. II should be amended to state this condition, and the authors should either add a misspecification test or restrict the claim to rate calibration on a known hypergraph support.","section":"Sec. III D; Sec. II"},{"comment":"In the triplet equations for the color code, the paper lists the seven possible error events p0, p1, p2, p01, p02, p12, p012 and states that keeping terms to O(p^7) 'makes the equations exact since there are 7 unknown quantities in total.' This presumes that no error event involving one of the three detectors and a detector outside the triplet (e.g., a two-point edge between D0 and an external detector, or an additional hyperedge sharing D0) can affect the marginal configuration probabilities P(x0, x1, x2). In the color code, each detector participates in multiple hyperedges and neighboring two-point edges, so the closure of the triplet region is not obvious. The paper should justify this assumption or show that outside contributions are negligible for the simulated error rates.","section":"Sec. III D"}],"minor_comments":[{"comment":"The text refers to 'the denominator of Eq. (4)' but the equation contains a product, not a denominator; the notation should be clarified, for example by making the product symbol explicit in the typeset formula.","section":"Sec. II A, Eq. (4)"},{"comment":"The text says the estimation uses N = 8 x 10^5 shots, while the Fig. 1 caption reports N = 5 x 10^6 shots; these numbers should be reconciled.","section":"Sec. II A and Fig. 1"},{"comment":"The text states that Fig. 9(d) uses p = 0.001 and then says Fig. 9(e) uses 'the same physical error rate of p = 0.01'; either the rates are different and should be stated separately, or the second value is a typo.","section":"Sec. III D and Fig. 9"},{"comment":"Refs. [57] and [58] are the same paper; the duplicate citation should be removed or replaced with a distinct resource.","section":"References"},{"comment":"The caption contains a typo: 'hyperdeges' should be 'hyperedges.'","section":"Fig. 10 caption"},{"comment":"The phrase 'given that the DEM contains a single event that flips at most m detectors simultaneously' is confusing because the method is intended for multiple hyperedges; rephrase to indicate the maximum order of any hyperedge in the region.","section":"Sec. II B"}],"recommendation":"major_revision","confidential_remarks":"The graph-like portion is a solid contribution and the hypergraph rate-calibration results are useful, but the advertised novelty is currently overstated because the hypergraph support is supplied by the exact circuit-level DEM rather than learned from syndrome statistics. If the authors can either demonstrate support learning from data or clearly scope the paper as calibrating rates on a known DEM support, the manuscript could be suitable for publication. The duplicate reference [57]/[58] should also be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: the paper is a useful and mostly sound contribution to the QEC noise-estimation toolkit. It takes the known two-point syndrome-statistics method (Spitz et al.) and extends it to hypergraph DEMs, with solid numerical evidence that reconstructed logical error rates match noise-aware decoding across repetition, surface, and color codes. The improvement numbers under fluctuating error rates are believable: 4-18% typically, up to 44% in the d=5 surface case.\n\nWhat's new: the multi-point configuration-probability approach for hypergraphs, applied to color-code DEMs and Steane-style repetition codes, plus the calibration-to-fluctuations results. The numerical validation is genuinely good: relative errors under a few percent, logical error rates reproduced over a wide range including above threshold, and the code is public. I also appreciate that they checked the concurrent Remm et al. closed-form expressions and found matching.\n\nThe main caveat is one the paper partly glosses over. In the hypergraph sections, the support of the hypergraph — which triplets are actual error events — is taken from Stim's exact DEM, not learned from syndrome statistics. The abstract and Sec II say 'using only syndrome statistics' and 'without relying on prior information.' That's true for the rates but not for the structure. If the true DEM has an unsuspected hyperedge (say, from crosstalk or a misidentified hook error), this pipeline would silently miss it. This is a real, but addressable, gap: in practice the circuit and its likely error mechanisms are known, and for graph-like DEMs the support is fixed by the circuit anyway. The authors should either walk back the 'no prior information' claim or show a procedure to infer or verify hyperedge support from data.\n\nAlso worth noting: the hypergraph extension is numerically validated but lacks a formal identifiability proof. That's not fatal given the scope, but it is a limitation. The independent-Bernoulli-noise assumption is stated up front, and the paper defers correlated/coherent noise to future work; fine as a first step.\n\nWho this is for: experimentalists and decoder people who want to calibrate a decoder using syndrome data they already have. It deserves a serious referee. I'd recommend accepting for peer review, with the support issue flagged for revision.","headline":"Solid, practical extension of syndrome-based DEM calibration to hypergraphs; main caveat is that hyperedge support is oracle-supplied from Stim, so 'no prior information' overstates the case.","tokens_in":21812,"tokens_out":3450,"would_cite":true,"duration_ms":33156,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"A memory QEC experiment's detector error model can be reconstructed from syndrome click statistics alone: two-point correlators for graph-like codes, multi-point correlators for hypergraphs, with the shot count independent of code size.","keywords":["detector error model","quantum error correction","noise estimation","decoding graph","decoding hypergraph","syndrome statistics","color code","surface code"],"falsifier":"Simulate a distance-3 surface code memory whose only noise is a coherent Z rotation through a known angle on each data qubit each round; estimate edge probabilities from two-point detector correlators using the paper's closed forms and compare with the exact Pauli-twirled DEM of the same circuit. A systematic, angle-dependent mismatch between reconstructed and twirled edge weights would pinpoint where the independent-Bernoulli assumption fails.","tokens_in":20804,"feed_emoji":"⚛️","tokens_out":9126,"duration_ms":86793,"temperature":0.7,"pith_summary":"The paper claims that the detector error model (the weighted decoding graph or hypergraph a decoder uses) can be learned directly from syndrome statistics, without any prior noise model from tomography or calibration. For graph-like codes, two-point detector correlations are enough; for hypergraph codes, multi-point correlations solve a small system of equations per local region. The number of experimental shots needed does not grow with code size. If true, this makes noise-aware decoding a routine by-product of the memory experiment itself, and it is demonstrated to match exact-model decoding across below- and above-threshold error rates.","feed_headline":"Syndrome clicks alone rebuild the decoder's error model","feed_subtitle":"Two-point clicks recover graph-like noise; multi-point clicks recover hyperedges, matching full decoding.","key_machinery":"The central object is the detector error model (DEM): a graph or hypergraph whose nodes are detectors (parities of ancilla outcomes that flag errors) and whose weighted edges or hyperedges are independent error mechanisms, with edge weight $w=-\\ln(p/(1-p))$. The estimation engine is a hierarchical correlator expansion: two-point closed forms for graph-like DEMs; and for hypergraph DEMs, a system of $2^m-1$ configuration-probability equations $P(x_0,x_1,\\ldots)$ per local region, truncated in the error rate and solved by least squares. The load-bearing bookkeeping step is subtraction: the recovered high-order event probabilities are subtracted out of lower-order edge probabilities so that each final weight corresponds to a single independent error event.","core_discovery":"The central discovery is that the error probabilities of a detector error model are identifiable from detector firing statistics alone, with a closed form for graph edges and a least-squares system for hyperedges. Bulk edge probabilities come from two-point coincidences via $p_{ij} = \\frac{1}{2} - \\sqrt{\\frac{1}{4} - \\frac{\\langle v_iv_j\\rangle - \\langle v_i\\rangle\\langle v_j\\rangle}{1-2(\\langle v_i\\rangle+\\langle v_j\\rangle)+4\\langle v_iv_j\\rangle}}$, and boundary edges are then fixed by subtracting all incident bulk contributions. When a single error can flip three or four detectors, the paper writes the probability of each detector-outcome configuration as a polynomial in the unknown DEM rates, solves by least squares, and then renormalizes lower-order edges with $p_{\\mathrm{new}}=(p_{\\mathrm{old}}-p_{ijkl})/(1-2p_{ijkl})$ to remove the high-order contribution. On repetition, surface, and color-code memories this recovers the DEM closely enough that decoding with the reconstructed graph reproduces the logical error rate of decoding with the exact model, including a color-code case where the hyperedges are recovered with relative error below about one percent.","pith_inferences":["The same local least-squares and subtraction recipe should transfer to other hypergraph codes, including small hypergraph low-density parity-check memories, wherever the maximum correlation order is a bounded constant; the paper's own partitioning argument suggests cost grows with the number of local regions, not with region size.","Because the method uses raw click statistics and no twirling, a natural on-line extension is drift tracking: re-estimating edge weights periodically could follow slowly varying stochastic noise and keep the decoder matched to current conditions.","A direct experimental test would compare edge weights reconstructed this way with weights derived from Pauli-twirled gate estimates on the same device; systematic disagreements would localize crosstalk, leakage, or coherent effects to specific regions of the decoding graph.","The d=3 color-code degeneracy between a detector error and an error-plus-logical-flip suggests a general recipe: any DEM region with such a degeneracy can be resolved by adding configuration equations that count logical-observable flips, not just detector states."],"forward_implications":["A memory experiment can supply its own decoder weights from accumulated click statistics, removing the need for a separate noise-characterization calibration step.","The number of shots for a fixed estimation accuracy stays bounded as code distance grows, as long as the maximum number of detectors flipped by one error remains a small constant.","Decoding with the reconstructed DEM reproduces the logical error rate of exact-model decoding both below and above threshold for repetition, surface, and color-code memories.","When physical error rates vary across qubits and gates, a decoder calibrated to the reconstructed fluctuating rates outperforms a decoder built from fixed mean rates, with up to about 44% better logical error suppression in the surface-code example."],"supporting_citations":[{"why":"Supplies the two-point correlator closed-form estimation method on which the graph-like part of this work is built.","marker":"[31]"},{"why":"Raises the identifiability limitation for noise estimation on the color code that the paper addresses by restricting attention to the DEM.","marker":"[32]"},{"why":"Provides the stabilizer-circuit simulator and exact detector error models used as ground truth and syndrome-data source for all demonstrations.","marker":"[35]"},{"why":"Provides the sparse minimum-weight perfect matching decoder used to compare logical error rates of the reconstructed and exact DEMs.","marker":"[36]"},{"why":"Provides the minimum-weight perfect matching implementation used for decoding the repetition and surface code reconstructed graphs.","marker":"[39]"},{"why":"Supplies the color-code decoder whose separation into color-restricted and color-only lattices requires the multi-point hypergraph estimation developed here.","marker":"[57]"},{"why":"Provides closed-form expressions for higher-order syndrome correlations that the paper verifies against its numerical hypergraph solutions.","marker":"[15]"},{"why":"Describes a scalable noise-characterization alternative whose cost and eigenvalue-learning constraints motivate learning the DEM directly from syndrome statistics.","marker":"[26]"}],"fun_headline_variants":["Detector clicks reveal decoder error models directly","Decode error models from syndrome click statistics alone","Recover QEC decoder graphs from detector firing rates","Click statistics alone identify decoder error probabilities","From syndrome clicks to decoder noise: direct estimation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every error event is an independent, fixed-probability Bernoulli draw, so correlated, coherent, crosstalk, leakage, or time-varying noise is outside what the reconstruction can faithfully represent.","fun_headline_variants_meta":{"raw":{"variants":["Detector clicks reveal decoder error models directly","Decode error models from syndrome click statistics alone","Recover QEC decoder graphs from detector firing rates","Click statistics alone identify decoder error probabilities","From syndrome clicks to decoder noise: direct estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2787,"prompt_tokens":995,"completion_tokens":1792,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":1723}},"tokens_in":611,"tokens_out":1792,"duration_ms":12320,"temperature":1.0,"reasoning_tokens":1723,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:35:09.046719+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a distance-3 surface code memory whose only noise is a coherent Z rotation through a known angle on each data qubit each round; estimate edge probabilities from two-point detector correlators using the paper's closed forms and compare with the exact Pauli-twirled DEM of the same circuit. A systematic, angle-dependent mismatch between reconstructed and twirled edge weights would pinpoint where the independent-Bernoulli assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the two-point correlator closed-form estimation method on which the graph-like part of this work is built."},{"cited_title":"Wagner, H","cited_arxiv_id":null,"evidence_quote":"Raises the identifiability limitation for noise estimation on the color code that the paper addresses by restricting attention to the DEM."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes a scalable noise-characterization alternative whose cost and eigenvalue-learning constraints motivate learning the DEM directly from syndrome statistics."}],"review_version":1}