{"id":"a9254949-a4e0-49ee-beb7-00c7b5231b8b","arxiv_id":"2502.05992","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Qutrit and ququint versions of the smallest 5-qudit error-correcting code achieve circuit-level error thresholds comparable to the qubit version, around 10^-4, when a flag qudit suppresses hook errors.","lead":"This paper simulates tiny quantum error-correcting codes that use multi-level quantum systems, or qudits, instead of two-level qubits. It reports that with a helper 'flag' qudit, codes of dimension 2, 3 and 5 all reach error thresholds near 10^-4, suggesting qudits could be practical for fault-tolerant quantum computing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Threshold values are fixed points of a single-level power-law fit, not measured concatenated thresholds; no l=2 simulation anchors the recursion.","rationale":"I read the paper in good faith. The decoder adaptations, the flag-qudit construction, and the level-1 circuit-level simulations are concrete and described in enough detail to be useful. The central quantitative claim, however, depends entirely on the concatenation extrapolation in Section 5.4. The reader identified this as the weakest assumption, and my stress test sharpens it: because the extrapolated curves are forced to intersect at the fixed point of the fitted power law, the threshold values are an algebraic consequence of the fit rather than a simulated phenomenon. No l=2 or l=3 data are presented, and the paper explicitly says such simulations are computationally infeasible. This is a genuine model risk, not a stylistic preference. The q=5 no-flag threshold of 4.36e-11 illustrates how sensitive the method is to the fitted exponent b. I also noticed a minor numerical inconsistency: the q=3 flag threshold is quoted as 3.24e-4 in Table I and 3.29e-4 in Figure 11, but this is not load-bearing. The appropriate verdict remains CONDITIONAL because the adaptation work is credible but the headline threshold claim needs at least one direct concatenated simulation. My recommended concrete test for q=2 is feasible with existing stabilizer tools and would settle whether the recursive method is sound.","tokens_in":15833,"tokens_out":6157,"duration_ms":66658,"concrete_test":"Run a stabilizer-based simulation of the two-level concatenated 5-qubit code (q=2, 25 data qubits plus ancilla and flag overhead) with the same circuit-level noise model and BM decoder, at ps = 2e-4, 5e-4, and 8e-4. Compare the measured level-2 logical error rate with the recursion prediction a (a ps^b)^b using a=766 and b=1.873. If the predicted and measured curves disagree by more than a factor of about 2, or if the measured crossing with ps differs materially from 4.95e-4, the recursive threshold estimator is invalid and the headline thresholds are not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on Section 5.4, where the authors fit PL(ps) = a ps^b to level-1 circuit-level data and then recursively compose that fitted curve to represent higher concatenation levels. Algebraically, with b > 1, all extrapolated levels intersect at the fixed point p* = a^{-1/(b-1)}. Therefore the headline thresholds in Table I (4.95e-4, 3.24e-4, 2.32e-4) are not measured thresholds of concatenated codes; they are determined entirely by the two parameters a and b extracted from a single-level simulation. The paper provides no level-2 or level-3 simulation and no evidence that the level-1 logical error channel behaves like an independent depolarizing channel at the next level. Correlated logical errors across the five sub-blocks, or failure of the power-law form outside the fitted ps range, would invalidate the recursion and shift the claimed thresholds. The extreme sensitivity of this procedure is visible in the q=5 no-flag threshold of 4.36e-11, which results from b=1.149. The claim of comparable 10^-4 thresholds is therefore a model-based extrapolation rather than an empirically established threshold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the smallest perfect quantum error-correcting code (five qudits) for q = 2, 3, and 5 under both standard depolarizing and circuit-level noise. The authors construct encoding and syndrome-extraction circuits valid for any prime qudit dimension, adapt the MWPM and belief-matching decoders to higher-dimensional codes, and introduce a flag qudit to handle hook errors. Their main quantitative claim is that, after the flag qudit is added, the error thresholds for q = 3 and q = 5 are comparable to the qubit threshold, on the order of 10^-4, with Table I reporting 4.95e-4, 3.24e-4, and 2.32e-4. These thresholds are not obtained from direct simulation of concatenated codes; instead, the paper fits the level-1 logical error probability to the power law P_L(p_s) = a p_s^b (Eq. 16) and recursively composes this fitted curve to represent higher concatenation levels, identifying the threshold as the intersection point of the resulting curves.","tokens_in":16005,"tokens_out":5427,"duration_ms":52045,"significance":"If the threshold comparison were firmly established, the paper would provide concrete evidence that qudit codes can achieve circuit-level thresholds comparable to qubits despite a noise model whose per-qudit error probability grows with dimension, which is of real practical interest for qudit-based platforms. The paper's strengths include a publicly available simulation extension on GitHub, a systematic generalization of matching-graph construction to prime qudit dimensions, and a careful level-1 circuit-level simulation with flag-qudit handling. The clean observation that the BM decoder outperforms MWPM on hyperedge errors at higher dimensions is also a useful contribution. The headline threshold claim, however, rests on a model-based extrapolation and is not directly simulated; this is the main weakness and the reason the paper requires revision.","major_comments":[{"comment":"The reported thresholds are not measured concatenated thresholds. The text fits P_L(p_s) = a p_s^b to level-1 data and then defines higher concatenation levels by composing this same function: P_L^{(l+1)} = P_L ∘ P_L^{(l)}. For b > 1, all such curves intersect at the fixed point p* = a^{-1/(b-1)}, which is exactly the threshold listed in Table I. Thus the threshold is mathematically forced by the two fit parameters a and b; the intersection of the three curves carries no independent information. The paper provides no level-2 or level-3 simulation and no test of the assumption that the level-1 logical error channel behaves like an independent per-step depolarizing channel at the next concatenation level. Because the abstract's central claim of 'comparable error thresholds of the order of 10^-4' depends on this extrapolation, the claim is not empirically established. I recommend either adding a direct simulation of at least one level of concatenation (even for a few physical error rates and one dimension, e.g., qubit with flag) to validate the recursion, or substantially reframing the result as an extrapolated estimate whose systematic uncertainty is discussed explicitly.","section":"Sec. 5.4, Eq. (16), Table I"},{"comment":"The no-flag q=5 threshold of 4.36e-11 is an artifact of the fitted exponent b = 1.149 being close to 1. Since p* = a^{-1/(b-1)}, a small change in b produces an enormous change in the fixed point; the reported uncertainty of ±2.9e-11 from Monte Carlo propagation of the fit parameters reflects only the covariance of the two-parameter fit and not the structural sensitivity of the fixed-point formula to the chosen functional form. This extreme fragility is itself evidence that composing Eq. (16) is unreliable for threshold estimation in this regime. At minimum, the authors should check whether the power-law fit is valid over the range of p_s values that contribute to the fixed point, or they should omit this row from the headline comparison.","section":"Sec. 5.4, Table I, q=5 no-flag row"},{"comment":"The abstract states that 'comparable error thresholds of the order of 10^-4 are obtained', and the conclusion repeats that 'we found that the thresholds for qutrits and ququints were comparable to those of qubits'. Given that the thresholds are extrapolations from a single-level fit rather than direct simulation results, this wording overstates the status of the numbers. The manuscript should either add validation or change the phrasing to something like 'estimated thresholds under the level-by-level power-law assumption' in both the abstract and the conclusion.","section":"Abstract and Sec. 6"}],"minor_comments":[{"comment":"The text in Sec. 5.2 says the error bars represent 90% confidence intervals, while the Figure 7 caption says they correspond to 99% confidence intervals. Please make these consistent.","section":"Sec. 5.2 and Fig. 7 caption"},{"comment":"The text in Sec. 5.3 says 'Error bars represent 99% confidence intervals', but the Figure 9 caption says 'Error bars indicate 90% confidence intervals'. One of these is a typo; please correct it.","section":"Sec. 5.3 and Fig. 9 caption"},{"comment":"Table I lists the qutrit flag threshold as 3.24e-4, while Figure 11 (right panel) labels the threshold star as 3.29e-4. The discrepancy should be resolved.","section":"Table I vs. Fig. 11"},{"comment":"The sentence describing the parameter A ('A is assigned a value of 20 for qubits, 5 for qutrits, and 2 for all data points corresponding to ququints with a flag qudit, except for the left-most point') is difficult to parse. Please clarify which data points use which number of samples and how A was chosen to balance statistics and runtime.","section":"Sec. 5.3, sample-count description"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent in Section 5.4 that the thresholds are obtained by a level-by-level extrapolation, and the code is publicly available, which is a credit. The central problem is that the abstract and conclusion present those extrapolated numbers as measured thresholds, and the fixed-point nature of the power-law recursion means the thresholds carry no independent evidence beyond the two fit parameters. If the authors can add even one direct level-2 concatenation simulation point, or explicitly reframe the result as an extrapolation with systematic uncertainties, the paper could be acceptable; as it stands, the headline comparison is not established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Main take: the paper delivers a clean, reproducible set of qudit decoder tools and circuit-level simulations, and the flag-qudit fix for hook errors is legitimate. The claim that q=3 and q=5 perfect codes have thresholds within a factor of two of the qubit version is plausible, but the specific numbers in Table I are not directly measured. They come from fitting a power law to level-1 data and recursively composing that curve. No level-2 or level-3 concatenated simulation anchors the recursion. The stress-test is right: with b>1, the threshold is p* = a^{-1/(b-1)}, so the table entries are just functions of two fitted parameters. The q=5 no-flag 'threshold' of 4.4e-11 is a red flag: b=1.149 means the logical error rate barely outpaces the physical error rate, and calling that a threshold is misleading.\n\nWhat is genuinely new: the general construction of matching graphs for prime-dimension qudit codes, the belief-matching adaptation, and the flag-qudit circuit for the 5-qudit code. Appendix A is clear and the GitHub code makes the simulations reproducible. The depolarising-noise comparison of BM vs MWPM decoders across q=2,3,5,7 is a useful benchmark.\n\nWhere it is soft: the threshold estimation section. The authors cite [44] and are transparent about using a level-by-level approach, but the abstract states 'comparable error thresholds ... are obtained' without the caveat that these are extrapolated. The assumption that the level-1 logical error channel behaves like an independent depolarising channel at the next level is unvalidated. Correlated logical errors or a power-law form that fails outside the fitted range would shift the numbers. A single level-2 simulation for one dimension would have gone a long way.\n\nAlso, the circuit-level sample sizes are modest (A=2 for q=5 with flag), and the fits are likely dominated by higher-ps points. Still, the qualitative pattern—flag qudits closing the gap between qubits and qudits—is visible directly in the level-1 data, so the main conclusion may survive a more careful threshold calculation.\n\nWho this is for: anyone implementing qudit decoders or evaluating qudit hardware for QEC. It deserves a serious referee. I would send it out, with a request to either run at least one concatenated simulation or reframe the threshold claims as estimates from a single-level extrapolation, and to treat the no-flag q=5 number as an artifact rather than a threshold.","headline":"Useful qudit decoder engineering, but the headline thresholds are fixed points of a power-law fit, not measured concatenated thresholds.","tokens_in":16610,"tokens_out":3064,"would_cite":true,"duration_ms":31859,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P73","81P68"],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"This paper shows that the five-qudit perfect code, when protected by a flag qudit and decoded with belief matching, sustains circuit-level error thresholds around $10^{-4}$ for qudits of dimension 2, 3, and 5, placing qutrits and ququints…","keywords":["qudit","quantum error correction","perfect code","flag qudit","circuit-level noise","error threshold","belief matching","minimum-weight perfect matching"],"falsifier":"Simulate the two-level concatenated 5-qudit code under circuit-level noise for $q=3$ (and ideally $q=5$) at physical error rates around the reported thresholds, then compare the observed logical error rate with the prediction obtained by composing the fitted power law; a significant mismatch would falsify the threshold claim.","tokens_in":15587,"feed_emoji":"⚛️","tokens_out":8208,"duration_ms":70480,"temperature":0.7,"pith_summary":"The paper asks whether qudits—quantum systems with more than two levels—can be used for error correction without paying a large noise penalty. It constructs and simulates the smallest perfect quantum error correction code, which uses five data qudits, for prime dimensions $q=2$, $3$, and $5$, under depolarizing and circuit-level noise. After adding a flag qudit that catches hook errors propagating from noisy ancillas, the simulated logical error rates yield comparable thresholds of $4.95\\times10^{-4}$ (qubit), $3.24\\times10^{-4}$ (qutrit), and $2.32\\times10^{-4}$ (ququint). The paper thus argues that higher-dimensional qudits remain viable for fault-tolerant computation despite a noise model whose Pauli error set grows like $q^2$.","feed_headline":"Flag qudit lifts qudit error thresholds to 10^-4","feed_subtitle":"In the 5-qudit perfect code, q=3 and q=5 land within a factor of two of qubits under circuit-level noise.","key_machinery":"The load-bearing object is the five-qudit perfect code, the smallest stabilizer code that corrects an arbitrary single qudit error, generalized to every prime dimension $q$. Three mechanisms carry the argument: a general encoding circuit derived from the code's parity-check matrix; detector matching graphs in which each ancilla is replaced by $q-1$ nodes, one per nonzero eigenvalue, so that minimum-weight perfect matching and belief matching decoders can process qudit syndromes; and a flag qudit coupled to each ancilla that flags hook errors, with corrections stored in a lookup table. The threshold estimate itself rests on the power-law model $P_L(p_s)=a p_s^b$ for the circuit-level logical error rate, recursively applied to approximate concatenation levels $l=2$ and $l=3$ and intersected to read off the threshold.","core_discovery":"The central claim is that the circuit-level error threshold of the 5-qudit perfect code does not degrade sharply with dimension once hook errors are suppressed. With a flag qudit and the adapted Belief Matching decoder, the estimated thresholds are $4.95\\times10^{-4}$ for $q=2$, $3.24\\times10^{-4}$ for $q=3$, and $2.32\\times10^{-4}$ for $q=5$. Hook errors—correlated errors spread from a noisy ancilla back onto data qudits—were the main obstacle; the flag qudit detects their occurrence and a brute-force lookup table supplies the correction. Without the flag qudit the thresholds are much worse, particularly for $q=5$ where the estimate drops to $4.36\\times10^{-11}$. The threshold estimates are obtained by fitting a power law $P_L(p_s)=a p_s^b$ to one-level circuit-level simulations and composing it to represent concatenation levels two and three, rather than by directly simulating concatenated codes.","pith_inferences":["Editorial inference: if the flag-qudit thresholds survive direct concatenation checks, qudit codes could encode more logical information per physical system without a proportional drop in threshold; a concrete test is to compare logical error rate per encoded qubit-equivalent for the $q=5$ code against the $q=2$ code at the same physical error rate.","Editorial inference: the disconnected structure of the $q=5$ matching graph suggests the cheaper minimum-weight perfect matching decoder should become more competitive at higher dimensions, which could be checked by simulating MWPM with the flag qudit at $q=5$ and $q=7$.","Editorial inference: the level-by-level concatenation extrapolation can be validated at modest cost for $q=3$ by simulating the two-level concatenated code at physical error rates near $3\\times10^{-4}$; agreement would confirm the power-law recursion, while disagreement would indicate correlated logical errors.","Editorial inference: the flag qudit is assumed to be measured and reinitialized fast enough not to add idle time on data qudits; hardware with slow mid-circuit measurement would need the threshold re-derived with a finite flag-reset duration included."],"forward_implications":["A physical implementation of the 5-qudit code with $q=3$ or $q=5$ would need a per-step error probability near $3\\times10^{-4}$ or $2\\times10^{-4}$ to sit below threshold, a target within roughly a factor of two of the qubit requirement.","Two-qudit gates acting on qutrits or ququints can be noisier per step than their qubit counterparts by up to about a factor of two and still yield comparable logical error rates for this code.","Without a flag qudit, the $q=5$ code is effectively unusable (threshold near $10^{-11}$), so hook-error correction is not an optional optimization for higher-dimensional qudit codes.","The decoder construction using $q-1$ nodes per ancilla applies to any prime-dimensional stabilizer code, giving a general route to simulate and decode larger qudit codes with existing matching-based decoders.","Smaller threshold gaps between dimensions imply that the overhead advantage promised by qudits is not erased by circuit-level noise, so comparisons at fixed logical error rate and fixed encoded information become the relevant next benchmark."],"supporting_citations":[{"why":"Defines the perfect five-qubit code that this paper generalizes to arbitrary prime dimension $q$.","marker":"[23]"},{"why":"Supplies the qudit stabilizer formalism and the fault-tolerant gate set used to build the logical operations.","marker":"[28]"},{"why":"Provides the method for deriving an efficient encoding circuit from a parity-check matrix, used to generate the general encoding circuit.","marker":"[32]"},{"why":"The minimum-weight perfect matching implementation whose decoding graphs the paper adapts to qudits.","marker":"[34]"},{"why":"The belief-matching algorithm that the paper adapts to qudit syndromes and uses to obtain its main logical error rates.","marker":"[38]"},{"why":"The flag-qubit error correction technique whose extension to flag qudits is used to catch hook errors.","marker":"[42]"},{"why":"The level-by-level concatenation threshold estimation approach that the paper adopts in place of direct concatenated simulations.","marker":"[44]"}],"fun_headline_variants":["Qudit error correction hits qubit-level thresholds with flag qudit","Flag qudit boosts qudit error thresholds to 10^-4","Higher-dimensional error correction: qudits reach qubit parity","Qudit code survives circuit noise: thresholds ~10^-4","Flag qudit tames hook errors, lifts qudit thresholds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline thresholds rest on the assumption that the logical error probability of one concatenation level can be modeled by composing the same fitted power law $P_L(p_s)=a p_s^b$ with itself, rather than by actually simulating concatenated codes; if first-level logical errors are correlated or the power law fails outside the fitted range, the threshold estimates are not established.","fun_headline_variants_meta":{"raw":{"variants":["Qudit error correction hits qubit-level thresholds with flag qudit","Flag qudit boosts qudit error thresholds to 10^-4","Higher-dimensional error correction: qudits reach qubit parity","Qudit code survives circuit noise: thresholds ~10^-4","Flag qudit tames hook errors, lifts qudit thresholds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000417,"raw_usage":{"total_tokens":2120,"prompt_tokens":882,"completion_tokens":1238,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1148}},"tokens_in":498,"tokens_out":1238,"duration_ms":9332,"temperature":1.0,"reasoning_tokens":1148,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T17:08:52.978099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the two-level concatenated 5-qudit code under circuit-level noise for $q=3$ (and ideally $q=5$) at physical error rates around the reported thresholds, then compare the observed logical error rate with the prediction obtained by composing the fitted power law; a significant mismatch would falsify the threshold claim.","supporting_citations":[{"cited_title":"Field and T","cited_arxiv_id":null,"evidence_quote":"Defines the perfect five-qubit code that this paper generalizes to arbitrary prime dimension $q$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the qudit stabilizer formalism and the fault-tolerant gate set used to build the logical operations."},{"cited_title":"Bianchetti, S","cited_arxiv_id":null,"evidence_quote":"Provides the method for deriving an efficient encoding circuit from a parity-check matrix, used to generate the general encoding circuit."},{"cited_title":"Fernndez de Fuentes, T","cited_arxiv_id":null,"evidence_quote":"The minimum-weight perfect matching implementation whose decoding graphs the paper adapts to qudits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The belief-matching algorithm that the paper adapts to qudit syndromes and uses to obtain its main logical error rates."},{"cited_title":"Luo and X","cited_arxiv_id":null,"evidence_quote":"The flag-qubit error correction technique whose extension to flag qudits is used to catch hook errors."},{"cited_title":"Gottesman, Fault-tolerant quantum computation with higher-dimensional systems, inQuantum Computing and Quantum Communications, edited by C","cited_arxiv_id":null,"evidence_quote":"The level-by-level concatenation threshold estimation approach that the paper adopts in place of direct concatenated simulations."}],"review_version":1}