{"id":"5662be2c-9919-4d84-af03-cbc17d63808a","arxiv_id":"2505.06165","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A daily, per-qubit choice of surface code distance, driven by calibration error rates, can cut QEC qubit overhead by more than half while retaining most qubits.","lead":"The authors propose choosing a quantum error correction code distance per qubit and per day based on IBM calibration data, instead of using one fixed code for everything. They report cutting physical qubit overhead by over 50 percent on a 127-qubit IBM processor while keeping most qubits usable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported >50% overhead saving is computed against a worst-case distance-13 baseline, not against a fixed-but-reasonable distance-9 policy, so the central quantitative claim does not yet demonstrate an advantage of per-qubit, per-day adaptation.","rationale":"The paper's engineering intuition is sound: error rates vary across qubits and days, and fixed QEC can be wasteful. I read the adaptive rule as: for each candidate qubit, pick the smallest distance whose simulated threshold is met, and exclude qubits that need distance greater than 9. The central numerical claim is the >50% reduction in physical-qubit overhead per logical qubit. That number, however, is derived in Section IV.B by comparing a distance-9 cap (81 qubits) against a worst-case distance-13 baseline (169 qubits). This does not measure the benefit of adaptation. A fixed distance-9 policy with the same exclusion step would deliver the same 85% usability and the same 81-qubit overhead. To support the headline, the paper needs to report what the adaptive rule actually chooses day by day and compare the average overhead to a fixed d=9 (or d=11) policy that also excludes high-error qubits. The ibm_brisbane/sherbrooke 71% figure is similarly just 169 vs 49, with no evidence that adaptive assignment beats fixed d=7 for the 80%+ usable qubits. The reader's weakest assumption about mapping single-qubit Pauli-X error rates to a per-gate depolarizing probability is also a genuine concern, because CNOT errors are typically much larger and the simulated thresholds could shift. But that concern applies to the threshold values; the baseline concern undermines the headline even if all thresholds are perfectly correct. I therefore recommend keeping the CONDITIONAL verdict: the concept is plausible, but the quantitative claim requires a corrected comparison and a published distance histogram. I agree only partially with the reader because the baseline mismatch is, in my view, the more immediately falsifiable weak point in the central claim.","tokens_in":7587,"tokens_out":9647,"duration_ms":95152,"concrete_test":"Recompute Section IV.B with two fixed baselines that use the same exclusion rule: (i) fixed distance 9, excluding qubits above its 1e-3 threshold, and (ii) fixed distance 11 with its threshold. Then record the adaptive rule's chosen distance per qubit for each calibration day and compute the resulting average physical-qubit cost per logical qubit. If the adaptive average is not below 81 (the fixed d=9 cost) on ibm_kyiv, or below 49 on the other devices, the >50% and 71% savings are artifacts of comparing to distance 13 rather than evidence of adaptivity. Publishing the per-day distance histograms as a table would settle the issue directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section IV.B the baseline for the 52% saving is 'a single worst-case error rate across all qubits and days' forcing distance 13 (169 physical qubits per logical qubit), while the adaptive scenario is capped at distance 9 (81 qubits). That comparison isolates the choice of maximum distance, not adaptivity: a fixed distance-9 policy that excludes qubits above the distance-9 threshold would give the same 85% usable-qubit fraction and the same 81 qubits per logical qubit. The paper never reports the histogram of distances actually assigned by the proposed rule over the 12 days, nor the average physical-qubit overhead per logical qubit under adaptive assignment. Without that, the claim that day-by-day, per-qubit adaptation, rather than simply choosing a moderate fixed distance and excluding bad qubits, creates the >50% saving is unsupported. The same issue affects the ibm_brisbane/ibm_sherbrooke result: Section IV.B says that because over 80% of qubits remain usable even with distance-7, 'we can consider using this stricter distance,' which is again a fixed-distance choice, and the 71% saving is just the ratio 169/49. Section IV.C lists limitations about calibration reliability and decoder adjustment, but does not flag this baseline mismatch or the absence of the actual adaptive distance distribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 12 days of calibration data from IBM's 127-qubit device ibm_kyiv, and additional data from ibm_brisbane and ibm_sherbrooke, showing temporal and spatial variation in Pauli-X and CNOT error rates. It proposes an adaptive QEC strategy that selects a rotated surface-code distance per qubit and per day, using Stim/PyMatching simulations to determine the smallest distance that meets a target logical error rate of 10^-6 and excluding qubits that would require a distance above an allowed maximum. The central claim is that this adaptive assignment reduces physical-qubit overhead by more than 50% per logical qubit on ibm_kyiv and up to 71% on the other two devices, while preserving access to 80-100% of usable qubits.","tokens_in":7852,"tokens_out":3169,"duration_ms":33578,"significance":"If the quantitative claims were properly supported, the paper would make a useful practical contribution: dynamic, calibration-aware selection of code distance is a plausible resource-optimization strategy for near-term QEC, and the paper grounds it in real device data rather than synthetic noise models. The strengths are the use of publicly available calibration data, the concrete pipeline of distance assignment from simulated logical-error thresholds, and the explicit target logical error rate. However, the headline savings are currently computed against a comparison that isolates the maximum allowed distance rather than adaptivity, and the noise-model mapping from calibration Pauli-X rates to the circuit-level depolarizing simulation is not justified. These issues are load-bearing for the central claim, so the paper needs substantial revision before the results can be accepted.","major_comments":[{"comment":"The reported 52% saving on ibm_kyiv compares a worst-case distance-13 baseline (169 physical qubits per logical qubit) with an adaptive scheme capped at distance-9 (81 physical qubits). This comparison measures the choice of maximum distance, not the benefit of day-by-day, per-qubit adaptation: a fixed distance-9 policy that simply excludes qubits with error rate above the distance-9 threshold would yield the same usable-qubit fraction (~85%) and the same 81 physical qubits per logical qubit. The same issue affects the ibm_brisbane/ibm_sherbrooke result, where the 71% saving is essentially the ratio 169/49 from choosing distance-7 instead of distance-13. To support the claim that adaptation creates the savings, the paper should report the actual distribution of assigned distances per day, the average physical-qubit overhead per logical qubit under the adaptive rule, and a comparison against fixed-distance policies at each candidate distance (e.g., d=7, d=9, d=11) with the same qubit-exclusion rule.","section":"Section IV.B"},{"comment":"The paper states that 'we restrict our analysis and logical error simulations to only single-qubit Pauli-X error rates' (Section III.A), but the simulation in Section III.B is described as a circuit-level model with symmetric depolarizing errors applied before and after Clifford gates, resets, and measurements. Since the calibration data shows CNOT error rates in the 10^-2 to 10^-1 range, substantially above the simulated thresholds quoted in Section IV.B, the mapping from the daily Pauli-X calibration value to the per-gate depolarizing probability p in Stim is not justified. If the simulation includes CNOT errors at the same p, then the thresholds cannot be correct for the real hardware; if it does not, the claim that the simulated thresholds govern usability of real qubits is unsupported. The authors should either simulate a noise model that actually uses only the single-qubit Pauli-X error rates with no CNOT/readout errors, or justify explicitly why excluding CNOT, resets, and measurement errors does not change which qubits are classified as usable.","section":"Section III.A and III.B"},{"comment":"The threshold values used for qubit usability (7x10^-4 for d=7, 10^-3 for d=9, 2x10^-3 for d=11, 7x10^-3 for d=13) are read from the simulated curves in Fig. 4, but the paper does not report the number of simulation shots, the statistical uncertainty in the logical error rates, or the exact criterion used to extract each threshold from the crossing of the p_L = 10^-6 line. Because many calibration error-rate values in Fig. 3(a) lie very close to these thresholds, small simulation uncertainties could change the day-by-day usable-qubit percentages and hence the final overhead savings. The authors should provide error bars or confidence intervals on the simulated logical error rates and state how the thresholds were obtained.","section":"Section IV.B and Fig. 5"}],"minor_comments":[{"comment":"There is a typo in the first paragraph of the introduction: 'In the the rest of the paper' should be 'In the rest of the paper'.","section":"Section I"},{"comment":"There are formatting issues in the text: 'distanced typically requires' should be 'distance d typically requires', and mathematical expressions such as 'd2 qubits' and '132 = 169' should be typeset properly as d^2 and 13^2 = 169 respectively.","section":"Section II.A"},{"comment":"The text '92 = 81 physical qubits' should be written as 9^2 = 81 to avoid confusion.","section":"Section IV.B"},{"comment":"The description of the Stim simulation should state which gates are included (e.g., Hadamard, S, CNOT, measurements, resets) and whether the depolarizing error rate for single-qubit gates equals that for two-qubit gates; this information is needed to reproduce the thresholds.","section":"Section III.B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a timely and practical question, and the raw calibration data analysis is potentially useful. However, the central quantitative claim of an adaptive advantage is not currently supported: the headline savings are an artifact of comparing distance-13 with distance-9/7 under fixed-distance reasoning, and the noise-model mapping from Pauli-X calibration rates to the circuit-level Stim simulation is inconsistent with the stated simulation model. These are fixable within the manuscript's scope by rerunning the analysis with fair baselines and a justified noise model, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a sensible engineering idea with a misleading headline number. The >50% overhead saving is computed against a worst-case distance-13 baseline, so it mostly reflects the choice of maximum distance rather than the adaptive rule. A fixed distance-9 policy that excludes bad qubits would give the same usable-qubit fraction and the same 81 qubits per logical qubit. The paper never reports the actual distribution of assigned distances, so the advantage of day-by-day adaptation over a moderate fixed distance is unproven.\n\nWhat is genuinely useful: the authors assemble 12 days of calibration data from ibm_kyiv, show that Pauli-X and CNOT error rates fluctuate significantly across qubits and time, and then run standard Stim/PyMatching simulations of rotated surface codes to extract per-distance thresholds. That part is reproducible and the direction of the argument is plausible. Applying the same procedure to two other 127-qubit devices is a nice consistency check. If a user is willing to reconfigure the QEC stack daily, per-qubit distance selection is a reasonable resource-management knob.\n\nThe soft spots are real and not minor. First, the baseline comparison is unfair. Baseline distance-13 is chosen as the worst-case over all qubits and days, while adaptive is capped at distance-9. That isolates the maximum-distance choice, not adaptivity. The 71% saving on Brisbane/Sherbrooke is just the ratio 169/49 for fixed distances; it has nothing to do with the adaptive rule. Second, the noise model mismatch: the calibration Pauli-X error rate is used directly as the depolarizing probability p in a circuit-level noise model that includes CNOTs, resets, and measurements. These are not the same quantity, and the paper does not justify the mapping. This could shift the thresholds and therefore the usable-qubit counts. Third, the evaluation choices (max distance 9, target 1e-6, cutoff 8e-3) look post hoc; there is no sensitivity analysis.\n\nThe paper should not be desk-rejected because the idea is testable and the data are real. But the current quantitative claims need major revision. The authors should report the actual histogram of distances assigned per day, compare adaptive assignment against a fixed distance-9 (and 11) policy with the same exclusion rule, and either validate the noise mapping or switch to a more honest per-gate error model.\n\nI'd send this to peer review with a request for major revision. It is a solid workshop-level contribution as is, but with a redesigned evaluation it could be a genuinely useful reference for NISQ-era resource optimization.","headline":"The adaptive-QEC idea is sensible, but the headline savings number is an artifact of an unfavorable baseline, so the paper needs a redesigned evaluation before its quantitative claims can be trusted.","tokens_in":8385,"tokens_out":2266,"would_cite":false,"duration_ms":21916,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["03.67.Pp"],"model":"deepseek-v4-flash","headline":"Per-qubit, per-day code distance selection halves the physical-qubit cost of quantum error correction on 127-qubit devices.","keywords":["quantum error correction","surface code","adaptive code distance","temporal error variation","qubit quality variability","calibration data","logical error rate","resource overhead"],"falsifier":"Run the proposed assignment on a real device for one day: assemble distance-9 logical qubits from qubits whose daily Pauli-X error is near $10^{-3}$ and measure the logical error rate; the method predicts it should meet the $10^{-6}$ target, so observing a logical error rate substantially above that level would show the simulation thresholds are too optimistic for real hardware.","tokens_in":7316,"feed_emoji":"⚛️","tokens_out":8435,"duration_ms":78487,"temperature":0.7,"pith_summary":"Qubit error rates on real superconducting processors drift from day to day and differ across qubits, so a single fixed surface-code distance is either too weak for noisy qubits or wasteful for stable ones. The paper claims that selecting the code distance per qubit and per calibration day, using the smallest distance whose simulated logical error rate still meets the target, avoids both failure modes. On 12 days of calibration data from a 127-qubit device, this adaptive assignment keeps 85-100% of qubits usable while cutting physical qubit overhead by over 50% per logical qubit; on two other 127-qubit devices the savings reach up to 71%. The practical point is that QEC can be configured from data already being collected, instead of reserving a worst-case error rate.","feed_headline":"Per-qubit code distances cut error-correction overhead in half","feed_subtitle":"Daily calibration picks the right code distance for each qubit, keeping 85-100% of qubits usable on 127-qubit hardware.","key_machinery":"The carrying object is the simulated logical-error-rate curve for rotated surface codes, which converts a physical Pauli-X error rate into a minimum code distance that still meets a target logical error rate such as $10^{-6}$. This curve supplies the per-distance cutoff thresholds used by the adaptive assignment rule: sort qubits by the day's calibration error, read off the smallest distance whose simulated logical error rate stays below the target, and drop qubits needing a distance above the practical maximum. The rule is what turns raw calibration data into a concrete resource allocation, and the quadratic qubit count of the rotated surface code, about $d^2$ qubits per logical qubit, is what makes the savings material.","core_discovery":"The central claim is that a resource-efficient QEC configuration should be a per-qubit, per-day decision rather than a global constant. Using daily Pauli-X calibration error rates, the paper simulates rotated surface codes of distances $d=3$ through $d=21$ to find, for each distance, the physical error threshold that still reaches a logical error rate of $10^{-6}$; the thresholds are $7\\times10^{-4}$ for $d=7$, $10^{-3}$ for $d=9$, $2\\times10^{-3}$ for $d=11$, and $7\\times10^{-3}$ for $d=13$. The adaptive rule assigns each qubit the smallest distance meeting the target from that day's calibration, and excludes qubits whose error rate would require a distance above the allowed maximum (for example, above 9). Compared with a worst-case fixed distance-13 configuration using 169 physical qubits per logical qubit, the adaptive scheme with maximum distance 9 uses about 81 qubits per logical qubit while keeping roughly 85% of the device's qubits usable, giving the reported overhead reduction of over 50%, with up to 71% on two other devices where distance-7 can be used.","pith_inferences":["The paper's thresholds come from a symmetric depolarizing model; extending the same calibration-driven assignment to CNOT error rates, noise bias, or correlated errors would likely shift the usable-qubit fractions and could make the savings device-specific.","A system-level scheduler might reasonably trade the logical error rate target for more logical qubits: relaxing the target from $10^{-6}$ to $10^{-5}$ would admit more qubits at distance 7 and could cut overhead further than reported.","The method's daily refresh assumes calibration remains valid for the whole execution window; a direct test would be to compare morning and evening calibrations on the same qubits and measure whether the assigned distance still meets the logical target.","The per-qubit independence assumption ignores crosstalk and competing resource demands; a global optimizer over all logical qubits simultaneously is a natural next step."],"forward_implications":["QEC configuration can be refreshed from calibration data that hardware providers already publish, so day-to-day drift no longer requires a permanent worst-case code distance.","A compiler or scheduler can use the per-day distance map to place logical qubits only on qubits whose error rate supports the target, reducing wasted physical qubits.","On the devices studied, the adaptive rule keeps 80-100% of physical qubits available for encoding, so practical logical-qubit capacity does not collapse when a few qubits degrade.","High-error outlier qubits are excluded rather than encoded at prohibitive distances, concentrating the error budget on qubits that can actually meet the logical target."],"supporting_citations":[{"why":"Defines the surface code framework, the exponential logical-error scaling, and the $10^{-6}$ target that the adaptive thresholds are built around.","marker":"[3]"},{"why":"Establishes topological quantum memory error suppression and the threshold behavior that motivates distance selection.","marker":"[5]"},{"why":"Documents that qubits vary in quality in real NISQ devices, the empirical premise for per-qubit adaptation.","marker":"[12]"},{"why":"Provides the rotated surface code layout used throughout the paper and its low-distance noise behavior.","marker":"[14]"},{"why":"Supplies the simulation methodology for logical error rates as a function of physical error rate that generates the distance thresholds.","marker":"[15]"},{"why":"Earlier adaptive surface code work that adapts to defects rather than temporal quality variation, the contrast for this paper's contribution.","marker":"[16]"}],"fun_headline_variants":["Daily qubit error rates drive code distance, cutting overhead 50%","Adaptive code distance per qubit halves quantum error-correction cost","Time-varying qubit quality? Adaptive QEC reduces overhead by half","Per-day calibration chooses code distance, slashing overhead 50%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a qubit's daily Pauli-X calibration error rate can stand in for the full per-gate error probability used in the logical-error simulation, even though the simulation includes all gates, resets, and measurements while the calibration analysis uses only single-qubit Pauli-X values.","fun_headline_variants_meta":{"raw":{"variants":["Daily qubit error rates drive code distance, cutting overhead 50%","Adaptive code distance per qubit halves quantum error-correction cost","Time-varying qubit quality? Adaptive QEC reduces overhead by half","Per-day calibration chooses code distance, slashing overhead 50%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3376,"prompt_tokens":1079,"completion_tokens":2297,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":2221}},"tokens_in":695,"tokens_out":2297,"duration_ms":17890,"temperature":1.0,"reasoning_tokens":2221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:47:26.751559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed assignment on a real device for one day: assemble distance-9 logical qubits from qubits whose daily Pauli-X error is near $10^{-3}$ and measure the logical error rate; the method predicts it should meet the $10^{-6}$ target, so observing a logical error rate substantially above that level would show the simulation thresholds are too optimistic for real hardware.","supporting_citations":[{"cited_title":"Not all qubits are created equal: A case for variability-aware policies for nisq-era quantum computers,","cited_arxiv_id":null,"evidence_quote":"Documents that qubits vary in quality in real NISQ devices, the empirical premise for per-qubit adaptation."},{"cited_title":"Low-distance surface codes under realistic quantum noise,","cited_arxiv_id":null,"evidence_quote":"Provides the rotated surface code layout used throughout the paper and its low-distance noise behavior."},{"cited_title":"Q-pandora unboxed: Character- izing resilience of quantum error correction codes under biased noise,","cited_arxiv_id":null,"evidence_quote":"Supplies the simulation methodology for logical error rates as a function of physical error rate that generates the distance thresholds."},{"cited_title":"Adaptive surface code for quantum error correction in the presence of temporary or permanent defects,","cited_arxiv_id":null,"evidence_quote":"Earlier adaptive surface code work that adapts to defects rather than temporal quality variation, the contrast for this paper's contribution."}],"review_version":1}