{"id":"5902cb0d-3f74-473d-8583-695379e4ecca","arxiv_id":"2607.22276","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Fault-tolerant quantum computing is reframed as an empirical question about watts per decade of suppressed logical error, with a concrete two-measurement protocol proposed.","lead":"This paper argues that the decades-long debate over whether quantum computers can be made fault-tolerant reduces to one measurable number: how much extra power it costs to suppress logical errors by a factor of ten. It proposes that labs run two simple power-meter measurements, which would turn a philosophical dispute about noise models into an empirical one.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The watts-per-decade test lacks a derived null prediction: under the threshold theorem's own polylog overhead, the marginal power per decade grows with code distance, so 'nearly flat' is not a definite falsifiable prediction.","rationale":"The reader's verdict is CONDITIONAL, and my concern supports that verdict without moving it. The reader identifies the circuit-to-wattage bridge as asserted rather than derived. I sharpen that concern: even granting a monotone translation from gate/qubit counts to power, the theorem's own polylog overhead produces a watts-per-decade slope that grows with code distance. This means the proposed experiment lacks a pre-defined null hypothesis: the paper says the theorem predicts 'nearly flat' and the inventory predicts 'climbs', but never states what quantitative slope would count as flat or climbing. Without that, the two measurements cannot settle the debate as claimed. This is a load-bearing gap because the paper's central contribution is precisely to reduce the feasibility question to an empirical measurement. The concern is addressable: the author could derive a quantitative prediction from the threshold theorem and explicit power-per-qubit assumptions, or explicitly downgrade the claim to a finite-scale engineering benchmark rather than a test of the theorem. The paper has genuine independent support in the published hardware record and in its explicit, symmetrical falsification proposal (Objection 6), so the issue is not fatal. The appropriate verdict remains CONDITIONAL: accept conditionally on providing the missing derivation or narrowing the claim. I mark partial agreement with the reader because the reader's stated weakest assumption concerns non-monotonicity or fixed-cost dominance, whereas my concern holds even in the idealized monotone, fixed-cost-free case; the underlying bridge issue is the same, but the specific failure mode I identify is different and more fundamental.","tokens_in":12271,"tokens_out":7811,"duration_ms":68862,"concrete_test":"Derive the theoretical watts-per-decade curve from the surface-code overhead formulas (Ref. [16]) and the Willow parameters (Ref. [9]): take N(d) = 2d^2 - 1, ε_L(d) = C Λ^{-(d+1)/2}, and P(d) = P0 N(d). Plot the marginal power per decade, ΔP(d) = P0[N(d + Δd) - N(d)] with Δd = 2 ln 10 / ln Λ, for d from 3 to 30. If this curve is monotonically increasing by more than a factor of 2 over that range, then the paper's 'nearly flat' claim fails even under the ideal theorem resource count, and the proposed flat-vs-climbing dichotomy is not well defined. Alternatively, specify a threshold slope δ such that the theorem's derived slope is below δ and the inventory's is above δ at the eight-hour scale; if no such δ can be found, the falsifier is vacuous.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical dichotomy requires that the threshold theorem predicts a nearly flat watts-per-decade curve (Section 3) and that the inventory predicts a climbing curve (Sections 4 and 8). But the theorem's prediction is never derived. A minimal derivation from the surface-code formulas the paper itself cites undermines 'nearly flat'. For a distance-d surface code, physical qubits N = 2d^2 - 1 and logical error per cycle scales as Λ^{-(d+1)/2}. One decade of error suppression therefore requires Δd = 2 ln 10 / ln Λ distance steps. With constant power per physical qubit P0, the marginal power per decade is ΔP(d) = P0[N(d + Δd) - N(d)] ≈ 4 P0 d Δd, which grows linearly with d. Thus even an idealized threshold-theorem machine, with no calibration, decoding, or entropy-flush costs, would show a rising watts-per-decade slope, not a flat one. The paper offers no quantitative threshold separating 'flat' from 'climbing'. Without such a threshold, any measured slope can be post-hoc rationalized as supporting either side, and the proposed two measurements cannot settle the FTQC debate. This is distinct from the fixed-cost concern (already addressed by defining a marginal slope in Objection 5) and from the asymptotic-classification concern (Objection 12). It is a missing null hypothesis at the heart of the empirical test.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript applies the Hagar–Sergioli resource-bounded notion of objective probability to fault-tolerant quantum computing. The threshold theorem is recast as a classification claim: below threshold, error-corrected logical states belong to the class of cheaply realizable states whose probability remains near one as the machine grows. The author proposes a measurable quantity—watts per decade of suppressed logical error—argues that the threshold theorem predicts a nearly flat watts-per-decade curve, and lists four resources unpriced in the original theorem (calibration, decoding, coherence, entropy flush) that are expected to make the physically realized curve climb. Section 8 proposes two measurements, one on quantum hardware and one on a classical competitor, that would settle the three-decade feasibility debate.","tokens_in":12661,"tokens_out":5849,"duration_ms":58993,"significance":"The paper has a genuinely attractive core: it turns an often-metaphysical debate about fault-tolerance assumptions into a request for a concrete, marginal power measurement, and it honestly confronts several obvious objections, including the fixed-cost objection. The compilation of published drift, decoder-latency, and cryogenic-power figures provides useful evidence that the four unpriced resources are real engineering costs. However, the central quantitative claim—the null prediction attributed to the threshold theorem—is never derived, and the proposed experiment is given no decision rule. As it stands, the paper does not yet establish that the two measurements would settle anything, because no falsifiable dichotomy is defined.","major_comments":[{"comment":"The sentence “The theorem predicts this curve is nearly flat” is asserted, not derived. Using the surface-code resource count cited in the same section (N = 2d^2 − 1 physical qubits) and the standard logical-error scaling Λ^{−(d+1)/2}, one decade of error suppression requires Δd = 2 ln 10 / ln Λ distance steps. With constant power per physical qubit P0, the marginal power per decade is approximately P0[N(d+Δd) − N(d)] ≈ 4 P0 d Δd, which grows linearly with d. Thus an idealized threshold-theorem machine with instantaneous decoding, no calibration, and free ancillas would show a rising watts-per-decade slope, not a flat one, unless additional assumptions or a different meaning of “nearly flat” are supplied. The null curve for the proposed experiment is therefore missing.","section":"Section 3"},{"comment":"No calculation connects the measure P = |A|/|S| to measured watts. Equation (1) is never instantiated for a concrete code, a budget, or a target logical error. The paper moves directly from circuit-complexity overhead (polylog(1/ε)) to “the price is denominated in watts,” but the mapping from gate/qubit counts to a real machine’s power draw is nontrivial—fixed costs, control electronics, and cryogenic overhead are not monotone simple functions of circuit size. The bridge may exist, but it is not shown; without it, “the theorem predicts this curve” is a slogan rather than a theorem.","section":"Eq. (1) / Section 3"},{"comment":"The decision rule is unspecified. The text repeatedly says a flat marginal slope across two generations settles for the theorem and a climbing slope for the inventory, but it never states what numerical slope counts as “flat,” what error bars or confidence criterion are required, or how engineering improvements “already on the books” are to be separated from the resource costs. A measurement with no pre-registered threshold cannot adjudicate a dispute; any observed slope can be rationalized post hoc as either “still nearly flat” or “already climbing.” This is a load-bearing gap in the empirical claim.","section":"Section 8 / Objection 6"},{"comment":"The paper concedes that a finite slope “classifies nothing” in the asymptotic Poly/Exp sense, but the reply does not replace the asymptotic classification with a finite-scale quantitative prediction. Naming the eight-hour benchmark as the scale of interest does not derive the expected watts-per-decade slope under the threshold theorem at that scale. The same gap recurs in Objection 11: Λ is called the “numerator” of the 2011 ratio, but the denominator is never computed from the theorem. The central dichotomy thus remains unsupported at exactly the point where the paper needs a falsifiable null.","section":"Section 7, Objection 12"},{"comment":"The four unpriced resources are documented with published figures, but they are never assembled into even a toy model of the marginal watts-per-decade curve. The cited numbers (e.g., 41.6% and 135.5% error increases after eight hours without calibration; 63 µs decoder latency against 1.1 µs cycle time; 6.25 W per physical qubit in the RAND estimate) support the existence of these costs, not the shape or slope of the curve. To claim that the inventory “predicts” a climbing curve, the paper needs at least a scaling argument showing how these terms enter the marginal power and why they dominate the threshold theorem’s own overhead.","section":"Section 4"}],"minor_comments":[{"comment":"The second proposed measurement, the “symmetric classical curve,” is not defined. The accounting rules for the classical competitor (which classical machine, which error model, which drift and calibration costs) are left unspecified, so the crossing point the text refers to cannot actually be located.","section":"Section 8"},{"comment":"The active-vs-passive discussion is presented as a corollary of the measure, but the historical claim that the field “entered where entry was cheap” is anecdotal and unsupported. This section could be shortened or reframed as an illustration rather than an empirical explanation.","section":"Section 5"},{"comment":"The term “watts per decade” should be defined consistently as a marginal slope, not a single power reading. The distinction is made only in Objection 5 and Section 8; it belongs in the abstract and in the first definition, since the whole proposal rests on it.","section":"Abstract / Section 3"},{"comment":"Several references are 2026 arXiv preprints (e.g., [17], [26], [28]) and are used to support specific empirical claims. The manuscript should mark these as preprints and, where possible, cite peer-reviewed versions, so readers can assess the maturity of the evidence.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The proposal is interesting and the paper is honest about many of its own limitations, but the central experimental dichotomy is not yet defined. The most serious issue is that a simple calculation from the surface-code formulas the paper itself cites suggests the threshold theorem predicts a rising, not flat, watts-per-decade marginal slope. The author needs to derive the predicted null curve (or replace the flat-vs-climbing dichotomy with a quantitative comparison), and specify a pre-registered decision rule. These are substantial but fixable within the manuscript's scope. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: the paper's central empirical test doesn't have a derived null prediction. It claims the threshold theorem predicts a nearly flat watts-per-decade curve (Section 3), but no calculation backs that up. If you take the surface-code formulas the paper itself cites—N = 2d^2 − 1, logical error per cycle scaling as Λ^−(d+1)/2—one decade of suppression requires Δd ≈ 2 ln 10 / ln Λ distance steps, so the marginal power per decade grows roughly as 4 P0 d Δd, linearly in d. Even a perfectly ideal threshold machine with no calibration, decoding, or entropy costs would show a rising slope. That is a load-bearing gap: without a quantitative threshold separating 'flat' from 'climbing,' any measured slope can be rationalized after the fact, and the two measurements don't settle anything.\n\nWhat's good: the reframing itself is genuinely useful. It gives the FTQC debate a common currency—watts per decade—and a concrete, cheap experiment. The paper is unusually self-aware: twelve objections with replies, including the fixed-cost concern (Objection 5) and the asymptotic classification worry (Objection 12). The four unpriced resources are documented with real numbers from the published record. The author explicitly states the meter is independent of the 2011 probability interpretation, so a reader who rejects the philosophy can still run the measurement. That's good scientific hygiene.\n\nSoft spots: besides the missing null prediction, the paper concedes in Objection 12 that a finite slope cannot prove Poly or Exp. That means even with the measurement, the interpretation is open unless a quantitative threshold is defined. The bridge from circuit complexity to wattage is asserted, not modeled. And some cited references (e.g., arXiv:2606.30805) look very new; not a problem per se, but worth checking.\n\nBottom line: this is a serious paper, honestly argued, and it should go to peer review. But the author needs to either derive the predicted curve from a concrete overhead model or explicitly downgrade the claim to a finite-scale engineering criterion. As written, the central test is not yet well-posed. I'd bring it to reading group; I wouldn't cite it in my own work until the prediction is sharpened.","headline":"A smart, honest proposal that reframes FTQC feasibility as a watts-per-decade measurement—but the central 'flat curve' prediction is asserted, not derived, and a simple surface-code calculation gives a rising curve instead.","tokens_in":13067,"tokens_out":2766,"would_cite":false,"duration_ms":23661,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper recasts the threshold theorem as a claim about watts per decade of suppressed logical error, and says two power-meter measurements can settle the debate.","keywords":["fault-tolerant quantum computing","threshold theorem","objective probability","resource-bounded realizability","watts per decade","quantum error correction","calibration cost","energy accounting"],"falsifier":"Measure total facility power—dilution refrigerator, control electronics, and decoder—while running surface-code error correction at distances 3, 5, and 7 on the same processor, with all recalibration downtime included in the energy total, and plot watts per decade of suppressed logical error. Repeat across two successive hardware generations spanning about ten times the qubit count. If the marginal slope is statistically flat, the theorem's prediction survives; if it climbs, the omitted inventory wins. The paper identifies this as a measurement that can be performed on existing hardware with a","tokens_in":12156,"feed_emoji":"⚡","tokens_out":6483,"duration_ms":56631,"temperature":0.7,"pith_summary":"This paper argues that the three-decade dispute over whether fault-tolerant quantum computing is physically feasible is not fundamentally about noise models. Under a resource-bounded interpretation of objective probability, a state's probability is the fraction of affordable dynamical paths to it within an energy-time budget, and the threshold theorem becomes a classification claim: error-corrected logical states are asserted to be cheaply realizable, so their probability stays near 1 as the machine grows. The author points out that the theorem's original resource inventory charged zero for four real costs: calibration of drifting hardware, decoding inside the correction cycle, coherence as a finite time budget, and entropy flushing through fresh ancillas. The paper's central proposal is to measure watts per decade of suppressed logical error, and to treat the slope of that curve as the empirical question: nearly flat if the theorem is right, climbing if the missing costs matter. A sympathetic reader would care because the proposal turns a long philosophical debate into a power-meter experiment that can be run on near-term hardware.","feed_headline":"One power meter can settle the fault-tolerance debate","feed_subtitle":"New metric: watts per decade of suppressed logical error; theorem predicts a flat slope, unpriced costs predict a climb.","key_machinery":"The load-bearing object is the resource-bounded objective probability P = |A|/|S|: over the finite set S of dynamical evolutions that could carry a system to a target state within a fixed energy-over-time budget, A is the subset whose cost fits that budget. The paper's operational counterpart is the watts-per-decade slope, the marginal facility power needed to lower logical error by one factor of ten as code distance and machine size grow. These two objects work together: the ratio turns a theorem about circuit combinatorics into a statement about which complexity class a state occupies, and the slope turns that classification into something a power meter can check. A secondary ingredient is","core_discovery":"On the paper's own terms, the discovery is that the threshold theorem, read through the 2011 resource-bounded probability measure, asserts membership of error-corrected logical states in the cheap complexity class Poly: the probability of realizing a logical state at error rate epsilon stays near 1 because the resources to hold it there grow only polylogarithmically in 1/epsilon and polynomially in size. The theorem derived that membership from an inventory that priced calibration, decoding, coherence time, and ancilla entropy-flush at zero; each of these is now known to consume measurable power. The paper's central proposal is to track the marginal power cost per decade of suppressed logica","pith_inferences":["The watts-per-decade meter is architecture-neutral in principle, so it could be run on photonic, trapped-ion, or neutral-atom machines; the paper's examples are drawn from superconducting hardware, but the unit transfers.","If power scaling is dominated by fixed costs such as refrigerator load, a single-machine reading will look flat; the paper anticipates this by defining the quantity as a marginal slope across generations, but a reader should expect the first one-machine data to be noisy on that axis.","Applying the same resource accounting to the classical competitor, as the paper suggests, would give quantum advantage claims a testable crossing point: the two curves meeting at a named machine size and power budget.","The calibration line item suggests a testable sub-prediction: machines with active in-situ calibration loops should show a flatter watts-per-decade slope than machines that halt computation for recalibration, and the difference should grow with machine size."],"forward_implications":["If the theorem's classification claim is read in this way, no noise-model assumption can settle feasibility by itself; the resource curve is the arbiter.","The four zero-priced resources—calibration, decoding, coherence time, and ancilla entropy-flush—must appear on the meter, and each already has published cost data.","A watts-per-decade slope is measurable on a single current processor: distances 3, 5, and 7 have been run in one campaign, so the per-distance marginal cost can be computed now.","Two successive hardware generations spanning an order of magnitude in qubit count, with recalibration downtime and decoder power included, are enough to tell whether the slope is flat or climbing.","Under this accounting, 'i.i.d. noise' stops being a free assumption and becomes a maintained burden whose cost appears on the electricity bill."],"fun_headline_variants":["Fault-tolerance debate settles by watts per decade","Threshold theorem's zero-price resources exposed","Quantum error correction's hidden power cost","Measure watts to end the FTQC argument"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The weakest load-bearing premise is that the threshold theorem's resource overhead translates monotonically into a real machine's power draw: if a machine's electricity bill is dominated by fixed costs or by effects that do not track gate and qubit counts, then the watts-per-decade curve measures the hardware, not the theorem.","fun_headline_variants_meta":{"raw":{"variants":["Fault-tolerance debate settles by watts per decade","Threshold theorem's zero-price resources exposed","Quantum error correction's hidden power cost","Measure watts to end the FTQC argument"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000384,"raw_usage":{"total_tokens":1917,"prompt_tokens":838,"completion_tokens":1079,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":1036}},"tokens_in":582,"tokens_out":1079,"duration_ms":8364,"temperature":1.0,"reasoning_tokens":1036,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:15:24.333680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure total facility power—dilution refrigerator, control electronics, and decoder—while running surface-code error correction at distances 3, 5, and 7 on the same processor, with all recalibration downtime included in the energy total, and plot watts per decade of suppressed logical error. Repeat across two successive hardware generations spanning about ten times the qubit count. If the marginal slope is statistically flat, the theorem's prediction survives; if it climbs, the omitted inventory wins. The paper identifies this as a measurement that can be performed on existing hardware with a","supporting_citations":[],"review_version":1}