{"id":"4c3a3656-60fa-4da2-bc13-158477bc4486","arxiv_id":"2608.12858","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For Rydberg-atom maximum independent set optimization, shots to reach near-exact solutions grow exponentially with system size, while relaxed targets need only a few shots, with quantum annealing showing a consistent but target-dependent advantage over a density-matched random baseline.","lead":"This paper measures how many runs of a neutral-atom quantum optimizer are needed to reach a given solution quality, and finds that near-perfect answers grow exponentially harder while approximate answers need almost no extra runs. It also isolates a genuine quantum concentration effect by comparing against random outputs with the same excitation density.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Near-exact STS claim rests on a one-parameter shell model that is not validated at the exact shell: direct counts exceed model STS by up to ~3x, so the exponential reduction is extrapolated, not measured.","rationale":"","tokens_in":20682,"tokens_out":11929,"duration_ms":125813,"concrete_test":"Compute direct-count STS(r=1) for both the annealing outputs and the excitation-matched Bernoulli baseline on all 96 instances, using 10x shot statistics (5000 shots per instance) or, failing that, a bias-corrected estimate from the existing 500 shots; compare p0 = empirical fraction in shell 0 for each arm and form the quantum-to-baseline STS ratio with a jackknife confidence interval. If the direct-count ratio is positive and comparable to the shell-model prediction, the model extrapolation is not load-bearing. If the ratio is within error of unity, the exact-target advantage is an artifact of Eq. (5) and the near-exact claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline near-exact regime is controlled by pi0, the per-shot probability of landing in the MIS shell j=0. The paper's Eq. (5) is asserted to describe the postprocessed distribution with one beta valid at all shell depths, and this is what converts the fitted beta difference into the claimed exponential shot-cost reduction. On the actual annealing data this is the shell where the model is least secure: Sec. IIIE reports that at J=0 the model overestimates pi0 by up to a factor of about 3, giving shell-model STS(r=1) medians of 11/27/63/140 versus direct-count values 14/49/190/459 for the four size groups, with the xlarge direct count resting on only 5 hits in 500 shots. Appendix D further shows that the D_KL<0.1 goodness-of-fit criterion is met by >97% of the synthetic mock datasets (Figs. 4e-h) but by only 78% of the 96 annealing instances, and that no single beta reproduces both the bulk shells and pi0. The near-exact quantum advantage is therefore not a direct measurement; it is an extrapolation from a model known to be wrong at the exact target. The authors flag this and call shell-model STS(r=1) a lower bound, but the central claim of an exponential shot-cost reduction by annealing at r~1 still rests on that model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a shots-to-approximate-solution metric, STS(r), for evaluating neutral-atom quantum optimization. Postprocessed bitstrings from Rydberg annealing on King's-lattice MIS instances are modeled by a one-parameter degeneracy-weighted shell distribution, Eq. (5), with an effective quality parameter β. An excitation-matched random baseline is constructed using Bernoulli(pexc) inputs put through the same postprocessing pipeline, and a positive residual Δβ_ann = β_ann − β_rand is reported in 86 of 96 instances (bins 1–4). From the fitted β, the paper derives two regimes: near-exact targets (r ≈ 1) show an exponential-in-N shot cost that is reduced by annealing, while relaxed targets (r = 0.9) require order-unity shots. The authors provide direct count checks at both targets, exact degeneracy computations, a classical reference cost, and open data/code.","tokens_in":20870,"tokens_out":4886,"duration_ms":50323,"significance":"The paper is valuable for proposing an operational, target-dependent benchmark and for explicitly separating excitation-density effects from genuine structural concentration. It is commendably transparent: direct counts are reported alongside shell-model values, the classical tractability of the graph family is acknowledged, and no end-to-end quantum speedup is claimed. If the shell model were validated at the exact shell, the near-exact exponential shot-cost reduction would be a meaningful physics statement. As it stands, the central near-exact claim relies on a model that the paper itself shows is inaccurate at j = 0, so the significance of the headline result is currently limited by that gap.","major_comments":[{"comment":"The near-exact STS claim rests on a shell model that is not validated at the exact target. The paper reports that the model overestimates π0 by up to a factor of ≈3, giving shell-model STS(r=1) medians of 11/27/63/140 versus direct-count values 14/49/190/459, and that no single β reproduces both the bulk shells and the j=0 population (Appendix D). Because the claimed exponential shot-cost reduction by annealing at r≈1 is computed from Eq. (5) rather than from direct counts, that central claim is an extrapolation from a model known to fail at the exact shell. The paper should either replace this part of the claim with direct-count comparisons between annealing and the matched random baseline, or supply the 1-swap-stable degeneracy counts that the pipeline actually samples.","section":"Sec. IIIE and Appendix D, Eq. (5)"},{"comment":"The validation narrative in the main text is misleading. The text states that 'over 97% of tested instances satisfy D_KL < 0.1', but Figs. 4(e)–(h) display mock data only; Appendix D reports that for the 96 annealing instances the median D_KL is 0.05 and only 78% fall below the same 0.1 threshold. The experimental fit quality should be stated in Sec. IIIE alongside the mock result, and the implications of the 22% failure rate for the shell-model conversion of measured outputs into STS(r) should be discussed.","section":"Sec. IIIE and Fig. 4"},{"comment":"The statement that the shell model is 'expected to be most accurate precisely in the low-j region that controls STS(r) for near-exact targets' (Sec. IIB) is contradicted by the empirical finding in Sec. IIIE that the largest model discrepancy occurs at j=0. The O(j²/N) correction derived in Appendix C does not capture the systematic overcounting of maximum independent sets versus 1-swap-stable sets, which is the dominant error at the exact shell. The text should reconcile the locality-based expectation with the observed j=0 failure and state clearly that Eq. (5) is reliable only for j≥1.","section":"Sec. IIB and Appendix C"}],"minor_comments":[{"comment":"The direct-count median at the exact target for the small group is reported as 146.5 per 500 shots; this non-integer median should be defined more precisely, e.g., as the average of the 24th and 25th order statistics.","section":"Sec. IIIE"},{"comment":"The maximum-entropy motivation for Eq. (5) is described as 'maximizing the entropy of π relative to the degeneracy measured', which is slightly imprecise; the quantity being maximized is the negative relative entropy S[π∥d] with the sign convention shown in Eq. (6).","section":"Sec. IIB"},{"comment":"In panel (b), the filled diamonds and open diamonds are distinguished only by color in the legend; adding different marker shapes or a clearer caption would improve accessibility.","section":"Fig. 5"},{"comment":"The statement that a fixed ratio r=0.9 'admits a one-atom deficit in the small group but four in xlarge' is an important caveat and should be repeated in the caption of Fig. 5(b) to avoid over-generalizing the constant-cost regime.","section":"Sec. IVC"},{"comment":"The sentence 'the paired points are slightly offset horizontally for visibility' in the Fig. 8 caption should specify which points are paired and offset, as it is not obvious from the figure.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually transparent about its own limitations, and the authors' willingness to report direct counts alongside model fits is a strength. However, the headline near-exact exponential shot-cost reduction is not directly measured; it is derived from a shell model that fails at the exact shell. I would ask for either direct-count quantum-versus-baseline STS at r=1 or stable-set degeneracy counts before publication. The discrepancy between the mock-based 97% D_KL pass rate and the experimental 78% should also be fixed in the main text. These are substantive but fixable issues, so major revision is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read of arXiv:2608.12858. The paper is worth a referee's time. It defines STS(r), the shots needed to reach approximation ratio r, and builds an excitation-matched Bernoulli baseline, so postprocessed Rydberg outputs are compared against a density-controlled classical null. The data, code, and exact-solver timing are all openly available, and the authors state their own limitations more candidly than most. That last point matters; this is not a hype paper.\n\nThe clean, robust result is the two-regime structure. At relaxed targets r=0.9, direct counting, with no shell model, gives STS about 2 shots in every size group, for both anneal and baseline. At the exact target, direct counts also grow exponentially with size, medians 14/49/190/459 shots for small to xlarge. So exact is hard and relaxed is easy on this graph family. That part holds.\n\nThe soft spot is the paper's central quantitative claim about annealing: that the fitted beta advantage, Delta_beta about 0.23-0.40, produces an exponential reduction in shot count at r=1. That reduction comes from inserting the fitted beta into the shell model, Eq. (5), and the model is least reliable exactly at the shell that controls r=1. Section IIIE reports that the model overestimates pi_0 by up to a factor of about 3, making the shell-model STS(r=1) a lower bound. Comparing direct counts to model values, 14/49/190/459 versus 11/27/63/140, the gap grows with N. So the exponential reduction relative to the random baseline is an extrapolation, not a measurement. The xlarge direct count rests on 5 hits in 500 shots, so its uncertainty is large. The authors flag all of this themselves, which is to their credit, but it means the headline near-exact advantage should be read as a model prediction.\n\nThe shell model is well supported for the bulk shells, with median D_KL 0.05 on annealing instances, but Appendix D shows that no single beta reproduces both the bulk and j=0. That points to the degeneracy d_alpha-j needing to be replaced by counts of 1-swap-stable independent sets, which the authors themselves suggest as a refinement. I would want that correction before trusting the near-exact numbers.\n\nMy recommendation: send it to peer review. The methodology, the metric, and the honest classical baseline are genuine contributions, and the two-regime conclusion is directly supported. Ask the authors to report direct-count STS curves for both anneal and baseline at the exact target, and to implement the stable-set degeneracy correction. I would cite this for the metric and the baseline even while remaining skeptical of the extrapolated near-exact speedup.","headline":"A serious, unusually honest benchmarking paper whose two-regime conclusion survives direct counting, but whose headline near-exact exponential advantage is a shell-model extrapolation, not a measurement.","tokens_in":21474,"tokens_out":3644,"would_cite":true,"duration_ms":35223,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that postprocessed Rydberg annealing outputs follow a one-parameter shell distribution and that annealing beats an excitation-matched random baseline, yielding exponential shot-cost reduction for near-exact targets…","keywords":["Rydberg atom arrays","maximum independent set","quantum annealing","shots-to-approximate-solution","shell distribution","approximation ratio","excitation-matched baseline","large deviations"],"falsifier":"Run N≈120–125 instances with enough annealing shots, say 5000 or more, that the direct count of exact-optimum hits is statistically solid, and compare the directly counted STS(r=1) with the shell-model value; if the direct STS remains systematically about three times the shell-model estimate, the near-exact exponential shot-cost reduction would not be directly established. A second check is to extend exact degeneracy counts and measurements to N≈260 and see whether the rate function I(δ_c) develops the predicted non-analytic kink at δ_c=δ^⋆, whose absence would undercut the large-deviation derivation of the two regimes.","tokens_in":2077,"feed_emoji":"⚛️","tokens_out":3135,"duration_ms":110611,"temperature":0.7,"pith_summary":"This paper tries to establish an operational measure of how many experimental shots a Rydberg-atom optimizer needs to reach a solution of a given approximation quality, and to determine whether quantum annealing genuinely concentrates probability near optimal solutions. The authors model postprocessed outputs as a degeneracy-weighted exponential shell distribution with one effective quality parameter β, and compare annealing against a random baseline matched to the same excitation density. On King's-lattice maximum-independent-set instances up to 125 sites, the residual advantage Δβ_ann is positive in 86 of 96 instances (0.23–0.40), which they interpret as genuine concentration beyond density effects. The payoff is a two-regime answer: exact targets need exponentially many shots in system size, with annealing reducing that cost at the same exponential level within the model, while relaxed targets need only order-unity shots where classical greedy already succeeds.","feed_headline":"Quantum annealing beats random chance on near-optimal sets","feed_subtitle":"Annealing outperforms an excitation-matched random baseline, but only near-exact targets gain an exponential shot-cost cut.","key_machinery":"The load-bearing object is the STS metric, STS(r;N) = ⌈ln(1−p_req)/ln(1−p_r)⌉, defined via the per-shot probability p_r of landing in shells j ≤ J(r)=α−⌈rα⌉. The distribution of postprocessed outputs is summarized by the degeneracy-weighted shell model π_j = d_{α−j} $e^{{−βj}}$ / Σ_u d_{α−u} $e^{{−βu}}$, where d_{α−j} are exact near-optimal independent-set counts from a transfer-matrix dynamic program and β is the effective quality parameter. The shell form is motivated by maximum entropy and derived from algorithmic locality: for product-measure inputs and local postprocessing the same exponential form appears with corrections O(j²/N), so β_rand(p_exc) is the quality attainable by any spatially uncorrelated input at the measured excitation density. The combination of the fitted β_ann, the matched baseline β_rand, and the residual Δβ_ann = β_ann − β_rand is what turns raw shot data into the two-regime shot-cost scaling.","core_discovery":"The central claim is that on site-diluted King's-lattice MIS instances with up to 125 sites, the postprocessed outputs of Rydberg quantum annealing are well described by a degeneracy-weighted exponential shell distribution, π_j ∝ d_{α−j} $e^{{−βj}}$, with a single effective parameter β, and that the fitted annealing parameter exceeds the excitation-matched random baseline by Δβ_ann = 0.23–0.40 (positive in 86 of 96 instances). Because the same postprocessing maps any product-measure input to the same shell form, this residual certifies correlations in the quantum output, beyond one-point statistics, that suppress postprocessing-irreparable defects. Translating β into a shots-to-approximate-solution cost yields two regimes: for near-exact targets (r≈1) the shot count grows exponentially with N with an intensive rate I≈0.039, and annealing reduces the prefactor at the same exponential level within the shell model; for relaxed targets (r=0.9) the shot cost is order unity, and the classical randomized greedy baseline alone reaches the target in one or two passes. The paper explicitly frames the shot counts as a diagnostic of probability concentration, not as evidence of an end-to-end quantum speedup, and reports exact classical solves in 0.6–97 ms per instance.","pith_inferences":["Editorial inference: because the shell model overestimates the exact-hit probability by up to a factor of about three, the near-exact shot-cost reduction should be read as a model-based lower bound, and direct counting with several thousand shots per instance at N≈120–125 would test whether the hardware's j=0 concentration matches the fitted β.","Editorial connection: the locality derivation implies the same degeneracy-weighted exponential form for any product-measure input and any local postprocessor, so the two-regime STS structure is likely a generic property of local postprocessing plus sharp concentration; testing it on other unit-disk graph families would separate what is special to Rydberg annealing from what is generic.","Editorial testable extension: replacing the independent-set degeneracy d_{α−j} in Eq. (5) with the count of 1-swap-stable independent sets actually returned by the pipeline should remove most of the j=0 overestimate; this is computable by extending the transfer-matrix dynamic program to track pipeline stability, and would give an exact-target shot-cost prediction not requiring the fitted β."],"forward_implications":["For near-exact targets the required number of shots grows exponentially with system size, at a measured intensive rate I≈0.039 per site, and annealing lowers the prefactor within the shell model without removing the exponential growth.","For relaxed targets such as r=0.9, the shot cost saturates near order unity over the whole size range, and the excitation-matched random baseline alone reaches the target in one to two greedy passes, so this target has little discriminatory power.","A positive excitation-matched advantage Δβ_ann = 0.23–0.40 (86 of 96 instances) indicates that Rydberg annealing concentrates probability toward the MIS manifold beyond what the raw excitation density alone explains.","The classical reference is cheap: exact transfer-matrix solves run in 0.6–97 ms per instance and are subexponential, 2^Θ(√N), so the paper's shot counts are a diagnostic of concentration, not an end-to-end speedup claim.","The rate function at the exact target is size-independent to within 8%, implying NI grows linearly with N and the sharp two-regime kink should become resolvable near N≈260 with larger arrays."],"supporting_citations":[{"why":"Supplies the Rydberg MIS benchmark setup, the King's-lattice instances, the hardness proxy H(G), and the vertex-reduction/addition postprocessing lineage.","marker":"[14]"},{"why":"Introduces per-sample approximation-ratio evaluation of QAOA, which motivates the target-dependent shot metric.","marker":"[22]"},{"why":"Shows approximation-ratio-dependent sampling cost in quantum optimization, the gate-based analogue of the two-regime result.","marker":"[23]"},{"why":"Provides the swap-based local-improvement postprocessing idea used in the pipeline.","marker":"[24]"},{"why":"Gives the maximum-entropy derivation of the degeneracy-weighted exponential shell form.","marker":"[25]"},{"why":"Identifies the hardware platform on which all annealing data were taken.","marker":"[32]"},{"why":"Is the open dataset and code that define the instances and reproduce the figures.","marker":"[33]"},{"why":"Supplies the Laplace/large-deviation principle that converts the shell model into the two-regime shot-cost scaling.","marker":"[34]"},{"why":"Supplies the ETH-tight subexponential lower bound that the classical transfer-matrix solver matches.","marker":"[35]"}],"fun_headline_variants":["Quantum annealing cuts shot cost only for near-exact solutions","Neutral-atom quantum beats random, but only at high approximation ratios","Quantum optimization's edge fades for relaxed targets","Shots-to-solution: quantum helps only when targets are tight","Quantum annealing's concentration edge limited to near-optimal"],"cache_read_input_tokens":23552,"weakest_assumption_plain":"The load-bearing premise is that one fitted value of β describes the whole postprocessed shell distribution, in particular the probability of hitting an exact optimum, so that shot counts near r=1 can be extrapolated from the model; the paper's own direct counts show this model overestimates the exact-hit probability by up to roughly a factor of three.","fun_headline_variants_meta":{"raw":{"variants":["Quantum annealing cuts shot cost only for near-exact solutions","Neutral-atom quantum beats random, but only at high approximation ratios","Quantum optimization's edge fades for relaxed targets","Shots-to-solution: quantum helps only when targets are tight","Quantum annealing's concentration edge limited to near-optimal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000809,"raw_usage":{"total_tokens":3613,"prompt_tokens":1068,"completion_tokens":2545,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":2476}},"tokens_in":684,"tokens_out":2545,"duration_ms":18141,"temperature":1.0,"reasoning_tokens":2476,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:48:44.259580+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run N≈120–125 instances with enough annealing shots, say 5000 or more, that the direct count of exact-optimum hits is statistically solid, and compare the directly counted STS(r=1) with the shell-model value; if the direct STS remains systematically about three times the shell-model estimate, the near-exact exponential shot-cost reduction would not be directly established. A second check is to extend exact degeneracy counts and measurements to N≈260 and see whether the rate function I(δ_c) develops the predicted non-analytic kink at δ_c=δ^⋆, whose absence would undercut the large-deviation derivation of the two regimes.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Rydberg MIS benchmark setup, the King's-lattice instances, the hardness proxy H(G), and the vertex-reduction/addition postprocessing lineage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows approximation-ratio-dependent sampling cost in quantum optimization, the gate-based analogue of the two-regime result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the swap-based local-improvement postprocessing idea used in the pipeline."},{"cited_title":"Benedetti, J","cited_arxiv_id":null,"evidence_quote":"Identifies the hardware platform on which all annealing data were taken."},{"cited_title":"Marshall, E","cited_arxiv_id":null,"evidence_quote":"Is the open dataset and code that define the instances and reproduce the figures."},{"cited_title":"Vuffray, C","cited_arxiv_id":null,"evidence_quote":"Supplies the Laplace/large-deviation principle that converts the shell model into the two-regime shot-cost scaling."},{"cited_title":"Weidemüller, Suppression of excitation and spec- tral broadening induced by interactions in a cold gas of Rydberg atoms, J","cited_arxiv_id":null,"evidence_quote":"Supplies the ETH-tight subexponential lower bound that the classical transfer-matrix solver matches."}],"review_version":1}