{"id":"b0a5e096-37ce-4fce-b98e-efa236aef1d6","arxiv_id":"2411.16656","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Optimal Bayesian-designed annealing parameters transfer across similar maximum independent set graph instances, enabling a fixed protocol to be applied from small training graphs to unseen 100-qubit graphs on a Rydberg atom processor.","lead":"Researchers tuned a quantum control program on small test problems and reused it unchanged on larger problems, up to 100 atoms on a real quantum processor. The same program found good answers to an electric vehicle charging schedule problem, though very large runs needed classical cleanup.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"At N=100, the experimental claim is not yet isolated from classical post-processing: raw P(MIS)≈0.05%, and no baseline shows that the App. H repair step would fail on trivial or random input distributions.","rationale":"The core transferability claim is partially supported: the noiseless emulation on 33 unseen smart-charging graphs (N=9-23) with average MIS probability 98.3(4)% is a genuine out-of-sample test, and the training protocol on 50 graphs with m=3 parameters is a reasonable construction. I therefore do not dispute transferability within the stated triangular unit-disk graph family. However, the strongest scale claim in the abstract—experimental validation up to 100 qubits—rests on Fig. 5(b), where raw P(MIS) is 0.05% at N=100, and on the post-processed curves of Fig. 5(c) and Table I. The post-processing step in App. H is powerful: it repairs non-IS bitstrings by deleting on average about 5 nodes and then performing depth-2 greedy completion. The paper provides no control experiment showing that this repair would fail to find a MIS on distributions not produced by the trained protocol, such as random bitstrings with matched marginals or shots from a trivial ramp. If the post-processor is that effective, the N=100 experimental data validates classical repair plus any blockade-like distribution, not specifically the transferable annealing schedule. This is a concrete, falsifiable gap. The reader's verdict already conditions on lack of code/data and overstatement of raw hardware results; this concern is a sharper form of that overstatement, so it does not move the verdict, but it should be an explicit condition: provide control baselines for the post-processor and MPS convergence checks for the N>25 emulation. The geometry limitation is real but explicitly scoped and honestly stated, so it is less damaging in my view. The issue is missing controls, not misrepresentation.","tokens_in":25164,"tokens_out":9441,"duration_ms":93910,"concrete_test":"Apply the App. H post-processor to 1000 shots generated by a trivial control (e.g., an instantaneous ramp to the final detuning of the trained schedule) on the same ten N=100 graphs, and separately to 1000 random bitstrings matched to the measured marginal excitation rate. Count the fraction of graphs for which at least one MIS is recovered. If either baseline matches the post-processed experimental recovery rate, the 100-atom validation does not isolate the transferable protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV C reports that on the Orion Alpha device the transferable protocol yields raw P(MIS)=0.05% at N=100, i.e. about one expected MIS in the ~1000-shot sample. The conclusion that 'we can successfully sample MIS configurations up to N=100' is therefore entirely mediated by the classical post-processing of App. H, which removes an average of 4.75 constraint-violating nodes and greedily adds up to 2 nodes. The paper does not provide any control baseline for this post-processor: it is not applied to (i) random bitstrings with the measured marginal excitation density, (ii) shots from a trivial/constant or untrained schedule, or (iii) the output of a purely classical independent-set sampler. Without such a baseline, the N=100 hardware data cannot distinguish a genuinely transferable quantum schedule from a classical repair algorithm that would retrieve a MIS from almost any input distribution that already respects Rydberg blockade correlations. The abstract's statement that experimental results 'validate the effectiveness of our approach, scaling to problems with up to 100 qubits' is thus stronger than the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a transferable variational quantum annealing protocol for maximum independent set (MIS) problems on unit-disk graphs. Using Bayesian optimization, a single VQAA schedule is trained on 50 small triangular-lattice graphs and then applied, without further optimization, to unseen smart-charging instances (9–23 nodes) and to larger triangular graphs up to N=100. Noiseless emulation yields 98.3(4)% average MIS probability on the industrial test set and about 37% at N=100; hardware experiments yield 29% on the small instances and 0.05% raw MIS probability at N=100, with MIS configurations recovered after classical post-processing. The paper also fits exponential decay laws for MIS-k probabilities and uses them to extrapolate sampling requirements at larger sizes.","tokens_in":25447,"tokens_out":5303,"duration_ms":53399,"significance":"If substantiated, this is a useful demonstration that a single annealing schedule can transfer across instances sharing a common geometry, reducing the variational optimization overhead that currently limits practical quantum optimization. The out-of-sample design — training on small graphs and testing on unseen, larger graphs — avoids the most common circularity failure, and the paper deserves credit for benchmarking emulated noiseless, noisy-emulated, and experimental results, for explicitly modeling detection errors and decoherence, and for building on the open-source Pulser framework. The main caveats are that the headline N=100 hardware claim currently rests on classical post-processing without control baselines, and that the scaling extrapolation is fitted to the same data it is used to predict.","major_comments":[{"comment":"The N=100 experimental claim is not isolated from classical post-processing. Raw experimental P(MIS) is reported as 0.05% at N=100, i.e. about one expected MIS in the ~1000-shot sample, and the statement that the authors 'are able to experimentally find the MIS at least once for each graph' refers to distributions after the App. H repair procedure, which removes an average of 4.75 constraint-violating nodes and greedily adds up to 2 nodes. No control baseline is provided for this post-processor: it is not applied to random bitstrings with the measured marginal excitation density, to shots from a trivial or untrained schedule, or to a purely classical independent-set sampler. Without such a baseline, the N=100 data cannot distinguish a genuinely transferable quantum schedule from a classical repair algorithm that would retrieve a MIS from almost any input distribution. The abstract's statement that experimental results 'validate the effectiveness of our approach, scaling to problems with up to 100 qubits' is therefore stronger than the evidence presented and should either be qualified to the post-processed hybrid pipeline or supported by the missing controls.","section":"Sec. IV C and App. H"},{"comment":"The exponential decay model in Eq. (3) is fitted to the same emulated and experimental data that it is then used to extrapolate, for example in the N=500 sampling estimates in Sec. IV B. The piecewise form has free thresholds b_k and decay constants N_k, no goodness-of-fit or uncertainty is reported, and the experimental fits use only five sizes (N=30, 50, 70, 80, 100). The text should present Eq. (3) as an empirical interpolation over the measured range rather than as a predictive scaling law, and any extrapolation should be accompanied by error bars and a clear statement of model risk.","section":"Sec. IV B, Eq. (3), Table I"}],"minor_comments":[{"comment":"The time-to-solution comparison is stated inconsistently: Sec. IV A says the VQAA time to solution is 'around three orders of magnitude higher than with CPLEX', while Sec. IV C says it is 'sill three orders of magnitude below the state-of-the-art CPLEX method'; these cannot both be correct, and the intended comparison (including the typo 'sill') should be fixed.","section":"Sec. IV A and IV C"},{"comment":"The threshold parameters b_k are listed only for the experimental and post-processed fits, not for the emulated fits; please clarify whether the emulated fits use fixed thresholds, b_k=0, or separately fitted values.","section":"Eq. (3) and Table I"},{"comment":"The text states that QAOA-like local minima 'do not approach 0 in value' and that VQAA landscapes reach 0.04–0.06; a sentence explaining why the averaged normalized approximation ratio saturates above zero would help the reader interpret the concentration plots.","section":"Sec. II B"},{"comment":"The paper acknowledges that cliques larger than 3 cannot be encoded on the triangular layout and that the industrial test set is restricted to 2D unit-disk instances; this limitation should be reflected more explicitly in the abstract's claim that the method applies to 'real-world scenarios'.","section":"Sec. III B and Conclusion"},{"comment":"The experimental points in Figs. 5(b) and 5(c) are shown without error bars; since each size appears to be represented by a single graph with ~1000 shots, the sampling uncertainty should be displayed or at least stated in the caption.","section":"Fig. 5"}],"recommendation":"major_revision","confidential_remarks":"The core transferability result is methodologically sound and out-of-sample, so I would not reject the paper. However, the N=100 hardware scaling claim needs either additional control experiments for the post-processor or a substantially softened statement, and the scaling extrapolation in Sec. IV B should be reframed as empirical interpolation. If the authors add the missing baselines and correct the overclaim in the abstract, the paper would be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the out-of-sample transfer result on small graphs is real, and the smart-charging test set is a concrete, honest demonstration. The N=100 hardware claim, however, is not yet supported by the evidence as presented — it rests entirely on a classical post-processing step with no baseline, and the abstract overstates what the hardware shows.\n\nWhat's new: the paper takes the known QAOA parameter-transfer idea and shows it works for continuous annealing schedules (VQAA). A single schedule, trained by Bayesian optimization on 50 small triangular graphs, gives 98.3(4)% MIS probability on unseen 9–23 node smart-charging instances in noiseless emulation, and around 29% raw on the device. That is a genuine extension, and the phase-diagram analysis showing the optimizer finds a path through the MIS phase while avoiding small gaps is sensible. The authors are also upfront about the geometry restriction: only graphs embeddable as subgraphs of the triangular lattice, no cliques larger than 3.\n\nThe soft spots. First, the N=100 result. Raw P(MIS) at that size is 0.05%, and the conclusion that the protocol 'successfully samples MIS configurations up to N=100' is entirely mediated by the App. H repair step, which deletes an average of 4.75 violating nodes and greedily adds up to 2. The paper gives no control: we don't see what that post-processor does on random bitstrings, on shots from a trivial schedule, or on a purely classical independent-set sampler. Without that baseline, the hardware data can't distinguish a transferable quantum schedule from a classical repair that would fix almost any input that respects blockade correlations. Second, the scaling extrapolation in Sec. IV B is an exponential fit to the same emulated/experimental data, so it's a summary, not an independent prediction. Third, no code or data are shipped, though the protocol is described well enough to be reimplemented.\n\nI think the central small-graph transfer result holds. The large-N claim needs a proper baseline before it can be taken at face value, and the abstract should be toned down. The paper deserves a serious referee — this is the kind of empirical work that's useful to the neutral-atom and hybrid quantum-classical community. It needs revision, not rejection. I'd cite the transferable schedule result.","headline":"The small-graph transfer result is real and worth citing, but the N=100 hardware claim needs a classical post-processing baseline before it can be taken at face value.","tokens_in":25924,"tokens_out":1937,"would_cite":true,"duration_ms":17748,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single annealing schedule, trained on 50 small triangular graphs, transfers to larger unseen graphs of the same geometry and solves a real smart-charging problem.","keywords":["transferable annealing protocols","variational quantum annealing","Bayesian optimization","maximum independent set","Rydberg atom arrays","unit disk graphs","smart charging","quantum optimization"],"falsifier":"Run the trained protocol, with no re-optimization, on a fresh set of triangular-embeddable graphs of sizes between 10 and 100 and compare the noiseless-emulated average MIS probability to the paper's fitted exponential decay (roughly 37% at size 100). If the success probability at intermediate sizes drops far faster than that decay, or if on hardware after detection-error correction no maximum independent set is found at size 100, the size-transfer claim fails.","tokens_in":1,"feed_emoji":"⚡","tokens_out":13101,"duration_ms":187285,"temperature":0.7,"pith_summary":"This paper tries to establish that one quantum annealing schedule can be trained once on small problems and then reused on larger, unseen problems that share the same geometry, removing the per-instance parameter tuning that makes variational quantum optimization expensive. The authors optimize the schedule with Bayesian methods on 50 small maximum-independent-set instances drawn from a triangular atomic layout, then apply the same schedule to smart-charging instances and to larger triangular graphs. In noiseless emulation the transferred schedule prepares maximum independent sets with an average probability of 98.3(4)% on unseen 9- to 23-node smart-charging instances, and on neutral-atom hardware it retrieves maximum independent sets up to 100 atoms once measured bitstrings are repaired by a classical post-processing step. If these results hold, variational annealing becomes a one-time training cost per graph geometry rather than a per-instance optimization cost.","feed_headline":"One pulse trained on 50 tiny graphs solves 100-atom problems","feed_subtitle":"It transfers to larger unseen instances and tackles a real EV smart-charging problem on a Rydberg device.","key_machinery":"The load-bearing object is a VQAA schedule: a time-dependent annealing drive parameterized by the Rabi frequency and detuning at a small number of fixed time points plus the total duration, with monotonic cubic spline interpolation between points. The drive is trained by Bayesian optimization on a cost that averages the MIS-preparation error over a family of graphs and penalizes spread, which pushes the optimizer toward schedules that work uniformly across the family. The geometric encoding is supplied by the Rydberg blockade effect, in which atoms closer than a blockade radius cannot both be excited, so a unit-disk graph can be embedded as an atom arrangement on a triangular lattice. The mechanism that makes transfer possible is parameter concentration: for VQAA-like drives the individual cost landscapes of graphs in a family overlap enough that one path through the phase diagram, ending in the maximum-independent-set phase and avoiding small-gap regions, prepares MIS configurations for all graphs in the family.","core_discovery":"The central claim is that optimal variational annealing parameters concentrate for graph families with a shared lattice geometry, so a single control protocol generalizes within the family and from small to large graphs. Specifically, a variational quantum annealing protocol with two control fields (Rabi frequency and detuning) sampled at fixed times and interpolated by monotonic splines is trained by Bayesian optimization on 50 triangular-lattice graphs of 5 to 9 nodes, with the cost defined as the average MIS-preparation error over the training set plus its standard deviation. The optimized schedule starts in the independent-set phase and ends inside the reduced maximum-independent-set phase of the phase diagram for the two control parameters, steering clear of regions with vanishing spectral gaps. On the training set it reaches a MIS probability of about 99.7%; on unseen smart-charging graphs of 9 to 23 nodes it reaches 98.3(4)% in noiseless emulation; and on hardware, after classical repair of constraint-violating bitstrings, it finds a maximum independent set at least once for every graph tested up to 100 atoms. The paper also reports that at large sizes the probability of sampling a maximum independent set decays roughly exponentially with graph size, while near-optimal solutions remain substantially more likely.","pith_inferences":["In our reading, the same training recipe should transfer to other regular layouts: training on square-lattice or Shastry-Sutherland graphs should produce a family-specific schedule, with the optimal ending detuning shifting as the geometry changes, as the paper's landscape comparison already hints.","The use of the standard deviation of the cost over the training family as a regularizer is a simple and possibly general idea; a testable extension would be to apply the same cost construction to QAOA parameter sets and see whether it tightens the weaker concentration the paper observes there.","If the graph-embedding gadgets mentioned in the outlook are placed on a regular lattice, the transferable-schedule approach could be applied to denser, non-unit-disk industrial constraints, making the method a candidate warm-start provider for classical solvers rather than only a standalone sampler.","The fitted exponential decay is itself a quantitative prediction: running the same protocol at intermediate sizes such as 40, 60, or 90 nodes and comparing the observed cumulative MIS probabilities to the fitted curves would test whether the decay law holds and how much of it is hardware-limited."],"forward_implications":["Once a schedule is trained for a lattice geometry, new instances of that geometry can be solved with a single fixed pulse sequence, skipping the closed-loop optimization that would otherwise cost hundreds or thousands of shots per instance.","For smart-charging instances that fit a triangular layout, the end-to-end pipeline turns a maximum-independent-set problem into a small number of experimental shots: the first MIS is typically found within a few shots, and with classical post-processing all MISs are found at least once for graphs up to 100 atoms.","At sizes beyond the training range, the protocol still acts as a useful low-energy sampler: even where the MIS probability has decayed to about 37% on average at size 100 in noiseless emulation, MIS or MIS-1 configurations are sampled frequently enough that roughly 14 shots suffice for a near-optimal solution at size 500 under the fitted decay.","Hardware noise, especially detection errors, degrades the raw distributions, but the classical repair step (removing conflicting nodes and greedily adding nodes up to depth 2) restores maximum independent sets, making the hybrid quantum-classical pipeline practical at current coherence times."],"supporting_citations":[{"why":"supplies the transferability concept for optimal QAOA parameters that this paper extends to annealing schedules.","marker":"[15]"},{"why":"supplies the Bayesian optimization method for designing quantum annealing schedules used to train the protocol.","marker":"[19]"},{"why":"introduces the Rydberg-atom encoding of maximum independent set problems that the experimental mapping relies on.","marker":"[26]"},{"why":"provides the smart-charging industrial use case whose instances form the test set.","marker":"[44]"},{"why":"provides the tensor-network emulation method used to evaluate the protocol on larger graphs.","marker":"[51]"},{"why":"introduces the classical repair of independent-set bitstrings used to post-process experimental data.","marker":"[63]"},{"why":"supports the claim that optimal variational parameters concentrate for graph families with shared local structure.","marker":"[17]"},{"why":"defines the adiabatic variational annealing framework that the VQAA parametrization generalises.","marker":"[35]"}],"fun_headline_variants":["Single annealing pulse from 50 tiny graphs solves 100-atom problems","Transferable quantum annealing schedule handles real EV smart-charging","Quantum pulse from 5-node graphs generalizes to 100-atom EV grids","One annealing protocol trained on small graphs works on 100-qubit problems","Small-graph training yields transferable annealing for 100-atom optimization"],"cache_read_input_tokens":28160,"weakest_assumption_plain":"The method rests on every target problem being drawable as a graph on the same triangular grid used in training, with edges only between nearby grid points, and on the physical atom array reproducing those connections faithfully.","fun_headline_variants_meta":{"raw":{"variants":["Single annealing pulse from 50 tiny graphs solves 100-atom problems","Transferable quantum annealing schedule handles real EV smart-charging","Quantum pulse from 5-node graphs generalizes to 100-atom EV grids","One annealing protocol trained on small graphs works on 100-qubit problems","Small-graph training yields transferable annealing for 100-atom optimization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001358,"raw_usage":{"total_tokens":5520,"prompt_tokens":965,"completion_tokens":4555,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":4459}},"tokens_in":581,"tokens_out":4555,"duration_ms":28625,"temperature":1.0,"reasoning_tokens":4459,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:54:49.008147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained protocol, with no re-optimization, on a fresh set of triangular-embeddable graphs of sizes between 10 and 100 and compare the noiseless-emulated average MIS probability to the paper's fitted exponential decay (roughly 37% at size 100). If the success probability at intermediate sizes drops far faster than that decay, or if on hardware after detection-error correction no maximum independent set is found at size 100, the size-transfer claim fails.","supporting_citations":[{"cited_title":"Qualifying quantum approaches for hard industrial optimization problems","cited_arxiv_id":null,"evidence_quote":"provides the smart-charging industrial use case whose instances form the test set."},{"cited_title":"Cloud on-demand emulation of quantum dynamics with tensor networks, 2023","cited_arxiv_id":null,"evidence_quote":"provides the tensor-network emulation method used to evaluate the protocol on larger graphs."},{"cited_title":"Quantum optimization of maximum inde- pendent set using rydberg atom arrays","cited_arxiv_id":null,"evidence_quote":"introduces the classical repair of independent-set bitstrings used to post-process experimental data."},{"cited_title":"Transferability of optimal QAOA parameters between random graphs","cited_arxiv_id":"2106.07531","evidence_quote":"supports the claim that optimal variational parameters concentrate for graph families with shared local structure."}],"review_version":1}