{"id":"e97e14bf-da70-4dff-a5ae-2be1dd0d7ca4","arxiv_id":"2607.14282","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A quantum-annealing schedule built from worldline susceptibility fluctuations measured during simulated annealing outperforms linear and often gap-based schedules on small Sherrington–Kirkpatrick instances.","lead":"The authors test a new way to schedule quantum annealing runs by measuring simple fluctuations in a classical simulation instead of computing the hard-to-get quantum energy gap. On small spin-glass problems, this cheap schedule often beats both linear annealing and a gap-based schedule, although the gap-based baseline may be missing a term.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'exact Roland–Cerf' baseline is defined by Eq. (4) with a transition matrix element, but the benchmark/analysis uses Eq. (24)'s gap-only cumulative allocation; the central comparison is not yet established.","rationale":"The reader identified the same load-bearing weakness: the Roland–Cerf baseline is effectively treated via a gap-only cumulative allocation (Eq. 24) that omits the transition matrix element present in Eq. (4). This is the most serious issue because it directly undermines the paper's strongest and most surprising claim—that a susceptibility-based surrogate outperforms the exact local-adiabatic schedule. The two proposed failure modes, boundary-gap trap and oscillatory instability, are then attributed to a schedule that may not be the exact Roland–Cerf schedule. Recomputing with the full matrix element is a concrete, feasible fix that would settle the issue. I am not raising this as a rejection: the surrogate method itself, the SQA-based observable, the open-source implementation, and the 370-instance statistical validation are all valuable and largely independent of the exact-RC comparison. The paper should be accepted conditionally on this re-benchmark or on explicitly relabeling the baseline as a gap-only schedule and adjusting the claims accordingly.","tokens_in":20771,"tokens_out":5268,"duration_ms":57332,"concrete_test":"Recompute the Roland–Cerf benchmark using the full Eq. (4) for the same 370-instance ensemble: for each instance, compute M(s)=|<E1(s)|∂sH|E0(s)>| by exact diagonalization, integrate ds/dt = ε Δ(s)^2/M(s) with ε chosen so that s(0)=0 and s(1)=1 for each total time T, and repeat the PGS calculations of Figs. 3–4 and the failure-class labels of Sec. III.D. In addition, for the representative n=10 and n=12 instances, plot the full cumulative allocation F_full(s)=∫_0^s |M|/Δ^2 ds' / ∫_0^1 |M|/Δ^2 ds' and compare it with Eq. (24). If the non-monotonic PGS curves and boundary-gap suppression persist and the surrogate advantage survives, the concern is resolved; if they weaken or disappear, the central comparison must be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim requires that the schedule being outperformed is the true Roland–Cerf local-adiabatic schedule of Eq. (4): ds/dt = ε Δ(s)^2 / |<E1|∂sH|E0>|. However, Eq. (24) and Fig. 5 define the Roland–Cerf cumulative time allocation as F(s)=∫Δ^{-2}ds'/∫Δ^{-2}ds', which is the local-adiabatic allocation only if the matrix element M(s)=|<E1|∂sH|E0>| is constant. For the transverse-field Ising Hamiltonian H(s)=A(s)H_D+B(s)H_P, ∂sH = A'(s)H_D+B'(s)H_P and M(s) has no general reason to be constant; it can vary substantially across the anneal. If the benchmarked trajectories were generated from Eq. (24) rather than Eq. (4), then the 'exact Roland–Cerf' baseline is actually a gap-only schedule, and both claimed failure modes—the boundary-gap trap and the oscillatory instability—are properties of that approximate baseline, not of exact local-adiabatic scheduling. Appendix A.5 states the RC trajectories were constructed from Eq. (4), making the discrepancy internal: either the full matrix element was not used, or Eq. (24) does not describe the schedule actually evolved. Either way, the central claim that the surrogate surpasses the exact Roland–Cerf schedule is not currently supported by the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a surrogate annealing schedule for quantum annealing constructed from the worldline magnetization susceptibility χ_m(s) measured in simulated quantum annealing. The schedule is defined by ds/dt = 1/[T(χ_m(s)+χ_0)] (Eq. 14), with cumulative mapping Eq. (11). Using exact diagonalization of Sherrington–Kirkpatrick instances (n=10–20), the authors compare the surrogate against linear annealing and the Roland–Cerf local-adiabatic schedule. They report that the surrogate achieves the highest disorder-averaged ground-state probability at all system sizes and finite times studied, and they attribute this to two finite-time failure modes of exact local-adiabatic scheduling: a boundary-gap trap and an oscillatory instability. The manuscript includes an open-source implementation (qanneal), decoherence checks under Lindblad dephasing, a Trotter-parameter robustness study, and a held-out logistic validation of a gap-based failure threshold.","tokens_in":21135,"tokens_out":7235,"duration_ms":73518,"significance":"If the findings hold, this is a practically relevant result: a low-cost equilibrium observable can schedule annealing better than a spectral-gap-optimal schedule at realistic finite times. The paper's strengths include the open-source framework, the use of exact diagonalization only for validation, the systematic disorder ensembles, and the supplementary checks (dephasing, Trotter refinement, held-out classification). The central empirical comparison, however, depends on the correct implementation of the exact Roland–Cerf baseline and on the validity of the proposed failure-mode explanation; both need the revisions below.","major_comments":[{"comment":"The cumulative Roland–Cerf time-allocation F(s) in Eq. (24) is defined as ∫Δ^{-2}ds'/∫Δ^{-2}ds', but the Roland–Cerf velocity in Eq. (4) contains the transition matrix element M(s)=|<E1|∂sH|E0>|. The actual cumulative allocation is F_M(s)=∫M(s')Δ^{-2}ds'/∫M(s')Δ^{-2}ds'. The paper neither shows M(s) is constant nor bounds its variation. Consequently, Fig. 5 and the quantitative statements built on it (e.g., 'half the total annealing time after s≈0.995' for the n=12 instance) are not established for the exact RC schedule. Appendix A.5 states that RC trajectories were generated from Eq. (4); hence either the figure is mislabeled as 'RC cumulative weight' or the benchmark actually used a gap-only schedule, which would invalidate the headline comparison. Please recompute the cumulative allocation with M(s), or state explicitly that F(s) is an approximation and verify the failure-mode conclus","section":"III.C, Eq. (24); Appendix A.5"},{"comment":"The oscillatory instability is attributed to 'coherent multilevel interference' on the basis of (i) a period mismatch with two-level LZS (T_LZS≈312 vs T_obs≈20) and (ii) survival under dephasing up to γ=0.01 (Appendix B.1). Neither diagnostic demonstrates multilevel interference: a period mismatch could arise from other causes, and dephasing robustness shows the effect is coherent but does not identify the subspace involved. Please provide direct evidence, e.g., populations of excited eigenstates during the RC-scheduled evolution, or a comparison with a truncated few-level model. Without this, the abstract's claim to have 'demonstrated' the mechanism is too strong.","section":"III.C.2"},{"comment":"At n≥14 the number of instances is 50, 50, and 20, and the plotted SEM bands appear to overlap substantially in several panels. The claim that the surrogate schedule 'consistently achieves the highest average ground-state probability' across all system sizes is not supported by a formal statistical analysis. Report pairwise differences with confidence intervals or significance tests (e.g., paired tests over instances) for surrogate vs Roland–Cerf and surrogate vs linear at each n and T. This is central to the generalization claim.","section":"III.D, Figs. 7–9"}],"minor_comments":[{"comment":"The text says 'Ground-state probabilities [Eq. (19) of the main text]' but the relevant definition is Eq. (20) (or Eq. 16). Please correct the cross-reference.","section":"Appendix A.6"},{"comment":"The sentence 'Although both schedules satisfy exactly the same local-adiabatic condition, they produce two qualitatively different dynamical pathologies' is confusing because the two panels of Fig. 5 show two different instances, not two different schedules. Rephrase to describe the two instances.","section":"III.C"},{"comment":"The classification of instances as boundary-gap (B), oscillatory (O), or conventional (C) is used for the base rates in Figs. 7–9 and 14, but the criteria are not defined in the main text. State the classification rule (e.g., location of s*_ED and number of direction changes in PGS(T)) so the quoted rates are reproducible.","section":"III.D and Appendix B.5"},{"comment":"The parameter ε in Eq. (4) is not discussed. For fixed total time T, how is ε determined for each instance when generating the RC trajectories? Specify this implementation detail, which is important for reproducibility.","section":"II.A, Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is potentially publishable after the RC-baseline inconsistency is fixed. The authors should either recompute Fig. 5 with the full matrix element or demonstrate that it is nearly constant, and they should provide direct evidence for multilevel interference. The statistical support for the ensemble-wide claim also needs quantification. The open-source framework and supplementary checks are clear strengths that should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: the surrogate scheduling idea is genuinely useful, and the paper has built a serious open-source framework around it. But the headline comparison—surrogate beats the exact Roland–Cerf schedule—rests on a baseline that may not be exact Roland–Cerf. That needs to be resolved before the comparative claims can be trusted.\n\nWhat's new and good: constructing the annealing schedule from the worldline magnetization susceptibility measured during SQA is a practical, inexpensive route to locate the critical region, and the authors test it carefully. They run 370 Sherrington–Kirkpatrick instances across n=10–20, check Trotter-slice and inverse-temperature dependence, add Lindblad dephasing controls, and even do a held-out logistic validation of the gap threshold. The Qanneal package is open source and the reproducibility story is unusually complete. The surrogate consistently beats linear annealing, which is a solid, useful result on its own.\n\nThe soft spot is the Roland–Cerf benchmark. The local-adiabatic condition in Eq. (4) involves the transition matrix element M(s)=|<E1|∂sH|E0>|, but the cumulative time-allocation function in Eq. (24) and Figure 5 uses only Δ(s)^{-2}. Unless M(s) is exactly constant along the anneal—nothing in the Hamiltonian suggests it should be—the baseline being outperformed is a gap-only schedule, not the exact Roland–Cerf schedule. Appendix A.5 says the RC trajectories were built from Eq. (4), so either the implementation dropped the matrix element or the analysis does not describe the actual schedule. Either way, the claimed failure modes (boundary-gap trap, oscillatory instability) are not yet demonstrated for true local-adiabatic evolution. The oscillatory mechanism is also asserted more than shown: the LZS period mismatch rules out two-level interference, but no direct evidence of multilevel populations is presented.\n\nThis is fixable. Re-run the comparison with the full Eq. (4), or honestly relabel the baseline as a gap-only schedule and soften the conclusions. The surrogate-vs-linear advantage stands; the interesting questions about when exact spectral information hurts can then be asked properly.\n\nWho should read it: anyone building annealing schedules for quantum annealers or SQA. It deserves a serious referee, but with a request for major revision on the baseline issue. I'd want that resolved before citing the comparative claims.","headline":"A practical surrogate schedule with good engineering, but the 'beats exact Roland–Cerf' claim is undermined by a gap-only baseline that omits the transition matrix element.","tokens_in":21599,"tokens_out":5045,"would_cite":true,"duration_ms":51915,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81P68","82B20","82B80"],"pacs":["03.67.-a","05.10.Ln","75.10.Nr"],"model":"deepseek-v4-flash","headline":"A worldline-susceptibility schedule, built from classical Monte Carlo during simulated quantum annealing, beats both linear and exact local-adiabatic annealing at finite time.","keywords":["quantum annealing","simulated quantum annealing","adiabatic schedule","Roland-Cerf schedule","worldline susceptibility","Sherrington-Kirkpatrick","Monte Carlo","finite-time dynamics"],"falsifier":"Recompute the Roland-Cerf schedule on the same instances including the full matrix element |<E1|partial_s H|E0>| in the velocity integral; if the surrogate no longer beats it on a substantial fraction of instances, the central claim fails. A cheaper check: evaluate the matrix element for the n=10 and n=12 instances and test whether the boundary-gap trap and oscillations persist.","tokens_in":20585,"feed_emoji":"⚛️","tokens_out":3771,"duration_ms":38200,"temperature":0.7,"pith_summary":"This paper tries to establish that the hard part of a quantum anneal can be found without computing the quantum spectrum: the equilibrium fluctuations of the worldline magnetization in simulated quantum annealing, chi_m(s), peak near the minimum spectral gap. Feeding chi_m into the annealing velocity, ds/dt = 1/[T(chi_m+chi_0)], yields a smooth schedule that outperforms linear annealing on Sherrington-Kirkpatrick spin glasses and, for a substantial fraction of instances, also beats the exact Roland-Cerf local-adiabatic schedule. The paper explains this counterintuitive result by identifying two finite-time failure modes of gap-based local-adiabatic scheduling: a boundary-gap trap, where the minimum gap sits at the end of the anneal so time is wasted where the transverse field is dead, and an oscillatory instability from over-localizing time around an interior gap minimum. A reader should care because it suggests that cheap, observable-driven scheduling can be more robust than exact spectral strategies under realistic finite-time conditions, and it scales beyond exact-diagonalization reach.","feed_headline":"Measured susceptibility outperforms exact-gap annealing schedules","feed_subtitle":"A spin-fluctuation signal from simulated annealing locates the critical region, avoiding two finite-time traps that sink the Roland-Cerf sch","key_machinery":"The central object is the worldline magnetization susceptibility chi_m(s) = NM(<m^2>-<|m|>^2), computed from equilibrium Suzuki-Trotter Monte Carlo samples of the transverse-field Ising model; near criticality it scales as 1/Delta(s)^2, making it a measurable surrogate for the inverse-square gap that drives the Roland-Cerf condition. The schedule is ds/dt = 1/[T(chi_m(s)+chi_0)], which allocates time by the susceptibility's cumulative weight.","core_discovery":"Using exact diagonalization as ground truth for Sherrington-Kirkpatrick instances (n=10-20), the authors show that a schedule constructed from the worldline magnetization susceptibility measured during SQA consistently gives the highest disorder-averaged ground-state probability, beating linear ramps and, for a large fraction of instances, the exact Roland-Cerf schedule. They attribute this to two mechanisms: the boundary-gap trap, in which the true local-adiabatic schedule concentrates runtime at s=1 where the transverse field has vanished, and an oscillatory instability, in which extreme runtime localization around an interior gap produces multilevel coherent interference. The susceptibili","pith_inferences":["If the transition-matrix element in the true Roland-Cerf condition is not constant, the claimed 'exact' benchmark is a gap-only schedule; a fairer comparison might narrow the reported advantage, though it would likely preserve the surrogate's edge over linear annealing.","The same susceptibility signal could be used to time pauses or reverse ramps in hardware annealers, since it identifies when quantum fluctuations are dynamically relevant.","The Delta* of about 0.016 failure threshold suggests a cheap predictor: instances with minimum spectral gaps below this value are the ones where gap-based scheduling is most likely to backfire."],"forward_implications":["Schedule construction no longer requires diagonalizing the quantum Hamiltonian; any problem amenable to SQA can be scheduled from the same Monte Carlo run that solves it.","Exact spectral-gap strategies can fail at finite time in two identifiable ways, and both are avoided by smoothing the time allocation over the critical region.","The fraction of instances exhibiting the boundary-gap trap grows with system size, so the advantage of observable-based scheduling should widen for larger problems.","The method comes with an open-source implementation, making the schedule reproducible and directly applicable to QUBO/Ising optimization workflows."],"fun_headline_variants":["Susceptibility schedule beats exact-gap annealing","Worldline susceptibility outperforms Roland-Cerf schedule","Spin-fluctuation annealing schedule defeats ideal adiabatic","Cheap observable beats exact spectral gap in annealing","Susceptibility-based annealing trumps exact local-adiabatic"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmarked 'exact Roland-Cerf schedule' is built from the inverse-square gap alone, omitting the transition-matrix element |<E1|partial_s H|E0>| of the true local-adiabatic condition; if that element varies along the anneal, the schedule being outperformed is not the exact one.","fun_headline_variants_meta":{"raw":{"variants":["Susceptibility schedule beats exact-gap annealing","Worldline susceptibility outperforms Roland-Cerf schedule","Spin-fluctuation annealing schedule defeats ideal adiabatic","Cheap observable beats exact spectral gap in annealing","Susceptibility-based annealing trumps exact local-adiabatic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1461,"prompt_tokens":734,"completion_tokens":727,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":650}},"tokens_in":478,"tokens_out":727,"duration_ms":7674,"temperature":1.0,"reasoning_tokens":650,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:33:17.995030+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Roland-Cerf schedule on the same instances including the full matrix element |<E1|partial_s H|E0>| in the velocity integral; if the surrogate no longer beats it on a substantial fraction of instances, the central claim fails. A cheaper check: evaluate the matrix element for the n=10 and n=12 instances and test whether the boundary-gap trap and oscillations persist.","supporting_citations":[],"review_version":1}