{"id":"10a6c421-89bd-4113-94a1-79a484663e74","arxiv_id":"2501.13205","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Numerical ground-state data for the p=3 mean-field spin glass show ln N/N energy corrections and Gumbel statistics, matching the random energy model and suggesting all p>=3 spin glasses share REM-like zero-temperature corrections.","lead":"This paper computes approximate ground states of a three-spin glass model for up to 256 spins and reports that its finite-size corrections match the random energy model, not the two-spin Sherrington-Kirkpatrick model. A generalist might care because it suggests a sharp boundary in spin glass behavior: all p-spin glasses with p greater or equal to 3 may share the same zero-temperature scaling, while p=2 is special.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uncertified EO heuristic bias could be comparable to the claimed ln N/N corrections, so the REM-like conclusion needs an optimality benchmark.","rationale":"The paper's central claim is that p=3 ground-state finite-size corrections are REM-like (ln N/N for the mean, 1/N for the width) and that this extends to all p>=3. The most load-bearing condition is that the reported 'approximate optima' from the EO heuristic are accurate enough that their systematic error does not masquerade as the small finite-size corrections being measured. The reader identified exactly this weakness, and I agree. The paper has real strengths: extensive sample counts at small N, anchoring to the exactly known 1RSB limit, a clear demonstration that a power-law form with fitted exponent fails to hit that limit, and a collapse of the q=0.25 diluted data onto the undiluted curve. These make the empirical result plausible, but they do not certify optimality. A heuristic bias that decays as slowly as ln N/N would reproduce the observed trend while having nothing to do with the true physics; the exact-limit anchor only constrains the combined signal. The universal statement for all p>=3 is also an extrapolation from a single p value, but it is phrased as a suggestion and is secondary to the numerical reliability issue. Because the reader already returned CONDITIONAL, and the proposed exact benchmark is the natural next step, the verdict should remain unchanged rather than being hardened or relaxed.","tokens_in":7865,"tokens_out":10862,"duration_ms":125750,"concrete_test":"For the same instances used in Table I (or a random subset with at least 1000 samples per N), compute certified ground states for N=16,18,20,22,24 by exhaustive search over the 2^N spin configurations, using the Hamiltonian in Eq. (1) with a branch-and-bound or meet-in-the-middle implementation; for N=24 this is about 1.7e7 configurations and is feasible. Compare the average gap Δ_N = <E_EO - E_exact>/N between the EO result and the certified optimum. If Δ_N decays faster than ln N/N and remains below 10% of the measured ln N/N correction at each N, the heuristic-bias concern is resolved; if Δ_N is comparable to or decays slower than ln N/N, the fitted corrections in Fig. 1 are contaminated and the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the extremal-optimization heuristic returns configurations close enough to true ground states that its systematic error decays faster than the measured ln N/N and 1/N corrections. This premise is not certified: Table I labels all entries 'approximate optima', Fig. 5 refers to the 'presumed ground state energy', and the adaptive stopping rule (twice the spin flips needed for the latest new optimum) supplies no lower bound on the true ground-state energy. Runtimes escalate so steeply that N is capped at 256. At N=256 the measured correction to the exact limit e_inf=-0.8132 is about 0.0135 in energy density, so even a heuristic bias of a few times 10^-3 at the largest sizes would materially shift the fitted amplitude A in <e0>_N = e_inf + A ln N/N. Anchoring the extrapolation to the known 1RSB limit and rejecting a power-law fit that extrapolates to the wrong value are genuinely helpful consistency checks, but they constrain only the sum of the true correction and the bias; they cannot separate the two. Without an independent optimality check, the claim that p=3 corrections are REM-like remains an unverified consequence of a heuristic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies ground-state properties of the fully connected p=3 Ising spin glass with all-to-all triplet bonds, using a GPU-vectorized extremal optimization (EO) heuristic for system sizes N=16 to 256, for the complete model and for a bond-diluted version with 25% of bonds present. From tabulated mean ground-state energies and standard deviations, the authors argue that the finite-size correction to the mean energy is C_N = ln N/N (Eq. 4), that the standardized ground-state energy distribution is close to a Gumbel form with m=0.87(5) (Eq. 5), that the standard deviation decays as 1/N, and that the local-field distribution develops a pseudogap. The extrapolation is anchored to the exactly known 1RSB thermodynamic limit <e_0>_∞ = -0.8132, which is used to reject a power-law correction that extrapolates to -0.8151. The results are contrasted with SK (p=2) and used to propose that all p-spin models with p>=3 share REM-like ground-state corrections.","tokens_in":8084,"tokens_out":8050,"duration_ms":87505,"significance":"If the central claim is correct, the paper establishes a sharp T=0 distinction between p=2 and p>=3 spin glasses and predicts a qualitative contrast in dilution dependence, which would be an important step toward a unified picture of finite-size corrections in mean-field spin glasses. The paper's strengths are its extensive tabulated data, its use of the exactly known thermodynamic limit as a model-selection constraint, and its explicit comparison of candidate correction forms. The central weakness is that all reported states are approximate optima of a heuristic whose systematic error is not quantified, and the exact-limit anchor constrains only the sum of the true correction and the heuristic bias. The stress-test concern therefore lands: the REM-like conclusion is plausible and well presented, but it is not yet certified.","major_comments":[{"comment":"The inference that C_N = ln N/N rests on extrapolating data that are explicitly labeled as approximate. Table I calls all entries 'approximate optima', Fig. 5 refers to the 'presumed ground state energy', and the adaptive stopping rule (use twice as many flips as were needed for the latest new optimum) provides no lower bound on the true ground-state energy. At N=256 the measured deviation from -0.8132 is about 0.0135 in energy density; a heuristic bias of a few times 10^-3 at that size would change the fitted amplitude A in <e0>_N = e_∞ + A ln N/N by tens of percent. Anchoring to the exact 1RSB limit and rejecting the power-law fit that extrapolates to -0.8151 are genuinely useful consistency checks, but they constrain only the combined effect of true corrections and bias and cannot separate the two. The authors should add an independent optimality benchmark, for example exact branch-and-bound or a second heuristic with a different landscape bias for smaller N, and report how any residual bias scales with N relative to ln N/N.","section":"Eq. (3)-(4), Fig. 1, Table I, Fig. 5"},{"comment":"The claim of no dilution dependence is supported by only one diluted value, q=0.25. The collapse of this single series onto the q=1 data after the 1/sqrt(q) rescaling from Ref. [26] is suggestive, but it cannot establish the abstract's stronger assertion that no variation with q is found, nor the broader conclusion that the absence of q-dependence is a general property of p>=3. The SK inset in Fig. 1 shows many q values, whereas the p=3 panel shows only q=1 and q=0.25. At least a few dilution values spanning different regimes are needed before claiming a qualitative contrast with SK, especially since the diluted data are produced by the same heuristic and may share the same systematic bias.","section":"Table I, Fig. 1"},{"comment":"The variance scaling claim is internally inconsistent with the stated REM comparison. The text says the observed behavior 'matches more closely with REM (which has sigma ~ ln N/N)' than with SK, but Fig. 3 is presented as showing 'simply a series in 1/N'. These are parametrically different leading behaviors, and a logarithmic factor cannot be dismissed as hard to discern when the claimed distinction between candidate forms is the main message. The authors should fit sigma(e0) N against ln N, or a two-term form with both 1/N and ln N/N, and state whether the data actually favor a log-free 1/N law or a REM-like ln N/N law. This affects the central claim that p>=3 has exactly the same corrections as REM.","section":"Fig. 3 and accompanying text"},{"comment":"The Gumbel conclusion is based on pooling all sizes after standardizing by sigma and on a four-parameter fit, with fitted m=0.87(5). The parameter m is about 2.6 standard errors below the pure Gumbel value m=1, so the evidence for an exact Gumbel form is weaker than the abstract implies. The paper should report a goodness-of-fit measure and examine how the fitted parameters vary with N, since the pooled fit is dominated by the many small-N samples. Without such tests, 'consistent with a Gumbel distribution' is a reasonable summary, but it is not a quantitative confirmation of the REM-like distributional claim.","section":"Eq. (5), Fig. 2"}],"minor_comments":[{"comment":"There is a typo: 'there its very little variation' should read 'there is very little variation'.","section":"Paragraph discussing Fig. 2"},{"comment":"The local-field pseudogap fit in Eq. (7) reports alpha ~ 0.6, A ~ 0.24, and B ~ 0.06 without uncertainties or a goodness-of-fit statement; please add fit details or label the values as preliminary.","section":"Fig. 4 and Eq. (7)"},{"comment":"The GPU implementation is deferred to a future paper via Ref. [31]; since the numerical data are the main product of the manuscript, a short description of the vectorized update rule and the bit-packed bond representation would improve reproducibility.","section":"Section IV, implementation paragraph"}],"recommendation":"major_revision","confidential_remarks":"The referee's main concern is the absence of any optimality benchmark for the EO heuristic, not the choice of model-selection methodology. The exact thermodynamic-limit anchor is a real strength and makes the manuscript worth pursuing. I would ask the authors to quantify heuristic bias on smaller sizes, add at least one or two more dilution values, and resolve the sigma ~ 1/N versus sigma ~ ln N/N tension before the paper is accepted. These are fixable within the scope of a revision; I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper delivers a genuinely new numerical result: the first systematic finite-size data for p=3 spin-glass ground states, and the observation that the diluted (q=0.25) data collapses onto the undiluted curve when rescaled by 1/sqrt(q). That collapse is the strongest piece of news, because SK (p=2) shows strong dependence on bond density while p=3 appears not to. The paper also argues, plausibly, that the mean-energy correction is ln N/N rather than a power law, using the exactly known 1RSB limit to reject a power-law fit that extrapolates to the wrong value. That is a principled model-selection step and deserves credit.\n\nThe soft spots are real but not fatal. The largest is the unquantified heuristic bias. Table I says \"approximate optima\", Fig. 5 calls the target the \"presumed ground state energy\", and the stopping rule (running twice as many flips as needed to find the last new optimum) gives no lower bound. At N=256 the claimed ln N/N correction to e_inf is about 0.0135; if the EO bias is a few times 10^-3, it shifts the fitted amplitude materially. The exact-limit anchoring constrains the sum of correction plus bias, not the correction alone. An independent optimality check (e.g., a different solver on the same instances, or exact branch-and-bound for N up to ~40) would tighten this.\n\nThe q-independence claim rests on a single dilution value, q=0.25. That is a modest evidential base for a universality-style claim across all p>=3. Also, the width-scaling statement has an internal inconsistency: the abstract says 1/N corrections for the distribution, but the text says sigma~ln N/N for REM, and Fig. 3 admits logarithmic factors. Minor, but should be cleaned up. No code or raw distributions are released, which is a shame for a benchmark paper.\n\nNone of this is fatal. The central empirical result—p=3 corrections look REM-like, unlike SK—is plausible and directly relevant to the spin-glass and combinatorial optimization communities. The paper deserves a serious referee, and the referee should ask for an optimality benchmark, a second dilution, and the artifacts.\n\nI would bring it to a reading group and would cite it, with the caveat that the heuristic bias question is open.\n\nBest.","headline":"A plausible and genuinely new numerical result—p=3 ground-state corrections look REM-like, unlike SK—but the unquantified heuristic bias and single-dilution evidence base keep it from being airtight.","tokens_in":8640,"tokens_out":1856,"would_cite":true,"duration_ms":17660,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82B44","82D30","60G70"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that ground states of the mean-field 3-spin Ising spin glass have finite-size corrections, distributions, and dilution behavior matching the random energy model (REM), making p=2 (SK) the exceptional case.","keywords":["mean-field spin glass","p-spin model","ground state energy","finite-size corrections","random energy model","extremal optimization","Gumbel distribution","Sherrington-Kirkpatrick model"],"falsifier":"An exact or rigorously certified computation of ground states for, say, N=128 or 256 (or a demonstration that the EO systematic error itself decays as ln N/N) would settle the claim: if the true energies depart from the ln N/N extrapolation, or if the 1/N width scaling breaks down, the REM-like conclusion fails.","tokens_in":7622,"feed_emoji":"⚛️","tokens_out":4135,"duration_ms":39217,"temperature":0.7,"pith_summary":"The paper tries to establish that, at zero temperature, the mean-field spin glass with three-spin couplings (all triplets of N Ising spins coupled) is governed by the same finite-size corrections as the random energy model, the p=∞ limit. Using an extremal-optimization heuristic on GPUs to reach N=256, the authors find the average ground-state energy density approaches its thermodynamic limit with a ln N/N correction, the width of its distribution decays as 1/N, and the distribution is close to Gumbel. Bond dilution does not change the leading corrections, unlike the Sherrington-Kirkpatrick (p=2) model where the correction exponent varies with dilution. If true, all p-spin models with p≥3 behave REM-like at T=0, and SK's anomalous corrections (like the $N^{-5}$/6 width) are the exception rather than the rule.","feed_headline":"3-spin glass ground states mimic REM at T=0","feed_subtitle":"Energy corrections scale as ln N/N and fluctuations as 1/N, unlike the 2-spin Sherrington-Kirkpatrick model.","key_machinery":"The load-bearing machinery is the extremal optimization (EO) heuristic, run in a GPU-vectorized, bit-packed form, which produces approximate ground states up to N=256; its reliability is judged by extrapolation plots to the exact thermodynamic limit. The analytic anchor is the exactly known ground-state energy density ⟨e0⟩_{p=3,∞} = −0.8132... from 1-step replica symmetry breaking calculations, which the fits are forced to reproduce. The finite-size correction ansatz C_N ~ ln N / N, borrowed from REM, is what discriminates between competing correction forms, and the overlap gap condition for p≥3 explains why the search is harder than for SK.","core_discovery":"The central claim is that the p=3 mean-field spin glass is REM-like in its zero-temperature ground-state statistics. Concretely: fixing the thermodynamic-limit energy to the exactly known value from 1-step replica calculations, the ensemble-averaged ground-state energy density follows ⟨e0⟩_N ~ ⟨e0⟩_∞ + A ln N / N, rather than a power-law; a power-law fit with exponent ω≈4/5 would miss the known limit by about 20σ. The standard deviation of the ground-state energy distribution scales as 1/N (with possible undetectable logarithmic factors), and the rescaled distribution is a generalized Gumbel with m≈0.87, close to the pure Gumbel of REM and far from SK's m≈5. The paper further claims that dilution to 25% bond density leaves these corrections essentially unchanged once energies are rescaled by 1/√q, in stark contrast to SK. The conclusion drawn is that p=2 is the special model, and all p≥3 share REM's ground-state corrections.","pith_inferences":["Editorial extension: a natural testable next step is to run the same extrapolation analysis for p=4 and p=5 to see whether the same ln N/N and Gumbel forms hold, which would corroborate the p≥3 universality without relying on the p→∞ limit.","Editorial extension: the claimed REM-like behavior at T=0 might connect to the finite-temperature random first-order transition; one could check whether the finite-size corrections cross over to SK-like behavior near the transition temperature.","Editorial extension: the heuristic's slower convergence for p=3 than for SK, despite similar system sizes, could be repurposed as a practical benchmark for the computational hardness implied by the overlap gap condition."],"forward_implications":["If correct, the zero-temperature ground-state properties of all p-spin models with p≥3 are universal and coincide with REM's, so p=3 data can serve as a benchmark for testing heuristics.","The overlap gap condition makes local search hard, but the observed ln N/N corrections imply that even for p=3, near-optima extrapolate reliably, extending the usefulness of extrapolation-based evaluation of heuristics.","The dilution independence gives a sharper distinction between p=2 and p≥3: any theory of SK's q-dependent correction exponent (ω(q)) must explain why p=3 has no such dependence.","The measured local-field pseudo-gap with P(0) ~ N^−0.6 provides a concrete signature for future mean-field or replica calculations to reproduce."],"supporting_citations":[{"why":"Establishes the overlap gap condition for the 3-spin model, motivating why local search is harder than for SK.","marker":"[14]"},{"why":"Provides the bond-diluted SK data with q-dependent correction exponent that the 3-spin results are contrasted against.","marker":"[18]"},{"why":"Supplies the extremal optimization method and SK ground-state extrapolation technique that this study adapts.","marker":"[21]"},{"why":"Gives the exactly known thermodynamic-limit ground-state energy for p=3 via 1-step replica calculations, used as the anchor for fits.","marker":"[24]"},{"why":"Confirms the thermodynamic-limit energy value used to fix extrapolations.","marker":"[25]"},{"why":"Defines the random energy model and its exact corrections, the template for the claimed scaling.","marker":"[15]"},{"why":"Provides the REM analysis that supplies the ln N/N correction form and Gumbel distribution.","marker":"[16]"},{"why":"Gives the exact SK result σ ~ N^−5/6, contrasted with the 1/N width scaling found for p=3.","marker":"[27]"}],"fun_headline_variants":["3-spin glass ground states mirror REM","p=3 spin glass: REM-like energy scaling","T=0 3-spin glass behaves like random energy model","3-spin glass ground state: Gumbel and 1/N","All p≥3 spin glasses share REM scaling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand on the assumption that the configurations the extremal-optimization heuristic returns are close enough to true ground states that the heuristic's systematic error decays faster than the measured ln N/N and 1/N corrections; there is no certificate of optimality.","fun_headline_variants_meta":{"raw":{"variants":["3-spin glass ground states mirror REM","p=3 spin glass: REM-like energy scaling","T=0 3-spin glass behaves like random energy model","3-spin glass ground state: Gumbel and 1/N","All p≥3 spin glasses share REM scaling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1514,"prompt_tokens":1099,"completion_tokens":415,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":335}},"tokens_in":715,"tokens_out":415,"duration_ms":4703,"temperature":1.0,"reasoning_tokens":335,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:22:27.930970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An exact or rigorously certified computation of ground states for, say, N=128 or 256 (or a demonstration that the EO systematic error itself decays as ln N/N) would settle the claim: if the true energies depart from the ln N/N extrapolation, or if the 1/N width scaling breaks down, the REM-like conclusion fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the overlap gap condition for the 3-spin model, motivating why local search is harder than for SK."},{"cited_title":"Boettcher, Physical Review Letters 124, 177202 (2020)","cited_arxiv_id":null,"evidence_quote":"Supplies the extremal optimization method and SK ground-state extrapolation technique that this study adapts."},{"cited_title":"Boettcher, The European Physical Journal B46, 501 (2005)","cited_arxiv_id":null,"evidence_quote":"Gives the exactly known thermodynamic-limit ground-state energy for p=3 via 1-step replica calculations, used as the anchor for fits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Confirms the thermodynamic-limit energy value used to fix extrapolations."},{"cited_title":"Montanari and F","cited_arxiv_id":null,"evidence_quote":"Gives the exact SK result σ ~ N^−5/6, contrasted with the 1/N width scaling found for p=3."}],"review_version":1}