{"id":"fe97c5f1-55bd-458a-9106-b3f206dafdcc","arxiv_id":"2507.02574","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"A reproducible pedagogical comparison showing pseudo-likelihood inference matches or beats mean-field inversion across phase-transitioning Ising, Potts, and Blume-Capel models.","lead":"This paper explains and tests three statistical mechanics methods for recovering hidden interaction parameters from observed configurations of Ising, Potts, and Blume-Capel systems near phase transitions. It provides a GitHub repository to reproduce the numerical experiments and experiment with new systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that pseudo-likelihood matches or beats mean field at every temperature is supported only by single-realization error curves with no error bars; the observed ordering may be sampling noise.","rationale":"The reader's weakest-assumption concern about representative sampling is real and is explicitly acknowledged in the manuscript, but the most load-bearing issue for the paper's strongest comparative claim is more specific: even granting perfectly equilibrated data, the reported γ_J curves come from single datasets and single disorder/graph realizations, with no error bars. The statement that pseudo-likelihood is 'equally or better' than mean field at every temperature therefore overstates what the evidence can establish; the differences could be within statistical uncertainty. This does not undermine the pedagogical purpose of the paper, and the derivations in Secs. III.B-III.C are standard and mostly sound. The GitHub repository is a genuine asset for reproducibility and could be used to run the replication test proposed above. Because the reader already issued a conditional verdict, my concern does not move the verdict; it sharpens the condition that is needed before the comparative claim can be taken quantitatively.","tokens_in":32253,"tokens_out":2628,"duration_ms":32961,"concrete_test":"For each model and topology in Figs. 5-8, generate R=20 or more independent datasets (fresh disorder realizations for the disordered models, fresh MCMC seeds otherwise) at each plotted temperature, using the same parameters as in the paper. Compute γ_J for Mean Field and Pseudo-Likelihood on every replicate, then report the mean and standard error of the paired difference γ_MF - γ_PL and the fraction of replicates in which PL is not worse. If at any temperature the difference is within error bars or the sign flips across replicates, the 'every temperature' claim should be softened to a qualitative observation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central comparative claim, stated in Sec. IV A 1 ('the Max-Pseudo-Likelihood approach ... performs equally or better than the Mean Field one at every temperature') and repeated for the Potts and Blume-Capel models, rests entirely on the reconstruction-error curves γ_J(T) defined by Eq. (85). As presented in Figs. 5-8, these curves have no error bars and no averaging over independent datasets or disorder realizations; for the disordered Ising models the couplings are a single random draw, and for the Potts model the caption of Fig. 6 states that 'the same graph has been used for all data points.' With M=20000 (or 15000 Wolff steps) configurations per temperature, γ_J is a random quantity whose run-to-run fluctuations can easily exceed the apparent gap between the pseudo-likelihood and mean-field curves, especially near criticality where correlation times and sample-to-sample fluctuations are large. The paper itself flags the deeper fragility: in Sec. III B the replacement of the full ensemble by M configurations is called 'the fundamental hypothesis' and it is noted that 'simple equilibrium might not be sufficient.' In low-temperature spin-glass or first-order-transition regimes, incomplete sampling can bias both methods differently, so the ranking could change. Additionally, the lasso strength is set ad hoc (λ=1, 0.01, 10^-4 for the three model families) without cross-validation; while the central claim is nominally about plain pseudo-likelihood, the figures interleave lasso results, making the comparison less clean. In short, the claim is plausible but not quantitatively supported as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops and compares three inverse-problem methods—maximum likelihood, maximum pseudo-likelihood (with and without lasso regularization), and naive mean-field—for reconstructing couplings of Ising, q=4 Potts-clock, and Blume-Capel models defined on lattices and Erdős–Rényi graphs. Analytic derivations in Sec. III are followed by numerical estimates of reconstruction error γ_J and rank plots from synthetic equilibrium Monte Carlo data in Sec. IV, with a companion GitHub repository. The central comparative claim is that pseudo-likelihood performs equally or better than mean-field at every temperature.","tokens_in":32545,"tokens_out":7442,"duration_ms":81944,"significance":"The manuscript is a useful didactic synthesis that brings together standard inference techniques and applies them across ferromagnetic, spin-glass, first-order, and tricritical settings. Its strengths include a careful presentation of the pseudo-likelihood derivation, explicit formulas for the Potts and Blume-Capel generalizations, and a public repository that should allow reproduction. The comparative claim, however, is currently supported only by single-realization error curves without uncertainty quantification, so the quantitative benchmark is not yet at the level needed for the stated 'every temperature' conclusion. If the comparison is confirmed with proper statistics, the paper would be a valuable reference for practitioners.","major_comments":[{"comment":"The statement in Sec. IV A 1 that 'the Max-Pseudo-Likelihood approach of Sec. III B performs equally or better than the Mean Field one at every temperature' is not supported by the evidence as presented. The γ_J(T) curves in Figs. 5-8 appear to be single realizations with no error bars, no averaging over independent datasets or initial conditions, and no disorder averaging for the disordered models; the caption of Fig. 6 states that the same graph is used for all data points. With M = 20000 configurations (or 15000 Wolff steps), γ_J is a random quantity whose run-to-run fluctuations, especially near criticality, could be comparable to the apparent gap between the pseudo-likelihood and mean-field curves. Please provide error bars from bootstrapping over configurations or from repeated independent simulations, and state explicitly how many disorder realizations are used, before drawing a conclusion that is uniform in temperature.","section":"Sec. IV A, Eq. (85), Figs. 5-8"},{"comment":"The lasso results are advertised as further improvements ('Lasso further decreases the error...' in Sec. IV A 1 and analogous statements for Potts and Blume-Capel), but the regularization strength is set ad hoc: λ = 1, 0.01, and 10^-4 for the Ising, Potts, and Blume-Capel cases, respectively. The paper itself notes in Sec. III B 4 that systematic methods such as cross-validation exist. Since λ controls the bias-variance trade-off, the reported lasso improvement may be an artifact of the chosen λ rather than a robust property of the method. At minimum, include a λ-sensitivity analysis or a cross-validated choice for one representative model, and temper the claims about lasso accordingly.","section":"Sec. III B 4 and Sec. IV A, Figs. 5-7"},{"comment":"The critical temperatures for Erdős–Rényi graphs are computed with BP equations derived for random regular graphs, with the text asserting that 'when a second order transition happens both RR and ER have the same critical temperature.' This is not generally correct as stated. For a ferromagnetic Ising model on an RR graph of degree c, the linear stability of the paramagnetic fixed point in Eqs. (C5)-(C6) gives (c-1) tanh(βJ) = 1, whereas for an ER graph of mean degree c the condition is c tanh(βJ) = 1, Eq. (7). For c = 4 these give different temperatures (approximately 2.885 and 3.912). The paper either needs to clarify that d denotes the mean residual degree of the ER graph (so that d = c, not c-1) or correct the RR/ER equivalence claim; as written, the ER critical lines in Table I, Figs. 2 and 16, and the dashed lines in Figs. 5-8 are not derived consistently.","section":"Appendix C, Eqs. (C1)-(C2)"}],"minor_comments":[{"comment":"The text 'we take M = 2c/N = 2N' appears to be a typo; with average connectivity c = 4 and N spins, the correct relation is M = cN/2 = 2N.","section":"Sec. II C"},{"comment":"The ℓ1 regularizer for the fields is written as λ_h h_i, but a sparsity-promoting penalty should use |h_i|; please correct the expression, or explain the intended signed penalty.","section":"Sec. III B 4"},{"comment":"There is a typo in 'phaenomenologu ` ıy' (should be 'phenomenology'), and several other spelling and grammar errors (e.g., 'refereed to' in Sec. III B 4) should be corrected.","section":"Sec. II A 2"},{"comment":"In the caption of Fig. 18, 'µc ∈ (1.960, 1970)' should read 'µc ∈ (1.960, 1.970)'.","section":"Appendix B"},{"comment":"The rank plots would be easier to interpret if the number of disorder realizations and the temperature at which data are sampled were stated in every caption; currently only some captions give T and β.","section":"Sec. IV B, Figs. 9-12"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a pedagogical contribution with a welcome code release; the headline comparison and the Appendix C consistency issue are fixable within the scope of a revision. I do not see grounds for rejection. The journal should ensure that the revised version reports uncertainty estimates for the central comparison and corrects or clarifies the RR/ER critical-temperature derivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful tutorial, not a research advance. It walks through Max-Likelihood, Pseudo-Likelihood, and Mean-Field inference on Ising, Potts, and Blume-Capel models on lattices and random graphs, with a GitHub repo to reproduce everything. If you need a single place to point a student learning inverse statistical mechanics, this is a decent candidate.\n\nWhat it does well: the derivations are standard and I did not catch errors. The presentation is careful, with clear motivation for the pseudo-likelihood approximation and an honest discussion of when sampling fails. The code availability is real, and the repository link is live. The comparison of PL versus MF across several models is useful as a benchmark, even if the headline claim is softer than the text suggests.\n\nThe soft spots: the central comparative claim, that PL \"performs equally or better than Mean Field at every temperature,\" is stated repeatedly, but the evidence is a set of single-realization error curves with no error bars. For the disordered Ising models the couplings are one random draw; the Potts figure caption notes the same graph is used for all data points. With M=20000 configurations, the gap between curves could easily be sampling noise, especially near criticality. The paper also uses ad hoc lasso parameters without cross-validation. That said, the paper itself flags the key limitation in Section III B: the sampling must be representative, and equilibrium may not be sufficient. So the authors are not hiding it; they just do not let it temper the conclusion enough.\n\nThe other limitation is by design: this is not new methodology, and the paper says so plainly. If you want new theory or a resolution of an open problem, this is not it. But as a teaching tool and a codebase for benchmarking, it is solid.\n\nI would send it to review. A referee can ask for error bars and a more careful statement of the comparative claim, and the result will be a stronger tutorial. For a reading group, I would use it as an introductory reference, but I would not build a research agenda on it.","headline":"A careful, genuinely useful tutorial on inverse statistical mechanics with real code; the headline comparative claim needs error bars, but the paper deserves peer review.","tokens_in":33076,"tokens_out":2264,"would_cite":false,"duration_ms":26526,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pseudo-likelihood inference beats mean-field at every temperature","keywords":["inverse problems","maximum pseudo-likelihood","mean-field inference","phase transitions","Ising model","Potts clock model","Blume-Capel model","reconstruction error"],"falsifier":"Generate a controlled dataset for a spin glass on a random graph using a deliberately short, poorly equilibrated Monte Carlo run (e.g., far fewer Monte Carlo sweeps than the autocorrelation time) and check whether the pseudo-likelihood reconstruction error rises sharply relative to the equilibrium-sampling baseline, or, conversely, find any temperature in the paper's own models where the mean-field reconstruction error falls strictly below the pseudo-likelihood one — either result would contradict the claimed uniform superiority.","tokens_in":32071,"feed_emoji":"⚛","tokens_out":6076,"duration_ms":62872,"temperature":0.7,"pith_summary":"The paper asks how interaction strengths of a statistical-physics model can be recovered from measured configurations when the model itself is poised near a phase transition, where data take qualitatively different forms. It develops and tests three inference routines — maximum likelihood, maximum pseudo-likelihood, and naive mean-field inversion — on ordered and disordered Ising models, the four-state Potts clock, and the Blume-Capel model, on both lattices and random graphs. The central finding is that max-pseudo-likelihood reconstructs couplings equally well or better than mean-field at every temperature studied, with lasso regularization giving further gains when the true couplings are sparse. A sympathetic reader would care because pseudo-likelihood requires no partition-function evaluation and works from raw spin configurations, making it a practical default for equilibrium data.","feed_headline":"Pseudo-likelihood beats mean-field at every temperature","feed_subtitle":"Coupling inference stays accurate across phase transitions in Ising, Potts, and Blume-Capel systems.","key_machinery":"The load-bearing object is the single-site pseudo-likelihood, which replaces the full Boltzmann-Gibbs conditional average over the other $N-1$ spins with an empirical average over the $M$ measured configurations, turning inference into a logistic-regression problem with convex structure per row. The update rules for couplings and fields are steepest-descent equations whose fixed points match empirical magnetizations and correlations; in the Ising case the conditional probability is $P(s_i|\\{s_{\\setminus i}\\}) = (1 + e^{-z_i})^{-1}$, so maximizing the pseudo-likelihood is a sigmoid/logistic regression. The mean-field comparator is the inverse of the empirical covariance matrix, $\\beta J^{\\rm MF}_{ij} = -(\\Gamma^{-1})_{ij}$. The two are judged through the normalized reconstruction error $\\gamma_J$ and through rank plots of the sorted inferred couplings.","core_discovery":"On its own terms, the paper establishes that the max-pseudo-likelihood approach of Sec. III B performs equally or better than the mean-field approach at every temperature. This is shown through the reconstruction error $\\gamma_J$ computed over temperature for the Ising model (ordered and bond-disordered, on a square lattice and on Erdős–Rényi graphs), for the $q=4$ Potts clock model (on a cubic lattice and on random graphs), and for the Blume-Capel model (on a square lattice and on random graphs, across both second-order and first-order transition regions). The finding holds across data of 15,000–20,000 Monte Carlo configurations; introducing lasso regularization further lowers the error in the ordered, sparse-coupling cases. The paper also derives the pseudo-likelihood inference as a row-by-row logistic regression and identifies the sampling assumption — that the measured configurations represent the full Boltzmann-Gibbs ensemble — as the premise on which the method is built.","pith_inferences":["The empirical claim is established only on the specific suite of models, temperatures, and dataset sizes used; whether the 'equally or better at every temperature' ordering survives for other topologies (e.g., scale-free graphs), other spin alphabets, or data with measurement noise is a natural testable extension.","A practical diagnostic suggested by the paper's own caveat: compute an equilibration indicator such as autocorrelation time or the Binder parameter of the input configurations and flag low-confidence inferences where sampling is poor.","The same pseudo-likelihood machinery, with an appropriate conditional distribution, could be carried over to inverse problems with non-Boltzmann or non-equilibrium data, but the equivalence to logistic regression would no longer guarantee convexity."],"forward_implications":["Near a phase transition, pseudo-likelihood remains a reliable coupling-reconstruction routine, so practitioners can use it without first knowing whether their data come from the ordered or disordered side.","When the true interaction network is sparse and ferromagnetic, lasso-regularized pseudo-likelihood is the best of the tested options, producing sharper sorted-coupling plots.","The rank-plot diagnostic gives a graphical way to check inference quality as a function of dataset size, without needing a known ground truth in application settings.","Since the method needs only raw configurations and no partition function, it scales to systems where exact maximum likelihood is infeasible.","The accompanying repository lets a user reproduce the comparisons and apply the pipelines to new models."],"supporting_citations":[{"why":"Supplies the pseudo-likelihood and logistic-regression framework that the paper's best-performing method is built on.","marker":"[10]"},{"why":"Supplies the mean-field inference framework against which pseudo-likelihood is compared.","marker":"[9]"},{"why":"Provides the general inverse-problem statistical-mechanics background connecting mean-field and pseudo-likelihood methods.","marker":"[11]"},{"why":"Provides the exact critical temperature of the 2D Ising model used to place data relative to the phase transition.","marker":"[41]"},{"why":"Defines the disordered Ising (Edwards-Anderson) model used as the spin-glass test case.","marker":"[39]"},{"why":"Defines the Potts clock model used in the vector Potts experiments.","marker":"[2]"},{"why":"Defines the Blume-Capel model with its tricritical and first-order transitions.","marker":"[3]"},{"why":"Supplies the Wolff algorithm used to generate equilibrium Potts data.","marker":"[7]"},{"why":"Supplies the Belief Propagation equations used to locate critical lines on Erdős–Rényi graphs.","marker":"[94]"}],"fun_headline_variants":["Pseudo-likelihood outperforms mean-field at every temperature","Inverse problems: pseudo-likelihood wins across phase transitions","Spin coupling inference: pseudo-likelihood beats mean-field","Pseudo-likelihood tops mean-field from order to disorder","Across transitions, pseudo-likelihood beats mean-field"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's results stand on the assumption that the $M$ measured Monte Carlo configurations are a statistically representative sample of the full Boltzmann-Gibbs equilibrium ensemble; if the sampling misses entire regions of configuration space — as in glassy low-temperature or first-order-transition regimes — the inferred couplings are biased.","fun_headline_variants_meta":{"raw":{"variants":["Pseudo-likelihood outperforms mean-field at every temperature","Inverse problems: pseudo-likelihood wins across phase transitions","Spin coupling inference: pseudo-likelihood beats mean-field","Pseudo-likelihood tops mean-field from order to disorder","Across transitions, pseudo-likelihood beats mean-field"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000756,"raw_usage":{"total_tokens":3327,"prompt_tokens":881,"completion_tokens":2446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":2363}},"tokens_in":497,"tokens_out":2446,"duration_ms":24271,"temperature":1.0,"reasoning_tokens":2363,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:25:25.217214+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a controlled dataset for a spin glass on a random graph using a deliberately short, poorly equilibrated Monte Carlo run (e.g., far fewer Monte Carlo sweeps than the autocorrelation time) and check whether the pseudo-likelihood reconstruction error rises sharply relative to the equilibrium-sampling baseline, or, conversely, find any temperature in the paper's own models where the mean-field reconstruction error falls strictly below the pseudo-likelihood one — either result would contradict the claimed uniform superiority.","supporting_citations":[{"cited_title":"low dimension","cited_arxiv_id":null,"evidence_quote":"Supplies the pseudo-likelihood and logistic-regression framework that the paper's best-performing method is built on."},{"cited_title":"(35), as a logistic regression function","cited_arxiv_id":null,"evidence_quote":"Supplies the mean-field inference framework against which pseudo-likelihood is compared."},{"cited_title":"In formulas, this amount to have vanishing connected correlation functions, Eq","cited_arxiv_id":null,"evidence_quote":"Provides the general inverse-problem statistical-mechanics background connecting mean-field and pseudo-likelihood methods."},{"cited_title":"Here we have explicitly written the likelihood in terms of the Boltzmann-Gibbs distribution Eq","cited_arxiv_id":null,"evidence_quote":"Defines the Potts clock model used in the vector Potts experiments."},{"cited_title":"(13), and the theoretical average magnetization and correlation, Eqs","cited_arxiv_id":null,"evidence_quote":"Defines the Blume-Capel model with its tricritical and first-order transitions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Wolff algorithm used to generate equilibrium Potts data."}],"review_version":1}