{"id":"4ffe2bb8-3d67-420b-9c10-8202c405250c","arxiv_id":"2505.07719","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Training neural control variates on autocorrelated MCMC samples can yield better variance reduction than using only decorrelated samples when the total chain length is fixed.","lead":"This paper asks whether correlated Monte Carlo samples, which are normally treated as redundant, can still help train neural networks that reduce noise in physics simulations. In simple lattice models, the author reports that using such samples improves variance reduction, which could make expensive simulations cheaper.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The correlation-specific benefit is confounded with training-set size: no experiment fixes sample count while varying autocorrelation, so the Section V claim rests on an untested mechanism.","rationale":"The reader's weakest-assumption analysis correctly identifies the central confound: correlated conditions in all figures have more training samples than decorrelated conditions. This is the single most load-bearing concern because the paper's advertised contribution is not merely that more training data helps, but that autocorrelated configurations carry structural information useful for NCV training. The experimental design cannot distinguish these explanations. The proposed equal-sample-count control would settle the matter directly. I agree with the reader's CONDITIONAL verdict: the practical guidance about not discarding correlated samples is plausible and may be useful, but the evidence as presented does not establish the mechanism claimed. The paper does have independent strengths: the underlying NCV construction preserves unbiasedness through Stein's identity, and the results span multiple models and dimensions, which makes the practical direction credible. However, without code, data, or seed statistics, and with the systematic confound, the central claim remains unverified. The correct next step is a revised experiment with sample count held fixed; the verdict should remain conditional until that is provided.","tokens_in":9052,"tokens_out":6127,"duration_ms":68652,"concrete_test":"Re-run the three-dimensional scalar-field experiment (Section IV.B, Figure 6) with two training sets of equal size, N = 10^4: (i) consecutive correlated samples from one chain, and (ii) decorrelated samples obtained by thinning the same chain with a measurement interval of order tau_int, requiring a longer chain to reach N samples, or independent samples if exact sampling is available (e.g., the factorized U(1) model in Section IV.A). Train with identical architecture, optimizer, initialization seeds, and total gradient steps, and compare the test-set variance reduction averaged over at least 10 seeds with error bars. If condition (ii) matches or beats condition (i), the correlation-specific mechanism is not supported and the result reduces to 'more training data helps.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section V) is that using correlated samples to train neural control variates can lead to more effective variance reduction for a fixed Monte Carlo chain length, with the proposed mechanism (Section III) being that autocorrelated configurations encode structural information about the target distribution. The experiments in Section IV, however, never isolate this mechanism. In each comparison, the condition with higher autocorrelation also has more training samples: for a fixed chain, measuring every sweep yields 10^4 training samples, while thinning to decorrelated samples yields 10^3 or fewer. Thus the correlated condition always enjoys a larger dataset and, with minibatch SGD, more gradient updates. The observed improvement in test loss and variance reduction could therefore be driven entirely by the well-known benefit of more training data, not by Markov-chain correlations carrying additional structural information. The paper's own summary in Section IV.A acknowledges this, stating 'more samples while being highly correlated can achieve a better training result,' which conflates the two variables. Because no experiment holds the number of training samples fixed while varying autocorrelation (or vice versa), the central conceptual claim is unsupported. At best, the results demonstrate the trivial point that training on more samples, even correlated ones, helps; they provide no evidence for the stronger claim that correlation itself is beneficial.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript argues that autocorrelated MCMC configurations, which are normally discarded for error estimation, can be used profitably to train neural control variates (NCVs). After reviewing the control-variate construction via Stein's identity and the neural parametrization in Section II, the paper develops a qualitative argument in Section III that correlated samples encode structural information about the target distribution. Section IV presents empirical studies in U(1) gauge theory and scalar field theory in 2-4 dimensions, comparing training on 10^4 samples measured every sweep with training on thinned decorrelated samples, all within a fixed Monte Carlo chain length. Figures 2, 3, 5-7 show loss curves and variance-reduction ratios on a decorrelated test set. Section V concludes that correlated samples should be retained for NCV training.","tokens_in":9356,"tokens_out":5664,"duration_ms":56137,"significance":"If the central claim were established, the paper would provide practical guidance for lattice field theory and other MCMC-intensive applications where independent configurations are scarce. The empirical scope (two theories, several dimensions, and observables with signal-to-noise problems) is appropriate for a first study, and the paper is clearly written. However, as detailed below, the current experimental design does not separate the effect of autocorrelation from the effect of training-set size, so the distinctive conceptual claim is not yet supported. The paper also omits statistical uncertainty estimates for all learning curves, which is essential in this context. The NCV construction itself is standard and the paper does not relabel fitted constants as predictions, so circularity is not a concern.","major_comments":[{"comment":"The main comparison conflates sample count with autocorrelation. For a fixed chain length, measuring every sweep gives 10^4 training samples, while the 'decorrelated' conditions are obtained by thinning and therefore contain at most 10^3 samples (Section IV.A). Consequently, the correlated condition also receives more minibatch gradient updates under fixed hyperparameters. The text in Section IV.A explicitly states 'more samples while being highly correlated can achieve a better training result,' which concedes the confound. To support the conclusion in Section V that correlation itself is beneficial, the authors need experiments in which the number of training samples and gradient steps are held fixed while the integrated autocorrelation time is varied (e.g., by subsampling a correlated chain to the same N_train, or by comparing chains with different proposal step sizes at identical N_train).","section":"Section IV.A, Figures 2, 3, 5-7"},{"comment":"Every learning curve appears to be a single training run; no error bars, standard deviations, or averaging over random seeds, network initializations, or independent MCMC chains are shown. Some displayed differences, for example between the curves in the left panels of Figure 6, appear modest and could be within run-to-run variability. The authors should provide seed-averaged curves with uncertainty bands and a quantitative summary table of final error-reduction ratios, including the number of seeds, so the reader can assess statistical significance.","section":"Figures 2, 3, 5, 6, and 7"},{"comment":"The central 'structural information' mechanism is asserted rather than tested. The statement that local updates 'encode structural information about the distribution beyond what is captured by a sparse set of decorrelated samples' (Section III) is not operationalized: no quantity measures this information, and the experiments in Section IV cannot distinguish the proposed mechanism from the alternative that a larger correlated training set merely supplies more gradient updates. Even after deconfounding sample count, a positive result would not uniquely confirm the mechanism; the paper should either test the mechanism more directly or reframe the claim as a practical observation.","section":"Section III"}],"minor_comments":[{"comment":"The values of the integrated autocorrelation times used in each figure are not stated; report them (for example, in a table) so the independent variable is quantified and the results are reproducible.","section":"Section IV.A and figures"},{"comment":"The signal-to-noise ratios in Figures 1 and 4 are shown without error bars or a description of how the errors were estimated; please add this information.","section":"Figures 1 and 4"},{"comment":"The caption of Figure 2 refers to 'the same MC chain sampled with different integrated autocorrelation times,' but the text does not explain how the autocorrelation time is varied (measurement interval, proposal step size, or acceptance rate). State the protocol explicitly.","section":"Figure captions"},{"comment":"The sentence 'To isolate this effect, we use unusually large acceptance rates' is unclear because the acceptance rate is not varied between the compared curves; clarify what this choice isolates.","section":"Section IV.A"},{"comment":"The manuscript would benefit from a statement on code/data availability for reproducibility, especially since the training curves depend on random seeds and network initialization.","section":"Throughout"},{"comment":"There are minor typographical artifacts, e.g., 'parametrizeg' in Section II.A and a line break in 'T raining' in Section III; please proofread the text.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the author's prior NCV work is cited appropriately. The main issue is the experimental confound described in major comment 1: no experiment isolates autocorrelation from training-set size. I recommend requesting a revised version with deconfounding experiments and proper uncertainty quantification; I do not see grounds for rejection because the proposed experiments are feasible and the practical question is well motivated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this paper asks a question the NCV literature skipped—should you keep autocorrelated MCMC configurations for training neural control variates instead of thinning to decorrelated samples?—and gives the first systematic empirical look at it. The practical answer it reaches is probably right: with a fixed chain length, using all configurations for training reduces estimator variance more than using a thinned set. That matters for lattice QCD and other MCMC-heavy calculations where independent samples are scarce.\n\nWhat is genuinely good: the control-variate construction is standard and correct, the test models (U(1) gauge theory and scalar field theory) are sensible and exhibit real signal-to-noise problems, and the unbiasedness of the final estimator is untouched by the training-data choice. The author also honestly flags in the discussion that the optimal measurement frequency was not investigated, which is exactly the right open question.\n\nThe soft spot is exactly what the stress-test note says, and it is not minor. Every comparison confounds sample count with autocorrelation: measuring every sweep yields 10^4 correlated training samples, while thinning to decorrelated samples yields fewer. So the observed improvement could simply be the well-known benefit of more training data and more gradient steps. The Section III claim that correlated configurations encode 'structural information' is an appealing hypothesis, but no experiment isolates it. To support it you would need to hold the number of training samples fixed while varying autocorrelation—for example, compare 10^4 thinned samples against 10^4 correlated samples, or equal gradient budgets. The paper's own wording in Section IV.A ('more samples while being highly correlated') shows the two variables moving together. I would also want seed averaging and error bars on the training curves; the figures show single trajectories, so it is hard to tell whether the differences are significant. The unusually high acceptance rates (0.9–0.95) are another uncontrolled choice that could affect the comparison.\n\nWho gets value from this: lattice practitioners looking for cheap ways to improve NCV training from existing ensembles. The practical hint is likely sound even if the mechanism is unproven. But as a conceptual paper it overreaches. The fix is straightforward: add a control experiment that separates sample count from autocorrelation, report seed statistics, and soften or substantiate the structural-information mechanism.\n\nRecommendation: send it to peer review, but require the control experiment in the revision. The question deserves referee time, and the experimental design can be repaired.","headline":"Worth a referee's time: the practical tip to reuse correlated MCMC samples for NCV training is likely sound, but the paper's central mechanism claim is not isolated from a sample-count confound.","tokens_in":9798,"tokens_out":3121,"would_cite":false,"duration_ms":30759,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training neural control variates on correlated MCMC configurations reduces estimator variance even when the Monte Carlo chain length is fixed.","keywords":["neural control variates","variance reduction","Markov chain Monte Carlo","autocorrelated configurations","lattice field theory","U(1) gauge theory","scalar field theory","Monte Carlo training data"],"falsifier":"Run two NCV training experiments matched in architecture, learning rate, optimizer settings, number of training samples, and number of gradient steps: one trained on consecutive configurations from a short chain, and one trained on every $k$-th configuration from a chain long enough that the sample counts match. If the variance reductions are equal, the claim that autocorrelation itself carries useful structural information is falsified; if the consecutive-configuration training wins, the claim survives.","tokens_in":8861,"feed_emoji":"📉","tokens_out":11008,"duration_ms":98142,"temperature":0.7,"pith_summary":"The paper argues that Markov-chain configurations usually discarded as autocorrelated are not wasted for training neural control variates. With a fixed-length Monte Carlo chain, feeding the network the closely spaced consecutive configurations can produce a lower-variance estimator than training only on a sparse set of decorrelated samples. The proposed reason is that local Metropolis updates make successive configurations trace smooth paths through configuration space, and those paths expose structure in the distribution and in the observable's fluctuations. The claim is demonstrated on Wilson loops in $U(1)$ gauge theory and on two-point correlators in two-, three-, and four-dimensional scalar field theory, with training on $10^4$ correlated samples and testing on $10^3$ decorrelated samples. If correct, large lattice simulations with only about a thousand independent configurations can repurpose the many correlated configurations in between to build better control variates at no extra sampling cost.","feed_headline":"Correlated MCMC samples cut estimator variance at no extra cost","feed_subtitle":"The autocorrelated configurations normally discarded improve neural control variate training in lattice field theory.","key_machinery":"The central object is a Markov chain run with an unusually high Metropolis acceptance rate (0.9–0.95), so consecutive configurations are strongly correlated and trace smooth trajectories in configuration space. The control variate is built from Stein's identity as $f=\\sum_i(\\partial g_i/\\partial\\phi_i - g_i\\,\\partial S/\\partial\\phi_i)$, with the vector-valued function $g$ realized by a neural network and trained by minimizing $L(w,\\mu)=\\langle (O-f-\\mu)^2\\rangle$ on minibatches drawn from the correlated chain. The integrated autocorrelation time $\\tau_{\\mathrm{int}}=1+2\\sum_{i=1}^{\\infty}\\mathrm{Cov}(\\theta_k,\\theta_{k+i})/\\sigma^2$ is the dial that sets how correlated the training samples are. This machinery works because more closely spaced samples give the optimizer more examples of the observable's fluctuation patterns, while unbiasedness is preserved by evaluating the final mean and error on a separate decorrelated test set.","core_discovery":"The central claim is that using correlated samples to train neural control variates can lead to more effective variance reduction, even when the total length of the Monte Carlo chain is fixed. The control variate under study is $f=\\sum_i(\\partial g_i/\\partial\\phi_i - g_i\\,\\partial S/\\partial\\phi_i)$ obtained from Stein's identity, with $g$ parametrized by a neural network and trained by minimizing $L(w,\\mu)=\\langle (O-f-\\mu)^2\\rangle$. Correlated samples are statistically redundant for error estimation but are argued to be informative for training: correlations arise from smooth local transitions in configuration space, which carry structural information about the distribution, and the improved estimator $\\tilde O=O-f$ can have an autocorrelation time shorter than that of $O$, making samples called correlated for $O$ effectively independent for $\\tilde O$. Numerically, training on $10^4$ correlated samples yields better variance reduction than decorrelated training for $U(1)$ Wilson loops of area $4$ and $16$ and for scalar field theory correlators in two, three, and four dimensions.","pith_inferences":["The paper's comparisons do not separate the number of training samples from their autocorrelation: for a fixed chain length, the correlated training sets also contain more configurations and more gradient updates, so the benefit attributed to correlation could in part be a sample-count effect.","A matched experiment that thins the correlated chain to the same number of training samples and gradient steps as the decorrelated chain would test the paper's mechanism directly; the paper reports no such control.","The unusually high acceptance rates (0.9–0.95) make the correlated configurations especially informative; production chains with lower acceptance may show a smaller benefit, so the practical gain is likely regime-dependent.","The paper's autocorrelation-time argument suggests a tuning rule the author leaves untested: thin the training data by the autocorrelation time of $\\tilde O$, not of $O$, to recover even more usable samples."],"forward_implications":["Training neural control variates no longer requires generating extra independent configurations; the correlated configurations already produced in a fixed-length production chain can supply the training data.","The standard rule that NCV training data must be decorrelated should be relaxed, with autocorrelation properties evaluated per observable: samples considered correlated for $O$ may be nearly independent for $\\tilde O=O-f$.","In large-scale lattice simulations that possess only about $10^3$ independent configurations, the intervening correlated configurations become a cheap resource for variance reduction.","The paper identifies correlated-sample training as an open avenue for other neural-network-based Monte Carlo methods, including contour deformation, neural network quantum states, and self-learning Monte Carlo."],"supporting_citations":[{"why":"Supplies the lattice-field-theory neural control variate construction and the translation-covariant ansatz the paper extends.","marker":"[3]"},{"why":"Introduces neural control variates for Monte Carlo variance reduction, including the loss-function formulation used here.","marker":"[4]"},{"why":"Provides theoretical guarantees for neural control variates in MCMC, the framework whose training data is under study.","marker":"[10]"},{"why":"Gives the stochastic-optimization construction of control variates and the universal form motivating the general ansatz of Eq. (4).","marker":"[12]"},{"why":"Represents the standard practice of using decorrelated configurations for lattice-field-theory control variates that this paper challenges.","marker":"[17]"},{"why":"Supplies the $U(1)$ gauge theory in plaquette variables and the Wilson-loop signal-to-noise setup used for the first numerical test.","marker":"[19]"},{"why":"Shows that a perfect control variate exists for the $U(1)$ model, providing the reference point for the general neural-network construction.","marker":"[20]"}],"fun_headline_variants":["Use normally discarded MCMC samples for better control variates","Correlated samples cut variance in neural control variate training","Discarded autocorrelated data improves neural control variates","Training on correlated samples yields better control variates","Reuse redundant MCMC samples for faster variance reduction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The experiments assume the improvement comes from the correlations between successive configurations, but every correlated training set in the paper also contains more samples and more gradient updates than the decorrelated one, so the benefit could come from having more data rather than from autocorrelation structure.","fun_headline_variants_meta":{"raw":{"variants":["Use normally discarded MCMC samples for better control variates","Correlated samples cut variance in neural control variate training","Discarded autocorrelated data improves neural control variates","Training on correlated samples yields better control variates","Reuse redundant MCMC samples for faster variance reduction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000598,"raw_usage":{"total_tokens":2804,"prompt_tokens":963,"completion_tokens":1841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":1763}},"tokens_in":579,"tokens_out":1841,"duration_ms":11314,"temperature":1.0,"reasoning_tokens":1763,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:10:04.044304+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run two NCV training experiments matched in architecture, learning rate, optimizer settings, number of training samples, and number of gradient steps: one trained on consecutive configurations from a short chain, and one trained on every $k$-th configuration from a chain long enough that the sample counts match. If the variance reductions are equal, the claim that autocorrelation itself carries useful structural information is falsified; if the consecutive-configuration training wins, the claim survives.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces neural control variates for Monte Carlo variance reduction, including the loss-function formulation used here."},{"cited_title":"Belomestny, A","cited_arxiv_id":null,"evidence_quote":"Provides theoretical guarantees for neural control variates in MCMC, the framework whose training data is under study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the stochastic-optimization construction of control variates and the universal form motivating the general ansatz of Eq. (4)."},{"cited_title":"Control variates with neural networks","cited_arxiv_id":"2501.14614","evidence_quote":"Shows that a perfect control variate exists for the $U(1)$ model, providing the reference point for the general neural-network construction."}],"review_version":1}