{"id":"3da1edb9-8428-490b-be5f-4b49e72001d5","arxiv_id":"2502.02804","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Path optimization with machine learning reproduces analytic results in the 1D lattice Thirring model, and dropping the Jacobian from the learning step still works.","lead":"Researchers applied the path optimization method with a neural network to the one-dimensional lattice Thirring model, a testbed for the sign problem caused by the fermion determinant. The deformed integration path reduced statistical errors and reproduced the analytic results, and a cheap Jacobian approximation gave consistent answers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The deformed contour's boundary conditions for the compact auxiliary field A_n are never imposed or checked; without them Cauchy's theorem does not guarantee the deformed path integral equals the original.","rationale":"The reader's weakest assumption was HMC ergodicity on the deformed path, specifically the risk of missing thimbles at µ=1.75. I agree that is a real concern, but the more load-bearing condition is that the deformed contour be homologous to the original compact integration domain. The paper complexifies the compact field A_n and defines the deformed path as the graph of a neural network over the real parts, yet never states or enforces boundary conditions on v_I at the edges of the compact interval. The cost function Eq. (12) and the observable reweighting Eq. (13) both assume the deformed integral equals Z; if the boundary contributions do not vanish, that assumption is false and the method is not computing the original partition function. This is a sharp, concrete gap that can be settled by the proposed boundary-integral check. The Jacobian approximation and the agreement with analytic results are useful empirical evidence, but they do not address this exact issue. The appropriate verdict remains CONDITIONAL, with the condition being a demonstration — or an explicit imposition — of the boundary/homology condition, not merely the ergodicity clarification requested by the reader. The reader and I partially agree because both concerns involve the deformed path failing to represent the full integral, but the boundary issue is logically prior and more severe.","tokens_in":9845,"tokens_out":12295,"duration_ms":128028,"concrete_test":"For a trained network at a strongly sign-problematic point (e.g., β=1, µ=1.75, L=16), compute the boundary term B = Σ_j [∫_{x_j=π} J e^{-S} d^{L-1}x − ∫_{x_j=−π} J e^{-S} d^{L-1}x] on the deformed contour, or, more directly, evaluate ∫_deformed J e^{-S} dv_R by deterministic quadrature on a smaller lattice (L=4 or L=8) and compare with the analytic Z from Eqs. (3)–(4). If |B|/|Z| is non-negligible (beyond statistical error), the deformed path is not homologous and the central claim fails. A complementary check: estimate Im⟨e^{iθ}⟩ on the learned path; for real Z it must be zero, so a nonzero imaginary part indicates boundary contamination.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires the deformed integration path to be homologous to the original contour. In Sec. II A the model is defined for a compact field A_n (S_B = β Σ (1-cos A_n), with the text explicitly saying the cosine makes A_n compact), so the original integral runs over one period, e.g., A_n ∈ [-π,π]^L. In Sec. II B the deformed path is v' = v_R + i v_I(v_R), where v_R is the real part of the compact variable and v_I is the output of a tanh neural network. For Cauchy's theorem to hold, the graph over the closed hypercube must have vanishing boundary contribution: either v_I = 0 on ∂[-π,π]^L or v_I is 2π-periodic in each input. No such boundary condition is stated, imposed, or verified; the cost function in Eq. (12) is an integral over the interior v_R only and contains no boundary term, so nothing in the training enforces homology. If v_I(π) ≠ v_I(-π) for any direction, the boundary integrals do not cancel and ∫_deformed J e^{-S} dv_R ≠ Z. Then the reweighting formula Eq. (13) is biased even with infinite statistics, and the agreement with analytic results in Sec. IV would be accidental or an artifact of small boundary weight. This is more fundamental than the ergodicity concern at µ=1.75, which is a sampling issue; the boundary issue is an exact mathematical gap in the method as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the machine-learning path optimization method (POM) to the one-dimensional massive lattice Thirring model at finite chemical potential, where the sign problem originates from the fermion determinant. The auxiliary field is complexified and its imaginary part is represented by a single-hidden-layer tanh neural network; the network is trained self-supervisedly with a cost function related to the average phase factor, and observables are computed by reweighting on the deformed path. Numerical results for the fermion condensate and number density are compared with known analytic expressions, and the authors also test an approximation in which the Jacobian is replaced by the identity during the learning step. The central claims are that POM reduces statistical errors and reproduces the analytic results, and that the identity-Jacobian approximation gives consistent results.","tokens_in":10095,"tokens_out":7269,"duration_ms":79592,"significance":"If the central claims hold, the paper provides a useful benchmark: it demonstrates that POM can handle a sign problem of the same origin as in QCD in a model with analytic control, and it proposes a cheap Jacobian approximation that could reduce the cost of learning in more realistic theories. The model choice is appropriate, the analytic results from Ref. [28] are used only as a benchmark and not as training input, the cost function in Eq. (12) is explicitly target-independent, and the APF scaling data in Fig. 10 are a useful diagnostic. The paper does not provide code or data, so the numerical claims are not independently checkable in this form, but the clean benchmark makes them checkable in principle. The main mathematical gap concerns the boundary behavior of the deformed contour for the compact auxiliary field, which is not addressed anywhere in the manuscript; this must be resolved before the numerical agreement can be taken as evidence for the method.","major_comments":[{"comment":"The original integration domain for the compact auxiliary field A_n is a torus: Sec. II A states that the cosine in S_B makes A_n compact, so the integral in Eq. (3) runs over one period in each direction. The deformed contour in Sec. II B is the graph v' = v_R + i v_I(v_R), with v_I the output of a tanh network. For Cauchy's theorem to guarantee equality of the deformed integral with the original one, the graph over the closed hypercube must have a vanishing boundary contribution: either v_I must vanish on the boundary of [-π,π]^L or v_I must be 2π-periodic in each input direction. Neither condition is stated, imposed, or verified. The cost function in Eq. (12) integrates only over the interior v_R and contains no boundary term, so nothing in the training enforces the homology condition. Consequently, the reweighting formula Eq. (13) can be biased even with infinite statistics if the learned v_I takes different values on opposite faces. The agreement with the analytic results in Sec. IV cannot rule this out unless the boundary contribution is explicitly checked. I ask the authors to impose a periodic output layer or a boundary-pinning term, or to demonstrate numerically and analytically that the learned path has a vanishing boundary contribution.","section":"Sec. II A and II B, Eqs. (7)-(13)"},{"comment":"The authors state that at µ=1.75 the histogram on the modified path has a less clear peak and that several thimbles contribute, and they propose parallel tempering as future work. Since the HMC sampling is performed on a single connected manifold parameterized by the real parts of the fields, a multi-modal effective weight can make the chain non-ergodic over the contributing regions. If one contributing thimble is missed, the agreement with the analytic curves at µ=1.75 would be accidental rather than a demonstration of the method. The paper should either present evidence that the sampling covers all relevant regions at this µ (for example, multiple chains with different initial conditions, replica-exchange moves, or a comparison of histograms from independent runs), or explicitly exclude µ=1.75 from the claim of reproducing the analytic results.","section":"Sec. IV, Fig. 7 and µ=1.75 discussion"}],"minor_comments":[{"comment":"The text and several captions use 'AFP' where 'APF' is intended; please correct this typo consistently.","section":"Sec. IV, Figs. 4-7"},{"comment":"The scaling law APF ∼ e^{-αV} is claimed from only three lattice sizes (L=4, 8, 16) and no fitted values of α are reported; please provide the fitted exponents with uncertainties so that the improvement on the modified path is quantitative.","section":"Fig. 10"},{"comment":"The numerical setup omits several details needed for reproducibility: the HMC trajectory length and step size, the number of thermalization sweeps, the AdamW learning rate, the initialization of the network weights, and the random seed. Please report these quantities, and also state how many independent Jackknife blocks remain after binning with bin size 50 on 1000 configurations.","section":"Sec. III and Appendix A"},{"comment":"The sentence stating that the worldvolume approach reduces the cost from O(N_d^3) to O(N_d^{1-2}) is ambiguous and appears to contain a typo; the intended scaling should be stated clearly.","section":"Sec. II B, cost comparison paragraph"},{"comment":"No code or data are provided; given that this is a machine-learning-based numerical study, making the trained network parameters and raw histograms available would materially help readers verify the central claims.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the benchmark choice is sound. The largest concern is the unaddressed boundary/homology condition for the compact auxiliary field; this is a correctness issue in the method as presented rather than a stylistic one. The authors' own prior work on POM is heavily cited, which is understandable in a specialist line of research, but the novelty relative to Refs. [14] and [17] should be stated more explicitly. The paper would also benefit from a data-availability statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time if you work on path optimization or sign problems. The genuinely new piece is the application of the established POM to the 1D massive lattice Thirring model, where the sign problem comes from the fermion determinant, and a check that dropping the Jacobian in the learning step (J=1) still gives the right observables. That last point is practically useful, since Jacobian cost dominates training. The numerics look careful: analytic benchmarks, APF improvement, volume scaling, Jackknife errors, and an honest discussion of the restart procedure and the µ=1.75 multi-thimble worry.\n\nThe soft spot that bothers me is more fundamental. The auxiliary field A_n is compact, so the original integral runs over one period, say [-π,π]^L. The deformed path is v_R + i v_I(v_R) with v_I from a tanh network, and the paper never states or enforces periodic boundary conditions on v_I. Without v_I(π)=v_I(-π), the deformed contour has a boundary contribution and Cauchy's theorem doesn't guarantee the deformed integral equals the original. The cost function integrates only over the interior, so training cannot fix this. This is an exact mathematical gap, not a sampling issue. The agreement with analytic results might still hold if the boundary weight is small, but the paper doesn't check.\n\nMinor points: no code/data released; 1000 configurations with bin 50 is modest; the restart rule is a mild selection bias; the QCD extrapolation in Sec. IV is speculative. None of these are load-bearing.\n\nI'd send it to a referee—the topic is relevant and the method is promising—but the referee should put the homology/boundary condition first. If the authors can show the deformed path is periodic or that boundary terms vanish, the paper becomes a solid demonstration. As it stands, I wouldn't cite it.","headline":"Useful numerics on the Thirring model, but the compact-field contour deformation lacks a boundary condition, leaving the central equivalence unproven.","tokens_in":10740,"tokens_out":3395,"would_cite":false,"duration_ms":31633,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that machine-learning path optimization, which deforms the integration contour of the complexified auxiliary field into the complex plane, reduces the fermion-determinant sign problem in the one-dimensional massive…","keywords":["sign problem","path optimization method","lattice Thirring model","fermion determinant","average phase factor","machine learning","chemical potential","Jacobian approximation"],"falsifier":"Repeat the $L=16$, $\\beta=1$ calculation at $\\mu=1.75$ and at $L=32$ with an independent method that explicitly sums over all contributing steepest-descent regions; if the summed expectation values disagree with the deformed-path condensate and number density beyond the quoted errors, the single connected contour missed a contributing region and the analytic agreement there would be accidental. A cheaper check is to add parallel tempering to the same path optimization and see whether the phase histogram at $\\mu=1.75$ sharpens and whether the observables shift.","tokens_in":9560,"feed_emoji":"⚛️","tokens_out":13080,"duration_ms":114664,"temperature":0.7,"pith_summary":"The paper sets out to show that the path optimization method with machine learning can cure the sign problem that comes specifically from a complex fermion determinant. It works in the one-dimensional massive lattice Thirring model, where exact results for the fermion condensate and number density are known, and finds that a neural-network-deformed integration contour raises the average phase factor and reproduces those analytic values with small statistical errors, where the undeformed path has huge errors. It also claims that replacing the Jacobian with the identity during the learning phase does not change the expectation values, which removes the most expensive part of training. The reason to care is that QCD at finite density has the same determinant origin of its sign problem, and this model provides a controlled testing ground for contour-deformation strategies before moving to more QCD-like theories.","feed_headline":"Machine-learned path tames the determinant sign problem","feed_subtitle":"A deformed path cuts phase noise and matches analytic values in a 1D Thirring test.","key_machinery":"The central object is the deformed integration contour: a neural network with one hidden layer (64 tanh units) maps each real auxiliary-field configuration $v_R$ to an imaginary part $v_I$, so the modified contour is a continuous connected surface in complexified field space. Cauchy's integral theorem guarantees the partition function is unchanged by the deformation. Training minimizes the average-phase-factor cost function $$F=\\frac{1}{2}\\int dv_R\\, |$e^{{i\\theta(v_R)}}$-$e^{{i\\theta_0}}$|^2\\,|J(v_R)$e^{{-S(v')}}$|=|Z|\\left(\\langle $e^{{i\\theta}}$\\rangle_{pq}^{-1}-1\\right),$$ with $\\theta=\\arg(e^{-S+\\ln J})$, using Hybrid Monte Carlo configurations and backpropagation; phase reweighting then estimates observables. The exact fermion determinant of the model is $\\det D=\\frac{1}{2^{L-1}}[\\cosh(L\\hat\\mu+i\\sum_n A_n)+\\cosh(L\\hat m)]$, and the closed-form analytic results (5) and (6) provide the benchmark. The identity-Jacobian approximation, $J\\to 1$ during learning, is the cost-reduction device whose validity is tested.","core_discovery":"On the original integration path, the complex weight from the fermion determinant makes the average phase factor nearly vanish, and the fermion condensate and number density carry huge errors; at $\\beta=1$ the condensate does not even agree with the analytic curve. Once the integral is deformed along a single-hidden-layer neural-network contour, the average phase factor is enhanced, the phase histograms become localized, and phase-reweighted expectation values of both observables follow the analytic formulas (5) and (6) at $\\beta=1,2$ on $L=16$ lattices. The paper additionally shows that training with the Jacobian set to the identity gives expectation values consistent with full-Jacobian training, so the $O(N^3)$ Jacobian can be omitted from the learning loop. The modified-path phase histogram at $\\mu=1.75$ has a less distinct peak, which the authors interpret as several thimbles contributing, and they name parallel tempering as a possible future refinement.","pith_inferences":["If identity-Jacobian training stays unbiased in more QCD-like models, the Jacobian becomes a measurement-stage cost only, so the practical bottleneck shifts to the HMC sampling itself; the paper's own next-model proposal could be tested immediately with this approximation.","The single connected neural contour is the piece most likely to fail at larger chemical potential, precisely where the paper's phase histogram at $\\mu=1.75$ loses its sharp peak; a natural extension is to compare one-contour training with a multi-contour or tempered variant on the same model before going to QCD-like theories.","Because exact answers are known in this model, it can be used to separate how much of the improvement comes from the contour deformation itself and how much from phase reweighting and the jackknife binning, by running the same reweighting on the undeformed path with matched statistics."],"forward_implications":["On the modified path, the fermion condensate and number density at $L=16$, $\\beta=1$ and $2$ match the analytic formulas (5) and (6) with small jackknife errors, while the original path gives huge errors and, at $\\beta=1$, disagreement with the analytic condensate.","The average phase factor is enhanced on the deformed path, and its volume scaling remains exponential, $\\text{APF}\\sim e^{-\\alpha V}$, but with a smaller exponent $\\alpha$ than on the original path; the paper stresses this does not fully solve the curse of dimensionality.","Training with the Jacobian replaced by the identity gives expectation values consistent with full-Jacobian training, so the $O(N^3)$ Jacobian can be omitted from the learning loop; the paper suggests pre-training with the approximation and full learning afterwards when Jacobian effects are not weak.","At $\\mu=1.75$ the modified-path phase histogram has a less distinct peak, indicating several thimbles contribute; the paper identifies combining path optimization with parallel tempering as a possible improvement.","The method reproduces analytic results in a determinant-origin sign problem, and the paper states the next step is applying the same approach to a more QCD-like theory."],"supporting_citations":[{"why":"Original formulation of the path optimization method by contour deformation.","marker":"[4]"},{"why":"Introduced machine learning into path optimization and defined the average-phase-factor cost function used here.","marker":"[5]"},{"why":"Earlier path-optimization study of the determinant-induced sign problem in 0+1 dimensional QCD, the direct precedent this work extends.","marker":"[14]"},{"why":"Source of the identity-Jacobian approximation applied during learning.","marker":"[17]"},{"why":"Defines the 1D massive lattice Thirring model and supplies the exact determinant plus the analytic condensate and number density used as benchmarks.","marker":"[28]"},{"why":"Introduced the exponential moving average used to stabilize the cost function and the discussion of parallel tempering for ergodicity.","marker":"[15]"},{"why":"Hybrid Monte Carlo algorithm used to generate configurations for training and for final measurements.","marker":"[54]"}],"fun_headline_variants":["Neural contour quells determinant sign noise","Path learning defeats sign problem in Thirring model","Deformed path solves sign problem, matches exact results","Machine-learned contours fix fermion sign trouble","ML path optimization beats fermion sign problem"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the neural-network contour is a single connected surface and that Hybrid Monte Carlo sampling on it reaches every region that contributes to the integral; the paper itself notes at $\\mu=1.75$ that several steepest-descent regions contribute and that parallel tempering may be needed for ergodicity, but no such safeguard is used in the present runs.","fun_headline_variants_meta":{"raw":{"variants":["Neural contour quells determinant sign noise","Path learning defeats sign problem in Thirring model","Deformed path solves sign problem, matches exact results","Machine-learned contours fix fermion sign trouble","ML path optimization beats fermion sign problem"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2604,"prompt_tokens":802,"completion_tokens":1802,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":418,"completion_tokens_details":{"reasoning_tokens":1732}},"tokens_in":418,"tokens_out":1802,"duration_ms":14704,"temperature":1.0,"reasoning_tokens":1732,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T11:02:11.602669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the $L=16$, $\\beta=1$ calculation at $\\mu=1.75$ and at $L=32$ with an independent method that explicitly sums over all contributing steepest-descent regions; if the summed expectation values disagree with the deformed-path condensate and number density beyond the quoted errors, the single connected contour missed a contributing region and the analytic agreement there would be accidental. A cheaper check is to add parallel tempering to the same path optimization and see whether the phase histogram at $\\mu=1.75$ sharpens and whether the observables shift.","supporting_citations":[{"cited_title":"Toward solving the sign problem with path optimization method","cited_arxiv_id":"1705.05605","evidence_quote":"Original formulation of the path optimization method by contour deformation."},{"cited_title":"Path optimization in 0+1 dimensional QCD at finite density","cited_arxiv_id":"1904.11140","evidence_quote":"Earlier path-optimization study of the determinant-induced sign problem in 0+1 dimensional QCD, the direct precedent this work extends."},{"cited_title":"Path optimization for $U(1)$ gauge theory with complexified parameters","cited_arxiv_id":"2007.04167","evidence_quote":"Introduced the exponential moving average used to stabilize the cost function and the discussion of parallel tempering for ergodicity."}],"review_version":1}