REVIEW 2 major objections 5 minor 1 cited by
Path optimization method for the sign problem caused by fermion determinant
T0 review · 2 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper claims that machine-learning path optimization, which deforms the integration contour of the complexified auxiliary field into the complex plane, reduces the fermion-determinant sign problem in the one-dimensional massive…
desk verdict Useful numerics on the Thirring model, but the compact-field contour deformation lacks a boundary condition, leaving the central equivalence unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the deformed integration contour: a neural network with one hidden layer (64 tanh units) maps each real auxiliary-field configuration $v_R$ to an imaginary part $v_I$, so the modified contour is a continuous connected surface in complexified field space. Cauchy's integral theorem guarantees the partition function is unchanged by the deformation. Training minimizes the average-phase-factor cost function $$F=\frac{1}{2}\int dv_R\, |$e^{{i\theta(v_R)}}$-$e^{{i\theta_0}}$|^2\,|J(v_R)$e^{{-S(v')}}$|=|Z|\left(\langle $e^{{i\theta}}$\rangle_{pq}^{-1}-1\right),$$ with $\theta=\arg(e^{-S+\ln J})$, using Hybrid Monte Carlo configurations and backpropagation; phase reweighting then estimates observables. The exact fermion determinant of the model is $\det D=\frac{1}{2^{L-1}}[\cosh(L\hat\mu+i\sum_n A_n)+\cosh(L\hat m)]$, and the closed-form analytic results (5) and (6) provide the benchmark. The identity-Jacobian approximation, $J\to 1$ during learning, is the cost-reduction device whose validity is tested.
What would settle it
Repeat the $L=16$, $\beta=1$ calculation at $\mu=1.75$ and at $L=32$ with an independent method that explicitly sums over all contributing steepest-descent regions; if the summed expectation values disagree with the deformed-path condensate and number density beyond the quoted errors, the single connected contour missed a contributing region and the analytic agreement there would be accidental. A cheaper check is to add parallel tempering to the same path optimization and see whether the phase histogram at $\mu=1.75$ sharpens and whether the observables shift.
Extended reading notes
Core claim
On the original integration path, the complex weight from the fermion determinant makes the average phase factor nearly vanish, and the fermion condensate and number density carry huge errors; at $\beta=1$ the condensate does not even agree with the analytic curve. Once the integral is deformed along a single-hidden-layer neural-network contour, the average phase factor is enhanced, the phase histograms become localized, and phase-reweighted expectation values of both observables follow the analytic formulas (5) and (6) at $\beta=1,2$ on $L=16$ lattices. The paper additionally shows that training with the Jacobian set to the identity gives expectation values consistent with full-Jacobian training, so the $O(N^3)$ Jacobian can be omitted from the learning loop. The modified-path phase histogram at $\mu=1.75$ has a less distinct peak, which the authors interpret as several thimbles contributing, and they name parallel tempering as a possible future refinement.
Load-bearing premise
The load-bearing premise is that the neural-network contour is a single connected surface and that Hybrid Monte Carlo sampling on it reaches every region that contributes to the integral; the paper itself notes at $\mu=1.75$ that several steepest-descent regions contribute and that parallel tempering may be needed for ergodicity, but no such safeguard is used in the present runs.
Editorial extensions
If this is right
- On the modified path, the fermion condensate and number density at $L=16$, $\beta=1$ and $2$ match the analytic formulas (5) and (6) with small jackknife errors, while the original path gives huge errors and, at $\beta=1$, disagreement with the analytic condensate.
- The average phase factor is enhanced on the deformed path, and its volume scaling remains exponential, $\text{APF}\sim e^{-\alpha V}$, but with a smaller exponent $\alpha$ than on the original path; the paper stresses this does not fully solve the curse of dimensionality.
- Training with the Jacobian replaced by the identity gives expectation values consistent with full-Jacobian training, so the $O(N^3)$ Jacobian can be omitted from the learning loop; the paper suggests pre-training with the approximation and full learning afterwards when Jacobian effects are not weak.
- At $\mu=1.75$ the modified-path phase histogram has a less distinct peak, indicating several thimbles contribute; the paper identifies combining path optimization with parallel tempering as a possible improvement.
- The method reproduces analytic results in a determinant-origin sign problem, and the paper states the next step is applying the same approach to a more QCD-like theory.
Reading between the lines
- If identity-Jacobian training stays unbiased in more QCD-like models, the Jacobian becomes a measurement-stage cost only, so the practical bottleneck shifts to the HMC sampling itself; the paper's own next-model proposal could be tested immediately with this approximation.
- The single connected neural contour is the piece most likely to fail at larger chemical potential, precisely where the paper's phase histogram at $\mu=1.75$ loses its sharp peak; a natural extension is to compare one-contour training with a multi-contour or tempered variant on the same model before going to QCD-like theories.
- Because exact answers are known in this model, it can be used to separate how much of the improvement comes from the contour deformation itself and how much from phase reweighting and the jackknife binning, by running the same reweighting on the undeformed path with matched statistics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper applies the machine-learning path optimization method (POM) to the one-dimensional massive lattice Thirring model at finite chemical potential, where the sign problem originates from the fermion determinant. The auxiliary field is complexified and its imaginary part is represented by a single-hidden-layer tanh neural network; the network is trained self-supervisedly with a cost function related to the average phase factor, and observables are computed by reweighting on the deformed path. Numerical results for the fermion condensate and number density are compared with known analytic expressions, and the authors also test an approximation in which the Jacobian is replaced by the identity during the learning step. The central claims are that POM reduces statistical errors and reproduces the analytic results, and that the identity-Jacobian approximation gives consistent results.
Significance. If the central claims hold, the paper provides a useful benchmark: it demonstrates that POM can handle a sign problem of the same origin as in QCD in a model with analytic control, and it proposes a cheap Jacobian approximation that could reduce the cost of learning in more realistic theories. The model choice is appropriate, the analytic results from Ref. [28] are used only as a benchmark and not as training input, the cost function in Eq. (12) is explicitly target-independent, and the APF scaling data in Fig. 10 are a useful diagnostic. The paper does not provide code or data, so the numerical claims are not independently checkable in this form, but the clean benchmark makes them checkable in principle. The main mathematical gap concerns the boundary behavior of the deformed contour for the compact auxiliary field, which is not addressed anywhere in the manuscript; this must be resolved before the numerical agreement can be taken as evidence for the method.
major comments (2)
- [Sec. II A and II B, Eqs. (7)-(13)] The original integration domain for the compact auxiliary field A_n is a torus: Sec. II A states that the cosine in S_B makes A_n compact, so the integral in Eq. (3) runs over one period in each direction. The deformed contour in Sec. II B is the graph v' = v_R + i v_I(v_R), with v_I the output of a tanh network. For Cauchy's theorem to guarantee equality of the deformed integral with the original one, the graph over the closed hypercube must have a vanishing boundary contribution: either v_I must vanish on the boundary of [-π,π]^L or v_I must be 2π-periodic in each input direction. Neither condition is stated, imposed, or verified. The cost function in Eq. (12) integrates only over the interior v_R and contains no boundary term, so nothing in the training enforces the homology condition. Consequently, the reweighting formula Eq. (13) can be biased even with infinite statistics if the learned v_I takes different values on opposite faces. The agreement with the analytic results in Sec. IV cannot rule this out unless the boundary contribution is explicitly checked. I ask the authors to impose a periodic output layer or a boundary-pinning term, or to demonstrate numerically and analytically that the learned path has a vanishing boundary contribution.
- [Sec. IV, Fig. 7 and µ=1.75 discussion] The authors state that at µ=1.75 the histogram on the modified path has a less clear peak and that several thimbles contribute, and they propose parallel tempering as future work. Since the HMC sampling is performed on a single connected manifold parameterized by the real parts of the fields, a multi-modal effective weight can make the chain non-ergodic over the contributing regions. If one contributing thimble is missed, the agreement with the analytic curves at µ=1.75 would be accidental rather than a demonstration of the method. The paper should either present evidence that the sampling covers all relevant regions at this µ (for example, multiple chains with different initial conditions, replica-exchange moves, or a comparison of histograms from independent runs), or explicitly exclude µ=1.75 from the claim of reproducing the analytic results.
minor comments (5)
- [Sec. IV, Figs. 4-7] The text and several captions use 'AFP' where 'APF' is intended; please correct this typo consistently.
- [Fig. 10] The scaling law APF ∼ e^{-αV} is claimed from only three lattice sizes (L=4, 8, 16) and no fitted values of α are reported; please provide the fitted exponents with uncertainties so that the improvement on the modified path is quantitative.
- [Sec. III and Appendix A] The numerical setup omits several details needed for reproducibility: the HMC trajectory length and step size, the number of thermalization sweeps, the AdamW learning rate, the initialization of the network weights, and the random seed. Please report these quantities, and also state how many independent Jackknife blocks remain after binning with bin size 50 on 1000 configurations.
- [Sec. II B, cost comparison paragraph] The sentence stating that the worldvolume approach reduces the cost from O(N_d^3) to O(N_d^{1-2}) is ambiguous and appears to contain a typo; the intended scaling should be stated clearly.
- [General] No code or data are provided; given that this is a machine-learning-based numerical study, making the trained network parameters and raw histograms available would materially help readers verify the central claims.
Circularity Check
No circularity found: the analytic benchmark and cost function are independent of the observables.
full rationale
The paper's central claim is that the path optimization method reduces the sign problem and reproduces analytic results for the one-dimensional Thirring model. The analytic results for the fermion condensate and number density are taken from an independent reference, Ref. [28], and are used only as a benchmark for comparison in Sec. IV; they are not used as input to the training. The cost function in Eq. (12) depends on the average phase factor, i.e., on the phase of the complex Boltzmann weight and the Jacobian, and not on the physical observables s or n. Therefore, agreement with the analytic curves is not enforced by construction. The reweighting formula in Eq. (13) is the standard phase reweighting and is applied equally to original and modified paths. The Jacobian approximation discussed in Sec. IV is validated by comparing approximate-learning results against full-Jacobian results, both of which are then compared with the independent analytic results. The paper cites earlier work by the same authors for the path optimization method and related techniques, but these citations establish the method and specific algorithmic choices; they do not supply the target result or forbid alternatives. The noted concerns about boundary conditions and ergodicity are correctness or rigor issues, not circularity, because they do not involve fitting the observables or deriving the conclusion from its own premise. Overall, the derivation chain is self-contained with respect to the claim tested: the observables are measured after training with a cost function that is blind to their values.
Assumptions & free parameters
free parameters (2)
- Neural network architecture (hidden units, depth) =
64 units, single hidden layer
- Training schedule (learning steps, batches per step, batch size) =
100 steps, 30 updates per step, batch size 64
assumptions (5)
- standard math Cauchy's integral theorem applies to the deformed integration path, and there are no contributions from infinity.
- standard math A neural network with one hidden layer can represent a sufficient path deformation.
- domain assumption HMC sampling on the deformed path is ergodic over all relevant thimbles.
- domain assumption The analytic determinant and observable formulas (Eqs. 4-6) from Ref. [28] are correct.
- domain assumption The partition function is real, so the phase offset θ0 = 0.
Cite this review
Pith. "Pith review of Path optimization method for the sign problem caused by fermion determinant." pith.science (2026). https://pith.science/paper/V4SXKMW3
@misc{pith2026250202804,
author = {Pith},
title = {Pith review of: Path optimization method for the sign problem caused by fermion determinant},
year = {2026},
howpublished = {\url{https://pith.science/paper/V4SXKMW3}},
note = {Machine review of arXiv:2502.02804}
}
read the original abstract
The path optimization method with machine learning is applied to the one-dimensional massive lattice Thirring model, which has the sign problem caused by the fermion determinant. This study aims to investigate how the path optimization method works for the sign problem. We show that the path optimization method successfully reduces statistical errors and reproduces the analytic results. We also examine an approximation of the Jacobian calculation in the learning process and show that it gives consistent results with those without an approximation.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
Path optimization method for the sign problem: Insights from random matrix models
Path optimization improves the average phase factor in the Stephanov model at high chemical potential but not at low chemical potential or in the chiral random matrix model, pointing to the global sign problem as the ...
Reference graph
Works this paper leans on
- [28]
-
[1]
de Forcrand, PoS LA T2009, 010 (2009), arXiv:1005.0539 [hep-lat]
P. de Forcrand, PoS LA T2009, 010 (2009), arXiv:1005.0539 [hep-lat]
arXiv 2009
-
[2]
A. Alexandru, G. Basar, P. F. Bedaque, and N. C. Warrington, Rev. Mod. Phys. 94, 015006 (2022), arXiv:2007.05436 [hep-lat]
arXiv 2022
-
[3]
Nagata, 素 粒 子 論 研 究 (Soryusironkenkyu) 31, 1 (2020); Prog
K. Nagata, 素 粒 子 論 研 究 (Soryusironkenkyu) 31, 1 (2020); Prog. Part. Nucl. Phys. 127, 103991 (2022), arXiv:2108.12423 [hep-lat]
arXiv 2020
-
[4]
Y. Mori, K. Kashiwa, and A. Ohnishi, Phys. Rev. D96, 111501 (2017), arXiv:1705.05605 [hep-lat]
work page Pith review arXiv 2017
-
[5]
Y. Mori, K. Kashiwa, and A. Ohnishi, PTEP 2018, 023B04 (2018), arXiv:1709.03208 [hep-lat]
arXiv 2018
-
[6]
A. Alexandru, P. F. Bedaque, H. Lamm, and S. Lawrence, Phys. Rev. D 97, 094510 (2018), arXiv:1804.00697 [hep-lat]
arXiv 2018
-
[7]
E. Witten, AMS/IP Stud. Adv. Math. 50, 347 (2011), arXiv:1001.2933 [hep-th]
arXiv 2011
Show all 61 references
-
[8]
Cristoforetti, F
M. Cristoforetti, F. Di Renzo, and L. Scorzato (Aurora- Science Collaboration), Phys.Rev. D86, 074506 (2012), arXiv:1205.3996 [hep-lat]
2012 arXiv
-
[9]
Fujii, D
H. Fujii, D. Honda, M. Kato, Y. Kikukawa, S. Komatsu, and T. Sano, JHEP 1310, 147 (2013), arXiv:1309.4371 [hep-lat]
2013 arXiv
-
[10]
Lawrence and Y
S. Lawrence and Y. Yamauchi, Phys. Rev. D110, 014508 (2024), arXiv:2311.13002 [hep-lat]
2024 arXiv
-
[11]
Kashiwa, Y
K. Kashiwa, Y. Mori, and A. Ohnishi, Phys. Rev. D 99, 014033 (2019), arXiv:1805.08940 [hep-ph]
2019 arXiv
- [12]
-
[13]
Kashiwa, Y
K. Kashiwa, Y. Mori, and A. Ohnishi, Phys. Rev. D 99, 114005 (2019), arXiv:1903.03679 [hep-lat]
2019 arXiv
-
[14]
Y. Mori, K. Kashiwa, and A. Ohnishi, PTEP 2019, 113B01 (2019), arXiv:1904.11140 [hep-lat]
2019 arXiv
-
[15]
Kashiwa and Y
K. Kashiwa and Y. Mori, Phys. Rev. D 102, 054519 (2020), arXiv:2007.04167 [hep-lat]
2020 arXiv
-
[16]
Namekawa, K
Y. Namekawa, K. Kashiwa, A. Ohnishi, and H. Takase, Phys. Rev. D 105, 034502 (2022), arXiv:2109.11710 [hep- lat]
2022 arXiv
-
[17]
Namekawa, K
Y. Namekawa, K. Kashiwa, H. Matsuda, A. Ohnishi, and H. Takase, Phys. Rev. D 107, 034509 (2023), arXiv:2210.05402 [hep-lat]
2023 arXiv
-
[18]
Giordano, K
M. Giordano, K. Kapas, S. D. Katz, A. Pasztor, and Z. Tulipant, Phys. Rev. D 106, 054512 (2022), arXiv:2202.07561 [hep-lat]
2022 arXiv
-
[19]
Rodekamp, E
M. Rodekamp, E. Berkowitz, C. G¨ antgen, S. Krieg, T. Luu, and J. Ostmeyer, Phys. Rev. B 106, 125139 8 (2022), arXiv:2203.00390 [physics.comp-ph]
2022 arXiv
-
[20]
Rodekamp, E
M. Rodekamp, E. Berkowitz, M. Dinc˘ a, C. G¨ antgen, S. Krieg, and T. Luu, PoS LA TTICE2023, 031 (2024), arXiv:2311.18312 [cond-mat.str-el]
2024 arXiv
-
[21]
Kanwar, A
G. Kanwar, A. Lovato, N. Rocco, and M. Wagman, Phys. Rev. C 109, 034317 (2024), arXiv:2304.03229 [nucl-th]
2024 arXiv
-
[22]
Y. Lin, W. Detmold, G. Kanwar, P. E. Shanahan, and M. L. Wagman, PoS LA TTICE2023, 043 (2024), arXiv:2309.00600 [hep-lat]
2024 arXiv
-
[23]
Detmold, G
W. Detmold, G. Kanwar, M. L. Wagman, and N. C. Warrington, Phys. Rev. D 102, 014514 (2020), arXiv:2003.05914 [hep-lat]
2020 arXiv
-
[24]
Detmold, G
W. Detmold, G. Kanwar, H. Lamm, M. L. Wagman, and N. C. Warrington, Phys. Rev. D 103, 094517 (2021), arXiv:2101.12668 [hep-lat]
2021 arXiv
-
[25]
P. F. Bedaque and H. Oh, Phys. Rev. D 109, 094519 (2024), arXiv:2312.08228 [hep-lat]
2024 arXiv
-
[26]
W. E. Thirring, Annals of Physics 3, 91 (1958)
1958
-
[27]
Fujii, S
H. Fujii, S. Kamata, and Y. Kikukawa, JHEP 11, 078 (2015), [Erratum: JHEP 02, 036 (2016)], arXiv:1509.08176 [hep-lat]
2015 arXiv
-
[29]
Alexandru, G
A. Alexandru, G. Basar, and P. Bedaque, Phys. Rev. D 93, 014504 (2016), arXiv:1510.03258 [hep-lat]
2016 arXiv
-
[30]
Alexandru, G
A. Alexandru, G. Basar, P. F. Bedaque, G. W. Ridg- way, and N. C. Warrington, JHEP 05, 053 (2016), arXiv:1512.08764 [hep-lat]
2016 arXiv
-
[31]
Fukuma and N
M. Fukuma and N. Umeda, PTEP 2017, 073B01 (2017), arXiv:1703.00861 [hep-lat]
2017 arXiv
-
[32]
Di Renzo and K
F. Di Renzo and K. Zambello, Phys. Rev. D 105, 054501 (2022), arXiv:2109.02511 [hep-lat]
2022 arXiv
-
[33]
Alexandru, P
A. Alexandru, P. F. Bedaque, H. Lamm, S. Lawrence, and N. C. Warrington, Phys. Rev. Lett. 121, 191602 (2018), arXiv:1808.09799 [hep-lat]
2018 arXiv
-
[34]
Lawrence and Y
S. Lawrence and Y. Yamauchi, Phys. Rev. D107, 114505 (2023), arXiv:2212.14606 [hep-lat]
2023 arXiv
-
[35]
M. A. Stephanov, Phys. Rev. Lett. 76, 4472 (1996), arXiv:hep-lat/9604003
1996 arXiv
-
[36]
A. M. Halasz, A. D. Jackson, R. E. Shrock, M. A. Stephanov, and J. J. M. Verbaarschot, Phys. Rev. D 58, 096007 (1998), arXiv:hep-ph/9804290
1998 arXiv
-
[37]
Bloch, J
J. Bloch, J. Glesaaen, J. J. M. Verbaarschot, and S. Zafeiropoulos, JHEP 03, 015 (2018), arXiv:1712.07514 [hep-lat]
2018 arXiv
-
[38]
Fukuma and N
M. Fukuma and N. Matsumoto, PTEP 2021, 023B08 (2021), arXiv:2012.08468 [hep-lat]
2021 arXiv
-
[39]
Giordano, A
M. Giordano, A. Pasztor, D. Pesznyak, and Z. Tulipant, Phys. Rev. D 108, 094507 (2023), arXiv:2301.12947 [hep- lat]
2023 arXiv
-
[40]
Alexandru, P
A. Alexandru, P. F. Bedaque, H. Lamm, and S. Lawrence, Phys. Rev. D96, 094505 (2017), arXiv:1709.01971 [hep-lat]
2017 arXiv
-
[41]
J. M. Pawlowski and C. Zielinski, Phys. Rev. D 87, 094503 (2013), arXiv:1302.1622 [hep-lat]
2013 arXiv
-
[42]
W. S. McCulloch and W. Pitts, The bulletin of mathe- matical biophysics 5, 115 (1943)
1943
-
[43]
The organization of behavior: A neuropsy- chological theory,
D. O. Hebb, “The organization of behavior: A neuropsy- chological theory,” (2005), psychology press
2005
-
[44]
Rosenblatt, Psychological review 65, 386 (1958)
F. Rosenblatt, Psychological review 65, 386 (1958)
1958
-
[45]
G. E. Hinton and R. R. Salakhutdinov, science 313, 504 (2006)
2006
-
[46]
Alexandru, G
A. Alexandru, G. Basar, P. F. Bedaque, G. W. Ridgway, and N. C. Warrington, Phys. Rev. D 93, 094514 (2016), arXiv:1604.00956 [hep-lat]
2016 arXiv
-
[47]
Alexandru, G
A. Alexandru, G. Basar, P. F. Bedaque, and G. W. Ridg- way, Phys. Rev. D 95, 114501 (2017), arXiv:1704.06404 [hep-lat]
2017 arXiv
-
[48]
Cybenko, Mathematics of Control, Signals, and Sys- tems (MCSS) 2, 303 (1989)
G. Cybenko, Mathematics of Control, Signals, and Sys- tems (MCSS) 2, 303 (1989)
1989
-
[49]
Hornik, Neural networks 4, 251 (1991)
K. Hornik, Neural networks 4, 251 (1991)
1991
-
[50]
R. H. Swendsen and J.-S. Wang, Physical review letters 57, 2607 (1986)
1986
-
[51]
C. J. Geyer, Computing science and statistics: Proceedings of 23rd Symposium on the Interface Interface Foundation, Fairfax Station, 1991, , 156 (1991)
1991
-
[52]
Hukushima and K
K. Hukushima and K. Nemoto, Journal of the Physical Society of Japan 65, 1604 (1996)
1996
-
[53]
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, nature 323, 533 (1986)
1986
-
[54]
Duane, A
S. Duane, A. D. Kennedy, B. J. Pendleton, and D. Roweth, Phys. Lett. B 195, 216 (1987)
1987
-
[55]
A. M. Ferrenberg and R. H. Swendsen, Phys. Rev. Lett. 61, 2635 (1988)
1988
-
[56]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al., Advances in neural information processing systems 32 (2019)
2019
-
[57]
Loshchilov and F
I. Loshchilov and F. Hutter, in International Conference on Learning Representations (2019) arXiv:1711.05101 [cs.LG]
2019 arXiv
-
[58]
Bottou, Online learning in neural networks (1998)
L. Bottou, Online learning in neural networks (1998)
1998
-
[59]
Fukuma, N
M. Fukuma, N. Matsumoto, and Y. Namekawa, PTEP 2021, 123B02 (2021), arXiv:2107.06858 [hep-lat]
2021 arXiv
-
[60]
Fukuma, PTEP 2024, 053B02 (2024), arXiv:2311.10663 [hep-lat]
M. Fukuma, PTEP 2024, 053B02 (2024), arXiv:2311.10663 [hep-lat]
2024 arXiv
-
[61]
Splittorff and J
K. Splittorff and J. J. M. Verbaarschot, Phys. Rev. D 75, 116003 (2007), arXiv:hep-lat/0702011
2007 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.