REVIEW 1 major objections 8 minor 1 cited by
A deep neural network approach to solve the Dirac equation
T0 review · 1 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A deep neural network with an inverse-Hamiltonian loss solves the radial Dirac equation for ground and low-lying excited states, bypassing the variational collapse caused by the Dirac sea.
desk verdict A credible proof-of-concept for DNN-based Dirac solvers, with a real but fixable gap: the log-mesh loss function is only variational under an unstated weighted inner product. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the inverse Hamiltonian H_inv = (ε' - H'_Dr)^{-1}, whose spectrum is inverted relative to the original Dirac Hamiltonian, placing the desired bound state at the bottom; the DNN outputs the trial function f(r) = F(r)/r, from which the small component G(r) is reconstructed through the radial Dirac equation, and the loss is the expectation value of H_inv. Two excited-state mechanisms hang off this object: shifting ε' between consecutive eigenvalues, or constructing a state orthogonal to all lower states before applying the inverse-Hamiltonian loss.
What would settle it
Compute the discrete spectrum of H_inv on the log mesh with the same derivative matrix and boundary conditions, using a dense linear solver; if spurious eigenvalues appear below the target state energy, the loss would converge to the wrong state on some initializations, directly contradicting the paper's consistency claim.
Extended reading notes
Core claim
The central claim is that the inverse Hamiltonian method, when used as an unsupervised loss function for a fully connected deep neural network, turns the target bound state of the radial Dirac equation into the lowest eigenstate of the modified operator H_inv = (ε' - H'_Dr)^{-1}, so that energy minimization no longer falls into the Dirac sea. For excited states, the paper demonstrates two routes: adjusting the parameter ε' to lie between adjacent eigenvalues, or projecting out all lower states via orthonormalization before minimizing the inverse-Hamiltonian expectation value. Both routes reproduce benchmark energies and wave functions for the Coulomb potential and for Woods-Saxon potentials, with the inverse Hamiltonian method being simpler to use and the orthonormal method converging faster but accumulating error from lower states.
Load-bearing premise
The discretized radial Dirac Hamiltonian on the logarithmic mesh is treated as a Hermitian matrix in the inverse-Hamiltonian loss, but the paper never specifies a discrete inner product that would make the finite-difference derivative matrix Hermitian, so the variational guarantee that should protect against collapse is not rigorously established for those calculations.
Editorial extensions
If this is right
- The method provides a parameter-free variational route to relativistic bound states without the need for kinetic balance or special basis sets.
- The computational cost scales as O(X M^2) with epochs X and mesh points M, potentially outpacing the O(M^3) eigen-solvers for large meshes.
- The same inverse-Hamiltonian loss can be adapted to relativistic two-body Dirac equations, as the paper notes, since those reduce to a Schrödinger-like eigenvalue problem in the center-of-mass frame.
- The orthonormal method's accuracy depends on the accuracy of lower states, so improving lower-state precision automatically improves higher excited states.
Reading between the lines
- If the log-mesh discretization is made Hermitian by introducing a suitable weighted inner product, the variational guarantee of the inverse-Hamiltonian loss would rest on a rigorous foundation; otherwise the observed agreement with benchmarks is empirical rather than guaranteed.
- The approach could be extended to non-spherical potentials by letting a single DNN output both radial components, but the paper's Appendix shows that small-component accuracy suffers when it contributes little to the norm, suggesting a balanced loss would be needed.
- Counting nodes in the large component, a technique the paper uses to identify principal quantum numbers, could be automated inside the loss to make the method fully self-tuning for excited states.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript extends a previously proposed unsupervised deep-neural-network variational method to the radial Dirac equation with spherical symmetry. To avoid variational collapse induced by the Dirac sea, it minimizes the Rayleigh quotient of the inverse Hamiltonian (ε′_nκ − H′_Dr)^{-1}; low-lying excited states are obtained either by choosing the shift ε′_nκ between neighboring eigenvalues or by projecting out lower states via an orthonormalization procedure. The method is applied to hydrogen on a logarithmic mesh and to Woods-Saxon potentials for 16O and 208Pb on a uniform mesh, with energies and wave functions compared against the analytic hydrogen solution and an independent general pseudospectral calculation.
Significance. If the validation is accepted, the paper contributes a practical unsupervised-DNN route to relativistic single-particle bound states and low-lying excited states, circumventing variational collapse in a way that does not fit the benchmark energies. The paper has genuine strengths: the loss function is the inverse-Hamiltonian Rayleigh quotient rather than a fit to known energies; the hydrogen comparison uses closed-form results and the Woods-Saxon comparison uses an independent pseudospectral method; two complementary excited-state strategies are compared; and Appendix A documents failed architectures, providing useful negative results. The main limitation is the unspecified discrete inner product, which currently leaves the logarithmic-mesh hydrogen part of the validation on weaker footing than the uniform-mesh Woods-Saxon part.
major comments (1)
- [Sec. III A, Eqs. (15) and (30)] The manuscript never specifies the discrete inner product used in the Rayleigh quotient of Eq. (15), in the normalization condition of Eq. (12), or in the orthonormalization of Eqs. (16) and (17). This is load-bearing because the log-mesh derivative matrix in Eq. (30) is not anti-symmetric under the Euclidean metric; the discretized H′_Dr is Hermitian only with respect to a weighted metric such as W = diag(e^{x_i}), with boundary terms vanishing under the Dirichlet conditions. Unless the implementation uses that weighted inner product, the minimized object is not the Rayleigh quotient of a Hermitian inverse Hamiltonian, and the variational selection of the target eigenstate has no rigorous foundation. The issue is not purely formal: the hydrogen inverse-Hamiltonian relative energy errors in Table I are uniformly about 5×10^{-4} to 8×10^{-4} for n=2–6, whereas the uniform-mesh Woods-Saxon results in Tables II and III, where Eq. (32) is anti-symmetric, reach about 10^{-5}–10^{-6} for the lowest states. The authors should state the discrete inner product explicitly and, if unweighted sums were used, repeat the hydrogen calculations with the correct metric; without this, the hydrogen part of the central validation is incomplete.
minor comments (8)
- [Sec. V C and Sec. VI] The statement in Sec. V C that all DNN energies are 'identical (within 0.002 %)' to the benchmark conflicts with Table III, where the largest relative error is 3.52×10^{-4} (0.035 %). The summary statement in Sec. VI that the differences are 'up to 0.15 %' also conflicts with Table I, where the n=6 orthonormal result has relative error 2.16×10^{-3} (0.216 %). These numbers should be harmonized.
- [Sec. IV C and Table I] The sentence that the inverse Hamiltonian method 'reaches at least 1 × 10^{-4} accuracy' is ambiguous and, taken literally, is not supported by Table I, where the n=2–6 inverse-Hamiltonian relative errors are 5.68×10^{-4} to 7.42×10^{-4}. Please state the per-state accuracy more precisely.
- [Sec. III B and Sec. IV C] The stopping criterion for training is never given. Figure 6 shows energy errors as functions of epochs and Tables I–III report final energies, but the reader cannot tell how convergence was decided, what tolerance was used, or how many independent runs produced the displayed values. This should be specified for reproducibility.
- [Sec. III B] The statement that two hidden layers of 16 units are 'confirmed to be sufficient' is not supported by any reported sensitivity study. A brief description of the tested architectures or hyperparameter variations would make the claim reproducible.
- [Sec. V C] The complexity comparison O(XM^2) for the DNN omits the one-time cost of constructing or factorizing the inverse Hamiltonian matrix. If the inverse is precomputed, there is an initial O(M^3) cost; if it is not, each epoch requires a linear solve. Since the comparison with the GPS O(M^3) is given as motivation, this should be clarified.
- [Table I] The row labeled ε′_{n−1} is confusing: Eq. (14) defines the shift as ε′_{nκ}, and the table appears to list the chosen shifts. Please make the notation consistent with the text.
- [Sec. I] The introduction contains a duplicated sentence: 'A fundamental challenge in variational treatments of the Dirac equation is variational collapse' appears twice in consecutive lines. This should be removed.
- [Appendix A 3] The observation that the fully connected DNN minimizes H′_Dr directly and converges to the hydrogen ground-state energy should be reconciled with the general claim in Sec. II B that naive energy minimization leads to variational collapse. The mechanism of trapping at a local minimum is interesting and should be stated in the main text rather than only in the appendix.
Circularity Check
No significant circularity: the reported energies are produced by minimizing physics-based loss functions and are compared with independent benchmarks; self-citations are contextual only.
full rationale
The derivation chain is self-contained. The DNN trial function fnκ(r) is fed into the discretized radial Dirac Hamiltonian (Eqs. (30) and (32)), and the loss is the Rayleigh quotient of the inverse Hamiltonian, Eq. (15): εnκ = ε′nκ − (min ⟨φ|Hnκ_inv|φ⟩/⟨φ|φ⟩)^−1. Minimizing this quotient is an eigenvalue computation for the assembled matrix; it does not use the benchmark energies as training labels. The hydrogen results are checked against the analytic expressions in Eq. (24), and the Woods-Saxon results are checked against the independent GPS method (Refs. [40,41]); those benchmarks appear only in comparison tables and figures. The hand-chosen shift ε′ selects which state lies at the bottom of the inverted spectrum, but the final eigenvalue is obtained from the minimization and is independent of ε′ as long as ε′ lies in the stated interval, so this is a spectral-window parameter, not a fitted prediction. The orthonormal method (Eqs. (16)–(18)) is a recursive Gram-Schmidt construction that removes lower computed states; its dependence on earlier states is an error-accumulation property, not a logical circularity. The citations to the authors' own Refs. [28] and [29] are contextual: the orthonormal-condition equations and the DNN architecture are fully specified in the present paper, and no uniqueness or existence theorem is imported solely from those references to force the reported numbers. The missing Hermitian-metric specification for the log-mesh discretization is a numerical-rigor concern, but it is not circular: the loss still evaluates the assembled discrete inverse Hamiltonian rather than a benchmark-derived target, so it does not reduce to the paper's inputs. The appendix's failed direct-minimization test (Fig. 16) further shows that the inverse-Hamiltonian construction is not cosmetic. Overall, the reported agreement with benchmarks is an independent numerical validation, not a consequence of the construction.
Assumptions & free parameters
free parameters (4)
- Inverse-Hamiltonian shift parameter ε'_nκ =
Examples: -0.51 to -0.015 Hartree for hydrogen; -45 to -20 MeV for Woods-Saxon
- Box size r_max =
20, 40, 40, 60, 90, 100 a.u. for hydrogen n=1..6; 20 fm for Woods-Saxon
- DNN hyperparameters =
2 hidden layers of 16 units; softplus; Adam lr=0.001; M=1700 or 2000 mesh points
- Number of training epochs =
Varies, roughly 8e3 to 8e4 (from Fig. 6)
assumptions (5)
- domain assumption The Dirac Hamiltonian with spherical symmetry reduces to the 2x2 radial system of Eq. (8).
- standard math For ε' between eigenvalues ε_{n-1} and ε_n, the inverse operator (ε' - H)^(-1) has the n-th bound state as its lowest eigenstate.
- ad hoc to paper The discretized derivative matrix of Eq. (30) with log mesh yields a Hermitian Hamiltonian under the inner product used in the loss function.
- ad hoc to paper The fully connected DNN with softplus activation can represent the target wave functions accurately enough, and Adam optimization finds the relevant minimum.
- domain assumption The small component G can be computed from F via Eq. (11) using the eigenvalue from the previous epoch.
Cite this review
Pith. "Pith review of A deep neural network approach to solve the Dirac equation." pith.science (2026). https://pith.science/paper/VZL7CVV6
@misc{pith2026241203090,
author = {Pith},
title = {Pith review of: A deep neural network approach to solve the Dirac equation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VZL7CVV6}},
note = {Machine review of arXiv:2412.03090}
}
read the original abstract
We extend the method from [Naito, Naito, and Hashimoto, Phys. Rev. Research 5, 033189 (2023)] to solve the Dirac equation not only for the ground state but also for low-lying excited states using a deep neural network and the unsupervised machine learning technique. The variational method fails because of the Dirac sea, which is avoided by introducing the inverse Hamiltonian method. For low-lying excited states, two methods are proposed, which have different performances and advantages. The validity of this method is verified by the calculations with the Coulomb and Woods-Saxon potentials.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
Towards constraining QCD phase transitions in neutron star interiors: Bayesian Inference with TOV linear response analysis
A Bayesian framework with analytical TOV linear-response gradients and a neural-network equation of state reconstructs neutron star EoSs and constrains first-order phase transition parameters from simulated mass-radius data.
Reference graph
Works this paper leans on
-
[1]
11, whereFnκ and Gnκ are gener- ated by two output units, respectively
Employing a non-fully connected neural network Besides the fully-connected DNN used in the main text, we also test a non-fully connected deep neural net- work to generate theF- andG-components, whose struc- ture is shown in Fig. 11, whereFnκ and Gnκ are gener- ated by two output units, respectively. For each output, the units of the hidden layers are full...
-
[2]
IIIA, a trial wave functionfnκ is used as the DNN output to calculateFnκ and Gnκ
Usage of the trial wave function fnκ In Sec. IIIA, a trial wave functionfnκ is used as the DNN output to calculateFnκ and Gnκ. The denominator r improves the accuracy of Fnκ and Gnκ close to the origin. If Fnκ is used as the DNN output, Gnκ of a hydro- gen atom diverges at the origin as shown in Fig. 14. The ground-state wave functions of208Pb are shown i...
-
[3]
Minimizing without the inverse Hamiltonian Because of the existence of the Dirac sea, the ground state is no longer the lowest energy of the Dirac Hamil- tonian H ′ Dr. We employ the fully connected DNN in the main text and the non-fully connected DNN in Ap- pendix A1 to minimize the energy expectation value of H ′ Dr εD = min ⟨ϕ|H ′ Dr|ϕ⟩ ⟨ϕ|ϕ⟩ (A1) as t...
-
[4]
Carleo, I
G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Rev. Mod. Phys.91, 045002 (2019)
2019
-
[5]
G. Carleo and M. Troyer, Solving the quantum many- body problem with artificial neural networks, Science 15 0 2 4 6 8 10 0.0 0.1 0.2 0.3 0.4 0.5 r (fm) Exact DNN 0 2 4 6 8 10 −0.04 −0.03 −0.02 −0.01 0.00 r (fm) Exact DNN0.0 0.2 0.4 −0.0008 −0.0004 0.0000 FIG. 15. Same as Fig. 14 but for the Woods-Saxon potential for 208Pb. 355, 602 (2017)
work page 2017
-
[6]
Nomura, A
Y. Nomura, A. S. Darmawan, Y. Yamaji, and M. Imada, Restricted Boltzmann machine learning for solving strongly correlated quantum systems, Phys. Rev. B96, 205152 (2017)
2017
-
[7]
Carleo, Y
G. Carleo, Y. Nomura, and M. Imada, Constructing ex- act representations of quantum many-body systems with deep neural networks, Nat. Commun.9, 5322 (2018)
2018
-
[8]
K. Choo, G. Carleo, N. Regnault, and T. Neupert, Symmetries and Many-Body Excitations with Neural- Network Quantum States, Phys. Rev. Lett.121, 167204 (2018)
2018
Show all 46 references
-
[9]
Nomura, Machine Learning Quantum States - Exten- sions to Fermion-Boson Coupled Systems and Excited- State Calculations, J
Y. Nomura, Machine Learning Quantum States - Exten- sions to Fermion-Boson Coupled Systems and Excited- State Calculations, J. Phys. Soc. Jpn.89, 054706 (2020)
2020
-
[10]
Stokes, J
J. Stokes, J. R. Moreno, E. A. Pnevmatikakis, and G. Carleo, Phases of two-dimensional spinless lattice fermions with first-quantized deep neural-network quan- tum states, Phys. Rev. B102, 205122 (2020)
2020
-
[11]
J. R. Moreno, G. Carleo, A. Georges, and J. Stokes, Fermionic wave functions from neural-network con- strained hidden states, Proc. Natl. Acad. Sci. USA119, 0 500 1000 −38000 −36000 −34000 −32000 −30000 −28000 −26000 −24000 −22000 −20000 eD (hartree) Epoch (a) Non-fully connec...
2022
-
[12]
Yoshino, Spatially heterogeneous learning by a deep student machine, Phys
H. Yoshino, Spatially heterogeneous learning by a deep student machine, Phys. Rev. Res.5, 033068 (2023)
2023
-
[13]
Saito, Method to Solve Quantum Few-Body Problems with Artificial Neural Networks, J
H. Saito, Method to Solve Quantum Few-Body Problems with Artificial Neural Networks, J. Phys. Soc. Jpn.87, 074002 (2018)
2018
-
[14]
Pescia, J
G. Pescia, J. Han, A. Lovato, J. Lu, and G. Carleo, Neural-network quantum states for periodic systems in continuous space, Phys. Rev. Res.4, 023138 (2022)
2022
-
[15]
D. Pfau, J. S. Spencer, A. G. D. G. Matthews, and W. M. C. Foulkes,Ab initio solution of the many-electron Schrödinger equation with deep neural networks, Phys. Rev. Res. 2, 033429 (2020)
2020
-
[16]
Cassella, H
G. Cassella, H. Sutterud, S. Azadi, N. D. Drummond, D. Pfau, J. S. Spencer, and W. M. C. Foulkes, Discov- ering Quantum Phase Transitions with Fermionic Neural 16 Networks, Phys. Rev. Lett.130, 036401 (2023)
2023
-
[17]
W. T. Lou, H. Sutterud, G. Cassella, W. M. C. Foulkes, J. Knolle, D. Pfau, and J. S. Spencer, Neural Wave Func- tions for Superfluids, Phys. Rev. X14, 021030 (2024)
2024
-
[18]
M.Ruggeri, S.Moroni,andM.Holzmann,NonlinearNet- work Description for Many-Body Quantum Systems in Continuous Space, Phys. Rev. Lett.120, 205302 (2018)
2018
-
[19]
Hermann, Z
J. Hermann, Z. Schätzle, and F. Noé, Deep-neural- network solution of the electronic Schrödinger equation, Nat. Chem. 12, 891 (2020)
2020
-
[20]
X. Li, Z. Li, and J. Chen,Ab initio calculation of real solids via neural network ansatz, Nat. Commun.13, 7895 (2022)
2022
-
[21]
Adams, G
C. Adams, G. Carleo, A. Lovato, and N. Rocco, Varia- tional Monte Carlo Calculations ofA ≤ 4 Nuclei with an Artificial Neural-Network Correlator Ansatz, Phys. Rev. Lett. 127, 022502 (2021)
2021
-
[22]
Gnech, C
A. Gnech, C. Adams, N. Brawand, G. Carleo, A. Lovato, andN.Rocco,NucleiwithUpto A = 6NucleonswithAr- tificial Neural Network Wave Functions, Few-Body Syst. 63, 7 (2022)
2022
-
[23]
Y. L. Yang and P. W. Zhao, A consistent description of the relativistic effects and three-body interactions in atomic nuclei, Phys. Lett. B835, 137587 (2022)
2022
-
[24]
Y. L. Yang and P. W. Zhao, Deep-neural-network ap- proach to solving the ab initio nuclear structure problem, Phys. Rev. C107, 034320 (2023)
2023
-
[25]
B. Fore, J. M. Kim, G. Carleo, M. Hjorth-Jensen, A. Lovato, and M. Piarulli, Dilute neutron star matter from neural-network quantum states, Phys. Rev. Res.5, 033062 (2023)
2023
-
[26]
B. Fore, J. Kim, M. Hjorth-Jensen, and A. Lovato, Inves- tigating the crust of neutron stars with neural-network quantum states (2024), arXiv:2407.21207 [nucl-th]
2024 arXiv
-
[27]
Lovato, C
A. Lovato, C. Adams, G. Carleo, and N. Rocco, Hidden- nucleons neural-network quantum states for the nuclear many-body problem, Phys. Rev. Res.4, 043178 (2022)
2022
-
[28]
Wilson, S
M. Wilson, S. Moroni, M. Holzmann, N. Gao, F. Wu- darski, T. Vegge, and A. Bhowmik, Neural network ansatz for periodic wave functions and the homogeneous electron gas, Phys. Rev. B107, 235139 (2023)
2023
-
[29]
Hermann, J
J. Hermann, J. Spencer, K. Choo, A. Mezzacapo, W. M. C. Foulkes, D. Pfau, G. Carleo, and F. Noé, Ab initio quantum chemistry with neural-network wavefunc- tions, Nat. Rev. Chem.7, 692 (2023)
2023
-
[30]
K. Choo, A. Mezzacapo, and G. Carleo, Fermionic neural-network states for ab-initio electronic structure, Nat. Commun. 11, 2368 (2020)
2020
-
[31]
Naito, H
T. Naito, H. Naito, and K. Hashimoto, Multi-body wave function of ground and low-lying excited states using un- ornamented deep neural networks, Phys. Rev. Res. 5, 033189 (2023)
2023
-
[32]
C. Wang, T. Naito, J. Li, and H. Liang, A neural network approach for two-body systems with spin and isospin de- grees of freedom (2024), arXiv:2403.16819 [nucl-th]
2024 arXiv
-
[33]
Lorin and X
E. Lorin and X. Yang, Time-dependent dirac equation with physics-informed neural networks: Computation and properties, Computer Physics Communications280, 108474 (2022)
2022
-
[34]
Armstrong, Lloyd, Relativistic Effects in Atomic Fine Structure, J
J. Armstrong, Lloyd, Relativistic Effects in Atomic Fine Structure, J. Math. Phys.7, 1891 (1966)
1966
-
[35]
Ernzerhof and F
M. Ernzerhof and F. Goyer, Conjugated Molecules De- scribed by a One-Dimensional Dirac Equation, J. Chem. Theory Comput. 6, 1818 (2010)
2010
-
[36]
B.D.Serot,BuildingAtomicNucleiwiththeDiracEqua- tion, Int. J. Mod. Phys. A19, 107 (2004)
2004
-
[37]
Hagino and Y
K. Hagino and Y. Tanimura, Iterative solution of a Dirac equation with an inverse Hamiltonian method, Phys. Rev. C 82, 057301 (2010)
2010
-
[38]
I. P. Grant, Relativistic quantum theory of atoms and molecules: theory and computation (Springer New York, NY, 2007)
2007
-
[39]
https://www.tensorflow.org/about/bib?hl=en
-
[40]
D. P. Kingma and J. Ba, Adam: A Method for Stochastic Optimization, in 3rd International Conference on Learn- ing Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , edited by Y. Bengio and Y. LeCun (2015) arXiv:1412.6980 [cs.LG]
2015 arXiv
-
[41]
J.J.Sakurai, Advanced quantum mechanics (PearsonEd- ucation India, 1967)
1967
-
[42]
Koepf and P
W. Koepf and P. Ring, The spin-orbit field in superde- formed nuclei: a relativistic investigation, Z. Phys. A 339, 81 (1991)
1991
-
[43]
Yao and S.-I
G. Yao and S.-I. Chu, Generalized pseudospectral meth- ods with mappings for bound and resonance state prob- lems, Chem. Phys. Lett.204, 381 (1993)
1993
-
[44]
L. G. Jiao, Y. Y. He, A. Liu, Y. Z. Zhang, and Y. K. Ho, Development of the kinetically and atomically balanced generalized pseudospectral method, Phys. Rev. A 104, 022801 (2021)
2021
-
[45]
H. W. Crater and P. Van Alstine, Two-body Dirac equa- tions, Ann. Phys.148, 57 (1983)
1983
-
[46]
H. W. Crater and P. Van Alstine, Two-body dirac equa- tions for relativistic bound states of quantum field theory, arXiv preprint hep-ph/9912386 (1999)
1999 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.