Pith. sign in

REVIEW 3 major objections 4 minor 65 references

An autoencoder trained on Fermi-Hubbard ground-state observables discovers a sharp minimal latent dimension of d=L−1 and turns its decoder into a differentiable variational ansatz for direct energy minimization in latent space.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 16:46 UTC pith:EVFZB4MH

load-bearing objection A well-executed demonstration that autoencoders find the L-1 dimensional ground-state manifold of 1D Fermi-Hubbard models, but the N-representability claim is stronger than the evidence. the 3 major comments →

arxiv 2512.11767 v1 pith:EVFZB4MH submitted 2025-12-12 quant-ph cond-mat.str-elcs.LG

Learning Minimal Representations of Fermionic Ground States

classification quant-ph cond-mat.str-elcs.LG
keywords autoencoderFermi-Hubbard modellatent spacevariational ansatzN-representabilityrepresentation learningground statesdimensionality reduction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to show that a plain autoencoder, trained only on Hamiltonian-term expectation values of 1D Fermi-Hubbard ground states, discovers the family's true intrinsic dimensionality: reconstruction error drops sharply exactly at d=L−1 latent dimensions, matching the L−1 independent parameters (L site potentials minus a constant shift). At that critical dimension the decoder defines a smooth, bounded manifold of physical expectation values, and gradient descent on the latent coordinates minimizes the energy for new, unseen potentials without ever explicitly enforcing N-representability. The authors argue this makes learned latent spaces both a probe for how many degrees of freedom a ground-state family really has and a practical variational ansatz, with the same threshold reappearing when the input is the larger two-body reduced density matrix.

Core claim

The central claim is empirical and structural: for the family of half-filled 1D Hubbard models with random local potentials, an unsupervised autoencoder trained to reconstruct Hamiltonian expectation values exhibits a sharp reconstruction-quality transition at d=L−1 latent dimensions—the exact number of independent physical degrees of freedom (L potentials minus the global constant shift). Below this dimension information is lost; above it, additional latent directions are redundant for reconstruction and actively harmful for energy optimization, since the optimizer exploits weakly constrained, unphysical directions. At d=L−1 the latent manifold is a smooth L−1-dimensional ball, the decoder

What carries the argument

The load-bearing object is an autoencoder whose bottleneck defines the latent space: an encoder maps the observables ω (on-site densities, nearest-neighbor hopping correlators, double occupancies) to z in R^d, and a decoder maps back. Reconstruction loss is augmented with a radial well that confines latent vectors to a sphere, a contrastive repulsion that prevents latent collapse and preserves approximate isometry, and Lipschitz smoothness on both networks. Because energy is the linear contraction E(z;μ)=h(μ)·D(z), the decoder turns latent space into a differentiable energy landscape that L-BFGS can minimize. The critical role of the regularizers is shown by ablations: without them, reconstr

Load-bearing premise

That the decoder's output stays on or near the set of physical, N-representable expectation values during latent-space energy optimization, so the optimized energies remain valid variational upper bounds; the training only enforces reconstruction fidelity and latent regularity, not physicality of decoded states.

What would settle it

Take a trained decoder at d=L−1, pick a test potential, minimize E(z;μ) in latent space, and check whether any accepted optimum either (a) yields an energy below the exact diagonalization ground-state energy or (b) decodes to a 2-RDM that violates known N-representability conditions (such as positivity or appropriate eigenvalue bounds). Finding either would show the optimizer exploited unphysical extrapolation, contradicting the claim of implicit N-representability.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The sharp threshold at d=L−1 gives a data-driven way to count the intrinsic degrees of freedom of a ground-state family without prior physics knowledge; this is a direct corollary of the paper's compression results.
  • The decoder at the critical dimension works as a variational ansatz: for many unseen potentials, latent-space optimization reaches ground-state energies with small RMSE, while larger d degrades performance.
  • Compressing the full two-body reduced density matrix gives the same L−1 threshold, implying the extra two-body correlation information is redundant for energy and for the manifold structure of these ground states.
  • Encoder Jacobian analysis shows the latent representation is driven mainly by local density and on-site interaction terms, with nearest-neighbor correlators playing a minor role, connecting the learned features to density-functional-style variables.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the threshold behavior generalizes beyond the Hubbard family, the method could be used as a diagnostic on data from quantum simulators or tensor-network calculations to reveal how many effective variables actually control a ground-state manifold; the paper gestures at this but does not demonstrate it.
  • The variational upper-bound property is inherited from the decoder's physicality, not enforced. A stress test would be to scan optimized latent points for violations of known N-representability conditions (e.g., 2-RDM positivity) or energies below the exact ground-state energy; the paper's rejection rule only bounds the latent norm.
  • The ring-shaped latent geometry observed for odd-L, near-degenerate systems suggests latent topology could be used as a fingerprint for ground-state degeneracy and disorder-induced symmetry breaking, a testable extension to other models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents an unsupervised autoencoder framework for one-dimensional Fermi-Hubbard ground states. The input is the vector of Hamiltonian-term expectation values ω (hopping, density, and double-occupancy terms on each site), and the encoder maps ω to a latent vector z∈ℝ^d while the decoder reconstructs ω. Training combines reconstruction loss with radial-well, contrastive-repulsion, and Lipschitz regularization. The authors report a sharp drop in reconstruction RMSE at d=L−1 across L=4–14, interpret this as the intrinsic dimensionality of the ground-state manifold, and then use the frozen decoder as a differentiable variational ansatz: they minimize E(z;μ)=h(μ)·D(z) over z for unseen potentials. They find the lowest energy error at d=L−1, with degraded stability for d≥L. The same threshold is reported for 2-RDM inputs (L=4,6,8). The central advertised claim is that this latent-space optimization 'circumvents the N-representability problem' because the learned manifold implicitly restricts the search to physical states.

Significance. If the N-representability claim could be substantiated, the method would be a useful data-driven variational ansatz and a practical diagnostic of ground-state-manifold dimensionality. The paper has real strengths: the d=L−1 threshold is an emergent, unfitted empirical result observed over a controlled range of system sizes; the ablation study is informative; and the 2-RDM experiment provides a nontrivial cross-check. However, the manuscript currently does not verify the load-bearing assumption that decoder outputs during energy optimization remain N-representable. The energy results are promising, but they cannot be interpreted as variational upper bounds unless physicality of the decoded states is demonstrated. The paper is therefore conditionally acceptable in principle, but the advertised central claim needs either direct evidence or a substantial reformulation.

major comments (3)
  1. [Sec. II (Energy optimization) and Sec. III] The variational interpretation of Eq. (10)-(11) requires that D(z) map into the set of physical N-electron expectation values for every z reached by the optimizer. The paper itself states that the decoder is not explicitly constrained and that one only 'can result in' the optimization remaining near the physical manifold. The rejection rule ∥z*∥≤r_opt does not test physicality; it only bounds the latent norm. The d≥L data in Fig. 4 show that the optimizer does exploit unphysical extrapolation, pushing z to the boundary and producing many rejected trajectories. Nothing rules out smaller, still unphysical excursions at d=L−1. Please report signed energy-error distributions (not only RMSE), and provide a direct check on physicality of the optimized decoder outputs—for example, density bounds 0≤n_i≤2, double-occupancy inequalities, consistency between the decoded ω and an exact-diagonalizati
  2. [App. I] The statement that the 2-RDM autoencoder is 'implicitly satisfying the N-representability conditions' is asserted without verification. A high reconstruction accuracy on the training set does not imply that every decoder output, or every point along an optimization trajectory, lies in the N-representable set. The 2-RDM decoder is even more complex (up to ~5×10^7 parameters) and less constrained. Please provide an explicit test of N-representability of decoded 2-RDMs (e.g., checking the known necessary conditions such as P-, Q-, and G-conditions, or comparing against exact 2-RDMs for the optimized states), or weaken the claim substantially.
  3. [Sec. III and App. I] The interpretation of d=L−1 as the 'intrinsic number of independent degrees of freedom' of the ground-state manifold is partly built into the data-generation model: the potentials μ have L parameters, and the constant-shift degeneracy leaves exactly L−1 independent degrees of freedom. The autoencoder's threshold therefore largely recovers the generative parameter count. The 2-RDM experiment shows that this number is not an artifact of the particular ω representation, which is valuable, but it does not establish that L−1 is an intrinsic property of the quantum states independent of the chosen Hamiltonian parameterization. I recommend stating this more cautiously and, if possible, testing a family of Hamiltonians whose generative parameter dimension differs from L−1 to distinguish these interpretations.
minor comments (4)
  1. [Eq. (8)] The row-normalization formula appears to contain a typo: as written, 'min[1, softplus(ci) P_k |(Wi)rk|]' multiplies by the row sum, whereas the surrounding text describes normalization by the row sum. The intended expression is likely min[1, softplus(ci)/Σ_k |(Wi)rk|] Wi. Please correct and define the row index.
  2. [App. H] The text says 'with the ring topology visible in Fig. 3' when discussing the (L=5,N=4) degenerate system; the ring should appear in Fig. 8(a). Please fix the cross-reference.
  3. [Sec. III (Energy optimization)] The description of the rejection procedure is confusing: an 'absorbing potential' that enforces ∥z∥≤r_opt is described, but then solutions with ∥z∥>r_opt are discarded. Clarify whether the soft constraint is used during optimization, after optimization, or both, and report how often the constraint is active.
  4. [Code and data availability] The statement that code and data 'will be made available upon publication' makes the results difficult to reproduce during review. Please provide an anonymized repository or describe how the key results can be reproduced from the method details alone.

Circularity Check

0 steps flagged

No circularity: the L−1 threshold is an empirical benchmark and the energy optimization uses held-out potentials.

full rationale

The paper’s derivation chain is empirical rather than definitional. The central threshold d=L−1 emerges from scanning latent dimensions and measuring held-out reconstruction loss; it is not a parameter fitted to the target statement, and the paper explicitly benchmarks it against the analytically known L−1 degrees of freedom of the Hubbard family. The energy-optimization experiments use new test potentials and compute energies from the Hamiltonian coefficients h(μ) contracted with decoder outputs, so there is no leakage of training labels into the reported errors. The self-citations (Refs. [38] and [43]) concern representation-learning methodology and the random-potential sampling distribution; they are not used to justify the threshold or to import a uniqueness theorem. The main limitation, acknowledged by the authors in Sec. II, is that the decoder is not explicitly constrained to produce N-representable states, so the claim of ‘circumventing the N-representability problem’ is an unverified hope rather than a circular step. This is a correctness/evidence concern, not circularity. Therefore no significant circularity is found.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The paper introduces no new physical entities. Its free parameters are ML hyperparameters and data-generation choices; the central L−1 threshold is not fitted but follows from the parameterization of the Hamiltonian family. The most consequential assumption is the ad hoc hope that the decoder output remains N-representable during optimization, which is not proven.

free parameters (6)
  • Latent dimension d = L−1 (observed optimal; scanned d ∈ [L−3, L+2])
    Central hyperparameter of the autoencoder; the paper's main claim is that the optimal compression occurs at d=L−1.
  • Radial well radius r = ≈2 (inferred from Fig. 3; exact value not stated)
    Hand-picked bound on latent norm; shapes the spherical latent geometry and affects optimization behavior.
  • Energy-optimization bound r_opt = 3
    Used to reject divergent latent-space optimizations (∥z*∥>r_opt); directly affects reported energy errors and acceptance percentages.
  • Regularization weights α,β,γ,δ = 1e−7, 1e−7, 1e−9, 1e−8
    Hand-set/hyperparameter-tuned weights in the total loss; affect reconstruction and optimization quality but not the threshold position.
  • Potential strength range W = [0.005t, 2.5t]
    Uniform sampling range for the disorder strength; determines which ground states appear in the training set.
  • Potential acceptance threshold σ(µ)<0.4t = 0.4t
    Data-generation rejection rule from Ref. [43]; ensures a spread of potential strengths and shapes the distribution of the ground-state manifold.
axioms (6)
  • domain assumption The ground-state manifold of the sampled Hubbard family has intrinsic dimension exactly L−1 because a global potential shift μ_i→μ_i+c leaves the ground state unchanged.
    Used in Sec. III to interpret the reconstruction threshold at d=L−1 as matching intrinsic degrees of freedom.
  • standard math A global potential shift adds a constant cN to the Hamiltonian and therefore does not alter the ground state.
    Elementary operator identity; basis for the L−1 counting argument.
  • ad hoc to paper Reconstruction loss as a function of latent dimension is a reliable estimator of the intrinsic dimensionality of the data manifold.
    The sharp drop at d=L−1 is taken as evidence of the minimal latent dimension; no formal proof or comparison to other intrinsic-dimension estimators is given.
  • ad hoc to paper For the optimized latent vectors, the decoder output remains within or close to the N-representable physical manifold, so the minimized energy is a valid variational upper bound.
    Explicitly acknowledged in Sec. II as not guaranteed; central to the claim of circumventing N-representability.
  • domain assumption The neural network autoencoder, regularized by well, repulsion, and Lipschitz terms, generalizes from 1e5 training potentials to held-out potentials in the same distribution.
    Required for the held-out energy-optimization results to support the method; empirically checked but not proven.
  • domain assumption Exact diagonalization (PySCF FCI) provides exact ground states and 2-RDMs for L≤14.
    Data-generation premise; standard for these system sizes but not independently verified in the paper.

pith-pipeline@v1.3.0-alltime-deepseek · 19071 in / 17790 out tokens · 168995 ms · 2026-08-03T16:46:10.195729+00:00 · methodology

0 comments
read the original abstract

We introduce an unsupervised machine-learning framework that discovers optimally compressed representations of quantum many-body ground states. Using an autoencoder neural network architecture on data from $L$-site Fermi-Hubbard models, we identify minimal latent spaces with a sharp reconstruction quality threshold at $L-1$ latent dimensions, matching the system's intrinsic degrees of freedom. We demonstrate the use of the trained decoder as a differentiable variational ansatz to minimize energy directly within the latent space. Crucially, this approach circumvents the $N$-representability problem, as the learned manifold implicitly restricts the optimization to physically valid quantum states.

Figures

Figures reproduced from arXiv: 2512.11767 by Emiel Koridon, Felix Frohnert, Stefano Polla.

Figure 1
Figure 1. Figure 1: Overview: workflow for learning minimal repre￾sentations of fermionic ground states. A) The process starts by generating instances of the Hubbard Hamiltonian Hˆ (µ) from random local potentials µ. For each instance, the ex￾pectation values of the Hamiltonian terms ω are computed via exact diagonalization. B) The resulting dataset is used to train a neural network–based autoencoder that compresses ω into a … view at source ↗
Figure 2
Figure 2. Figure 2: visualizes the test reconstruction loss as a func￾tion of latent dimension d. We plot the mean and stan￾dard deviation over three independent models with ran￾dom initialization. Across all system sizes, the loss de￾creases with increasing d, exhibiting a sharp drop at d = L − 1, followed by saturation for d ≥ L − 1. This sudden improvement marks a compression threshold: for d < L − 1 the latent space is in… view at source ↗
Figure 3
Figure 3. Figure 3: Visualizing the learned representation: la￾tent space of the autoencoder for (L = N = 6), color-coded by the strength of the potential µ. The circular structure in the feature-wise projections results from the well constraint and uniform coverage enforced by contrastive repulsion, while the smooth radial color gradient shows that states with sim￾ilar potential strengths cluster together, indicating that th… view at source ↗
Figure 4
Figure 4. Figure 4: Energy optimization threshold: Error (RMSE) of optimized energies versus latent dimension d. Energies are obtained via gradient-based minimization in the learned la￾tent space, using the decoder as a differentiable variational ansatz. A distinct error minimum occurs at d = L−1, match￾ing the intrinsic degrees of freedom of the half-filled Hubbard ground state. For d ≥ L, performance degrades as excess late… view at source ↗
Figure 5
Figure 5. Figure 5: Ablation study: Performance comparison of model variants on systems with L = N = 6 and latent di￾mensions d = L − 1. We evaluate six distinct model con￾figurations in terms of reconstruction error and optimization performance. Error bars show the standard deviation over five independently trained models with different random ini￾tializations. The full model consistently outperforms all ab￾lated variants, i… view at source ↗
Figure 6
Figure 6. Figure 6: Training data efficiency: root mean squared er￾ror reconstruction loss as a function of the size of the training and validation dataset (accounting for symmetry augmenta￾tion) for different system sizes at the optimal latent dimension d = L − 1. The curves indicate how the data requirement grows with system size to achieve high-quality compression and display the best out of three training instances [PITH… view at source ↗
Figure 7
Figure 7. Figure 7: Interpreting the trained encoder: (a) Average Jacobian matrix and variability. Visualization of the mean Ja￾cobian matrix J(ω) of the encoder for the L = N = 6 model (latent dimension d = L − 1) computed over 104 input sam￾ples. The color map represents the mean magnitude (µ), while the overlaid numbers represent the sample standard deviation (σ), scaled by a factor of 103 . The consistently small variabil… view at source ↗
Figure 8
Figure 8. Figure 8: Compression of degenerate systems: (a) La￾tent space of the autoencoder for (L = 5, N = 4), color-coded by the potential strength µ. Unlike the L = N systems, the data does not exhibit the same circular structure in the feature-wise projections, indicating qualitatively different un￾derlying physics. The ring-like structure illustrates an un￾derlying degeneracy in the ground state manifold. (b) Test recons… view at source ↗
Figure 9
Figure 9. Figure 9: Compression and energy optimization of two-body reduced density matrices: (Top) Test recon￾struction RMSE (mean and standard deviation over three models) versus latent dimension d. A sharp drop at d = L−1 confirms that the 2-RDM shares the same intrinsic degrees of freedom as the Hamiltonian terms, despite its larger dimen￾sionality. (Middle) Energy optimization error using the de￾coder as a variational an… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 18 linked inside Pith

  1. [1]

    R. P. Feynman, Simulating physics with computers, International Journal of Theoretical Physics21, 467 (1982)

  2. [2]

    Tindall, M

    J. Tindall, M. Fishman, E. M. Stoudenmire, and D. Sels, Efficient Tensor Network Simulation of IBMs Eagle Kicked Ising Experiment, PRX Quantum5, 010308 (2024), publisher: American Physical Society

  3. [3]

    Patra, S

    S. Patra, S. S. Jahromi, S. Singh, and R. Or´ us, Efficient tensor network simulation of IBMs largest quantum pro- cessors, Physical Review Research6, 013326 (2024), pub- lisher: American Physical Society

  4. [4]

    Menczer, M

    A. Menczer, M. van Damme, A. Rask, L. Hunting- ton, J. Hammond, S. S. Xantheas, M. Ganahl, and O. Legeza, Parallel Implementation of the Density Ma- trix Renormalization Group Method Achieving a Quarter petaFLOPS Performance on a Single DGX-H100 GPU Node, Journal of Chemical Theory and Computation20, 8397 (2024), publisher: American Chemical Society

  5. [5]

    Brabec, J

    J. Brabec, J. Brandejs, K. Kowalski, S. Xantheas, O. Leg- eza, and L. Veis, Massively parallel quantum chemical 7 density matrix renormalization group method (2020), arXiv:2001.04890 [physics]

  6. [6]

    P. R. Nagy and M. K´ allay, Approaching the Basis Set Limit of CCSD(T) Energies for Large Molecules with Lo- cal Natural Orbital Coupled-Cluster Methods, Journal of Chemical Theory and Computation15, 5275 (2019), publisher: American Chemical Society

  7. [7]

    Schollwoeck, The density-matrix renormalization group in the age of matrix product states, Annals of Physics326, 96 (2011), arXiv:1008.3477 [cond-mat]

    U. Schollwoeck, The density-matrix renormalization group in the age of matrix product states, Annals of Physics326, 96 (2011), arXiv:1008.3477 [cond-mat]

  8. [8]

    G. K.-L. Chan and S. Sharma, The density matrix renor- malization group in quantum chemistry, Annual Review of Physical Chemistry62, 465 (2011), publisher: Annual Reviews

  9. [9]

    Or´ us, Tensor networks for complex quantum systems, Nature Reviews Physics1, 538 (2019), publisher: Nature Publishing Group

    R. Or´ us, Tensor networks for complex quantum systems, Nature Reviews Physics1, 538 (2019), publisher: Nature Publishing Group

  10. [10]

    Helgaker, P

    T. Helgaker, P. Jørgensen, and J. Olsen,Molecular Electronic-Structure Theory, 1st ed. (Wiley, 2000)

  11. [11]

    C. D. Sherrill, Frontiers in electronic structure theory, The Journal of Chemical Physics132, 110902 (2010)

  12. [12]

    S. R. White, Density matrix formulation for quantum renormalization groups, Physical Review Letters69, 2863 (1992), publisher: American Physical Society

  13. [13]

    S. R. White, Density-matrix algorithms for quantum renormalization groups, Physical Review B48, 10345 (1993), publisher: American Physical Society

  14. [14]

    Verstraete, J

    F. Verstraete, J. I. Cirac, and V. Murg, Matrix Prod- uct States, Projected Entangled Pair States, and vari- ational renormalization group methods for quantum spin systems, Advances in Physics57, 143 (2008), arXiv:0907.2796 [quant-ph]

  15. [15]

    Eisert, M

    J. Eisert, M. Cramer, and M. B. Plenio, Colloquium: Area laws for the entanglement entropy, Reviews of Mod- ern Physics82, 277 (2010), publisher: American Physical Society

  16. [16]

    Burke and L

    K. Burke and L. O. Wagner, DFT in a nutshell, Interna- tional Journal of Quantum Chemistry113, 96 (2013)

  17. [17]

    Burke, Perspective on density functional theory, The Journal of Chemical Physics136, 150901 (2012)

    K. Burke, Perspective on density functional theory, The Journal of Chemical Physics136, 150901 (2012)

  18. [18]

    Mardirossian and M

    N. Mardirossian and M. Head-Gordon, Thirty years of density functional theory in computational chem- istry: an overview and extensive assessment of 200 density functionals, Molecular Physics115, 2315 (2017), publisher: Taylor & Francis eprint: https://doi.org/10.1080/00268976.2017.1333644

  19. [19]

    R. J. Bartlett and M. Musia l, Coupled-cluster theory in quantum chemistry, Reviews of Modern Physics79, 291 (2007), publisher: American Physical Society

  20. [20]

    T. D. Crawford and H. F. Schaefer III, An introduction to coupled cluster theory for computational chemists, Re- views in computational chemistry14, 33 (2007)

  21. [21]

    I. Y. Zhang and A. Gr¨ uneis, Coupled Cluster The- ory in Materials Science, Frontiers in Materials6, 10.3389/fmats.2019.00123 (2019), publisher: Frontiers

  22. [22]

    Georges, Strongly Correlated Electron Materials: Dy- namical Mean-Field Theory and Electronic Structure, in AIP Conference Proceedings, Vol

    A. Georges, Strongly Correlated Electron Materials: Dy- namical Mean-Field Theory and Electronic Structure, in AIP Conference Proceedings, Vol. 715 (2004) pp. 3–74, iSSN: 0094243X arXiv:cond-mat/0403123

  23. [23]

    Dawid, J

    A. Dawid, J. Arnold, B. Requena, A. Gresch, M. P lodzie´ n, K. Donatella, K. A. Nicoli, P. Stornati, R. Koch, M. B¨ uttner, R. Oku la, G. Mu˜ noz-Gil, R. A. Vargas-Hern´ andez, A. Cervera-Lierta, J. Carrasquilla, V. Dunjko, M. Gabri´ e, P. Huembeli, E. v. Nieuwenburg, F. Vicentini, L. Wang, S. J. Wetzel, G. Carleo, E. Gre- plov´ a, R. Krems, F. Marquardt,...

  24. [24]

    Rocchetto, E

    A. Rocchetto, E. Grant, S. Strelchuk, G. Carleo, and S. Severini, Learning hard quantum distributions with variational autoencoders, npj Quantum Information4, 28 (2018), publisher: Nature Publishing Group

  25. [25]

    Dunjko and H

    V. Dunjko and H. J. Briegel, Machine learning & artifi- cial intelligence in the quantum domain: a review of re- cent progress, Reports on Progress in Physics81, 074001 (2018), publisher: IOP Publishing

  26. [26]

    Carleo, I

    G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a, Machine learning and the physical sciences, Reviews of Modern Physics91, 045002 (2019), publisher: American Physical Society

  27. [27]

    J. Zang, M. Medvidovi´ c, D. Kiese, D. D. Sante, A. M. Sengupta, and A. J. Millis, Machine learning-based com- pression of quantum many body physics: PCA and au- toencoder representation of the vertex function (2024), arXiv:2403.15372 [cond-mat]

  28. [28]

    L. M. Sager-Smith and D. A. Mazziotti, Reducing the Quantum Many-Electron Problem to Two Electrons with Machine Learning, Journal of the American Chemical So- ciety144, 18959 (2022), publisher: American Chemical Society

  29. [29]

    Carleo and M

    G. Carleo and M. Troyer, Solving the quantum many- body problem with artificial neural networks, Science 355, 602 (2017), publisher: American Association for the Advancement of Science

  30. [30]

    Sajjan, J

    M. Sajjan, J. Li, R. Selvarajan, S. H. Sureshbabu, S. S. Kale, R. Gupta, V. Singh, and S. Kais, Quantum machine learning for chemistry and physics, Chemical Society Re- views51, 6475 (2022), publisher: The Royal Society of Chemistry

  31. [31]

    Bengio, A

    Y. Bengio, A. Courville, and P. Vincent, Representa- tion Learning: A Review and New Perspectives (2014), arXiv:1206.5538 [cs]

  32. [32]

    Alshammari, J

    S. Alshammari, J. Hershey, A. Feldmann, W. T. Free- man, and M. Hamilton, I-Con: A Unifying Framework for Representation Learning (2025), arXiv:2504.16929 [cs]

  33. [33]

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, A Simple Framework for Contrastive Learning of Visual Representations (2020), arXiv:2002.05709 [cs]

  34. [34]

    Costa, G

    E. Costa, G. Scriva, and S. Pilati, Solving deep-learning density functional theory via variational autoencoders (2024), arXiv:2403.09788

  35. [35]

    S. J. D. Prince,Understanding Deep Learning(MIT Press, 2023) google-Books-ID: rvyxEAAAQBAJ

  36. [36]

    Chen and W

    S. Chen and W. Guo, Auto-Encoders in Deep Learn- ing—A Review with New Perspectives, Mathematics11, 1777 (2023), publisher: Multidisciplinary Digital Pub- lishing Institute

  37. [37]

    Møller, G

    F. Møller, G. Fern´ andez-Fern´ andez, T. Schweigler, P. d. Schoulepnikoff, J. Schmiedmayer, and G. Mu˜ noz- Gil, Learning Minimal Representations of Many-Body Physics from Snapshots of a Quantum Simulator (2025), arXiv:2509.13821 [quant-ph]

  38. [38]

    Frohnert and E

    F. Frohnert and E. van Nieuwenburg, Explainable rep- resentation learning of small quantum states, Machine Learning: Science and Technology5, 015001 (2024), pub- lisher: IOP Publishing. 8

  39. [39]

    Hubbard, Electron correlations in narrow energy bands, Proceedings of the Royal Society of London

    J. Hubbard, Electron correlations in narrow energy bands, Proceedings of the Royal Society of London. Series A. Mathematical and Physical Sciences276, 238 (1997), publisher: Royal Society

  40. [40]

    D. P. Arovas, E. Berg, S. A. Kivelson, and S. Raghu, The Hubbard Model, Annual Review of Condensed Matter Physics13, 239 (2022), publisher: Annual Reviews

  41. [41]

    L. W. Cheuk, M. A. Nichols, K. R. Lawrence, M. Okan, H. Zhang, E. Khatami, N. Trivedi, T. Paiva, M. Rigol, and M. W. Zwierlein, Observation of spatial charge and spin correlations in the 2D Fermi-Hubbard model, Sci- ence353, 1260 (2016), publisher: American Association for the Advancement of Science

  42. [42]

    J. P. F. LeBlanc, A. E. Antipov, F. Becca, I. W. Bu- lik, G. K.-L. Chan, C.-M. Chung, Y. Deng, M. Ferrero, T. M. Henderson, C. A. Jim´ enez-Hoyos, E. Kozik, X.- W. Liu, A. J. Millis, N. V. Prokof’ev, M. Qin, G. E. Scuseria, H. Shi, B. V. Svistunov, L. F. Tocchio, I. S. Tupitsyn, S. R. White, S. Zhang, B.-X. Zheng, Z. Zhu, and E. Gull, Solutions of the Two...

  43. [43]

    Koridon, F

    E. Koridon, F. Frohnert, E. Prehn, E. van Nieuwen- burg, J. Tura, and S. Polla, Learning density functionals from noisy quantum data, Machine Learning: Science and Technology6, 025020 (2025)

  44. [44]

    A. J. Coleman, Structure of Fermion Density Matrices, Reviews of Modern Physics35, 668 (1963), publisher: American Physical Society

  45. [45]

    L¨ owdin, Quantum Theory of Many-Particle Sys- tems

    P.-O. L¨ owdin, Quantum Theory of Many-Particle Sys- tems. I. Physical Interpretations by Means of Density Matrices, Natural Spin-Orbitals, and Convergence Prob- lems in the Method of Configurational Interaction, Phys- ical Review97, 1474 (1955)

  46. [46]

    L. H. Delgado-Granados, L. M. Sager-Smith, K. Tri- fonova, and D. A. Mazziotti, Machine Learning of Two-Electron Reduced Density Matrices for Many-Body Problems, The Journal of Physical Chemistry Letters16, 2231 (2025), publisher: American Chemical Society

  47. [47]

    Akiba, S

    T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, Optuna: A Next-generation Hyperparameter Opti- mization Framework, inProceedings of the 25th ACM SIGKDD International Conference on Knowledge Dis- covery & Data Mining, KDD ’19 (Association for Com- puting Machinery, New York, NY, USA, 2019) pp. 2623– 2631

  48. [48]

    H.-T. D. Liu, F. Williams, A. Jacobson, S. Fidler, and O. Litany, Learning Smooth Neural Functions via Lips- chitz Regularization (2022), arXiv:2202.08345 [cs]

  49. [49]

    Nakata, B

    M. Nakata, B. J. Braams, K. Fujisawa, M. Fukuda, J. K. Percus, M. Yamashita, and Z. Zhao, Variational calculation of second-order reduced density matrices by strong N-representability conditions and an accurate semidefinite programming solver, The Journal of Chem- ical Physics128, 164113 (2008)

  50. [50]

    D. A. Mazziotti, Variational minimization of atomic and molecular ground-state energies via the two-particle re- duced density matrix, Physical Review A65, 062511 (2002)

  51. [51]

    D. C. Liu and J. Nocedal, On the limited memory BFGS method for large scale optimization, Mathematical Pro- gramming45, 503 (1989)

  52. [52]

    Luise, C.-W

    G. Luise, C.-W. Huang, T. Vogels, D. P. Kooi, S. Ehlert, S. Lanius, K. J. H. Giesbertz, A. Karton, D. Gunceler, M. Stanley, W. P. Bruinsma, L. Huang, X. Wei, J. G. Torres, A. Katbashev, R. C. Zavaleta, B. M´ at´ e, S.-O. Kaba, R. Sordillo, Y. Chen, D. B. Williams-Young, C. M. Bishop, J. Hermann, R. v. d. Berg, and P. Gori-Giorgi, Accurate and scalable exc...

  53. [53]

    Kirkpatrick, B

    J. Kirkpatrick, B. McMorrow, D. H. P. Turban, A. L. Gaunt, J. S. Spencer, A. G. D. G. Matthews, A. Obika, L. Thiry, M. Fortunato, D. Pfau, L. R. Castellanos, S. Pe- tersen, A. W. R. Nelson, P. Kohli, P. Mori-S´ anchez, D. Hassabis, and A. J. Cohen, Pushing the frontiers of density functionals by solving the fractional electron problem, Science374, 1385 (2...

  54. [54]

    J. C. Snyder, M. Rupp, K. Hansen, K.-R. M¨ uller, and K. Burke, Finding Density Functionals with Machine Learning, Physical Review Letters108, 253002 (2012), publisher: American Physical Society

  55. [55]

    Nelson, R

    J. Nelson, R. Tiwari, and S. Sanvito, Machine learning density functional theory for the Hubbard model, Phys- ical Review B99, 075132 (2019), publisher: American Physical Society

  56. [56]

    Q. Sun, T. C. Berkelbach, N. S. Blunt, G. H. Booth, S. Guo, Z. Li, J. Liu, J. D. McClain, E. R. Sayfut- yarova, S. Sharma, S. Wouters, and G. K.-L. Chan, PySCF: the Python-based simulations of chemistry framework, WIREs Computational Molecular Science8, e1340 (2018)

  57. [57]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Rai- son, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, PyTorch: An Imperative Style, High-Performance Deep Learning Library, inAdvances in Neural Information Processing Sy...

  58. [58]

    D. P. Kingma and J. Ba, Adam: A Method for Stochastic Optimization (2017), arXiv:1412.6980 [cs]

  59. [59]

    Smilkov, N

    D. Smilkov, N. Thorat, B. Kim, F. Vi´ egas, and M. Wat- tenberg, SmoothGrad: removing noise by adding noise (2017), arXiv:1706.03825 [cs]

  60. [60]

    Simonyan, A

    K. Simonyan, A. Vedaldi, and A. Zisserman, Deep Inside Convolutional Networks: Visualising Image Classifica- tion Models and Saliency Maps (2014), arXiv:1312.6034 [cs]

  61. [61]

    Sundararajan, A

    M. Sundararajan, A. Taly, and Q. Yan, Axiomatic at- tribution for deep networks, CoRRabs/1703.01365 (2017), 1703.01365

  62. [62]

    Rozemberczki, L

    B. Rozemberczki, L. Watson, P. Bayer, H.-T. Yang, O. Kiss, S. Nilsson, and R. Sarkar, The Shapley Value in Machine Learning (2022), arXiv:2202.05594 [cs]

  63. [63]

    Lundberg and S.-I

    S. Lundberg and S.-I. Lee, A Unified Approach to Inter- preting Model Predictions (2017), arXiv:1705.07874 [cs]

  64. [64]

    D. A. Mazziotti, Structure of Fermionic Density Matri- ces: Complete n-Representability Conditions, Physical Review Letters108, 263002 (2012), publisher: American Physical Society

  65. [65]

    D. A. Mazziotti, Two-Electron Reduced Density Matrix as the Basic Variable in Many-Electron Quantum Chem- istry and Physics, Chemical Reviews112, 244 (2012), publisher: American Chemical Society. 9 Appendix A: Dataset Details To generate training data for the neural-network-based autoencoder, we sample Hubbard model ground states at varying random potenti...