{"id":"c55a62c6-aab5-4eaf-9b4d-7b951e6598b8","arxiv_id":"2507.10287","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Grassmann Variational Monte Carlo generalizes neural-network excited-state optimization to subspaces and accurately reproduces low-lying spectra of the 2D Heisenberg model.","lead":"A team at EPFL reformulated a method for computing excited quantum states using the geometry of linear subspaces, called Grassmannians, and applied it with neural network wave functions. The approach computes many excited-state energies and observables of the 2D Heisenberg magnet accurately, which could help simulations of quantum materials.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Determinant-sampling estimator in Eq. (13) lacks a conditioning/well-posedness guarantee; near-singular Φ(S) could bias the reported energies and V-scores.","rationale":"The paper's central claim is that Grassmann VMC gives accurate excited-state energies and observables on the 6×6 and 10×10 Heisenberg model. The 6×6 exact-diagonalization agreement is strong independent evidence that the estimator pipeline works in the tested regime, and the 10×10 V-scores are consistent with that. I therefore do not see a demonstrated fatal flaw. The load-bearing weak point is the absence of a well-posedness/robustness argument for the determinant-sampling inversion in Eq. (13): the entire numerical pipeline (OEMs, OVMs, QGT, overlap matrices) is built on inverting Φ(S), and a systematic near-singularity problem would corrupt all reported quantities. The paper neither bounds the estimator variance nor reports conditioning, and the claim of avoiding ill-conditioned inversions concerns the global Gram matrix, not the per-sample Φ(S). This is exactly the gap the reader flagged. It is addressable with a conditioning diagnostic and/or a finite-variance proof, which supports keeping the verdict CONDITIONAL rather than moving to ACCEPT or REJECT.","tokens_in":28046,"tokens_out":19477,"duration_ms":221225,"concrete_test":"Instrument the 6×6 benchmark to record the condition number κ(Φ(S)) for every sampled tuple during a representative optimization, and recompute the OEM/OVM estimates after truncating samples with κ > 10^6 (or after replacing Φ^{-1} by a pseudoinverse with threshold). If the energies and structure-factor values shift by more than the reported error bars, or if near-singular samples occur at measurable frequency, the estimator's stability is not established; if no high-κ samples appear and the truncated recomputation reproduces the published values, the concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec. III.B.3, the estimator eA(S) = Φ^{-1}(S) A(S) (Eq. 13) requires the N×N overlap matrix Φ(S) = [[S|Φ]] to be invertible for every sampled tuple S. The sampling distribution PΦ(S) ∝ |det Φ(S)|^2 (Eq. 11) assigns zero probability to exactly singular tuples, but it does not control near-singular ones: a tuple with det ~ ε contributes O(ε^2) to the normalization but O(1/ε) to the local matrix, giving O(1) per-sample variance contributions that can accumulate if near-singular tuples are frequent. The paper neither proves the estimator has finite variance over the finite configuration space nor reports conditioning diagnostics; the claimed 'avoids ill-conditioned Gram inversion' (Sec. I) is not the same as avoiding ill-conditioned Φ(S). If near-singular samples occur in the 10×10 runs, the V-scores and spin structure factors in Sec. V, which rely on the covariance estimators Eq. (14), could be biased or overly noisy. This is load-bearing because the central numerical claim depends on these estimators being reliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalizes the excited-state variational framework of Pfau et al. in terms of Grassmann geometry of Hilbert space. It promotes scalar quantum values (expectation values, variances, covariances, fidelities) to N×N matrices with basis-independent eigenvalues, defines a Grassmann quantum geometric tensor, and extends stochastic reconfiguration to the simultaneous optimization of N variational wave functions. The method is tested on the 2D Heisenberg model: on a 6×6 lattice the first four excited states in many momentum/spin-flip sectors are compared with exact diagonalization and found to have relative energy errors around or below 10^-4, and on a 10×10 lattice six excited states are reported with V-scores of order 10^-3 or below. Spin structure factors are also computed for both system sizes, with the ground-state values checked against a quantum Monte Carlo reference.","tokens_in":28282,"tokens_out":4330,"duration_ms":53228,"significance":"If the claims hold, this is a valuable and timely contribution: it gives a clean geometric language for multi-state variational Monte Carlo, generalizes stochastic reconfiguration and operator-variance estimators to subspaces, and provides a practical route to excited states of two-dimensional spin systems where tensor-network or sign-problem-based methods struggle. The paper's strengths are its careful derivations of determinant-sampling averages in the appendices, the clean 6×6 exact-diagonalization benchmark across many symmetry sectors, and the explicit comparison of the 10×10 ground-state energy and structure factor with quantum Monte Carlo. The main gap is that the 10×10 excited states are validated only through internal V-scores rather than independent benchmarks, and the determinant-sampling estimator has no conditioning analysis.","major_comments":[{"comment":"The local operator matrix eA(S)=Φ^{-1}(S)·A(S) requires the sampled N×N overlap matrix Φ(S) to be invertible, and the sampling probability PΦ(S)∝|det Φ(S)|^2/(N! det G) only excludes exactly singular tuples. A tuple with det Φ(S)~ε contributes probability O(ε²) but local matrix elements O(1/ε), so the estimator can have very large fluctuations or undefined values if near-singular tuples occur with non-negligible frequency. The paper neither proves finite variance over the finite configuration space nor reports condition-number or determinant diagnostics from the runs. Because the OVM/OCM estimates in Eq. (14), the V-scores in Fig. 3/Table II, and the spin structure factors in Sec. V all rely on these local matrices, this is load-bearing for the central numerical claim. Please add a conditioning analysis, a regularized inverse, or empirical determinant/condition-number diagnostics, or prove that the estimator has finite variance for the models considered.","section":"Sec. III.B.3, Eq. (13)"},{"comment":"The 10×10 excited-state results are supported only by V-scores, which are an internal consistency measure. The statement that for a given V-score the actual energy relative error is typically an order of magnitude smaller is a heuristic from Ref. [42] and is not established for excited states of this model; the ground state is checked against QMC, but no excited state has an independent benchmark. To support the abstract claim of 'highly accurate energies and physical observables for a large number of excited states', please provide external validation for at least a subset of the 10×10 excited states (e.g., DMRG or a finite-size-scaling comparison), or explicitly temper the accuracy claim to what V-scores can certify.","section":"Sec. V, Fig. 3 and Table II"},{"comment":"The claim that the method 'avoids to invert the possibly ill-conditioned Gram matrix of the basis states' is overstated. Equation (13) still requires inversion of the N×N overlap matrix Φ(S) for every sampled tuple, and near-singularity of Φ(S) is exactly the kind of ill-conditioning that the text says is bypassed. Please clarify the precise sense in which the method avoids Gram-matrix inversion and discuss the relationship between the conditioning of G and the conditioning of Φ(S).","section":"Sec. I vs. Sec. III.B.3"}],"minor_comments":[{"comment":"The software statement says the code 'will be made public in a later revision'; for a numerical-methods paper, full reproducibility would be helped by releasing the code or at least providing per-run hyperparameters, sample counts, and random seeds in the final version.","section":"Sec. VII"},{"comment":"The sentence 'the ground state of H with fermionic symmetry is the Slater determinant of the first N excited states' is ambiguous; it should read 'the N lowest-energy eigenstates' (including the ground state) to avoid implying that the ground state is built from the N-th through 2N-th levels.","section":"Sec. III.B.5"},{"comment":"The figures mix two spin-flip sectors (qsf=0,1) but the marker legend is implicit; please add explicit labels or a legend so the reader can identify which marker corresponds to which sector.","section":"Fig. 2 and Fig. 3"},{"comment":"There is a typo in 'the the bra matrix of its dual basis'.","section":"Sec. II.B"},{"comment":"The summation notation over S^{N-M} is introduced only in passing; please state explicitly whether the sums run over ordered tuples or sets and how the (N-M)! prefactor in Eq. (34) arises from that convention.","section":"Appendix 2, Eq. (34)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a quantum-simulation journal and the authors appropriately acknowledge concurrent work by Ma et al. and by Khahn et al. The main risk is the conditioning of the determinant-sampling estimator and the reliance on V-scores for the 10×10 excited states; both are addressable with additional analysis and validation. I do not see a fatal flaw in the derivations, but the central claim needs the proposed strengthening before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nQuick take: this is a solid, worthwhile paper that cleans up the geometric foundations of the Pfau et al. excited-state VMC and adds useful estimators. The 6x6 exact-diagonalization benchmark is convincing. The main soft spot is the unexamined conditioning of the per-sample overlap matrices, which the authors wave away with a claim that is slightly overstated. It is not fatal on the evidence here, but it needs an honest paragraph before publication.\n\nWhat is new: the Grassmann formulation is not just recasting. It gives a natural subspace generalization of stochastic reconfiguration, and the multidimensional variance/overlap estimators (OVM/OCM, fidelity matrices) are genuinely new relative to Pfau et al. The appendix derivations of the sampling averages are detailed and, as far as I can tell on a careful read, correct. The 6x6 results are clean: relative energy errors around 1e-4 or better across many momentum sectors, and the structure factor matches ED well. That is real evidence the framework works.\n\nWhere I would press: the determinant-sampling estimator in Eq. (13) requires inverting the N x N overlap matrix Phi(S) for every sampled tuple. The sampling distribution puts zero weight on exactly singular tuples, but nothing controls near-singular ones, and the paper gives no conditioning diagnostics or regularization. The claim that the method avoids ill-conditioned Gram inversion is true only in a narrow sense: it avoids inverting the N x N Gram matrix G, but it inverts a different N x N matrix at every sample. If near-singular samples are frequent, the covariance estimators in Eq. (14) could be noisy or biased, and the 10x10 V-scores are exactly the kind of quantity that would be affected. The 6x6 agreement with ED suggests that in practice the instability does not bite there, but the 10x10 results rest on V-scores alone. That is an addressable gap - report conditioning numbers or add a regularized inverse - not a fatal flaw.\n\nAlso, the code is not released. For a methods paper of this type, that is a real minus for reproducibility, and the 'later revision' promise should be enforced.\n\nWho it is for: people working on neural quantum states for excited states, and anyone who wants a clean geometric framing of multi-wavefunction VMC. It will be cited. It deserves a serious referee, with the conditioning question as the main thing to send back.","headline":"A worthwhile geometric formalization of excited-state VMC with clean 6x6 benchmarks; the unexamined conditioning of per-sample overlap matrices is the main soft spot, but not a fatal one.","tokens_in":28788,"tokens_out":2102,"would_cite":true,"duration_ms":24329,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that excited states of the two-dimensional Heisenberg model can be computed to about $10^{-4}$ relative error by performing variational Monte Carlo over a whole linear subspace of Hilbert space.","keywords":["excited states","Grassmannian","neural quantum states","variational Monte Carlo","stochastic reconfiguration","Heisenberg model","quantum many-body"],"falsifier":"Run Grassmann VMC on an $8\\times8$ Heisenberg lattice in a degenerate momentum sector while recording the smallest singular value of each sampled overlap matrix $\\Phi(S)$; if a noticeable fraction of samples are singular or ill-conditioned, the energy and variance estimators would become unstable or biased, which would show the claimed estimator is not well-defined as stated.","tokens_in":27845,"feed_emoji":"🧲","tokens_out":11283,"duration_ms":120933,"temperature":0.7,"pith_summary":"This paper claims that the accurate computation of excited states no longer requires symmetry constraints or penalty terms if the variational object is promoted from a single wavefunction to an $N$-dimensional subspace. It formalizes the determinant-sampling excited-state framework of [1] in the language of Grassmann geometry, turning the search for several excited states into one simultaneous geometric optimization. It generalizes Stochastic Reconfiguration to multiple wavefunctions and introduces matrix versions of variances, overlaps, and fidelities, all estimated by a single determinant sampling scheme. On the 2D Heisenberg model the method yields relative energy errors around $10^{-4}$ on a $6\\times6$ lattice and V-scores of order $10^{-3}$ on $10\\times10$, with spin structure factors matching benchmarks.","feed_headline":"Excited Heisenberg spectra computed to one part in 10,000","feed_subtitle":"Neural wave functions on whole subspaces reproduce exact excited-state energies and magnetic correlations.","key_machinery":"The central object is the Grassmannian $\\mathrm{Gr}_N(H)$, the manifold of $N$-dimensional linear subspaces of the Hilbert space, realized by the span of $N$ neural wavefunctions $\\Phi=(\\phi_1,\\dots,\\phi_N)$. The carrying identity is the wedge-product overlap $\\det[[\\Psi|\\Phi]]=\\langle\\langle\\Psi|\\Phi\\rangle\\rangle$, which turns a subspace into a single antisymmetric wavefunction and motivates determinant sampling, $P_\\Phi(S)\\propto |\\det[[S|\\Phi]]|^2/(N!\\det G)$. The local operator matrix $\\tilde A(S)=\\Phi^{-1}(S)A(S)$ is the estimator that avoids constructing and inverting the Gram matrix directly. The geometric optimization uses the Grassmann Quantum Geometric Tensor $S_{\\mu\\nu}=\\frac{1}{N}\\mathrm{Tr}\\big(G^{-1}[[\\partial_\\mu\\Phi|\\hat P_{V^\\perp}|\\partial_\\nu\\Phi]]\\big)$, whose inversion defines the Stochastic Reconfiguration update; operator covariance matrices built the same way supply variances and observables such as the spin structure factor.","core_discovery":"Central claim: Grassmann Variational Monte Carlo yields accurate excited states of many-body systems by optimizing a linear subspace of the Hilbert space, not a single wavefunction. The subspace $V$ is represented by a basis of $N$ neural wavefunctions $\\Phi=(\\phi_1,\\dots,\\phi_N)$, and the wedge-product identity $\\det[[\\Psi|\\Phi]]=\\langle\\langle\\Psi|\\Phi\\rangle\\rangle$ lets the whole subspace be sampled as a single antisymmetric state over $N$-tuples of configurations with probability $P_\\Phi(S)\\propto |\\det[[S|\\Phi]]|^2/(N!\\det G)$. All quantum values are promoted from scalars to matrices: the operator expectation matrix is $\\tilde A(\\Phi)=G^{-1}[[\\Phi|A|\\Phi]]$, its eigenvalues are the approximate excited-state energies, and its Monte Carlo estimator is the average of local matrices $\\tilde A(S)=\\Phi^{-1}(S)A(S)$. The same covariance estimators yield operator variance matrices, overlap and fidelity matrices between subspaces, and the Grassmann Quantum Geometric Tensor that defines Stochastic Reconfiguration. On the $6\\times6$ Heisenberg model the first four excited states in tested momentum and spin-flip sectors reach relative energy errors at or below about $10^{-4}$; on $10\\times10$ the V-scores are of order $10^{-3}$ and the ground-state energy density $-0.671544(4)$ agrees with quantum Monte Carlo.","pith_inferences":["I infer that the determinant-sampling estimator will need a practical singularity monitor in production use: one should track the smallest singular value of the sampled overlap matrix $\\Phi(S)$, and near-degenerate sectors may require regularization or a pseudo-inverse, since the paper does not address this case.","I expect the same subspace geometry to transfer to ab initio electronic structure, where the wedge-product picture matches Slater determinants; a natural test is whether simultaneous excited-state optimization with neural backflow surpasses single-reference quantum chemistry on small molecules.","I see a natural route to real-time dynamics through the new subspace overlap and fidelity matrices: rather than evolving a single state, one could project time evolution onto a moving Grassmann subspace and monitor leakage through the operator variance, a direction the authors mention through subspace-expansion methods.","I read the main computational bottleneck as the $N^2$ forward passes needed for the overlap matrix entries when using separate networks; shared-backbone architectures mitigate this, and further architectural sharing would determine how large $N$ can practically become."],"forward_implications":["On a $6\\times6$ Heisenberg lattice, the first four excited states in every tested momentum and spin-flip sector have relative energy errors at or below about $10^{-4}$ against exact diagonalization.","On a $10\\times10$ lattice, six excited states per sector are obtained with V-scores of order $10^{-3}$ or better, and the ground-state energy and $S(\\pi,\\pi)$ agree with quantum Monte Carlo.","Excited-state spin structure factors are read off directly from the diagonal of the operator variance matrix, so a single optimization run supplies both energies and correlation observables.","Stochastic Reconfiguration extends naturally to simultaneous optimization of several wavefunctions, so multiple eigenstates are produced in one run without penalty terms or hand-imposed symmetry constraints.","The formulation is written in second quantization and purely in terms of the Hamiltonian, which opens the same treatment to lattice fermions and frustrated spin systems."],"supporting_citations":[{"why":"Introduces the excited-state-as-ground-state-of-an-enlarged-Hamiltonian approach and the determinant sampling scheme that this paper formalizes in Grassmann-geometric terms.","marker":"[1]"},{"why":"Supplies the single-state Stochastic Reconfiguration method that the paper generalizes to simultaneous optimization of several variational wavefunctions.","marker":"[26]"},{"why":"Introduces neural quantum states, the variational ansatz class from which the basis wavefunctions are built.","marker":"[7]"},{"why":"Provides the foundational formulation of variational Monte Carlo that the Grassmann determinant sampling extends.","marker":"[25]"},{"why":"Provides the quantum Monte Carlo ground-state energy and spin-structure-factor benchmarks on the $10\\times10$ lattice.","marker":"[41]"},{"why":"Defines the V-score used to validate the accuracy of the $10\\times10$ excited-state energies.","marker":"[42]"},{"why":"Provides the exact diagonalization benchmarks against which the $6\\times6$ excited-state energies and structure factors are compared.","marker":"[52]"}],"fun_headline_variants":["Grassmann VMC computes excited Heisenberg spectra with neural wavefunctions","Accurate excited states: Grassmann VMC with neural subspaces on Heisenberg","Subspace neural wave functions solve Heisenberg excited states to 1e-4","Grassmann geometry in VMC: neural wave functions tackle excited states"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every sampled N-tuple of configurations gives an invertible overlap matrix, because the local operator estimator divides by that matrix, and the sampling distribution does not exclude near-singular tuples.","fun_headline_variants_meta":{"raw":{"variants":["Grassmann VMC computes excited Heisenberg spectra with neural wavefunctions","Accurate excited states: Grassmann VMC with neural subspaces on Heisenberg","Subspace neural wave functions solve Heisenberg excited states to 1e-4","Grassmann geometry in VMC: neural wave functions tackle excited states"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000447,"raw_usage":{"total_tokens":2276,"prompt_tokens":983,"completion_tokens":1293,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":1219}},"tokens_in":599,"tokens_out":1293,"duration_ms":13880,"temperature":1.0,"reasoning_tokens":1219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:34:44.837562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Grassmann VMC on an $8\\times8$ Heisenberg lattice in a degenerate momentum sector while recording the smallest singular value of each sampled overlap matrix $\\Phi(S)$; if a noticeable fraction of samples are singular or ill-conditioned, the energy and variance estimators would become unstable or biased, which would show the claimed estimator is not well-defined as stated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the excited-state-as-ground-state-of-an-enlarged-Hamiltonian approach and the determinant sampling scheme that this paper formalizes in Grassmann-geometric terms."},{"cited_title":"Nomura, A","cited_arxiv_id":null,"evidence_quote":"Supplies the single-state Stochastic Reconfiguration method that the paper generalizes to simultaneous optimization of several variational wavefunctions."},{"cited_title":"Its corresponding un- normalized amplitudes are given by: ⟨⟨S|Φ⟩⟩ = det [ [S|Φ] ]","cited_arxiv_id":null,"evidence_quote":"Introduces neural quantum states, the variational ansatz class from which the basis wavefunctions are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the foundational formulation of variational Monte Carlo that the Grassmann determinant sampling extends."},{"cited_title":"Park and M","cited_arxiv_id":null,"evidence_quote":"Defines the V-score used to validate the accuracy of the $10\\times10$ excited-state energies."}],"review_version":1}