{"id":"9d0db504-e082-4287-afad-523959659fe7","arxiv_id":"2504.21459","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Minimizing the trace of S^{-1}H over non-orthogonal variational states simultaneously finds several low-lying eigenstates, demonstrated with matrix product states, quantics tensor trains, and quantum circuits.","lead":"A new method finds multiple low-energy quantum states at once by minimizing a single function built from the overlaps and energies of a set of trial states. It is tested on spin chains, a molecular vibration, and a small Hubbard model with different ansatzes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The variational principle is sound, but the practical claim depends on optimization robustness and on S staying nonsingular; the paper's own Discussion concedes both, so the unconditional wording of the central claim is not yet supported.","rationale":"The reader's weakest_assumption identifies optimizer convergence and S non-singularity, which matches the main practical vulnerability. I partially agree because the concern can be sharpened: the domain of L excludes singular S, and the loss can have finite, path-dependent limits on the singular boundary, so the optimization problem is not well-posed without regularization or a proof that the minimum is interior. This is more than a generic 'optimization may fail' caveat; it directly affects whether the method can be used as an unconditional variational principle. However, the variational principle itself is mathematically correct, and the three demonstrations show that the approach works in the tested regimes. The novelty overstatement relative to Ref. [43] is a scope issue rather than a correctness issue. The absence of released code weakens reproducibility but does not invalidate the central claim. Therefore the appropriate verdict remains CONDITIONAL, unchanged from the reader's verdict: the central claim should be accepted conditionally on demonstrating robust optimization and S stability across more initializations, larger systems, and near-degenerate starting conditions.","tokens_in":16684,"tokens_out":11866,"duration_ms":137540,"concrete_test":"Numerically: rerun the 1D MPS benchmark (N=16, Ns=32, chi=16) but initialize two of the 32 MPS nearly identically, e.g., the same random tensors plus a relative perturbation of 1e-6. Track cond(S) and the relative errors of all 32 levels over 200k L-BFGS steps. If cond(S) grows beyond about 1e8 or the energy errors degrade by more than an order of magnitude relative to the random initialization runs, the nonsingularity/optimization assumption is load-bearing and explicit regularization is required. Analytically, use H=diag(0,1), Ns=2 with |psi1>=cos(theta)|0>+sin(theta)|1> and |psi2>=cos(theta)|0>+sin(theta)e^{i epsilon}|1>; compute the limit of L as epsilon approaches 0 to verify that the boundary is a degenerate minimizer with path-dependent behavior, demonstrating the need for a well-posedness constraint.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The mathematical core is correct: for any full-column-rank set of states, L = Tr(S^{-1}H) equals the sum of the Ritz values of the spanned subspace, and the minimax argument in SM I gives E_alpha >= E_exact_alpha, so the minimum over all subspaces is exactly the lowest Ns eigenvalues. The load-bearing gap is that the optimization over L is a non-convex problem whose feasible set is open: L is only defined when S is nonsingular. As the variational states approach linear dependence, L can have finite, path-dependent limits on the boundary (e.g., two states both approaching the ground state along different first-order directions), so the loss is not coercive and can develop flat directions or near-singular basins. The paper supplies no explicit regularization or proof that the global minimum is attained away from the singular boundary, and the Hubbard application reports only best-of-10 trials, while Supplement VI notes that the highest S condition number (Morse case, O(10^4)) correlates with occasional training instability. The paper's own Discussions state that convergence to the global minimum is not guaranteed and that S needs careful monitoring. Thus the strongest claim, phrased as 'the optimization drives the variational states to span the lowest-energy subspace,' is contingent on L-BFGS navigating this boundary landscape successfully.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes minimizing the loss L = Tr(S^{-1}H) over a set of N_s non-orthogonal, unnormalized variational states in order to simultaneously approximate the lowest N_s energy levels of a quantum Hamiltonian without explicit orthogonality constraints or penalty terms. The authors prove, via the minimax principle, that the Ritz values of the spanned subspace upper-bound the exact eigenvalues, so minimizing the sum of Ritz values targets the low-energy subspace. The method is demonstrated on three problems: a 16-site Heisenberg spin chain with periodic matrix product states, a Morse oscillator spectrum with quantics tensor trains, and a 2x3 Hubbard model with hardware-efficient variational quantum circuits, the last benchmarked against subspace VQE. The Supplementary Material contains the minimax proof, a Cholesky reduction of the generalized eigenproblem, a discussion of relations to prior subspace methods, and additional numerical results and implementation details.","tokens_in":16944,"tokens_out":14637,"duration_ms":157291,"significance":"The mathematical core is correct and standard: for any linearly independent set of states, L equals the sum of the Ritz values of the spanned subspace, and the minimax argument in SM I establishes the upper-bound property. If the optimization can be made reliable, the loss is an attractive ansatz-agnostic objective that avoids penalty terms and sequential orthogonality constraints, with natural compatibility with automatic differentiation. The numerical demonstrations span three different ansatz families and include checks against exact diagonalization, which is a genuine strength. The main unaddressed issues are the gap between the unconditional wording of the central claim and the actual optimizer behavior, and the limited statistical reporting of the quantum-circuit comparison. These are fixable within the scope of the manuscript.","major_comments":[{"comment":"The variational statement is mathematically correct for any nonsingular S, but the advertised behavior ('the optimization drives the variational states to span the subspace corresponding to the lowest possible sum') is a statement about the global minimizer, not about L-BFGS on a non-convex loss whose feasible set is open. Near the singular boundary of S the generalized eigenvalue problem becomes ill-conditioned, and the authors themselves note that convergence to the global minimum is not guaranteed and that S requires careful monitoring, with SM VI reporting condition numbers up to O(10^4) and occasional training instability. I ask that the main text state explicitly that the central claim holds when the optimization remains inside the nonsingular region and reaches a sufficiently good global minimum, and that the authors report the smallest eigenvalue or condition number of S along the optimization trajectory for each application. If no regularization is used, a short explanation of why the reported L-BFGS runs avoid the singular boundary would make the practical claim concrete.","section":"Method (Eq. 3) and Discussions"},{"comment":"The comparison with subspace VQE reports only the best of 10 trials in the energy-error insets and displays raw loss curves without aggregate statistics, so the sentence 'the variance of the loss across different trials is also smaller in our case' is not quantitatively supported. Similarly, SM VI mentions that the Morse condition number O(10^4) 'correlated with occasional training instability' but does not state how often or how severely the training failed. Please report the success rate and the mean/median and spread of converged losses and energy errors over all independent initializations for each application, and for the Hubbard benchmark show the full trial distribution for both methods. This is necessary to support the robustness and advantage claims.","section":"Hubbard results (Fig. 4) and SM VI"}],"minor_comments":[{"comment":"Reference [3] contains stray annotative text after the bibliographic entry ('1010.1992The work give the rebirth ...'); this should be removed.","section":"References"},{"comment":"Reference [72] lists 'G. K. lic Lic Chan'; the author name should be corrected to G. K.-L. Chan.","section":"References"},{"comment":"In the caption of Fig. 4, 'hardware effcient ansatz' should read 'hardware-efficient ansatz'.","section":"Fig. 4 caption"},{"comment":"The phrase 'the loss is shifted by the sum of exact eigenvalues' is ambiguous about the sign of the shift; specify that the plotted quantity is L minus the sum of the exact eigenvalues.","section":"Fig. 2(a) and Fig. 3 captions"},{"comment":"The statement 'All 2NsNchi^2 trainable parameters' would be clearer as 2 N_s N chi^2, with the power of chi indicated explicitly.","section":"Spin-chain results"}],"recommendation":"major_revision","confidential_remarks":"The mathematical content is sound and the benchmarks are useful, but the paper's abstract and introduction state the practical claim more strongly than the optimization evidence warrants. The lack of trial statistics for the Hubbard comparison and the absence of explicit S-regularity diagnostics are the main barriers. I would support publication after a revision that qualifies the central claim and adds the requested diagnostics; I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Martin—the headline: the variational principle in this paper is not new. Minimizing Tr(S^{-1}H) over a subspace of non-orthogonal states is Rayleigh-Ritz, and Pfau et al. (Ref [43]) already used a mathematically equivalent loss for neural quantum states. The paper admits this in the supplement and in the Discussions, which is good. What is actually new is the systematic application of this loss across three very different ansatze: periodic MPS for spin chains, quantics tensor trains for a Morse oscillator, and hardware-efficient VQE circuits for a 2x3 Hubbard model. That demonstration is clean and useful. The results are strong: relative errors on the order of 1e-7 for the 1D chain with bond dimension 16, 1e-10 with chi=24, 1e-5 for the Morse potential, and the Hubbard comparison shows better loss and smaller variance than subspace VQE. The math is standard and correct: the minimax proof in SM I is rigorous, and the generalized eigenvalue post-processing is straightforward.\n\nSoft spots: the abstract calls the principle 'novel,' which is an overstatement given the prior literature; that language should be dialed back. The central claim, that minimizing L drives the states to span the lowest-energy subspace, is stated unconditionally, but the paper itself concedes that the optimization is non-convex and that S can become ill-conditioned. In practice L-BFGS handles the tested cases, and the condition numbers they report in SM VI are mostly benign (O(1e2), with O(1e4) for Morse, which they note correlates with occasional instability). That is a real caveat, but it is not a defect in the variational principle; it is the usual heuristic nature of variational optimization. The Hubbard comparison reports only the best of 10 trials in the inset, though the training curves for all 10 runs are shown. A median and spread would be better than a best, but this is minor. There is also no standalone code/data release, though the implementation uses the open-source TensorCircuit-NG.\n\nVerdict: this paper deserves a serious referee. It is honest, the numerics are reproducible by anyone with the library, and the cross-ansatz demonstration is a useful reference for anyone wanting to compute excited states. I would ask for minor revisions: temper the novelty claim, add error bars or full trial distributions, and add a short paragraph on robustness (condition number evolution and multiple seeds) in the Hubbard section. I'd bring it to the reading group as a nice practical example of a known principle.","headline":"A mathematically correct but known variational principle, made useful by clean demonstrations across MPS, quantics tensor train, and VQE; the main gap is the overclaimed novelty and the heuristic robustness of the optimizer.","tokens_in":17497,"tokens_out":3520,"would_cite":true,"duration_ms":36997,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Minimizing $\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$ over non-orthogonal states simultaneously returns the lowest $N_s$ energy levels of a Hamiltonian, with no orthogonality constraints.","keywords":["quantum excited states","variational principle","non-orthogonal variational states","generalized eigenvalue problem","matrix product states","quantics tensor train","variational quantum eigensolver","low-energy spectrum"],"falsifier":"On a small exactly solvable system, run many optimizations with an ansatz that provably contains the lowest $N_s$ eigenstates; if the minimum of $\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$ found over all restarts stays strictly above the sum of the exact $N_s$ eigenvalues, or if the computed states have nonzero energy variance, the claim that this loss locates the low-energy subspace is not holding in practice.","tokens_in":16465,"feed_emoji":"⚛️","tokens_out":12656,"duration_ms":117308,"temperature":0.7,"pith_summary":"The paper proposes that $\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$, built from a set of deliberately non-orthogonal variational states, is a variational target whose minimization returns the lowest $N_s$ energy levels of a quantum Hamiltonian in a single optimization run. Because the loss equals the sum of the Ritz values of the subspace generalized eigenproblem, minimizing it pulls the whole spanned subspace down toward the lowest $N_s$ exact eigenstates, with the minimax principle guaranteeing these variational energies stay above the exact ones. The authors demonstrate the principle on a 1D spin chain with matrix product states, a Morse oscillator with quantics tensor trains, and a 2D Hubbard model with variational quantum circuits, recovering multiple low-lying levels accurately in each case. The appeal is that excited states, usually obtained by state-specific optimizations or penalty terms, come from one unconstrained loss that works for any variational wavefunction family.","feed_headline":"One loss function returns a quantum spectrum in a single run","feed_subtitle":"Minimizing the inverse-overlap trace skips orthogonality penalties and reproduces spectra on three platforms.","key_machinery":"The load-bearing object is the loss $L=\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$ formed from the $N_s\\times N_s$ overlap and Hamiltonian matrices of non-orthogonal variational states. When the states are linearly independent, $\\mathbf{S}$ is Hermitian positive definite, so the trace is well defined and equals the sum of the generalized eigenvalues (Ritz values) of $\\mathbf{H}\\mathbf{c}=E\\mathbf{S}\\mathbf{c}$. The minimax, or Rayleigh-Ritz, principle supplies the variational guarantee that these Ritz values are upper bounds on the exact low-lying eigenvalues, so reducing their sum pushes the span of the states toward the exact low-energy subspace. Around this core, the paper's machinery includes tensor-network contractions and automatic differentiation for matrix product states and quantics tensor trains, measurement protocols for circuits on quantum hardware, and a final generalized diagonalization to read out orthonormal eigenstates.","core_discovery":"The central claim is that a set of $N_s$ independent, non-orthogonal, unnormalized variational states can be optimized simultaneously to approximate the lowest $N_s$ eigenstates of a Hamiltonian by minimizing $L=\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$, where $\\mathbf{H}_{ij}=\\langle\\psi_i|H|\\psi_j\\rangle$ and $\\mathbf{S}_{ij}=\\langle\\psi_i|\\psi_j\\rangle$. The loss is exactly the sum $E_1+\\cdots+E_{N_s}$ of the Ritz values obtained from the generalized eigenvalue problem $\\mathbf{H}\\mathbf{c}=E\\mathbf{S}\\mathbf{c}$ in the variational subspace, and each Ritz value is an upper bound on the corresponding exact eigenvalue by the minimax principle. The optimization therefore drags the entire low-energy subspace downward instead of targeting one state at a time, and after convergence the final orthonormal eigenstates are obtained by one post-optimization diagonalization of the generalized eigenproblem. The paper's numerical demonstrations on spin chains, a Morse potential, and the Hubbard model are offered as evidence that the principle is accurate and transferable across ansatzes.","pith_inferences":["Because the loss weights every targeted level equally, a weighted variant $\\mathrm{Tr}(\\mathbf{W}\\mathbf{S}^{-1}\\mathbf{H})$ could rebalance accuracy between the lowest and highest states in the window; this is a testable extension the paper does not discuss.","The reliance on $\\mathbf{S}^{-1}$ suggests that near-degenerate spectra or larger $N_s$ will make the overlap matrix increasingly ill-conditioned, so explicit regularization or periodic subspace reorganization may be needed outside the tested regimes.","Building the spectrum incrementally by adding one variational state at a time and warm-starting from the converged subspace would turn the method into an iterative spectrum-construction algorithm, an extension not explored here."],"forward_implications":["One optimization run replaces $N_s$ separate state-specific calculations, because the inverse overlap couples all variational states into a single loss.","No penalty coefficients or orthogonality constraints need tuning; the tested runs kept the overlap matrix well conditioned without explicit regularization.","The same loss transfers across wavefunction families: matrix product states, quantics tensor trains, and parameterized quantum circuits all converge to accurate low-lying spectra.","On quantum hardware, the required $\\mathbf{H}$ and $\\mathbf{S}$ entries can be obtained with standard non-orthogonal measurement protocols, making simultaneous multi-state excited-state computation a single variational loop.","The framework is positioned to connect with neural quantum states through an equivalent variational-Monte-Carlo formulation, which the paper identifies as a promising direction."],"supporting_citations":[{"why":"Contains the minimax-principle proof that each Ritz value from the variational subspace is an upper bound on the exact eigenvalue, the mathematical foundation of the loss.","marker":"[58]"},{"why":"Supplies the trace-of-inverse-overlap-times-Hamiltonian loss form from machine-learning contexts that the paper adapts to quantum excited states.","marker":"[41, 42]"},{"why":"Shows a mathematically equivalent, more efficient version of this loss in neural-network variational Monte Carlo, cited as the path to neural quantum states.","marker":"[43]"},{"why":"Provides differentiable programming of tensor networks, used to compute gradients of the loss for matrix product states and quantics tensor trains.","marker":"[59]"},{"why":"Gives the non-orthogonal measurement protocols for the $\\mathbf{H}$ and $\\mathbf{S}$ matrix elements used in the variational quantum circuit application.","marker":"[67]"},{"why":"Defines the subspace-VQE baseline whose loss landscape and excited-state errors are compared in the Hubbard model benchmark.","marker":"[23]"},{"why":"Supplies the periodic matrix-product-state dispersion benchmark whose reported accuracy the method matches at larger bond dimension.","marker":"[6]"}],"fun_headline_variants":["Trace inverse overlap to get many excited states at once","Variational principle computes whole spectrum in one go","No more orthogonality penalties: one loss for all excited states","Simultaneous excited states via Tr(S^-1H) minimization","Minimize Tr(S^-1H) for entire low-energy spectrum"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's practical success rests on gradient-based optimization finding a good minimum of a non-convex loss while the overlap matrix stays invertible, a convergence and stability behavior the paper observes empirically but does not guarantee.","fun_headline_variants_meta":{"raw":{"variants":["Trace inverse overlap to get many excited states at once","Variational principle computes whole spectrum in one go","No more orthogonality penalties: one loss for all excited states","Simultaneous excited states via Tr(S^-1H) minimization","Minimize Tr(S^-1H) for entire low-energy spectrum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000906,"raw_usage":{"total_tokens":3961,"prompt_tokens":1075,"completion_tokens":2886,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":691,"completion_tokens_details":{"reasoning_tokens":2802}},"tokens_in":691,"tokens_out":2886,"duration_ms":21160,"temperature":1.0,"reasoning_tokens":2802,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:03:11.530229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a small exactly solvable system, run many optimizations with an ansatz that provably contains the lowest $N_s$ eigenstates; if the minimum of $\\mathrm{Tr}(\\mathbf{S}^{-1}\\mathbf{H})$ found over all restarts stays strictly above the sum of the exact $N_s$ eigenvalues, or if the computed states have nonzero energy variance, the claim that this loss locates the low-energy subspace is not holding in practice.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contains the minimax-principle proof that each Ritz value from the variational subspace is an upper bound on the exact eigenvalue, the mathematical foundation of the loss."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a mathematically equivalent, more efficient version of this loss in neural-network variational Monte Carlo, cited as the path to neural quantum states."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the subspace-VQE baseline whose loss landscape and excited-state errors are compared in the Hubbard model benchmark."},{"cited_title":"Pirvu, J","cited_arxiv_id":null,"evidence_quote":"Supplies the periodic matrix-product-state dispersion benchmark whose reported accuracy the method matches at larger bond dimension."}],"review_version":1}