REVIEW 4 major objections 4 minor 18 references
Direct unconstrained optimization of excited states in density functional theory
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A continuous log-det penalty converts excited-state DFT into a guaranteed-convergent unconstrained optimization, preventing variational collapse.
desk verdict A genuinely new unconstrained excited-state optimization scheme with closed-form gradients, but its accuracy claims outrun the evidence on non-minimum states and one misleading MAD. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the interstate penalty $\Omega^P = -C_P \sum_{\tau} \ln \det(\Phi_\tau \Phi_{\tau d}^{-1})$, built from the matrix $\Phi_\tau$ whose entries are the determinants of overlap matrices between occupied orbitals of different states. This penalty is zero for orthogonal states, becomes positively infinite when any two states become linearly dependent, and its strength $C_P$ is multiplied by a factor $f_{\mathrm{upd}}$ in an outer loop until the deviation from orthogonality falls below a threshold. Its analytical gradient contains the adjugate of the state-overlap matrix, which is evaluated through the singular value decomposition so that ill-conditioned, nearly orthogonal overlaps remain numerically stable. Combined with the intrastate gradient from prior VM SCF work [54], this makes the orbital coefficients free variables of a well-behaved unconstrained minimization, which is what guarantees convergence without collapse.
What would settle it
Take a converged VM TIDFT solution for a state with a large residual energy gradient (e.g., He $1s^2\rightarrow 1s2s$, gradient $7.9\times10^{-2}$ Ha), then set the interstate penalty strength $C_P$ to zero and re-optimize the orbitals; if the state collapses to a lower state or its energy shifts by more than the expected DFT error, the reported excited-state energy depends on the penalty rather than being an intrinsic property of the DFT functional.
Extended reading notes
Core claim
The central claim is that the variational collapse of excited states in orbital-optimized DFT can be prevented by a simple continuous penalty that drives the states to orthogonality without ever imposing it as a hard constraint. The loss for state $I$ is $\Omega = \bar{E}_I + \Omega^P$, where the intrastate term $\bar{E}_I$ contains the energy and a determinant-based penalty for linear independence of orbitals, and the interstate term $\Omega^P = -C_P \sum_{\tau=\alpha,\beta} \ln \det(\Phi_\tau \Phi_{\tau d}^{-1})$ is zero for orthogonal states and positively infinite for linearly dependent states. With this penalty, molecular orbital coefficients become independent variables; the gradient has closed-form expressions, with the interstate gradient involving the adjugate of the MO overlap matrix computed stably via SVD, so any unconstrained optimization algorithm can be used. The paper shows with a preconditioned conjugate gradient optimizer that the procedure converges robustly, producing accurate energies for ordinary excitations and, unlike TDDFT, for charge-transfer and double-electron excitations. Importantly, the fully optimized states are stationary points of the loss, not necessarily minima of the DFT energy; the paper argues, citing the Perdew–Levy extrema work [75], that such stationary points still correspond to physically meaningful excited states.
Load-bearing premise
The load-bearing premise is that a stationary point of the penalized loss functional is still a physically meaningful excited state even when it is not a minimum of the DFT energy; if that premise fails, the reported energies for high-gradient states would be artifacts of the penalty rather than true DFT predictions.
Editorial extensions
If this is right
- VM TIDFT reproduces the long-range Coulomb asymptote of charge-transfer excitation energies in H$_2$ and HeH$^+$, where TDDFT underestimates the tail.
- Double-electron excitations, which TDDFT cannot describe, are obtained directly, e.g., He $1s^2 \rightarrow 2s^2$ at 57.68 eV versus 57.84 eV experimental.
- Because the gradient is analytic and closed-form, the method can in principle use quasi-Newton and trust-region optimizers in addition to the preconditioned conjugate gradient used here, with the same guarantee against collapse.
- The analytic gradient makes excited-state atomic forces straightforward, facilitating computation of emission spectra and photophysical and photochemical dynamics.
- The fully optimized states are stationary points of the loss; for systems where the energy gradient is nonzero, the penalty exactly compensates it, so the states are kept from collapsing even though they are not DFT energy minima.
Reading between the lines
- The penalty strategy is not tied to DFT: any variational method that suffers from collapse to lower solutions, such as multiconfigurational SCF or nonorthogonal configuration interaction, could adopt the same log-determinant penalization to stabilize excited-state searches.
- For states with large residual energy gradients (e.g., Table 2, He $1s^2\rightarrow 1s2s$ with gradient $7.9\times10^{-2}$ Ha), the reported energies depend on the assumption that the stationary point of the penalized loss is a physical state; a direct test is to turn off the penalty at the converged point and see whether the state remains stationary or collapses.
- The outer-loop penalty schedule (doubling $C_P$ until DFO $< 10^{-3}$) could be replaced by a dynamic update inside the optimizer, potentially avoiding the sharp gradient spikes observed when the penalty is abruptly strengthened.
- Because the interstate cost grows as $O(S^2 B^2 N)$, applications to dense excited-state manifolds with hundreds of states would need an alternative scaling or a sparsity approximation for the state-overlap matrices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces variable-metric time-independent DFT (VM TIDFT), an orbital-optimized DFT method for excited states that avoids variational collapse by penalizing non-orthogonality between the target state and previously optimized states. The loss functional (Eq. 1) combines the state energy with an intrastate log-determinant penalty and an interstate log-determinant penalty (Eqs. 2-3). Analytical gradients for the interstate penalty are derived (Eq. 10), and a preconditioned conjugate-gradient optimizer is used (Eq. 13, Table 1). The method is implemented in CP2K and tested on atoms and molecules, including charge-transfer curves for H2 and HeH+ and double excitations in He. The paper claims accurate excitation energies for well-behaved, charge-transfer, and double-electron excitations, with MADs of 1.02 eV (first six systems) and 0.88 eV (all listed systems with experimental values).
Significance. If the central accuracy claims hold, VM TIDFT is a valuable addition to the OODFT toolbox: it provides closed-form gradients, permits direct unconstrained optimization with standard algorithms, is straightforwardly extensible to spin-purified ROKS, and reproduces the correct long-range behavior of CT excitations where TDDFT fails. The analytical gradient derivation, the SVD-based adjugate algorithm for ill-conditioned overlaps, and the convergence tests across many systems are concrete strengths. The paper also deposits coordinates, which aids reproducibility. However, the significance is tempered by the incomplete support for states that are not minima of the DFT energy functional and by reporting inconsistencies that affect the claimed accuracy.
major comments (4)
- [§Accuracy, Table 3] The text reporting MAD = 0.88 eV for 'all test systems with available experimental excitation energies' is only correct if the helium row is excluded from the average; with the error of 5.41 eV for helium included, the MAD over the 16 listed systems with experimental values is about 1.16 eV. The table lists helium with an experimental value and no exclusion is stated, so the reported MAD is misleading and should be corrected or explicitly justified.
- [Double-electron excitations] The claim that the VM TIDFT energies for the He 1s2→2s2 and 1s2→2s12p1 transitions are 'close to the experimental values' is not supported by the numbers: the reported values of 67.63 eV (QZV3P) and 65.79 eV (aug-cc-pV5Z) differ from the experimental 60.15 eV by 5.6–7.5 eV, which is not close. This overstatement should be removed, and the large error for this high-gradient state should be acknowledged and analyzed.
- [Penalty strength adjustment / Results (penalty effect)] For states that are not stationary points of the DFT energy E (He in Table 2; N-phenylpyrrole, NH3...BF3, and pyrrole-pyrazine in Table 3), the converged loss minimum satisfies ∇E = −∇ΩP, not ∇E = 0. The paper invokes Perdew–Levy (Ref. 75) to assert that such points are physically meaningful, but it does not demonstrate that the penalty gradient approximates the exact constraint force or that the finite-DFO solutions approach the orthogonality-constrained OCDFT solution in the limit E0→0, CP→∞. A numerical demonstration (e.g., comparing VM TIDFT with OCDFT for He, or showing convergence of the reported energy as E0 is tightened) is needed to support the accuracy claim for non-minimum states.
- [Abstract / Conclusions] The abstract states that the method 'guarantees convergence of the excited-state optimization.' As the Conclusions note, the optimization finds a stationary point in the basin of the initial guess and does not guarantee the lowest-energy excitation; moreover, convergence of the loss does not guarantee convergence to a physical excited state for non-minimum states. Please soften or qualify this claim.
minor comments (4)
- [Table 3] The 'Molecules' column contains a typo: 'formadehyde' should be 'formaldehyde'.
- [Fig. 1 caption] Please identify which He transition is the 1s→2s curve and which is the 1s2→2s2 curve, as the text refers to both without a clear mapping in the figure.
- [Table 1 / Fig. 1 discussion] Table 1 lists the default DFO threshold E0 as 10−3, but the text following Fig. 1 suggests E0 = 10−4 is acceptable; please reconcile these values.
- [Computational cost] The text contains a typo: 'prportional' should be 'proportional', and 'a large number of excited state' should be 'a large number of excited states'.
Circularity Check
No significant circularity: the excited-state energies are outputs of an unconstrained penalty optimization benchmarked against external methods, not fitted to them; the self-citations are building blocks and are not load-bearing.
full rationale
The core derivation is not circular. The loss functional Ω = E^I + Ω^P (Eq. 1) is a standard penalty construction: Ω^P drives interstate orthogonality, and the reported excitation energies are the DFT energy E^I at the converged orbitals, not parameters adjusted to reproduce the benchmark values. The penalty parameters (C_P^(0), f_upd, E_0 in Table 1, Eqs. 14–15) are convergence controls; Fig. 1 explicitly studies how excitation energy varies with the allowed deviation from orthogonality rather than tuning to experiment. The formulas taken from the authors' prior work (Eq. 8, Eq. 13, Refs. 52–54) are restated in the paper and serve as computational building blocks; they do not contain the excited-state orthogonality result and are not used to force agreement with the external benchmarks. The physical-interpretation argument for high-gradient states relies on the external Perdew–Levy result (Ref. 75), and the paper explicitly identifies the cases where the converged point is not a minimum of E alone (Tables 2–3), disclosing rather than concealing the penalty balance. Comparisons to TDDFT, OCDFT, MP2, CCSD, and experiment are external and independent. Concerns about the meaning of energies at non-stationary points of E, and about the reported MADs, are scientific-validity and transparency issues, not circularity. The only notable citation pattern is the use of the authors' own variable-metric formulation, but it is not load-bearing for the central excited-state claim.
Assumptions & free parameters
free parameters (4)
- Interstate penalty strength CP =
10 Ha initial, updated by factor 2
- Deviation-from-orthogonality threshold E0 =
1e-3 default; 1e-5 used for Table 2
- Intrastate penalty coefficient c_p
- Preconditioner regularization kappa =
user-adjustable
assumptions (4)
- domain assumption Kohn-Sham DFT with PBE and an unrestricted single-determinant wavefunction provides an adequate approximate energy surface for the tested excited states.
- domain assumption Stationary points of the penalized loss functional, even when not stationary points of the electronic energy, correspond to physically meaningful excited states.
- ad hoc to paper As CP increases, the minimum of the penalized loss approaches the orthogonality-constrained excited state.
- standard math The adjugate and SVD linear-algebra identities in Eq. (11) remain valid and stable for ill-conditioned overlap matrices.
Cite this review
Pith. "Pith review of Direct unconstrained optimization of excited states in density functional theory." pith.science (2026). https://pith.science/paper/BBOZ3LSU
@misc{pith2026250110907,
author = {Pith},
title = {Pith review of: Direct unconstrained optimization of excited states in density functional theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/BBOZ3LSU}},
note = {Machine review of arXiv:2501.10907}
}
read the original abstract
Orbital-optimized density functional theory (DFT) has emerged as an alternative to time-dependent (TD) DFT capable of describing difficult excited states with significant electron density redistribution, such as charge-transfer, Rydberg, and double-electron excitations. Here, a simple method is developed to solve the main problem of the excited-state optimization -- the variational collapse of the excited states onto the ground state. In this method, called variable-metric time-independent DFT (VM TIDFT), the electronic states are allowed to be nonorthogonal during the optimization but their orthogonality is gradually enforced with a continuous penalty function. With nonorthogonal electronic states, VM TIDFT can use molecular orbital coefficients as independent variables, which results in a closed-form analytical expression for the gradient and allows to employ any of the multiple unconstrained optimization algorithms that guarantees convergence of the excited-state optimization. Numerical tests on multiple molecular systems show that the variable-metric optimization of excited states performed with a preconditioned conjugate gradient algorithm is robust and produces accurate energies for well-behaved excitations and, unlike TDDFT, for more challenging charge-transfer and double-electron excitations.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[4]
(47) Barca, G. M.; Gilbert, A. T.; Gill, P. M. Simple Mod- els for Difficult Electronic Excitations. J. Chem. Theory Comput. 2018, 14, 1501–1509. (48) Barca, G. M.; Gilbert, A. T.; Gill, P. M. Excitation Num- ber: Characterizing Multiply Excited States. J. Chem. Theory Comput. 2018, 14, 9–13. (49) Mewes, J. M.; Jovanovi´ c, V.; Marian, C. M.; Dreuw, A. On...
work page 2018
-
[5]
Line search steps, one per PCG step, are not included in the iteration count
Convergence rate for the VM TIDFT procedure for the ground states (GS) and selected excited states (ES). Line search steps, one per PCG step, are not included in the iteration count. The dashed lines show the SCF convergence threshold. ment the full variational ROKS optimization of spin- purified open-shell singlet states 54–56, which are nec- essary for ...
work page 2017
-
[24]
Direct unconstrained optimization of excited states in density functional theory
There has been a number of methods proposed to avoid variational collapse of high-energy excited states to the states with lower energies27–44. These methods have been recently reviewed6,30,45,46 and only a brief description of some of them is presented here. In the maximum overlap method (MOM) 28 and ini- tial maximum overlap method (IMOM) 47,48, the ele...
work page Pith review arXiv 2025
-
[49]
Quantum chemistry of the excited state: 2005 overview
(2) Serrano-Andr´ es, L.; Merch´ an, M. Quantum chemistry of the excited state: 2005 overview. J. Mol. Struct. Theochem. 2005, 729, 99–108. (3) Schlag, E. W.; Schneider, S.; Fischer, S. F. Lifetimes in Excited States. Annu. Rev. Phys. Chem.1971, 22, 465–
work page 2005
-
[370]
(31) Ziegler, T.; Krykunov, M.; Cullen, J. The implementa- tion of a self-consistent constricted variational density functional theory for the description of excited states. J. Chem. Phys. 2012, 136, 124107. (32) Evangelista, F. A.; Shushkov, P.; Tully, J. C. Orthogo- nality Constrained Density Functional Theory for Elec- tronic Excited States. J. Phys. C...
work page 2012
-
[526]
Single-reference ab initio methods for the calculation of excited states of large molecules
(4) Dreuw, A.; Head-Gordon, M. Single-reference ab initio methods for the calculation of excited states of large molecules. Chem. Rev. 2005, 105, 4009–4037. (5) Runge, E.; Gross, E. K. Density-Functional Theory for Time-Dependent Systems. Phys. Rev. Letters1984, 52,
work page 2005
-
[997]
Orbital Optimized Den- sity Functional Theory for Electronic Excited States
(6) Hait, D.; Head-Gordon, M. Orbital Optimized Den- sity Functional Theory for Electronic Excited States. J. Phys. Chem. Lett.2021, 12, 4517–4529. (7) Ullrich, C. A. Time-Dependent Density-Functional The- ory: Concepts and Applications ; Oxford University Press,
work page 2021
-
[1703]
Pseudopotentials for H to Kr Optimized for Gradient-Corrected Exchange-Correlation Functionals
(72) Krack, M. Pseudopotentials for H to Kr Optimized for Gradient-Corrected Exchange-Correlation Functionals. Theor. Chem. Acc.2005, 114, 145–152. (73) Genovese, L.; Deutsch, T.; Neelov, A.; Goedecker, S.; Beylkin, G. Efficient solution of Poisson’s equation with free boundary conditions. J. Chem. Phys. 2006, 125, 074105. (74) Genovese, L.; Deutsch, T.; ...
work page 2005
Show all 18 references
-
[1721]
A.; Gwin, E.; Neuscamman, E
(51) Shea, J. A.; Gwin, E.; Neuscamman, E. A Generalized Variational Principle with Applications to Excited State Mean Field Theory. J. Chem. Theory Comput2020, 16, 1526–1540. (52) Luo, Z.; Khaliullin, R. Z. Direct Unconstrained Variable- Metric Localization of One-Electron Or...
2020
-
[2006]
(58) A Survey of Nonlinear Conjugate Gradient Methods. Pac. J. Optim.2006, 2, 35–58. (59) Stewart, G. W. On the adjugate matrix. Linear Algebra Its Appl. 1998, 283, 151–164. (60) VandeVondele, J.; Hutter, J. An efficient orbital trans- formation method for electronic structure...
2006
-
[2011]
(8) Levy, M.; ´Agnes Nagy Variational density-functional theory for an individual excited state. Phys. Rev. Lett. 1999, 83,
1999
-
[2821]
T.; Besley, N
(28) Gilbert, A. T.; Besley, N. A.; Gill, P. M. Self-consistent field calculations of excited states using the maximum overlap method (MOM). J. Phys. Chem. A2008, 112, 13164–13171. (29) Baruah, T.; Pederson, M. R. DFT calculations on charge-transfer states of a carotenoid-porp...
2009
-
[3651]
(45) Ziegler, T.; Krykunov, M.; Seidu, I.; Park, Y. C. In Con- stricted variational density functional theory approach to the description of excited states; Michael, Nicolas, H.-R. M. F., Filatov, Eds.; Springer International Publishing, 2015; Vol. 368; pp 61–95. (46) Cernatic...
2015
-
[4361]
Variational density-functional theory for degenerate excited states
(9) Nagy; Levy, M. Variational density-functional theory for degenerate excited states. Phys. Rev. A: At. Mol. Opt. Phys. 2001, 63, 052502. (10) Ayers, P. W.; Levy, M. Time-independent (static) density-functional theories for pure excited states: Ex- tensions and unification. ...
2001 arXiv
-
[4365]
Generalized Gra- dient Approximation Made Simple
(61) K¨ uhne, T. D. et al. CP2K: An Electronic Structure and Molecular Dynamics Software Package -Quickstep: Effi- cient and Accurate Electronic Structure Calculations. J. Chem. Phys. 2020, 152, 194103. (62) Borˇ stnik, U.; VandeVondele, J.; Weber, V.; Hutter, J. Sparse matrix...
1996
-
[4462]
M.; McDonald, D
(89) Behlen, F. M.; McDonald, D. B.; Sethuraman, V.; Rice, S. A. Fluorescence spectroscopy of cold and warm naphthalene molecules: Some new vibrational assign- ments. J. Chem. Phys.1981, 75, 5685–5693. (90) Druzhinin, S. I.; Kovalenko, S. A.; Senyushkina, T. A.; Demeter, A.; Z...
1981 arXiv
-
[5581]
(54) Pham, H. D. M.; Khaliullin, R. Z. Direct Unconstrained Optimization of Molecular Orbital Coefficients in Den- sity Functional Theory.J. Chem. Theory Comput.2024, 20, 7936–7947. (55) Frank, I.; Hutter, J.; Marx, D.; Parrinello, M. Molecu- lar dynamics in low-spin excited s...
2024
-
[8947]
V.; Levi, G.; J´ onsson, E.; J´ onsson, H
(41) Ivanov, A. V.; Levi, G.; J´ onsson, E.; J´ onsson, H. Method for Calculating Excited Electronic States Using Density Functionals and Direct Orbital Optimization with Real Space Grid or Plane-Wave Basis Set. J. Chem. Theory Comput. 2021, 17, 5034–5049. (42) MacEtti, G.; Ge...
2021
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.