Pith. sign in

REVIEW 3 major objections 4 minor 63 references

Solvers for Large-Scale Electronic Structure Theory: ELPA and ELSI

T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper reports that the 2024.05 release of ELPA fully ports the matrix-matrix multiplication step of the generalized eigensolver to GPUs, yielding a three-to-fourfold speedup over the best CPU configuration.

desk verdict ELPA 2024.05's GPU port of the matrix multiply is a real, clearly reported speedup, but the performance claim rests on a thin single-node benchmark. read the letter →

arxiv 2502.02460 v1 pith:4YNK3GNY submitted 2025-02-04 cond-mat.mtrl-sci physics.comp-ph

classification cond-mat.mtrl-sciphysics.comp-ph PACS 71.15.-m02.70.-c
keywords ELPAELSIgeneralizedeigenvalueproblemGPUaccelerationmatrix-matrixmultiplicationKohn-Shamdensityfunctionaltheoryparalleleigensolverelectronicstructuresoftware
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports that the generalized eigensolver in the ELPA library, a core linear-algebra engine for Kohn-Sham density functional theory, no longer leaves its most time-consuming step on the CPU. The 2024.05 release fully ports the matrix-matrix multiplication used to reduce a generalized eigenproblem to a standard one, closing a bottleneck that previously dominated the runtime. On a single-node benchmark with a $40960 \times 40960$ matrix, the new GPU path solves the complete generalized eigenproblem in 134 seconds versus 669 seconds for the previous release, and it is roughly three to four times faster than the best CPU-only setup. The paper also describes ELSI, an interface layer that lets electronic structure codes call ELPA and other solvers through one API, so the speedup is accessible without rewriting application code.

What carries the argument

The load-bearing mechanism is the standard reduction of a symmetric or Hermitian generalized eigenproblem to a standard one: Cholesky factorize the overlap matrix $S = U^T U$, invert the triangular factor, form $\tilde H = (U^{-1})^T H U^{-1}$, solve the standard eigenproblem, and backtransform the eigenvectors with $C = U^{-1}\tilde C$. The formation and backtransformation steps are dense matrix-matrix multiplications, which previously used ELPA's own parallel implementations (SUMMA and Cannon's algorithm) and were not fully offloaded to GPUs. The 2024.05 port moves these multiplications onto GPUs with direct intra-node GPU collective communication, which is what cuts the wall-clock time so sharply.

What would settle it

Repeat the Table 1 benchmark on the same node while varying the matrix dimension (for example 8192, 20480, and 81920) and the number of GPUs (1, 2, 8, and 16); if the GPU path is not consistently about three to four times faster than the all-CPU run, or if the one-process-per-GPU setup is not faster than the multi-process-per-GPU setup, the claimed speedup does not generalize as stated.

Watch

Extended reading notes

Core claim

The central claim is that the 2024.05 release of ELPA removes the last major CPU bottleneck in its GPU-accelerated generalized eigensolver by fully porting the matrix-matrix multiplication step to GPUs. In the generalized eigenproblem $HC = \varepsilon S C$, reducing to a standard problem requires the multiplications $\tilde H = (U^{-1})^T H U^{-1}$ and the backtransformation $C = U^{-1}\tilde C$; these steps previously ran largely on CPUs through parallel multiplication algorithms that were not fully GPU-ported. With the full port and direct GPU collective communication, the 'Multiply' step for a $40960 \times 40960$ matrix drops from 294 to 18.8 seconds in the one-MPI-process-per-GPU configuration, and the total solution time drops from 669 to 134 seconds, about five times faster than the prior GPU release and roughly three to four times faster than the best CPU-only configuration. This brings the generalized eigensolver's GPU speedup in line with that of the standard eigensolver.

Load-bearing premise

The load-bearing premise is that a single $40960 \times 40960$ all-eigenpair generalized eigenproblem on one node with four GPUs, measured with single runs and no error bars, is representative enough of production workloads that the observed three-to-fourfold speedup transfers to other matrix sizes, GPU counts, and node counts.

Editorial extensions

If this is right

  • Electronic structure codes that use ELPA through ELSI can obtain the GPU-accelerated generalized eigensolution by linking a newer ELPA release, without changing application code.
  • Because the multiplication step is no longer CPU-bound, the one-MPI-process-per-GPU code path becomes the more promising route, and future ELPA work is expected to move the tridiagonal solve and backtransformation onto GPU collective communication as well.
  • For production calculations where ELPA dominates runtime, the complete generalized eigenproblem solution on a GPU node can be up to four times faster overall, and individual steps up to ten times faster, substantially reducing wall-clock time.
  • The well-conditioned-overlap requirement for the Cholesky-reduction path remains; ill-conditioned overlap matrices still require the filtered overlap-eigenbasis path supported by ELSI.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the single-node result generalizes to multi-node runs, the GPU multiplication port could make direct diagonalization competitive at system sizes previously assigned to linear-scaling density-matrix solvers, since GPU nodes typically offer fewer CPU cores for the CPU-bound fallback.
  • The paper's own future-plan discussion implies an untested claim: the one-process-per-GPU collective-communication path will beat the multi-process-per-GPU path once more steps use GPU collectives; a head-to-head comparison across node counts would settle this.
  • The same fully GPU-ported multiplication logic applies to dense generalized eigenproblems outside electronic structure, so the speedup is likely generic to matrix multiplication rather than specific to Kohn-Sham DFT.
  • Because the reported timings are single runs without error bars, the run-to-run variance is unknown; reporting repeated timings would show how much of the three-to-fourfold speedup is stable gain rather than measurement spread.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript is a software overview of the ELPA eigensolver library and the ELSI interface layer, both tightly integrated with the FHI-aims electronic structure code. The authors describe the Kohn-Sham eigenproblem, the Cholesky-based reduction of the generalized eigenproblem, and the alternative density-matrix solvers available through ELSI. The main novel technical content is a benchmark (Table 1) comparing ELPA releases 2023.11 and 2024.05 on a single node with four NVIDIA A100 GPUs; the 2024.05 release ported the matrix-matrix multiplication step to GPUs, reducing that step from 294 s to 18.8 s in the NCCL configuration and the total generalized-eigenproblem time from 669 s to 134 s, with an even lower total of 97.8 s in the MPS-4 configuration. The paper claims a three-to-fourfold speedup of the GPU generalized eigensolver over the best CPU configuration and recommends ELPA-GPU for production use, while also discussing future GPU collective-communication development and Intel GPU support.

Significance. If the reported speedup is representative, the result is practically significant: ELPA is one of the most widely used parallel eigensolvers in DFT, and the 2024.05 GPU port of the multiply step removes a known bottleneck in the generalized eigenproblem. The manuscript also provides a useful, current overview of the ELPA/ELSI ecosystem, including tutorials, supported solvers, and architecture plans. The strongest point is the concrete, side-by-side timing table for the two ELPA releases, which directly demonstrates a large improvement in the Multiply step; the weakest point is that this demonstration rests on a single configuration and single runs, so the general "3-4x" claim is not yet established with the precision the text uses. The open-source nature of the software and the availability of the ELPA manual and tutorial materials are assets for reproducibility.

major comments (3)
  1. [Table 1 and the 'ELPA - Usability' discussion] The central performance claim—that the 2024.05 GPU port of the matrix-matrix multiplication brings the generalized eigensolver to a three-to-fourfold speedup over CPU—is supported only by a single benchmark: one 40960×40960 matrix, all eigenpairs, one node, four A100 GPUs, with one reported run per configuration and no error bars or reproducibility artifacts. The same text generalizes this into "up to a 4x speedup for the complete solution of standard and generalized eigenproblems" and a recommendation to use ELPA-GPU. Because the two GPU configurations in Table 1 themselves differ by a factor of 1.4 in total time (134 s for NCCL vs 97.8 s for MPS-4), the reported speedup is demonstrably sensitive to configuration details; a single tuning point is insufficient to establish a general speedup claim. I am not disputing that the port improves performance—the Multiply reduction from 294 to 18.8 s is striking—but the paper should either restrict the claim to the measured configuration or add repeat runs, error bars, and additional matrix sizes and GPU counts.
  2. [Table 1] Small internal inconsistencies in Table 1 make the benchmark protocol hard to assess. For the 2024.05 MPS-4 row, the listed components sum to 97.5 s rather than the printed total of 97.8 s; for the CPU row they sum to 392.0 s rather than 393.0 s. These differences are small enough to be rounding effects, but they are not explained, and they leave open whether the components and totals came from separate runs or from a single controlled measurement. Please clarify the protocol and report the rounding convention.
  3. [Table 1 and the 'ELPA - Usability' section] The comparison protocol is under-specified in ways that affect the headline factor. The CPU baseline is ELPA2 with 72 MPI processes, while the GPU runs use ELPA1 with four A100 GPUs and "tuned from 1 to 18 CPU cores per GPU", but the actual number of CPU cores used in the reported GPU rows is not given. It is also unclear whether the CPU baseline is the "best-performing ELPA-CPU configuration" or merely the one tested. Since the speedup ratio is the paper's main quantitative result, the CPU-core counts, BLAS/threading settings, and tuning procedure should be documented for each row.
minor comments (4)
  1. [Title page and Table 1] The rendering of the author list shows "P¨oppl" with an unrendered umlaut, and Table 1 entries such as "5 .7" and "18 .0" contain stray spacing artifacts from LaTeX; these should be fixed in the final version.
  2. [ELSI section] The sentence "ELSI supports ELPA-GPU (NVIDIA only) directly in ELPA's 2020 release, and as an externally compiled library for later versions and other GPU types" is awkward and ambiguous; it should be reworded to clarify which ELPA versions and GPU types are supported through which linkage mode.
  3. [ELPA - Usability section] The claim of "up to a 10x speedup for individual solution steps" appears inconsistent with Table 1, where the backward multiplication drops from 265 s to 6.8 s (about 39x) and the forward multiply from 294 s to 18.8 s (about 16x); the statement and the table should be reconciled.
  4. [ELPA - Usability section] The phrase "full support for NVIDIA and AMD GPUs" is later qualified by "ELSI supports ELPA-GPU (NVIDIA only)"; the distinction between ELPA's native GPU support and ELSI's interface support should be stated more explicitly to avoid apparent contradiction.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the reported GPU speedups are direct benchmark measurements, not derivations from fitted inputs or self-cited premises.

full rationale

This paper is a software overview and performance report, not a derivation. The central claim—that the ELPA 2024.05 GPU port of the matrix-matrix multiplication yields a roughly three- to four-fold speedup for the generalized eigensolver—is supported directly by wall-clock timings in Table 1 (e.g., the NCCL 'Multiply' step dropping from 294 to 18.8 seconds and the total from 669 to 134 seconds). These are empirical measurements of the authors' own software, which introduces a possible self-evaluation bias, but bias is not circular reasoning. No parameter is fitted to the reported outcome and then renamed a prediction; no equation is defined in terms of the result it is supposed to establish; and no load-bearing premise is justified solely by a self-citation. The paper does cite prior work by the same authors for background and for earlier GPU developments, but those citations are not used to force the present conclusion: the 2024.05 performance claims stand on the presented benchmark table rather than on the cited references. One could question whether a single 40960x40960 matrix on one node with four A100 GPUs is representative, and the paper itself notes configuration dependence (MPS versus NCCL), but that is a question of empirical generalizability, not of circularity. The recommendation to use ELPA-GPU follows from the measured timings and is explicitly qualified to the tested setup. Accordingly, no circular step is present, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The benchmark is an empirical measurement; no free parameters are fitted and no new entities are postulated. The central claim relies on the standard linear algebra transformation and the representativeness of the hardware setup.

assumptions (2)
  • domain assumption The overlap matrix S is positive definite and well-conditioned, a requirement for the Cholesky-based transformation used to reduce the generalized eigenproblem to standard form.
    Stated in the introduction where the transformation steps are described.
  • domain assumption Kohn-Sham density functional theory provides a valid description of the electronic structure problem.
    The entire software stack is built on the Kohn-Sham formalism.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Solvers for Large-Scale Electronic Structure Theory: ELPA and ELSI." pith.science (2026). https://pith.science/paper/4YNK3GNY

@misc{pith2026250202460,
  author       = {Pith},
  title        = {Pith review of: Solvers for Large-Scale Electronic Structure Theory: ELPA and ELSI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4YNK3GNY}},
  note         = {Machine review of arXiv:2502.02460}
}
read the original abstract

In this contribution, we give an overview of the ELPA library and ELSI interface, which are crucial elements for large-scale electronic structure calculations in FHI-aims. ELPA is a key solver library that provides efficient solutions for both standard and generalized eigenproblems, which are central to the Kohn-Sham formalism in density functional theory (DFT). It supports CPU and GPU architectures, with full support for NVIDIA and AMD GPUs, and ongoing development for Intel GPUs. Here we also report the results of recent optimizations, leading to significant improvements in GPU performance for the generalized eigenproblem. ELSI is an open-source software interface layer that creates a well-defined connection between "user" electronic structure codes and "solver" libraries for the Kohn-Sham problem, abstracting the step between Hamilton and overlap matrices (as input to ELSI and the respective solvers) and eigenvalues and eigenvectors or density matrix solutions (as output to be passed back to the "user" electronic structure code). In addition to ELPA, ELSI supports solvers including LAPACK and MAGMA, the PEXSI and NTPoly libraries (which bypass an explicit eigenvalue solution), and several others.

Figures

Figures reproduced from arXiv: 2502.02460 by the authors.

Figure 1
Figure 1. Tasks performed by the ELSI interface software, connecting different eigenvalue and density matrix solvers to electronic structure codes including FHI-aims, Siesta, DFTB+, NWChemEx and others. ELSI provides a uniform interface that is callable in Fortran, C, C++, and Python, and handles matrix format conversion between user codes and solver libraries. The solvers presented here include eigenvalue and density matrix … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 63 canonical work pages

  1. [1]

    Kohn and L.J

    W. Kohn and L.J. Sham, Self-consistent equations including exchange and correlation effects. Phys. Rev., 140, A1133–A1138 (1965)

  2. [2]

    Seidl, A

    A. Seidl, A. G ¨orling, P . Vogl, J.A. Majewski and M. Levy,Generalized Kohn-Sham schemes and the band-gap problem, Phys. Rev. B53, 3764–3774 (1996)

  3. [3]

    Burns, T

    Volker Blum, Ryoji Asahi, Jochen Autschbach, Christoph Bannwarth, Gustav Bihlmayer, Stefan Bl¨ugel, Lori A. Burns, T. Daniel Crawford, William Dawson, Wibe Albert de Jong et al.,Roadmap on software for electronic structure-based simulations in chemistry and materials, Electronic Structure 6, 042501 (2024)

  4. [4]

    https://wordpress.elsi-interchange.org/

  5. [5]

    V. W.-z. Yu, F. Corsetti, A. Garc´ıa, W. P . Huhn, M. Jacquelin, W. Jia, B. Lange, L. Lin, J. Lu, W. Mi et al., ELSI: A Unified Software Interface for Kohn-Sham Electronic Structure Solvers. Computer Physics Communications 222, 267-285 (2018)

  6. [6]

    V. W.-z. Yu, C. Campos, W. Dawson, A. Garc ´ıa, V. Havu, B. Hourahine, W. P . Huhn, M. Jacquelin, W. Jia, M. Kec ¸eli, R. Laasner, Y . Li, L. Lin, J. Lu, J. Moussa, J. E. Roman, ´A. V ´azquez-Mayagoitia, C. Yang, and V. Blum,ELSI – An Open Infrastructure for Electronic Structure Solvers. Computer Physics Communications 256, 107459 (2020)

  7. [7]

    V. Blum, R. Gehrke, F. Hanke, P . Havu, V. Havu, X. Ren, K. Reuter, M. Scheffler, Ab initio molec- ular simulations with numeric atom-centered orbitals . Computer Physics Communications 180, 2175–2196 (2009)

  8. [8]

    Hourahine, B

    B. Hourahine, B. Aradi, V. Blum, F. Bonaf ´e, A. Buccheri, C. Camacho, C. Cevallos, M.Y . Deshaye, T. Dumitrica, A. Dominguez et al., DFTB+, a software package for efficient approximate density functional theory based atomistic simulations. The Journal of Chemical Physics152, 124101 (2020)

Show all 63 references
  1. [9]

    Garc ´ıa, N

    A. Garc ´ıa, N. Papior, A. Akhtar, E. Artacho, V. Blum, E. Bosoni, P Brandimarte, M. Brandbyge, J. I. Cerd´a, F. Corsetti, et al. Siesta: Recent developments and applications . The Journal of Chemical Physics 152, 204108 (2020)

  2. [10]

    NWChemEx Community, https://github.com/NWChemEx (2024)

  3. [11]

    https://www.netlib.org/lapack/

  4. [12]

    Anderson, Z

    E. Anderson, Z. Bai, C. Bischof, S. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney and D. Sorensen, LAPACK Users’ Guide, SIAM (1999)

  5. [13]

    https://elpa.mpcdf.mpg.de/

  6. [14]

    Marek, V

    A. Marek, V. Blum, R. Johanni, V. Havu, B. Lang, T. Auckenthaler, A. Heinecke, H.-J. Bungartz, and H. Lederer, The ELPA library: scalable parallel eigenvalue solutions for electronic structure theory and computational science. J. Phys.: Condens. Matter 26, 213201 (2014)

  7. [15]

    Marek, P

    A. Marek, P . Karpov, T. Melson. ELPA Manual: User’s Guide and Best Practices (2024). Available online at https://elpa.mpcdf.mpg.de/elpa_userguide.pdf 8

  8. [16]

    L. Lin, J. Lu, L. Ying, R. Car, W. E, Fast algorithm for extracting the diagonal of the inverse matrix with application to the electronic structure analysis of metallic systems , Comm. Math. Sci. 7, 755 (2009)

  9. [17]

    L. Lin, M. Chen, C. Yang, L. He, Accelerating atomic orbital-based electronic structure calculation via pole expansion and selected inversion, J. Phys. Condens. Matter 25 295501 (2013)

  10. [18]

    Dawson, T

    W. Dawson, T. Nakajima, Massively Parallel Sparse Matrix Function Calculations with NTPoly, Com- puter Physics Communications 225 154-165 (2018)

  11. [19]

    Imamura, S

    T. Imamura, S. Yamada, and M. Machida, Development of a High-Performance Eigensolver on a Peta-Scale Next-Generation Supercomputer System, Progress in Nuclear Science and Technology2 643-650, (2011)

  12. [20]

    Dongarra, M

    J. Dongarra, M. Gates, A. Haidar, J. Kurzak, P . Luszczek, S. Tomov, and I. Yamazaki, Accelerating numerical dense linear algebra calculations with GPUs , in: Kindratenko, V. (ed), Numerical Com- putations with GPUs, 1–26 (Springer, Cham, 2014)

  13. [21]

    Winkelmann, P

    J. Winkelmann, P . Springer, and E. Di Napoli, ChASE: a Chebyshev Accelerated Subspace iteration Eigensolver for sequences of Hermitian eigenvalue problems , ACM Transaction on Mathematical Software 45, 21 (2019)

  14. [22]

    X. Wu, D. Davidovi ´c, S. Achilles,E. Di Napoli, ChASE: a distributed hybrid CPU-GPU eigensolver for large-scale hermitian eigenvalue problems , Proceedings of the Platform for Advanced Scientific Computing Conference (PASC22) (2022). DOI:10.1145/3539781.3539792

  15. [23]

    Keceli, F

    M. Keceli, F. Corsetti, C. Campos, J. E. Roman, H. Zhang, ´A. V ´azquez-Mayagoitia, P . Zapol, and A. F. Wagner,SIESTA-SIPs: Massively Parallel Spectrum-Slicing Eigensolver for an Ab Initio Molecular Dynamics Package, Journal of Computational Chemistry 39 (2018), 1806-1814

  16. [24]

    Corsetti, The Orbital Minimization Method for Electronic Structure Calculations with Finite-Range Atomic Basis Sets, Computer Physics Communications 185 (2014), 873-883

    F. Corsetti, The Orbital Minimization Method for Electronic Structure Calculations with Finite-Range Atomic Basis Sets, Computer Physics Communications 185 (2014), 873-883

  17. [25]

    M. Shao, F. H. de Jornada, C. Yang, J. Deslippe, and S. G. Louie, Structure Preserving Parallel Algo- rithms for Solving the Bethe–Salpeter Eigenvalue Problem, Linear Algebra and its Applications488 (2016), 1481-67

  18. [26]

    Gonze, B

    X. Gonze, B. Amadon, G. Antonius, F. Arnardi, L. Baguet, J.-M. Beuken, J. Bieder, F. Bottin, J. Bouchet, E. Bousquet, et al., The ABINIT project: Impact, environment and recent developments . Computer Physics Communications 248, 107042 (2020)

  19. [27]

    Deslippe, G

    J. Deslippe, G. Samsonidze, D.A. Strubbe, M. Jain, M.L. Cohen, S.G. Louie, BerkeleyGW: A mas- sively parallel computer package for the calculation of the quasiparticle and optical properties of materials and nanostructures. Computer Physics Communications 183, 1269–1289 (2012)

  20. [28]

    K ¨uhne, M

    T.D. K ¨uhne, M. Iannuzzi, M. Del Ben, V. V. Rybkin, P . Seewald, F. Stein, T. Laino, R. Z. Khaliullin, O. Sch¨utt, F. Schiffmann et al., CP2K: An electronic structure and molecular dynamics software pack- age – Quickstep: Efficient and accurate electronic structure calculations ...

  21. [29]

    Kl ¨offel, G

    T. Kl ¨offel, G. Mathias, B. Meyer, Integrating state of the art compute, communication, and auto- tuning strategies to multiply the performance of ab initio molecular dynamics on massively parallel multi-core supercomputers. Computer Physics Communications 260, 107745 (2021)

  22. [30]

    Enkovaara, C

    J. Enkovaara, C. Rostgaard, J. J. Mortensen, J. Chen, M. Dulak, L. Ferrighi, J. Gavnholt, C. Glinsvad, V. Haikola, H. A. Hansen et al.,Electronic structure calculations with GPAW: A real-space implementa- tion of the projector augmented-wave method . Journal of Physics: Conden...

  23. [31]

    Apr `a, E

    E. Apr `a, E. J. Bylaska, W. A. de Jong, N. Govind, K. Kowalski, T. P . Straatsma, M. Valiev, H. J. J. van Dam, Y . Alexeev, J. Anchell et al., NWChem: Past, present, and future . The Journal of Chemical Physics 152, 184102 (2020)

  24. [32]

    Tancogne-Dejean, M.J.T

    N. Tancogne-Dejean, M.J.T. Oliveira, X. Andrade, H. Appel, C.H. Borca, G. Le Breton, F. Buchholz, A. Castro, S. Corni, A. A. Correa et al.,Octopus, a computational framework for exploring light-driven phenomena and quantum dynamics in extended and finite systems. The Journal of...

  25. [33]

    Smidstrup, T

    S. Smidstrup, T. Markussen, P . Vancraeyveld, J. Wellendorff, J. Schneider, T. Gunst, B. Verstichel, D. Stradi, P . A. Khomyakov, U. G. VejHansen et al.,QuantumATK: An integrated platform of electronic and atomic-scale modelling tools. Journal of Physics: Condensed Matter 32, 0...

  26. [34]

    Giannozzi, S

    P . Giannozzi, S. Baroni, N. Bonini, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, G. L. Chiarotti, M. Cococcioni, I. Dabo, et al., QUANTUM ESPRESSO: a modular and open-source software project for quantum simulations of materials. J.Phys.: Condens. Matter 21, 395502 (2009)

  27. [35]

    Giannozzi, O

    P . Giannozzi, O. Andreussi, T. Brumme, O. Bunau, M. Buongiorno Nardelli, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, M. Cococcioni, et al. Advanced capabilities for materials modelling with Quantum ESPRESSO. J.Phys.: Condens. Matter 29, 465901 (2017)

  28. [36]

    Kresse, J

    G. Kresse, J. Furthm ¨uller, Efficient iterative schemes for ab initio total- energy calculations using a plane-wave basis set. Physical Review B 54 11169–11186 (1996)

  29. [37]

    Blaha, K

    P . Blaha, K. Schwarz, F. Tran, R. Laskowski, G.K.H. Madsen, L.D. Marks,WIEN2k: An APW+lo program for calculating the properties of solids , The Journal of Chemical Physics 152, 074101 (2020)

  30. [38]

    P . Kus, H. Lederer, A. Marek. GPU optimization of large-scale eigenvalue solver. In: F. Radu, K. Kumar, I. Berre, J. Nordbotten, I. Pop, editors. Numerical mathematics and advanced applications ENUMATH 2017. Lecture Notes in Computational Science and Engineering. Vol. 126. Ch...

  31. [39]

    V.W.-z. Yu , J. Moussa , P . Kus, A. Marek, P . Messmer, M. Yoon, H. Lederer, V. BlumGPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and Hermitian eigenproblems. Comput Phys Commun. 262:107808 (2021)

  32. [40]

    Wlazlowski, M

    G. Wlazlowski, M. Forbes, S.R. Sarkar, A. Marek, M. Szpindle,Fermionic quantum turbulence: Push- ing the limits of high-performance computing . PNAS Nexus 3, pgae160 (2024)

  33. [41]

    https://www.netlib.org/scalapack/

  34. [42]

    Blackford, J

    L.S. Blackford, J. Choi, A. Cleary, E. D’ Azevedo, J. Demmel, I. Dhillon, J. Dongarra, S. Hammarling, G. Henry, A. Petitet, et al., ScaLAPACK users’ guide, SIAM (1997)

  35. [43]

    https://developer.nvidia.com/nccl

  36. [44]

    https://rocm.docs.amd.com/projects/rccl

  37. [45]

    https://www.intel.com/content/www/us/en/developer/tools/oneapi/ data-parallel-c-plus-plus.html

  38. [46]

    https://www.intel.com/content/www/us/en/developer/tools/oneapi/onemkl.html

  39. [47]

    Lang, A parallel algorithm for reducing symmetric banded matrices to tridiagonal form

    B. Lang, A parallel algorithm for reducing symmetric banded matrices to tridiagonal form . SIAM Journal on Scientific Computing 14, 1320–1338 (1993)

  40. [48]

    Bischof, X

    C. Bischof, X. Sun, B. Lang, Parallel tridiagonalization through two-step band reduction . Proceed- ings of IEEE Scalable High Performance Computing Conference, pp. 23–27 (1994) 10

  41. [49]

    Auckenthaler, V

    T. Auckenthaler, V. Blum, H.-J. Bungartz, T. Huckle, R. Johanni, L. Kr ¨amer, B. Lang, H. Lederer and P . R. Willems,Parallel solution of partial symmetric eigenvalue problems from electronic structure calculations. Parallel Computing 37, 783-794 (2011)

  42. [50]

    P . Kus, A. Marek, S.S. K ¨ocher, H.-H. Kowalski, C. Carbogno, Ch. Scheurer, K. Reuter, M. Scheffler, H. Lederer, Optimizations of the eigensolvers in the ELPA library. Parallel Comput. 85, 167–177 (2019)

  43. [51]

    van de Geijn, J

    R.A. van de Geijn, J. Watts, SUMMA: scalable universal matrix multiplication algorithm. Concur- rency: Practice and Experience 9, 255-274 (1997)

  44. [52]

    Chtchelkanova, J

    A. Chtchelkanova, J. Gunnels, G. Morrow, J. Overfelt, R.A. van de Geijn Parallel implementation of BLAS: general techniques for Level 3 BLAS. Concurrency: Practice and Experience 9, 837–857 (1997)

  45. [53]

    Cannon, A cellular computer to implement the Kalman filter algorithm

    L.E. Cannon, A cellular computer to implement the Kalman filter algorithm . PhD thesis, Carnegie- Mellon University (1969)

  46. [54]

    https://docs.nvidia.com/deploy/mps

  47. [55]

    https://docs.mpcdf.mpg.de/doc/computing/raven-user-guide.html

  48. [56]

    Karpov, ScaLAPACK + ELPA tutorial (2023), GitHub repository https://github.com/ karpov-peter/elpa-tutorial

    P . Karpov, ScaLAPACK + ELPA tutorial (2023), GitHub repository https://github.com/ karpov-peter/elpa-tutorial

  49. [57]

    https://www.amd.com/de/products/accelerators/instinct/mi300/mi300a.html

  50. [58]

    https://github.com/riscvarchive/riscv-v-spec

  51. [59]

    Europe 4, 165 (2024)

    Rogeli Grima Torres, Pablo Vizca´ıno, Filippo Mantovani, Jos´e Julio Guti´errez Moreno, Co-designing ab initio electronic structure methods on a RISC-V vector architecture , Open Res. Europe 4, 165 (2024)

  52. [60]

    Distributed Linear Algebra from the Future, https://github.com/eth-cscs/DLA-Future

  53. [61]

    Solc `a, M

    R. Solc `a, M. Simberg, R. Meli, A. Invernizzi, A. Reverdell, J. Biddiscombe, DLA-Future: A Task- Based Linear Algebra Library Which Provides a GPU-Enabled Distributed Eigensolver , in: P . Diehl, J. Schuchart, P . Valero-Lara, G. Bosilca, George (eds)Asynchronous Many-Task Sy...

  54. [62]

    Pederson, J

    R. Pederson, J. Kozlowski, R. Song, J. Beall, M. Ganahl, M. Hauru, A. G.M. Lewis, Y . Yao, S. Basu Mallick, V. Blum, G. Vidal,Large Scale Quantum Chemistry with Tensor Processing Units, Journal of Chemical Theory and Computation 19 25-32 (2023)

  55. [63]

    V. W.-z. Yu, J. Moussa, V. Blum, Accurate Frozen Core Approximation for All-Electron Density- Functional Theory, The Journal of Chemical Physics 154, 224107 (2021). 11

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.