Pith. sign in

REVIEW 2 major objections 5 minor 59 references

Efficient implementation of the quasiparticle self-consistent $GW$ method on GPU

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read One GPU node achieves over a tenfold speedup for quasiparticle self-consistent $GW$ self-energy calculations, and a mixed-precision mode adds a further threefold, enabling large interface and alloy systems.

desk verdict A credible and useful QSGW-on-GPU implementation paper whose cleanest claim is the mixed-precision speedup; the headline GPU/CPU ratio is inflated by a weak CPU baseline, but the core software contribution deserves peer review. read the letter →

arxiv 2506.03477 v1 pith:THMJG72R submitted 2025-06-04 physics.comp-ph cond-mat.mtrl-sci

classification physics.comp-phcond-mat.mtrl-sci
keywords QSGWmulti-GPUOpenACCmixedprecisionecaljGWapproximationInAs/GaSbsuperlatticeNi2MnGa
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the quasiparticle self-consistent $GW$ (QSGW) method, a high-accuracy first-principles approach for electronic excitations, can be made dramatically faster by porting its four dominant kernels to GPUs without sacrificing maintainability. The authors implemented a multi-GPU version inside the ecalj package using OpenACC directives, GPU math libraries, and a module-based Fortran design, with one MPI process per GPU. On benchmarks, one GPU node ran the self-energy computation over 10 times faster than four CPU nodes, and a mixed-precision mode (TF32-based matrix multiplications with double precision where needed) added more than a factor of three, keeping the mean self-energy difference around 6–9 meV. The paper further demonstrates that the accelerated code produces physically meaningful results for (InAs)$_n$(GaSb)$_n$ superlattices and Ni$_2$MnGa martensite phases, systems that were previously impractical at these sizes with QSGW.

What carries the argument

The load-bearing object is the set of matrix elements $Z^{q-k n', q n}_{k\mu}=\langle E^k_\mu \Psi^{q-k n'}\,|\,\Psi^{q n}\rangle$ built from the mixed product basis $\{E^k_\mu\}$ and the PMT wave functions, where PMT is the paper's mixed basis of muffin-tin orbitals and augmented plane waves. All four QSGW kernels reduce to sums of products of these $Z$ elements with $v_\mu$, $W^c_{\mu\nu}$, and frequency weights, i.e. to GEMM operations; the implementation restructures arrays and loop order so that these GEMMs run on the GPU, with segmented batching over intermediate states to keep the large $Z$ and $W$ arrays within memory, and the mixed-precision mode applies TF32 to the matrix products while retaining double precision for inversion and overlap-type steps.

What would settle it

Run one full QSGW self-consistency cycle for (InAs)$_{10}$(GaSb)$_{10}$ on the same CPU hardware but with the screened Coulomb interaction $W$ kept resident in memory and with the MPI communication over the product-basis index avoided, then compare its wall time with the GPU run; if the GPU/CPU ratio drops close to the ratio of theoretical FLOPs (about 3.8), the claimed 'over 10 times' is mostly an artifact of the baseline's I/O and MPI overhead rather than of GPU compute.

Watch

Extended reading notes

Core claim

The central claim is that QSGW can be decomposed into complex matrix multiplications over the mixed product basis and intermediate-state indices, and that this decomposition lets the four expensive kernels — core-exchange self-energy, exchange self-energy, screened Coulomb interaction $W$, and static correlation self-energy $\tilde\Sigma_c$ — run efficiently on GPUs via cuBLAS/cuSOLVER, with OpenACC handling the loops and one MPI process per GPU. The paper reports that for the largest benchmark systems, one GPU node achieves a 10-fold speedup over a four-CPU-node reference, and the mixed-precision version reaches up to 38-fold, with the difference between mixed-precision and double-precision self-energies having a mean of about 6–9 meV. It also claims that the resulting electronic structures — a band-gap maximum near $n=3$ in (InAs)$_n$(GaSb)$_n$ and a Fermi level sitting in a density-of-states valley for the modulated Ni$_2$MnGa phases — demonstrate that the accelerated QSGW code yields reliable physics at previously impractical system sizes.

Load-bearing premise

The headline speedups assume the four-CPU-node run is a fair, well-optimized baseline, but that baseline uses 128 MPI processes with file I/O for $W$ and MPI communication that the paper itself says already suffers from bottlenecks below linear scaling, so an improved CPU baseline would shrink the reported GPU/CPU ratios.

Editorial extensions

If this is right

  • QSGW self-energy calculations for systems with roughly 40 atoms per cell, such as (InAs)$_{10}$(GaSb)$_{10}$ and Ni$_2$MnGa 14M, complete on a single GPU node in times previously associated with much smaller systems, broadening QSGW to interfaces, surfaces, and superlattices.
  • The mixed-precision mode halves the memory footprint of the large $Z$ and $W$ arrays, roughly doubling the accessible system size at a given GPU memory and adding about a factor of three in speed without moving the self-energy by more than ~9 meV in the tested cases.
  • Because the speedup nearly saturates for the largest systems, GPU compute rather than data transfer dominates, so the same design should extend to multiple GPU nodes and even larger unit cells.
  • The reproduced band-gap maximum near $n=3$ in (InAs)$_n$(GaSb)$_n$ and the suggested DOS-valley electronic stabilization for modulated Ni$_2$MnGa are physical consequences that became reachable because the GPU implementation made the large-cell QSGW cycles affordable.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported acceleration partly conflates GPU raw compute with the removal of MPI and file-I/O overhead: the CPU baseline's below-linear scaling and file reading of $W$ (acknowledged in Sec. 4.2) inflate the GPU/CPU ratio, so a stricter comparison with a memory-resident CPU run would likely yield a smaller but still substantial speedup.
  • The mixed-precision accuracy was verified only on the first QSGW cycle for two systems; the 6–9 meV mean difference is a property of those cases, and testing mixed precision to full self-consistency across a broader set of semiconductors and metals would show whether the error stays below typical target accuracies.
  • The module-based OpenACC design with unified CPU/GPU interfaces could be transferred to other all-electron GW codes, provided those codes share the same GEMM-friendly decomposition of kernels over product-basis indices.
  • The DOS-valley stabilization mechanism suggested for Ni$_2$MnGa is testable: calorimetric or elastic measurements that perturb the modulated phases should detect a narrower electronic susceptibility when the Fermi level sits in the valley, as the QSGW DOS predicts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports a multi-GPU implementation of the quasiparticle self-consistent GW (QSGW) method inside the ecalj package, using modern Fortran modules, OpenACC directives, and GPU-accelerated math libraries. After summarizing the QSGW formalism and the four expensive kernels (core-exchange, exchange, screened Coulomb interaction, and correlation self-energy), the authors describe their MPI+GPU parallelization, batch processing of the large Z arrays, and a mixed-precision (TF32) variant. Benchmarks on (InAs)n(GaSb)n superlattices and Ni2MnGa martensite phases report substantial speedups: one GPU node is claimed to be about 10x faster than four CPU nodes, and the mixed-precision GPU version up to 38x faster, with mean diagonal self-energy differences of 5-9 meV relative to double precision. The paper also presents QSGW80 electronic structure results for the superlattice band gaps and for the Ni2MnGa DOS, ending with a suggested stabilization mechanism for the modulated phases.

Significance. If the headline speedups were measured against a well-optimized CPU baseline, this work would make QSGW calculations dramatically more accessible for large supercells, interfaces, and magnetic systems. The paper has clear strengths: the source code is publicly available on GitHub, the parallelization strategy is described in detailed pseudo-code (Algorithms 1 and 2), the choice of one MPI process per GPU is justified, and the mixed-precision comparison GPUMP vs. GPU is a well-controlled measurement of the TF32/FP64 arithmetic trade-off, giving more than a 3x speedup at mean self-energy differences of 5-9 meV. The physical QSGW80 results reproduce the experimental band-gap trend of the InAs/GaSb superlattice and give a plausible qualitative description of the Ni2MnGa DOS. However, the central GPU/CPU performance claim is weakened by the CPU baseline's known scalability problems, which the authors themselves acknowledge, and the mixed-precision accuracy test is limited to first-cycle, diagonal, Gamma-point self-energies. The GPUMP-vs-GPU result is the most robust part of the performance claim and deserves to be foregrounded.

major comments (2)
  1. [Sec. 4.2, Fig. 4] The headline GPU/CPU speedup is contaminated by the choice of CPU baseline. The CPU run uses 128 MPI processes with four-thread MKL parallelization, whose measured thread efficiency is only 1.6x over flat MPI (Sec. 4.1), and the authors state that the CPU results were already below linear scaling due to file I/O and MPI communication, while the GPU run uses only four MPI processes with more memory per process. Consequently, ratios such as "over 10 times" and "equivalent to 40 CPU nodes" include reductions in MPI and I/O overhead, not just GPU compute capability. Because this is the central performance claim, the manuscript should either benchmark against a more carefully optimized CPU baseline (for example, fewer MPI ranks with better thread scaling, or a second independent CPU code) or reframe the primary claim around the GPUMP-vs-GPU comparison, which isolates the TF32 vs. FP64 arithmetic trade-off.
  2. [Sec. 4.3, Sec. 3.4] The mixed-precision accuracy test is too narrow to support "preserving numerical accuracy" as a general claim. The comparison in Fig. 7 is limited to two materials, the first QSGW cycle, the Gamma point, and diagonal self-energy matrix elements; off-diagonal elements, self-consistency convergence, and final quasiparticle energies or band gaps are not compared between GPUMP and GPU. Since QSGW uses off-diagonal self-energy elements to update H0 through Eq. (3), the 5-9 meV diagonal differences do not guarantee that converged QSGW80 gaps agree to similar accuracy. The authors should add at least one converged QSGW80 comparison of band gaps or DOS between GPUMP and GPU, or explicitly restrict the accuracy claim to first-cycle diagonal self-energies.
minor comments (5)
  1. [Fig. 4 caption] The sentence "×12 means that the computational speed of the GPU (1 node) is worth (128×4)×12 cores (48 nodes)" is unclear: 128×4 is 512 threads, and 48 nodes is inconsistent with the four-node CPU baseline; please clarify the intended core-count or node-count equivalence.
  2. [Fig. 6 caption] The caption states that both vertical and horizontal axes are logarithmic, but the horizontal axes in panels (a-d) appear to be linear in n; please correct the caption or the axes.
  3. [Eq. (2)] The denominator in Eq. (2) should be written with explicit parentheses as ω − (H0 + Σ − Vxc) to avoid the appearance of a scalar denominator.
  4. [Sec. 5.0.2] The statement that the valley in the DOS stabilizes the modulated structures is presented as a direct implication, but the paper does not compute total energies or vibrational properties; please label this as a hypothesis to be tested in future work.
  5. [Abstract and Highlights] The abstract and highlights say "A GPU node achieved 10-fold speedup over four CPU nodes," while Fig. 4 reports 12x for the largest (InAs)n(GaSb)n and 14M cases; please harmonize the wording.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified: the implementation and benchmark claims are self-contained and the physical predictions use fixed external inputs rather than fitted targets.

full rationale

The paper's main claim is an implementation/performance result: QSGW kernels are reformulated as tensor contractions and offloaded to GPUs, with benchmark timings compared against the authors' own CPU code. The speedup ratios are empirical elapsed-time measurements, not quantities fitted to a target; the CPU baseline's scalability limitations (128 MPI processes, file I/O, MPI communication) are explicitly disclosed in Sec. 4.2 and the Fig. 4 caption, making this a benchmark-fairness concern rather than a circular derivation. The mixed-precision gain is validated by direct comparison of MP versus DP self-energies (Sec. 4.3, Fig. 7), reporting mean differences of 5-9 meV; this is a controlled numerical comparison, not a fitted agreement. The physical results use QSGW80, an empirically fixed 80/20 mixing ratio inherited from earlier work (Ref. [2]), but QSGW80 is an input to the calculations rather than a quantity this paper derives or claims to predict, so its use is not circular. The Ni2MnGa valley-at-Fermi-level interpretation is a speculative physics conclusion, not a derivation from the implementation claim. Self-citations to Refs. [2, 8, 20] support the underlying QSGW formalism and ecalj method with stated assumptions that do not include the benchmark or mixed-precision results, so they are not load-bearing circularity. No equation-level reduction of a prediction to a fitted parameter or self-defined quantity was found.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central performance claim introduces no free parameters or new entities. The only fitted or ad hoc numerical input is the QSGW80 mixing ratio inherited from prior work. The implementation relies on standard domain assumptions: time-reversal symmetry, frozen cores, RPA screening, and the completeness of the mixed product basis.

free parameters (1)
  • QSGW80 mixing ratio = 0.8
    Empirical hybrid of 80% QSGW and 20% LDA exchange-correlation, adopted from Ref. [2] to compensate for missing vertex corrections; used for all band-structure results in Sec. 5. It is an input to this paper, not derived here.
assumptions (4)
  • domain assumption Time-reversal symmetry of the Hamiltonian is assumed when building the polarization function.
    Introduced in Sec. 2.3 after Eq. (13) to simplify Im[Pi] to the positive-energy term in Eq. (14); the paper notes a careful examination is needed for non-time-reversal systems.
  • domain assumption Frozen-core approximation with core contributions to W neglected.
    Stated in Sec. 2.1; cores affect valence electrons only through the core-exchange self-energy and frozen density.
  • domain assumption The GW approximation itself, with RPA screening and no vertex corrections, gives accurate quasiparticle energies.
    The entire QSGW framework relies on this; the paper uses QSGW80 as a partial empirical correction for the known overestimate of exchange.
  • domain assumption The mixed product basis can accurately expand products of eigenfunctions.
    Underlies Eqs. (8)-(24); the paper expects the basis to express products well and uses a finite APW cutoff around 3 Ry (Sec. 2.2, Table 1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficient implementation of the quasiparticle self-consistent $GW$ method on GPU." pith.science (2026). https://pith.science/paper/THMJG72R

@misc{pith2026250603477,
  author       = {Pith},
  title        = {Pith review of: Efficient implementation of the quasiparticle self-consistent $GW$ method on GPU},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/THMJG72R}},
  note         = {Machine review of arXiv:2506.03477}
}
abstract

We have developed a multi-GPU version of the quasiparticle self-consistent $GW$ (QSGW), a cutting-edge method for describing electronic excitations in a first-principles approach. While the QSGW calculation algorithm is inherently well-suited for GPU computation due to its reliance on large-scale tensor operations, achieving a maintainable and extensible implementation is not straightforward. Addressing this, we have developed a GPU version within the \texttt{ecalj} package, utilizing module-based programming style in modern Fortran. This design facilitates future development and code sustainability. Following the summary of the QSGW formalism, we present our GPU implementation approach and the results of benchmark calculations for two types of systems to demonstrate the capability of our GPU-supported QSGW calculations.

Figures

Figures reproduced from arXiv: 2506.03477 by the authors.

Figure 1
Figure 1. Self-consistent cycle of QSGW. The name of the program that performs each task is noted at the top right corner of [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Parallelization scheme for QSGW calculations. [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Structures of (InAs)8(GaSb)8(a) and Ni2MnGa 14M (b), where In, As, Ga, Sb, Mn, and Ni are represented by pink, green, red, yellow, blue, and gray spheres, respectively. The solid black lines represent the unit cell, and the dashed orange lines in (b) represent the boundaries of the nano twins. However, since TF32 operations reduce computational precision, they can affect the accuracy of the results. Thus, a verifica… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Computational time of Σ. The meaning of the key is as follows. CPU (4 nodes): This is the reference to measure the speed of GPU computations. The total number of cores is 128×4. GPU (1 node): Use one node with four GPUs. GPUMP (1 node): Use one node with four GPUs with…
Figure 5
Figure 5. Figure 5: Computational time of the predominant four tasks. For the acceleration ratio, see the left (right) vertical axis for [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Computational time of kernel calculation in the predominant four tasks and those of acceleration ratio between GPU [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Self-energy and its components differences between MP and DP at [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Band gap (a) with respect to n. The dashed lines are a guide for the eye. Band dispersion of n = 3 (b) and n = 10 (c). Ref. [Otsuka] is [9], Exp. [PL] and Exp. [Abs] are experimental values from photoluminescence and optical absorption. [49, 50] These experimental data…
Figure 9
Figure 9. Figure 9: Total density of states (DOS) (a) and its enlarged plots near the Fermi level (b) of the NM, 10M, and 14M structures [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

59 extracted references · 40 canonical work pages

  1. [1]

    Aryasetiawan, O

    F. Aryasetiawan, O. Gunnarsson, The GW method, Rep. Prog. Phys. 61 (3) (1998) 237.doi:10.1088/ 0034-4885/61/3/002

  2. [2]

    Deguchi, K

    D. Deguchi, K. Sato, H. Kino, T. Kotani, Accurate energy bands calculated by the hybrid quasiparticle self-consistentGWmethod implemented in the ecalj package, Jpn. J. Appl. Phys. 55 (5) (2016) 051201. doi:10.7567/JJAP.55.051201

  3. [3]

    Hedin, S

    L. Hedin, S. Lundqvist, Effects of Electron-Electron and Electron-Phonon Interactions on the One- Electron States of Solids, in: Solid State Phys., Vol. 23, Elsevier, 1970, pp. 1–181.doi:10.1016/ S0081-1947(08)60615-3

  4. [4]

    M. S. Hybertsen, S. G. Louie, Electron correlation in semiconductors and insulators: Band gaps and quasiparticle energies, Phys. Rev. B 34 (1986) 5390.doi:10.1103/PhysRevB.34.5390

  5. [5]

    Usuda, N

    M. Usuda, N. Hamada, T. Kotani, M. van Schilfgaarde, All-electron GW calculation based on the LAPW method: Application to wurtzite ZnO, Phys. Rev. B 66 (12) (2002) 125101.doi:10.1103/ PhysRevB.66.125101

  6. [6]

    Usuda, H

    M. Usuda, H. Hamada, K. Siraishi, A. Oshiyama, Band structures of wurtzite InN and Ga1-xInxN by all-electron GW calculation, Jpn. J. Appl. Phys. Lett. (part 2) 43 (2004) L407

  7. [7]

    van Schilfgaarde, T

    M. van Schilfgaarde, T. Kotani, S. Faleev, Quasiparticle self-consistentGWtheory, Phys. Rev. Lett. 96 (2006) 226402.doi:10.1103/PhysRevLett.96.226402. 22

  8. [8]

    Kotani, M

    T. Kotani, M. van Schilfgaarde, S. V. Faleev, Quasiparticle self-consistentGWmethod: A basis for the independent-particle approximation, Phys. Rev. B 76 (2007) 165106.doi:10.1103/PhysRevB.76. 165106

Show all 59 references
  1. [9]

    Otsuka, T

    J. Otsuka, T. Kato, H. Sakakibara, T. Kotani, Band structures for short-period (InAs)n(GaSb)n super- lattices calculated by the quasiparticle self-consistent GW method, Jpn. J. Appl. Phys. 56 (2) (2017) 021201.doi:10.7567/JJAP.56.021201

  2. [10]

    S. V. Faleev, M. van Schilfgaarde, T. Kotani, All-electron self-consistent GW approximation: Applica- tion to Si, MnO, and NiO, Phys. Rev. Lett. 93 (12) (2004) 126406.doi:10.1103/PhysRevLett.93. 126406

  3. [11]

    Bruneval, N

    F. Bruneval, N. Vast, L. Reining, M. Izquierdo, F. Sirotti, N. Barrett, Exchange and correlation effects in electronic excitations ofCu2O, Phys. Rev. Lett. 97 (2006) 267601.doi:10.1103/PhysRevLett.97. 267601

  4. [12]

    S. W. Jang, T. Kotani, H. Kino, K. Kuroki, M. J. Han, Quasiparticle self-consistent GW study of cuprates: electronic structure, model parameters, and the two-band theory for Tc, Sci. Rep. 5 (2015) 12050.doi:10.1038/srep12050

  5. [13]

    https://github.com/tkotani/ecalj

  6. [14]

    H. Yoon, S. W. Jang, J. H. Sim, T. Kotani, M. J. Han, Magnetic force theory combined with quasi- particle self-consistent GW method, J. Phys. Condens. Matter 31 (2019) 405503.doi:10.1088/ 1361-648x/ab2b7e

  7. [15]

    Obata, T

    M. Obata, T. Kotani, T. Oda, Intrinsic instability to martensite phases in ferromagnetic shape memory alloy Ni2MnGa: Quasiparticle self-consistentGWinvestigation, Phys. Rev. Mater. 7 (2023) 024413. doi:10.1103/physrevmaterials.7.024413

  8. [16]

    Luštinec, M

    J. Luštinec, M. Obata, R. Majumder, K. Hyodo, T. Kotani, L. Kalvoda, T. Oda, Comparative study of electronic structure in ferromagnetic heusler alloys Ni2MnX (X = Al, Ga, In) using the quasi-particle self-consistentGWmethod, J. Magn. Soc. Jpn. 48 (6) (2024) 94–107.doi:10.3379/...

  9. [17]

    Marzari, A

    N. Marzari, A. A. Mostofi, J. R. Yates, I. Souza, D. Vanderbilt, Maximally localized Wannier functions: Theory and applications, Reviews of Modern Physics 84 (4) (2012) 1419–1475.doi: 10.1103/RevModPhys.84.1419

  10. [18]

    Okumura, K

    H. Okumura, K. Sato, T. Kotani, Spin-wave dispersion of3dferromagnets based on quasiparticle self- consistentGWcalculations, Phys. Rev. B 100 (2019) 054419.doi:10.1103/PhysRevB.100.054419

  11. [19]

    H.Okumura, K.Sato, T.Kotani, Nonlinearextensionofthedynamicallinearresponseofspin: Extended heisenberg model, J. Phys. Soc. Jpn. 90 (9) (2021) 094710.doi:10.7566/JPSJ.90.094710

  12. [20]

    Kotani, Quasiparticle self-consistentGWmethod based on the augmented plane-wave and muffin-tin orbital method, J

    T. Kotani, Quasiparticle self-consistentGWmethod based on the augmented plane-wave and muffin-tin orbital method, J. Phys. Soc. Jpn. 83 (9) (2014) 094711.doi:10.7566/JPSJ.83.094711

  13. [21]

    Takano, T

    S. Takano, T. Kotani, M. Obata, H. Fujii, K. Sato, H. Saito, T. Oda, in preparation

  14. [23]

    Kotani, M

    T. Kotani, M. van Schilfgaarde, Fusion of the LAPW and LMTO methods: The augmented plane wave plus muffin-tin orbital method, Phys. Rev. B 81 (2010) 125117.doi:10.1103/PhysRevB.81.125117. 23

  15. [24]

    Sakakibara, T

    H. Sakakibara, T. Kotani, M. Obata, T. Oda, Finite electric-field approach to evaluate the vertex correction for the screened Coulomb interaction in the quasiparticle self-consistentGWmethod, Phys. Rev. B 101 (2020) 205120.doi:10.1103/PhysRevB.101.205120

  16. [25]

    Landau, On the conservation laws for weak interactions, Nucl

    L. Landau, On the conservation laws for weak interactions, Nucl. Phys. 3 (1) (1957) 127.doi:10.1016/ 0029-5582(57)90061-5

  17. [26]

    Fujita, K

    R. Fujita, K. Konaga, Y. Ueoka, Y. Kamakura, N. Mori, T. Kotani, Analysis of anisotropic ionization coefficient in bulk 4H-SiC with full-band Monte Carlo simulation, in: 2017 International Conference on Simulation of Semiconductor Processes and Devices (SISPAD), IEEE, Kamakura...

  18. [27]

    Shishkin, M

    M. Shishkin, M. Marsman, G. Kresse, Accurate quasiparticle spectra from self-consistentGWcalcu- lations with vertex corrections, Phys. Rev. Lett. 99 (2007) 246403.doi:10.1103/PhysRevLett.99. 246403

  19. [28]

    W. Chen, A. Pasquarello, Accurate band gaps of extended systems via efficient vertex corrections in GW, Phys. Rev. B 92 (2015) 041115.doi:10.1103/PhysRevB.92.041115

  20. [29]

    Kotani, M

    T. Kotani, M. van Schilfgaarde, All-electron GW approximation with the mixed basis expansion based on the full-potential LMTO method, Solid State Commun. 121 (9) (2002) 461.doi:https://doi.org/ 10.1016/S0038-1098(02)00028-5

  21. [30]

    Introduction to the LAPW Method, in: D. J. Singh, L. Nordström (Eds.), Planewaves, Pseu- dopotentials and the LAPW Method, Springer US, Boston, MA, 2006, pp. 43–52.doi:10.1007/ 978-0-387-29684-5_4

  22. [31]

    Aryasetiawan, O

    F. Aryasetiawan, O. Gunnarsson, Product-basis method for calculating dielectric matrices, Phys. Rev. B 49 (1994) 16214.doi:10.1103/PhysRevB.49.16214

  23. [32]

    Friedrich, S

    C. Friedrich, S. Blügel, A. Schindlmayr, Efficient implementation of theGWapproximation within the all-electron FLAPW method, Phys. Rev. B 81 (2010) 125102.doi:10.1103/PhysRevB.81.125102

  24. [33]

    A. L. Fetter, J. D. Walecka, Quantum Theory of Many-Particle Systems, Dover Publications, 1971

  25. [34]

    J. Rath, A. J. Freeman, Generalized magnetic susceptibilities in metals: Application of the analytic tetrahedronlinearenergymethodtosc, Phys.Rev.B11(1975)2109.doi:10.1103/PhysRevB.11.2109

  26. [35]

    Kotani, Contracted plane wave satisfying periodic gauge, arXiv preprint arXiv:2010.03959 (2020)

    T. Kotani, Contracted plane wave satisfying periodic gauge, arXiv preprint arXiv:2010.03959 (2020)

  27. [36]

    W. Jia, J. Fu, Z. Cao, L. Wang, X. Chi, W. Gao, L.-W. Wang, Fast plane wave density functional theory molecular dynamics calculations on multi-GPU machines, J. Comput. Phys. 251 (2013) 102. doi:https://doi.org/10.1016/j.jcp.2013.05.005

  28. [37]

    S. Das, P. Motamarri, V. Subramanian, D. M. Rogers, V. Gavini, DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization, Comput. Phys. Commun. 280 (2022) 108473.doi:https://doi.org/10.1016/j.cpc.2022.108473

  29. [38]

    Pashov, S

    D. Pashov, S. Acharya, W. R. Lambrecht, J. Jackson, K. D. Belashchenko, A. Chantis, F. Jamet, M. van Schilfgaarde, Questaal: A package of electronic structure methods based on the linear muffin- tin orbital technique, Comput. Phys. Commun. 249 (2020) 107065.doi:https://doi.org...

  30. [39]

    Barker, D

    J. Barker, D. Pashov, J. Jackson, Electronic structure and finite temperature magnetism of yttrium iron garnet, Electron. Struct. 2 (4) (2021) 044002.doi:10.1088/2516-1075/abd097

  31. [40]

    V. W.-z. Yu, M. Govoni, GPU acceleration of large-scale full-frequency GW calculations, J. Chem. Theory Comput. 18 (8) (2022) 4690–4707.doi:10.1021/acs.jctc.2c00241. 24

  32. [41]

    B. Li, T. Patel, S. Samsi, V. Gadepally, D. Tiwari, MISO: exploiting multi-instance GPU capability on multi-tenant GPU clusters, in: Proceedings of the 13th Symposium on Cloud Computing, SoCC ’22, Association for Computing Machinery, New York, NY, USA, 2022, p. 173.doi:10.1145...

  33. [42]

    Gajdoš, K

    M. Gajdoš, K. Hummer, G. Kresse, J. Furthmüller, F. Bechstedt, Linear optical properties in the projector-augmented wave methodology, Phys. Rev. B 73 (2006) 045112.doi:10.1103/PhysRevB.73. 045112

  34. [43]

    Niquet, C

    Y.-M. Niquet, C. Delerue, Band offsets, wells, and barriers at nanoscale semiconductor heterojunctions, Phys. Rev. B 84 (2011) 075478.doi:10.1103/PhysRevB.84.075478

  35. [44]

    Ullakko, J

    K. Ullakko, J. K. Huang, C. Kantner, R. C. O’Handley, V. V. Kokorin, Large magnetic-field-induced strains in Ni2MnGa single crystals, Appl. Phys. Lett. 69 (1996) 1966.doi:10.1063/1.117637

  36. [45]

    Heczko, V

    O. Heczko, V. Kopecký, A. Sozinov, L. Straka, Magnetic shape memory effect at 1.7 K, Appl. Phys. Lett. 103 (2013) 072405.doi:10.1063/1.4817941

  37. [46]

    Sozinov, N

    A. Sozinov, N. Lanska, A. Soroka, W. Zou, 12% magnetic field-induced strain in Ni-Mn-Ga-based non-modulated martensite, Appl. Phys. Lett. 102 (2013) 021902.doi:10.1063/1.4775677

  38. [47]

    D. R. Baigutlin, V. V. Sokolovskiy, O. N. Miroshkina, M. A. Zagrebin, J. o. Nokelainen, Electronic structure beyond the generalized gradient approximation forNi2MnGa, Phys. Rev. B 102 (2020) 045127. doi:10.1103/PhysRevB.102.045127

  39. [48]

    P. J. Webster, K. R. Ziebeck, S. L. Town, M. S. Peak, Magnetic order and phase transformation in Ni2MnGa, Philos. Mag. B 49 (1984) 295.doi:10.1080/13642817408246515

  40. [49]

    A. P. Ongstad, R. Kaspi, C. E. Moeller, M. L. Tilton, D. M. Gianardi, J. R. Chavez, G. C. Dente, Spec- tral blueshift and improved luminescent properties with increasing GaSb layer thickness in InAs–GaSb type-II superlattices, J. Appl. Phys. 89 (4) (2001) 2185.doi:10.1063/1.1337918

  41. [50]

    Klein, E

    B. Klein, E. Plis, M. N. Kutty, N. Gautam, A. Albrecht, S. Myers, S. Krishna, Varshni parameters for InAs/GaSb strained layer superlattice infrared photodetectors, J. Phys. D: Appl. Phys. 44 (7) (2011) 075102.doi:10.1088/0022-3727/44/7/075102

  42. [51]

    Taghipour, E

    Z. Taghipour, E. Shojaee, S. Krishna, Many-body perturbation theory study of type-II InAs/GaSb superlattices within the GW approximation, J. Phys.: Condens. Matter 30 (32) (2018) 325701.doi: 10.1088/1361-648X/aacdce

  43. [52]

    Vurgaftman, J

    I. Vurgaftman, J. R. Meyer, L. R. Ram-Mohan, Band parameters for III–V compound semiconductors and their alloys, J. Appl. Phys. 89 (11) (2001) 5815.doi:10.1063/1.1368156

  44. [53]

    Mukherjee, A

    S. Mukherjee, A. Singh, A. Bodhankar, B. Muralidharan, Carrier localization and miniband modeling of inas/gasb based type-ii superlattice infrared detectors, J. Phys. D: Appl. Phys. 54 (34) (2021) 345104. doi:10.1088/1361-6463/ac0702

  45. [54]

    J. P. Perdew, K. Burke, M. Ernzerhof, Generalized gradient approximation made simple, Phys. Rev. Lett. 77 (1996) 3865.doi:10.1103/PhysRevLett.77.3865

  46. [55]

    Zelený, L

    M. Zelený, L. Straka, A. Sozinov, O. Heczko,Ab initioprediction of stable nanotwin double layers and 4O structure in Ni2MnGa, Phys. Rev. B 94 (2016) 224108.doi:10.1103/PhysRevB.94.224108

  47. [56]

    Ooiwa, K

    K. Ooiwa, K. Endo, A. Shinogi, A structural phase transition and magnetic properties in a Heusler alloy Ni2MnGa, J. Magn. Magn. Mater. 104 (PART 3) (1992) 2011.doi:10.1016/0304-8853(92)91645-A. 25

  48. [57]

    M. Khan, J. Brock, I. Sugerman, Anomalous transport properties of Ni2Mn1−xCrxGa Heusler alloys at the martensite-austenite phase transition, Phys. Rev. B 93 (5) (2016) 054419.doi:10.1103/PhysRevB. 93.054419

  49. [58]

    Janovec, M

    J. Janovec, M. Zelený, O. Heczko, A. Ayuela, Localization versus delocalization of d-states within the Ni2MnGa Heusler alloy, Sci. Rep. 12 (1) (2022) 20577.doi:10.1038/s41598-022-23575-1

  50. [59]

    Sponza, P

    L. Sponza, P. Pisanti, A. Vishina, D. Pashov, C. Weber, M. van Schilfgaarde, S. Acharya, J. Vidal, G. Kotliar, Self-energies in itinerant magnets: A focus on Fe and Ni, Phys. Rev. B 95 (2017) 041112(R). doi:10.1103/PhysRevB.95.041112

  51. [60]

    Momma, F

    K. Momma, F. Izumi,VESTA 3for three-dimensional visualization of crystal, volumetric and morphol- ogy data, J. Appl. Crystallogr. 44 (6) (2011) 1272.doi:10.1107/S0021889811038970. 26

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.