{"id":"8393a00d-a303-4bc2-880b-3e833798b287","arxiv_id":"2607.14025","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"WEST achieves scalable excited-state simulations for thousand-atom systems by avoiding explicit virtual states across GW, TDDFT, BSE, and QDET.","lead":"This paper describes WEST, an open-source code that computes excited-state properties of materials using GW, TDDFT, Bethe-Salpeter, and quantum-embedding methods without explicitly computing empty electronic states. It summarizes benchmarks and applications showing the code can handle systems with more than a thousand atoms on GPU supercomputers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"New strong-scaling benchmark (Fig. 2) lacks raw wall-clock data and absolute timings, leaving the headline 'near-ideal scaling to thousands of GPUs' claim unverifiable as presented.","rationale":"The paper is a credible overview: the code is open source, the methods are well-cited, and several results (NV− lifetime, PL spectra, perovskite emission) match experiment. I am not arguing the scaling claim is false. But the claim that the code achieves near-ideal scaling to 4,096 GPUs is presented in Fig. 2 as new evidence, and it is the least supported piece in the manuscript: the reader sees only slopes, not numbers. The reader's weakest_assumption — systematic convergence of the PDEP/Wannier/TDA approximations across diverse materials — is a substantive scientific concern, but the manuscript explicitly lists TDA and single-excitation restrictions as known limitations, defers to previous publications for detailed benchmarks, and the new data most directly tied to the headline 'scalable platform for predictive excited-state simulations' are the GPU scaling results. Thus I place the load-bearing concern on the missing benchmark data. If the benchmarks reproduce with high parallel efficiency, the CONDITIONAL verdict can move to ACCEPT; if not, the scaling claim should be weakened. This does not change the reader's CONDITIONAL verdict, hence UNCHANGED. My agreement is partial because the reader's weakest_assumption points to a different (also valid) issue, while the missing raw scaling data appears in the reader's rationale but not as the primary concern.","tokens_in":23442,"tokens_out":8676,"duration_ms":88328,"concrete_test":"Request or obtain from the WEST GitHub repository the exact input files and job scripts used for Fig. 2; rerun the four benchmarks (QDET, G0W0, BSE, TDDFT) for the 999-atom NV− diamond supercell on Perlmutter at 64, 256, 1024, and 4096 GPUs. Extract wall times from output logs, compute parallel efficiency E_P = T_64 / ((P/64) * T_P), and compare absolute G0W0 and BSE wall times with the y-axis values of Fig. 2. If E_4096 ≥ 90% and the log timings match the figure within noise, the scaling claim is confirmed; otherwise the manuscript should report measured efficiencies and absolute runtimes, and the abstract's scaling claim should be revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that WEST 'deliver[s] near-ideal strong scaling to thousands of GPUs' is supported in this paper almost entirely by Fig. 2, a log-log plot of wall time vs GPU count for four 999-atom NV− calculations on Perlmutter. The figure reports no raw wall-clock values, no absolute timings, no parallel-efficiency numbers, and no repeated runs or error bars. The caption says only that 'minor deviations from ideal scaling at large GPU counts are attributed to I/O and inter-node communication overhead,' which is not a quantitative statement. Without these data, a reader cannot verify that the slopes correspond to, e.g., ≥90% parallel efficiency at 4,096 GPUs, nor assess whether the absolute time per calculation is short enough to make 'more than a thousand atoms' simulations practical. Because this scaling claim is the headline capability and Fig. 2 appears to be the main new evidence in this manuscript (the other scalability citations are to prior work), the evidential gap is load-bearing. This is a reproducibility/evidence concern, not an accusation of error; the code is open source and the benchmarks should be reproducible.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents the WEST code, an open-source plane-wave pseudopotential package for large-scale excited-state materials simulations. It reviews the theoretical foundations (full-frequency G0W0, QDET, BSE, TDDFT), the algorithmic strategies that avoid explicit empty-state summations (PDEP, DMPT, Wannier localization, ACE), the hierarchical CPU/GPU parallelization, and interoperability with external packages. It reports a new strong-scaling benchmark (Fig. 2) on up to 4,096 GPUs for 999-atom NV− calculations, and illustrates capabilities through applications: NV− optical cycle, SiV0 finite-size extrapolation, NV centers near dislocations in 1,727-atom supercells, spin defects in oxides/nitrides/2D materials/molecular qubits, self-trapped excitons in perovskites, and water/ice optical response. The central claim is that WEST delivers near-ideal scaling to thousands of GPUs and enables accurate excited-state simulations of systems with more than a thousand atoms across diverse material classes.","tokens_in":23737,"tokens_out":2634,"duration_ms":28921,"significance":"If the claims are substantiated, WEST is a significant community resource: one open-source code providing quasiparticle energies, neutral excitation energies, spectra, excited-state forces, and non-adiabatic couplings without explicit virtual states, with demonstrated applications to experimentally relevant systems. The manuscript is a code overview rather than a new methodological derivation; most algorithmic advances are published elsewhere. Its strengths are the broad and coherent set of validated applications (GW100 benchmarks and experimental comparisons for NV−, SiV0, perovskites, water/ice), the open-source release with documentation and continuous integration, and the clear acknowledgment of approximation caveats (TDA, static screening, single-excitation restriction, frequency-independent QDET Hamiltonian). The main new evidence is the GPU scaling figure, which is also the weakest link: it lacks absolute timings, raw wall-clock values, parallel-efficiency numbers, and error bars, making the headline scalability claim difficult to verify.","major_comments":[{"comment":"The strong-scaling claim of 'near-ideal scaling to thousands of GPUs' rests almost entirely on Fig. 2, which plots wall time vs. GPU count without reporting absolute wall-clock values, per-point timings, parallel efficiency percentages, or repeated-run statistics. The caption attributes deviations to I/O and communication overhead, but no quantitative evidence is given. This is a load-bearing claim for the abstract and introduction ('thousands of GPUs', 'more than a thousand atoms'). Please provide a table of timings (or in-figure labels), compute parallel efficiencies relative to the 64-GPU point, and specify whether these are single runs. The benchmarks should be reproducible from the open-source code and stated parameters.","section":"Fig. 2 and Sec. III.B"},{"comment":"The finite-size extrapolation of the bound-exciton VEE uses Eq. (29) with fitting parameters A, B, and D, where D is varied over [10,40] Å, yielding E_BE(∞) ∈ [1.33,1.50] eV. The paper reports this as 'close agreement' with experiment (1.39 eV), but the 0.17 eV spread is comparable to the spread in the other VEE comparisons and is not propagated to the conclusion. Since D is a screening length with a plausible but not uniquely determined range, the uncertainty in the extrapolated value should be discussed explicitly and the sensitivity to D shown, rather than presenting a point-like agreement.","section":"Eq. (29) and Sec. IV.B.1"},{"comment":"The manuscript explicitly acknowledges key approximations (TDA, neglect of dynamical screening, restriction to single excitations in TDDFT/BSE; frequency-independent effective Hamiltonian in QDET) but does not provide a systematic convergence or error analysis across the claimed range of applications. Given the breadth of systems (spin defects, perovskites, water/ice, 2D materials), the reader is asked to rely on selected benchmarks. Adding a concise table of convergence parameters (N_PDEP, k-points, supercell size) and, where possible, estimated errors from these approximations for representative cases would materially strengthen the accuracy claim.","section":"Sec. IV.A.2 / Sec. II.C"}],"minor_comments":[{"comment":"Typo: 'Owing to the complexity of the ISC and ISC processes' should read 'ISC and IC processes'.","section":"Sec. IV.A.4"},{"comment":"The caption states 'minor deviations from ideal scaling' but no error bars or repeated runs are shown; a statement about wall-clock time per QDET/G0W0/BSE/TDDFT calculation (e.g., in minutes) would help judge practical utility.","section":"Fig. 2"},{"comment":"The claim of 'excellent scalability to 25,920 GPUs' cites reference 20 but is not shown or quantified in this manuscript; clarify that this is from prior work and refer to the specific data.","section":"Sec. III.B"},{"comment":"The number of Coulomb integrals reduced 'by more than an order of magnitude' by Wannier localization is stated without a quantitative example or reference for the 999-atom system; a concrete number would strengthen the efficiency argument.","section":"Eqs. (10)-(11)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the WEST paper. It's a code overview rather than a new-methods paper; almost all the equations and applications were published by the group between 2015 and 2025. What's genuinely new is the consolidated picture: one open-source framework covering full-frequency G0W0, QDET, BSE, and TDDFT without explicit empty states, plus a strong-scaling benchmark on a 999-atom NV− supercell.\n\nThe paper does several things well. It's clearly written, the algorithmic ideas (PDEP, density-matrix perturbation theory, Wannier localization) are coherent, and the stated scalings are plausible. The applications are representative and mostly cross-checked against experiment: GW100, NV− ZPL and lifetimes, perovskite emission, water/ice absorption. The code is open source with documentation and continuous integration. That's real, reproducible infrastructure, and the external benchmarks give credibility.\n\nThe soft spot is the new scaling claim. Figure 2 plots wall time vs GPU count for QDET, G0W0, BSE, and TDDFT on 4,096 GPUs, but there are no raw wall-clock values, no parallel-efficiency numbers, no error bars, no repeated runs. The caption attributes deviations to I/O and communication overhead without quantifying them. Since 'near-ideal scaling to thousands of GPUs' is the headline capability and this figure is the main new evidence, that's a load-bearing gap. I don't doubt the code scales well—prior work cites 25,920 GPUs—but as presented, the claim isn't verifiable. The fix is straightforward: release input files and raw timing data, add efficiency curves, maybe a table of absolute times.\n\nSmaller quibbles: a few spectra are shifted to align with experiment (e.g., Cs4SnBr6 emission), which is fine for line-shape comparison but should be labeled more prominently as alignment. The SiV0 bound-exciton extrapolation uses a screening length D chosen in [10,40] Å with fitted A and B; the resulting range brackets experiment, but it's not fully first-principles. None of this undermines the core value.\n\nVerdict: this is a useful reference for anyone working on excited-state simulations of large or heterogeneous systems, and it makes a solid case for WEST as a platform. The scaling data need to be supplied, and the spectral alignments flagged. I would send it to peer review—a good referee will ask for the benchmark raw data and a convergence check on N_PDEP and localization approximations, and the authors should be able to provide both.\n\nFor your reading group, maybe; worth a look if you care about GPU scaling of MBPT.","headline":"Useful consolidated overview of a mature excited-state code, but the headline scaling claim rests on a figure with no raw timings.","tokens_in":24224,"tokens_out":1956,"would_cite":true,"duration_ms":20825,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One open-source code, WEST, unifies full-frequency GW, quantum-defect embedding, BSE, and TDDFT in a framework that skips virtual electronic states, so excited-state simulations of more than one thousand atoms become practical.","keywords":["WEST code","GW approximation","Bethe-Salpeter equation","time-dependent density functional theory","quantum defect embedding","excited-state forces","GPU scaling","plane-wave pseudopotential"],"falsifier":"Compare WEST's BSE and TDDFT excitation energies against a conventional full BSE calculation on a set of small molecules, intentionally varying the PDEP truncation count and the Wannier interaction cutoff, and check for monotonic drift or a sudden onset of large errors as the approximations are loosened.","tokens_in":23344,"feed_emoji":"⚛️","tokens_out":4252,"duration_ms":43610,"temperature":0.7,"pith_summary":"The paper presents WEST, an open-source plane-wave code that implements four major excited-state methods — full-frequency GW, quantum-defect embedding theory, the Bethe-Salpeter equation, and time-dependent density functional theory — within a single algorithmic framework. Its central claim is that by eliminating explicit sums over empty electronic states and compressing the dielectric response into a low-rank form, these methods scale to supercells of a thousand or more atoms and to thousands of GPUs with near-ideal strong scaling. If that claim holds, one code can deliver quasiparticle energies, neutral excitation energies, optical and photoluminescence spectra, excited-state forces, non-adiabatic couplings, and decay rates for complex, heterogeneous materials at experimentally relevant sizes. The paper supports this through benchmarks and applications to spin defects, metal-halide perovskites, and water and ice.","feed_headline":"One code base covers GW, BSE, TDDFT and quantum embedding at scale","feed_subtitle":"WEST avoids empty-state sums, making excited-state simulations practical for 1,000+-atom systems.","key_machinery":"The load-bearing object is the projective dielectric eigenpotential (PDEP) expansion of the density–density response function, which supplies a low-rank representation of dielectric screening that both G0W0 and the Bethe-Salpeter equation draw upon without constructing large dielectric matrices or empty-state sums. Around it, the code builds a common algorithmic core: density-matrix perturbation theory recasts BSE/TDDFT diagonalization as an action on occupied wave functions, Wannier localization prunes the Coulomb integrals to spatially overlapping pairs, and the adaptively compressed exchange (ACE) method keeps hybrid-functional calculations at semilocal cost. These components are unified","core_discovery":"The central discovery, on the paper's own terms, is that the dominant cost of excited-state calculations — the explicit enumeration of empty electronic states — can be removed across four distinct formalisms and replaced by a common set of scalable operations. The projective dielectric eigenpotential (PDEP) technique iteratively diagonalizes the symmetrized density–density response function, expressing the screened Coulomb interaction as a low-rank sum of a few eigenpotentials (Eq. 5). The Bethe-Salpeter and TDDFT problems are reformulated in terms of the density-matrix perturbation theory, so the Liouville superoperator acts only on occupied states, and Wannier localization of the occupied","pith_inferences":["If the scalability and accuracy claims hold, the most natural next stress test is metallic or small-gap systems, where the dielectric response is more delocalized and the PDEP low-rank form may require many more eigenpotentials — the code's own approximations predict where the approach will get harder.","Since energies, forces, and non-adiabatic couplings are all available at the same level of theory, the framework should enable ab initio non-adiabatic molecular dynamics in extended heterogeneous systems, a step the paper describes as crucial but does not itself demonstrate.","The Wannier-localization speedup suggests a testable scaling law: for sparse systems, wall time should grow roughly with the number of overlapping occupied-state pairs rather than the square of the number of atoms, a prediction a user could verify with controlled supercell sizes.","The claim that empty states can be avoided altogether could be probed by blind comparisons against a conventional full-BSE implementation on a shared set of small molecules, reusing identical geometries and pseudopotentials."],"forward_implications":["Quasiparticle energies, neutral excitation energies, optical and photoluminescence spectra, excited-state forces, non-adiabatic couplings, and radiative/non-radiative decay rates become available for supercells of 1,000+ atoms, including environments with dislocations, surfaces, and interfaces.","Because all methods share one framework and one set of numerical parameters, results across different levels of theory (TDDFT vs BSE vs QDET) become directly comparable on identical geometries, removing a common source of inconsistency in the literature.","The convergence of response functions is controlled by a single parameter — the number of PDEP eigenpotentials — rather than a separate empty-state cutoff, making convergence checks more systematic.","The same first-principles output can be piped into quantum chemistry solvers for embedded multi-configurational states, vibronic coupling calculations, and quantum-computing diagonalization workflows, extending the reach of each individual method.","The scalability demonstrated opens the door to generating large, high-fidelity excited-state datasets for machine-learning potentials and high-throughput screening campaigns."],"fun_headline_variants":["WEST code skips empty states for 1000+ atom excitations","Excited states without empty states: WEST scales to 1000 atoms","WEST: one code, four formalisms, no virtual orbitals","WEST avoids empty states, making GW/BSE/TDDFT practical at scale","WEST electron code scales to 1000 atoms by skipping virtual states"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The scalability and accuracy claims rest on the assumption that the low-rank and localization approximations do not systematically degrade results across the diverse materials discussed: a limited number of PDEP eigenpotentials and Wannier-truncated Coulomb integrals must preserve accuracy for spin defects, perovskites, and water and ice alike.","fun_headline_variants_meta":{"raw":{"variants":["WEST code skips empty states for 1000+ atom excitations","Excited states without empty states: WEST scales to 1000 atoms","WEST: one code, four formalisms, no virtual orbitals","WEST avoids empty states, making GW/BSE/TDDFT practical at scale","WEST electron code scales to 1000 atoms by skipping virtual states"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000708,"raw_usage":{"total_tokens":3043,"prompt_tokens":776,"completion_tokens":2267,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":2168}},"tokens_in":520,"tokens_out":2267,"duration_ms":14740,"temperature":1.0,"reasoning_tokens":2168,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:57:24.817734+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare WEST's BSE and TDDFT excitation energies against a conventional full BSE calculation on a set of small molecules, intentionally varying the PDEP truncation count and the Wannier interaction cutoff, and check for monotonic drift or a sudden onset of large errors as the approximations are loosened.","supporting_citations":[],"review_version":1}