{"id":"e7f76065-2e7b-4168-bdf3-be4c9704e5d5","arxiv_id":"2505.24807","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PySEQM 2.0 implements GPU-accelerated CIS and TDHF for semiempirical Hamiltonians and computes 20 excited states for a ~1000-atom carbon nanotube in about 45 seconds.","lead":"This paper adds excited-state calculations (CIS and TDHF) to the GPU-accelerated semiempirical package PySEQM, and reports that 20 excited states for a molecule with nearly 1000 atoms take about 45 seconds on a single A100 GPU. The significance is that large-scale excited-state dynamics, which usually require expensive TDDFT or coupled-cluster runs, might now be screened with fast semiempirical methods on GPUs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Convergence tolerance is never stated, so the 45 s CN-100 timing and 0.18 meV ORCA agreement lack a reproducible accuracy target; the memory-triggered restarts in Fig. 4B make this omission material.","rationale":"The paper is a software capability report; its central claim is that PySEQM produces correct CIS/TDHF semiempirical excited states for ~1000 atoms in under a minute, with energies matching ORCA to ~0.18 meV. For that claim to hold, the benchmark must be reproducible and the GPU/CPU/ORCA comparisons must target the same accuracy. The weakest link is the unspecified Davidson residual threshold in Section 3.3 Step 5. Since the same paragraph in Section 4 both claims \"over an order of magnitude speedup\" and reports 45 s vs 380 s (8.4x), and since Fig. 4B shows memory-triggered Krylov collapses with nonmonotonic timing for the same large-molecule regime as the headline, the benchmark as written cannot be independently checked. The ORCA validation on 28 small Thiel molecules is a genuine positive control and makes fundamental implementation error unlikely; it does not cover the memory-pressure regime where the large-system timing claim is made. This is fixable reporting and reproducibility work, not a fundamental flaw, so the conditional verdict is right. No change needed beyond requiring the tolerance, timing repeats, and ORCA input details.","tokens_in":14925,"tokens_out":10624,"duration_ms":117233,"concrete_test":"Re-run the CN-100 20-state CIS calculation on the same A100 at three Davidson residual-norm tolerances (e.g., 1e-4, 1e-6, 1e-8 a.u.) on GPU and CPU, recording wall time and the 20 lowest excitation energies. If the 45 s and 380 s timings are stable to <10% and all excitation energies agree to <0.01 meV across tolerances, the missing number is not material. If the 45 s figure is only obtained at a tolerance that shifts energies by >0.1 meV or differs from ORCA's default, the manuscript must state the tolerance and qualify the speedup.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3, Step 5, defines convergence only as \"the norm of every residual vector is below tolerance\" without giving the tolerance value. This value is load-bearing for both quantitative claims in Section 4. The 45 s GPU / 380 s CPU CN-100 timing is only a well-defined benchmark if both runs converge to the same threshold; if the GPU run uses a looser threshold, the speedup is inflated. The 0.18 meV (CIS) / 0.19 meV (RPA) agreement with ORCA is likewise only meaningful if PySEQM and ORCA target the same root accuracy. The paper also reports that for molecules with >500 atoms, GPU memory exhaustion forces Krylov subspace collapses and restarts with nonmonotonic timing. That large-system regime is exactly where the headline \"under a minute\" claim lives, and it is not covered by the ORCA validation, which is on small Thiel molecules. Without the tolerance and without a large-system accuracy check, the central claim is not reproducible. The same paragraph's \"order of magnitude\" speedup is inconsistent with its own 8.4x number, reinforcing that the benchmark details need tightening.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a GPU-accelerated implementation of CIS and TDHF/RPA excited-state calculations in the PySEQM semiempirical framework, built on PyTorch. The authors describe a batched Davidson algorithm with zero-padded subspaces and a modified Fock-like contraction for non-symmetric density matrices, and they benchmark it on carbon nanotubes, dendrimers, and perylene-diimide stacks. They report mean absolute excitation-energy deviations of 0.18 meV (CIS) and 0.19 meV (RPA) against ORCA on 28 small Thiel molecules, and a headline timing of about 45 s for 20 CIS states of a 996-atom carbon nanotube on a single A100 GPU. They also demonstrate batch-mode speedups for ensembles of geometries and highlight the software's machine-learning interface.","tokens_in":15083,"tokens_out":4189,"duration_ms":48310,"significance":"If correct, the implementation is a useful contribution: it extends an open-source GPU-accelerated semiempirical code to excited states using standard CIS/TDHF equations with no newly fitted parameters, and the ORCA comparison provides strong evidence that the core equations are implemented correctly. The batched Davidson mode is a practical advance for surface-hopping dynamics and ML-driven workflows, and the open-source availability is a clear strength. However, the reproducibility of the central performance and accuracy claims is weakened by missing convergence-tolerance specifications, absent timing statistics, and an inconsistency in the reported speedup.","major_comments":[{"comment":"The Davidson residual tolerance is never given; the text only states that molecules are converged when 'the norm of every residual vector is below tolerance.' This value is load-bearing for both the 0.18/0.19 meV agreement with ORCA and the CN-100 timing in Section 4. Please state the exact thresholds used for PySEQM and for ORCA, and report the per-molecule maximum absolute deviation along with the number of states included in the error statistics.","section":"Section 3.3, Step 5"},{"comment":"The headline 'about 45 s' for CN-100 is a single timing in a regime where Fig. 4B reports memory-triggered Krylov subspace collapses and nonmonotonic runtimes. No repetitions or variance are reported, and the CPU and GPU runs are not shown to be converged to the same threshold. Please provide repeated timings on the stated hardware, state the GPU memory configuration, and explicitly verify that the restarted runs converge to the same states satisfying the same criterion.","section":"Section 4, Fig. 4A and 4B"},{"comment":"The claim that PySEQM robustly handles up to 100 excited states in the memory-restart regime is a performance claim without an accuracy check in that regime. The validation against ORCA is limited to 5 states of small molecules; for CN-60/80/100, Dendrimer-4, and PDI-12 no reference comparison is given. I ask for a convergence check, such as final residual norms and a comparison of 20-state energies obtained with a tighter tolerance, to confirm that restarts do not alter the reported results.","section":"Section 4, Fig. 4B"}],"minor_comments":[{"comment":"The text says 'over an order of magnitude speedup' while the quoted numbers give 380/45 ≈ 8.4× and Fig. 4A says 'more than 8× speed-up'; please make these descriptions consistent.","section":"Section 4"},{"comment":"There are misspellings of 'Hartree' as 'Hartee' and 'Hatree', and 'neuclobase' appears for 'nucleobase' in Section 4; a proofread of these terms is needed.","section":"Introduction and Section 3.2"},{"comment":"The 'Supporting Information Available' section contains placeholder text ('This will usually read something like: ...') that should be replaced with an actual description of the supporting information.","section":"Supporting Information"},{"comment":"The batched Davidson description explains zero-padding of the subspace vectors but does not specify how the orthogonalization in Step 6 is performed for molecules with different subspace sizes; a brief statement of the batched orthogonalization procedure would improve reproducibility.","section":"Section 3.3, Steps 2 and 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is publishable after the benchmark reporting is brought to a reproducible standard. The missing residual tolerance and the single-timing measurements are the main obstacles; the underlying equations are standard and the ORCA agreement strongly suggests the implementation is correct. I did not find evidence of overclaiming beyond the 'order of magnitude' wording and the unverified robustness statement in the large-system regime."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, this is a solid software-capability paper, and the central claim is plausible: batched, zero-padded Davidson on GPUs gets semiempirical CIS/TDHF excited states for ~1000-atom molecules in under a minute, with excitation energies matching ORCA's semiempirical implementation to about 0.18 meV. Second, the paper has a real but fixable reporting gap: it never states the Davidson residual tolerance, and that omission sits right underneath both the 45-second timing and the ORCA agreement.\n\nWhat is actually new is the batched GPU Davidson implementation with zero-padding and its demonstrated scaling to ~1000 atoms and 100 states. The CIS/TDHF equations are textbook, and PySEQM already had GPU ground-state SCF and dynamics, so the contribution is the excited-state solver plus benchmarks. The validation against ORCA over 28 Thiel molecules is genuinely good evidence of numerical correctness, and there are no fitted parameters or circular claims here. The batch-mode scaling results are useful for people who want ensemble nonadiabatic dynamics.\n\nThe soft spots are mostly in the reporting. Section 3.3, step 5 defines convergence only as \"the norm of every residual vector is below tolerance\" with no number. That makes the 45 s versus 380 s comparison and the 0.18 meV MAE hard to reproduce: if GPU and CPU runs or PySEQM and ORCA used different tolerances, the speedup and agreement would not be apples-to-apples. Also, the paper says \"over an order of magnitude speedup\" but its own numbers give 8.4x, which is just wrong. The large-system timings in Fig. 4B show memory-triggered Krylov subspace collapses and nonmonotonic timing, so the headline 45-second number lives in exactly the regime that is not covered by the ORCA accuracy check. Minor issues: no timing repetitions, no raw data, no commit hash. None of these are load-bearing flaws in the implementation; they are missing details that a revision can supply.\n\nWho is this for? Anyone doing semiempirical excited-state dynamics, materials screening, or ML-parameterized Hamiltonians. The paper deserves a serious referee. My recommendation: send it to peer review, but require the authors to state the convergence tolerance, correct the speedup language, and add per-molecule error and timing statistics. With those, I'd trust the numbers.","headline":"A genuinely useful GPU implementation of semiempirical CIS/TDHF with credible ORCA validation, but the missing Davidson tolerance and overstated speedup language need fixing before I'd trust the headline numbers.","tokens_in":15713,"tokens_out":1449,"would_cite":true,"duration_ms":18189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PySEQM 2.0 runs CIS and TDHF excited-state calculations for molecules of nearly 1,000 atoms on a single GPU in under a minute.","keywords":["semiempirical quantum chemistry","excited states","CIS","TDHF","GPU acceleration","PyTorch","NDDO","batched Davidson eigensolver"],"falsifier":"Reproduce the CN-100 calculation (20 CIS states, AM1 Hamiltonian) on the same A100 GPU with the residual tolerance explicitly set, and record total time and excitation energies; if with a stated tolerance (for example, $10^{{-4}}$ a.u.) the runtime exceeds one minute, or the excitation energies deviate by more than about 0.2 meV from ORCA under identical tolerance, then the headline claim would be shown to be tolerance- or hardware-specific rather than intrinsic.","tokens_in":14682,"feed_emoji":"⚡","tokens_out":5997,"duration_ms":61667,"temperature":0.7,"pith_summary":"The paper claims that semiempirical excited-state calculations—configuration interaction singles (CIS) and time-dependent Hartree-Fock (TDHF)—can be made fast enough on GPUs to handle molecules of nearly one thousand atoms in about 45 seconds. Sitting on top of the PySEQM ground-state code and PyTorch, the implementation builds a batched Davidson iterative eigensolver that never forms the full CIS/RPA matrices. If the reported numbers hold, this turns excited-state electronic structure into a practical tool for nonadiabatic dynamics and ML-driven parameter optimization on large photoactive molecules. The paper also reports agreement with ORCA's semiempirical CIS and RPA excitation energies to about 0.18 meV on 28 benchmark molecules, which supports the correctness of the GPU implementation. The practical stakes are that large-scale excited-state dynamics and photophysics simulations become possible on a single GPU, where they previously required hours or were out of reach.","feed_headline":"GPU code computes excited states of 1,000-atom molecules in 45 seconds","feed_subtitle":"PySEQM's CIS/TDHF excited states match ORCA to 0.18 meV, opening large-scale nonadiabatic dynamics.","key_machinery":"The load-bearing object is the batched Davidson eigensolver implemented as 3-D tensor operations on the GPU. Instead of forming the A and B matrices of the CIS/RPA eigenvalue equations (size (N_occ N_vir)^2), it repeatedly computes matrix-vector products of the form (AV)_{ia} = (epsilon_a - epsilon_i) V_{ia} + sum_{mu nu} C_{mu i} \\tilde{F}_{mu nu} C_{nu a}, where \\tilde{F} is built by contracting NDDO two-electron integrals with a nonsymmetric density matrix. Zero-padded subspace vectors enforce uniform tensor widths so all molecules in a batch multiply in a single fused step; Krylov subspace size is adapted to available GPU memory, with collapse-and-restart when memory is exceeded. This machinery is what converts the formal O($N^{4}$) CIS cost into the observed near-O($N^{3}$) scaling on GPU.","core_discovery":"On the paper's own terms, the central claim is that a PyTorch-based implementation of CIS and TDHF for NDDO semiempirical Hamiltonians makes excited-state calculations scale to nearly 1,000 atoms in under a minute on a single GPU (20 excited states for CN-100 with ~996 atoms in ~45 s, versus ~380 s on CPU and ~78 min in ORCA). The same implementation is verified by reproducing ORCA's semiempirical CIS and RPA excitation energies to mean absolute errors of 0.18 meV and 0.19 meV over 28 Thiel benchmark molecules, and it includes a batched mode that processes ensembles of molecules concurrently. The authors would state the discovery as: GPU-native linear algebra plus NDDO integral sparsity reduces semiempirical excited-state cost to roughly O($N^{3}$) or lower, making thousand-atom excited-state simulations routine enough for ML-enabled nonadiabatic dynamics.","pith_inferences":["If one fixed a common residual tolerance and benchmarked on the same GPU, the 45-second figure would be a hardware- and memory-dependent data point; at larger state counts the reported nonmonotonic timings suggest the practical limit is GPU memory as much as arithmetic cost.","The 0.18 meV agreement with ORCA validates internal consistency of the two implementations, not ground-truth accuracy of AM1 excitation energies, since both codes use the same semiempirical Hamiltonian and reference.","A testable extension: computing excited-state gradients and nonadiabatic couplings in the same batched Davidson framework would let the method drive ab initio surface-hopping trajectories directly, which is the stated next step in the paper.","The batch zero-padding trick could transfer to other iterative eigensolvers in quantum chemistry (e.g., TDDFT with Tamm-Dancoff) wherever multiple systems need low-lying states simultaneously."],"forward_implications":["Nonadiabatic molecular dynamics with surface-hopping ensembles becomes feasible for systems of hundreds to about 1,000 atoms at semiempirical cost on a single GPU, since batch mode directly serves trajectory ensembles.","Excitation energies and transition properties for carbon nanotubes, dendrimers, and perylene-diimide stacks—materials relevant to light harvesting—can be generated in seconds rather than hours, enabling rapid screening of optoelectronic candidates.","Because the implementation is in PyTorch, the gradient chain can connect excited-state energies to SEQM parameter re-optimization and neural-network Hamiltonian corrections, extending machine-learning workflows to excited states.","Even CPU-only PySEQM beats a general quantum-chemistry package (45 s versus 78 min for CN-100), suggesting that the algorithmic structure, not just the hardware, carries much of the gain.","Memory-triggered Krylov collapses for large molecules point to a concrete next target: low-memory algorithms such as Lanczos or shift-based restarts for large state counts."],"supporting_citations":[{"why":"Supplies the ground-state NDDO engine and batched Born-Oppenheimer dynamics framework that this excited-state implementation extends.","marker":"[61]"},{"why":"Provides the 28-molecule Thiel benchmark set used to verify PySEQM's CIS and RPA excitation energies against ORCA.","marker":"[34]"},{"why":"Identifies ORCA 5.0.4 as the reference implementation for the reported excitation-energy comparison and the 78-minute CPU baseline.","marker":"[76]"},{"why":"Defines the AM1 semiempirical Hamiltonian used for all benchmark calculations in the paper.","marker":"[64]"},{"why":"Supplies the CIS and TDHF/RPA equations that the implementation and batched Davidson solver are built on.","marker":"[73]"},{"why":"Provides the Davidson iterative diagonalization method that the batched eigensolver adapts for GPU tensor operations.","marker":"[74]"},{"why":"Supports the simultaneous-expansion variant of Davidson used when computing several lowest eigenstates at once.","marker":"[75]"},{"why":"Establishes the machine-learning interface for re-optimizing semiempirical Hamiltonian parameters with neural networks, which the excited-state capability now connects to.","marker":"[62]"},{"why":"Provides the HIP-NN neural network model that PySEQM can interface with for learned Hamiltonian parameters.","marker":"[72]"}],"fun_headline_variants":["GPU-accelerated excited states: 1,000 atoms in 45 seconds","PySEQM 2.0: 1,000-atom excited states in under a minute","CIS/TDHF on GPU matches ORCA at 0.18 meV for 1,000 atoms","Excited-state calculations for 1,000 atoms speed up to 45s on GPU","Thousand-atom excited states on a single GPU in 45 seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmarks assume that all compared runs—PySEQM GPU, PySEQM CPU, and ORCA—use the same Davidson residual convergence tolerance, but the paper never states a numerical tolerance value, so the speedup and meV accuracy numbers are only comparable if the convergence criteria actually match.","fun_headline_variants_meta":{"raw":{"variants":["GPU-accelerated excited states: 1,000 atoms in 45 seconds","PySEQM 2.0: 1,000-atom excited states in under a minute","CIS/TDHF on GPU matches ORCA at 0.18 meV for 1,000 atoms","Excited-state calculations for 1,000 atoms speed up to 45s on GPU","Thousand-atom excited states on a single GPU in 45 seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00032,"raw_usage":{"total_tokens":1754,"prompt_tokens":845,"completion_tokens":909,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":792}},"tokens_in":461,"tokens_out":909,"duration_ms":9895,"temperature":1.0,"reasoning_tokens":792,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:12:57.500838+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reproduce the CN-100 calculation (20 CIS states, AM1 Hamiltonian) on the same A100 GPU with the residual tolerance explicitly set, and record total time and excitation energies; if with a stated tolerance (for example, $10^{{-4}}$ a.u.) the runtime exceeds one minute, or the excitation energies deviate by more than about 0.2 meV from ORCA under identical tolerance, then the headline claim would be shown to be tolerance- or hardware-specific rather than intrinsic.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ground-state NDDO engine and batched Born-Oppenheimer dynamics framework that this excited-state implementation extends."},{"cited_title":"R.; Thiel, W","cited_arxiv_id":null,"evidence_quote":"Provides the 28-molecule Thiel benchmark set used to verify PySEQM's CIS and RPA excitation energies against ORCA."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the AM1 semiempirical Hamiltonian used for all benchmark calculations in the paper."},{"cited_title":"Single- Reference ab Initio Methods for the Calculation of Excited States of Large Molecules","cited_arxiv_id":null,"evidence_quote":"Supplies the CIS and TDHF/RPA equations that the implementation and batched Davidson solver are built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Davidson iterative diagonalization method that the batched eigensolver adapts for GPU tensor operations."},{"cited_title":"The simultaneous expansion method for the iterative solution of several of the lowest eigenvalues and corresponding eigenvectors of large real-symmetric matrices","cited_arxiv_id":null,"evidence_quote":"Supports the simultaneous-expansion variant of Davidson used when computing several lowest eigenstates at once."},{"cited_title":"Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics","cited_arxiv_id":null,"evidence_quote":"Establishes the machine-learning interface for re-optimizing semiempirical Hamiltonian parameters with neural networks, which the excited-state capability now connects to."},{"cited_title":"S.; Barros, K","cited_arxiv_id":null,"evidence_quote":"Provides the HIP-NN neural network model that PySEQM can interface with for learned Hamiltonian parameters."}],"review_version":1}