{"id":"f7cf262a-d3d6-47ca-af8c-1a47eae0f170","arxiv_id":"2507.05480","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MBFormer maps DFT mean-field states to GW quasiparticle energies and BSE exciton properties, achieving 0.16 eV and 0.20 eV MAE on held-out 2D materials and extrapolating from coarse to fine k-grids.","lead":"A new transformer-based machine learning model called MBFormer predicts quasiparticle and exciton properties of 2D materials directly from DFT inputs, achieving errors of 0.16 to 0.20 eV. If the results hold, it could make excited-state calculations cheap enough for high-throughput materials screening.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The BSE output head is unspecified: the decoder returns one scalar per electron-hole pair, yet exciton energies are eigenvalues of a global Hamiltonian; the reported 0.20 eV exciton MAE is therefore not verifiable from the paper as written.","rationale":"The reader's weakest assumption concerns label convergence and domain generalization, which is a valid empirical concern but would affect all ML supervisions. My review identified a more internal, architectural issue: for the BSE application—the paper's most novel contribution—the mapping from the transformer's per-query scalar outputs to the reported exciton eigenvalues is never specified. In the stated architecture (Sec. II.B), the decoder outputs one scalar per target basis state (electron-hole pair); exciton energies are eigenvalues of the full BSE Hamiltonian (Eq. 5) and, unlike GW corrections, are not naturally associated with individual pairs. Without a described sorting/diagonalization mechanism, the parity plots in Fig. 4(a) and the R2=0.999 fine-grid result cannot be reproduced or interpreted. This is not a question of label noise but of whether the model's output space can encode the target. The authors defer to an unavailable SI, so the claim is currently unverifiable. I therefore keep the CONDITIONAL verdict, but with an explicit condition: clarify the BSE output head or release the code/SI. The GW task, which has a clear per-state output, is less affected.","tokens_in":12395,"tokens_out":15741,"duration_ms":192856,"concrete_test":"Obtain the SI or code and identify the exact BSE output parameterization. If the model outputs N' per-pair scalars, verify that they are trained as a permutation-invariant set of eigenvalues (e.g., with a sorted-L1 or Chamfer loss) or that a predicted BSE Hamiltonian is diagonalized; if neither, the exciton energies in Fig. 4(a) cannot be produced by the described architecture. Decisive check: apply the described BSE-MBFormer to a minimal two-pair model with known eigenvalue splitting; if the per-pair outputs cannot represent the off-diagonal coupling, the 0.20 eV exciton MAE does not demonstrate learning of the interaction.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Sec. II.B, the model output is defined as a scalar vector of length N_input (the number of target basis states). For BSE, the target basis are electron-hole pairs |cvk> (Sec. II.D), so the output is one scalar per pair. However, exciton energies in Eq. 5 are eigenvalues of the BSE Hamiltonian—global quantities that have no intrinsic alignment with individual pairs. The paper never specifies how these per-pair scalars are converted into the exciton spectrum in Fig. 4(a) or Fig. 5(b): no sorted-eigenvalue loss, no diagonalization of a predicted kernel, and no assignment rule are stated. The only related statement is that the attention row can represent |Phi_alpha|, which is a wavefunction, not an energy. The fine-k-grid R2=0.999 result is likewise affected. The details are deferred to an unavailable SI (Notes 1 and 4), making it impossible to determine whether the model actually learns the electron-hole interaction kernel or is fitting a simpler scalar. This is the most load-bearing gap in the central claim of predicting two-particle excitonic properties.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MBFormer, an encoder-decoder transformer that takes Kohn-Sham DFT states as input tokens and predicts many-body properties derived from GW and GW-BSE calculations. The method is demonstrated on 721 two-dimensional materials from C2DB, with reported MAEs of 0.16 eV for G0W0 quasiparticle corrections and 0.20 eV for exciton energies, together with predictions of oscillator strengths and exciton wavefunction amplitudes. A separate experiment on monolayer hBN claims coarse-to-fine k-grid inference for excitonic properties. The authors argue that the attention mechanism is essential for capturing many-body correlations, based on an ablation that removes self-attention.","tokens_in":12717,"tokens_out":4920,"duration_ms":57474,"significance":"If the claims are substantiated, MBFormer would represent a meaningful step toward machine-learning models that map ground-state DFT data to excited-state many-body observables, including two-particle quantities that are difficult to obtain at DFT cost. The held-out test-set protocol, with the model trained on external GW-BSE labels rather than fitted to the target properties, gives the work a plausible generalization claim, and the inclusion of exciton wavefunction and oscillator-strength prediction goes beyond single-property ML models. However, the manuscript as submitted does not specify how the model converts per-electron-hole-pair decoder outputs into exciton eigenvalues, and it lacks baseline comparisons, error bars, and reference-data convergence analysis. The significance is therefore conditional on the missing BSE decoding specification and on the availability of the supplemental details.","major_comments":[{"comment":"The manuscript does not specify how the per-target-state scalar output of BSE-MBFormer is converted into exciton energies. In Sec. II.B the decoder returns a scalar vector in R^{N' x 1} with one entry per input target state; for the BSE task the target states are electron-hole pairs |cvk>. Exciton energies in Eq. (5) are eigenvalues of the N' x N' BSE Hamiltonian, and no mapping from per-pair scalars to the spectrum is given: there is no sorted-eigenvalue loss, no diagonalization of a predicted kernel, and no assignment rule. The attention row representing |Phi_alpha| is a wavefunction, not an energy. Because the reported exciton MAE of 0.20 eV, the R2 = 0.966 parity plot, and the fine-k-grid R2 = 0.999 result all rest on this unspecified step, the central two-particle claim is not verifiable from the submitted manuscript; the relevant SI Notes are not available to the reader.","section":"Sec. II.D, Eq. (5), Fig. 4(a)"},{"comment":"The description of the fine-k-grid experiment is ambiguous. The text says the model is trained 'on 6 coarse k-grids with 84 data augmentations (10% of the dataset)' from a total of 900 BSE Hamiltonians; 10% of 900 is 90, not 84, and 'data augmentations' is not defined. It is also unclear whether the training set consists of 6 grids, 84 Hamiltonians, or 84 augmentations of 6 grids, and whether region 1 in Fig. 5(a) is entirely excluded from training. These details are essential for evaluating the extrapolation claim that the model can infer excitonic properties on fine k-grids from coarse-grid training data.","section":"Sec. II.E, first paragraph"},{"comment":"The ablation that supports the claim that attention is crucial for BSE is not documented as a controlled experiment. The figure compares a 5-layer BSE-MBFormer without and with self-attention, but the text does not state that the two models are matched in parameter count, embedding dimension, training budget, or optimizer settings. Without this information, the observed 75% reduction in validation MSE cannot be attributed specifically to the self-attention module. In addition, the assertion that self-attention within the electron-hole pair basis is sufficient is justified only by a physical argument and by an unavailable SI Note 1; no cross-attention ablation is shown for the BSE model.","section":"Sec. II.D, Fig. 4(b)"},{"comment":"The claimed state-of-the-art performance is not supported by baseline comparisons or reference-data uncertainty analysis. The paper does not compare against existing GW-ML models such as the per-state representation model of Ref. [22], nor against simple baselines like linear regression on KS eigenvalues or a GNN-based predictor. The reported results come from a single random train/test split without error bars or multiple seeds, and the GW-BSE labels are assumed consistent and converged across the 721 C2DB materials, but no convergence tests or label-error analysis are presented. Consequently, the reported MAEs may not reflect true generalization error across different DFT setups, pseudopotentials, or material classes.","section":"General, Sec. III"}],"minor_comments":[{"comment":"The phrase 'determinant coefficient' should be 'coefficient of determination', and R2 is described as a 'determinant coefficient' in Fig. 4(a).","section":"Abstract"},{"comment":"The captions contain the typo 'applicaiton' instead of 'application'.","section":"Figure captions 3 and 4"},{"comment":"The text says 'scaler value prediction' when it means 'scalar value prediction'.","section":"Sec. II.E"},{"comment":"Essential technical details are repeatedly deferred to SI Notes 1, 2, 4, and 5, including the E2-VAE wavefunction embedder, the positional encoding of k-points and band indices, and the BSE energy decoding. At least the definitions needed to reproduce the core claims should appear in the main text or in a supplement available to the reviewers.","section":"Sec. II.B and II.D"},{"comment":"No data or code availability statement is provided; given the complexity of the pipeline and the reliance on an unpublished supplement, this hinders reproducibility.","section":"General"},{"comment":"The region labeled 'area 1' is used in the text and figure caption but is only implicitly defined; a precise description of which k-grids belong to area 1 and how it is selected should be stated.","section":"Fig. 5(a) and Sec. II.E"},{"comment":"The arrow notation in Eq. (2) represents a mapping whose input and output dimensionalities depend on the material, so the notation should be clarified to avoid implying a fixed-dimensional function.","section":"Eq. (2)"}],"recommendation":"major_revision","confidential_remarks":"The key issue is the missing BSE decoding specification in Sec. II.D: without it, the central two-particle claim cannot be assessed. If SI Notes 1 and 4 cannot be supplied, or if they do not resolve how per-pair scalar outputs become exciton eigenvalues, the paper's main result would not be established. I also note the absence of baseline comparisons and error bars, and the need for a data/code availability statement. The topic is suitable for the journal, and the issues appear fixable in a revision, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on MBFormer (arXiv:2507.05480). The GW part is genuinely promising; the exciton-energy claim is not verifiable from the paper as written, and that gap sits at the center of the paper's most interesting claims.\n\nWhat's new: a transformer encoder-decoder that treats Kohn-Sham states as tokens, with cross-attention between target states and a mean-field 'context' embedding, plus an E2-VAE wavefunction embedder. The demonstrations—GW corrections, exciton energies, oscillator strengths, and exciton wavefunction amplitudes—go beyond earlier ML GW work, and the coarse-to-fine k-grid extrapolation for hBN is a useful idea. The GW numbers (0.16 eV MAE, R²≈0.97 on held-out materials) are plausible, and the target-state definition for GW is clear: one scalar per KS state, matching the self-energy correction. That part holds up.\n\nThe BSE output is the problem. Section II.B defines the decoder output as a length-N' scalar vector—one value per electron-hole pair basis state. But exciton energies are eigenvalues of the BSE Hamiltonian; nothing in the paper explains how per-pair scalars become an exciton spectrum. The attention-row story covers |Φ_S|, not energies. With the details deferred to an unavailable SI, the 0.20 eV exciton MAE and the R²=0.999 fine-grid numbers could reflect a sorted-scalar fitting shortcut rather than learning the electron-hole interaction kernel. That is load-bearing, not a cosmetic footnote.\n\nOther soft spots are addressable: no baselines against existing GW-ML models, no error bars or multiple splits, no code or data release, and the interpretability claim rests on an ablation whose fairness isn't documented. The 'foundation model' language overshoots a 721-material, single-workflow, 2D-only dataset.\n\nMy take: the paper deserves refereeing, but the referee needs to demand a precise specification of the BSE output head, baselines, and data/code release. The core idea is useful; the exciton-energy claim is currently a promissory note. For a reading group it would be a good case study in evaluation gaps. I wouldn't cite the exciton results yet.","headline":"A useful architecture with a verifiable GW result, but the exciton-energy prediction is under-specified to the point of being unverifiable as written.","tokens_in":13180,"tokens_out":5181,"would_cite":false,"duration_ms":55987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A symmetry-aware transformer learns the many-body hierarchy directly from mean-field DFT inputs, predicting GW quasiparticle and exciton energies to within about 0.2 eV.","keywords":["transformer","many-body perturbation theory","GW approximation","Bethe-Salpeter equation","exciton","quasiparticle","machine learning","two-dimensional materials"],"falsifier":"Train or evaluate MBFormer on the same 721 materials but recompute the reference exciton energies with a denser k-grid and more empty bands, or with a different pseudopotential; if held-out MAE changes by more than roughly the claimed 0.2 eV tolerance, the reported accuracy does not transfer beyond the reference workflow.","tokens_in":12212,"feed_emoji":"⚛️","tokens_out":5033,"duration_ms":50090,"temperature":0.7,"pith_summary":"MBFormer is a transformer-based architecture that treats each Kohn-Sham state from a DFT calculation as a token, embeds the full set of mean-field states as context, and decodes them into many-body observables. The paper claims that this single paradigm can predict GW quasiparticle corrections, exciton energies, oscillator strengths, and exciton wavefunction amplitudes across hundreds of 2D materials, with mean absolute errors around 0.16-0.20 eV for energies. The goal is to show that the whole post-mean-field hierarchy can be learned end-to-end directly from ground-state inputs, bypassing the expensive construction of self-energy and interaction kernels. If true, this would turn a cheap DFT run into a fast estimator of excited-state properties, useful for high-throughput screening of optical and electronic materials.","feed_headline":"Transformer predicts excitons and quasiparticle gaps from DFT alone","feed_subtitle":"A single model learns GW and BSE corrections to within about 0.2 eV across 721 2D materials.","key_machinery":"The central object is a grid-free encoder-decoder transformer. Input mean-field Bloch states (wavefunction squared, energy, band index, k-point) are mapped by an equivariant variational autoencoder into compact latent tokens; a basis-assembly module combines those single-particle tokens into electron-hole pair tokens via concatenation. Self-attention in the encoder captures correlations among all mean-field states, cross-attention lets query states interact with the ground-state context, and the last decoder layer projects the mixed representation to scalars (energies, oscillator strengths) or to the distribution of exciton wavefunction amplitudes.","core_discovery":"On its own terms, the paper establishes that a single, symmetry-aware transformer, trained on the mean-field Kohn-Sham states of 721 two-dimensional semiconductors, can reproduce state-level GW-BSE results. For G0W0 self-energy corrections the test-set MAE is 0.16 eV with R2=0.969; for exciton energies the MAE is 0.20 eV, dropping to 0.15 eV for the first 100 bound excitons. The model also predicts exciton dipole oscillator strengths to a normalized MAE of 0.003 and exciton wavefunction amplitudes, and it can infer exciton properties on fine k-grids after coarse-grid training. The authors argue that the attention mechanism is what carries the many-body physics: removing self-attention raises the BSE validation MSE from 0.16 to 0.63 eV2, while simply deepening the network cannot compensate.","pith_inferences":["If the learned functional is as universal as claimed, a natural extension is to 3D materials, heterostructures, or defect states, but the current reference-label quality and the single database impose an unknown transfer penalty across DFT setups.","The reported errors are lower bounds in the sense that they are measured against labels from one consistent workflow; recomputed labels with different pseudopotentials or k-grids would likely show larger deviations.","For degenerate or nearly degenerate exciton states, the random rotation over the degenerate subspace acts as label noise; this is visible in the larger JSD for higher-energy excitons and could be treated by equivariant losses rather than raw state matching."],"forward_implications":["A single trained model can replace GW and BSE calculations for a range of materials, reducing the cost of excited-state screening from many-body perturbation theory to a forward pass through DFT inputs.","Because the model is grid-free and accepts variable numbers of source and query states, the same architecture can be retrained for other many-body targets, such as self-energy operators or higher-order Green's functions.","Fine-k-grid inference means the costly BSE convergence step can be skipped at prediction time: a model trained on coarse grids can resolve exciton dispersions, oscillator strengths, and wavefunction distributions on dense grids.","The explicit attention-matrix output provides a direct view of which mean-field states each exciton or quasiparticle draws from, offering interpretable access to the many-body correlation."],"supporting_citations":[{"why":"Supplies the transformer encoder-decoder and attention mechanism that the model is built on.","marker":"[26]"},{"why":"Defines the G0W0 self-energy and quasiparticle equation that the GW prediction task aims at.","marker":"[14]"},{"why":"Gives the Bethe-Salpeter equation Hamiltonian and electron-hole kernel that the exciton prediction task mirrors.","marker":"[15]"},{"why":"Provides the dataset of 721 two-dimensional materials whose DFT and GW-BSE labels are used for training and testing.","marker":"[27, 28]"},{"why":"Supplies the unsupervised VAE representation learning of Kohn-Sham states that the embedder extends with equivariance.","marker":"[39]"},{"why":"Documents how strongly GW quasiparticle energies depend on the number of empty bands, the convergence axis the model approximates via the number of source states.","marker":"[42]"}],"fun_headline_variants":["Attention learns many-body corrections for 2D materials","MBFormer: One transformer for GW-BSE predictions from DFT","From mean-field to excitons: A transformer learns many-body physics","Self-attention captures GW-BSE physics in 2D semiconductors","Transformer predicts exciton and quasiparticle energies to 0.2 eV"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 0.16-0.20 eV errors are only meaningful if the reference GW-BSE calculations across the 721 materials are accurate and mutually consistent; the paper gives no convergence tests of those labels, so the model may be learning workflow noise rather than physics.","fun_headline_variants_meta":{"raw":{"variants":["Attention learns many-body corrections for 2D materials","MBFormer: One transformer for GW-BSE predictions from DFT","From mean-field to excitons: A transformer learns many-body physics","Self-attention captures GW-BSE physics in 2D semiconductors","Transformer predicts exciton and quasiparticle energies to 0.2 eV"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00089,"raw_usage":{"total_tokens":3860,"prompt_tokens":990,"completion_tokens":2870,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2781}},"tokens_in":606,"tokens_out":2870,"duration_ms":21269,"temperature":1.0,"reasoning_tokens":2781,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:25:39.540649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train or evaluate MBFormer on the same 721 materials but recompute the reference exciton energies with a denser k-grid and more empty bands, or with a different pseudopotential; if held-out MAE changes by more than roughly the claimed 0.2 eV tolerance, the reported accuracy does not transfer beyond the reference workflow.","supporting_citations":[{"cited_title":"Data-driven Low-rank Approximation for Electron-hole Kernel and Acceleration of Time-dependent GW Calculations","cited_arxiv_id":"2502.05635","evidence_quote":"Supplies the transformer encoder-decoder and attention mechanism that the model is built on."},{"cited_title":"Electron correlation in semiconductors and insulators: Band gaps and quasiparticle energies,","cited_arxiv_id":null,"evidence_quote":"Defines the G0W0 self-energy and quasiparticle equation that the GW prediction task aims at."},{"cited_title":"General e (2)-equivariant steerable cnns,","cited_arxiv_id":null,"evidence_quote":"Supplies the unsupervised VAE representation learning of Kohn-Sham states that the embedder extends with equivariance."},{"cited_title":"Atomic positional embedding-based transformer model for predicting the density of states of crys- talline materials,","cited_arxiv_id":null,"evidence_quote":"Documents how strongly GW quasiparticle energies depend on the number of empty bands, the convergence axis the model approximates via the number of source states."}],"review_version":1}