{"id":"eab92bdf-b24d-45c4-b9b8-af87733805a2","arxiv_id":"2501.11019","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A symmetry-adapted kernel couples the angular momentum of atomic environments with the vector field of the density response, yielding rotation-equivariant and data-efficient predictions.","lead":"The paper introduces a symmetry-adapted machine-learning kernel that predicts how electron densities respond to electric fields, with full rotational equivariance for the resulting vector field. This makes accurate polarizability calculations feasible for metal nanoparticles with more than 2000 atoms, at a small fraction of the cost of quantum simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The >2000-atom scaling claim rests on five unrelaxed cuboctahedra with no DFT reference, so the headline extrapolation is unvalidated.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the 2057-atom scaling statement is validated only against DFT up to 201 atoms, with the five largest points having no reference. I agree this is the principal soft spot. The mathematical core—Eq. (2), the Clebsch-Gordan decomposition, and the O(3) symmetry assignment—is a standard and plausible group-theoretic construction, and the rotated-water equivariance test directly verifies the central symmetry claim. The liquid-water and naphthalene benchmarks provide credible evidence of data efficiency over SALTER, though error bars are absent. The damping schedule Mλ = M0 exp(−0.05λ²) is a pragmatic heuristic but is not the central claim; it affects accuracy but is not obviously fatal. The most defensible verdict is CONDITIONAL: the equivariance and benchmark results merit publication, while the headline extrapolation to >2000 atoms should be either supported by at least one reference calculation or explicitly labeled as an unvalidated prediction. Since the reader already reached CONDITIONAL for essentially this reason, my independent read does not move the verdict.","tokens_in":8926,"tokens_out":3023,"duration_ms":35577,"concrete_test":"Compute a reference DFPT or finite-field DFT polarizability for the 309-atom unrelaxed cuboctahedron (the smallest of the five extrapolation targets) using the same computational settings as the 201-atom benchmark set, and compare the resulting α0 (and, if feasible, the real-space density response profile) with the model prediction. If the error is comparable to the reported 1.2% collective RMSE, the extrapolation claim is substantially supported. If the error is significantly larger, the linear scaling of the five largest points is not established, and the abstract/headline should be softened to describe a physically motivated extrapolation rather than a reproduced scaling law. As a secondary check, include unrelaxed small clusters in the training size range to separate size extrapolation from relaxation-state extrapolation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's flagship application claim—reproducing the polarizability scaling law 'up to systems with more than 2000 atoms'—depends on predictions for five unrelaxed cuboctahedra (309, 561, 923, 1415, and 2057 atoms) for which the text explicitly states 'no reference value is computed because of the large computational cost.' The model is trained exclusively on relaxed nanoparticles of 28–139 atoms, covering icosahedra, decahedra, fcc-truncated octahedra, and randomly defected relaxed clusters. The five extrapolation targets therefore differ in both size and relaxation state from the training distribution. The observed linear α0-vs-natoms trend for these points is consistent with classical metallic-sphere electrostatics, but it is not a validated prediction of the ML model: the red points in Fig. 3 could lie on the line while the underlying density-response fields are inaccurate, if the unrelaxed cuboctahedral surfaces induce a different surface-charge distribution than the relaxed training geometries. This is the load-bearing part of the paper's 'accurate and inexpensive strategy ... up to systems with more than 2000 atoms' claim. The central equivariance construction in Eq. (2) and the water/naphthalene benchmarks are not affected by this concern, but the extrapolation claim is not yet supported by reference data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper derives a symmetry-adapted kernel for the machine-learning prediction of the electron-density response to an electric field, a vector-field property. The central result, Eq. (2), expresses the kernel in the λ⊗1 tensor-product space as a sum over spherical-tensor kernels of order l = |λ−1|, λ, λ+1 using Clebsch-Gordan coefficients, thereby enforcing rotational equivariance of the predicted vector field. The method is tested on rotated water molecules, liquid water and naphthalene supercells, and gold nanoparticles. For the molecular and periodic benchmarks, the equivariant model is reported to outperform the component-wise SALTER baseline in accuracy and data efficiency. For gold nanoparticles, the model is applied to relaxed clusters of 28–139 atoms and then extrapolated to larger sizes, including five unrelaxed cuboctahedra up to 2057 atoms, with the claim that the predicted polarizability reproduces the classical linear scaling of α0 with atom number.","tokens_in":9228,"tokens_out":3587,"duration_ms":41734,"significance":"If the results hold, this is the first fully equivariant kernel model for vector-valued electron-density responses, and it is an elegant extension of established spherical-tensor kernels: the construction in Eq. (2) is mathematically standard but is applied here to a target whose rotational symmetry has previously been handled only approximately (component-wise SALTER). The water/naphthalene benchmarks show a clear and consistent improvement over the non-equivariant baseline, and the availability of open code and reference data on Zenodo/GitHub is a concrete strength that supports reproducibility. The gold-nanoparticle application is interesting and potentially impactful, but the headline claim of reproducing the polarizability scaling law 'up to systems with more than 2000 atoms' is not backed by reference data in that size window. The central derivation is sound; the issue is the scope of the extrapolation claim.","major_comments":[{"comment":"The abstract and concluding statements claim that the method reproduces the polarizability scaling law 'up to systems with more than 2000 atoms'. This claim rests entirely on five unrelaxed cuboctahedra (309, 561, 923, 1415, and 2057 atoms) for which the text explicitly states 'no reference value is computed because of the large computational cost'. The model is trained on relaxed nanoparticles of 28–139 atoms, and reference validation extends only to 201 atoms. The five extrapolation targets therefore differ from the training distribution in both size and relaxation state. The observed linear trend is consistent with classical metallic-sphere electrostatics, but it is not a validated prediction of the ML model: the predicted points could lie on the line even if the underlying density-response fields are inaccurate. To support the headline claim, I ask for at least one reference DFPT or finite-field calculation in this size range (or an indirect validation, e.g., convergence of the predicted response field with respect to the basis); otherwise the claim should be softened to 'predicted scaling law consistent with classical electrostatics'.","section":"Gold nanoparticles section, Fig. 3, and abstract"},{"comment":"The accuracy and data-efficiency comparisons are based on single train/test splits without error bars or repeated resampling. For Table I, the differences between SALTER and the equivariant method are large, but the reader cannot assess the variability of these numbers. More seriously, the text reports 'A collective error of only 1.2% RMSE' for 'the whole set of 90 test structures', yet five of those structures have no reference values. The 1.2% figure must refer to the 85 structures with DFT references, or to the 55 in-distribution plus 30 extrapolation structures; this needs to be stated precisely. Reporting the error separately for the 28–139 atom test set and the 140–201 atom extrapolation set, with uncertainties over multiple splits, would materially strengthen the central accuracy claim.","section":"Table I and text near 'A collective 15% RMSE' and '1.2% RMSE'"},{"comment":"The atom-centered basis expansion in Eq. (1) must be able to represent the nonlocal, surface-dominated metallic polarization that is the main challenge of the Au nanoparticle application. The paper mitigates this by adding LODE features to the kernel, and Fig. 2 shows a good agreement for a 55-atom particle, but the basis-completeness question for larger metallic particles is not directly addressed. Since the headline extrapolation to >2000 atoms depends on the model's ability to represent the collective surface response, I ask for a test of basis convergence (e.g., dependence of α0 on λ_max and on the radial basis size for a representative extrapolated particle, even if only for a mid-size unrelaxed cuboctahedron where a reference is affordable).","section":"Eq. (1) and the gold-nanoparticle application"}],"minor_comments":[{"comment":"Typo: 'orthogonalizion' should be 'orthogonalization'.","section":"Section 2, text before Eq. (2)"},{"comment":"The text says the dataset has '100 rigid configurations that are randomly rotated, plus 900 that have a distorted geometry', but Fig. 1(b) refers to 'five water conformers that are only rotated with respect to one another'. Please clarify whether these five are a subset of the 100 rigid rotations and how the 900 distorted structures are distributed; this will make the equivariance test easier to reproduce.","section":"Dataset description in 'We test our method on a dataset of 1000 isolated water molecules' and Fig. 1(b)"},{"comment":"The gray line is described as a linear fit of α0 on the reference DFT values, but the caption does not state which points enter the fit (presumably only the structures with DFT references). Please specify this explicitly, and also state which points are used for the quoted 1.2% collective RMSE.","section":"Figure 3 caption and text"},{"comment":"The Gaussian damping schedule is introduced as a practical way to reduce the RKHS size, but no sensitivity analysis is provided. Reporting the dependence of the validation error on the damping exponent (e.g., 0.02, 0.05, 0.10) would show that the choice is not load-bearing for the reported accuracy.","section":"Section on Gaussian damping, Eq. (Mλ = M0 e^{-0.05 λ^2})"},{"comment":"The reported speedup of '>10^3' compared to DFPT or finite-field DFT would be more informative if the DFT reference cost and the basis set / k-point settings were itemized in the SI; as written, the comparison is difficult to reproduce.","section":"Computational timings paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for the journal and the core method is sound. My main concern is that the abstract and conclusion overstate the validation of the >2000-atom extrapolation; this is fixable either by computing a reference for at least one large cuboctahedron or by qualifying the claim. The absence of error bars on the benchmark numbers is a second issue that should be addressed in revision. I recommend major revision rather than rejection, because the central equivariant-kernel derivation and the molecular/liquid benchmarks are convincing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe takeaway: this paper earns its keep on the kernel construction, and stumbles only where it reaches beyond its validation. Eq. (2) — the Clebsch-Gordan coupling λ⊗1 in the spherical tensor kernel — is genuinely new for density-response learning, and the benchmarks on water and naphthalene show a clear, consistent accuracy gain over the component-wise SALTER baseline. That part is solid.\n\nWhat's good: the equivariance argument is clean, the derivation is a natural extension of established spherical tensor kernels (Refs. 33, 14), and the method is tested on a sensible progression: rotated monomers, liquid water, molecular crystals, and then metallic nanoparticles. The code and data are on GitHub and Zenodo, which makes the results reproducible. The efficiency claim — 40 training configurations beating 200 for SALTER — holds up in the table.\n\nThe soft spots are where the reported promise outruns the evidence. The flagship statement, reproducing polarizability scaling 'up to more than 2000 atoms', rests on five unrelaxed cuboctahedra (309 to 2057 atoms) for which the authors explicitly write 'no reference value is computed because of the large computational cost.' That is a real limitation. The model is trained on relaxed clusters of 28–139 atoms; the five targets are unrelaxed and much larger. The observed linear α0-vs-N trend is consistent with classical metallic-sphere electrostatics, but nothing in the paper verifies that the underlying density-response field is accurate at those sizes. If the surface-charge distribution differs for unrelaxed cuboctahedra, the red dots could fall on the line for the wrong reason. The stress-test note has this right. To be fair, the paper is transparent about the missing references — it says so in the text and in the figure caption — so this is an overstated extrapolation claim, not a concealment.\n\nTwo other things worth flagging, both minor in comparison. The accuracy numbers come from single train/test splits with no error bars, so the 9.26 vs 13.22 %RMSE differences could be noisier than they look. And the Gaussian damping Mλ = M0 exp(-0.05λ²) is a hand-set schedule; the paper doesn't report sensitivity to it.\n\nThe core equivariant construction and the molecular benchmarks hold up. I'd send this to review, with the request that the authors either compute references for at least a couple of the larger clusters or soften the 'up to 2057 atoms' claim to what the data actually support. It's a useful paper for anyone working on ML for electronic responses, and the code availability raises its value.\n\nRegards.","headline":"A clean equivariant kernel extension for vector-field density response, with an extrapolation claim that outruns the reference data.","tokens_in":9756,"tokens_out":2203,"would_cite":true,"duration_ms":21360,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A rotation-aware kernel learns how electron density responds to electric fields, sharply cutting errors.","keywords":["equivariant machine learning","electron density response","polarizability","kernel methods","SALTED","LODE","vector field","nanoparticles"],"falsifier":"Compute a DFPT reference polarizability for one of the unrelaxed cuboctahedra, such as the 309- or 561-atom geometries, and compare it with the model's prediction; a substantial deviation would show the extrapolation fails.","tokens_in":1472,"feed_emoji":"⚛️","tokens_out":3321,"duration_ms":53067,"temperature":0.7,"pith_summary":"This paper tries to establish that the electrostatic response of the electron density, a continuous three-dimensional vector field, can be learned with fully rotation-equivariant kernel models. The authors derive a kernel that couples the angular momentum of atom-centered basis functions with the vector character of the response using a Clebsch-Gordan sum, and they show this symmetry-adapted model outperforms the component-wise SALTER baseline on water, liquid water, and naphthalene benchmarks. For gold nanoparticles, the model reproduces the classical cubic scaling of polarizability with radius and extends predictions to clusters beyond 2000 atoms at small computational cost.","feed_headline":"Rotation-aware kernel predicts electron response with 50x lower error","feed_subtitle":"Machine-learned density response stays equivariant under rotations, cutting data needs and scaling to 2000-atom gold clusters.","key_machinery":"The central object is the symmetry-adapted kernel of Eq. (2), which enforces rotational equivariance of the learned density-response coefficients by coupling the angular momentum $\\lambda$ of the basis functions with the angular momentum $1$ of the vector field. It is built from standard spherical tensor kernels $K^l$, so the additional symmetry costs only a Clebsch-Gordan sum. A Gaussian damping of the number of sparse environments with increasing $\\lambda$ keeps the learning problem tractable, and the use of LODE features injects long-range information needed for metallic polarization.","core_discovery":"The central claim is that the response of the electron density to a homogeneous electric field, written as an expansion in atom-centered spherical harmonics, can be regressed with kernels that are equivariant under rotations by construction. The key identity expresses the kernel for the coupled angular momentum state $\\lambda \\otimes 1$ as a sum over irreducible spherical tensor kernels $K^l$ weighted by Clebsch-Gordan coefficients: $K^{\\lambda\\otimes 1}_{\\mu k,\\mu'k'} = \\sum_{l=|\\lambda-1|}^{\\lambda+1} \\langle\\lambda\\mu,1k|lm\\rangle \\langle\\lambda\\mu',1k'|lm'\\rangle K^l_{mm'}$. Because the prediction target is the full vector field, not its derived scalar properties, the model learns a local but transferable response. The authors demonstrate equivariance on rotated water molecules, data efficiency on liquid water and naphthalene, and, when combined with long-distance equivariant (LODE) features, recovery of the linear polarizability scaling of gold nanoparticles up to 2057 atoms.","pith_inferences":["This kernel construction could plausibly be applied to other vector-valued response properties, such as spin densities or current densities, where rotational equivariance is equally crucial.","The observed linear polarizability scaling on unrelaxed cuboctahedra is an extrapolation claim; it would be more convincing if at least one reference DFPT value were computed above 140 atoms.","The Gaussian damping of sparse environments for high angular momentum may limit accuracy for strongly oscillating density responses; a different decay schedule could be tested.","Because the response is learned in real space, the model may also capture spatially inhomogeneous polarizations, which connects to tip-enhanced Raman and interface responses."],"forward_implications":["The method yields density-response predictions that are exactly equivariant under rotations, eliminating the need to augment training data with rotated structures.","Training on 40 configurations with the equivariant model beats training on 200 with the component-wise baseline, indicating a large gain in data efficiency.","Polarizability tensors can be computed directly from the predicted $\\lambda=0$ and $\\lambda=1$ response coefficients, enabling cheap extrapolation to large metallic clusters.","For gold nanoparticles, only the LODE-augmented model reproduces the classical linear scaling of polarizability with atom count, while near-sighted SOAP features fail.","Prediction times of a few seconds for clusters of hundreds to thousands of atoms make the model a practical surrogate for DFPT calculations."],"supporting_citations":[{"why":"Defines the SALTER baseline that learns each Cartesian component independently and is the direct comparison for this work.","marker":"[28]"},{"why":"Supplies the spherical tensor kernel formalism for molecular tensors that the irreducible components $K^l$ are based on.","marker":"[33]"},{"why":"Provides the SALTED formulation for equivariant density learning whose low-dimensional RKHS recasting this work extends.","marker":"[21]"},{"why":"Earlier SALTED electron-density model that established the atom-centered basis expansion and kernel approach for densities.","marker":"[18]"},{"why":"Introduces the long-distance equivariant (LODE) features needed to capture nonlocal metallic polarization.","marker":"[38]"},{"why":"Shows tensor-product angular momentum decomposition for molecular properties, motivating the $\\lambda \\otimes 1$ construction.","marker":"[31]"},{"why":"Provides the DFPT implementation used to generate reference density-response data.","marker":"[11]"}],"fun_headline_variants":["Rotation-equivariant kernels predict electron response with less data","Symmetry-adapted kernels cut training data for electron response","Equivariant vector field kernel predicts electrostatic response","Scaling law for polarizability reproduced by equivariant ML","Rotation-aware kernels cut data needs for electron response"],"cache_read_input_tokens":11904,"weakest_assumption_plain":"The model is trained on relaxed gold nanoparticles with 28 to 139 atoms and then applied without retraining to unrelaxed cuboctahedra of up to 2057 atoms, where no reference values are computed, so the extrapolated linear polarizability scaling is asserted rather than verified.","fun_headline_variants_meta":{"raw":{"variants":["Rotation-equivariant kernels predict electron response with less data","Symmetry-adapted kernels cut training data for electron response","Equivariant vector field kernel predicts electrostatic response","Scaling law for polarizability reproduced by equivariant ML","Rotation-aware kernels cut data needs for electron response"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000917,"raw_usage":{"total_tokens":3914,"prompt_tokens":899,"completion_tokens":3015,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":515,"completion_tokens_details":{"reasoning_tokens":2939}},"tokens_in":515,"tokens_out":3015,"duration_ms":24673,"temperature":1.0,"reasoning_tokens":2939,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:42:30.214202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute a DFPT reference polarizability for one of the unrelaxed cuboctahedra, such as the 309- or 561-atom geometries, and compare it with the model's prediction; a substantial deviation would show the extrapolation fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the SALTER baseline that learns each Cartesian component independently and is the direct comparison for this work."},{"cited_title":"Grisafi, D","cited_arxiv_id":null,"evidence_quote":"Supplies the spherical tensor kernel formalism for molecular tensors that the irreducible components $K^l$ are based on."},{"cited_title":"Grisafi, A","cited_arxiv_id":null,"evidence_quote":"Provides the SALTED formulation for equivariant density learning whose low-dimensional RKHS recasting this work extends."},{"cited_title":"Nigam, M","cited_arxiv_id":null,"evidence_quote":"Shows tensor-product angular momentum decomposition for molecular properties, motivating the $\\lambda \\otimes 1$ construction."},{"cited_title":"Shang, N","cited_arxiv_id":null,"evidence_quote":"Provides the DFPT implementation used to generate reference density-response data."}],"review_version":1}