Pith. sign in

REVIEW 4 major objections 7 minor 22 references

Edge many-body products and radial rotary attention push SO(2) interatomic potentials past prior Matbench leaders.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 10:07 UTC pith:GB27VOGS

load-bearing objection Solid SO(2) methods paper with real operators and ablations; the Matbench SOTA is a 0.001 CPS edge and should not be the main reason you care. the 4 major comments →

arxiv 2607.10664 v1 pith:GB27VOGS submitted 2026-07-12 stat.ML cond-mat.mtrl-scics.LGphysics.chem-ph

Edge Cluster Expansion with Radial Rotary Attention for Interatomic Potentials

classification stat.ML cond-mat.mtrl-scics.LGphysics.chem-ph
keywords machine learning interatomic potentialsSO(2) equivarianceEdge Cluster ExpansionRadial Rotary Complex AttentionWigner-D matricesAtomic Cluster ExpansionMatbench Discovery
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Machine learning interatomic potentials need equivariant operations that are both accurate and cheap. The authors show that common SO(2) linear layers underperform full SO(3) tensor products when radial information is mishandled or when complex phases break reflection symmetry, and they fix the construction of the rotation matrices that move features into the local edge frame. They then put many-body expansion directly on edges via a generalized asymmetric contraction (Edge Cluster Expansion) and replace ordinary attention with Radial Rotary Complex Attention, whose logits mix node similarity with learned radial phase and bias. Together with cleaner Atomic Cluster Expansion, residual, and normalization choices, the resulting TECE-OAM-RRA-1.0 model trained on OMat24, sAlex, and MPTrj reaches the highest overall Matbench Discovery score among the models they compare. A sympathetic reader cares because the work both diagnoses why SO(2) models often extrapolate poorly and supplies concrete operators that raise body order and radial inductive bias without discarding the efficiency of the local frame.

Core claim

Conventional SO(2) Linear is weaker than Clebsch–Gordan tensor products mainly because of design choices (path-wise radial weights, real versus complex weights, edge nonlinearities), not because of SO(2) equivariance itself; Edge Complex Product Basis via generalized asymmetric contraction raises effective body order on edges, and Radial Rotary Complex Attention improves extrapolation over prior attention-vector schemes, yielding TECE-OAM-RRA-1.0 with state-of-the-art overall Matbench Discovery performance after training on OMat24, sAlex, and MPTrj.

What carries the argument

Edge Cluster Expansion (ECE): generalized asymmetric contraction that builds higher-order product bases on edge features with SO(2) coupling coefficients, paired with Radial Rotary Complex Attention whose logits are the real part of a complex QK product scaled by a radial phase and shifted by a radial bias.

Load-bearing premise

Modules are kept or discarded mainly by whether they keep force error below a fixed high-temperature molecular threshold, treating that single check as a reliable stand-in for how well a universal materials model will extrapolate.

What would settle it

Train the same TECE stack with ECE and RRA ablated (or replaced by standard symmetric contraction and Equiformer-style attention vectors) on the same OMat24/sAlex/MPTrj mix and re-score Matbench Discovery; if CPS, F1, and kappa_SRME no longer lead, the claimed gains do not come from those blocks.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The manuscript systematically analyzes SO(2) equivariant operators for machine learning interatomic potentials, contrasts uuSO2Linear with uvSO2Linear and real versus complex SO(2) weights, and proposes two constructions of Wigner-D matrices (direct Cartesian and recursive Clebsch–Gordan). It introduces Edge Cluster Expansion (ECE) via generalized asymmetric contraction on edges and Radial Rotary Complex Attention (RRA), plus ACE improvements (group linear, GLU-style nonlinear coefficients, residual/LayerNorm choices). Controlled molecular ablations (3BPA/AcAc, M-fold symmetry, Pozdnyakov body-order graphs) support the design choices. Models trained on OMat24, sAlex, and MPTrj yield TECE-OAM-RRA-1.0, reported as state-of-the-art on Matbench Discovery (CPS 0.908).

Significance. If the methodological claims hold, the paper is a useful consolidation of SO(2) practice for MLIPs: the O(2) analysis of complex weights (Eq. 14), the uu/uv path distinction, the recursive Wigner-D construction with favorable scaling (Fig. 1), and the completeness arguments for ECE (M-fold math and Table 3 body-order tests) are concrete contributions. RRA’s radial phase/bias design is a clear inductive-bias improvement over attention-vector baselines on 3BPA (Table 2). Code and model releases strengthen reproducibility. The Matbench ranking is of practical interest to the materials community, though the reported margin is very small. Overall significance is primarily architectural and theoretical rather than a decisive leap in materials accuracy.

major comments (4)
  1. Table 4 and Appendix Table 7: TECE-OAM-RRA-1.0 is ranked SOTA by CPS 0.908 vs EquFlashV2 0.907. The two models share F1=0.929 and RMSD=0.058; the only reported difference is κSRME 0.093 vs 0.094. No multi-seed variance, bootstrap, or uncertainty on CPS is given. A 0.001 CPS edge is not, by itself, a robust SOTA claim. Either quantify ranking uncertainty, report multiple independent runs, or temper the abstract/introduction language to “competitive / among top models” unless stronger evidence is added.
  2. Section 3.6: architecture modules are retained only if 3BPA 1200K force RMSE stays ≤65 meV/Å. This single molecular high-T threshold is treated as the gate for a universal materials model. The manuscript does not show that modules ranked by this gate also rank by Matbench F1/κSRME/RMSD, nor that ECE/RRA (vs ablated variants) improve Matbench metrics. Without materials-side ablations or a demonstrated correlation between the 3BPA gate and Matbench, the causal link from ECE/RRA to the Matbench ranking remains under-supported.
  3. Section 3.11 states Adam is best for extrapolation and that Muon/SOAP can compromise it; Appendix A reports the released TECE-OAM-RRA-1.0 was trained with Muon for speed. For a paper whose narrative emphasizes extrapolation (RRA, real weights, 3BPA high-T), this is a load-bearing inconsistency. Please either retrain/fine-tune a key checkpoint with AdamW and report Matbench/3BPA deltas, or clearly qualify that the SOTA model may not reflect the authors’ preferred extrapolation recipe and discuss the risk.
  4. Tables 1–3 and the M-fold derivation establish ECE/RRA on small molecular tasks, but the large-scale claim (Abstract, §3.12) attributes Matbench SOTA to “these advances” without component ablations at OAM scale (ECE on/off, RRA vs EquiformerV3 attention, uu vs uv under matched params). At minimum, a smaller matched-budget materials ablation or intermediate TACE-OAM-L-style comparison that isolates ECE/RRA would make the contribution narrative falsifiable rather than confounded with data, width, DeNS, and optimizer.
minor comments (7)
  1. Abstract and §1: “propose direct Cartesian construction and recursive Clebsch-Gordan construction” — spacing/typos (“proposedirect”, “Complex At- tention”) should be cleaned throughout the arXiv text.
  2. §3.7: sentence fragment “two key factors that are critical to the a (Joshi et al., 2023) identifies” needs repair.
  3. Eq. (35): θ_{m,h}(r) is written with an m index but the surrounding text often uses θ_h; clarify whether phase is m-dependent and how it is shared across heads/orders.
  4. Table 5: TECE OOMs earlier than EquFlashV2; the discussion correctly notes fusion advantages of SO(3), but a brief note on peak memory vs parameter count (222M vs 44.9M) would help readers interpret speed fairly.
  5. §3.5 / Table 2: “w1 w1 w1 w1 w2” column headers are hard to parse; a compact legend (w1 | w1=w2 | w1+iw2) would improve readability.
  6. §4.1 diatomic discussion is valuable; consider moving a short quantitative diatomic table into the main text or appendix so the “poor diatomic / high Matbench” trade-off is documented rather than only narrated.
  7. References and arXiv dates in the bibliography include 2026 entries; ensure consistency of citation keys and that concurrent work (EquiformerV3, EquFlashV2, DPA4) is fairly scoped in related work.

Circularity Check

0 steps flagged

No significant circularity: SOTA and ablation claims are external-benchmark evaluations of proposed operators, not tautologies of fitted inputs or self-citation uniqueness.

full rationale

The paper proposes architectural operators (Cartesian/recursive Wigner-D construction, Edge Cluster Expansion via generalized asymmetric contraction, Radial Rotary Complex Attention, and ACE group-linear/GLU-style improvements) and evaluates them on held-out molecular sets (3BPA temperatures/dihedrals, AcAc), geometric completeness tests (M-fold SO(2) structures; Pozdnyakov k-body counterexamples), and the external Matbench Discovery leaderboard after training on OMat24/sAlex/MPTrj. None of these results is obtained by fitting a parameter to a quantity and then reporting that same quantity as a prediction, nor by defining an operator in terms of the metric it is said to derive. Self-citations to prior TACE Cartesian work supply background ACE/Cartesian operators and residual/norm design context; they do not import a uniqueness theorem that forces the Matbench ranking or the ECE/RRA claims. The 3BPA-1200K 65 meV/Å module gate and the Muon-vs-Adam optimizer choice are design/selection decisions that may affect causal attribution of SOTA gains, but they are not circular reductions of outputs to inputs. Score 0 with empty steps is therefore the correct finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 4 invented entities

The work sits on standard SO(3)/SO(2) representation theory and public materials datasets; free parameters are the usual MLIP hyperparameters plus an ad hoc 65 meV/Å exclusion rule. Invented entities are architectural modules, not new physical particles. No claim depends on an unmeasured physical constant invented here.

free parameters (5)
  • 3BPA-1200K force RMSE exclusion threshold = 65 meV/Å
    Section 3.6 excludes any module with force RMSE > 65 meV/Å on 3BPA 1200K; this hand-chosen gate selects the final architecture.
  • Interaction / product channel widths and Lmax/mmax = 64 / 256 / 4 / 2
    Table 6: interaction channel 64, product channel 256, Lmax=mmax=4, correlation order 2—capacity knobs fitted for Matbench performance.
  • RRA temperature bounds and radial phase form = τ∈[0.25,4.0]
    τ_min=0.25, τ_max=4.0 and optional mπ·tanh radial phase (Eqs. 35–37) are design choices controlling attention sharpness.
  • Muon learning-rate schedule and DeNS/stochastic-depth rates = see Table 6
    Appendix training recipe (max LR 5e-3→1e-4, DeNS on direct stage, stochastic depth 0.05) is tuned for convergence, not derived.
  • Cutoff radius = 6 Å
    6 Å neighborhood cutoff for all TECE-OAM-RRA stages (Table 6).
axioms (5)
  • standard math SO(2) complex multiplication with m3=m1±m2 is equivariant under planar rotations (Eqs. 1–4).
    Standard representation theory of the circle group; used throughout SO(2) Linear/TP and ECE.
  • standard math Wigner-D matrices implement SO(3) action on spherical tensors and can be obtained from Cartesian rotations via ICTD or recursive CG coupling (Eqs. 10, 13).
    Classical angular-momentum theory; paper contributes constructions, not the existence claim.
  • domain assumption Local edge frames from bond directions make global SO(3) features into local SO(2) features without breaking equivariance when D-matrices are applied correctly.
    Core eSCN-style modeling assumption (Section 2.1); discontinuities if frames jump are acknowledged but not fully eliminated.
  • domain assumption OMat24, sAlex, and MPTrj labels are sufficiently consistent DFT targets for ranking universal MLIPs on Matbench Discovery.
    Section 3.12 and Discussion note Hubbard-U discontinuities and rattle artifacts, yet SOTA is still claimed on this benchmark stack.
  • ad hoc to paper Real-valued SO(2) weights (w1) are preferred because complex weights break O(2) reflection equivariance unless w2=0 (Eq. 14).
    Derived for O(2), then elevated to a design rule; Table 2 supports it on 3BPA but large-data regimes may differ as the authors note.
invented entities (4)
  • Edge Cluster Expansion (ECE) / Edge Complex Product Basis via Generalized Asymmetric Contraction no independent evidence
    purpose: Build higher-order many-body interactions directly on edges in the local SO(2) frame to raise angular resolution and body order.
    New architectural block relative to node ACE; independent evidence is only the paper’s ablations and Matbench model, not an external physical observable.
  • Radial Rotary Complex Attention (RRA) no independent evidence
    purpose: Define attention logits from Re(e^{iθ} Q K̄) with radial phase/bias so attention depends on distance and feature similarity.
    Novel attention variant combining SO(2) complex products with RoPE-like radial phase; validated only inside this paper’s tables.
  • uuSO2Linear vs uvSO2Linear distinction no independent evidence
    purpose: Separate path-wise radial-weighted SO(2) linear maps from fully connected channel-mixing SO(2) linear maps.
    Naming/analysis contribution; Table 1 is the only external handle.
  • Recursive Clebsch-Gordan Wigner-D construction (TACE-Recursive) independent evidence
    purpose: Compute D^{(ℓ)} from rotation matrices with polynomial intermediates instead of 3^ℓ Cartesian tensors or Euler angles.
    Implementation method with runtime evidence in Figure 1; not a new physical entity.

pith-pipeline@v1.1.0-grok45 · 29810 in / 4085 out tokens · 41653 ms · 2026-07-14T10:07:05.741902+00:00 · methodology

0 comments
read the original abstract

In this paper, we provide a systematic investigation of SO(2) theory to machine learning interatomic potentials (MLIPs) and identify the limitations of conventional SO(2) Linear architectures relative to SO(3) Clebsch-Gordan Tensor Products (CGTP). Building on these insights, we propose direct Cartesian construction and recursive Clebsch-Gordan construction of Wigner D-matrices and introduce two novel interaction building blocks. First, we propose the Edge Complex Product Basis based on Generalized Asymmetric Contraction, a new formulation for many-body expansion that directly constructs higher-order interactions on edges through complex-valued equivariant multiplications. Second, we introduce Radial Rotary Complex Attention(RRA), which enhances extrapolation performance and surpasses existing attention vector formulations. We also introduce several improvements to the Atomic Cluster Expansion module. Building on these advances, we train our models on OMat24, sAlex, and MPTrj, and introduce TECE-OAM-RRA-1.0, which achieve state-of-the-art (SOTA) performance on the Matbench Discovery.

Figures

Figures reproduced from arXiv: 2607.10664 by P.Hu, Wenbo Xie, Zemin Xu.

Figure 1
Figure 1. Figure 1: Runtime comparison of different Wigner-D matrix con￾struction methods from rotation matrices or quaternion. The bench￾mark is performed on an NVIDIA RTX 4090 GPU using 10000 edges, torch.float32 precision, and averages the runtime over re￾peated forward evaluations. The recursive Clebsch-Gordan con￾struction shows better scaling with increasing ℓmax (construct from 0 to ℓmax) compared with other methods. T… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 16 linked inside Pith

  1. [1]

    M., Dzamba, M., Gao, M., Rizvi, A., Uyttendaele, M., Zit- nick, C

    Barros-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Uyttendaele, M., Zit- nick, C. L., and Ulissi, Z. W. The open materials 2024 (omat24) inorganic materials dataset and models.Nature Computational Science, pp. 1–11,

  2. [2]

    Dauphin, Y

    URL https://arxiv.org/abs/ 2601.16195. Dauphin, Y . N., Fan, A., Auli, M., and Grangier, D. Lan- guage modeling with gated convolutional networks. In Precup, D. and Teh, Y . W. (eds.),Proceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Re- search, pp. 933–941. PMLR, 06–11 Aug

  3. [3]

    Harari, G., Zimmermann, Y ., Kulseng, O

    URL https: //arxiv.org/abs/2607.03433. Harari, G., Zimmermann, Y ., Kulseng, O. T., Zichi, L., Tan, C. W., Descoteaux, M. L., and Kozinsky, B. Beyond adam: Soap and muon for faster, label-efficient training of machine learning interatomic potentials,

  4. [4]

    He, K., Zhang, X., Ren, S., and Sun, J

    URL https://arxiv.org/abs/2607.02499. He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learn- ing for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778,

  5. [5]

    Huang, G., Sun, Y ., Liu, Z., Sedra, D., and Weinberger, K

    URL https://arxiv.org/abs/2603.08630. Huang, G., Sun, Y ., Liu, Z., Sedra, D., and Weinberger, K. Deep networks with stochastic depth,

  6. [6]

    Huang, L., Huang, C., Wang, Z., Du, Y ., Wang, C., Lu, H., Li, Y ., Liu, X., Jiang, A., and Zhang, J

    URL https://arxiv.org/abs/1603.09382. Huang, L., Huang, C., Wang, Z., Du, Y ., Wang, C., Lu, H., Li, Y ., Liu, X., Jiang, A., and Zhang, J. E2former- v2: On-the-fly equivariant attention with linear activation memory.arXiv preprint arXiv:2601.16622,

  7. [7]

    Joshi, C

    URL https://arxi v.org/abs/2401.04088. Joshi, C. K., Bodnar, C., Mathis, S. V ., Cohen, T., and Lio, P. On the expressive power of geometric graph neural networks

  8. [8]

    Kondor, R., Lin, Z., and Trivedi, S

    URL https://arxiv.org/ab s/1412.6980. Kondor, R., Lin, Z., and Trivedi, S. Clebsch–gordan nets: a fully fourier space spherical convolutional neural network. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.),Advances in Neural Information Processing Systems, volume

  9. [9]

    Li, T., Li, W., Peng, A., Xue, J., Zhang, L., Zhang, D., and Wang, H

    URL https://openreview.net /forum?id=wiQe95BPaB. Li, T., Li, W., Peng, A., Xue, J., Zhang, L., Zhang, D., and Wang, H. Dpa4: Pushing the accuracy-cost frontier of interatomic potentials with emfa so(2) convolution, 2026a. URLhttps://arxiv.org/abs/2606.02419. Li, Y ., Huang, L., Ding, Z., Wei, X., Wang, C., Yang, H., Wang, Z., Liu, C., Shi, Y ., Jin, P., Q...

  10. [10]

    Liao, Y .-L., Hoffman, A

    URL https: //arxiv.org/abs/2403.09549. Liao, Y .-L., Hoffman, A. J., Shen, S. C., Duval, A., Nor- wood, S. W., and Smidt, T. Equiformerv3: Scaling ef- ficient, expressive, and general se (3)-equivariant graph attention transformers.arXiv preprint arXiv:2604.09130,

  11. [11]

    Loshchilov, I

    URL https://arxiv.org/abs/ 2502.16982. Loshchilov, I. and Hutter, F. Decoupled weight decay regu- larization,

  12. [12]

    Luo, S., Chen, T., and Krishnapriyan, A

    URL https://arxiv.org/abs/ 1711.05101. Luo, S., Chen, T., and Krishnapriyan, A. S. Enabling effi- cient equivariant operations in the fourier basis via gaunt tensor products. InThe Twelfth International Confer- ence on Learning Representations,

  13. [13]

    Orb-v3: atomistic simulation at scale.arXiv preprint arXiv:2504.06231,

    Rhodes, B., Vandenhaute, S., ˇSimkus, V ., Gin, J., Godwin, J., Duignan, T., and Neumann, M. Orb-v3: atomistic simulation at scale.arXiv preprint arXiv:2504.06231,

  14. [14]

    Team, K., Chen, G., Zhang, Y ., Su, J., Xu, W., Pan, S., Wang, Y ., Wang, Y ., Chen, G., Yin, B., et al

    URL https://arxiv.org/ab s/2104.09864. Team, K., Chen, G., Zhang, Y ., Su, J., Xu, W., Pan, S., Wang, Y ., Wang, Y ., Chen, G., Yin, B., et al. Attention residuals.arXiv preprint arXiv:2603.15031,

  15. [15]

    Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds.arXiv preprint arXiv:1802.08219,

    Thomas, N., Smidt, T., Kearnes, S., Yang, L., Li, L., Kohlhoff, K., and Riley, P. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds.arXiv preprint arXiv:1802.08219,

  16. [16]

    Unke, O. T. and Maennel, H. E3x: E(3)-equivariant deep learning made easy.arXiv preprint arXiv:2401.07595,

  17. [17]

    Warford, T., Thiemann, F

    URLhttps://arxiv.org/abs/2409.11321. Warford, T., Thiemann, F. L., and Cs´anyi, G. Better without u: Impact of selective hubbard u correction on founda- tional mlips,

  18. [18]

    15 Weiler, M., Geiger, M., Welling, M., Boomsma, W., and Cohen, T

    URL https://arxiv.org/ab s/2601.21056. 15 Weiler, M., Geiger, M., Welling, M., Boomsma, W., and Cohen, T. S. 3d steerable cnns: Learning rotationally equivariant features in volumetric data. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.),Advances in Neural Informa- tion Processing Systems, volume

  19. [19]

    Xie, Y ., Daigavane, A., Kotak, M., and Smidt, T

    URL https://arxiv.org/abs/ 2506.23971. Xie, Y ., Daigavane, A., Kotak, M., and Smidt, T. Asymp- totically fast clebsch-gordan tensor products with vector spherical harmonics,

  20. [20]

    org/abs/2602.21466

    URL https://arxiv. org/abs/2602.21466. Xu, Z., Wu, C., Xie, W., and Hu, P. A cartesian-3j frame- work for machine learning interatomic potentials, 2026a. URLhttps://arxiv.org/abs/2512.16882. Xu, Z., Xie, W., and Hu, P. Spectral/spatial tensor atomic cluster expansion with universal embeddings in cartesian space.arXiv preprint arXiv:2509.14961, 2026b. Yu, ...

  21. [21]

    Training Details All TECE models were trained in torch.float32 precision using 32 NVIDIA H20 GPUs

    16 A. Training Details All TECE models were trained in torch.float32 precision using 32 NVIDIA H20 GPUs. During pretraining, we employed DeNS (Liao et al., 2024)and stochastic depth (Huang et al.,

  22. [22]

    DeNS was mainly used with the direct model to accelerate convergence, rather than to improve the final accuracy

    as additional regularization techniques. DeNS was mainly used with the direct model to accelerate convergence, rather than to improve the final accuracy. With appropriately chosen hyperparameters, the direct and conservative models achieved comparable performance upon convergence. The DeNS hyperparameters were kept consistent with those used in Equiformer...