{"id":"eb869db9-f282-4001-bf5c-6dcd499de3a8","arxiv_id":"2601.16195","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Unconstrained non-equivariant and direct-force neural interatomic potentials scale to 730M parameters and match or beat equivariant state-of-the-art models on several atomistic benchmarks.","lead":"This paper shows that machine-learned interatomic potentials that skip built-in physical symmetries can match or beat symmetry-preserving models in accuracy and speed when trained on large datasets. It releases large pretrained models and practical rules for fixing symmetry-related errors in geometry optimization and lattice dynamics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'superior in accuracy and speed' headline rests on an unmatched SPICE comparison: PET-L got 3x the eSEN training epochs and eSEN timings were taken from another paper, so the superiority claim is not yet established.","rationale":"The reader's conditional verdict is appropriate. The paper has real strengths: public pretrained models, multiple benchmarks, transparent appendices, and honest caveats about non-conservative forces in phonon workflows. The central 'on par with equivariant models at scale' claim is plausibly supported by the matbench, LAMbench, and MADBench results. However, the abstract's stronger phrase 'superior in accuracy and speed' goes beyond what the evidence shown supports, and the most direct evidence for it — the SPICE comparison — is not head-to-head. I do not think this warrants rejection; it warrants the conditional framing the reader already chose, with the claim sharpened to 'on par' and the speed comparison backed by matched wall-clock measurements. I therefore leave the verdict unchanged. The reader's identified weakest assumption (the ~20,000-orientation learnability estimate) is a reasonable theoretical worry, but it is backed by direct raw-vs-averaged benchmark comparisons and the empirical scaling results, making it less load-bearing than the unmatched SPICE protocol. That is why my agreement is partial rather than full.","tokens_in":24513,"tokens_out":10004,"duration_ms":103664,"concrete_test":"Run a matched SPICE comparison: (1) train PET-L from scratch on the same SPICE split for 100 epochs (the eSEN budget) with identical loss/batch settings as Table II, and report per-subset energy/force MAEs over at least 3 seeds; (2) re-measure eSEN-L inference wall-clock on the same A100 with the same evaluation code path used for PET and MACE, and record total training wall-clock for both models. If PET-L at 100 epochs no longer holds the lowest MAE on a majority of SPICE subsets, the 'superior accuracy' claim is not established; if it still leads, the concern is resolved. This also quantifies whether the 3x epoch handicap is compensated by inference speed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires a fair comparison against equivariant models. The strongest direct evidence is the SPICE benchmark (Sec. IV B, Table II, Fig. 3), but that comparison is not matched. PET-L (190M params) was trained for 300 epochs, three times the number reported for eSEN (Sec. IV B a). The inference timings for eSEN in Fig. 3 are taken from Ref. [31] rather than re-measured on the same hardware/software path used for PET and MACE, and Table II reports a single run without error bars or seed variation. The paper acknowledges eSEN training times are not available and speculates that the two models required 'around the same amount of compute to train' based on inference timings; that is not a substitute for a wall-clock measurement. If PET-L needs 3x epochs to reach those accuracies, the claim 'unconstrained models can be superior in accuracy and speed' is an artifact of unequal training budget. The rotational-learnability heuristic in Sec. III A c is not the load-bearing point: the benchmark results stand independently of that estimate, and Table V shows raw-vs-SO(3)-averaged differences are modest, so the empirical comparison is where the headline claim lives or dies.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether machine-learned interatomic potentials that do not enforce rotational equivariance or energy conservation can be scaled to large, diverse datasets and still match or beat physically constrained, equivariant models. The authors introduce a scaled-up version of the PET architecture and train models on MPtrj, subsampled Alexandria, OMat24, SPICE, and OMol-1. They benchmark these models on matbench-discovery, LAMBench, MADBench, and SPICE, and they test geometry optimization and phonon workflows. The central claim is that fully unconstrained models can be trained at scale and reach accuracy on par with state-of-the-art equivariant architectures, while offering competitive inference speed; they also argue that rotational symmetry breaking can be corrected at inference time by SO(3) averaging. The manuscript is candid about cases where non-conservative, direct-force models underperform, particularly for phonons and elastic constants.","tokens_in":24813,"tokens_out":4735,"duration_ms":49926,"significance":"If the central claim is established, this is a significant result for the MLIP community. It suggests that expensive equivariant constraints may be unnecessary at scale, enabling simpler, faster architectures without sacrificing accuracy. The paper also provides open-source models and reproducible workflows, which is a strength. The material-benchmark results and the honest reporting of non-conservative failure modes are valuable. The main uncertainty is whether the headline 'superior in accuracy and speed' holds under matched training budgets and controlled timing comparisons; the current evidence supports 'on par under favorable evaluation choices' more than it supports a general superiority claim.","major_comments":[{"comment":"The claim that unconstrained models are 'superior in accuracy and speed' rests heavily on the SPICE comparison, but the comparison is not matched. PET-L was trained for 300 epochs, three times the number reported for eSEN, and the eSEN accuracies are taken from Ref. [31] rather than re-evaluated. Table II reports a single run with no seed variation or error bars. The statement that the two models required 'around the same amount of compute to train' is inferred from inference timings, not measured wall-clock training time. This is load-bearing for the abstract's superiority claim. Please either run eSEN with the same training budget under controlled conditions, report PET-L at the eSEN epoch budget, or explicitly restrict the claim to 'can match' with the unequal budget stated as a caveat.","section":"Sec. IV B a, Table II, Fig. 3"},{"comment":"The accuracy-speed Pareto front is not a controlled measurement. Appendix F states that eSEN timings are taken from Ref. [31], while PET and MACE timings are evaluated by the authors, with PET additionally using torchscript compilation. The manuscript argues that differences are minimal, but a Pareto plot that mixes measurement setups does not establish a speed advantage. Please provide timings for eSEN, MACE, and PET on the same hardware and software stack, or remove eSEN from the timing comparison and restrict the speed claim accordingly.","section":"Sec. IV B b, App. F, Fig. 3"},{"comment":"The matbench-discovery headline comparison uses evaluation choices that favor the proposed models. All results in Table I are reported after SO(3) rotational averaging, while baselines from the leaderboard may not use this post-processing. In addition, Table VI and Appendix H show that central finite differences and displacement size materially affect the phonon metric κSRME for non-equivalent models (e.g., 0.157 at 0.01 Å versus 0.110 at 0.05 Å). If leaderboard baselines were evaluated with forward differences and no averaging, Table I is not an apples-to-apples comparison. Please report raw and averaged metrics side by side, specify the evaluation protocol for every baseline, and avoid ranking conclusions that depend on protocol choices.","section":"Sec. IV A b, Table I, Table VI, App. H"},{"comment":"The phrase 'unconstrained models can be superior in accuracy and speed' is broader than the paper's own evidence. The non-conservative direct-force heads consistently underperform conservative models on static workflows: κSRME is 0.197–0.216 for non-conservative versus 0.119 for conservative PET-OAM, and elastic moduli are markedly worse (Table IX). The paper's nuanced discussion correctly recommends non-conservative models primarily as pretraining or multiple-time-stepping components. The abstract and introduction should be aligned with that nuance, so that 'unconstrained' clearly refers to rotationally unconstrained conservative models when accuracy on static tasks is claimed.","section":"Sec. V, Table V, Table VI"}],"minor_comments":[{"comment":"The displayed locality equation appears garbled: the same term V({r_i,a_i}_{i≠n}) is subtracted twice, and the text's mathematical rendering is hard to follow. Please correct the formula so that the near-sightedness condition is stated cleanly.","section":"Sec. II A c"},{"comment":"The model listed as 'PET' in Table I is not fully specified in the table caption. Please state which checkpoint (PET-OAM conservative, with or without SO(3) averaging) is used, and whether the other entries are reproduced from the leaderboard or re-evaluated with the same protocol.","section":"Table I"},{"comment":"The learnability estimate of ~20,000 orientation samples assumes a 0.1 Å resolution and a 4 Å environment radius. This is a heuristic with no sensitivity analysis. It is not the main empirical claim, but the manuscript should label it as an order-of-magnitude argument rather than a quantitative bound.","section":"Sec. III A c"},{"comment":"The text says the stable high-temperature phase of Ti is BCC but that it is stabilized by entropic effects and static calculations find HCP preferred. This is physically coherent, but the wording is easy to misread. Please rephrase to distinguish static enthalpy from free energy.","section":"Sec. IV A a, Fig. 2"},{"comment":"The footnote about different GPU power settings and torchscript compilation is appropriate, but the main-text sentence 'the two models required around the same amount of compute' should be flagged as an inference, not a measurement, in the main text as well.","section":"App. F"}],"recommendation":"major_revision","confidential_remarks":"This is a strong and honest empirical study, and the open release of models and workflows is commendable. The main risk is the headline claim of superiority, which currently rests on an unmatched SPICE training budget and mixed timing sources. The material benchmarks are extensive and support 'on par' or 'competitive', but protocol differences (SO(3) averaging, central finite differences) should be disclosed more prominently. I believe the paper can be made acceptable by either providing matched comparisons or scaling back the abstract's claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about arXiv:2601.16195. First, the core empirical result is solid: a 730M-parameter rotationally unconstrained PET model, trained on MPtrj+Alexandria+OMat24, matches or beats equivariant state-of-the-art on matbench-discovery and holds its own on LAMbench and MADBench. Weights are released. That is a genuine scaling result, and the architectural tweaks (4x node features, skip connections, adaptive cutoff) are sensible and well documented. Second, the abstract's \"superior in accuracy and speed\" is stronger than the evidence. The accuracy superiority comes almost entirely from the SPICE comparison, and that comparison is not matched: PET-L got 300 epochs vs eSEN's 100, and the eSEN inference timings in the Pareto plot are borrowed from Ref [31] rather than measured on the same hardware path. The paper acknowledges the epoch gap and speculates the training compute is similar based on inference speed, but that is not a wall-clock measurement. So the fair conclusion right now is \"on par in accuracy, faster at inference, more expensive to train\"—still a useful result, but not the headline claim.\n\nThe paper does several things well. The matbench-discovery table is credible and includes multiple models plus the raw/SO(3)-averaged comparison. The phonon and geometry-optimization section honestly shows where unconstrained models break: symmetry detection tolerances need loosening, non-conservative forces give iffy phonons, and the finite-difference displacement matters. The adaptive cutoff and the OMol conditioning are useful contributions. Data and code are available.\n\nSoft spots, in proportion: (1) The unmatched SPICE comparison is the load-bearing piece for \"superior in accuracy\"; it needs a matched re-run or a sharpened claim. (2) The κSRME phonon metric involves a finite-difference displacement that appears to have been selected after inspecting the benchmark results; the effect is small (0.119 vs 0.110) and the paper reports the sweep, so this is minor, but the headline number should state the displacement. (3) The Sec III A c rotational-learnability estimate (20,000 orientation samples) is hand-wavy—fine as motivation, not as a bound, and the empirical results don't depend on it. Your weakest-assumption point is right: that heuristic is not load-bearing.\n\nOverall: the paper is a serious empirical study, honestly reported, with reproducible artifacts. The claims just need to be sharpened to match the actual evidence. I'd send it to peer review, and I'd tell the authors to fix the abstract and either match the SPICE training budget or present the comparison as accuracy-parity with faster inference.","headline":"Useful scaling study with a real result—unconstrained 730M-parameter PET matches equivariant SOTA on matbench—but the abstract's 'superior in accuracy and speed' rests on a SPICE comparison that is not matched (3x epochs, borrowed timings), so the strong form of the claim is not yet established.","tokens_in":25332,"tokens_out":2665,"would_cite":true,"duration_ms":25927,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper sets out to show that machine-learned interatomic potentials can be scaled to large datasets without hard-coding rotational symmetries or energy conservation, and that such unconstrained models match or beat equivariant, symmetry","keywords":["machine-learned interatomic potentials","unconstrained architectures","rotational invariance","equivariant neural networks","non-conservative forces","materials discovery benchmarks","geometry optimization","lattice dynamics"],"falsifier":"Train the same architecture on the same large datasets but limit rotational augmentation to a single axis, then evaluate on test structures rotated around an orthogonal axis. A sharp force-error degradation would confirm the paper's estimate that roughly 20,000 dense orientation samples are needed to learn rotational invariance; little degradation would falsify it, implying the model relies on other inductive biases. A practical counterpart: on a perfectly symmetric BCC titanium cell from a trained conservative model, measure the residual forces that should vanish under the space group without","tokens_in":24369,"feed_emoji":"⚛️","tokens_out":10913,"duration_ms":162797,"temperature":0.7,"pith_summary":"The paper's central claim is that fully unconstrained machine-learned interatomic potentials — architectures that do not hard-code rotational symmetry and may output forces directly instead of deriving them from an energy — can be trained on very large, diverse datasets and reach the accuracy of state-of-the-art equivariant models, with faster inference. The demonstration is built around a transformer-based graph neural network scaled to hundreds of millions of parameters, trained on large materials and molecular datasets, and tested on discovery benchmarks as well as geometry optimization and phonon calculations. If the claim holds, hard-coding physical symmetries is not a prerequisite for accurate large-scale atomistic simulation; simple inference-time modifications — averaging over rotations, looser symmetry tolerances, symmetrizing force constants — can restore symmetry-consistent observables. The paper also finds that direct-force non-conservative versions are faster but less trustworthy for static property calculations, so they are best used to accelerate or pre-train conservative models.","feed_headline":"Symmetry-free ML force fields match equivariant rivals at scale","feed_subtitle":"Trained on large datasets, rotation-blind models run faster and match benchmark accuracy.","key_machinery":"The central object is PET, a graph neural network that processes each atom's neighborhood as a set of edge tokens through a transformer without equivariant features. For scale, the paper quadruples node features relative to edge features, carries node features across layers, and uses modern transformer components (RMSNorm, SwiGLU, pre-normalization) and a smooth adaptive cutoff, reaching 730M parameters at low inference cost. Rotational invariance is learned from data augmentation during training; at inference it can be restored approximately by averaging over a Lebedev grid of rotations, projecting forces onto a known space group, or loosening symmetry-detection tolerances. A second head em","core_discovery":"Fully unconstrained architectures — no hard-coded rotation symmetry, and optionally direct-force outputs — can be scaled to large, diverse datasets and match or exceed state-of-the-art equivariant neural networks in accuracy while running faster at inference. The demonstration uses a transformer-based graph neural network scaled to 730M parameters on materials data and 190M parameters on molecular data, reaching competitive results on discovery and molecular benchmarks, with benchmark numbers reported after rotational symmetrization. Conservative unconstrained models are safe for geometry optimization and phonons when symmetry detection is loosened or predictions are rotationally averaged. N","pith_inferences":["If rotational invariance is genuinely learned from orientation diversity in data, the accuracy gap between equivariant and unconstrained architectures may close further as datasets grow, making dataset design — not architecture — the dominant factor; the paper leaves this scaling prediction implicit.","A direct test of the paper's learnability estimate would be to train the same architecture with rotational augmentation restricted to a single axis and evaluate on rotations around an orthogonal axis; a sharp degradation would confirm the ~20,000-orientation estimate, while little degradation would point to other inductive biases.","The hybrid workflow suggested by the results — non-conservative head for pretraining and most force evaluations, conservative head for static and vibrational properties — is a natural next benchmark to quantify energy drift and computational savings on realistic molecular dynamics.","The symmetry-breaking that lets an unstable BCC titanium cell relax toward close-packed structures could be turned into a feature for automated structure search, where exactly equivariant models sometimes get trapped in high-symmetry metastable states."],"forward_implications":["Unconstrained models are viable at scale: a 730M-parameter materials model and a 190M-parameter molecular model reach state-of-the-art benchmark accuracy, in several cases exceeding equivariant baselines.","Inference is faster: avoiding equivariant operations and using direct-force heads gives a 2–3x speedup, which matters for long molecular dynamics runs; the extra training epochs needed from scratch are offset by cheaper per-epoch evaluation.","Fine-tuning transfers well: a model pre-trained on a large open materials dataset can be fine-tuned to a much smaller dataset in fewer than 1/20th of the epochs needed from scratch, roughly halving validation errors.","Static workflows can be made reliable: geometry optimization works with unconstrained models, and symmetry breaking can even help an unstable high-symmetry cell escape to a lower-energy structure; phonon calculations are safe with the conservative model when the Hessian is symmetrized and central finite differences are used.","Non-conservative forces are not a free lunch: despite comparable test-set accuracy, direct-force models systematically trail on downstream tasks such as phonon-derived properties and elastic constants, so the paper positions them as pretraining engines or partners in multiple-time-stepping dynamics."],"fun_headline_variants":["Unconstrained ML force fields beat equivariant rivals when trained big","Symmetry-free MLIPs outperform equivariant models on large datasets","Scaling unconstrained MLIPs to 730M params beats constrained accuracy","Dropping symmetry constraints in MLIPs proves faster and more accurate at scale","Large-data MLIPs without hard-coded symmetry match or beat equivariant nets"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is the informal estimate that a rotationally unconstrained model can learn rotational invariance from roughly 20,000 orientation samples on a sphere of radius about 4 Å at 0.1 Å resolution, and that a low-order Lebedev grid at inference suppresses symmetry breaking enough for geometry optimization and phonons.","fun_headline_variants_meta":{"raw":{"variants":["Unconstrained ML force fields beat equivariant rivals when trained big","Symmetry-free MLIPs outperform equivariant models on large datasets","Scaling unconstrained MLIPs to 730M params beats constrained accuracy","Dropping symmetry constraints in MLIPs proves faster and more accurate at scale","Large-data MLIPs without hard-coded symmetry match or beat equivariant nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1391,"prompt_tokens":705,"completion_tokens":686,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":602}},"tokens_in":449,"tokens_out":686,"duration_ms":9155,"temperature":1.0,"reasoning_tokens":602,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:39:13.389271+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same architecture on the same large datasets but limit rotational augmentation to a single axis, then evaluate on test structures rotated around an orthogonal axis. A sharp force-error degradation would confirm the paper's estimate that roughly 20,000 dense orientation samples are needed to learn rotational invariance; little degradation would falsify it, implying the model relies on other inductive biases. A practical counterpart: on a perfectly symmetric BCC titanium cell from a trained conservative model, measure the residual forces that should vanish under the space group without","supporting_citations":[],"review_version":1}