Pith. sign in

REVIEW 2 cited by

Transferability of datasets between Machine-Learning Interaction Potentials

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05590 v1 pith:EIPNBGOH submitted 2024-09-09 physics.chem-ph

classification physics.chem-ph
keywords trainingdatamodelsdatasetsdifferentmodelalgorithmarchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the emergence of Foundational Machine Learning Interatomic Potential (FMLIP) models trained on extensive datasets, transferring data between different ML architectures has become increasingly important. In this work, we examine the extent to which training data optimised for one machine-learning forcefield algorithm may be re-used to train different models, aiming to accelerate FMLIP fine-tuning and to reduce the need for costly iterative training. As a test case, we train models of an organic liquid mixture that is commonly used as a solvent in rechargeable battery electrolytes, making it an important target for reactive MLIP development. We assess model performance by analysing the properties of molecular dynamics trajectories, showing that this is a more stringent test than comparing prediction errors for fixed datasets. We consider several types of training data, and several popular MLIPs - notably the recent MACE architecture, a message-passing neural network designed for high efficiency and smoothness. We demonstrate that simple training sets constructed without any ab initio dynamics are sufficient to produce stable models of molecular liquids. For simple neural-network architectures, further iterative training is required to capture thermodynamic and kinetic properties correctly, but MACE performs well with extremely limited datsets. We find that configurations designed by human intuition to correct systematic model deficiencies transfer effectively between algorithms, but active-learned data that are generated by one MLIP do not typically benefit a different algorithm. Finally, we show that any training data which improve model performance also improve its ability to generalise to similar unseen molecules. This suggests that trajectory failure modes are connected with chemical structure rather than being entirely system-specific.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-Tuning Universal Machine-Learned Interatomic Potentials: A Tutorial on Methods and Applications

    physics.comp-ph 2025-06 conditional novelty 4.0 of 10

    Fine-tuning universal MLIPs improves accuracy and data efficiency across electrolytes, defects, and interfaces, with some evidence of implicit long-range behavior that is not conclusive.

  2. A Study on the Fine-Tuning Performance of Universal Machine-Learned Interatomic Potentials (U-MLIPs)

    physics.comp-ph 2025-06 conditional novelty 4.0 of 10

    Fine-tuning universal MACE potentials on targeted datasets generally improves accuracy and convergence speed, though data selection, not the foundation model alone, determines success.

Pith tools