REVIEW 5 major objections 6 minor 11 references
Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials
T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A fully automated cloud-parallel crystal structure prediction protocol driven by the Lavo-NN neural network potential generates and correctly ranks all known $Z'=1$ polymorphs of a 49-molecule pharmaceutical benchmark at an average cost…
desk verdict A serious and useful CSP automation paper whose headline generalization claim needs a holdout check before it can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Lavo-NN is an equivariant message-passing neural network whose crystal energy is decomposed through the many-body expansion truncated at second order: an intramolecular term plus a sum of dimer interaction energies. Message passing runs only along covalent intramolecular edges, and the intermolecular readout augments a short-range neural term with a long-range polarizable force field, with some force-field parameters predicted by the network; this decomposition is what lets the model reach near-DFT accuracy at low cost. Around the potential sits a protocol that generates dense crystal candidates geometrically, refines them with a semi-local Monte Carlo search using interpolated moves, tracks convergence per space group with a Poisson-lognormal species-abundance model, and optionally re-ranks the lowest-energy structures with periodic PBE-D3(BJ) plus a monomer PBE0 correction.
What would settle it
Search the 10k-molecule training set and the extracted conformers and dimers for any of the 49 benchmark molecules or their close tautomers, then retrain Lavo-NN with those molecules strictly excluded and rerun the benchmark to see whether all 110 polymorphs are still generated and ranked near the bottom of the landscape.
Extended reading notes
Core claim
The central claim is that a single purpose-built neural network potential, trained once on a large corpus of pharmaceutical-like crystal structures and dimer interactions, can both enumerate and rank the experimentally observable polymorphs of new drug molecules without system-specific re-training or manual specification. Specifically, the paper reports generating structures that match all 110 known $Z'=1$ experimental polymorphs of its 49-molecule benchmark, ranking 87% of the matched structures within the top 50 of their predicted landscapes, and doing so at an average cost of about 8.4k CPU hours per molecule. The authors argue, through case studies and a semi-blinded exercise using only powder diffraction patterns, that the predicted landscapes can resolve ambiguities in experimental data and identify the structures of forms that lack single-crystal data.
Load-bearing premise
The benchmark molecules must be absent from Lavo-NN's training data; the paper trains on structures built from 10k SMILES drawn from a public molecular dataset and never states that the 49 test molecules were held out, so the headline result could be retrieval from training data rather than generalization.
Editorial extensions
If this is right
- CSP can be run routinely on real drug candidates: a typical molecule costs about 8.4k CPU hours and finishes in days of wall time, making solid-form screens practical during lead optimization.
- The protocol can flag thermodynamically metastable marketed forms before launch; the rotigotine case shows a form that later caused a recall being correctly ranked above the stable form.
- When only powder diffraction data exist, the protocol can propose full three-dimensional crystal structures and rank their stability, as demonstrated for three recent drugs.
- Lavo-NN's speed makes it suitable for the generation phase of CSP, where hundreds of millions of energy evaluations are needed, while its accuracy (54% Top-10 on the benchmark) means optional DFT re-ranking can be limited to single-point calculations.
- As ranking costs drop, the bottleneck shifts to structure generation, so future gains for multi-component systems ($Z'>1$, hydrates, salts, co-crystals) will need new generative methods.
Reading between the lines
- A prospective blind test on molecules guaranteed absent from the training corpus is the decisive check on generalization; the current retrospective benchmark, lacking a stated holdout split, cannot rule out memorization of some structures.
- If the second-order many-body truncation fails to cancel between polymorphs of some molecule, rankings could reverse; testing on molecules with strong three-body dispersion would probe this.
- The per-space-group convergence scheme transfers to other sampling problems where one wants to estimate missed low-energy configurations from repeat counts.
- If the cost claims hold in outside use, routine early-stage CSP could change how polymorph risk is priced in drug development, reducing late-appearing-form surprises like the one that forced reformulation of ritonavir.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces Lavo-NN, an equivariant neural network potential designed for pharmaceutical crystal structure prediction, and packages it into a cloud-based, near-automatic CSP workflow. The protocol is validated on a benchmark of 49 drug-like molecules, with claimed success in generating and ranking all 110 reported Z' = 1 experimental polymorphs at roughly 8.4k CPU hours per molecule, roughly 438k CPU hours in total. Additional case studies address rotigotine, mebendazole, fenofibrate, and progesterone, and a semi-blinded PXRD study assigns structures to polymorphs of three marketed drugs. The paper also compares Lavo-NN with other NNPs on a Top-10 ranking metric, reporting 54% Top-10 accuracy.
Significance. If the central claims are ultimately supported, this would be a substantial advance for pharmaceutical CSP: the benchmark is the largest of its kind, the reported cost reduction is striking, and the PXRD-to-structure case studies illustrate a practical route to solving powder-only forms. The paper's explicit architectural choices (intramolecular-only message passing, a polarizable force-field readout, and a truncated many-body expansion training target) are well motivated and clearly described. The inclusion of a multi-model comparison on a common ranking task is a useful contribution to the community. However, the validity of the headline claims currently depends on the resolution of several load-bearing issues: the possibility of training/benchmark overlap, the post-hoc exclusion of a known polymorph, an unexplained discrepancy in the total polymorph count, and an unvalidated convergence heuristic.
major comments (5)
- [§2.4 vs §5.1] The paper does not establish that the 49 benchmark molecules were held out from Lavo-NN's training data. Section 2.4 describes iterative training on roughly 10k SMILES strings sampled from the SPICE dataset, plus crystal structures, conformers, and dimers generated from those SMILES, with no statement that the benchmark molecules (or their close analogs, tautomers, or protonation states) were excluded. Because the benchmark consists largely of marketed drugs and SPICE is a drug-like dataset, overlap is plausible. The absence of a documented holdout means the claimed generalization to new pharmaceutical molecules is not supported by the presented evidence. The authors should either report an overlap analysis between the training SMILES and the benchmark molecules, or rerun the benchmark on a subset that is provably unseen during training; without this, the central 'successful generation of all experimental polymorphs' result may reflect memorization rather than generalization.
- [§5.1 vs Abstract] The 'all 110' claim is internally inconsistent. The abstract and Section 5 state that the benchmark contains 110 experimental Z' = 1 polymorphs across 49 molecules. Section 5.1 then reports 104 CSD-sourced polymorphs for 46 molecules, and Section 5.2 adds seven PXRD-only forms (omaveloxolone: 2, deucravacitinib: 2, zuranolone: 3), which sums to 111. Moreover, Section 5.1 says both that 'All 104 polymorphs are both generated and ranked' and that 'One polymorph, galunisertib Form I, is excluded from analysis.' The reader cannot determine the exact number of forms attempted and successfully generated. The authors must reconcile these numbers and present a single, unambiguous success count that accounts for all inclusions and exclusions.
- [§3.4] The completeness claim — that the protocol generates all low-energy polymorphs — rests on a convergence estimator that is not validated. Section 3.4 models the counts of unique crystal structures with a Poisson-lognormal distribution and uses the estimated number of unseen structures as the stopping criterion. No evidence is given that this species-abundance analogy is reliable for CSP landscapes, nor is the estimator tested against a known answer (e.g., by subsampling a fully enumerated landscape or by comparing predictions for molecules with exhaustive prior CSP studies). Because the headline 'all 110 polymorphs generated' requires that generation is truly complete, the authors should provide a validation of the convergence metric, for example by showing that the estimator correctly predicts the number of held-out known structures in a subsampling experiment, or by benchmarking against an independent exhaustive search for a small test molecule.
- [§5.1, §5.2, §3] The 'fully automated' claim is qualified by several manual interventions in the benchmark. Specifically, mebendazole required the user to identify three tautomers and run separate CSPs (§5.1.2), progesterone required separate runs in Sohncke and non-Sohncke space groups to cover enantiopure and racemic scenarios (§5.1.4), and the ripretinib landscape required manual addition of a Z' = 2 structure that is outside the stated Z' = 1 scope (Figure 6 caption). These cases contradict the protocol description in Section 3, which states that the only inputs are a SMILES string and a chirality choice, and they undermine the claim of minimal manual input. The authors should either quantify how many of the 49 molecules needed such expert intervention beyond the stated inputs, or explicitly limit the 'fully automated' claim to the subset that required no additional specification.
- [§5.1 vs §5.3] The strong ranking results reported in Section 5.1 are produced by the full protocol, which includes periodic PBE-D3(BJ) and a monomer PBE0 correction (Section 3.5), not by Lavo-NN alone. Section 5.3 shows that Lavo-NN's own Top-10 accuracy is 54%, whereas the benchmark statement 'All 104 polymorphs are both generated and ranked near the bottom' describes the DFT-requalified landscape. The abstract and conclusions should state clearly that the NNP is responsible for structure generation and preliminary ranking, and that the final polymorph rankings in the headline benchmark reflect the additional DFT re-ranking step. As written, the reader could reasonably attribute the ranking success to the NNP, which is not what the data show.
minor comments (6)
- [§2.4] The basis set 'def2-TZPPD' appears to be a typo; the standard basis set is def2-TZVPPD. Please confirm the correct basis set name.
- [Table 1] Compound names are inconsistently capitalized in the table: 'GSk-269984B', 'Mk-2022', 'Mk-8876', and 'TIk-301' should be 'GSK-269984B', 'MK-2022', 'MK-8876', and 'TIK-301'.
- [Figure 1 caption] The caption contains a typo: 'T op' should be 'Top'.
- [§5.1] The sentence 'a dramatic reduction in the typical the computational cost' contains a duplicated article and should be corrected.
- [§5.3] The Top-10 accuracy is evaluated on the landscapes generated by the protocol, which may themselves be biased by Lavo-NN used during generation. The metric is described as a ranking test, but the authors should acknowledge that generation and ranking are not fully decoupled in this evaluation.
- [Appendix A.1] The data availability statement says the landscapes will be made available 'upon publication.' For a benchmark paper, it would strengthen reproducibility if the CIFs and energy files were accessible to reviewers or available as supporting information at submission.
Circularity Check
No demonstrated circularity: the benchmark is external and the central claims do not reduce to training data, fitted parameters, or self-citations by construction.
full rationale
The paper's central derivation chain is self-contained with respect to the benchmark: Lavo-NN is trained on DFT labels for monomers and dimers, the CSP protocol generates and optimizes crystal structures with Lavo-NN and optionally re-ranks them with periodic PBE-D3(BJ) plus a monomer PBE0 correction, and the benchmark outcomes are comparisons against experimental CSD structures and experimental PXRD patterns. None of the stated equations (Eqs. 1-5) define the benchmark answer in terms of the training labels, and no fitted parameter is renamed as a prediction. The self-citations to AP-Net and Splinter are prior-architecture and training-data references, not load-bearing uniqueness theorems or ansatz-smuggling citations. The absence of an explicit holdout split between the 10k SPICE-derived training molecules and the 49 benchmark molecules is a legitimate generalization-risk concern, but the paper text does not establish that any benchmark molecule was actually in the training set, so this cannot be exhibited as a specific circular reduction. Similarly, the exclusions and manual additions (galunisertib Form I, ripretinib, mebendazole, progesterone) qualify the strength of the headline claims but do not make the derivation circular. The core validation is against external experimental structures, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (4)
- Ai, Bi, Ci force-field elemental parameters =
not specified
- Poisson-lognormal convergence parameters (mu, sigma) =
fit per space group per molecule
- RMSD20 similarity cutoff =
0.8 angstrom
- NNP weights =
learned from DFT datasets
assumptions (5)
- domain assumption Experimental crystal structures correspond to minima of the lattice energy landscape.
- domain assumption Second-order many-body expansion is sufficient for polymorph ranking.
- ad hoc to paper The Poisson-lognormal distribution models the distribution of undiscovered crystal structures.
- domain assumption PBE-D3(BJ) with monomer PBE0 correction is accurate enough for final polymorph ranking.
- domain assumption The benchmark molecules are representative of pharmaceutical CSP targets and are not in the training set.
Cite this review
Pith. "Pith review of Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials." pith.science (2026). https://pith.science/paper/Q2ZO2OWW
@misc{pith2026250716218,
author = {Pith},
title = {Pith review of: Toward Routine CSP of Pharmaceuticals: A Fully Automated Protocol Using Neural Network Potentials},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q2ZO2OWW}},
note = {Machine review of arXiv:2507.16218}
}
abstract
Crystal structure prediction (CSP) is a useful tool in pharmaceutical development for identifying and assessing risks associated with polymorphism, yet widespread adoption has been hindered by high computational costs and the need for both manual specification and expert knowledge to achieve useful results. Here, we introduce a fully automated, high-throughput CSP protocol designed to overcome these barriers. The protocol's efficiency is driven by Lavo-NN, a novel neural network potential (NNP) architected and trained specifically for pharmaceutical crystal structure generation and ranking. This NNP-driven crystal generation phase is integrated into a scalable cloud-based workflow. We validate this CSP protocol on an extensive retrospective benchmark of 49 unique molecules, almost all of which are drug-like, successfully generating structures that match all 110 $Z' = 1$ experimental polymorphs. The average CSP in this benchmark is performed with approximately 8.4k CPU hours, which is a significant reduction compared to other protocols. The practical utility of the protocol is further demonstrated through case studies that resolve ambiguities in experimental data and a semi-blinded challenge that successfully identifies and ranks polymorphs of three modern drugs from powder X-ray diffraction patterns alone. By significantly reducing the required time and cost, the protocol enables CSP to be routinely deployed earlier in the drug discovery pipeline, such as during lead optimization. Rapid turnaround times and high throughput also enable CSP that can be run in parallel with experimental screening, providing chemists with real-time insights to guide their work in the lab.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
International tables for crystallography: Space-group symmetry
(2016). International tables for crystallography: Space-group symmetry. http://dx.doi.org/10.1107/97809553602060000114 Adamo, C. & Barone, V. (1999). The Journal of Chemical Physics , 110(13), 6158–6170. http://dx.doi.org/10.1063/1.478522 Agarwal, P., Huckle, J., Newman, J. & Reid, D. L. (2022). Drug Discovery Today, 27(12), 103366. https://www.sciencedir...
-
[2]
Perdew, J. P., Burke, K. & Ernzerhof, M. (1996). Phys. Rev. Lett. 77, 3865–3868. https://link.aps.org/doi/10.1103/PhysRevLett.77.3865 Perdew, J. P. & Schmidt, K. (2001). AIP Conference Proceedings, 577(1), 1–20. https://doi.org/10.1063/1.1390175 Perry, C., Ramos, S., Phelps, M., Mueller, L. & Beran, G. (2025). ChemRxiv. https://doi.org/10.26434/chemrxiv-2...
-
[579]
M¨ uller, M., Hansen, A. & Grimme, S. (2023).Journal of Chemical Physics , 158(1), 014103. Nelson, P. M. & Sherrill, C. D. (2024). The Journal of Chemical Physics , 161(21). Neumann, M. A. (2008). The Journal of Physical Chemistry B , 112(32), 9810–9829. http://dx.doi.org/10.1021/jp710575h Newman, J. A., Iuzzolino, L., Tan, M., Orth, P., Bruhn, J. & Lee, ...
-
[1986]
Siam. Hellweg, A. & Rappoport, D. (2015). Physical Chemistry Chemical Physics , 17, 1010–1017. Heo, Y.-A. (2023). Drugs, 83(16), 1559–1567. Herman, K. M. & Xantheas, S. S. (2023). Journal of Physical Chemistry Letters , 14(4), 989–999. Hilfiker, R. e. (2006). Polymorphism in the Pharmaceutical Industry . Weinheim: Wiley-VCH. Chaps. 2–4 survey solvent/temp...
arXiv 2015
-
[2018]
https://patents.google.com/patent/WO2018018618A1/en Sidhu, G. & Tripp, J. (2023). Fenofibrate. Treasure Island (FL): StatPearls Publishing. Last update March 13,
work page 2023
-
[2023]
https://www.ncbi.nlm.nih.gov/books/NBK559219/ Simmonett, A. C., Pickard, F. C., Shao, Y., Cheatham, T. E. & Brooks, B. R. (2015). The Journal of chemical physics , 143(7). Smith, D. G. A., Burns, L. A., Simmonett, A. C., Parrish, R. M., Schieber, M. C., Galvelis, R., Kraus, P., Kruse, H., Di Remigio, R. et al. (2020). Journal of Chemical Physics , 152(18)...
-
[2025]
https://www.ccdc.cam.ac.uk/structures Campsteyn, H., Dupont, L. & Dideberg, O. (1972). Acta Crystallographica Section B , 28, 3032–
work page 1972
-
[2991]
38 Tang, K. & Toennies, J. P. (1986). Zeitschrift f¨ ur Physik D Atoms, Molecules and Clusters, 1(1), 91–101. Taylor, C. R., Butler, P. W. & Day, G. M. (2025 a). Faraday Discussions, 256, 434–458. Taylor, C. R., Butler, P. W. V. & Day, G. M. (2025 b). Faraday Discussions, 256, 434–458. Thirunahari, S., Aitipamula, S., Chow, P. S. & Tan, R. B. (2010). Jour...
arXiv 1986
Show all 11 references
-
[3042]
C., Kuang, J., Wang, L., Zhang, C., Carbone, M
https://doi.org/10.1107/S0567740872007393 Cao, C., Kingan, A., Hill, R. C., Kuang, J., Wang, L., Zhang, C., Carbone, M. R., van Dam, H., Yoo, S., Marschilok, A. C. & Lu, D. (2025). PRX Energy, 4, 023004. Catlow, C. R. A. (2023). IUCrJ, 10(2), 143–144. http://dx.doi.org/10.1107...
2025
-
[5148]
D., Markland, T
Saban´ es Zariquiey, F., Galvelis, R., Gallicchio, E., Chodera, J. D., Markland, T. E. & De Fabritiis, G. (2024). Journal of Chemical Information and Modeling , 64(5), 1481–1485. Sargent, C. T., Metcalf, D. P., Glick, Z. L., Borca, C. H. & Sherrill, C. D. (2023). The Journal o...
2024
-
[6615]
& L´ opez, N
Lian, Z., Dattila, F. & L´ opez, N. (2024).Nature Catalysis, 7, 401–411. Liang, Y. H., Ye, H.-Z. & Berkelbach, T. C. (2023). The Journal of Physical Chemistry Letters , 14(46), 10435–10441. Lin, L., Pan, L. & Liu, S. (2022). Computers in Industry , 141, 103718. https://www.sci...
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.