Pith. sign in

REVIEW 3 major objections 3 minor 2 cited by

NovoMolGen: Rethinking Molecular Language Model Pretraining

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper claims that standard pretraining metrics (loss, perplexity) predict molecular generation quality only weakly, and that its 1.5-billion-molecule NovoMolGen family beats prior molecular LLMs and specialized generative models on unco

desk verdict A solid, open empirical map of molecular LM pretraining choices, but the SOTA claim needs per-seed variance and clearer PMO tuning separation before I'd trust the headline numbers. read the letter →

arxiv 2508.13408 v2 pith:GGULXS7O submitted 2025-08-19 cs.LG

classification cs.LG
keywords denovomoleculegenerationmolecularlanguagemodelsSMILEStokenizationpretrainingchemicalspaceexplorationgoal-directedfoundation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the usual language-model pretraining signals are the wrong yardstick for molecule generation: across a family of transformer models trained on 1.5 billion molecules, loss and perplexity correlate only weakly with how well the models generate valid molecules with desired properties. It claims that its NovoMolGen family sets a new high-water mark for both unconstrained de novo generation and goal-directed generation, outperforming earlier molecular LLMs and specialized generative models. The practical consequence is that molecular-model builders should choose representations, tokenizers, sizes, and checkpoints by downstream task scores instead of pretraining loss. If the claim holds, molecular language modeling separates from general NLP training dynamics and gains an open pretrained foundation model for de novo design.

What carries the argument

The load-bearing object is a family of autoregressive transformers trained with next-token prediction on molecular strings, evaluated across four string encodings and two tokenization schemes. The analysis device is an ablation grid—representation $\times$ tokenizer $\times$ model scale $\times$ training data—that lets the paper separate pretraining behavior (loss, perplexity) from downstream behavior (validity, property scores, docking rewards). The weak correlation between those two measurement levels is the mechanism that overturns standard model-selection practice.

What would settle it

Rerun the same pretraining configurations and baseline comparisons across many independent seeds on the unconstrained and goal-directed benchmarks and report per-seed distributions. The central claims fail if the NovoMolGen margins shrink to noise, or if pretraining loss, across seeds, actually ranks downstream performance as well as downstream scores do.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that pretraining objective values are not a reliable proxy for generation ability in molecular language models. After systematically varying string representation (SMILES, SELFIES, DeepSMILES, SAFE), tokenization (atom-wise vs BPE), model size (32M to 300M), and dataset scale (up to 1.5 billion molecules), the model family it calls NovoMolGen delivers its best downstream generation results in configurations whose pretraining metrics would not have predicted the win. The paper reports state-of-the-art performance against prior Mol-LLMs and specialized generative models, with detailed analyses centered on a 32M-parameter SMILES-BPE variant.

Load-bearing premise

The load-bearing premise is that the measured benchmark scores reflect true model performance rather than run-to-run noise, since no per-seed variance is reported and example molecules come from a single run.

Editorial extensions

If this is right

  • Molecular LM comparisons should report downstream generation scores, not just pretraining loss or perplexity, because the paper finds those pretraining signals do not rank models correctly.
  • If the state-of-the-art results hold, NovoMolGen gives the field an open, pretrained transformer family that can serve as a base for constrained and goal-directed generation without task-specific architectures.
  • Tokenizer and representation choices are not interchangeable: the same model family changes downstream generation quality materially across SMILES, SELFIES, DeepSMILES, and SAFE and across atom-wise and BPE tokenization.
  • Pretraining at 1.5 billion molecules with simple string inputs is sufficient to move goal-directed generation scores beyond prior specialized graph and 3D generators.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Not stated in the paper: model and checkpoint selection should switch to downstream oracle scores; this is testable by comparing hit rates of downstream-selected versus perplexity-selected checkpoints.
  • Not stated in the paper: if the weak correlation is general, task-specific scaling curves (for example, docking reward versus training tokens) are needed rather than NLP-style loss scaling laws.
  • Not stated in the paper: because BPE tokens capture chemical substructures, the tokenizer comparison could be pushed further by measuring whether the winning tokenizer shifts generation toward synthesizable fragments.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. NovoMolGen is an empirical study introducing a family of transformer-based molecular language models pretrained on 1.5 billion molecules, spanning multiple string representations (SMILES, SAFE, SELFIES, DeepSMILES), tokenization schemes (AtomWise, BPE), and model sizes (32M, 157M, 300M). The authors report systematic ablations of these design choices, evaluate unconstrained and goal-directed (PMO) generation, and claim that pretraining metrics such as loss correlate only weakly with downstream generation performance. They further claim that their best configuration, NovoMolGen-SMILES-BPE, establishes new state-of-the-art results. Model weights and code are publicly released.

Significance. If the claims hold, this is a substantial contribution: a large-scale, openly released family of molecular pretraining models, an unusually broad comparison of representations and tokenizers, and a practical caution that standard pretraining metrics may not select the best generative model. The weak-correlation finding, if made quantitative, would be a useful methodological message for the Mol-LLM community. The release of code and weights is a clear strength. However, the headline SOTA claim currently depends on evaluation protocols whose run-to-run variability and selection procedure are not sufficiently documented in the reviewed material, so the significance is conditional on those issues being resolved.

major comments (3)
  1. [Appendix I, Figs. I23–I27] The PMO benchmark results are reported as point estimates without per-seed variance, confidence intervals, or repeated-run statistics. Figure I27 explicitly says the shown top-5 molecules come from 'a single run', and no evidence is provided that the aggregate/3K/10K scores in the tables are stable across seeds. Goal-directed generation scores are order statistics of a stochastic optimization trajectory, making them especially sensitive to seed, oracle budget, and starting checkpoint. The abstract's 'substantially outperforming' claim requires that the reported margins exceed run-to-run noise; with point estimates alone this is not established. Please report mean and standard deviation over at least 3–5 seeds, perform significance tests against each baseline, and state the oracle budget and stopping rule.
  2. [Appendix E, Figs. E10–E21] The parallel-coordinate plots appear to show that hyperparameters were selected by directly inspecting the PMO Aggregated Score, which is the same aggregate metric used in the goal-directed SOTA comparison. Selecting configurations on the evaluation metric can inflate apparent performance relative to baselines that did not undergo the same selection. Please clarify whether the PMO Aggregated Score was used as a validation criterion. If so, the claim of SOTA requires an honest protocol—for example, holdout PMO tasks, nested validation, or applying the same selection procedure to the baselines—or a quantitative estimate of selection bias.
  3. [Abstract / weak-correlation analysis] The paper's central 'rethinking' claim is that pretraining metrics correlate only weakly with downstream generation performance. The reviewed material does not report any numerical correlation coefficient, scatter plot, number of configurations compared, confidence interval, or statistical test. Without specifying which pretraining metrics (loss, perplexity, validity, etc.), which downstream scores, and how 'weak' is quantified, the claim is not testable. Please report Spearman/Pearson correlations with p-values or confidence intervals across the swept configurations for each pretraining metric against each downstream benchmark.
minor comments (3)
  1. [Fig. I27 caption] Please clarify whether the displayed top-5 molecules are representative examples or a selected best case, and include the success rate or validity rate of the generated set for context.
  2. [Figs. E10–E21] The parallel-coordinate plots are hard to read. Mark the chosen hyperparameter configuration for each model and provide a table of the tested ranges and final selected values.
  3. [Appendix K] Figure K30 shows BPE substructures but does not explain how the substructures were extracted or why they are chemically meaningful. A brief description of the tokenization vocabulary and examples of full SMILES encodings would help the reader interpret the tokenization comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical benchmark study with externally evaluated claims.

full rationale

This is an empirical benchmark study rather than a derivation, so the standard circularity patterns do not apply. The central weak-correlation finding is an observed statistical relationship between pretraining losses and downstream PMO scores; it does not define either quantity in terms of the other, and it is not used as an input to compute those scores. The SOTA claim is a measured outcome on an external benchmark (PMO), not a quantity reconstructed from the model's own premises or from a self-citation. Appendix E shows parallel-coordinate analyses of hyperparameters against the aggregated score; even if this implies model selection on the PMO metric, that would be a test-set-selection / evaluation-protocol concern rather than a construction-level circularity. Appendix I's 'single run' examples (Figure I27, J29) are stability limitations that caution against over-reading the headline margins, but they do not make any result equal to its inputs. No load-bearing self-citations, no imported uniqueness theorems, and no ansatz-as-citation steps are present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This paper is an empirical model-development and benchmark study, not a derivation, so the ledger is short. The central claim depends on assumptions that autoregressive language-model pretraining transfers to molecule design, that the evaluation oracles are meaningful proxies for useful molecules, and that the training corpus covers the relevant chemical space. The comparison also assumes single-run or few-run benchmark scores are stable enough to rank models, which the paper's own appendix (single-run example molecules) calls into question. No new physical entities are introduced, and no parameters are fitted in a mathematical derivation; the model hyperparameters are design choices investigated in the study.

assumptions (4)
  • domain assumption Autoregressive next-token prediction on molecular strings is a sufficient pretraining objective for downstream de novo generation.
    The whole NovoMolGen approach relies on standard LM pretraining transferring to molecule design; invoked throughout the pretraining description and the weak-correlation analysis.
  • domain assumption Benchmark scoring functions (QED, SA, logP, docking oracles) used in goal-directed tasks are valid proxies for drug-likeness and synthesizability.
    The goal-directed SOTA claim is measured with these oracles; if they are not faithful to real design objectives, the practical significance is reduced.
  • domain assumption The 1.5B-molecule pretraining corpus is representative of the relevant synthesizable chemical space.
    The abstract frames exploration of 10^23 to 10^60 candidates; generalization depends on the training distribution covering useful chemistry.
  • domain assumption Reported benchmark scores are stable enough to rank models, despite some reported runs being single-run.
    Figure I27 in Appendix I explicitly shows a single-run example, so the claim of clear SOTA margins assumes evaluation noise is small.

how reviews work

0 comments
Cite this review

Pith. "Pith review of NovoMolGen: Rethinking Molecular Language Model Pretraining." pith.science (2026). https://pith.science/paper/GGULXS7O

@misc{pith2026250813408,
  author       = {Pith},
  title        = {Pith review of: NovoMolGen: Rethinking Molecular Language Model Pretraining},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GGULXS7O}},
  note         = {Machine review of arXiv:2508.13408}
}
abstract

Designing de-novo molecules with desired property profiles requires efficient exploration of the vast chemical space ranging from $10^{23}$ to $10^{60}$ possible synthesizable candidates. While various deep generative models have been developed to design small molecules using diverse input representations, Molecular Large Language Models (Mol-LLMs) based on string representations have emerged as a scalable approach capable of exploring billions of molecules. However, there remains limited understanding regarding how standard language modeling practices such as textual representations, tokenization strategies, model size, and dataset scale impact molecular generation performance. In this work, we systematically investigate these critical aspects by introducing NovoMolGen, a family of transformer-based foundation models pretrained on 1.5 billion molecules for de-novo molecule generation. Through extensive empirical analyses, we identify a weak correlation between performance metrics measured during pretraining and actual downstream performance, revealing important distinctions between molecular and general NLP training dynamics. NovoMolGen establishes new state-of-the-art results, substantially outperforming prior Mol-LLMs and specialized generative models in both unconstrained and goal-directed molecular generation tasks, thus providing a robust foundation for advancing efficient and effective molecular modeling strategies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Augmenting Molecular Language Models with Local $n$-gram Memory

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    MolGram integrates a conditional n-gram memory module into molecular language models to address locality gaps in SMILES tokenization, improving performance on generation, forward prediction, and retrosynthesis while o...

  2. FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    FORGE reformulates molecular optimization as context-aware fragment ranking and replacement using mined low-to-high edit pairs, outperforming larger language models and graph methods on standard benchmarks.

Reference graph

Works this paper leans on

93 extracted references · 39 canonical work pages · cited by 2 Pith papers

  1. [1]

    Generative Pre - Training from Molecules , September 2021

    Sanjar Adilov. Generative Pre - Training from Molecules , September 2021. URL https://chemrxiv.org/engage/chemrxiv/article-details/6142f60742198e8c31782e9e

  2. [2]

    Fast, accurate, and reliable molecular docking with QuickVina 2

    Amr Alhossary, Stephanus Daniel Handoko, Yuguang Mu, and Chee-Keong Kwoh. Fast, accurate, and reliable molecular docking with QuickVina 2. Bioinformatics, 31 0 (13): 0 2214--2216, July 2015. ISSN 1367-4803. doi:10.1093/bioinformatics/btv082. URL https://doi.org/10.1093/bioinformatics/btv082

  3. [3]

    Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. MolGPT : Molecular Generation Using a Transformer - Decoder Model . Journal of Chemical Information and Modeling, 62 0 (9): 0 2064--2076, May 2022. ISSN 1549-9596. doi:10.1021/acs.jcim.1c00600. URL https://doi.org/10.1021/acs.jcim.1c00600. Publisher: American Chemical Society

  4. [4]

    Bemis and Mark A

    Guy W. Bemis and Mark A. Murcko. The Properties of Known Drugs . 1. Molecular Frameworks . Journal of Medicinal Chemistry, 39 0 (15): 0 2887--2893, January 1996. ISSN 0022-2623. doi:10.1021/jm9602928. URL https://doi.org/10.1021/jm9602928. Publisher: American Chemical Society

  5. [5]

    Acegen: Reinforcement learning of generative chemical agents for drug discovery

    Albert Bou, Morgan Thomas, Sebastian Dittert, Carles Navarro, Maciej Majewski, Ye Wang, Shivam Patel, Gary Tresadern, Mazen Ahmad, Vincent Moens, et al. Acegen: Reinforcement learning of generative chemical agents for drug discovery. Journal of Chemical Information and Modeling, 64 0 (15): 0 5900--5911, 2024

  6. [6]

    GNN - FiLM : Graph Neural Networks with Feature -wise Linear Modulation

    Marc Brockschmidt. GNN - FiLM : Graph Neural Networks with Feature -wise Linear Modulation . In Proceedings of the 37th International Conference on Machine Learning , pages 1144--1152. PMLR, November 2020. URL https://proceedings.mlr.press/v119/brockschmidt20a.html. ISSN: 2640-3498

  7. [7]

    MolGAN : An implicit generative model for small molecular graphs, September 2022

    Nicola De Cao and Thomas Kipf. MolGAN : An implicit generative model for small molecular graphs, September 2022. URL http://arxiv.org/abs/1805.11973. arXiv:1805.11973 [stat]

  8. [8]

    BARTSmiles: Generative Masked Language Models for Molecular Representations

    Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan, Nelly Babayan, Lusine Khondkaryan, Karen Hambardzumyan, Zaven Navoyan, Hrant Khachatrian, and Armen Aghajanyan. BARTSmiles : Generative Masked Language Models for Molecular Representations , November 2022. URL http://arxiv.org/abs/2211.16349. arXiv:2211.16349 [cs, q-bio]

Show all 93 references
  1. [9]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FLASHATTENTION : fast and memory-efficient exact attention with IO -awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS '22, pages 16344--16359, Red...

  2. [10]

    On the Art of Compiling and Using ' Drug - Like ' Chemical Fragment Spaces

    Jörg Degen, Christof Wegscheid-Gerlach, Andrea Zaliani, and Matthias Rarey. On the Art of Compiling and Using ' Drug - Like ' Chemical Fragment Spaces . ChemMedChem, 3 0 (10): 0 1503--1507, 2008. ISSN 1860-7187. doi:10.1002/cmdc.200800178. URL https://onlinelibrary.wiley.com/d...

  3. [11]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  4. [12]

    Durant, Burton A

    Joseph L. Durant, Burton A. Leland, Douglas R. Henry, and James G. Nourse. Reoptimization of MDL Keys for Use in Drug Discovery . Journal of Chemical Information and Computer Sciences, 42 0 (6): 0 1273--1280, November 2002. ISSN 0095-2338. doi:10.1021/ci010132r. URL https://do...

  5. [13]

    LIMO : Latent Inceptionism for Targeted Molecule Generation

    Peter Eckmann, Kunyang Sun, Bo Zhao, Mudong Feng, Michael Gilson, and Rose Yu. LIMO : Latent Inceptionism for Targeted Molecule Generation . In Proceedings of the 39th International Conference on Machine Learning , pages 5777--5792. PMLR, June 2022. URL https://proceedings.mlr...

  6. [14]

    Bioreason: Incentivizing multimodal biological reasoning within a dna-llm model

    Adibvafa Fallahpour, Andrew Magnuson, Purav Gupta, Shihao Ma, Jack Naimer, Arnav Shah, Haonan Duan, Omar Ibrahim, Hani Goodarzi, Chris J Maddison, et al. Bioreason: Incentivizing multimodal biological reasoning within a dna-llm model. arXiv preprint arXiv:2505.23579, 2025

  7. [15]

    Geometry-enhanced molecular representation learning for property prediction

    Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4 0 (2): 0 127--134, February 2022. ISSN 2522-5839. d...

  8. [16]

    Domain- Agnostic Molecular Generation with Chemical Feedback

    Yin Fang, Ningyu Zhang, Zhuo Chen, Lingbing Guo, Xiaohui Fan, and Huajun Chen. Domain- Agnostic Molecular Generation with Chemical Feedback . In The Twelfth International Conference on Learning Representations, October 2023. URL https://openreview.net/forum?id=9rPyHyjfwP

  9. [17]

    Neural scaling of deep chemical models

    Nathan C Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi, Rafael Gomez-Bombarelli, Connor W Coley, and Vijay Gadepally. Neural scaling of deep chemical models. Nature Machine Intelligence, 5 0 (11): 0 1297--1305, 2023

  10. [18]

    Wenhao Gao, Tianfan Fu, Jimeng Sun, and Connor W. Coley. Sample Efficiency Matters : A Benchmark for Practical Molecular Optimization , October 2022. URL http://arxiv.org/abs/2206.12411. arXiv:2206.12411

  11. [19]

    Niklas W. A. Gebauer, Michael Gastegger, and Kristof T. Schütt. Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules. In Proceedings of the 33rd International Conference on Neural Information Processing Systems , pages 7566--7578. Curran Associa...

  12. [20]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for Quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 , ICML '17, pages 1263--1272, Sydney, NSW, Australia, Aug...

  13. [21]

    Bidirectional Molecule Generation with Recurrent Neural Networks

    Francesca Grisoni, Michael Moret, Robin Lingwood, and Gisbert Schneider. Bidirectional Molecule Generation with Recurrent Neural Networks . Journal of Chemical Information and Modeling, 60 0 (3): 0 1175--1183, March 2020. ISSN 1549-9596. doi:10.1021/acs.jcim.9b00943. URL https...

  14. [22]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...

  15. [23]

    Saturn: Sample-efficient generative molecular design using memory manipulation

    Jeff Guo and Philippe Schwaller. Saturn: Sample-efficient generative molecular design using memory manipulation. arXiv preprint arXiv:2405.17066, 2024

  16. [24]

    Scaffold splits overestimate virtual screening performance

    Qianrong Guo, Saiveth Hernandez-Hernandez, and Pedro J Ballester. Scaffold splits overestimate virtual screening performance. In International Conference on Artificial Neural Networks, pages 58--72. Springer, 2024

  17. [25]

    Iyer, Yihong Ma, Olaf Wiest, Xiangliang Zhang, Wei Wang, Chuxu Zhang, and Nitesh V

    Zhichun Guo, Kehan Guo, Bozhao Nan, Yijun Tian, Roshni G. Iyer, Yihong Ma, Olaf Wiest, Xiangliang Zhang, Wei Wang, Chuxu Zhang, and Nitesh V. Chawla. Graph-based Molecular Representation Learning . In Proceedings of the Thirty-Second International Joint Conference on Artificia...

  18. [26]

    Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D

    Rafael Gómez-Bombarelli, Jennifer N. Wei, David Duvenaud, José Miguel Hernández-Lobato, Benjamín Sánchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D. Hirzel, Ryan P. Adams, and Alán Aspuru-Guzik. Automatic Chemical Design Using a Data - Driven Continuous...

  19. [27]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Proceedings of the 34th International Conference on Neural Information Processing Systems , NIPS '20, pages 6840--6851, Red Hook, NY, USA, December 2020. Curran Associates Inc. ISBN 978-1-7...

  20. [28]

    Equivariant Diffusion for Molecule Generation in 3D

    Emiel Hoogeboom, V\'ictor Garcia Satorras, Cl\'ement Vignac, and Max Welling. Equivariant Diffusion for Molecule Generation in 3D . In Proceedings of the 39th International Conference on Machine Learning , pages 8867--8887. PMLR, June 2022. URL https://proceedings.mlr.press/v1...

  21. [29]

    MDM : Molecular Diffusion Model for 3D Molecule Generation

    Lei Huang, Hengtong Zhang, Tingyang Xu, and Ka-Chun Wong. MDM : Molecular Diffusion Model for 3D Molecule Generation . Proceedings of the AAAI Conference on Artificial Intelligence, 37 0 (4): 0 5105--5112, June 2023. ISSN 2374-3468. doi:10.1609/aaai.v37i4.25639. URL https://oj...

  22. [30]

    Zinc: a free tool to discover chemistry for biology

    John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman. Zinc: a free tool to discover chemistry for biology. Journal of chemical information and modeling, 52 0 (7): 0 1757--1768, 2012

  23. [31]

    Chemformer: a pre-trained transformer for computational chemistry

    Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. Chemformer: a pre-trained transformer for computational chemistry. Machine Learning: Science and Technology, 3 0 (1): 0 015022, January 2022. ISSN 2632-2153. doi:10.1088/2632-2153/ac3ffb. URL https://dx.do...

  24. [32]

    Autonomous molecule generation using reinforcement learning and docking to develop potential novel inhibitors

    Woosung Jeon and Dongsup Kim. Autonomous molecule generation using reinforcement learning and docking to develop potential novel inhibitors. Scientific reports, 10 0 (1): 0 22104, 2020

  25. [33]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  26. [34]

    Junction Tree Variational Autoencoder for Molecular Graph Generation

    Wengong Jin, Regina Barzilay, and Tommi Jaakkola. Junction Tree Variational Autoencoder for Molecular Graph Generation . In Proceedings of the 35th International Conference on Machine Learning , pages 2323--2332. PMLR, July 2018. URL https://proceedings.mlr.press/v80/jin18a.ht...

  27. [35]

    Multi- Objective Molecule Generation using Interpretable Substructures

    Wengong Jin, Dr Regina Barzilay, and Tommi Jaakkola. Multi- Objective Molecule Generation using Interpretable Substructures . In Proceedings of the 37th International Conference on Machine Learning , pages 4849--4859. PMLR, November 2020 a . URL https://proceedings.mlr.press/v...

  28. [36]

    Hierarchical generation of molecular graphs using structural motifs

    Wengong Jin, Regina Barzilay, and Tommi Jaakkola. Hierarchical generation of molecular graphs using structural motifs. In Proceedings of the 37th International Conference on Machine Learning , volume 119 of ICML '20 , pages 4839--4848. JMLR.org, July 2020 b

  29. [37]

    Score-based Generative Modeling of Graphs via the System of Stochastic Differential Equations

    Jaehyeong Jo, Seul Lee, and Sung Ju Hwang. Score-based Generative Modeling of Graphs via the System of Stochastic Differential Equations . In Proceedings of the 39th International Conference on Machine Learning , pages 10362--10383. PMLR, June 2022. URL https://proceedings.mlr...

  30. [38]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. Semi- Supervised Classification with Graph Convolutional Networks . In International Conference on Learning Representations, February 2017. URL https://openreview.net/forum?id=SJU4ayYgl

  31. [39]

    Chemical space

    Peter Kirkpatrick and Clare Ellis. Chemical space. Nature, 432 0 (7019): 0 823--823, December 2004. ISSN 1476-4687. doi:10.1038/432823a. URL https://www.nature.com/articles/432823a. Publisher: Nature Publishing Group

  32. [40]

    Aditya Prakash, and Chao Zhang

    Lingkai Kong, Jiaming Cui, Haotian Sun, Yuchen Zhuang, B. Aditya Prakash, and Chao Zhang. Autoregressive Diffusion Model for Graph Generation . In Proceedings of the 40th International Conference on Machine Learning , pages 17391--17408. PMLR, July 2023. URL https://proceeding...

  33. [41]

    Frey, Pascal Friederich, Théophile Gaudin, Alberto Alexander Gayle, Kevin Maik Jablonka, Rafael F

    Mario Krenn, Qianxiang Ai, Senja Barthel, Nessa Carson, Angelo Frei, Nathan C. Frey, Pascal Friederich, Théophile Gaudin, Alberto Alexander Gayle, Kevin Maik Jablonka, Rafael F. Lameiro, Dominik Lemm, Alston Lo, Seyed Mohamad Moosavi, José Manuel Nápoles-Duarte, AkshatKumar Ni...

  34. [42]

    Subword Regularization : Improving Neural Network Translation Models with Multiple Subword Candidates

    Taku Kudo. Subword Regularization : Improving Neural Network Translation Models with Multiple Subword Candidates . In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , p...

  35. [43]

    MolGrow : A Graph Normalizing Flow for Hierarchical Molecular Generation

    Maksim Kuznetsov and Daniil Polykovskiy. MolGrow : A Graph Normalizing Flow for Hierarchical Molecular Generation . Proceedings of the AAAI Conference on Artificial Intelligence, 35 0 (9): 0 8226--8234, May 2021. ISSN 2374-3468. doi:10.1609/aaai.v35i9.17001. URL https://ojs.aa...

  36. [44]

    Neuraldecipher – reverse-engineering extended-connectivity fingerprints ( ECFPs ) to their molecular structures

    Tuan Le, Robin Winter, Frank Noé, and Djork-Arné Clevert. Neuraldecipher – reverse-engineering extended-connectivity fingerprints ( ECFPs ) to their molecular structures. Chemical Science, 11 0 (38): 0 10378--10389, October 2020. ISSN 2041-6539. doi:10.1039/D0SC03115A. URL htt...

  37. [45]

    Exploring Chemical Space with Score -based Out -of-distribution Generation

    Seul Lee, Jaehyeong Jo, and Sung Ju Hwang. Exploring Chemical Space with Score -based Out -of-distribution Generation . In Proceedings of the 40th International Conference on Machine Learning , pages 18872--18892. PMLR, July 2023. URL https://proceedings.mlr.press/v202/lee23f....

  38. [46]

    Molecule Generation with Fragment Retrieval Augmentation

    Seul Lee, Karsten Kreis, Srimukh Prasad Veccham, Meng Liu, Danny Reidenbach, Saee Gopal Paliwal, Arash Vahdat, and Weili Nie. Molecule Generation with Fragment Retrieval Augmentation . In The Thirty-eighth Annual Conference on Neural Information Processing Systems, November 20...

  39. [47]

    Drug discovery with dynamic goal-aware fragments

    Seul Lee, Seanie Lee, Kenji Kawaguchi, and Sung Ju Hwang. Drug discovery with dynamic goal-aware fragments. In International Conference on Machine Learning, pages 26731--26751. PMLR, 2024 b

  40. [48]

    Limits to depth-efficiencies of self-attention

    Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua. Limits to depth-efficiencies of self-attention. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS '20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546

  41. [49]

    Chemical reaction enhanced graph learning for molecule representation

    Anchen Li, Elena Casiraghi, and Juho Rousu. Chemical reaction enhanced graph learning for molecule representation. Bioinformatics, 40 0 (10): 0 btae558, October 2024. ISSN 1367-4811. doi:10.1093/bioinformatics/btae558. URL https://doi.org/10.1093/bioinformatics/btae558

  42. [50]

    GeomGCL : Geometric Graph Contrastive Learning for Molecular Property Prediction

    Shuangli Li, Jingbo Zhou, Tong Xu, Dejing Dou, and Hui Xiong. GeomGCL : Geometric Graph Contrastive Learning for Molecular Property Prediction . Proceedings of the AAAI Conference on Artificial Intelligence, 36 0 (4): 0 4541--4549, June 2022. ISSN 2374-3468. doi:10.1609/aaai.v...

  43. [51]

    A 3D generative model for structure-based drug design

    Shitong Luo, Jiaqi Guan, Jianzhu Ma, and Jian Peng. A 3D generative model for structure-based drug design. In Proceedings of the 35th International Conference on Neural Information Processing Systems , NIPS '21, pages 6229--6239, Red Hook, NY, USA, June 2024. Curran Associates...

  44. [52]

    Masked graph modeling for molecule generation

    Omar Mahmood, Elman Mansimov, Richard Bonneau, and Kyunghyun Cho. Masked graph modeling for molecule generation. Nature Communications, 12 0 (1): 0 3156, May 2021. ISSN 2041-1723. doi:10.1038/s41467-021-23415-2. URL https://www.nature.com/articles/s41467-021-23415-2. Publisher...

  45. [53]

    Provably Powerful Graph Networks

    Haggai Maron, Heli Ben-Hamu, Hadar Serviansky, and Yaron Lipman. Provably Powerful Graph Networks . In Proceedings of the 33rd International Conference on Neural Information Processing Systems , pages 2156--2167. Curran Associates Inc., Red Hook, NY, USA, December 2019

  46. [54]

    Molecule generation using transformers and policy gradient reinforcement learning

    Eyal Mazuz, Guy Shtar, Bracha Shapira, and Lior Rokach. Molecule generation using transformers and policy gradient reinforcement learning. Scientific Reports, 13 0 (1): 0 8799, May 2023. ISSN 2045-2322. doi:10.1038/s41598-023-35648-w. URL https://www.nature.com/articles/s41598...

  47. [55]

    Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe

    Christopher Morris, Martin Ritzert, Matthias Fey, William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and Leman Go Neural : Higher - Order Graph Neural Networks . Proceedings of the AAAI Conference on Artificial Intelligence, 33 0 (01): 0 4602--4...

  48. [56]

    Emmanuel Noutahi, Cristian Gabellini, Michael Craig, Jonathan S. C. Lim, and Prudencio Tossou. Gotta be SAFE : a new framework for molecular design. Digital Discovery, 3 0 (4): 0 796--804, 2024. doi:10.1039/D4DD00019F. URL https://pubs.rsc.org/en/content/articlelanding/2024/dd...

  49. [57]

    DeepSMILES : An Adaptation of SMILES for Use in Machine - Learning of Chemical Structures , September 2018

    Noel O'Boyle and Andrew Dalke. DeepSMILES : An Adaptation of SMILES for Use in Machine - Learning of Chemical Structures , September 2018. URL https://chemrxiv.org/engage/chemrxiv/article-details/60c73ed6567dfe7e5fec388d

  50. [58]

    Molecular de-novo design through deep reinforcement learning

    Marcus Olivecrona, Thomas Blaschke, Ola Engkvist, and Hongming Chen. Molecular de-novo design through deep reinforcement learning. Journal of Cheminformatics, 9 0 (1): 0 48, September 2017. ISSN 1758-2946. doi:10.1186/s13321-017-0235-x. URL https://doi.org/10.1186/s13321-017-0235-x

  51. [59]

    The jungle of generative drug discovery: Traps, treasures, and ways out

    R za \"O z c elik and Francesca Grisoni. The jungle of generative drug discovery: Traps, treasures, and ways out. arXiv preprint arXiv:2501.05457, 2024

  52. [60]

    A Deep Generative Model for Fragment - Based Molecule Generation

    Marco Podda, Davide Bacciu, and Alessio Micheli. A Deep Generative Model for Fragment - Based Molecule Generation . In Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics , pages 2240--2250. PMLR, June 2020. URL https://proceeding...

  53. [61]

    Molecular Sets ( MOSES ): A Benchmarking Platform for Molecular Generation Models

    Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling, Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, Artur Kadurin, Simon Johansson, Hongming Chen, Sergey Nikolenko, Alán Aspuru-Guzik, and A...

  54. [62]

    Extended- Connectivity Fingerprints

    David Rogers and Mathew Hahn. Extended- Connectivity Fingerprints . Journal of Chemical Information and Modeling, 50 0 (5): 0 742--754, May 2010. ISSN 1549-9596. doi:10.1021/ci100050t. URL https://doi.org/10.1021/ci100050t. Publisher: American Chemical Society

  55. [63]

    Large- Scale Chemical Language Representations Capture Molecular Structure and Properties , December 2022

    Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das. Large- Scale Chemical Language Representations Capture Molecular Structure and Properties , December 2022. URL http://arxiv.org/abs/2106.09553. arXiv:2106.09553 [cs, q-bio]

  56. [64]

    Hoffman, Vijil Chenthamarakshan, Youssef Mroueh, and Payel Das

    Jerret Ross, Brian Belgodere, Samuel C. Hoffman, Vijil Chenthamarakshan, Youssef Mroueh, and Payel Das. GP - MoLFormer : A Foundation Model For Molecular Generation , April 2024. URL http://arxiv.org/abs/2405.04912. arXiv:2405.04912 [q-bio]

  57. [65]

    Blum, and Jean-Louis Reymond

    Lars Ruddigkeit, Ruud van Deursen, Lorenz C. Blum, and Jean-Louis Reymond. Enumeration of 166 Billion Organic Small Molecules in the Chemical Universe Database GDB -17. Journal of Chemical Information and Modeling, 52 0 (11): 0 2864--2875, November 2012. ISSN 1549-9596. doi:10...

  58. [66]

    Proximal policy optimization algorithms

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017

  59. [67]

    Hunter, Costas Bekas, and Alpha A

    Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee. Molecular Transformer : A Model for Uncertainty - Calibrated Chemical Reaction Prediction . ACS Central Science, 5 0 (9): 0 1572--1583, September 2019. ISS...

  60. [68]

    Neural Machine Translation of Rare Words with Subword Units

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural Machine Translation of Rare Words with Subword Units . In Katrin Erk and Noah A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pages 1...

  61. [69]

    GraphVAE : Towards Generation of Small Graphs Using Variational Autoencoders , February 2018

    Martin Simonovsky and Nikos Komodakis. GraphVAE : Towards Generation of Small Graphs Using Variational Autoencoders , February 2018. URL http://arxiv.org/abs/1802.03480. arXiv:1802.03480 [cs] version: 1

  62. [70]

    Chemical language models enable navigation in sparsely populated chemical space

    Michael A Skinnider, R Greg Stacey, David S Wishart, and Leonard J Foster. Chemical language models enable navigation in sparsely populated chemical space. Nature Machine Intelligence, 3 0 (9): 0 759--770, 2021

  63. [71]

    Deep Unsupervised Learning using Nonequilibrium Thermodynamics

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep Unsupervised Learning using Nonequilibrium Thermodynamics . In Proceedings of the 32nd International Conference on Machine Learning , pages 2256--2265. PMLR, June 2015. URL https://proceedings.mlr...

  64. [72]

    Tazhigulov, Joshua Schiller, Jacob Oppenheim, and Max Winston

    Ruslan N. Tazhigulov, Joshua Schiller, Jacob Oppenheim, and Max Winston. Molecular Fingerprints for Robust and Efficient ML - Driven Molecular Generation , November 2022. URL http://arxiv.org/abs/2211.09086. arXiv:2211.09086 [cs]

  65. [73]

    Augmented hill-climb increases reinforcement learning efficiency for language-based de novo molecule generation

    Morgan Thomas, Noel M O’Boyle, Andreas Bender, and Chris De Graaf. Augmented hill-climb increases reinforcement learning efficiency for language-based de novo molecule generation. Journal of cheminformatics, 14 0 (1): 0 68, 2022

  66. [74]

    Promptsmiles: prompting for scaffold decoration and fragment linking in chemical language models

    Morgan Thomas, Mazen Ahmad, Gary Tresadern, and Gianni De Fabritiis. Promptsmiles: prompting for scaffold decoration and fragment linking in chemical language models. Journal of Cheminformatics, 16 0 (1): 0 77, 2024

  67. [75]

    Tingle, Khanh G

    Benjamin I. Tingle, Khanh G. Tang, Mar Castanon, John J. Gutierrez, Munkhzul Khurelbaatar, Chinzorig Dandarchuluun, Yurii S. Moroz, and John J. Irwin. ZINC -22- A Free Multi - Billion - Scale Database of Tangible Compounds for Ligand Discovery . Journal of Chemical Information...

  68. [76]

    Genetic algorithms are strong baselines for molecule generation

    Austin Tripp and Jos \'e Miguel Hern \'a ndez-Lobato. Genetic algorithms are strong baselines for molecule generation. arXiv preprint arXiv:2310.09267, 2023

  69. [77]

    cMolGPT : A Conditional Generative Pre - Trained Transformer for Target - Specific De Novo Molecular Generation

    Ye Wang, Honggang Zhao, Simone Sciabola, and Wenlu Wang. cMolGPT : A Conditional Generative Pre - Trained Transformer for Target - Specific De Novo Molecular Generation . Molecules, 28 0 (11): 0 4430, January 2023. ISSN 1420-3049. doi:10.3390/molecules28114430. URL https://www...

  70. [78]

    SMILES , a chemical language and information system

    David Weininger. SMILES , a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28 0 (1): 0 31--36, February 1988. ISSN 0095-2338. doi:10.1021/ci00057a005. URL https://doi.org/10.10...

  71. [79]

    Efficient multi-objective molecular optimization in a continuous latent space

    Robin Winter, Floriane Montanari, Andreas Steffen, Hans Briem, Frank Noé, and Djork-Arné Clevert. Efficient multi-objective molecular optimization in a continuous latent space. Chemical Science, 10 0 (34): 0 8016--8024, August 2019. ISSN 2041-6539. doi:10.1039/C9SC01928F. URL ...

  72. [80]

    Huggingface's transformers: State-of-the-art natural language processing

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R \'e mi Louf, Morgan Funtowicz, et al. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019

  73. [81]

    MoleculeNet: a benchmark for molecular machine learning

    Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. MoleculeNet: a benchmark for molecular machine learning . Chemical science, 9 0 (2): 0 513--530, 2018

  74. [82]

    Graph neural networks for automated de novo drug design

    Jiacheng Xiong, Zhaoping Xiong, Kaixian Chen, Hualiang Jiang, and Mingyue Zheng. Graph neural networks for automated de novo drug design. Drug Discovery Today, 26 0 (6): 0 1382--1393, June 2021. ISSN 1359-6446. doi:10.1016/j.drudis.2021.02.011. URL https://www.sciencedirect.co...

  75. [83]

    Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism

    Zhaoping Xiong, Dingyan Wang, Xiaohong Liu, Feisheng Zhong, Xiaozhe Wan, Xutong Li, Zhaojun Li, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, and Mingyue Zheng. Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism . Journal of M...

  76. [84]

    How Powerful are Graph Neural Networks ? In International Conference on Learning Representations, September 2018

    Keyulu Xu*, Weihua Hu*, Jure Leskovec, and Stefanie Jegelka. How Powerful are Graph Neural Networks ? In International Conference on Learning Representations, September 2018. URL https://openreview.net/forum?id=ryGs6iA5Km

  77. [85]

    Powers, Ron O

    Minkai Xu, Alexander S. Powers, Ron O. Dror, Stefano Ermon, and Jure Leskovec. Geometric Latent Diffusion Models for 3D Molecule Generation . In Proceedings of the 40th International Conference on Machine Learning , pages 38592--38610. PMLR, July 2023. URL https://proceedings....

  78. [86]

    Hit and lead discovery with explorative rl and fragment-based molecule generation

    Soojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu, and Sung Ju Hwang. Hit and lead discovery with explorative rl and fragment-based molecule generation. Advances in Neural Information Processing Systems, 34: 0 7924--7936, 2021

  79. [87]

    Hit and lead discovery with explorative RL and fragment-based molecule generation

    Soojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu, and Sung Ju Hwang. Hit and lead discovery with explorative RL and fragment-based molecule generation. In Proceedings of the 35th International Conference on Neural Information Processing Systems , NIPS '21, pages 7924--7936, ...

  80. [88]

    Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset

    Botao Yu, Frazier N Baker, Ziqi Chen, Xia Ning, and Huan Sun. Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. In First Conference on Language Modeling, 2024

  81. [89]

    MoFlow : An Invertible Flow Model for Generating Molecular Graphs

    Chengxi Zang and Fei Wang. MoFlow : An Invertible Flow Model for Generating Molecular Graphs . In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , KDD '20, pages 617--626, New York, NY, USA, August 2020. Association for Computi...

  82. [90]

    Chuxu Zhang, Dongjin Song, Chao Huang, Ananthram Swami, and Nitesh V. Chawla. Heterogeneous Graph Neural Network . In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , KDD '19, pages 793--803, New York, NY, USA, July 2019. Assoc...

  83. [91]

    ResGen is a pocket-aware 3D molecular generation model based on parallel multiscale modelling

    Odin Zhang, Jintu Zhang, Jieyu Jin, Xujun Zhang, RenLing Hu, Chao Shen, Hanqun Cao, Hongyan Du, Yu Kang, Yafeng Deng, Furui Liu, Guangyong Chen, Chang-Yu Hsieh, and Tingjun Hou. ResGen is a pocket-aware 3D molecular generation model based on parallel multiscale modelling. Natu...

  84. [92]

    BindGPT : A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning , June 2024

    Artem Zholus, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Daniil Polykovskiy, Sarath Chandar, and Alex Zhavoronkov. BindGPT : A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning , June 2024. URL http://arxiv.org/abs/2406.03686....

  85. [93]

    Uni- Mol : A Universal 3D Molecular Representation Learning Framework

    Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. Uni- Mol : A Universal 3D Molecular Representation Learning Framework . In The Eleventh International Conference on Learning Representations, September 2022. URL https://...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.