Pith. sign in

REVIEW 4 major objections 9 minor 87 references

Graph-structured Small Molecule Drug Discovery Through Deep Learning: Progress, Challenges, and Opportunities

T0 review · 4 major / 9 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims to be the first systematic and comprehensive review of deep learning for small-molecule drug discovery, organized around graph representations and six core tasks.

desk verdict A useful six-task taxonomy for DL drug discovery, but the 'first review' claim and undocumented method/dataset selection need fixing before it goes out. read the letter →

arxiv 2502.08975 v2 pith:KB53TTOE submitted 2025-02-13 cs.LG q-bio.BM

classification cs.LGq-bio.BM
keywords graphminingmoleculerepresentationdrugdiscoveryscreeningmoleculardatasetdeeplearningdrug-targetinteractionout-of-distributiongeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to be a systematic and comprehensive review of deep learning applied to small molecule drug discovery, organized around graph representations of molecules. It identifies six core tasks—drug–target interaction/affinity, drug–cell response, drug–drug interaction, molecular property prediction, molecular generation, and molecular optimization—and shows how they relate as stages of a single discovery pipeline. The survey compiles representative methods and commonly used datasets for each task, and it highlights open problems shared across tasks: poor interpretability, weak out-of-distribution generalization, the gap between model training and wet-lab validation, and the absence of fair benchmarking. If the map is accurate, it gives new researchers a reliable entry point into the field and clarifies where methodological progress is still needed.

What carries the argument

The organizing device is the graph representation of a molecule as $G=(X,A)$, with a 3D extension $G_{3D}=(X,A,R)$, together with a task taxonomy that splits the field into single-molecule property prediction, pairwise interaction prediction (drug–target, drug–cell, drug–drug), and molecule generation/optimization. This formalization lets the survey treat the three interaction tasks as one abstract pattern, and it lets generation/optimization be written as a search for a molecule minimizing a value function. The tables of representative methods and the dataset summary provide the evidence for the survey's coverage claim.

What would settle it

One could check a random sample of the dataset counts in Table VIII against the public sources, and scan the literature for a peer-reviewed survey published before February 2025 that already organizes the same six tasks with comparable methodological tables; either finding would undermine the paper's claims of accuracy and first systematic coverage.

Watch

Extended reading notes

Core claim

The paper's central claim is that it is the first review to cover recent deep-learning advances in small molecule drug discovery in a systematic, task-organized way, spanning six tasks with a unified graph-based problem formulation. It unifies the prediction tasks (molecular property prediction; and drug–target, drug–cell, and drug–drug interactions) under a shared 'drug+X' interaction format, and it treats molecular generation and optimization as search over a chemical space with a value function. It further claims that the field's main bottlenecks are interpretability, out-of-distribution generalization, the gap between model training and laboratory validation, and the lack of standardized benchmarking. The paper does not compare method performance directly; it instead organizes methods by architecture, representation, and technique.

Load-bearing premise

The survey's usefulness rests on the assumption that the methods and datasets it selected from the literature are representative and that the dataset statistics in Table VIII are accurate; if those selections are biased or the numbers are wrong, the summary's value as a systematic guide collapses.

Editorial extensions

If this is right

  • A researcher can locate any new method in the taxonomy and see which task it addresses and which datasets are standard.
  • The unified 'drug+X' formulation suggests that techniques developed for one interaction task may transfer to the others.
  • The challenges list gives a concrete agenda: work on OOD splits, interpretability, and benchmarking would address the field's stated bottlenecks.
  • The dataset table, with counts and access URLs, allows quick selection of benchmarks.
  • The survey's lack of performance comparison implies that fair benchmarking studies remain an open need.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The task taxonomy implies a natural progression from prediction to generation; one might expect future work to train a single foundation model that handles all six tasks.
  • The four-way OOD categorization (Seen-Both, Unseen-Drug, Unseen-X, Unseen-Both) could be adopted as a standard evaluation protocol for any 'drug+X' task.
  • If the survey's coverage is representative, the recent shift toward pre-training and diffusion models suggests these will dominate near-term methods.
  • The absence of direct performance comparison among methods may reflect a field-wide problem; a community benchmark built from the listed datasets could resolve it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 9 minor

Summary. This manuscript surveys deep learning for small-molecule drug discovery, organizing the area into six tasks — drug-target interaction/affinity (DTI/DTA), drug-cell response (DRP), drug-drug interaction (DDI), molecular property prediction (MPP), molecular generation (MG), and molecular optimization (MO) — with a unified problem formulation (Eqs. 1-3), method tables (Tables I-VII), a dataset catalog (Table VIII), and a challenges section on interpretability, out-of-distribution generalization, the training/lab validation gap, and benchmarking. The paper claims in Section I to be 'the first attempt to present a systematic and comprehensive review of recent DL advancements in small molecule drug discovery,' covering roughly the past three years, and it explicitly declines to compare methods on performance in Section III-D.

Significance. The survey fills a genuinely useful niche if its catalog is trustworthy: it connects six loosely coupled tasks through a common drug-X formulation, provides an OOD scenario taxonomy (Seen-Both / Unseen-Drug / Unseen-X / Unseen-Both), includes explicit mathematical notation for generative and optimization models in Tables VI-VII, and is honest in Section III-D that no performance comparison is attempted. However, the paper's value claim rests entirely on the 'systematic and comprehensive' selection, and that machinery is currently undocumented: inclusion criteria are absent, several table entries are the authors' own preprints or conference papers, the 'first attempt' claim is made without positioning against prior surveys, and the dataset table lacks provenance for its statistics despite promising URLs. These issues are fixable in revision, but they currently undermine the reliability of the paper as a reference map for the field.

major comments (4)
  1. [Section I (Introduction) and Section III-D (Molecule Property Prediction)] The paper's central value claim is stated in Section I: it is 'the first attempt to present a systematic and comprehensive review of recent DL advancements in small molecule drug discovery.' The word 'systematic' requires a documented selection procedure, but Section III-D states only that the authors 'selected representative methods,' naming no inclusion criteria, literature sources, search dates, or performance-comparison protocol. The consequences are visible in Tables I, II, and VII: at least five entries are works of the authors' own group, including the arXiv preprints SiamDTI [8], FMOP [80], and Xiong et al. [87] alongside the authors' IJCAI-24 papers CLDR [20] and MSDA [21]. That concentration can be legitimate if justified, but as written the reader cannot tell whether the tables represent the field or the authors' own portfolio. The 'first attempt' claim also needs support: earlier surveys on drug-target interaction prediction, cancer drug response, and deep learning for drug discovery are not cited or comparatively positioned. Please add a transparent selection protocol (databases, query terms, screening criteria, inclusion window) and a positioning discussion against the closest prior surveys, or soften the centrality of the 'first attempt' claim.
  2. [Table VIII and Section IV (Datasets)] Table VIII is a core deliverable of the 'comprehensive' claim, but its provenance is undocumented. Section IV promises 'access URLs for each dataset,' yet the table contains no URLs, only footnoted descriptions; the footnote that the statistics are 'accurate as of February 4, 2025' is unverifiable without a per-row collection procedure and retrieval dates. Moreover, the 'x / y' columns are semantically undefined: the reader cannot tell which number is drugs, cell lines, pairs, or activity values. Two entries appear mislabeled: OGB-biokg[22] is a heterogeneous knowledge graph with drug-gene-disease triples, not a drug-drug classification dataset, and ZhangDDI[23] is standardly a classification task over literature-derived DDI types, not a regression task. I note that the specific CCLE entry (24 drugs / 479 cell lines) is consistent with the original CCLE drug-sensitivity screen and the NCI-60 entry (50,000 screened compounds / 60 cell lines) with the DTP screen, so my concern is not those raw numbers; the problem is the absence of definitions and sources. Please define the count semantics, provide per-row sources or stable URLs, and document the collection and verification date.
  3. [Table V (DDI) and References [57], [62]] Reference [62], cited for DDKG in Table V, has the same title, journal, volume, issue, and page range as reference [57] (the DANN-DDI paper), differing only in the publication year. These are two distinct methods, so one of the two entries is a mis-citation or a duplicated reference. Because the method tables are the survey's evidence base, the full reference list needs to be re-verified against the table entries; the reader currently cannot trust the DDI table's attribution.
  4. [Table VI caption versus Section II (Problem Formulation)] Section II defines G3D = (X, A, R), explicitly stating that X is the node feature matrix and R is the set of coordinates. The caption of Table VI inverts this: 'For a 3D molecule, it can be represented as point clouds (X, R), where X is the atom coordinates matrix and R is the node feature matrix.' This direct contradiction in notation appears in the notation block intended to make the generation-formula table self-contained and should be corrected in one of the two places.
minor comments (9)
  1. [Section I (Introduction)] The statement that small-molecule drugs account for 'about 98% of the total number of commonly used drugs' is asserted without a citation; either add a source or delete the figure.
  2. [Section I (Introduction)] The sentence 'the representation methods for these objects are specific and vary considerably, customized approaches' is ungrammatical and should be rewritten.
  3. [Section II (Problem Formulation)] The phrase 'It is can be formulated as' contains a duplicated auxiliary verb.
  4. [Section III-C (Drug-Drug Interaction Prediction)] The phrase 'contribute significantly contribute to the drug discovery field' contains a duplicated word and should be corrected.
  5. [Abstract and Section I] The scope is described as 'recent years' in the abstract and 'the past three years' in Section I, but the tables include foundational work from 2017-2018 (e.g., JTVAE [74] and DeepDDI [61]); please define the inclusion window explicitly and apply it consistently.
  6. [Table VIII] The entry 'OBG-biogk' should be 'OGB-biokg,' and the sub-header 'Drugs / Total' is unclear for rows that report drug-pair counts rather than drug counts.
  7. [Section V-b] The four-scenario out-of-distribution taxonomy is repeated nearly verbatim from Section III-A; cross-reference the earlier definition instead of duplicating it.
  8. [Table VI] The notation paragraph under Table VI is extremely dense, running to roughly a page of definitions; moving it to an appendix or supplement would substantially improve readability.
  9. [Figure 1] Several labels in Figure 1 (task icons and representation panels) are difficult to read at the current resolution; a higher-resolution version is needed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey makes no derived predictions; representative-method selection is descriptive, not a fitted input renamed as a result.

full rationale

This paper is a survey, not a derivation chain: it organizes existing methods and datasets into six task categories and summarizes trends. There is no equation whose output is defined by its input, no parameter fitted to a subset and then relabeled as a prediction, and no uniqueness theorem imported from the authors' prior work to force a choice. The closest potential issue is that several 'representative methods' in Tables I, II, and VII are the authors' own papers or preprints (e.g., SiamDTI [8], CLDR [20], MSDA [21], FMOP [80], text-guided diffusion [87]). However, the paper explicitly states in Section III-D that it 'selected representative methods and provided detailed descriptions of molecular representation and feature encoding techniques without comparing performance,' so these entries function as descriptive examples rather than as evidence that is reduced to the authors' own claims. The central assertion of being 'the first attempt to present a systematic and comprehensive review' is a bibliographic and scope claim, not a result derived from the cited literature; whether that claim is accurate is a correctness or completeness question, not a circularity one. No step in the paper's reasoning reduces by construction to its own inputs, so no circularity is found.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities appear because this is a review, not a new model. The axioms are domain assumptions about task scope and representativeness.

assumptions (3)
  • domain assumption The six selected tasks (DTI/DTA, DRP, DDI, MPP, MG, MO) are the core of small molecule drug discovery.
    Section I declares these tasks as core and interconnected without a systematic justification or bibliometric basis.
  • domain assumption Graph-based molecular representations are inherently advantageous for small molecule drug discovery.
    Section I states graph representations offer 'inherent advantages' without comparative evidence; this frames the entire review.
  • ad hoc to paper The cited methods are representative of each task's landscape.
    Tables I-VII select methods with no inclusion criteria; several entries are the authors' own works, including preprints, which may skew representativeness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Graph-structured Small Molecule Drug Discovery Through Deep Learning: Progress, Challenges, and Opportunities." pith.science (2026). https://pith.science/paper/KB53TTOE

@misc{pith2026250208975,
  author       = {Pith},
  title        = {Pith review of: Graph-structured Small Molecule Drug Discovery Through Deep Learning: Progress, Challenges, and Opportunities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KB53TTOE}},
  note         = {Machine review of arXiv:2502.08975}
}
read the original abstract

Due to their excellent drug-like and pharmacokinetic properties, small molecule drugs are widely used to treat various diseases, making them a critical component of drug discovery. In recent years, with the rapid development of deep learning (DL) techniques, DL-based small molecule drug discovery methods have achieved excellent performance in prediction accuracy, speed, and complex molecular relationship modeling compared to traditional machine learning approaches. These advancements enhance drug screening efficiency and optimization and provide more precise and effective solutions for various drug discovery tasks. Contributing to this field's development, this paper aims to systematically summarize and generalize the recent key tasks and representative techniques in graph-structured small molecule drug discovery in recent years. Specifically, we provide an overview of the major tasks in small molecule drug discovery and their interrelationships. Next, we analyze the six core tasks, summarizing the related methods, commonly used datasets, and technological development trends. Finally, we discuss key challenges, such as interpretability and out-of-distribution generalization, and offer our insights into future research directions for small molecule drug discovery.

Figures

Figures reproduced from arXiv: 2502.08975 by the authors.

Figure 1
Figure 1. Molecular representation methods and their drug discovery applications. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 74 canonical work pages

  1. [8]

    A cross-field fusion strategy for drug-target interaction prediction,

    H. Zhang, X. Gong, S. Pan, J. Wu, B. Du, and W. Hu, “A cross-field fusion strategy for drug-target interaction prediction,” arXiv preprint arXiv:2405.14545, 2024

  2. [20]

    Contrastive learning drug response models from natural language supervision,

    K. Li, X. Gong, J. Wu, and W. Hu, “Contrastive learning drug response models from natural language supervision,” in Proceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, IJCAI-24, pp. 2126–2134, 2024

  3. [21]

    Zero-shot learning for preclinical drug screening,

    K. Li, W. Liu, Y . Luo, X. Cai, J. Wu, and W. Hu, “Zero-shot learning for preclinical drug screening,” in Proceedings of the Thirty-Third Interna- tional Joint Conference on Artificial Intelligence, IJCAI-24 , pp. 2117– 2125, 2024

  4. [80]

    Fragment-Masked Diffusion for Molecular Optimization

    K. Li, X. Cai, J. Wu, B. Du, and W. Hu, “Fragment-masked molecular optimization,” arXiv preprint arXiv:2408.09106 , 2024

  5. [87]

    Text-guided multi-property molecular optimization with a diffusion language model,

    Y . Xiong, K. Li, W. Liu, J. Wu, B. Du, S. Pan, and W. Hu, “Text-guided multi-property molecular optimization with a diffusion language model,” arXiv preprint arXiv:2410.13597 , 2024

  6. [22]

    A context-aware deconfounding autoencoder for robust prediction of personalized clinical drug response from cell-line compound screening,

    D. He, Q. Liu, Y . Wu, and L. Xie, “A context-aware deconfounding autoencoder for robust prediction of personalized clinical drug response from cell-line compound screening,”Nature Machine Intelligence, vol. 4, no. 10, pp. 879–892, 2022

  7. [23]

    Deep transfer learning of cancer drug responses by integrating bulk and single-cell rna-seq data,

    J. Chen, X. Wang, A. Ma, Q.-E. Wang, B. Liu, L. Li, D. Xu, and Q. Ma, “Deep transfer learning of cancer drug responses by integrating bulk and single-cell rna-seq data,” Nature Communications , vol. 13, no. 1, p. 6494, 2022

  8. [62]

    Enhancing drug-drug interaction prediction using deep attention neural networks,

    S. Liu, Y . Zhang, Y . Cui, Y . Qiu, Y . Deng, Z. Zhang, and W. Zhang, “Enhancing drug-drug interaction prediction using deep attention neural networks,” IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 20, no. 2, pp. 976–985, 2023

  9. [57]

    Enhancing drug-drug interaction prediction using deep attention neu- ral networks,

    S. Liu, Y . Zhang, Y . Cui, Y . Qiu, Y . Deng, Z. Zhang, and W. Zhang, “Enhancing drug-drug interaction prediction using deep attention neu- ral networks,” IEEE/ACM transactions on computational biology and Bioinformatics, vol. 20, no. 2, pp. 976–985, 2022

Show all 87 references
  1. [1]

    Drugclip: Contrasive protein-molecule representation learning for virtual screening,

    B. Gao, B. Qiang, H. Tan, Y . Jia, M. Ren, M. Lu, J. Liu, W.-Y . Ma, and Y . Lan, “Drugclip: Contrasive protein-molecule representation learning for virtual screening,” Advances in Neural Information Processing Systems, vol. 36, 2024

  2. [2]

    Uni-mol: A universal 3d molecular representation learning framework,

    G. Zhou, Z. Gao, Q. Ding, H. Zheng, H. Xu, Z. Wei, L. Zhang, and G. Ke, “Uni-mol: A universal 3d molecular representation learning framework,” in The Eleventh International Conference on Learning Representations, 2023

  3. [3]

    Adapting protein language models for rapid dti prediction,

    S. Sledzieski, R. Singh, L. Cowen, and B. Berger, “Adapting protein language models for rapid dti prediction,” bioRxiv, pp. 2022–11, 2022

  4. [4]

    Drug–target interaction predication via multi-channel graph neural networks,

    Y . Li, G. Qiao, K. Wang, and G. Wang, “Drug–target interaction predication via multi-channel graph neural networks,” Briefings in Bioinformatics, vol. 23, no. 1, p. bbab346, 2022

  5. [5]

    Hyperattentiondti: improving drug–protein interaction prediction by sequence-based deep learning with attention mechanism,

    Q. Zhao, H. Zhao, K. Zheng, and J. Wang, “Hyperattentiondti: improving drug–protein interaction prediction by sequence-based deep learning with attention mechanism,” Bioinformatics, vol. 38, no. 3, pp. 655–662, 2022

  6. [6]

    Predicting drug–protein interaction using quasi-visual question answering system,

    S. Zheng, Y . Li, S. Chen, J. Xu, and Y . Yang, “Predicting drug–protein interaction using quasi-visual question answering system,” Nature Ma- chine Intelligence, vol. 2, no. 2, pp. 134–140, 2020

  7. [7]

    Moltrans: molecular interaction transformer for drug–target interaction prediction,

    K. Huang, C. Xiao, L. M. Glass, and J. Sun, “Moltrans: molecular interaction transformer for drug–target interaction prediction,” Bioinfor- matics, vol. 37, no. 6, pp. 830–836, 2021

  8. [9]

    Dtiam: a unified framework for predicting drug- target interactions, binding affinities and drug mechanisms,

    Z. Lu, G. Song, H. Zhu, C. Lei, X. Sun, K. Wang, L. Qin, Y . Chen, J. Tang, and M. Li, “Dtiam: a unified framework for predicting drug- target interactions, binding affinities and drug mechanisms,” Nature Communications, vol. 16, no. 1, p. 2548, 2025

  9. [10]

    Interpretable bilinear atten- tion network with domain adaptation improves drug–target prediction,

    P. Bai, F. Miljkovi ´c, B. John, and H. Lu, “Interpretable bilinear atten- tion network with domain adaptation improves drug–target prediction,” Nature Machine Intelligence , vol. 5, no. 2, pp. 126–136, 2023

  10. [11]

    Psc-cpi: Multi-scale protein sequence-structure contrasting for efficient and generalizable compound-protein interaction prediction,

    L. Wu, Y . Huang, C. Tan, Z. Gao, B. Hu, H. Lin, Z. Liu, and S. Z. Li, “Psc-cpi: Multi-scale protein sequence-structure contrasting for efficient and generalizable compound-protein interaction prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol....

  11. [12]

    Cross-modality and self-supervised protein em- bedding for compound–protein affinity and contact prediction,

    Y . You and Y . Shen, “Cross-modality and self-supervised protein em- bedding for compound–protein affinity and contact prediction,” Bioin- formatics, vol. 38, no. Supplement 2, pp. ii68–ii74, 2022

  12. [13]

    Mgndti: A drug-target interaction prediction framework based on multimodal representation learning and the gating mechanism,

    L. Peng, X. Liu, M. Chen, W. Liao, J. Mao, and L. Zhou, “Mgndti: A drug-target interaction prediction framework based on multimodal representation learning and the gating mechanism,” Journal of Chemical Information and Modeling , vol. 64, no. 16, pp. 6684–6698, 2024

  13. [14]

    Perceiver cpi: a nested cross-attention network for compound–protein interaction prediction,

    N.-Q. Nguyen, G. Jang, H. Kim, and J. Kang, “Perceiver cpi: a nested cross-attention network for compound–protein interaction prediction,” Bioinformatics, vol. 39, no. 1, p. btac731, 2023

  14. [15]

    Deeptta: a transformer-based model for predicting cancer drug response,

    L. Jiang, C. Jiang, X. Yu, R. Fu, S. Jin, and X. Liu, “Deeptta: a transformer-based model for predicting cancer drug response,” Briefings in Bioinformatics, vol. 23, no. 3, p. bbac100, 2022

  15. [16]

    A subcomponent-guided deep learning method for interpretable cancer drug response prediction,

    X. Liu and W. Zhang, “A subcomponent-guided deep learning method for interpretable cancer drug response prediction,” PLOS Computational Biology, vol. 19, no. 8, p. e1011382, 2023

  16. [17]

    Tgsa: protein–protein association-based twin graph neural networks for drug response prediction with similarity augmentation,

    Y . Zhu, Z. Ouyang, W. Chen, R. Feng, D. Z. Chen, J. Cao, and J. Wu, “Tgsa: protein–protein association-based twin graph neural networks for drug response prediction with similarity augmentation,” Bioinformatics, vol. 38, no. 2, pp. 461–468, 2022

  17. [18]

    Improving drug response prediction via integrating gene relationships with deep learning,

    P. Li, Z. Jiang, T. Liu, X. Liu, H. Qiao, and X. Yao, “Improving drug response prediction via integrating gene relationships with deep learning,” Briefings in Bioinformatics , vol. 25, no. 3, p. bbae153, 2024

  18. [19]

    Graphcdr: a graph neural network method with contrastive learning for cancer drug response prediction,

    X. Liu, C. Song, F. Huang, H. Fu, W. Xiao, and W. Zhang, “Graphcdr: a graph neural network method with contrastive learning for cancer drug response prediction,” Briefings in Bioinformatics , vol. 23, no. 1, p. bbab457, 2022

  19. [24]

    Wiser: Weak supervision and supervised representation learning to improve drug response prediction in cancer,

    K. Shubham, A. Jayagopal, S. M. Danish, P. AP, and V . Rajan, “Wiser: Weak supervision and supervised representation learning to improve drug response prediction in cancer,” arXiv preprint arXiv:2405.04078 , 2024

  20. [25]

    Geometry-enhanced molecular representation learning for property prediction,

    X. Fang, L. Liu, J. Lei, D. He, S. Zhang, J. Zhou, F. Wang, H. Wu, and H. Wang, “Geometry-enhanced molecular representation learning for property prediction,” Nature Machine Intelligence , vol. 4, no. 2, pp. 127–134, 2022

  21. [26]

    Sliced denoising: A physics-informed molecular pre-training method,

    Y . Ni, S. Feng, W.-Y . Ma, Z.-M. Ma, and Y . Lan, “Sliced denoising: A physics-informed molecular pre-training method,” in The Twelfth International Conference on Learning Representations , 2024

  22. [27]

    Pre-training molecular graph representation with 3d geometry,

    S. Liu, H. Wang, W. Liu, J. Lasenby, H. Guo, and J. Tang, “Pre-training molecular graph representation with 3d geometry,” in International Conference on Learning Representations , 2022

  23. [28]

    One transformer can understand both 2d & 3d molecular data,

    S. Luo, T. Chen, Y . Xu, S. Zheng, T.-Y . Liu, L. Wang, and D. He, “One transformer can understand both 2d & 3d molecular data,” in The Eleventh International Conference on Learning Representations , 2023

  24. [29]

    Mole: a founda- tion model for molecular graphs using disentangled attention,

    O. M ´endez-Lucio, C. A. Nicolaou, and B. Earnshaw, “Mole: a founda- tion model for molecular graphs using disentangled attention,” Nature Communications, vol. 15, no. 1, p. 9431, 2024

  25. [30]

    Molspectra: Pre- training 3d molecular representation with multi-modal energy spectra,

    L. Wang, S. Liu, Y . Rong, D. Zhao, Q. Liu, and S. Wu, “Molspectra: Pre- training 3d molecular representation with multi-modal energy spectra,” in The Thirteenth International Conference on Learning Representa- tions, 2025

  26. [31]

    Multi-channel learning for integrating structural hierarchies into context-dependent molecular representation,

    Y . Wan, J. Wu, T. Hou, C.-Y . Hsieh, and X. Jia, “Multi-channel learning for integrating structural hierarchies into context-dependent molecular representation,” Nature Communications, vol. 16, no. 1, p. 413, 2025

  27. [32]

    Data-driven quantum chemical property prediction leveraging 3d conformations with uni- mol+,

    S. Lu, Z. Gao, D. He, L. Zhang, and G. Ke, “Data-driven quantum chemical property prediction leveraging 3d conformations with uni- mol+,” Nature communications, vol. 15, no. 1, p. 7104, 2024

  28. [33]

    Molecular con- trastive learning of representations via graph neural networks,

    Y . Wang, J. Wang, Z. Cao, and A. Barati Farimani, “Molecular con- trastive learning of representations via graph neural networks,” Nature Machine Intelligence, vol. 4, no. 3, pp. 279–287, 2022

  29. [34]

    Mul- timodal molecular pretraining via modality blending,

    Q. Yu, Y . Zhang, Y . Ni, S. Feng, Y . Lan, H. Zhou, and J. Liu, “Mul- timodal molecular pretraining via modality blending,” in The Twelfth International Conference on Learning Representations , 2024

  30. [35]

    Learning topology-specific experts for molecular property prediction,

    S. Kim, D. Lee, S. Kang, S. Lee, and H. Yu, “Learning topology-specific experts for molecular property prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, pp. 8291–8299, 2023

  31. [36]

    Exploring molecular pretraining model at scale,

    X. Ji, Z. Wang, Z. Gao, H. Zheng, L. Zhang, G. Ke, and W. E, “Exploring molecular pretraining model at scale,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024

  32. [37]

    Graphmae: Self-supervised masked graph autoencoders,

    Z. Hou, X. Liu, Y . Cen, Y . Dong, H. Yang, C. Wang, and J. Tang, “Graphmae: Self-supervised masked graph autoencoders,” in Proceed- ings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , pp. 594–604, 2022

  33. [38]

    Triplet interaction improves graph transformers: accurate molecular graph learning with triplet graph transformers,

    M. S. Hussain, M. J. Zaki, and D. Subramanian, “Triplet interaction improves graph transformers: accurate molecular graph learning with triplet graph transformers,” in Proceedings of the 41st International Conference on Machine Learning , ICML’24, JMLR.org, 2024

  34. [39]

    Instructor-inspired machine learning for robust molecular property prediction,

    F. Wu, S. Jin, S. Li, and S. Z. Li, “Instructor-inspired machine learning for robust molecular property prediction,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024

  35. [40]

    Spherical message passing for 3d molecular graphs,

    Y . Liu, L. Wang, M. Liu, Y . Lin, X. Zhang, B. Oztekin, and S. Ji, “Spherical message passing for 3d molecular graphs,” in International Conference on Learning Representations , 2022

  36. [41]

    Molecular set representation learning,

    M. Boulougouri, P. Vandergheynst, and D. Probst, “Molecular set representation learning,” Nature Machine Intelligence , vol. 6, no. 7, pp. 754–763, 2024

  37. [42]

    Comenet: Towards complete and efficient message passing for 3d molecular graphs,

    L. Wang, Y . Liu, Y . Lin, H. Liu, and S. Ji, “Comenet: Towards complete and efficient message passing for 3d molecular graphs,” Advances in Neural Information Processing Systems , vol. 35, pp. 650–664, 2022

  38. [43]

    Expressivity and generalization: Fragment-biases for molecular gnns,

    T. Wollschl ¨ager, N. Kemper, L. Hetzel, J. Sommer, and S. G ¨unnemann, “Expressivity and generalization: Fragment-biases for molecular gnns,” in International Conference on Machine Learning , pp. 53113–53139, PMLR, 2024

  39. [44]

    Equivariant transformers for neural network based molecular potentials,

    P. Th ¨olke and G. D. Fabritiis, “Equivariant transformers for neural network based molecular potentials,” in International Conference on Learning Representations, 2022

  40. [45]

    Representing molecules as random walks over interpretable grammars,

    M. Sun, M. Guo, W. Yuan, V . Thost, C. E. Owens, A. F. Grosz, S. Selvan, K. Zhou, H. Mohiuddin, B. J. Pedretti, et al., “Representing molecules as random walks over interpretable grammars,” in International Conference on Machine Learning , pp. 46988–47016, PMLR, 2024

  41. [46]

    A theoretically-principled sparse, connected, and rigid graph representation of molecules,

    S.-H. Wang, Y . Huang, J. M. Baker, Y .-E. Sun, Q. Tang, and B. Wang, “A theoretically-principled sparse, connected, and rigid graph representation of molecules,” in The Thirteenth International Conference on Learning Representations, 2025

  42. [47]

    Hierarchical grammar-induced geometry for data-efficient molecular property prediction,

    M. Guo, V . Thost, S. W. Song, A. Balachandran, P. Das, J. Chen, and W. Matusik, “Hierarchical grammar-induced geometry for data-efficient molecular property prediction,” in International Conference on Machine Learning, pp. 12055–12076, PMLR, 2023

  43. [48]

    Equiformer: Equivariant graph attention transformer for 3d atomistic graphs,

    Y .-L. Liao and T. Smidt, “Equiformer: Equivariant graph attention transformer for 3d atomistic graphs,” in The Eleventh International Conference on Learning Representations , 2023

  44. [49]

    Graph sampling-based meta-learning for molecular property prediction,

    X. Zhuang, Q. Zhang, B. Wu, K. Ding, Y . Fang, and H. Chen, “Graph sampling-based meta-learning for molecular property prediction,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, pp. 4729–4737, 2023

  45. [50]

    Enhancing geometric representations for molecules with equivariant vector-scalar interactive message passing,

    Y . Wang, T. Wang, S. Li, X. He, M. Li, Z. Wang, N. Zheng, B. Shao, and T.-Y . Liu, “Enhancing geometric representations for molecules with equivariant vector-scalar interactive message passing,” Nature Commu- nications, vol. 15, no. 1, p. 313, 2024

  46. [51]

    Efficient sharpness- aware minimization for molecular graph transformer models,

    Y . Wang, K. Zhou, N. Liu, Y . Wang, and X. Wang, “Efficient sharpness- aware minimization for molecular graph transformer models,” in The Twelfth International Conference on Learning Representations , 2024

  47. [52]

    Zeroddi: A zero-shot drug-drug interaction event prediction method with semantic enhanced learning and dual-modal uniform alignment,

    Z. Wang, Z. Xiong, F. Huang, X. Liu, and W. Zhang, “Zeroddi: A zero-shot drug-drug interaction event prediction method with semantic enhanced learning and dual-modal uniform alignment,” arXiv preprint arXiv:2407.00891, 2024

  48. [53]

    Mkg-fenn: A multimodal knowledge graph fused end-to-end neural network for accurate drug– drug interaction prediction,

    D. Wu, W. Sun, Y . He, Z. Chen, and X. Luo, “Mkg-fenn: A multimodal knowledge graph fused end-to-end neural network for accurate drug– drug interaction prediction,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 10216–10224, 2024

  49. [54]

    Bi-level graph neural networks for drug-drug interaction prediction,

    Y . Bai, K. Gu, Y . Sun, and W. Wang, “Bi-level graph neural networks for drug-drug interaction prediction,” arXiv preprint arXiv:2006.14002 , 2020

  50. [55]

    Conditional graph information bottleneck for molecular relational learning,

    N. Lee, D. Hyun, G. S. Na, S. Kim, J. Lee, and C. Park, “Conditional graph information bottleneck for molecular relational learning,” in Inter- national Conference on Machine Learning , pp. 18852–18871, PMLR, 2023

  51. [56]

    Phgl-ddi: A pre-training based hierarchical graph learning framework for drug-drug interaction prediction,

    Y . Yuan, J. Yue, R. Zhang, and W. Su, “Phgl-ddi: A pre-training based hierarchical graph learning framework for drug-drug interaction prediction,” Expert Systems with Applications , p. 126408, 2025

  52. [58]

    Dual-channel learning framework for drug-drug interaction prediction via relation- aware heterogeneous graph transformer,

    X. Su, P. Hu, Z.-H. You, S. Y . Philip, and L. Hu, “Dual-channel learning framework for drug-drug interaction prediction via relation- aware heterogeneous graph transformer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, pp. 249–256, 2024

  53. [59]

    Gognn: Graph of graphs neural network for predicting structured entity interactions,

    H. Wang, D. Lian, Y . Zhang, L. Qin, and X. Lin, “Gognn: Graph of graphs neural network for predicting structured entity interactions,” arXiv preprint arXiv:2005.05537 , 2020

  54. [60]

    Csgnn: Contrastive self-supervised graph neural network for molecular interaction predic- tion.,

    C. Zhao, S. Liu, F. Huang, S. Liu, and W. Zhang, “Csgnn: Contrastive self-supervised graph neural network for molecular interaction predic- tion.,” in IJCAI, pp. 3756–3763, 2021

  55. [61]

    Deep learning improves prediction of drug–drug and drug–food interactions,

    J. Y . Ryu, H. U. Kim, and S. Y . Lee, “Deep learning improves prediction of drug–drug and drug–food interactions,” Proceedings of the national academy of sciences , vol. 115, no. 18, pp. E4304–E4311, 2018

  56. [63]

    3dlinker: An e (3) equivariant variational autoencoder for molecular linker design,

    Y . Huang, X. Peng, J. Ma, and M. Zhang, “3dlinker: An e (3) equivariant variational autoencoder for molecular linker design,” in International Conference on Machine Learning , pp. 9280–9294, PMLR, 2022

  57. [64]

    Gf-vae: a flow-based variational autoencoder for molecule generation,

    C. Ma and X. Zhang, “Gf-vae: a flow-based variational autoencoder for molecule generation,” in Proceedings of the 30th ACM international conference on information & knowledge management , pp. 1181–1190, 2021

  58. [65]

    Molhf: a hierarchical normalizing flow for molecular graph generation,

    Y . Zhu, Z. Ouyang, B. Liao, J. Wu, Y . Wu, C.-Y . Hsieh, T. Hou, and J. Wu, “Molhf: a hierarchical normalizing flow for molecular graph generation,” in Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , pp. 5002–5010, 2023

  59. [66]

    Learning neural generative dynamics for molecular conformation generation,

    M. Xu, S. Luo, Y . Bengio, J. Peng, and J. Tang, “Learning neural generative dynamics for molecular conformation generation,” arXiv preprint arXiv:2102.10240, 2021

  60. [67]

    A de novo molecular generation method using latent vector based generative adversarial network,

    O. Prykhodko, S. V . Johansson, P.-C. Kotsias, J. Ar ´us-Pous, E. J. Bjer- rum, O. Engkvist, and H. Chen, “A de novo molecular generation method using latent vector based generative adversarial network,” Journal of Cheminformatics, vol. 11, pp. 1–13, 2019

  61. [68]

    Transformer- based objective-reinforced generative adversarial network to generate desired molecules,

    C. Li, C. Yamanaka, K. Kaitoh, and Y . Yamanishi, “Transformer- based objective-reinforced generative adversarial network to generate desired molecules,” in Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 (L. D. Raedt, ed.), ...

  62. [69]

    Score-based generative modeling of graphs via the system of stochastic differential equations,

    J. Jo, S. Lee, and S. J. Hwang, “Score-based generative modeling of graphs via the system of stochastic differential equations,” in Interna- tional conference on machine learning, pp. 10362–10383, PMLR, 2022

  63. [70]

    Geometric latent diffusion models for 3d molecule generation,

    M. Xu, A. S. Powers, R. O. Dror, S. Ermon, and J. Leskovec, “Geometric latent diffusion models for 3d molecule generation,” in International Conference on Machine Learning , pp. 38592–38610, PMLR, 2023

  64. [71]

    Exploring chemical space with score- based out-of-distribution generation,

    S. Lee, J. Jo, and S. J. Hwang, “Exploring chemical space with score- based out-of-distribution generation,” in International Conference on Machine Learning, pp. 18872–18892, PMLR, 2023

  65. [72]

    Decompdiff: Diffusion models with decomposed priors for structure-based drug design,

    J. Guan, X. Zhou, Y . Yang, Y . Bao, J. Peng, J. Ma, Q. Liu, L. Wang, and Q. Gu, “Decompdiff: Diffusion models with decomposed priors for structure-based drug design,” in International Conference on Machine Learning, pp. 11827–11846, PMLR, 2023

  66. [73]

    Graph diffusion transformers for multi-conditional molecular generation,

    G. Liu, J. Xu, T. Luo, and M. Jiang, “Graph diffusion transformers for multi-conditional molecular generation,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024

  67. [74]

    Junction tree variational autoen- coder for molecular graph generation,

    W. Jin, R. Barzilay, and T. Jaakkola, “Junction tree variational autoen- coder for molecular graph generation,” in International conference on machine learning, pp. 2323–2332, PMLR, 2018

  68. [75]

    Fflom: A flow-based autoregressive model for fragment- to-lead optimization,

    J. Jin, D. Wang, G. Shi, J. Bao, J. Wang, H. Zhang, P. Pan, D. Li, X. Yao, H. Liu, et al., “Fflom: A flow-based autoregressive model for fragment- to-lead optimization,” Journal of Medicinal Chemistry , vol. 66, no. 15, pp. 10808–10823, 2023

  69. [76]

    Leveraging language model for advanced multiproperty molecular optimization via prompt engineering,

    Z. Wu, O. Zhang, X. Wang, L. Fu, H. Zhao, J. Wang, H. Du, D. Jiang, Y . Deng, D. Cao, et al. , “Leveraging language model for advanced multiproperty molecular optimization via prompt engineering,” Nature Machine Intelligence, pp. 1–11, 2024

  70. [77]

    Mol-cyclegan: a generative model for molecular optimization,

    Ł. Maziarka, A. Pocha, J. Kaczmarczyk, K. Rataj, T. Danel, and M. War- choł, “Mol-cyclegan: a generative model for molecular optimization,” Journal of Cheminformatics , vol. 12, no. 1, p. 2, 2020

  71. [78]

    Decom- popt: Controllable and decomposed diffusion models for structure-based molecular optimization,

    X. Zhou, X. Cheng, Y . Yang, Y . Bao, L. Wang, and Q. Gu, “Decom- popt: Controllable and decomposed diffusion models for structure-based molecular optimization,” in The Twelfth International Conference on Learning Representations, 2024

  72. [79]

    A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets,

    L. Huang, T. Xu, Y . Yu, P. Zhao, X. Chen, J. Han, Z. Xie, H. Li, W. Zhong, K.-C. Wong, et al. , “A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets,” Nature Communications, vol. 15, no. 1, p. 2657, 2024

  73. [81]

    Mars: Markov molecular sampling for multi-objective drug discovery,

    Y . Xie, C. Shi, H. Zhou, Y . Yang, W. Zhang, Y . Yu, and L. Li, “Mars: Markov molecular sampling for multi-objective drug discovery,” in International Conference on Learning Representations , 2021

  74. [82]

    Molecule optimization by explainable evolution,

    B. Chen, T. Wang, C. Li, H. Dai, and L. Song, “Molecule optimization by explainable evolution,” in International conference on learning representation (ICLR), 2021

  75. [83]

    Molsearch: search-based multi-objective molecular generation and property opti- mization,

    M. Sun, J. Xing, H. Meng, H. Wang, B. Chen, and J. Zhou, “Molsearch: search-based multi-objective molecular generation and property opti- mization,” in Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining , pp. 4724–4732, 2022

  76. [84]

    Dif- ferentiable scaffolding tree for molecule optimization,

    T. Fu, W. Gao, C. Xiao, J. Yasonik, C. W. Coley, and J. Sun, “Dif- ferentiable scaffolding tree for molecule optimization,” in International Conference on Learning Representations , 2022

  77. [85]

    Sample-efficient multi-objective molecular optimization with gflownets,

    Y . Zhu, J. Wu, C. Hu, J. Yan, T. Hou, J. Wu, et al. , “Sample-efficient multi-objective molecular optimization with gflownets,” Advances in Neural Information Processing Systems , vol. 36, 2024

  78. [86]

    Dynamic many-objective molecular optimization: Unfolding complexity with ob- jective decomposition and progressive optimization,

    D.-H. Shin, Y .-H. Son, D.-J. Lee, J.-W. Han, and T.-E. Kam, “Dynamic many-objective molecular optimization: Unfolding complexity with ob- jective decomposition and progressive optimization,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intel...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.