Pith. sign in

REVIEW 4 major objections 4 minor 174 references

Investigating Graph Neural Networks and Classical Feature-Extraction Techniques in Activity-Cliff and Molecular Property Prediction

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This thesis argues that replacing hash-based folding with frequency-ranked substructure selection improves ECFPs and is entropy-optimal under stated assumptions.

arxiv 2411.13688 v1 pith:CLGQS6I3 submitted 2024-11-20 cs.LG q-bio.BMstat.ML

classification cs.LGq-bio.BMstat.ML MSC 68T0768T05
keywords extended-connectivityfingerprintssubstructurepoolingSort&Slicehashingmolecularpropertypredictionactivity-cliffgraphneuralnetworksQSAR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

At its core this thesis asks whether trainable graph neural networks actually beat classical molecular featurisations, and it answers with a qualified no for ordinary QSAR prediction while pointing to where each method wins. The central methodological contribution is Sort & Slice, a change to how extended-connectivity fingerprints (ECFPs) are built: instead of hashing every detected circular substructure into a fixed bit vector, sort the substructures by how often they occur in the training set and keep only the most frequent ones. The thesis proves that, under assumptions it states, this selection is entropy-optimal, and its experiments show it consistently outperforms hashing across data sets, regressors, splits, and ECFP hyperparameters. If true, a practically free modification improves both accuracy and interpretability of one of the most widely used molecular representations.

What carries the argument

Substructure pooling is the general operation that maps a multiset of enumerated circular substructures, the unordered output of an ECFP-style enumeration, to a real-valued vector; hashing is one instance, and Sort & Slice is another. Sort & Slice ranks the substructures by training-set frequency and truncates to the most frequent ones. This frequency ranking is also the load-bearing mechanism of the proof: with the stated assumption that substructure frequency tracks informativeness, the top-frequency slice is shown to be entropy-optimal, meaning no other fixed-size selection of substructures can carry more Shannon information about the target. The thesis also develops a pair-based data-splitting scheme and a twin neural network whose max-pooling and odd-MLP design hard-code order-invariance for activity-cliff labels and order-equivariance for potency-direction labels.

What would settle it

Construct a property-prediction benchmark whose target is determined almost entirely by a deliberately rare substructure present in under 1% of training molecules, then compare Sort & Slice with hashed ECFPs on that benchmark: if hashing wins, the frequency-informativeness premise fails, since Sort & Slice is designed to drop rare substructures.

Watch

Extended reading notes

Core claim

On the paper's own terms, the key discovery is that the hashing step in standard ECFP vectorisation is not a neutral technical detail: it is a lossy, collision-prone form of substructure pooling, and it can be replaced by a simpler frequency-ranked procedure that works better. Sort & Slice first sorts the enumerated circular substructures of a molecule according to their frequency in the training set and then keeps the most frequent substructures in the final vector, so each retained dimension corresponds to one concrete chemical substructure with no bit collisions. The accompanying mathematical argument shows that under the stated frequency-informativeness assumption, keeping the most frequent substructures is exactly the choice that maximises information about the target in an entropic sense. Empirically, the thesis reports that this simple pooling technique robustly beats hashing and two supervised substructure-selection schemes for molecular property prediction, with the advantage growing as the expected number of hash collisions increases. For the field's broader question, the thesis finds ECFPs still deliver the best QSAR predictions, while GIN-based features are competitive or better for activity-cliff classification.

Load-bearing premise

The proof depends on training-set substructure frequency being a faithful proxy for how informative a substructure is about the target, and on those frequencies persisting at test time; if a rare substructure carries the signal, Sort & Slice discards exactly the feature the model needs.

Editorial extensions

If this is right

  • Molecular-property pipelines can replace hashed ECFPs with Sort & Slice at no extra model cost and obtain higher predictive performance, with each retained bit tied to one explicit substructure.
  • The gap between Sort & Slice and hashing should widen when fingerprints are shorter or radii larger, because those settings raise the expected number of hash collisions.
  • For QSAR prediction, classical ECFP features remain at least as strong as the tested GNN features, so claims that message-passing GNNs categorically supersede fingerprints need to be conditioned on task and data.
  • For activity-cliff classification, GIN-based features and a twin network architecture can outperform repurposed QSAR baselines, giving practical baselines for future AC-prediction studies.
  • When the activity of one compound in a matched pair is known, QSAR models detect cliffs much better than when both activities are unknown, so cliff-prediction performance should be reported separately for these two settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The entropy-optimality result implies a testable design rule: choose the fingerprint dimension as the number of top-frequency substructures that cover most of the training-set mass, which could let practitioners set ECFP length without grid search.
  • Because the frequency ranking is task-agnostic, a natural extension is to make the selection differentiable, for instance with self-attention weights, so the model learns which frequency bands matter rather than assuming frequency equals informativeness.
  • The sharp activity-cliff sensitivity drop in the both-activities-unknown setting suggests that QSAR errors concentrate on the most informative pairs; a training loss that explicitly rewards sensitivity on predicted cliffs might improve models more than adding featurisation capacity.
  • The twin network's hard-coded symmetry properties transfer to other pairwise chemistry problems, such as predicting reaction outcomes or matched-molecular-pair transformations, where label reversal rules are known a priori.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript is a PhD thesis that investigates molecular featurisation methods for two tasks: quantitative structure-activity relationship (QSAR) prediction and activity-cliff (AC) prediction. Chapter 2 reviews physicochemical-descriptor vectors (PDVs), extended-connectivity fingerprints (ECFPs), and message-passing graph neural networks (GNNs), with an emphasis on graph isomorphism networks (GINs). Chapter 3 reports a computational study comparing nine combinations of featurisation and regression models on three pharmacological targets, and it introduces a pair-based data-splitting scheme for evaluating AC prediction. Chapter 4 proposes a twin neural network for AC and potency-direction (PD) classification, with proofs of built-in order-invariance and order-equivariance properties, and evaluates four versions of the model on a SARS-CoV-2 main protease data set. Chapter 5 introduces 'substructure pooling' as a general operation for vectorising structural fingerprints, proposes 'Sort & Slice' as an alternative to hashing that keeps the most frequent substructures, claims an entropy-optimality theorem under stated assumptions, and reports computational experiments suggesting that Sort & Slice outperforms hashing and other selection schemes for molecular property prediction. Chapter 6 outlines two future research directions.

Significance. If the Sort & Slice claim is correct, the manuscript identifies a simple, interpretable change to ECFP vectorisation that could replace hash-based folding in many molecular property prediction pipelines. The comparison against external baselines (hashing, standard QSAR models) rather than against the proposed models themselves is a methodological strength, and the symmetry proofs for the twin neural network (Propositions 4.1-4.3) are clean and machine-checkable in principle. The computational studies in the published chapters are careful in their use of cross-validation and hyperparameter optimisation. However, the central novelty of Chapter 5 is not yet supported to the standard claimed: the theoretical result relies on an unvalidated frequency-informativeness link, and the empirical robustness claim is not accompanied by code, full data, or significance tests. The featurisation comparisons in Chapter 3 also rest on a small number of targets and trials.

major comments (4)
  1. [Chapter 5, Section 5.2.2.2; abstract] The entropy-optimality proof for Sort & Slice is conditional on the assumption that the frequency of a circular substructure in the training set tracks its informativeness about the target property. The manuscript states this only as 'reasonable theoretical assumptions' and does not provide a formal, testable condition under which the proof holds. The method truncates to the most frequent substructures, so it discards rare substructures by construction; the activity-cliff example in Figure 3.1 shows that a rare substituent change can alter pKi by almost three orders of magnitude. Thus, the claim that Sort & Slice 'robustly leads to higher predictive performance than hashing' is not established for distributions in which rare substructures carry the predictive signal. The theoretical result collapses to a frequency filter if this assumption fails. Please either state and prove a label-dependent guarantee or explicitly restrict the scope of the optimality and robustness claims.
  2. [Chapter 5, Section 5.3 and Figures 5.2-5.7] The central empirical claim that Sort & Slice outperforms hashing 'across a large number of settings' is not supported by the information provided in the manuscript. No code or data are made available, the number of independent repetitions or seeds is not reported, and no statistical significance tests or confidence intervals are given for the comparisons in Figures 5.2-5.7. Table 5.2 lists hyperparameter ranges but not the per-model variability or the exact data-split protocol. For a claim of robust superiority, please provide the full experimental protocol, effect sizes with uncertainty, and a reproducibility package so that the comparison can be independently checked.
  3. [Chapter 3, Section 3.4 and Figures 3.7-3.12] The featurisation ranking (ECFPs best for QSAR, GINs best for AC classification) is based on only three pharmacological targets and six trials per model (2-fold cross-validation with three seeds), with no hypothesis tests. The error bars showing twice the standard deviation in Figures 3.7-3.12 indicate substantial variability relative to the reported differences, particularly for factor Xa in Figure 3.8. Statements such as 'ECFPs consistently deliver the best performance' and 'robust evidence' are stronger than the evidence supports. Please add significance tests (for example, paired bootstrap tests over the seeds) or additional datasets before making general claims about the relative merits of the featurisations.
  4. [Chapter 4, Section 4.3.2] The conclusion that the twin architecture 'outperforms standard QSAR models at AC-prediction in a variety of scenarios' is based on a single data set (SARS-CoV-2 main protease). In addition, the evaluation compares the twin model only to QSAR baselines, not to the existing tailored AC-prediction methods cited in Section 3.2 (for example, references [51], [104], and [111]), even though the chapter criticises those methods for lacking appropriate baselines. Please either add at least one additional target and include existing AC-prediction baselines, or explicitly limit the conclusion to the data set and baselines actually studied.
minor comments (4)
  1. [Throughout] There are several typos: 'activites' in Section 4.1, 'compunds' in Section 4.2.3, 'preferrable' in Section 2.3, 'alond' in Section 2.2, and 'T able' in the caption of Table 2.1. A proofreading pass is needed.
  2. [Section 1.1] The manuscript states that Chapter 3 is 'the first study that investigates the capabilities of QSAR models to classify between ACs and non-ACs,' but Section 3.2 cites related work that indirectly evaluates QSAR models on cliffy compounds. Please clarify the precise novelty claim.
  3. [Section 1.1] The text mentions 'one other work [48]' that investigates a technique similar to Sort & Slice, but the full reference is not visible in the bibliography excerpt. Please ensure the citation is complete.
  4. [Section 4.2.3] The notation `Mdouble` is introduced without a formal definition. Please define it explicitly, as the subsequent argument about containing one orientation of each MMP depends on it.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation established; Sort & Slice optimality proof is not inspectable and rests on an explicit assumption rather than a definitional circle.

full rationale

I walked the claimed derivation chain. The empirical studies in Chapters 3 and 4 are self-contained: models are trained on training folds and evaluated on disjoint test folds, and the twin-network contribution consists of architectural definitions plus symmetry proofs that do not assume the target results. The central new claim in Chapter 5 is that Sort & Slice, which keeps the most frequent circular substructures from the training set, 'robustly outperforms hash-based folding at molecular property prediction' and, 'under reasonable theoretical assumptions', 'only selects the most informative substructures from an entropic point of view'. The provided text does not include the equations of Section 5.2.2.2, so I cannot exhibit a specific reduction from the entropy-optimality theorem to its own assumptions. The frequency-informativeness link is an explicit assumption rather than a demonstrated identity; if the proof defined 'informative' as 'frequent' it would be self-definitional, but no such definition is available to quote. The self-citations in the thesis ([45], [46], [47]) point to the author's own previously published versions of the same empirical chapters and are not used to justify the Sort & Slice result, so they are not load-bearing. The absence of code and data for Chapter 5 and the unpublished status of that chapter are evidence/verification concerns, not circularity. I therefore find no significant circularity; the score of 2 reflects the presence of minor non-load-bearing self-citation within the normal 0-2 no-significant-circularity range.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims do not introduce new physical entities; the new artifacts are algorithms (Sort & Slice, twin network, pair-based splitting) and are covered under free parameters and axioms. The main unpaid assumptions are the frequency-informativeness link and train/test stability for substructures.

free parameters (2)
  • Sort & Slice top-k substructure count = not reported in abstract; optimized per data set
    The method must choose how many most frequent substructures to retain; this length controls the fingerprint dimensionality and predictive performance.
  • Activity-cliff threshold dcrit = 1.5 pK units
    Chosen as the midpoint between the non-AC (<=1) and AC (>=2) log-activity-difference intervals; all AC classification results depend on this hand-set threshold.
assumptions (4)
  • ad hoc to paper Substructure frequency in the training set approximates substructure informativeness for the target property
    Used in Chapter 5 to prove Sort & Slice selects the most informative substructures; the assumption is stated as 'reasonable' and only approximately true.
  • domain assumption Training-set substructure frequencies generalize to test molecules
    Needed for the empirical claim that selecting frequent training substructures improves test-set predictions; can fail under distribution shift.
  • domain assumption Activity cliffs are operationally defined by matched molecular pairs with >=100x activity difference, and non-cliffs by <=10x
    Standard in the AC literature, but the thresholds are choices that shape all AC and PD classification tasks.
  • domain assumption ChEMBL and COVID Moonshot activity measurements, after duplicate unification, are reliable enough for benchmarking
    The data cleaning removes conflicting duplicates; residual measurement noise is not modeled.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Investigating Graph Neural Networks and Classical Feature-Extraction Techniques in Activity-Cliff and Molecular Property Prediction." pith.science (2026). https://pith.science/paper/CLGQS6I3

@misc{pith2026241113688,
  author       = {Pith},
  title        = {Pith review of: Investigating Graph Neural Networks and Classical Feature-Extraction Techniques in Activity-Cliff and Molecular Property Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CLGQS6I3}},
  note         = {Machine review of arXiv:2411.13688}
}
read the original abstract

Molecular featurisation refers to the transformation of molecular data into numerical feature vectors. It is one of the key research areas in molecular machine learning and computational drug discovery. Recently, message-passing graph neural networks (GNNs) have emerged as a novel method to learn differentiable features directly from molecular graphs. While such techniques hold great promise, further investigations are needed to clarify if and when they indeed manage to definitively outcompete classical molecular featurisations such as extended-connectivity fingerprints (ECFPs) and physicochemical-descriptor vectors (PDVs). We systematically explore and further develop classical and graph-based molecular featurisation methods for two important tasks: molecular property prediction, in particular, quantitative structure-activity relationship (QSAR) prediction, and the largely unexplored challenge of activity-cliff (AC) prediction. We first give a technical description and critical analysis of PDVs, ECFPs and message-passing GNNs, with a focus on graph isomorphism networks (GINs). We then conduct a rigorous computational study to compare the performance of PDVs, ECFPs and GINs for QSAR and AC-prediction. Following this, we mathematically describe and computationally evaluate a novel twin neural network model for AC-prediction. We further introduce an operation called substructure pooling for the vectorisation of structural fingerprints as a natural counterpart to graph pooling in GNN architectures. We go on to propose Sort & Slice, a simple substructure-pooling technique for ECFPs that robustly outperforms hash-based folding at molecular property prediction. Finally, we outline two ideas for future research: (i) a graph-based self-supervised learning strategy to make classical molecular featurisations trainable, and (ii) trainable substructure-pooling via differentiable self-attention.

Figures

Figures reproduced from arXiv: 2411.13688 by the authors.

Figure 2.1
Figure 2.1. Generation of a Simplified Molecular-Input Line-Entry System (SMILES) string from the molecular graph of the antibiotic molecule Ciprofloxacin. First the molecular graph is reduced to its hydrogen-depleted version. Then cycles are broken to turn the graph into a spanning tree. Finally, a depth-first traversal of the spanning tree (here starting with the leftmost nitrogen atom as a root) produces the SMILES string wh… view at source ↗
Figure 2.2
Figure 2.2. The fingerprint hyperparameter R defines the maximum radius of any [PITH_FULL_IMAGE:figures/full_fig_p037_2_2.png] view at source ↗
Figure 2.3
Figure 2.3. Schematic overview of the molecular-featurisation mechanism of a message-passing graph neural network (GNN) with radius R = 2. All depicted functions may contain trainable deep-learning components. and standard gradient-based optimisation algorithms. 2.6.2 Graph Convolutional Networks We now describe an early GNN model that has been used frequently in the literature due to its relative simplicity and computational e… view at source ↗
Figures from the paper (25 more)
Figure 2.4
Figure 2.4. Figure 2.4: Example of two non-isomorphic graphs that cannot be distinguished by the 1-WL test if all nodes are assumed to have identical initial colourings. Image source: [90]. Theorem 2.1 (GNN-Conditions for 1-WL Power). A message-passing GNN can be shown to be maximally expre…
Figure 3.1
Figure 3.1. Figure 3.1: Example of an activity cliff (AC) for blood coagulation factor Xa. A small structural change in the upper compound leads to an increase in binding affinity of almost three orders of magnitude. Here binding affinity is quantified via the commonly-used pKi-value, which…
Figure 3.2
Figure 3.2. Figure 3.2: Protein structure of dopamine receptor D2. Extracted from the Research Collab￾oratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) [121]. PDB ID: 6CM4. 53 [PITH_FULL_IMAGE:figures/full_fig_p067_3_2.png]
Figure 3.3
Figure 3.3. Figure 3.3: Protein structure of factor Xa. Extracted from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) [121]. PDB ID: 2JKH. 54 [PITH_FULL_IMAGE:figures/full_fig_p068_3_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: Protein structure of SARS-CoV-2 main protease. Extracted from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) [121]. PDB ID: 6LU7. 55 [PITH_FULL_IMAGE:figures/full_fig_p069_3_4.png]
Figure 3.5
Figure 3.5. Figure 3.5: Illustration of our data splitting strategy for activity-cliff (AC) and potency￾direction (PD) classification. We distinguish between three sets of matched molecular pairs (MMPs), Mtrain,Minter and Mtest, depending on whether both MMP compounds are in Dtrain, one MMP…
Figure 3.6
Figure 3.6. Figure 3.6: Schematic showing the combinatorial experimental methodology used for the study. Each molecular featurisation method is systematically combined with each regression technique, giving a total of nine quantitative structure-activity relationship (QSAR) models. Each QSA…
Figure 3.7
Figure 3.7. Figure 3.7: QSAR-prediction and activity-cliff (AC) classification results for dopamine recep￾tor D2. For each plot, the x-axis corresponds to a combination of MMP set and AC-classification performance metric and the y-axis shows the QSAR-prediction performance on the molecular …
Figure 3.8
Figure 3.8. Figure 3.8: QSAR-prediction and activity-cliff (AC) classification results for factor Xa. For each plot, the x-axis corresponds to a combination of MMP set and AC-classification performance metric and the y-axis shows the QSAR-prediction performance on the molecular test set Dte…
Figure 3.9
Figure 3.9. Figure 3.9: QSAR-prediction and activity-cliff (AC) classification results for SARS CoV-2 main protease. For each plot, the x-axis corresponds to a combination of MMP set and AC-classification performance metric and the y-axis shows the QSAR-prediction performance on the molecul…
Figure 3.10
Figure 3.10. Figure 3.10: QSAR-prediction and potency-direction (PD) classification results for dopamine receptor D2. Each column corresponds to an upper plot and a lower plot for one of the MMP sets Minter, Mtest or Mcores. The x-axis of each upper plot indicates the PD-classification accur…
Figure 3.11
Figure 3.11. Figure 3.11: QSAR-prediction and potency-direction (PD) classification results for factor Xa. Each column corresponds to an upper plot and a lower plot for one of the MMP sets Minter, Mtest or Mcores. The x-axis of each upper plot indicates the PD-classification accuracy on the …
Figure 3.12
Figure 3.12. Figure 3.12: QSAR-prediction and potency-direction (PD) classification results for SARS-CoV-2 main protease. Each column corresponds to an upper plot and a lower plot for one of the MMP sets Minter, Mtest or Mcores. The x-axis of each upper plot indicates the PD-classification a…
Figure 4.1
Figure 4.1. Figure 4.1: Twin neural network model for activity-cliff (AC) and potency-direction (PD) clas￾sification. Changing the order of the input compounds leaves the predicted AC-classification label invariant but reverses the predicted PD-classification label. The second mapping corre…
Figure 4
Figure 4. Figure 4: that uses pre-trained ECFP-NFPs for the featurisation of individual [PITH_FULL_IMAGE:figures/full_fig_p111_4.png]
Figure 4.2
Figure 4.2. Figure 4.2: Activity-cliff (AC) classification results on the SARS-CoV-2 main protease data set for two baseline QSAR models and four newly developed twin neural network models with distinct input featurisations. The red and violet bars correspond to ECFP-based and GIN-based mod…
Figure 4.3
Figure 4.3. Figure 4.3: Potency-direction (PD) classification results on the SARS-CoV-2 main protease data set for two baseline QSAR models and four newly developed twin neural network models with distinct input featurisations. The red and violet bars correspond to ECFP-based and GIN-based …
Figure 5.1
Figure 5.1. Figure 5.1: Schematic overview of the four investigated substructure-pooling methods for the vectorisation of ECFPs. 121 [PITH_FULL_IMAGE:figures/full_fig_p135_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Predictive performance of the four substructure-pooling methods (indicated by colours) for the lipophilicity regression data set using varying data splitting techniques, prediction models and ECFP hyperparameters. Each coloured bar shows the average mean absolute err…
Figure 5.3
Figure 5.3. Figure 5.3: Predictive performance of the four substructure-pooling methods (indicated by colours) for the solubility regression data set using varying data splitting techniques, prediction models and ECFP hyperparameters. Each coloured bar shows the average mean absolute error …
Figure 5.4
Figure 5.4. Figure 5.4: Predictive performance of the four substructure-pooling methods (indicated by colours) for the SARS-CoV-2 main protease binding affinity regression data set using varying data splitting techniques, prediction models and ECFP hyperparameters. Each coloured bar shows t…
Figure 5.5
Figure 5.5. Figure 5.5: Predictive performance of the four substructure-pooling methods (indicated by colours) for the balanced mutagenicity classification data set using varying data splitting techniques, pre￾diction models and ECFP hyperparameters. Each coloured bar shows the average area…
Figure 5.6
Figure 5.6. Figure 5.6: Predictive performance of the four substructure-pooling methods (indicated by colours) for the imbalanced estrogen receptor alpha antagonism classification data set using varying data splitting techniques, prediction models and ECFP hyperparameters. Each coloured bar…
Figure 5.7
Figure 5.7. Figure 5.7: Overview of the predictive performance of the four investigated substructure￾pooling methods (indicated by colours) across regression and classification data sets, data split￾ting techniques and prediction models. Each boxplot visualises the performance of a substruc…
Figure 6.1
Figure 6.1. Figure 6.1: Step 1: Self-supervised pre-training of a graph neural network (GNN) to predict precomputed extended-connectivity fingerprints (ECFPs) from a large corpus of unlabelled molecular graphs. Step 2: Supervised training of a standard ECFP-based multilayer perceptron (MLP)…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

174 extracted references · 60 canonical work pages

  1. [51]

    Prediction of activity cliffs on the basis of images using convolutional neural networks

    Javed Iqbal, Martin Vogt, and J¨ urgen Bajorath. Prediction of activity cliffs on the basis of images using convolutional neural networks. Journal of Computer- Aided Molecular Design, 2021

  2. [104]

    Prediction of activity cliffs using support vector machines

    Kathrin Heikamp, Xiaoying Hu, Aixia Yan, and J¨ urgen Bajorath. Prediction of activity cliffs using support vector machines. Journal of Chemical Information and Modeling, 52(9):2354–2365, 2012

  3. [111]

    Prediction of activity cliffs using condensed graphs of reaction representations

    Dragos Horvath, Gilles Marcou, Alexandre Varnek, Shilva Kayastha, Antonio de la Vega de Le´ on, and J¨ urgen Bajorath. Prediction of activity cliffs using condensed graphs of reaction representations. Journal of Chemical Information and Modeling, 56(9):1631–1640, 2016

  4. [1]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. ImageNet classifica- tion with deep convolutional neural networks. In Advances in Neural Informa- tion Processing Systems, pages 1097–1105, 2012

  5. [2]

    Very deep convolutional networks for large-scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 , 2014

  6. [3]

    Zeiler and Rob Fergus

    Matthew D. Zeiler and Rob Fergus. Visualizing and understanding convolu- tional networks. In Proceedings of the European Conference on Computer Vi- sion, pages 818–833, 2014

  7. [4]

    Going deeper with convolutions

    Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabi- novich. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 1–9, 2015

  8. [5]

    Deep residual learn- ing for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learn- ing for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 770–778, 2016

Show all 174 references
  1. [6]

    A comprehensive comparison of molecular feature representations for use in predictive modeling

    Tomaˇ z Stepiˇ snik, Blaˇ zˇSkrlj, J¨ org Wicker, and Dragi Kocev. A comprehensive comparison of molecular feature representations for use in predictive modeling. Computers in Biology and Medicine , 130:104197, 2021

  2. [7]

    Large-scale comparison of machine learning methods for drug target prediction on ChEMBL

    Andreas Mayr, G¨ unter Klambauer, Thomas Unterthiner, Marvin Steijaert, J¨ org K Wegner, Hugo Ceulemans, Djork-Arn´ e Clevert, and Sepp Hochreiter. Large-scale comparison of machine learning methods for drug target prediction on ChEMBL. Chemical Science, 9(24):5441–5451, 2018

  3. [8]

    Could 163 graph neural networks learn better molecular representation for drug discov- ery? A comparison study of descriptor-based and graph-based models

    Dejun Jiang, Zhenxing Wu, Chang-Yu Hsieh, Guangyong Chen, Ben Liao, Zhe Wang, Chao Shen, Dongsheng Cao, Jian Wu, and Tingjun Hou. Could 163 graph neural networks learn better molecular representation for drug discov- ery? A comparison study of descriptor-based and graph-based ...

  4. [9]

    Mol- CLR: Molecular contrastive learning of representations via graph neural net- works

    Yuyang Wang, Jianren Wang, Zhonglin Cao, and Amir Barati Farimani. Mol- CLR: Molecular contrastive learning of representations via graph neural net- works. arXiv preprint arXiv:2102.10056 , 2021

  5. [10]

    Using domain-specific fingerprints generated through neural networks to enhance ligand-based virtual screening

    Janosch Menke and Oliver Koch. Using domain-specific fingerprints generated through neural networks to enhance ligand-based virtual screening. Journal of Chemical Information and Modeling , 61(2):664–675, 2021

  6. [11]

    ChemBERTa: Large-Scale self-supervised pretraining for molecular property prediction

    Seyone Chithrananda, Gabe Grand, and Bharath Ramsundar. ChemBERTa: Large-Scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885, 2020

  7. [12]

    Milios, and Axel J

    Mar ´ ıa Virginia Sabando, Ignacio Ponzoni, Evangelos E. Milios, and Axel J. Soto. Using molecular embeddings in QSAR modeling: Does it make a differ- ence? arXiv preprint arXiv:2104.02604 , 2021

  8. [13]

    Learn- ing continuous and data-driven molecular descriptors by translating equivalent chemical representations

    Robin Winter, Floriane Montanari, Frank No´ e, and Djork-Arn´ e Clevert. Learn- ing continuous and data-driven molecular descriptors by translating equivalent chemical representations. Chemical Science, 10(6):1692–1701, 2019

  9. [14]

    Handbook of Molecular Descriptors

    Roberto Todeschini and Viviana Consonni. Handbook of Molecular Descriptors. John Wiley & Sons, 2008

  10. [15]

    Fingerprints, and other molec- ular descriptions for database analysis and searching

    D´ avid Bajusz, Anita R´ acz, and K´ aroly H´ eberger. Fingerprints, and other molec- ular descriptions for database analysis and searching. 2017

  11. [16]

    Extended-connectivity fingerprints

    David Rogers and Mathew Hahn. Extended-connectivity fingerprints. Journal of Chemical Information and Modeling , 50(5):742–754, 2010

  12. [17]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. In Inter- national Conference on Machine Learning , pages 1263–1272. PMLR, 2017

  13. [18]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 , 2016. 164

  14. [19]

    Molecular graph convolutions: Moving beyond fingerprints

    Steven Kearnes, Kevin McCloskey, Marc Berndl, Vijay Pande, and Patrick Riley. Molecular graph convolutions: Moving beyond fingerprints. Journal of Computer-Aided Molecular Design, 30(8):595–608, 2016

  15. [20]

    Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Al´ an Aspuru-Guzik, and Ryan P

    David K. Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Al´ an Aspuru-Guzik, and Ryan P. Adams. Convolutional net- works on graphs for learning molecular fingerprints. In Advances in Neural Information Processing Systems, pages 2224–2232, 2015

  16. [21]

    Strategies for pre-training graph neural networks

    Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. Strategies for pre-training graph neural networks. arXiv preprint arXiv:1905.12265 , 2019

  17. [22]

    Analyzing learned molecular representations for property prediction

    Kevin Yang, Kyle Swanson, Wengong Jin, Connor Coley, Philipp Eiden, Hua Gao, Angel Guzman-Perez, Timothy Hopper, Brian Kelley, Miriam Mathea, et al. Analyzing learned molecular representations for property prediction. Journal of Chemical Information and Modeling , 59(8):3370–3...

  18. [23]

    Yu Philip

    Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S. Yu Philip. A comprehensive survey on graph neural networks. IEEE Trans- actions on Neural Networks and Learning Systems , 2020

  19. [24]

    A compact review of molecu- lar property prediction with graph neural networks

    Oliver Wieder, Stefan Kohlbacher, M´ elaine Kuenemann, Arthur Garon, Pierre Ducrot, Thomas Seidel, and Thierry Langer. A compact review of molecu- lar property prediction with graph neural networks. Drug Discovery Today: Technologies, 2020

  20. [25]

    Gated graph sequence neural networks

    Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. Gated graph sequence neural networks. arXiv preprint arXiv:1511.05493 , 2015

  21. [26]

    Interaction networks for learning about objects, relations and physics

    Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. In Ad- vances in Neural Information Processing Systems , pages 4502–4510, 2016

  22. [27]

    Convolutional neural networks on graphs with fast localized spectral filtering

    Micha¨ el Defferrard, Xavier Bresson, and Pierre Vandergheynst. Convolutional neural networks on graphs with fast localized spectral filtering. In Advances in Neural Information Processing Systems , pages 3844–3852, 2016. 165

  23. [28]

    Chemi-Net: A molecular graph convo- lutional network for accurate drug property prediction

    Ke Liu, Xiangyan Sun, Lei Jia, Jun Ma, Haoming Xing, Junqiu Wu, Hua Gao, Yax Sun, Florian Boulnois, and Jie Fan. Chemi-Net: A molecular graph convo- lutional network for accurate drug property prediction. International Journal of Molecular Sciences , 20(14):3389, 2019

  24. [29]

    How powerful are graph neural networks? arXiv preprint arXiv:1810.00826 , 2018

    Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826 , 2018

  25. [30]

    Towards deeper graph neural networks

    Meng Liu, Hongyang Gao, and Shuiwang Ji. Towards deeper graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , pages 338–348, 2020

  26. [31]

    Universal readout for graph convolutional neural networks

    Nicol` o Navarin, Dinh Van Tran, and Alessandro Sperduti. Universal readout for graph convolutional neural networks. In Proceedings of International Joint Conference on Neural Networks (IJCNN) , pages 1–7, 2019

  27. [32]

    Quantitative evaluation of explainable graph neural networks for molecular property predic- tion

    Jiahua Rao, Shuangjia Zheng, Yutong Lu, and Yuedong Yang. Quantitative evaluation of explainable graph neural networks for molecular property predic- tion. Patterns, page 100628, 2022

  28. [33]

    Geometric deep learning au- tonomously learns chemical features that outperform those engineered by do- main experts

    Patrick Hop, Brandon Allgood, and Jessen Yu. Geometric deep learning au- tonomously learns chemical features that outperform those engineered by do- main experts. Molecular Pharmaceutics, 15(10):4371–4377, 2018

  29. [34]

    Edge attention-based multi-relational graph convolutional networks

    Chao Shang, Qinqing Liu, Ko-Shin Chen, Jiangwen Sun, Jin Lu, Jinfeng Yi, and Jinbo Bi. Edge attention-based multi-relational graph convolutional networks. arXiv preprint arXiv: 1802.04944 , 2018

  30. [35]

    Learning graph-level representation for drug discovery

    Junying Li, Deng Cai, and Xiaofei He. Learning graph-level representation for drug discovery. arXiv preprint arXiv:1709.03741 , 2017

  31. [36]

    Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism

    Zhaoping Xiong, Dingyan Wang, Xiaohong Liu, Feisheng Zhong, Xiaozhe Wan, Xutong Li, Zhaojun Li, Xiaomin Luo, Kaixian Chen, Hualiang Jiang, et al. Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism. Journal of Medicinal Chem...

  32. [37]

    Exposing the limitations of molecular machine learning with activity cliffs

    Derek van Tilborg, Alisa Alenicheva, and Francesca Grisoni. Exposing the limitations of molecular machine learning with activity cliffs. ChemRxiv, 2022. doi: 10.26434/chemrxiv-2022-mfq52. 166

  33. [38]

    QSAR, rational approaches to the design of bioactive compounds

    Carlo Silipo and Antonio Vittoria. QSAR, rational approaches to the design of bioactive compounds. In Proceedings of European Symposium on Quantitative Structure-Activity Relationships. Distributors for the US and Canada, Elsevier Science, 1991

  34. [39]

    Maggiora

    Gerald M. Maggiora. On outliers and activity cliffs: Why QSAR often disap- points. Journal of Chemical Information and Modeling , 46(4):1535–1535, 2006

  35. [40]

    Sheridan, Prabha Karnachi, Matthew Tudor, Yuting Xu, Andy Liaw, Falgun Shah, Alan C

    Robert P. Sheridan, Prabha Karnachi, Matthew Tudor, Yuting Xu, Andy Liaw, Falgun Shah, Alan C. Cheng, Elizabeth Joshi, Meir Glick, and Juan Alvarez. Experimental error, kurtosis, activity cliffs, and methodology: What limits the predictivity of quantitative structure–activity ...

  36. [41]

    Medina-Franco, Yunierkis P´ erez-Castillo, Orazio Nicolotti, M

    Maykel Cruz-Monteagudo, Jos´ e L. Medina-Franco, Yunierkis P´ erez-Castillo, Orazio Nicolotti, M. Nat´ alia D. S. Cordeiro, and Fernanda Borges. Activity cliffs in drug discovery: Dr Jekyll or Mr Hyde? Drug Discovery Today, 19(8): 1069–1080, 2014

  37. [42]

    Recent progress in understanding activity cliffs and their utility in medicinal chem- istry: miniperspective

    Dagmar Stumpfe, Ye Hu, Dilyana Dimova, and J¨ urgen Bajorath. Recent progress in understanding activity cliffs and their utility in medicinal chem- istry: miniperspective. Journal of Medicinal Chemistry , 57(1):18–28, 2014

  38. [43]

    Evolving concept of ac- tivity cliffs

    Dagmar Stumpfe, Huabin Hu, and J¨ urgen Bajorath. Evolving concept of ac- tivity cliffs. ACS Omega, 4(11):14360–14368, 2019

  39. [44]

    Advances in exploring activity cliffs

    Dagmar Stumpfe, Huabin Hu, and J¨ urgen Bajorath. Advances in exploring activity cliffs. Journal of Computer-Aided Molecular Design , 34(9):929–942, 2020

  40. [45]

    Markus Dablander, Thierry Hanser, Renaud Lambiotte, and Garrett M. Morris. Exploring QSAR models for activity-cliff prediction. Journal of Cheminformat- ics, 15(1):47, 2023. URL https://doi.org/10.1186/s13321-023-00708-w

  41. [46]

    Mor- ris

    Markus Dablander, Thierry Hanser, Renaud Lambiotte, and Garrett M. Mor- ris. Exploring molecular machine learning models for activity-cliff predic- tion. Poster presentation at the 10th International Congress on Industrial and Applied Mathematics (ICIAM). In-person, Tokyo, 202...

  42. [47]

    Markus Dablander, Thierry Hanser, Renaud Lambiotte, and Garrett M. Morris. Siamese neural networks work for activity cliff prediction. Poster presentation at the 4th RSC-BMCS / RSC-CICAG Artificial Intelligence in Chemistry Sym- posium. Virtual, 2021. URL http://dx.doi.org/10....

  43. [48]

    (2022) Reduced collision fingerprints and pairwise molec- ular comparisons for explainable property prediction using deep learning

    Thomas MacDougall. (2022) Reduced collision fingerprints and pairwise molec- ular comparisons for explainable property prediction using deep learning. M.Sc. thesis. Universit´ e de Montr´ eal. URLhttps://hdl.handle.net/1866/26533. Accessed on 05.10.2023

  44. [49]

    Goh, Charles Siegel, Abhinav Vishnu, Nathan O

    Garrett B. Goh, Charles Siegel, Abhinav Vishnu, Nathan O. Hodas, and Nathan Baker. Chemception: A deep neural network with minimal chemistry knowl- edge matches the performance of expert-developed QSAR/QSPR models. arXiv preprint arXiv:1706.06689, 2017

  45. [50]

    Prediction of molecular properties using molecular topo- graphic map

    Atsushi Yoshimori. Prediction of molecular properties using molecular topo- graphic map. Molecules, 26(15):4475, 2021

  46. [52]

    Can one hear the shape of a molecule (from its Coulomb matrix eigenvalues)? Journal of Chemical Information and Modeling, 60(8):3804–3811, 2020

    Joshua Schrier. Can one hear the shape of a molecule (from its Coulomb matrix eigenvalues)? Journal of Chemical Information and Modeling, 60(8):3804–3811, 2020

  47. [53]

    Uni-Mol: A universal 3D molecu- lar representation learning framework

    Gengmo Zhou, Zhifeng Gao, Qiankun Ding, Hang Zheng, Hongteng Xu, Zhewei Wei, Linfeng Zhang, and Guolin Ke. Uni-Mol: A universal 3D molecu- lar representation learning framework. ChemRxiv, 2022. doi: 10.26434/ chemrxiv-2022-jjm0j-v2

  48. [54]

    Keith Lloyd, and Robin J

    Norman Biggs, E. Keith Lloyd, and Robin J. Wilson. Graph Theory, 1736-1936. Oxford University Press, 1986

  49. [55]

    Comparison of atom representations in graph neural networks for molecular property prediction

    Agnieszka Pocha, Tomasz Danel, and Lukasz Maziarka. Comparison of atom representations in graph neural networks for molecular property prediction. arXiv preprint arXiv:2012.04444 , 2020. 168

  50. [56]

    SMILES, a chemical language and information system

    David Weininger. SMILES, a chemical language and information system. Jour- nal of Chemical Information and Computer Sciences , 28(1):31–36, 1988

  51. [57]

    Weininger

    David Weininger, Arthur Weininger, and Joseph L. Weininger. Algorithm for generation of unique SMILES notation. Journal of Chemical Information and Computer Sciences, 29(2):97–101, 1989

  52. [58]

    Graphical depiction of chemical structures

    David Weininger. Graphical depiction of chemical structures. Journal of Chem- ical Information and Computer Sciences , 30(3):237–243, 1990

  53. [59]

    Image: Deriving the SMILES represen- tation of a chemical molecule, Shown example: ciprofloxacin, a fluoroquinolone antibiotic

    Fdardel (original) and DMacks (edited). Image: Deriving the SMILES represen- tation of a chemical molecule, Shown example: ciprofloxacin, a fluoroquinolone antibiotic. URL https://commons.wikimedia.org/wiki/File:SMILES.png. CC BY-SA 3.0 License, via Wikimedia Commons. Accessed...

  54. [60]

    InChI - the worldwide chemical structure identifier standard

    Stephen Heller, Alan McNaught, Stephen Stein, Dmitrii Tchekhovskoi, and Igor Pletnev. InChI - the worldwide chemical structure identifier standard. Journal of Cheminformatics , 5(1):1–9, 2013

  55. [61]

    Wei, David Duvenaud, Jos´ e Miguel Hern´ andez-Lobato, Benjam ´ ın S´ anchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D

    Rafael G´ omez-Bombarelli, Jennifer N. Wei, David Duvenaud, Jos´ e Miguel Hern´ andez-Lobato, Benjam ´ ın S´ anchez-Lengeling, Dennis Sheberla, Jorge Aguilera-Iparraguirre, Timothy D. Hirzel, Ryan P. Adams, and Al´ an Aspuru- Guzik. Automatic chemical design using a data-drive...

  56. [62]

    Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation

    Mario Krenn, Florian H¨ ase, Akshat-Kumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Machine Learning: Science and Technology , 1 (4):045024, 2020

  57. [63]

    DeepSMILES: An adaptation of SMILES for use in machine-learning of chemical structures

    Noel O’Boyle and Andrew Dalke. DeepSMILES: An adaptation of SMILES for use in machine-learning of chemical structures. ChemRxiv, 2018. doi: 10.26434/ chemrxiv.7097960.v1

  58. [64]

    Tomasz Puzyn, Jerzy Leszczynski, and Mark T. Cronin. Recent advances in QSAR studies: Methods and applications , volume 8. Springer Science & Busi- ness Media, 2010

  59. [65]

    Mold 2, molecular descriptors 169 from 2D structures for chemoinformatics and toxicoinformatics

    Huixiao Hong, Qian Xie, Weigong Ge, Feng Qian, Hong Fang, Leming Shi, Zhenqiang Su, Roger Perkins, and Weida Tong. Mold 2, molecular descriptors 169 from 2D structures for chemoinformatics and toxicoinformatics. Journal of Chemical Information and Modeling , 48(7):1337–1344, 2008

  60. [66]

    Molecular descriptors in chemoinformatics, computational combinatorial chemistry, and virtual screening

    Ling Xue and J¨ urgen Bajorath. Molecular descriptors in chemoinformatics, computational combinatorial chemistry, and virtual screening. Combinatorial Chemistry & High Throughput Screening , 3(5):363–372, 2000

  61. [67]

    Molecular descriptors

    Viviana Consonni and Roberto Todeschini. Molecular descriptors. In Recent Advances in QSAR Studies , pages 29–102. Springer, 2010

  62. [68]

    Lipinski, Franco Lombardo, Beryl W

    Christopher A. Lipinski, Franco Lombardo, Beryl W. Dominy, and Paul J. Feeney. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Advanced Drug De- livery Reviews, 23(1-3):3–25, 1997

  63. [69]

    Wildman and Gordon M

    Scott A. Wildman and Gordon M. Crippen. Prediction of physicochemical pa- rameters by atomic contributions. Journal of Chemical Information and Com- puter Sciences, 39(5):868–873, 1999

  64. [70]

    RDKit: Open-source cheminformatics

    Greg Landrum. RDKit: Open-source cheminformatics. 2006. URL http:// www.rdkit.org. Accessed on 05.10.2023

  65. [71]

    Alexandru T. Balaban. Highly discriminating distance-based topological index. Chemical Physics Letters , 89(5):399–404, 1982

  66. [72]

    Molecular representation learn- ing with language models and domain-relevant auxiliary tasks

    Benedek Fabian, Thomas Edlich, H´ el´ ena Gaspar, Marwin Segler, Joshua Mey- ers, Marco Fiscato, and Mohamed Ahmed. Molecular representation learn- ing with language models and domain-relevant auxiliary tasks. arXiv preprint arXiv:2011.13230, 2020

  67. [73]

    Harry L. Morgan. The generation of a unique machine description for chemical structures—A technique developed at chemical abstracts service. Journal of Chemical Documentation, 5(2):107–113, 1965

  68. [74]

    Durant, Burton A

    Joseph L. Durant, Burton A. Leland, Douglas R. Henry, and James G. Nourse. Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences , 42(6):1273–1280, 2002

  69. [75]

    URL https://ftp

    Online description of PubChem substructure fingerprints. URL https://ftp. ncbi.nlm.nih.gov/pubchem/specifications/pubchem_fingerprints.pdf. Accessed on 01.10.2023. 170

  70. [76]

    Lianyi Han, Yanli Wang, and Stephen H. Bryant. Developing and validating predictive decision tree models from mining chemical structural fingerprints and high-throughput screening data in PubChem. BMC Bioinformatics, 9:1–8, 2008

  71. [77]

    URL https:// www.daylight.com/dayhtml/doc/theory/theory.finger.html

    Online description of Daylight substructure fingerprints. URL https:// www.daylight.com/dayhtml/doc/theory/theory.finger.html. Accessed on 01.10.2023

  72. [78]

    Open-source platform to benchmark fin- gerprints for ligand-based virtual screening

    Sereina Riniker and Greg Landrum. Open-source platform to benchmark fin- gerprints for ligand-based virtual screening. Journal of Cheminformatics , 5(1): 26, 2013

  73. [79]

    Webel, Talia B

    Henry E. Webel, Talia B. Kimber, Silke Radetzki, Martin Neuenschwander, Marc Nazar´ e, and Andrea Volkamer. Revealing cytotoxic substructures in molecules using deep learning. Journal of Computer-aided Molecular Design , 34(7):731–746, 2020

  74. [80]

    Brown, and Mathew Hahn

    David Rogers, Robert D. Brown, and Mathew Hahn. Using extended- connectivity fingerprints with Laplacian-modified Bayesian analysis in high- throughput screening follow-up. Journal of Biomolecular Screening, 10(7):682– 686, 2005

  75. [81]

    Jonathan Alvarsson, Martin Eklund, Ola Engkvist, Ola Spjuth, Lars Carlsson, Jarl E. S. Wikberg, and Tobias Noeske. Ligand-based target prediction with signature fingerprints. Journal of Chemical Information and Modeling , 54(10): 2647–2653, 2014

  76. [82]

    Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst

    Michael M. Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst. Geometric deep learning: Going beyond Euclidean data. IEEE Signal Processing Magazine, 34(4):18–42, 2017

  77. [83]

    Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇ ckovi´ c

    Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇ ckovi´ c. Geomet- ric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021

  78. [84]

    Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson

    Vijay Prakash Dwivedi, Chaitanya K. Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. Benchmarking graph neural networks. arXiv preprint arXiv:2003.00982, 2020. 171

  79. [85]

    Continuous representa- tion of molecules using graph variational autoencoder

    Mohammadamin Tavakoli and Pierre Baldi. Continuous representa- tion of molecules using graph variational autoencoder. arXiv preprint arXiv:2004.08152, 2020

  80. [86]

    Open graph benchmark: Datasets for machine learning on graphs

    Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open graph benchmark: Datasets for machine learning on graphs. Advances in Neural Information Processing Systems, 33:22118–22133, 2020

  81. [87]

    Multilayer feedforward networks are universal approximators

    Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural Networks, 2(5):359–366, 1989

  82. [88]

    Understanding graph isomorphism net- work for rs-fMRI functional connectivity analysis

    Byung-Hoon Kim and Jong Chul Ye. Understanding graph isomorphism net- work for rs-fMRI functional connectivity analysis. Frontiers in Neuroscience, page 630, 2020

  83. [89]

    The reduction of a graph to canonical form and the algebra which appears therein

    Boris Weisfeiler and Andrei Lehman. The reduction of a graph to canonical form and the algebra which appears therein. NTI, Series , 2(9):12–16, 1968

  84. [90]

    Zafeiriou, and Michael Bron- stein

    Giorgos Bouritsas, Fabrizio Frasca, Stefanos P. Zafeiriou, and Michael Bron- stein. Improving graph neural network expressivity via subgraph isomorphism counting. IEEE Transactions on Pattern Analysis and Machine Intelligence , 2022

  85. [91]

    Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe

    Christopher Morris, Martin Ritzert, Matthias Fey, William L. Hamilton, Jan Eric Lenssen, Gaurav Rattan, and Martin Grohe. Weisfeiler and Leman go neural: Higher-order graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 33, pages 460...

  86. [92]

    Feature over- correlation in deep graph neural networks: A new perspective

    Wei Jin, Xiaorui Liu, Yao Ma, Charu Aggarwal, and Jiliang Tang. Feature over- correlation in deep graph neural networks: A new perspective. arXiv preprint arXiv:2206.07743, 2022

  87. [93]

    Evaluating deep graph neural networks

    Wentao Zhang, Zeang Sheng, Yuezihan Jiang, Yikuan Xia, Jun Gao, Zhi Yang, and Bin Cui. Evaluating deep graph neural networks. arXiv preprint arXiv:2108.00955, 2021

  88. [94]

    Gaunt, Alvaro Sanchez-Gonzalez, Yulia Rubanova, Petar Veliˇ ckovi´ c, James Kirkpatrick, and 172 Peter Battaglia

    Jonathan Godwin, Michael Schaarschmidt, Alexander L. Gaunt, Alvaro Sanchez-Gonzalez, Yulia Rubanova, Petar Veliˇ ckovi´ c, James Kirkpatrick, and 172 Peter Battaglia. Simple GNN regularisation for 3D molecular property pre- diction and beyond. In International Conference on Le...

  89. [95]

    Measuring and relieving the over-smoothing problem for graph neural networks from the topo- logical view

    Deli Chen, Yankai Lin, Wei Li, Peng Li, Jie Zhou, and Xu Sun. Measuring and relieving the over-smoothing problem for graph neural networks from the topo- logical view. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34, pages 3438–3445, 2020

  90. [96]

    Fast graph representation learning with PyTorch Geometric

    Matthias Fey and Jan Eric Lenssen. Fast graph representation learning with PyTorch Geometric. arXiv preprint arXiv:1903.02428 , 2019

  91. [97]

    Data set modelability by QSAR

    Alexander Golbraikh, Eugene Muratov, Denis Fourches, and Alexander Trop- sha. Data set modelability by QSAR. Journal of Chemical Information and Modeling, 54(1):1–4, 2014

  92. [98]

    Leadley et al

    Jr. Leadley et al. Coagulation factor Xa inhibition: Biological background and rationale. Current Topics in Medicinal Chemistry , 1(2):151–159, 2001

  93. [99]

    From activity cliffs to activity ridges: Informative data structures for SAR analysis

    Martin Vogt, Yun Huang, and J¨ urgen Bajorath. From activity cliffs to activity ridges: Informative data structures for SAR analysis. Journal of Chemical Information and Modeling , 51(8):1848–1856, 2011

  94. [100]

    Activity cliff clusters as a source of structure–activity relationship information

    Dilyana Dimova, Dagmar Stumpfe, Ye Hu, and J¨ urgen Bajorath. Activity cliff clusters as a source of structure–activity relationship information. Expert Opinion on Drug Discovery , 10(5):441–447, 2015

  95. [101]

    Medina-Franco

    Jos´ e L. Medina-Franco. Activity cliffs: Facts or artifacts? Chemical Biology & Drug Design , 81(5):553–556, 2013

  96. [102]

    Maykel Cruz-Monteagudo, Jos´ e L. Medina-Franco, Yunier Perera-Sardi˜ na, Fer- nanda Borges, Eduardo Tejera, Cesar Paz-y Mino, Yunierkis P´ erez-Castillo, Aminael S´ anchez-Rodr ´ ıguez, Zuleidys Contreras-Posada, Nat´ alia DS Cordeiro, et al. Probing the hypothesis of SAR con...

  97. [103]

    Winkler and Tu C

    David A. Winkler and Tu C. Le. Performance of deep and shallow neural net- works, the universal approximation theorem, activity cliffs, and QSAR. Molec- ular Informatics , 36(1-2):1600118, 2017. 173

  98. [105]

    Ligand-based activ- ity cliff prediction models with applicability domain

    Shunsuke Tamura, Tomoyuki Miyao, and Kimito Funatsu. Ligand-based activ- ity cliff prediction models with applicability domain. Molecular Informatics, 39 (12):2000103, 2020

  99. [106]

    Prediction of compound potency changes in matched molecular pairs using support vector regression

    Antonio De la Vega de Le´ on and J¨ urgen Bajorath. Prediction of compound potency changes in matched molecular pairs using support vector regression. Journal of Chemical Information and Modeling , 54(10):2654–2663, 2014

  100. [107]

    Beck and Clayton Springer

    Jeremy M. Beck and Clayton Springer. Quantitative structure-activity relation- ship models of chemical transformations from matched pairs analyses. Journal of Chemical Information and Modeling , 54(4):1226–1234, 2014

  101. [108]

    Searching for coordinated activity cliffs using particle swarm optimization

    Vigneshwaran Namasivayam and J¨ urgen Bajorath. Searching for coordinated activity cliffs using particle swarm optimization. Journal of Chemical Informa- tion and Modeling , 52(4):927–934, 2012

  102. [109]

    Prediction of individual compounds forming activity cliffs using emerging chemical patterns

    Vigneshwaran Namasivayam, Preeti Iyer, and J¨ urgen Bajorath. Prediction of individual compounds forming activity cliffs using emerging chemical patterns. Journal of Chemical Information and Modeling , 53(12):3131–3139, 2013

  103. [110]

    Structure-based predictions of activity cliffs

    Jarmila Husby, Giovanni Bottegoni, Irina Kufareva, Ruben Abagyan, and An- drea Cavalli. Structure-based predictions of activity cliffs. Journal of Chemical Information and Modeling , 55(5):1062–1076, 2015

  104. [112]

    Predicting activity cliffs with free-energy pertur- bation

    Laura P´ erez-Benito, Nil Casajuana-Martin, Mireia Jim´ enez-Ros´ es, Herman van Vlijmen, and Gary Tresadern. Predicting activity cliffs with free-energy pertur- bation. Journal of Chemical Theory and Computation , 15(3):1884–1895, 2019

  105. [113]

    Prediction of an MMP-1 inhibitor activity cliff using the SAR matrix approach and its experimental validation

    Yasunobu Asawa, Atsushi Yoshimori, J¨ urgen Bajorath, and Hiroyuki Naka- mura. Prediction of an MMP-1 inhibitor activity cliff using the SAR matrix approach and its experimental validation. Scientific Reports, 10(1):14710, 2020. 174

  106. [114]

    PCAC: A new method for predicting compounds with activity cliff property in QSAR approach

    Mohammad Reza Keyvanpour, Mehrnoush Barani Shirzad, and Farhaneh Moradi. PCAC: A new method for predicting compounds with activity cliff property in QSAR approach. International Journal of Information Technology , 13(6):2431–2437, 2021

  107. [115]

    ACGCN: Graph convolutional networks for activity cliff prediction be- tween matched molecular pairs

    Junhui Park, Gaeun Sung, SeungHyun Lee, SeungHo Kang, and ChunKyun Park. ACGCN: Graph convolutional networks for activity cliff prediction be- tween matched molecular pairs. Journal of Chemical Information and Modeling, 2022

  108. [116]

    DeepAC—Conditional transformer-based chemical language model for the prediction of activity cliffs formed by bioactive compounds

    Hengwei Chen, Martin Vogt, and J¨ urgen Bajorath. DeepAC—Conditional transformer-based chemical language model for the prediction of activity cliffs formed by bioactive compounds. Digital Discovery, 2022

  109. [117]

    Con- densed graph of reaction: Considering a chemical reaction as one single pseudo molecule

    Frank Hoonakker, Nicolas Lachiche, Alexandre Varnek, and Alain Wagner. Con- densed graph of reaction: Considering a chemical reaction as one single pseudo molecule. International Journal on Artificial Intelligence Tools , 20(2):253–270, 2011

  110. [118]

    Machine learning of generic reactions: 1

    Philippe Jauffret, Thierry Hanser, Christian Tonnelier, and G´ erard Kaufmann. Machine learning of generic reactions: 1. Scope of the project; The GRAMS program. Tetrahedron Computer Methodology, 3(6):323–333, 1990

  111. [119]

    Dopamine receptors and the dopamine hypothesis of schizophre- nia

    Philip Seeman. Dopamine receptors and the dopamine hypothesis of schizophre- nia. Synapse, 1(2):133–152, 1987

  112. [120]

    The SARS-CoV-2 main protease as drug target

    Sven Ullrich and Christoph Nitsche. The SARS-CoV-2 main protease as drug target. Bioorganic & Medicinal Chemistry Letters , 30(17):127377, 2020

  113. [121]

    Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N

    Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne. The Protein Data Bank. Nucleic Acids Research, 28(1):235–242, 2000. URL https://www. rcsb.org

  114. [122]

    Jorissen, and Michael K

    Tiqing Liu, Yuhmei Lin, Xin Wen, Robert N. Jorissen, and Michael K. Gilson. BindingDB: A web-accessible database of experimentally determined protein- ligand binding affinities. Nucleic Acids Research, 35:D198–D201, 2007

  115. [123]

    Bobby, Juliane Brun, BVNBS Sarma, Mark 175 Calmiano, et al

    Hagit Achdout, Anthony Aimon, Elad Bar-David, Haim Barr, Amir Ben- Shmuel, James Bennett, Melissa L. Bobby, Juliane Brun, BVNBS Sarma, Mark 175 Calmiano, et al. COVID moonshot: Open science discovery of SARS-CoV-2 main protease inhibitors by combining crowdsourcing, high-throu...

  116. [124]

    Patr ´ ıcia Bento, Anne Hersey, Eloy F´ elix, Greg Landrum, Anna Gaulton, Francis Atkinson, Louisa J

    A. Patr ´ ıcia Bento, Anne Hersey, Eloy F´ elix, Greg Landrum, Anna Gaulton, Francis Atkinson, Louisa J. Bellis, Marleen de Veij, and Andrew R. Leach. An open source chemical structure curation pipeline using RDKit. Journal of Cheminformatics, 12(1):1–16, 2020

  117. [125]

    Kenny and Jens Sadowski

    Peter W. Kenny and Jens Sadowski. Structure modification in chemical databases. Chemoinformatics in Drug Discovery , 23:271–285, 2005

  118. [126]

    Extending the activity cliff concept: Structural categorization of activity cliffs and systematic identification of different types of cliffs in the ChEMBL database

    Ye Hu and J¨ urgen Bajorath. Extending the activity cliff concept: Structural categorization of activity cliffs and systematic identification of different types of cliffs in the ChEMBL database. Journal of Chemical Information and Modeling, 52(7):1806–1811, 2012

  119. [127]

    mmpdb: An open-source matched molecular pair platform for large multiproperty data sets

    Andrew Dalke, Jerome Hert, and Christian Kramer. mmpdb: An open-source matched molecular pair platform for large multiproperty data sets. Journal of Chemical Information and Modeling , 58(5):902–910, 2018

  120. [128]

    Exploring activity cliffs from a chemoinformatics perspective

    J¨ urgen Bajorath. Exploring activity cliffs from a chemoinformatics perspective. Molecular Informatics, 33(6-7):438–442, 2014

  121. [129]

    MMP-cliffs: Systematic identification of activity cliffs on the basis of matched molecular pairs

    Xiaoying Hu, Ye Hu, Martin Vogt, Dagmar Stumpfe, and J¨ urgen Bajorath. MMP-cliffs: Systematic identification of activity cliffs on the basis of matched molecular pairs. Journal of Chemical Information and Modeling , 52(5):1138– 1145, 2012

  122. [130]

    Scikit-learn: Machine learning in Python

    Fabian Pedregosa, Ga¨ el Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in Python. Jour- nal of Machine Learning Research , 12:2825–2830, 2011

  123. [131]

    PyTorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems , 32...

  124. [132]

    Optuna: A next-generation hyperparameter optimization framework

    Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowl- edge Discovery & Data Mining , pages 2623–2631, 2019

  125. [133]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. InProceedings of Machine Learning Research, pages 448–456, 2015

  126. [134]

    Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  127. [135]

    Dropout: A simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Rus- lan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958, 2014

  128. [136]

    Siamese neural networks: An overview

    Davide Chicco. Siamese neural networks: An overview. Artificial Neural Net- works, pages 73–94, 2021

  129. [137]

    Bentz, L´ eon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard S¨ ackinger, and Roopak Shah

    Jane Bromley, James W. Bentz, L´ eon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard S¨ ackinger, and Roopak Shah. Signature verification us- ing a “Siamese” time delay neural network. International Journal of Pattern Recognition and Artificial Intelligence, 7(04):669–...

  130. [138]

    Siamese neural net- works for one-shot image recognition

    Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. Siamese neural net- works for one-shot image recognition. In ICML Deep Learning Workshop , vol- ume 2. Lille, 2015

  131. [139]

    Deepface: Closing the gap to human-level performance in face verification

    Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf. Deepface: Closing the gap to human-level performance in face verification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 1701–1708, 2014

  132. [140]

    Predicting drug-drug interactions from molecular structure images

    Devendra Singh Dhami, Gautam Kunapuli, David Page, and Sriraam Natara- jan. Predicting drug-drug interactions from molecular structure images. In Proceedings of AAAI Fall Symposium on AI for Social Good , 2019

  133. [141]

    Graph-augmented convolutional networks on drug-drug interactions pre- diction

    Yi Zhong, Xueyu Chen, Yu Zhao, Xiaoming Chen, Tingfang Gao, and Zuquan Weng. Graph-augmented convolutional networks on drug-drug interactions pre- diction. arXiv preprint arXiv:1912.03702 , 2019. 177

  134. [142]

    AttentionDDI: Siamese attention-based deep learning method for drug-drug interaction predictions

    Kyriakos Schwarz, Ahmed Allam, Nicolas Andres Perez Gonzalez, and Michael Krauthammer. AttentionDDI: Siamese attention-based deep learning method for drug-drug interaction predictions. arXiv preprint arXiv:2012.13248 , 2020

  135. [143]

    Exploring a Siamese neural network architecture for one-shot drug discovery

    Luis Torres, Nelson Monteiro, Jos` e Oliveira, Joel Arrais, and Bernardete Ribeiro. Exploring a Siamese neural network architecture for one-shot drug discovery. In Proceedings of 20th International Conference on Bioinformatics and Bioengineering (BIBE) , pages 168–175, 2020

  136. [144]

    Baskin, Vladimir A

    Igor I. Baskin, Vladimir A. Palyulin, and Nikolai S. Zefirov. Neural networks in building QSAR models. In Artificial Neural Networks, pages 133–154. Springer, 2006

  137. [145]

    Alvarez and Jaime Pahissa

    Paulino A. Alvarez and Jaime Pahissa. QT alterations in psychopharmacology: Proven candidates and suspects. Current Drug Safety , 5(1):97–104, 2010

  138. [146]

    Multifaceted protein-protein interaction prediction based on Siamese residual RCNN

    Muhao Chen, Chelsea J-T Ju, Guangyu Zhou, Xuelu Chen, Tianran Zhang, Kai-Wei Chang, Carlo Zaniolo, and Wei Wang. Multifaceted protein-protein interaction prediction based on Siamese residual RCNN. Bioinformatics, 35 (14):i305–i314, 2019

  139. [147]

    Siamese recurrent neural network with a self-attention mechanism for bioactivity prediction

    Daniel Fern´ andez-Llaneza, Silas Ulander, Dea Gogishvili, Eva Nittinger, Hong- tao Zhao, and Christian Tyrchan. Siamese recurrent neural network with a self-attention mechanism for bioactivity prediction. ACS Omega, 6(16):11086– 11094, 2021

  140. [148]

    ReSimNet: Drug response similarity prediction using Siamese neural networks

    Minji Jeon, Donghyeon Park, Jinhyuk Lee, Hwisang Jeon, Miyoung Ko, Sunkyu Kim, Yonghwa Choi, Aik-Choon Tan, and Jaewoo Kang. ReSimNet: Drug response similarity prediction using Siamese neural networks. Bioinformatics, 35(24):5249–5256, 2019

  141. [149]

    Purushothama, Vishal T

    Nicholas Roberts, Poornav S. Purushothama, Vishal T. Vasudevan, Siddarth Ravichandran, Chen Zhang, William H. Gerwick, and Garrison W. Cottrell. Using deep Siamese neural networks to speed up natural products research. ICLR 2019 Conference Blind Submission , 2018

  142. [150]

    McHardy, and Mohammad R

    Esmaeil Nourani, Ehsaneddin Asgari, Alice C. McHardy, and Mohammad R. K. Mofrad. TripletProt: Deep representation learning of proteins based on Siamese networks. IEEE/ACM Transactions on Computational Biology and Bioinfor- matics, 19(6):3744–3753, 2021. 178

  143. [151]

    Interpretable drug-target prediction using deep neural representa- tion

    Kyle Yingkai Gao, Achille Fokoue, Heng Luo, Arun Iyengar, Sanjoy Dey, and Ping Zhang. Interpretable drug-target prediction using deep neural representa- tion. In Proceedings of International Joint Conference on Artificial Intelligence, volume 2018, pages 3371–3377, 2018

  144. [152]

    Towards sparse hierarchical graph classifiers

    C˘ at˘ alina Cangea, Petar Veliˇ ckovi´ c, Nikola Jovanovi´ c, Thomas Kipf, and Pietro Li` o. Towards sparse hierarchical graph classifiers. arXiv preprint arXiv:1811.01287, 2018

  145. [153]

    Self-attention graph pooling

    Junhyun Lee, Inyeop Lee, and Jaewoo Kang. Self-attention graph pooling. In International Conference on Machine Learning, pages 3734–3743. PMLR, 2019

  146. [154]

    Asap: Adaptive struc- ture aware pooling for learning hierarchical graph representations

    Ekagra Ranjan, Soumya Sanyal, and Partha Talukdar. Asap: Adaptive struc- ture aware pooling for learning hierarchical graph representations. In Pro- ceedings of the AAAI Conference on Artificial Intelligence , volume 34, pages 5470–5477, 2020

  147. [155]

    Path integral based convolution and pooling for graph neural networks

    Zheng Ma, Junyu Xuan, Yu Guang Wang, Ming Li, and Pietro Li` o. Path integral based convolution and pooling for graph neural networks. Advances in Neural Information Processing Systems , 33:16421–16433, 2020

  148. [156]

    A probabilistic molecular fingerprint for big data settings

    Daniel Probst and Jean-Louis Reymond. A probabilistic molecular fingerprint for big data settings. Journal of Cheminformatics , 10:1–12, 2018

  149. [157]

    Filtered circular fingerprints improve either prediction or runtime performance while retaining interpretability

    Martin G¨ utlein and Stefan Kramer. Filtered circular fingerprints improve either prediction or runtime performance while retaining interpretability. Journal of Cheminformatics, 8(1):1–16, 2016

  150. [158]

    Claude E. Shannon. A mathematical theory of communication. The Bell System Technical Journal, 27(3):379–423, 1948

  151. [159]

    Cover, Joy A

    Thomas M. Cover, Joy A. Thomas, et al. Entropy, relative entropy and mutual information. Elements of Information Theory , 2(1):12–13, 1991

  152. [160]

    Daylight Chemical Information Systems

    SMARTS Theory Manual. Daylight Chemical Information Systems . URL https://www.daylight.com/dayhtml/doc/theory/theory.smarts.html. Accessed on 05.10.2023

  153. [161]

    Karl Pearson. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. The London, 179 Edinburgh, and Dublin Philosophical Magazine and J...

  154. [162]

    LIT-PCBA: An unbiased data set for machine learning and virtual screening

    Viet-Khoa Tran-Nguyen, C´ elien Jacquemard, and Didier Rognan. LIT-PCBA: An unbiased data set for machine learning and virtual screening. Journal of Chemical Information and Modeling , 60(9):4263–4273, 2020

  155. [163]

    Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S

    Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande. MoleculeNet: A benchmark for molecular machine learning. Chemical Science, 9(2):513–530, 2018

  156. [164]

    AqSolDB, a cu- rated reference set of aqueous solubility and 2D descriptors for a diverse set of compounds

    Murat Cihan Sorkun, Abhishek Khetan, and S¨ uleyman Er. AqSolDB, a cu- rated reference set of aqueous solubility and 2D descriptors for a diverse set of compounds. Scientific Data , 6(1):143, 2019

  157. [165]

    Benchmark data set for in silico prediction of Ames mutagenicity

    Katja Hansen, Sebastian Mika, Timon Schroeter, Andreas Sutter, Antonius Ter Laak, Thomas Steger-Hartmann, Nikolaus Heinrich, and Klaus-Robert Muller. Benchmark data set for in silico prediction of Ames mutagenicity. Jour- nal of Chemical Information and Modeling , 49(9):2077–2...

  158. [166]

    Bemis and Mark A

    Guy W. Bemis and Mark A. Murcko. The properties of known drugs: Molecular frameworks. Journal of Medicinal Chemistry , 39(15):2887–2893, 1996

  159. [167]

    Fast binary feature selection with conditional mutual infor- mation

    Fran¸ cois Fleuret. Fast binary feature selection with conditional mutual infor- mation. Journal of Machine Learning Research , 5(9), 2004

  160. [168]

    Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geof- frey E. Hinton. Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems , 33:22243–22255, 2020

  161. [169]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems , 30, 2017

  162. [170]

    Set transformer: A framework for attention-based permutation- invariant neural networks

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation- invariant neural networks. In International Conference on Machine Learning , pages 3744–3753. PMLR, 2019. 180

  163. [171]

    Substructure-atom cross attention for molecular representation learning

    Jiye Kim, Seungbeom Lee, Dongwoo Kim, Sungsoo Ahn, and Jaesik Park. Substructure-atom cross attention for molecular representation learning. arXiv preprint arXiv:2210.08243, 2022

  164. [172]

    Nonadditivity in public and inhouse data: Implications for drug design

    Dea Gogishvili, Eva Nittinger, Christian Margreitter, and Christian Tyrchan. Nonadditivity in public and inhouse data: Implications for drug design. Journal of Cheminformatics , 13:1–18, 2021

  165. [173]

    Implications of additivity and nonadditivity for machine learning and deep learning models in drug design

    Karolina Kwapien, Eva Nittinger, Jiazhen He, Christian Margreitter, Alexey Voronov, and Christian Tyrchan. Implications of additivity and nonadditivity for machine learning and deep learning models in drug design. ACS Omega, 7 (30):26573–26581, 2022

  166. [174]

    Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems, 34, 2021

    Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems, 34, 2021. 181

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.