Pith. sign in

REVIEW 2 major objections 5 minor 121 references

Concepts and Applications of Conformal Prediction in Computational Drug Discovery

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Conformal prediction gives any drug-discovery model confidence regions whose coverage is guaranteed at a chosen level, provided the data are exchangeable.

desk verdict Useful review of conformal prediction in drug discovery, but the classification tutorial confuses a confidence filter with CP and overclaims the validity guarantee. read the letter →

arxiv 1908.03569 v1 pith:IXYYPHX3 submitted 2019-08-09 q-bio.QM cs.LG

classification q-bio.QMcs.LG
keywords conformalpredictiondrugdiscoveryvirtualscreeningQSARuncertaintyapplicabilitydomainexchangeabilityMondrian
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Conformal Prediction (CP) is a wrapper: given any machine-learning model, it turns a bare prediction into a confidence region whose coverage is guaranteed at a user-chosen level, provided the calibration and future data are exchangeable. This review argues that CP is the practical answer to the long-standing 'applicability domain' problem in drug discovery—at confidence level 80%, the true activity or label falls inside the predicted region in at least 80% of cases, with negligible extra compute and no retraining of the underlying algorithm. The authors work through an hERG potassium-channel dataset to show how an inductive conformal predictor can be globally valid yet locally unreliable for the minority active class, and they present Mondrian calibration as the fix that restores class-wise validity. They also survey aggregated and cross-conformal variants that use all labelled data, plus deep-learning recipes based on snapshot ensembles and test-time dropout. The review is candid that the guarantee is not assumption-free: when iterative screening breaks exchangeability, conformal predictors become 'useless,' so users must check the assumption rather than take coverage on faith.

What carries the argument

The load-bearing mechanism is the non-conformity score, a scalar that measures how unusual a new instance is relative to the training distribution—for example, the fraction of random-forest trees voting for the predicted class, or the absolute prediction residual scaled by the ensemble standard deviation. These scores are computed for a held-out calibration set and sorted; a test instance receives a P-value equal to the fraction of calibration scores at least as large as its own, and the prediction is accepted at significance level $\epsilon = 1 - \text{CL}$ when the P-value is at least $\epsilon$. The quantile comparison is what converts an arbitrary model output into a region with guaranteed coverage under exchangeability, and the choice of non-conformity measure controls efficiency, meaning the tightness of the intervals or the rate of single-class predictions. Mondrian conformal prediction refines the machinery by maintaining a separate sorted score list per class, which is what restores class-wise validity on imbalanced data.

What would settle it

Build an inductive conformal predictor at $\text{CL}=0.8$ on a drug-discovery dataset, hold out a truly exchangeable test set, and count how often the true activity falls inside the predicted interval; if the empirical coverage is materially below 80% on a large test set, the practical claim of validity fails, signalling either a broken non-conformity implementation or an unnoticed breach of exchangeability. The authors' own iterative-screening experiment, in which selected molecules are no longer exchangeable with the training set, is the concrete setting where this failure is expected and observed.

Watch

Extended reading notes

Core claim

The review's central claim is that conformal prediction converts any model used in computational drug discovery—random forests, SVMs, deep networks, matrix-factorization multitask models—into a predictor that reports a confidence region with a formal coverage guarantee: at confidence level CL, the true value (regression) or true class label (classification) is contained in the predicted region in at least CL of cases. The guarantee is distribution-free and does not depend on the choice of underlying algorithm or non-conformity measure; it follows from ranking a test instance's non-conformity score against the scores of a held-out calibration set. This is what distinguishes CP from earlier applicability-domain heuristics, which only correlate distance or ensemble variance with error but do not bound the error rate. The hERG case study in the review demonstrates the distinction between global and local validity: the model is well-calibrated overall, yet most active compounds get flagged as unreliable, motivating Mondrian (class-wise) conformal prediction as the standard treatment for imbalanced screening data.

Load-bearing premise

The load-bearing premise is exchangeability: the calibration set and the molecules to be predicted must come from the same underlying distribution without any ordering or selection effect — an assumption the review admits is 'not usually verified in practice' and is openly breached when iterative virtual screening selects which molecules to test next, in which case the coverage guarantee collapses.

Editorial extensions

If this is right

  • At a chosen confidence level such as 80%, a conformal predictor flags reliable predictions with a bounded false-positive rate, so a medicinal chemist can prioritize compounds knowing the error rate is capped rather than merely correlated with model confidence.
  • On imbalanced bioactivity sets, ordinary conformal validity is only global; the minority (active) class can receive mostly unreliable predictions, so Mondrian calibration is needed to guarantee per-class coverage.
  • Because CP adds negligible computational cost and no architectural changes, the same wrapper can be applied to random forests, SVMs, deep regression networks, and multitask matrix-factorization models, giving all of them comparable uncertainty statements.
  • If exchangeability is breached—as in iterative screening where model-selected molecules are tested next—the guarantee no longer holds and the predictors can become useless, so practitioners should diagnose exchangeability or re-calibrate at each screening round.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If exchangeability is routinely violated in prospective use, the durable contribution of the review may be negative: it clarifies that distribution-free coverage is not assumption-free, and that the practical value of CP hinges on developing cheap exchangeability diagnostics or online re-calibration methods.
  • The efficiency bottleneck—intervals spanning multiple pIC50 units—is where future progress will matter most; the review's own benchmark suggests that non-conformity measures exploiting model-internal variance (bagged or dropout variance) beat simple distance-based measures, pointing to a general recipe of using learner-specific confidence signals.
  • The same quantile-calibration idea extends naturally to clinical decision support, where the guarantee would be attractive for patient-level risk predictions; but the exchangeability caveat is even more severe there, since patient cohorts are selected, not sampled, so the review's caveat transfers with added force.
  • A testable extension suggested by the review's comparative tables: use the calibration slope (observed vs nominal coverage) as a dataset-quality screen—datasets with inconsistent labels or duplicated records should show flatter calibration curves, turning conformal prediction into a diagnostic tool rather than only a prediction tool.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This review introduces conformal prediction (CP) to the computational drug discovery community, covering the core concepts of validity, efficiency, nonconformity measures, and confidence levels; the main CP modalities (inductive, Mondrian, aggregated, cross-conformal, and deep-learning variants); a broad survey of applications in virtual screening and activity modelling; open-source implementations; and current limitations, especially the exchangeability assumption. The authors include a worked hERG classification example and a regression example to illustrate ICP and MCP, and they provide a large table summarizing 25+ CP studies in drug discovery.

Significance. The manuscript is a useful and timely review of a technique that is rapidly gaining adoption for uncertainty quantification in drug discovery. Its strengths include the comprehensive survey of CP applications (Table 2), the practical descriptions of CP variants and their corresponding software, and the candid discussion of exchangeability failure in iterative screening. The core mathematical promise of CP, namely coverage validity at a user-specified confidence level, is correctly stated in the introduction and in the regression sections. However, the classification tutorial incorrectly presents a model-confidence filter as conformal prediction and claims a precision guarantee that CP does not provide. Because the paper aims to teach practitioners how to apply CP correctly, this is a load-bearing error that should be fixed before the review can serve as a reliable guide.

major comments (2)
  1. [Inductive Conformal Prediction for Classification (pp. 15-18, Figure 2, Table 3)] The worked example does not implement conformal prediction as defined by Vovk et al. It computes a single nonconformity score per test instance (the RF majority-vote fraction), thresholds it at the 80th percentile of the calibration scores, and labels a prediction 'reliable' when its score exceeds the threshold. The statement that 'the mathematical validity of CP guarantees that at least 80% of the predictions considered reliable will be correct' is not a consequence of conformal prediction theory. Conformal validity is a coverage guarantee for the set of labels whose individual p-values exceed the significance level; it is not a precision guarantee for a confidence-filtered subset. Even under exchangeability, a model that always outputs 0.9 for class A when the true prevalence of A is 0.1 will pass the threshold almost everywhere while being correct on only 10% of the 'reliable' predictions. To be a valid conformal predictor, the example must compute per-class nonconformity scores and per-label p-values (e.g., comparing the vote fraction for each class with the calibration distribution for that class), or it should be explicitly presented as a heuristic confidence filter that does not inherit CP's validity guarantee.
  2. [Mondrian Conformal Prediction (pp. 20-21)] The statement that 'at least 1-ε of the predictions for the minority class will be correct' conflates class-conditional coverage with precision. Mondrian CP guarantees that for test objects whose true class is k, the prediction set contains k with probability at least 1-ε; it does not guarantee that among objects assigned the single label k, at least 1-ε truly belong to k, since the latter depends on the prior class frequencies. Table 4 reports class-conditional coverage, so the prose should be revised to state the coverage guarantee precisely rather than implying a precision guarantee on single-label predictions.
minor comments (5)
  1. [Table 1 and p. 21] The definition of classification efficiency in Table 1 ("the fraction of single-class predictions that are correct") is inconsistent with the definition in the Mondrian section ("the single-label prediction rate"); these are different quantities and should be aligned.
  2. [Figure 5 caption] The caption of Figure 5 repeats panel labels (b) and (c), misspells 'sorted' and 'corresponds', and references non-existent panels; the panel labels and caption text should be corrected.
  3. [Equations 1 and 2] The typesetting of Equations 1 and 2 is garbled (e.g., '∝0=', '>?@ABC6@D6 E6FB?@='), making the formulas hard to read; the notation for predicted values and standard deviations should be defined consistently for calibration and test instances.
  4. [p. 5-6 and pp. 31-32] The claim that "CP does not introduce more assumptions than those generally made when modelling bioactivity data" is in tension with the later acknowledgement that exchangeability is not usually verified in practice and is breached in iterative screening; consider softening the earlier claim or pointing forward to the limitations section.
  5. [p. 7] The statement that CP requires "no parameterization ... except for the selection of a non-conformity measure" overlooks the user-specified confidence level, which is also a parameter; this should be clarified.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review's coverage guarantee is imported from external conformal prediction theory, not derived from its own fitted inputs.

full rationale

This is a review/tutorial, so there is no self-contained derivation chain to be circular. The central claim that conformal predictors provide valid confidence regions with guaranteed coverage is attributed to the external conformal prediction literature (Vovk, Gammerman and Shafer; Shafer and Vovk), not to the authors' own prior work. The authors' self-citations, such as Deep Confidence, Test-Time Dropout, and the exchangeability-failure study in iterative screening, are cited as empirical applications or limitations rather than as the source of the validity theorem; they are also published, benchmarked studies rather than restatements of the current review. The worked ICP classification example uses RF vote fractions as nonconformity scores and compares them with a calibration percentile, which is a standard inductive conformal construction. Whether the described single-score majority-vote procedure inherits the full label-set coverage guarantee may be a correctness or interpretation concern, but it is not a circularity: the coverage guarantee is not defined into the example by construction, and no fitted parameter is renamed as a prediction. Consequently, no specific reduction of the paper's claims to its own inputs can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The review's central claims rest on the conformal prediction validity theorem from Vovk et al., which is cited rather than proved, and on the exchangeability assumption on the data. The only hand-chosen number in the worked example is the active/inactive threshold. No new entities are invented.

free parameters (1)
  • hERG active/inactive cutoff (pIC50 >= 7) = 7 (pIC50 units)
    Hand-chosen threshold in the illustrative example; it creates the 1:15 class imbalance that drives the demonstration of global versus local validity. Not fitted and not load-bearing for the review's central claims.
assumptions (2)
  • standard math Conformal prediction validity theorem: for any nonconformity measure, under exchangeability, the true label lies in the prediction set with probability at least 1 - epsilon.
    The paper's central claim about validity relies on this theorem, which it cites to Vovk et al. (refs 36, 37) rather than proving.
  • domain assumption Training and application data are exchangeable (or i.i.d.).
    The validity theorem requires it; the review acknowledges in the limitations section that this assumption is not usually verified in practice and that iterative screening can violate it (ref 41).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Concepts and Applications of Conformal Prediction in Computational Drug Discovery." pith.science (2026). https://pith.science/paper/IXYYPHX3

@misc{pith2026190803569,
  author       = {Pith},
  title        = {Pith review of: Concepts and Applications of Conformal Prediction in Computational Drug Discovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IXYYPHX3}},
  note         = {Machine review of arXiv:1908.03569}
}
read the original abstract

Estimating the reliability of individual predictions is key to increase the adoption of computational models and artificial intelligence in preclinical drug discovery, as well as to foster its application to guide decision making in clinical settings. Among the large number of algorithms developed over the last decades to compute prediction errors, Conformal Prediction (CP) has gained increasing attention in the computational drug discovery community. A major reason for its recent popularity is the ease of interpretation of the computed prediction errors in both classification and regression tasks. For instance, at a confidence level of 90% the true value will be within the predicted confidence intervals in at least 90% of the cases. This so called validity of conformal predictors is guaranteed by the robust mathematical foundation underlying CP. The versatility of CP relies on its minimal computational footprint, as it can be easily coupled to any machine learning algorithm at little computational cost. In this review, we summarize underlying concepts and practical applications of CP with a particular focus on virtual screening and activity modelling, and list open source implementations of relevant software. Finally, we describe the current limitations in the field, and provide a perspective on future opportunities for CP in preclinical and clinical drug discovery.

Figures

Figures reproduced from arXiv: 1908.03569 by the authors.

Figure 2
Figure 2. Steps for the generation of an Inductive Conformal Predictor in the context of classification. As in the main text, e denotes the significance level ( [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 5
Figure 5. Modelling the hERG data set using Mondrian Conformal Prediction. (a) Fraction of trees voting for the active class for the active compounds in the calibration set. (b) Fraction of trees voting for the active class for the active compounds in the test set. Unreliable predictions (i.e., those with P values below the confidence level selected, 80%) are shown in red. (c) Distribution of class predictions for the active … view at source ↗
Figure 6
Figure 6. Influence of the confidence level on the efficiency of MCP classification models trained on the [PITH_FULL_IMAGE:figures/full_fig_p023_6.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 77 canonical work pages

  1. [1]

    Mondrian Conformal Prediction No Norinder et al.59 2017 Binary classification of imbalanced data using SVM and the distance to the separating hyperplane as the non-conformity measure Data set extracted from Baba et al.60 encompassing 211 compounds Permeation rate (log Kp) through the human skin RF and SVM Regression Aggregated Mondrian Conformal Predictio...

  2. [2]

    Today, Mondrian CP modalities, including Mondrian ACP and CCP, have become the standard approach to model imbalanced data sets when using Conformal Prediction

    might also be considered in drug discovery applications as alternative methods to generate more efficient Conformal Predictors80,90. Today, Mondrian CP modalities, including Mondrian ACP and CCP, have become the standard approach to model imbalanced data sets when using Conformal Prediction. However, also the data sets themselves which are used to model c...

  3. [3]

    Hanser, T., Barber, C., Marchaland, J. F. & Werner, S. Applicability domain: towards a more formal definition. SAR QSAR Environ. Res. 27, 865–881 (2016)

  4. [4]

    It can be seen that the predictions are globally and locally valid, i.e

    Performance of the MCP model trained for the hERG data set. It can be seen that the predictions are globally and locally valid, i.e. across the model and also for individual classes. Confidence level Global validity Validity for actives Validity for inactives 0.1 0.14 0.14 0.14 0.5 0.5 0.41 0.51 0.8 0.81 0.77 0.81 0.9 0.92 0.9 0.92 0.99 0.99 1 0.99 ICP fo...

  5. [5]

    & Cohen, I

    Vayena, E., Blasimme, A. & Cohen, I. G. Machine learning in medicine: Addressing ethical challenges. PLOS Med. 15, e1002689 (2018)

  6. [6]

    Segall, M. D. & Champness, E. J. The challenges of making decisions using uncertain data. J. Comput. Aided. Mol. Des. 29, 809–16 (2015)

  7. [7]

    Netzeva, T. I. et al. Current Status of Methods for Defining the Applicability Domain of ( Quantitative ) Structure – Activity Relationships. 2, (2005)

  8. [8]

    V, Bruneau, P., Mewes, H.-W., Rohrer, D

    Tetko, I. V, Bruneau, P., Mewes, H.-W., Rohrer, D. C. & Poda, G. I. Can we estimate the accuracy of ADME-Tox predictions? Drug Discov. Today 11, 700–707 (2006)

Show all 121 references
  1. [9]

    & Gleeson, M

    Weaver, S. & Gleeson, M. P. The importance of the domain of applicability in QSAR modeling. J. Mol. Graph. Model. 26, 1315–1326 (2008)

  2. [10]

    Toplak, M. et al. Assessment of Machine Learning Reliability Methods for Quantifying the Applicability Domain of QSAR Regression Models. J. Chem. Inf. Model. 54, 431–441 (2014)

  3. [11]

    & Hinton, G

    Lecun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521, 436–444 (2015)

  4. [12]

    Sun, J. et al. ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics. J. Cheminform. 9, 17 (2017)

  5. [13]

    Rivas-Perea, P. et al. Support Vector Machines for Regression: A Succinct Review of Large-Scale and Linear Programming Formulations. Int. J. 3, (2013)

  6. [14]

    Cherkasov, A. et al. QSAR modeling: where have you been? Where are you going to? J. Med. Chem. 57, 4977–5010 (2014)

  7. [15]

    Random Forests

    Breiman, L. Random Forests. Mach. Learn. 45, 5–32 (2001)

  8. [16]

    & Consonni, V

    Todeschini, R. & Consonni, V. Handbook of Molecular Descriptors. (2008). doi:10.1002/9783527613106

  9. [17]

    Cortes-Ciriano, I. et al. Proteochemometric modeling in a Bayesian framework. J. Cheminf. 6, 35 (2014)

  10. [18]

    & Vapnik, V

    Cortes, C. & Vapnik, V. Support-Vector Networks. Mach. Learn. 20, 273–297 (1995)

  11. [19]

    Rasmussen, C. E. & Williams, C. K. I. Gaussian Processes for Machine Learning. (Mit Press, 2006)

  12. [20]

    Obrezanova, O., Csányi, G., Gola, J. M. R. & Segall, M. D. Gaussian Processes: A Method for Automatic QSAR Modeling of ADME Properties. J. Chem. Inf. Model. 47, 1847–1857 (2007)

  13. [21]

    Gal, Y., Ghahramani, Z., Uk, Z. A. & Ghahramani, Z. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. 2015, arXiv1506.02142 arXiv.org ePrint Arch. https//arxiv.org/abs/1506.02142 35 (accessed Jul 10, 2018). (2015)

  14. [22]

    & Segall, M

    Obrezanova, O. & Segall, M. D. Gaussian processes for classification: QSAR modeling of ADMET and target activity. J. Chem. Inf. Model. 50, 1053–1061 (2010)

  15. [23]

    Svetnik, V. et al. Random forest: a classification and regression tool for compound classification and QSAR modeling. J. Chem. Inf. Comp. Sci. 43, 1947–1958 (2003)

  16. [24]

    & Lee, A

    Zhang, Y. & Lee, A. A. Bayesian semi-supervised learning for uncertainty-calibrated prediction of molecular properties and active learning. (2019)

  17. [25]

    & Wallqvist, A

    Liu, R. & Wallqvist, A. Molecular Similarity-Based Domain Applicability Metric Efficiently Identifies Out-of-Domain Compounds. 59, 181–189 (2019)

  18. [26]

    Sheridan, R. P. Using Random Forest To Model the Domain Applicability of Another Random Forest Model. J. Chem. Inf. Model. 53, 2837–2850 (2013)

  19. [27]

    Sushko, I. et al. Applicability domain for in silico models to achieve accuracy of experimental measurements. J. Chemom. 24, 202–208 (2010)

  20. [28]

    & Yamanishi, Y

    Berenger, F. & Yamanishi, Y. A Distance-Based Boolean Applicability Domain for Classification of High Throughput Screening Data. J. Chem. Inf. Model. 59, 463–476 (2019)

  21. [29]

    J., Carlsson, L., Eklund, M., Norinder, U

    Wood, D. J., Carlsson, L., Eklund, M., Norinder, U. & Stålring, J. QSAR with experimental and predictive distributions: an information theoretic approach for assessing model quality. J. Comput. Aided Mol. Des. 27, 203–219 (2013)

  22. [30]

    Schroeter, T. S. et al. Estimating the domain of applicability for machine learning QSAR models: a study on aqueous solubility of drug discovery molecules. J. Comput. Mol. Des. 21, 485–498 (2007)

  23. [31]

    Sheridan, R. P. The Relative Importance of Domain Applicability Metrics for Estimating Prediction Errors in QSAR Varies with Training Set Diversity. J. Chem. Inf. Model. 55, 1098–1107 (2015)

  24. [32]

    Defining the Applicability Domain of QSAR models : An overview

    Sahigara, F. Defining the Applicability Domain of QSAR models : An overview. Mol. Descriptors. Free online Resour. 1–6 (2007)

  25. [33]

    Schwaighofer, A. et al. Accurate solubility prediction with error bars for electrolytes: a machine learning approach. J. Chem. Inf. Model. 47, 407–424 (2007)

  26. [34]

    P., Feasel, M

    Liu, R., Glover, K. P., Feasel, M. G. & Wallqvist, A. General Approach to Estimate Error Bars for Quantitative Structure–Activity Relationship Predictions of Molecular Activity. J. Chem. Inf. Model. 58, 1561–1575 (2018)

  27. [35]

    S., van Westen, G

    Cortes-Ciriano, I., Murrell, D. S., van Westen, G. J. P., Bender, A. & Malliavin, T. Prediction of the Potency of Mammalian Cyclooxygenase Inhibitors with Ensemble Proteochemometric Modeling. J. Cheminf. 7, 1 (2014)

  28. [36]

    Norinder, U. et al. Introducing Conformal Prediction in Predictive Modeling. A Transparent and Flexible Alternative To Applicability Domain Determination. J. Chem. Inf. Model. 54, 1596–1603 (2014)

  29. [37]

    & Vovk, V

    Shafer, G. & Vovk, V. A Tutorial on Conformal Prediction. J. Mach. Learn. Res. 9, 371–421 (2008)

  30. [38]

    & Gammerman, A

    Vovk, V., Fedorova, V., Nouretdinov, I. & Gammerman, A. Criteria of efficiency for conformal prediction. in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 9653, 36 23–39 (2016)

  31. [39]

    & Eklund, M

    Norinder, U., Carlsson, L., Boyer, S. & Eklund, M. Introducing conformal prediction in predictive modeling for regulatory purposes. A transparent and flexible alternative to applicability domain determination. Regul. Toxicol. Pharmacol. 71, 279–284 (2015)

  32. [40]

    & Shafer, G

    Vovk, V., Gammerman, A. & Shafer, G. Algorithmic learning in a random world. (Springer, 2005)

  33. [41]

    C., Bender, A

    Cortes-Ciriano, I., Firth, N. C., Bender, A. & Watson, O. Discovering highly potent molecules from an initial set of inactives using iterative screening. J. Chem. Inf. Model. 58, 2000–2014 (2018)

  34. [42]

    & Boström, H

    Johansson, U., Linusson, H., Löfström, T. & Boström, H. Interpretable regression trees using conformal prediction. Expert Syst. Appl. 97, 394–404 (2018)

  35. [43]

    Lampa, S. et al. Predicting Off-Target Binding Profiles With Confidence Using Conformal Prediction. Front. Pharmacol. 9, 1256 (2018)

  36. [44]

    Svensson, F. et al. Conformal Regression for Quantitative Structure–Activity Relationship Modeling—Quantifying Prediction Uncertainty. J. Chem. Inf. Model. 58, 1132–1140 (2018)

  37. [45]

    Wahlberg, E. et al. Family-wide chemical profiling and structural analysis of PARP and tankyrase inhibitors. Nat. Biotechnol. 30, 283–288 (2012)

  38. [46]

    & Malliavin, T

    Cortés-Ciriano, I., Bender, A. & Malliavin, T. Prediction of PARP Inhibition with Proteochemometric Modelling and Conformal Prediction. Mol. Inform. 34, 357–366 (2015)

  39. [47]

    & Benfenati, E

    Fjodorova, N., Vračko, M., Novič, M., Roncaglioni, A. & Benfenati, E. New public QSAR model for carcinogenicity. Chem. Cent. J. 4, S3 (2010)

  40. [48]

    & Carlsson, L

    Eklund, M., Norinder, U., Boyer, S. & Carlsson, L. The application of conformal prediction to the drug discovery process. Ann. Math. Artif. Intell. 74, 117–132 (2015)

  41. [49]

    & Carlsson, L

    Ahlberg, E., Spjuth, O., Hasselgren, C. & Carlsson, L. Interpretation of Conformal Prediction Classification Models. in 323–334 (Springer, Cham, 2015). doi:10.1007/978-3-319-17091-6_27

  42. [50]

    Cortés-Ciriano, I. et al. Improved large-scale prediction of growth inhibition patterns using the NCI60 cancer cell line panel. Bioinformatics 32, 85–95 (2016)

  43. [51]

    Cortes-Ciriano, I. et al. Cancer Cell Line Profiler (CCLP): a webserver for the prediction of compound activity across the NCI60 panel. bioRxiv 105478 (2017). doi:10.1101/105478

  44. [52]

    Altern. Lab. Anim. 33, 155–173 (2005)

  45. [53]

    & Gini, G

    Ferrari, T. & Gini, G. An open source multistep model to predict mutagenicity from statistical analysis and relevant structural alerts. Chem. Cent. J. 4, S2 (2010)

  46. [54]

    & Andersson, P

    Norinder, U., Rybacka, A. & Andersson, P. L. Conformal prediction to define applicability domain – A case study on predicting ER and AR binding. SAR QSAR Environ. Res. 27, 303–316 (2016)

  47. [55]

    M., Carchia, M., Irwin, J

    Mysinger, M. M., Carchia, M., Irwin, J. J. & Shoichet, B. K. Directory of useful decoys, enhanced (DUD-E): better ligands and decoys for better benchmarking. 37 J. Med. Chem. 55, 6582–6594 (2012)

  48. [56]

    Kuiper, G. G. J. M. et al. Comparison of the Ligand Binding Specificity and Transcript Tissue Distribution of Estrogen Receptors α and β. Endocrinology 138, 863–870 (1997)

  49. [57]

    O., Tarairah, M., Zalloum, H

    Taha, M. O., Tarairah, M., Zalloum, H. & Abu-Sheikha, G. Pharmacophore and QSAR modeling of estrogen receptor β ligands and subsequent validation and in silico search for new hits. J. Mol. Graph. Model. 28, 383–400 (2010)

  50. [58]

    Hansen, K. et al. Benchmark Data Set for in Silico Prediction of Ames Mutagenicity. J. Chem. Inf. Model. 49, 2077–2081 (2009)

  51. [59]

    & Boyer, S

    Norinder, U. & Boyer, S. Binary classification of imbalanced datasets using conformal prediction. J. Mol. Graph. Model. 72, 256–265 (2017)

  52. [60]

    & Bender, A

    Svensson, F., Norinder, U. & Bender, A. Improving Screening Efficiency through Iterative Screening Using Docking and Conformal Prediction. J. Chem. Inf. Model. 57, 439–444 (2017)

  53. [61]

    Sun, J. et al. Applying Mondrian Cross-Conformal Prediction To Estimate Prediction Confidence on Large Imbalanced Bioactivity Data Sets. J. Chem. Inf. Model. 57, 1591–1598 (2017)

  54. [62]

    & Bender, A

    Svensson, F., Norinder, U. & Bender, A. Modelling compound cytotoxicity using conformal prediction and PubChem HTS data. Toxicol. Res. (Camb). 6, 73–80 (2017)

  55. [63]

    & Bender, A

    Cortés-Ciriano, I., Bender, A., Cortes-Ciriano, I. & Bender, A. Deep Confidence: A Computationally Efficient Framework for Calculating Reliable Prediction Errors for Deep Neural Networks. J. Chem. Inf. Model. 59, 1269–1281 (2019)

  56. [64]

    & Mamitsuka, H

    Baba, H., Takahara, J. & Mamitsuka, H. In Silico Predictions of Human Skin Permeability using Nonlinear Quantitative Structure–Property Relationship Models. Pharm. Res. 32, 2360–2371 (2015)

  57. [65]

    & Norinder, U

    Lindh, M., Karlén, A. & Norinder, U. Predicting the Rate of Skin Penetration Using an Aggregated Conformal Prediction Framework. Mol. Pharm. 14, 1571–1576 (2017)

  58. [66]

    Honma, M. et al. Improvement of quantitative structure–activity relationship (QSAR) tools for predicting Ames mutagenicity: outcomes of the Ames/QSAR International Challenge Project. Mutagenesis 34, 3–16 (2019)

  59. [67]

    & Carlsson, L

    Norinder, U., Ahlberg, E. & Carlsson, L. Predicting Ames Mutagenicity Using Conformal Prediction in the Ames/QSAR International Challenge Project. Mutagenesis (2018). doi:10.1093/mutage/gey038

  60. [68]

    Huang, G. et al. Snapshot Ensembles: Train 1, get M for free. 2017, arXiv1704.00109 arXiv.org ePrint Arch. https//arxiv.org/abs/1704.00109 (accessed Jul 10, 2018)

  61. [69]

    Norinder, U. et al. Predicting Aromatic Amine Mutagenicity with Confidence: A Case Study Using Conformal Prediction. Biomolecules 8, 85 (2018)

  62. [70]

    M., Norinder, U

    Svensson, F., Afzal, A. M., Norinder, U. & Bender, A. Maximizing gain in high-throughput screening using conformal prediction. J. Cheminform. 10, 7 (2018)

  63. [71]

    & Borrebaeck, C

    Johansson, H., Lindstedt, M., Albrekt, A.-S. & Borrebaeck, C. A. A genomic biomarker signature can predict skin sensitizers using a cell-based in vitro alternative to animal tests. BMC Genomics 12, 399 (2011)

  64. [72]

    & Forsby, A

    Norinder, U., Mucs, D., Pipping, T. & Forsby, A. Creating an efficient screening model for TRPV1 agonists using conformal prediction. Comput. Toxicol. 6, 9–15 (2018)

  65. [73]

    A., Hughes, S

    Giblin, K. A., Hughes, S. J., Boyd, H., Hansson, P. & Bender, A. Prospectively Validated Proteochemometric Models for the Prediction of Small-Molecule Binding to Bromodomain Proteins. J. Chem. Inf. Model. 58, 1870–1888 (2018)

  66. [74]

    Lapins, M. et al. A confidence predictor for logD using conformal regression and a support-vector machine. J. Cheminform. 10, 17 (2018)

  67. [75]

    UCI Machine Learning Repository

    Dua, Dheeru and Graff, C. UCI Machine Learning Repository. University of California, Irvine, School of Information and Computer Sciences (2017). Available at: http://archive.ics.uci.edu/ml. (Accessed: 1st July

  68. [76]

    & Lindstedt, M

    Forreryd, A., Norinder, U., Lindberg, T. & Lindstedt, M. Predicting skin sensitizers 38 with confidence — Using conformal prediction to determine applicability domain of GARD. Toxicol. Vitr. 48, 179–187 (2018)

  69. [77]

    & Bender, A

    Ji, C., Svensson, F., Zoufir, A. & Bender, A. eMolTox: prediction of molecular toxicity with confidence. Bioinformatics 34, 2508–2509 (2018)

  70. [78]

    Bosc, N. et al. Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery. J. Cheminform. 11, 4 (2019)

  71. [79]

    & Gillet, V

    de la Vega de León, A., Chen, B. & Gillet, V. J. Effect of missing data on multitask prediction methods. J. Cheminform. 10, 26 (2018)

  72. [80]

    & Gauraha, N

    Spjuth, O., Carlsson, L. & Gauraha, N. Aggregating Predictions on Multiple Non-disclosed Datasets using Conformal Prediction. (2018)

  73. [81]

    & Spjuth, O

    Gauraha, N., Carlsson, L. & Spjuth, O. Conformal Prediction in Learning Under Privileged Information Paradigm with Applications in Drug Discovery. (2018)

  74. [82]

    & Carlsson, L

    Eklund, M., Norinder, U., Boyer, S. & Carlsson, L. Benchmarking Variable Selection in QSAR. Mol. Inf. 31, 173–179 (2012)

  75. [83]

    & Bender, A

    Cortés-Ciriano, I. & Bender, A. KekuleScope: prediction of cancer cell line sensitivity and compound potency using convolutional neural networks trained on compound images. J. Cheminform. 11, 41 (2019)

  76. [84]

    & Svensson, F

    Norinder, U. & Svensson, F. Multitask Modeling with Confidence Using Matrix Factorization and Conformal Prediction. J. Chem. Inf. Model. acs.jcim.9b00027 (2019). doi:10.1021/acs.jcim.9b00027

  77. [85]

    & Bender, A

    Cortés-Ciriano, I. & Bender, A. Reliable Prediction Errors for Deep Neural Networks Using Test-Time Dropout. J. Chem. Inf. Model. 59, 3330–3339 (2019)

  78. [86]

    & Bender, A

    Cortés-Ciriano, I. & Bender, A. How consistent are publicly reported cytotoxicity data? Large-scale statistical analysis of the concordance of public independent cytotoxicity measurements. ChemMedChem 11, 57–71 (2015)

  79. [87]

    & Gedeck, P

    Kalliokoski, T., Kramer, C., Vulpetti, A. & Gedeck, P. Comparability of mixed IC₅₀ data - a statistical analysis. PLoS One 8, e61007 (2013)

  80. [88]

    Vovk, V., Hoi, S. C. H. & Buntine, W. Conditional validity of inductive conformal predictors. Mach Learn 92, 349–376 (Springer US, 2013)

  81. [89]

    & Vulpetti, A

    Kramer, C., Kalliokoski, T., Gedeck, P. & Vulpetti, A. The experimental uncertainty of heterogeneous public K(i) data. J. Med. Chem. 55, 5165–5173 (2012)

  82. [90]

    & Gammerman, A

    Papadopoulos, H., Vovk, V. & Gammerman, A. Regression conformal prediction with nearest neighbours. J. Artif. Intell. Res. 40, 815–840 (2011)

  83. [91]

    & Vulpetti, A

    Kalliokoski, T., Kramer, C. & Vulpetti, A. Quality Issues with Public Domain Chemogenomics Data. Mol. Inform. 32, 898–905 (2013)

  84. [92]

    J., Johnson, W

    Roberts, G., Myatt, G. J., Johnson, W. P., Cross, K. P. & Blower, P. E. LeadScope : Software for Exploring Large Sets of Screening Data. J. Chem. Inf. Comput. Sci. 40, 1302–1314 (2000)

  85. [93]

    & Clark, T

    Beck, B., Breindl, A. & Clark, T. QM/NN QSPR Models with Error Estimation: Vapor Pressure and LogP. J. Chem. Inf. Comput. Sci. 40, 1046–1051 (2000)

  86. [94]

    Sheridan, R. P. Three useful dimensions for domain applicability in QSAR models using random forest. J. Chem. Inf. Model. 52, 814–823 (2012)

  87. [95]

    & Haralambous, H

    Papadopoulos, H. & Haralambous, H. Reliable prediction intervals with regression neural networks. Neural Networks 24, 842–851 (2011)

  88. [96]

    Cortés-Ciriano, I. et al. Improved large-scale prediction of growth inhibition 39 patterns on the NCI60 cancer cell-line panel. Bioinformatics 32, 85–95 (2016)

  89. [97]

    & Verdi, M

    Wainer, H., Gessaroli, M. & Verdi, M. Visual Revelations. CHANCE 19, 49–52 (2006)

  90. [98]

    & Norinder, U

    Carlsson, L., Eklund, M. & Norinder, U. Aggregated Conformal Prediction. in 231–240 (Springer, Berlin, Heidelberg, 2014). doi:10.1007/978-3-662-44722-2_25

  91. [99]

    Linusson, H. et al. On the Calibration of Aggregated Conformal Predictors. Proc. Mach. Learn. Res. 60, 1–20 (2017)

  92. [100]

    Cross-conformal predictors

    Vovk, V. Cross-conformal predictors. Ann. Math. Artif. Intell. 74, 9–28 (2015)

  93. [101]

    & Spjuth, O

    Gauraha, N. & Spjuth, O. Synergy Conformal Prediction. (2018)

  94. [102]

    Cortes-Ciriano, I. et al. Polypharmacology Modelling Using Proteochemometrics: Recent Developments and Future Prospects. Med. Chem. Comm. 6, 24 (2015)

  95. [103]

    A., Cohen, D

    Carpenter, K. A., Cohen, D. S., Jarrell, J. T. & Huang, X. Deep learning and virtual drug screening. Future Med. Chem. 10, 2557–2567 (2018)

  96. [104]

    Ha, T.-H. et al. TRPV1 antagonist with high analgesic efficacy: 2-Thio pyridine C-region analogues of 2-(3-fluoro-4-methylsulfonylaminophenyl)propanamides. Bioorg. Med. Chem. 21, 6657–6664 (2013)

  97. [105]

    Ahmed, L. et al. Efficient iterative virtual screening with Apache Spark and conformal prediction. J. Cheminform. 10, 8 (2018)

  98. [106]

    & Salakhutdinov, R

    Srivastava, N., Hinton, G., Krizhevsky, A. & Salakhutdinov, R. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 15, 1929–1958 (2014)

  99. [107]

    & Hersey, A

    Nowotka, M., Papadatos, G., Davies, M., Dedman, N. & Hersey, A. Want Drugs? Use Python. 2016, arXiv1607.00378 arXiv.org ePrint Arch. https//arxiv.org/abs/1607.00378 (accessed Jul 10, 2018)

  100. [108]

    & Blaschke, T

    Chen, H., Engkvist, O., Wang, Y., Olivecrona, M. & Blaschke, T. The rise of deep learning in drug discovery. Drug Discov. Today 23, 1241–1250 (2018)

  101. [109]

    & Caruana, R

    Niculescu-Mizil, A. & Caruana, R. Predicting good probabilities with supervised learning. in Proceedings of the 22nd international conference on Machine learning - ICML ’05 625–632 (ACM Press, 2005). doi:10.1145/1102351.1102430

  102. [110]

    Murrell, D. S. et al. Chemically Aware Model Builder (camb): an R package for property and bioactivity modelling of small molecules. J. Cheminform. 7, 45 (2015)

  103. [111]

    Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 12, 2825–2830 (2011)

  104. [112]

    & Kuhn, M

    Mente, S. & Kuhn, M. The use of the R language for medicinal chemistry applications. Curr. Top. Med. Chem. 12, 1957–1964 (2012)

  105. [113]

    Building Predictive Models in R Using the caret Package

    Kuhn, M. Building Predictive Models in R Using the caret Package. J. Stat. Softw. 28, 1–26 (2008)

  106. [114]

    Walters, W. P. Modeling, informatics, and the quest for reproducibility. J. Chem. 40 Inf. Model. 53, 1529–1530 (2013)

  107. [115]

    Landrum, G. A. & Stiefl, N. Is that a scientific publication or an advertisement? Reproducibility, source code and data in the computational chemistry literature. Future Med. Chem. 4, 1885–1887 (2012)

  108. [116]

    Paszke, A. et al. Automatic differentiation in PyTorch. in Advances in Neural Information Processing Systems 30 1–4 (2017)

  109. [117]

    CPSign Documentation — CPSign 0.7.8 documentation

    Arvidsson, S. CPSign Documentation — CPSign 0.7.8 documentation. (2016)

  110. [120]

    (Proceedings of the IASTED international conference on artificial intelligence and applications, AIA

    Normalized Nonconformity Measures for Regression Conformal Prediction. (Proceedings of the IASTED international conference on artificial intelligence and applications, AIA. ACTA Press, 2008)

  111. [121]

    & Gammerman, A

    Toccaceli, P. & Gammerman, A. Combination of inductive mondrian conformal predictors. Mach. Learn. 108, 489–510 (2019)

  112. [239]

    The resulting data set had an imbalance in the ratio of active to inactive compounds of ~1:15

    To generate a RF binary classification model we considered as active those compounds with a pIC50 value ³ 7 (n=332), and assigned the remaining compounds to the inactive class (n=4,875). The resulting data set had an imbalance in the ratio of active to inactive compounds of ~1...

  113. [2017]

    45, D945–D954 (2017)

    Nucleic Acids Res. 45, D945–D954 (2017)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.