REVIEW 3 major objections 5 minor 66 references
A Machine Learning Pipeline for Molecular Property Prediction using ChemXploreML
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that a 32-dimensional molecular embedding can rival a 300-dimensional one, cutting runtime by about tenfold while keeping prediction accuracy within a few points.
desk verdict Useful modular tool and a real benchmark, but the headline R^2 values are inflated by pre-split cleanlab pruning, and the Mol2Vec vs VICGAE comparison runs on different cleaned datasets. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the embedding-to-regressor pipeline with a head-to-head embedding comparison. Mol2Vec represents each molecule as a 300-dimensional vector by summing learned fragment embeddings; VICGAE is a GRU autoencoder trained on partially masked SELFIES strings with variance-invariance-covariance regularization, producing 32-dimensional vectors that resist latent-space collapse and keep chemically similar molecules close. Five-fold cross-validation with Bayesian hyperparameter search is what converts those vectors into the reported R², RMSE, and MAE values, and the per-model per-property tables make the comparison traceable.
What would settle it
Recompute every reported R² and RMSE by applying the outlier-cleaning step independently inside each cross-validation fold, and compare those numbers with scores from the uncleaned data; a material drop in melting-point R² from its reported 0.86 would show the headline accuracy is an artifact of data selection.
Extended reading notes
Core claim
The central claim is that ChemXploreML works as an end-to-end platform: given molecular strings, it builds embeddings, cleans the data, tunes four tree-based regressors, and returns cross-validated property predictions. The benchmark portion claims that on well-distributed datasets the results are strong, with CatBoost on Mol2Vec reaching R² = 0.931(7) for critical temperature, 0.925(8) for boiling point, and an RMSE near 36 °C for melting point, which the authors place at the level of published QSPR models. The more transferable discovery is the embedding comparison: VICGAE's 32-dimensional vectors keep accuracy within a few points of Mol2Vec's 300-dimensional vectors while giving roughly a 10-fold speedup for gradient boosting on the largest dataset. The authors therefore present the compact embedding as the practical choice for high-throughput screening and the modular pipeline as the reusable vehicle for such comparisons.
Load-bearing premise
The load-bearing premise is that removing outliers from the full dataset before splitting it into cross-validation folds does not make the test folds artificially easy, since the cleaning step is not repeated inside each fold.
Editorial extensions
If this is right
- Compact 32-dimensional embeddings can substitute for 300-dimensional ones in tree-based property models, cutting training time by roughly an order of magnitude on large datasets.
- Data distribution and sample size dominate achievable accuracy: near-normal, well-populated properties such as critical temperature reach R² around 0.93, while the small, heavily skewed vapor-pressure set stalls near R² = 0.4 regardless of embedding.
- Because the pipeline is modular, adding a new embedding or regressor does not require rearchitecting the workflow, so the same code can be pointed at classification tasks, larger libraries, or newly integrated representations.
- Melting-point RMSE around 36 °C matches published structure-property models, suggesting the platform can serve as a practical screening tool, not just a benchmark harness.
Reading between the lines
- Pith inference: if the 32-dimensional embeddings are truly within a few points of 300-dimensional ones, representation cost—not model capacity—is the practical bottleneck for high-throughput screening, and the same speed advantage should transfer to any regressor that consumes embeddings directly.
- Pith inference: the vapor-pressure failure suggests neither embedding encodes non-covalent intermolecular interactions such as hydrogen bonding; adding explicit interaction-aware descriptors and checking whether VP R² rises above 0.4 would test that diagnosis.
- Pith inference: because outlier removal happens before the cross-validation split, all reported metrics may be optimistic; a re-evaluation that cleans inside each fold, or reports on uncleaned data, is the direct way to establish the true ceiling.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents ChemXploreML, a modular desktop application for molecular property prediction, and validates it on five CRC Handbook properties (melting point, boiling point, vapor pressure, critical temperature, critical pressure). The pipeline combines SMILES/SELFIES handling, RDKit validation, two molecular embeddings (Mol2Vec, 300-D, and VICGAE, 32-D), UMAP/DBSCAN chemical-space exploration, cleanlab-based outlier removal, Optuna hyperparameter tuning, and four tree-based regressors (GBR, XGBoost, CatBoost, LightGBM). The best reported results are R² = 0.93 for critical temperature and 0.925 for boiling point with Mol2Vec, while VICGAE is reported as comparable at much lower dimensionality and with roughly 10x speedup on the melting-point dataset. The authors claim the platform's modular design makes advanced ML accessible and reproducible.
Significance. If the evaluation were unbiased, the contribution would be a practically useful open-source platform and a favorable efficiency/accuracy trade-off for VICGAE. The manuscript's concrete strengths include public release of code and data, installation packages, a documentation site, and a thorough exploratory analysis of chemical space. The UMAP/DBSCAN clustering analyses are informative and chemically plausible. However, the central quantitative claims rest on an evaluation protocol that removes model-dependent subsets of samples before cross-validation and that reuses the same CV folds for hyperparameter tuning and final reporting. Until the evaluation is made statistically sound, the reported R² values and the Mol2Vec/VICGAE comparison cannot be accepted as stated.
major comments (3)
- [Section 6, Table 1, Table 4] The evaluation protocol applies cleanlab once to each validated dataset before the 5-fold CV split, and the cleaned sets differ substantially by embedder (CT: Mol2Vec 819/819 retained, VICGAE 777/818; MP: Mol2Vec 6167/7476 retained, VICGAE 6030/7200). Because cleanlab scores samples using predictions from a regression model, pruning before the split can selectively remove hard-to-predict molecules, making the test folds in Table 4 artificially easy and non-representative of the original CRC data. The Mol2Vec versus VICGAE comparison is also uncontrolled because the two pipelines are evaluated on different molecule sets. No results on the uncleaned datasets or with cleaning nested inside each CV fold are reported, so the magnitude of this bias is unknown; the headline R² values (0.93 CT, 0.925 BP) and the speed comparison rest on this assumption.
- [Section 7.2, Section 8.1] Hyperparameter optimization is performed with 5-fold CV ('For each model, we performed extensive hyperparameter tuning using Optuna with 5-fold cross-validation'), and the reported metrics in Table 4 are also computed with 5-fold CV on the same cleaned data. No held-out test set or nested CV is described. Selecting hyperparameters by minimizing RMSE on the same folds used to report R² can inflate the reported performance and the ranking between Mol2Vec and VICGAE. Please report an independent final evaluation, e.g., a fixed held-out split or nested CV.
- [Section 8.2, Section 9, Table 4] The conclusion states that VICGAE 'even outperformed Mol2Vec for vapor pressure prediction,' but Section 8.2 says the VP difference is likely not statistically significant because uncertainty ranges overlap (e.g., CatBoost VP R² 0.4(2) versus 0.32(7)). This is an internal contradiction in one of the paper's stated advantages of VICGAE and should be corrected to say 'comparable' rather than 'outperformed.'
minor comments (5)
- [Figure 6] In most panels the inset labels appear to swap RMSE and MAE relative to Table 4 (e.g., MP GBR inset gives RMSE 30(2) and MAE 39(2), whereas Table 4 lists RMSE 39(2) and MAE 30(2)). Please correct the figure or the table.
- [Section 9] The sentence 'VICGAE showed R² values of 0.4(2) compared to Mol2Vec's 0.32(7)' omits the property name; specify vapor pressure (VP) and the model (CatBoost) for clarity.
- [Section 6] The text says the melting point dataset experienced an 18% reduction in Mol2Vec embeddings, while Table 1 shows 6167 of 7476 retained (17.5% reduction); please round consistently or state exact percentages.
- [Section 9] The claim that applicability domain (AD) analysis via leverage and Mahalanobis distance is implemented in ChemXploreML is not demonstrated anywhere in the paper; either provide evidence of this functionality or label it as planned.
- [Section 5] UMAP and DBSCAN parameters are described as optimized through visual assessment; since these choices affect the exploratory analysis (not the regression), the paper should state more explicitly that they are heuristic and not part of the quantitative evaluation.
Circularity Check
Pre-split cleanlab pruning makes reported R² partially self-referential and confounds the Mol2Vec vs VICGAE comparison.
-
fitted input called prediction
[Section 6 (Data Preprocessing Pipeline), Table 1, Table 4 caption]
"Post-embedding, for automated outlier detection, we leverage cleanlab 35–37 to identify and remove problematic data points. ... As shown in Table 1, while the melting point dataset experienced an 18% reduction in Mol2Vec embeddings ... only a minimal fraction of data is pruned, which inherently enhances data reliability for robust model training. ... All metrics are computed with 5-fold CV."
cleanlab's confident-learning framework detects label noise and outliers using predictions from regression models (refs 35–37), and the cleaning is applied once before the 5-fold split that produces Table 4. The test folds are therefore not an independent sample of the original CRC data: hard-to-predict molecules are preferentially removed by the same model family being evaluated, so the reported R² and RMSE partially measure the cleaning model's ability to select easy points rather than genuine generalization. The paper reports no metrics on uncleaned data and does not nest cleaning inside each fold, so the magnitude of the optimistic bias is unknown.
full rationale
The main derivation content is not circular: Mol2Vec and VICGAE are pre-existing external embedding methods, the regression models are standard external algorithms, and the property datasets are from the CRC Handbook. The pipeline's reported R² values are empirical benchmark results rather than analytically derived predictions, so the core 'ChemXploreML works' claim has independent content. The significant circularity is in the evaluation protocol: cleanlab, which uses model predictions to prune outliers, is run before the 5-fold CV, and the same post-clean data are then used to report prediction quality. This makes the reported MP, VP, CP, and cross-embedder performance partially self-referential, because the test folds have been selected using predictions from the same family of models. The CT and BP headline numbers are less affected (Mol2Vec CT retains all 819 molecules; BP retains 98%), but the Mol2Vec-vs-VICGAE comparison is still confounded by different cleaned sets. A milder, non-circular but related concern is that Optuna hyperparameter tuning and the final 5-fold CV evaluation share the same folds, which can optimistically bias the reported scores. Overall, this is partial evaluation circularity rather than a derivation that reduces to its inputs by construction.
Assumptions & free parameters
free parameters (6)
- UMAP n_neighbors =
25
- UMAP min_dist =
0.3
- DBSCAN eps =
0.7
- DBSCAN min_samples =
15
- cleanlab data-pruning decisions =
18% of MP data removed (Mol2Vec)
- Model hyperparameters (Optuna-tuned) =
See Table 3
assumptions (4)
- domain assumption CRC Handbook values are accurate experimental ground truth for the five properties.
- domain assumption The pre-trained Mol2Vec and VICGAE embeddings contain sufficient chemical information to predict the five properties.
- ad hoc to paper cleanlab identifies true label errors rather than hard-to-predict molecules, so removing them does not bias evaluation.
- standard math 5-fold cross-validation on the cleaned data estimates generalization to new molecules.
Cite this review
Pith. "Pith review of A Machine Learning Pipeline for Molecular Property Prediction using ChemXploreML." pith.science (2026). https://pith.science/paper/CT5PQMZK
@misc{pith2026250508688,
author = {Pith},
title = {Pith review of: A Machine Learning Pipeline for Molecular Property Prediction using ChemXploreML},
year = {2026},
howpublished = {\url{https://pith.science/paper/CT5PQMZK}},
note = {Machine review of arXiv:2505.08688}
}
abstract
We present ChemXploreML, a modular desktop application designed for machine learning-based molecular property prediction. The framework's flexible architecture allows integration of any molecular embedding technique with modern machine learning algorithms, enabling researchers to customize their prediction pipelines without extensive programming expertise. To demonstrate the framework's capabilities, we implement and evaluate two molecular embedding approaches - Mol2Vec and VICGAE (Variance-Invariance-Covariance regularized GRU Auto-Encoder) - combined with state-of-the-art tree-based ensemble methods (Gradient Boosting Regression, XGBoost, CatBoost, and LightGBM). Using five fundamental molecular properties as test cases - melting point (MP), boiling point (BP), vapor pressure (VP), critical temperature (CT), and critical pressure (CP) - we validate our framework on a dataset from the CRC Handbook of Chemistry and Physics. The models achieve excellent performance for well-distributed properties, with R$^2$ values up to 0.93 for critical temperature predictions. Notably, while Mol2Vec embeddings (300 dimensions) delivered slightly higher accuracy, VICGAE embeddings (32 dimensions) exhibited comparable performance yet offered significantly improved computational efficiency. ChemXploreML's modular design facilitates easy integration of new embedding techniques and machine learning algorithms, providing a flexible platform for customized property prediction tasks. The application automates chemical data preprocessing (including UMAP-based exploration of molecular space), model optimization, and performance analysis through an intuitive interface, making sophisticated machine learning techniques accessible while maintaining extensibility for advanced cheminformatics users.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Molecular representations in AI -driven drug discovery: a review and practical guide
David, L.; Thakkar, A.; Mercado, R.; Engkvist, O. Molecular representations in AI -driven drug discovery: a review and practical guide. Journal of Cheminformatics 2020, 12, 56
work page 2020
-
[2]
Janet, J. P.; Kulik, H. J. Machine Learning in Chemistry ; ACS In Focus ; American Chemical Society, 2020
work page 2020
-
[3]
Kulik, H. J. Making machine learning a useful tool in the accelerated discovery of transition metal complexes. WIREs Computational Molecular Science 2020, 10, e1439
work page 2020
-
[4]
Internally Consistent Prediction of Vapor Pressure and Related Properties
Korsten, H. Internally Consistent Prediction of Vapor Pressure and Related Properties . Industrial & Engineering Chemistry Research 2000, 39, 813--820, Publisher: American Chemical Society
work page 2000
-
[5]
Breedveld, G. J. F.; Prausnitz, J. M. Thermodynamic properties of supercritical fluids and their mixtures at very high pressures. AIChE Journal 1973, 19, 783--796
work page 1973
-
[6]
Marano, J. J.; Holder, G. D. General Equation for Correlating the Thermophysical Properties of n- Paraffins , n- Olefins , and Other Homologous Series . 2. Asymptotic Behavior Correlations for PVT Properties . Industrial & Engineering Chemistry Research 1997, 36, 1895--1907, Publisher: American Chemical Society
work page 1997
-
[7]
Prediction of Fluid Phase Behavior from Molecular Models
Lucas, K. Prediction of Fluid Phase Behavior from Molecular Models . AIP Conference Proceedings 2007, 963, 93--103
work page 2007
-
[8]
Whiteside, T. S.; Hilal, S. H.; Brenner, A.; Carreira, L. A. Estimating the melting point, entropy of fusion, and enthalpy of fusion of organic compounds via SPARC . SAR and QSAR in Environmental Research 2016, Publisher: Taylor & Francis
work page 2016
Show all 66 references
-
[9]
Dearden, J. C. Quantitative structure-property relationships for prediction of boiling point, vapor pressure, and melting point. Environmental Toxicology and Chemistry 2003, 22, 1696--1709
2003
-
[10]
W.; Craighead, J
Hay, M.; Thomas, D. W.; Craighead, J. L.; Economides, C.; Rosenthal, J. Clinical development success rates for investigational drugs. Nature Biotechnology 2014, 32, 40--51, Publisher: Nature Publishing Group
2014
-
[11]
Trends in clinical success rates and therapeutic focus
Dowden, H.; Munro, J. Trends in clinical success rates and therapeutic focus. Nature Reviews Drug Discovery 2019, 18, 495--496, Bandiera\_abtest: a Cg\_type: From The Analyst's Couch Publisher: Nature Publishing Group Subject\_term: Drug discovery
2019
-
[12]
Mol2vec: Unsupervised Machine Learning Approach with Chemical Intuition
Jaeger, S.; Fulle, S.; Turk, S. Mol2vec: Unsupervised Machine Learning Approach with Chemical Intuition . Journal of Chemical Information and Modeling 2018, 58, 27--35, Publisher: American Chemical Society
2018
-
[13]
Lee, K. L. K. Language models for astrochemistry. 2021; https://zenodo.org/records/7559628
2021
-
[14]
N.; Remijan, A
Scolati, H. N.; Remijan, A. J.; Herbst, E.; McGuire, B. A.; Lee, K. L. K. Explaining the Chemical Inventory of Orion KL through Machine Learning . The Astrophysical Journal 2023, 959, 108, Publisher: The American Astronomical Society
2023
-
[15]
Lee, K. L. K.; Patterson, J.; Burkhardt, A. M.; Vankayalapati, V.; McCarthy, M. C.; McGuire, B. A. Machine Learning of Interstellar Chemical Inventories . The Astrophysical Journal Letters 2021, 917, L6, Publisher: The American Astronomical Society
2021
-
[16]
Friedman, J. H. Greedy function approximation: A gradient boosting machine. The Annals of Statistics 2001, 29, 1189--1232, Publisher: Institute of Mathematical Statistics
2001
-
[17]
XGBoost : A Scalable Tree Boosting System
Chen, T.; Guestrin, C. XGBoost : A Scalable Tree Boosting System . Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . New York, NY, USA, 2016; pp 785--794
2016
-
[18]
LightGBM : A Highly Efficient Gradient Boosting Decision Tree
Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.-Y. LightGBM : A Highly Efficient Gradient Boosting Decision Tree . Advances in Neural Information Processing Systems . 2017
2017
-
[19]
V.; Gulin, A
Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A. V.; Gulin, A. CatBoost : unbiased boosting with categorical features. Advances in Neural Information Processing Systems . 2018
2018
-
[20]
Boldini, D.; Grisoni, F.; Kuhn, D.; Friedrich, L.; Sieber, S. A. Practical guidelines for the use of gradient boosting for molecular property prediction. Journal of Cheminformatics 2023, 15, 73
2023
-
[21]
Marimuthu, A. N. aravindhnivas/ ChemXploreML : v4.0.0. 2025; https://zenodo.org/doi/10.5281/zenodo.14977096
2025 doi
-
[22]
SMILES , a chemical language and information system
Weininger, D. SMILES , a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences 1988, 28, 31--36, Publisher: American Chemical Society
1988
-
[23]
O’Boyle, N. M. Towards a Universal SMILES representation - A standard method to generate canonical SMILES based on the InChI . Journal of Cheminformatics 2012, 4, 22
2012
-
[24]
Self-referencing embedded strings ( SELFIES ): A 100\
Krenn, M.; Häse, F.; Nigam, A.; Friederich, P.; Aspuru-Guzik, A. Self-referencing embedded strings ( SELFIES ): A 100\
-
[25]
Team, T. C. Tauri: Build smaller, faster, and more secure desktop applications with a web frontend. 2024; https://tauri.app, Version 2.1.1
2024
-
[26]
Harris, R.; Team, S. C. Svelte: Cybernetically enhanced web applications. 2024; https://svelte.dev, Version 4.2.1
2024
-
[27]
2019; https://www.python.org/, Version 3.12
Python Core Team Python: A dynamic, open source programming language . 2019; https://www.python.org/, Version 3.12
2019
-
[28]
Landrum, G. et al. rdkit/rdkit: 2024\_09\_2 ( Q3 2024) Release . 2024; https://zenodo.org/doi/10.5281/zenodo.591637
2024 doi
-
[29]
Pedregosa, F. et al. Scikit-learn: Machine Learning in Python . Journal of Machine Learning Research 2011, 12, 2825--2830
2011
-
[30]
Optuna: A Next -generation Hyperparameter Optimization Framework
Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A Next -generation Hyperparameter Optimization Framework . 2019; http://arxiv.org/abs/1907.10902, arXiv:1907.10902
2019 arXiv
-
[31]
Dask: Parallel Computation with Blocked algorithms and Task Scheduling
Rocklin, M. Dask: Parallel Computation with Blocked algorithms and Task Scheduling . Austin, Texas, 2015; pp 126--132
2015
-
[32]
R., Brunno, T
Rumble, J. R., Brunno, T. J., Doa, M. J., Eds. CRC handbook of chemistry and physics: a ready-reference book of chemical and physical data , 105th ed.; CRC handbook of chemistry and physics / Chemical Rubber Company 105th edition (2024); CRC Press: Boca Raton London New York, 2024
2024
-
[33]
A.; Cheng, T.; Zhang, J.; Gindulyte, A.; Bolton, E
Kim, S.; Thiessen, P. A.; Cheng, T.; Zhang, J.; Gindulyte, A.; Bolton, E. E. PUG - View : programmatic access to chemical annotations integrated in PubChem . Journal of Cheminformatics 2019, 11, 56
2019
-
[34]
https://github.com/mcs07/CIRpy
mcs07/ CIRpy : Python wrapper for the NCI Chemical Identifier Resolver ( CIR ). https://github.com/mcs07/CIRpy
-
[35]
2024; https://github.com/cleanlab/cleanlab, original-date: 2018-05-11T01:55:21Z
cleanlab. 2024; https://github.com/cleanlab/cleanlab, original-date: 2018-05-11T01:55:21Z
2024
-
[36]
Detecting Errors in Numerical Data via any Regression Model
Zhou, H.; Mueller, J.; Kumar, M.; Wang, J.-L.; Lei, J. Detecting Errors in Numerical Data via any Regression Model. ICML Workshop on Data-centric Machine Learning Research. 2023
2023
-
[37]
Model-agnostic label quality scoring to detect real-world label errors
Kuan, J.; Mueller, J. Model-agnostic label quality scoring to detect real-world label errors. ICML DataPerf Workshop. 2022
2022
-
[38]
Riniker, S.; Landrum, G. A. Open-source platform to benchmark fingerprints for ligand-based virtual screening. Journal of Cheminformatics 2013, 5, 26
2013
-
[39]
M.; Sayle, R
O’Boyle, N. M.; Sayle, R. A. Comparing structural fingerprints using a literature-based similarity benchmark. Journal of Cheminformatics 2016, 8, 36
2016
-
[40]
Riniker, S.; Fechner, N.; Landrum, G. A. Heterogeneous Classifier Fusion for Ligand - Based Virtual Screening : Or , How Decision Making by Committee Can Be a Good Thing . Journal of Chemical Information and Modeling 2013, 53, 2829--2836, Publisher: American Chemical Society
2013
-
[41]
DeepTox : Toxicity Prediction using Deep Learning
Mayr, A.; Klambauer, G.; Unterthiner, T.; Hochreiter, S. DeepTox : Toxicity Prediction using Deep Learning . Frontiers in Environmental Science 2016, 3, Publisher: Frontiers
2016
-
[42]
Profiling Prediction of Kinase Inhibitors : Toward the Virtual Assay
Merget, B.; Turk, S.; Eid, S.; Rippmann, F.; Fulle, S. Profiling Prediction of Kinase Inhibitors : Toward the Virtual Assay . Journal of Medicinal Chemistry 2017, 60, 474--485, Publisher: American Chemical Society
2017
-
[43]
A.; Fulle, S.; Merget, B
Sorgenfrei, F. A.; Fulle, S.; Merget, B. Kinome- Wide Profiling Prediction of Small Molecules . ChemMedChem 2018, 13, 495--499
2018
-
[44]
Efficient Estimation of Word Representations in Vector Space
Mikolov, T.; Chen, K.; Corrado, G.; Dean, J. Efficient Estimation of Word Representations in Vector Space . 2013; http://arxiv.org/abs/1301.3781, arXiv:1301.3781 [cs]
2013 arXiv
-
[45]
Extended- Connectivity Fingerprints
Rogers, D.; Hahn, M. Extended- Connectivity Fingerprints . Journal of Chemical Information and Modeling 2010, 50, 742--754, Publisher: American Chemical Society
2010
-
[46]
Repurposed drugs and nutraceuticals targeting envelope protein: A possible therapeutic strategy against COVID -19
Das, G.; Das, T.; Chowdhury, N.; Chatterjee, D.; Bagchi, A.; Ghosh, Z. Repurposed drugs and nutraceuticals targeting envelope protein: A possible therapeutic strategy against COVID -19. Genomics 2021, 113, 1129--1140
2021
-
[47]
Identifying Structure – Property Relationships through SMILES Syntax Analysis with Self - Attention Mechanism
Zheng, S.; Yan, X.; Yang, Y.; Xu, J. Identifying Structure – Property Relationships through SMILES Syntax Analysis with Self - Attention Mechanism . Journal of Chemical Information and Modeling 2019, 59, 914--923, Publisher: American Chemical Society
2019
-
[48]
Fried, Z. T. P.; Lee, K. L. K.; Byrne, A. N.; McGuire, B. A. Implementation of rare isotopologues into machine learning of the chemical inventory of the solar-type protostellar source IRAS 16293-2422. Digital Discovery 2023, 2, 952--966, Publisher: RSC
2023
-
[49]
VICReg : Variance - Invariance - Covariance Regularization for Self - Supervised Learning
Bardes, A.; Ponce, J.; LeCun, Y. VICReg : Variance - Invariance - Covariance Regularization for Self - Supervised Learning . 2022; http://arxiv.org/abs/2105.04906, arXiv:2105.04906 [cs]
2022 arXiv
-
[50]
Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism
Xiong, Z.; Wang, D.; Liu, X.; Zhong, F.; Wan, X.; Li, X.; Li, Z.; Luo, X.; Chen, K.; Jiang, H.; Zheng, M. Pushing the Boundaries of Molecular Representation for Drug Discovery with the Graph Attention Mechanism . Journal of Medicinal Chemistry 2020, 63, 8749--8760, Publisher: ...
2020
-
[51]
Analyzing Learned Molecular Representations for Property Prediction
Yang, K.; Swanson, K.; Jin, W.; Coley, C.; Eiden, P.; Gao, H.; Guzman-Perez, A.; Hopper, T.; Kelley, B.; Mathea, M.; Palmer, A.; Settels, V.; Jaakkola, T.; Jensen, K.; Barzilay, R. Analyzing Learned Molecular Representations for Property Prediction . Journal of Chemical Inform...
2019
-
[52]
Correction to Analyzing Learned Molecular Representations for Property Prediction
Yang, K.; Swanson, K.; Jin, W.; Coley, C.; Eiden, P.; Gao, H.; Guzman-Perez, A.; Hopper, T.; Kelley, B.; Mathea, M.; Palmer, A.; Settels, V.; Jaakkola, T.; Jensen, K.; Barzilay, R. Correction to Analyzing Learned Molecular Representations for Property Prediction . Journal of C...
2019
-
[53]
B.; Hodas, N
Goh, G. B.; Hodas, N. O.; Siegel, C.; Vishnu, A. SMILES2Vec : An Interpretable General - Purpose Deep Neural Network for Predicting Chemical Properties . 2018; http://arxiv.org/abs/1712.02034, arXiv:1712.02034 [stat]
2018 arXiv
-
[54]
ChemBERTa : Large - Scale Self - Supervised Pretraining for Molecular Property Prediction
Chithrananda, S.; Grand, G.; Ramsundar, B. ChemBERTa : Large - Scale Self - Supervised Pretraining for Molecular Property Prediction . 2020; http://arxiv.org/abs/2010.09885, arXiv:2010.09885
2020 arXiv
-
[55]
Mol- BERT : An Effective Molecular Representation with BERT for Molecular Property Prediction
Li, J.; Jiang, X. Mol- BERT : An Effective Molecular Representation with BERT for Molecular Property Prediction . Wireless Communications and Mobile Computing 2021, 2021, 7181815
2021
-
[56]
Large- Scale Chemical Language Representations Capture Molecular Structure and Properties
Ross, J.; Belgodere, B.; Chenthamarakshan, V.; Padhi, I.; Mroueh, Y.; Das, P. Large- Scale Chemical Language Representations Capture Molecular Structure and Properties . 2022; http://arxiv.org/abs/2106.09553, arXiv:2106.09553 [cs]
2022 arXiv
-
[57]
G.; Jung, G.; Cole, J
Jung, S. G.; Jung, G.; Cole, J. M. Gradient boosted and statistical feature selection workflow for materials property predictions. The Journal of Chemical Physics 2023, 159, 194106
2023
-
[58]
G.; Jung, G.; Cole, J
Jung, S. G.; Jung, G.; Cole, J. M. Automatic Prediction of Molecular Properties Using Substructure Vector Embeddings within a Feature Selection Workflow . Journal of Chemical Information and Modeling 2024, 65, 133--152
2024
-
[59]
Functionality pattern matching as an efficient complementary structure/reaction search tool: an open-source approach
Haider, N. Functionality pattern matching as an efficient complementary structure/reaction search tool: an open-source approach. Molecules (Basel, Switzerland) 2010, 15, 5079--5092
2010
-
[60]
A Density - Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise
Ester, M.; Kriegel, H.-P.; Sander, J.; Xu, X. A Density - Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise . In: Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining, Portland, OR, AAAI Press 1996, 226--231, In: P...
1996
-
[61]
P.; Xu, X
Schubert, E.; Sander, J.; Ester, M.; Kriegel, H. P.; Xu, X. DBSCAN Revisited , Revisited : Why and How You Should ( Still ) Use DBSCAN . ACM Trans. Database Syst. 2017, 42, 19:1--19:21
2017
-
[62]
Tree- Structured Parzen Estimator : Understanding Its Algorithm Components and Their Roles for Better Empirical Performance
Watanabe, S. Tree- Structured Parzen Estimator : Understanding Its Algorithm Components and Their Roles for Better Empirical Performance . 2023; http://arxiv.org/abs/2304.11127, arXiv:2304.11127
2023 arXiv
-
[63]
Snoek, J.; Larochelle, H.; Adams, R. P. Practical Bayesian Optimization of Machine Learning Algorithms . 2012; http://arxiv.org/abs/1206.2944, arXiv:1206.2944
2012 arXiv
-
[64]
Banchero, M.; Manna, L. Comparison between Multi - Linear - and Radial - Basis - Function - Neural - Network - Based QSPR Models for The Prediction of The Critical Temperature , Critical Pressure and Acentric Factor of Organic Compounds . Molecules : A Journal of Synthetic Che...
2018
-
[65]
V.; Sushko, Y.; Novotarskyi, S.; Patiny, L.; Kondratov, I.; Petrenko, A
Tetko, I. V.; Sushko, Y.; Novotarskyi, S.; Patiny, L.; Kondratov, I.; Petrenko, A. E.; Charochkina, L.; Asiri, A. M. How Accurately Can We Predict the Melting Points of Drug -like Compounds ? Journal of Chemical Information and Modeling 2014, 54, 3320--3329, Publisher: America...
2014
-
[66]
sWE䮨 &HHܳȌ犈? _o_|? G|W |W CK ׃/__ ۿm w s<S] ǥw k 5 ;s Ꮖ ]5y?=n hH ; ) =ۿHGG;)+BL=;4A|O &>
McDonagh, J. L.; van Mourik, T.; Mitchell, J. B. O. Predicting Melting Points of Organic Molecules : Applications to Aqueous Solubility Prediction Using the General Solubility Equation . Molecular Informatics 2015, 34, 715--724 mcitethebibliography main.tex00006640000000000000...
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.