Pith. sign in

REVIEW 1 cited by

Targeting Neurodegeneration: Three Machine Learning Methods for G9a Inhibitors Discovery Using PubChem and Scikit-learn

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16214 v2 pith:QVEWCRQC submitted 2025-03-20 q-bio.QM

classification q-bio.QM
keywords modelerrormodelsclassifierpubchemsamplestestedtrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In light of the increasing interest in G9a's role in neuroscience, three machine learning (ML) models, that are time efficient and cost effective, were developed to support researchers in this area. The models are based on data provided by PubChem and performed by algorithms interpreted by the scikit-learn Python-based ML library. The first ML model aimed to predict the efficacy magnitude of active G9a inhibitors. The ML models were trained with 3,112 and tested with 778 samples. The Gradient Boosting Regressor perform the best, achieving 17.81% means relative error (MRE), 21.48% mean absolute error (MAE), 27.39% root mean squared error (RMSE) and 0.02 coefficient of determination (R2) error. The goal of the second ML model called a CID_SID ML model, utilised PubChem identifiers to predict the G9a inhibition probability of a small biomolecule that has been primarily designed for different purposes. The ML models were trained with 58,552 samples and tested with 14,000. The most suitable classifier for this case study was the Extreme Gradient Boosting Classifier, which obtained 78.1% accuracy, 84.3% precision,69.1% recall, 75.9% F1-score and 8.1% Receiver-operating characteristic (ROC). The third ML model based on the Random Forest Classifier algorithm led to the generation of a list of descending-ordered functional groups based on their importance to the G9a inhibition. The model was trained with 19,455 samples and tested with 14,100. The probability of this rank was 70% accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Comparative analysis of computational approaches for predicting Transthyretin (TTR) transcription activators and human dopamine D1 receptor antagonists

    q-bio.QM 2025-06 reject novelty 3.0 of 10

    Adding PubChem molecular features to 13C NMR bins improves the authors' D1 antagonist classifier, but the headline TTR accuracy is a speculative extrapolation from a much smaller dataset.

Pith tools