REVIEW 6 cited by
SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In drug-discovery-related tasks such as virtual screening, machine learning is emerging as a promising way to predict molecular properties. Conventionally, molecular fingerprints (numerical representations of molecules) are calculated through rule-based algorithms that map molecules to a sparse discrete space. However, these algorithms perform poorly for shallow prediction models or small datasets. To address this issue, we present SMILES Transformer. Inspired by Transformer and pre-trained language models from natural language processing, SMILES Transformer learns molecular fingerprints through unsupervised pre-training of the sequence-to-sequence language model using a huge corpus of SMILES, a text representation system for molecules. We performed benchmarks on 10 datasets against existing fingerprints and graph-based methods and demonstrated the superiority of the proposed algorithms in small-data settings where pre-training facilitated good generalization. Moreover, we define a novel metric to concurrently measure model accuracy and data efficiency.
Forward citations
Cited by 6 Pith papers
-
TopoFormer: Topology Meets Attention for Graph Learning
Sliding-window interlevel Betti sequences (Topo-Scan) plus Transformers match or beat strong GNN and TDA baselines on graph classification and molecular property tasks while avoiding full persistence diagrams.
-
MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model
Querying a fingerprint-conditioned molecule-language model with a calibrated band of spectrum-induced posteriors beats point-threshold fingerprint decoding on NPLIB1 and MassSpecGym.
-
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...
-
SIGMA: Semantic Identifier Grouping for Molecular Autoregression
A same-suffix contrastive objective makes autoregressive molecular-string models more invariant to how a molecule is written, improving generation fidelity on the reported ZINC benchmark.
-
SmilesT5: Domain-specific pretraining for molecular language models
SmilesT5 shows that pretraining a T5 model to reconstruct Murcko scaffolds and predict molecular fragments beats masked-language pretraining on six molecular property classification benchmarks.
-
ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation
ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...
Discussion (0). Continue with ORCID to comment.