REVIEW 17 cited by
ChemBERTa-2: Towards Chemical Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
ChemBERTa-2: Towards Chemical Foundation Models
read the original abstract
Large pretrained models such as GPT-3 have had tremendous impact on modern natural language processing by leveraging self-supervised learning to learn salient representations that can be used to readily finetune on a wide variety of downstream tasks. We investigate the possibility of transferring such advances to molecular machine learning by building a chemical foundation model, ChemBERTa-2, using the language of SMILES. While labeled data for molecular prediction tasks is typically scarce, libraries of SMILES strings are readily available. In this work, we build upon ChemBERTa by optimizing the pretraining process. We compare multi-task and self-supervised pretraining by varying hyperparameters and pretraining dataset size, up to 77M compounds from PubChem. To our knowledge, the 77M set constitutes one of the largest datasets used for molecular pretraining to date. We find that with these pretraining improvements, we are competitive with existing state-of-the-art architectures on the MoleculeNet benchmark suite. We analyze the degree to which improvements in pretraining translate to improvement on downstream tasks.
Forward citations
Cited by 17 Pith papers
-
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
oMeBench and oMeS provide the first large-scale expert-annotated benchmark and dynamic scoring method for assessing LLM performance on organic mechanism elucidation and multi-step reasoning.
-
Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction
Chem-GMNet uses sphere-native embeddings, DualSKA attention, and SH-FFN layers to match or beat ChemBERTa-2 on MoleculeNet tasks with fewer parameters and sometimes no pretraining.
-
The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure
Screening of 52,000 bioRxiv preprints finds dual-use-adjacent content routinely present in open titles and abstracts, often exceeding risk thresholds.
-
What Does a Chemical Language Model Know About Molecules?
Sparse autoencoders on MolFormer reveal position-tracking latents in early layers and atom-in-substructure plus pharmacologically relevant features in later layers, with non-canonical SMILES causing greater representa...
-
The Metric Picks the Winner: Evaluation Choice Flips Model Rankings for Drug-Response Prediction in Unseen Chemistry
Under Bemis-Murcko scaffold split on THP-1 DRUG-seq data, inverse-variance proxy ranks linear Morgan fingerprint regression highest while contest wMSE ranks deep fusion models highest, with fusion beating linear by -0...
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
Benchmark across 78 endpoint-split entries finds classical ML winning 47.4% of best performances over pretrained models, GNNs, and LLMs, with performance depending on model-task-split fit rather than scale.
-
One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them
ROME and MEMIT knowledge edits share a common weight subset isolable by a compact binary mask that reverses ~70–80% of edits and is necessary for editing success.
-
Foundation Models for Discovery and Exploration in Chemical Space
MIST models up to 10x larger than prior work, fine-tuned on over 400 structure-property tasks, match or exceed SOTA on benchmarks and demonstrate zero-shot olfactory perception mapping consistent with hyperbolic geometry.
-
Thermodynamically consistent machine learning model for excess Gibbs energy
HANNA is a thermodynamically consistent ML model for predicting excess Gibbs energy from molecular structures, trained on various binary mixture data and extended to multi-component mixtures using geometric projection.
-
SciCore-Mol: Augmenting Large Language Models with Pluggable Molecular Cognition Modules
SciCore-Mol augments LLMs with three integrated modules for molecular perception, latent diffusion generation, and reaction reasoning, claiming an 8B open model competes with or exceeds proprietary systems on chemical tasks.
-
MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification
MSAlign aligns frozen DreaMS and ChemBERTa models with MLPs and candidate-based contrastive learning to outperform prior methods on molecule retrieval from MS/MS spectra while quantifying distribution shift in data splits.
-
Bolek: A Multimodal Language Model for Molecular Reasoning
Bolek injects Morgan fingerprint embeddings into an instruction-tuned text model, then fine-tunes on molecular alignment and synthetic chain-of-thought tasks to improve performance and grounding on 15 TDC binary class...
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
A benchmark across 156 comparisons finds classical ML models win 116 times while larger pretrained and LLM models win far fewer, showing predictive performance depends on model-task fit rather than scale.
-
A Systematic Survey and Benchmark of Deep Learning for Molecular Property Prediction in the Foundation Model Era
A systematic survey and benchmark of four deep learning paradigms for molecular property prediction that organizes the field, critiques current data practices, and outlines three future directions.
-
CVT Archives and Chemical Embedding Measures for Multi-Objective Quality Diversity in Molecular Design
CVT archives with learned chemical embeddings improve median global hypervolume and multi-objective quality diversity in NLO molecular design compared to grid-based archives.
-
Machine learning for smell: Ordinal odor strength prediction of molecular perfumery components
The authors compile an ordinal odor strength dataset for over 2,000 molecules from public sources and demonstrate supervised ML prediction of intensity categories, identifying molecular size, polarity, rings, and bran...
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
Large benchmark shows classical ML and GNNs outperform pretrained large models on most of 22 drug-discovery endpoints under strict cross-validation.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.