REVIEW 43 cited by
ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
read the original abstract
GNNs and chemical fingerprints are the predominant approaches to representing molecules for property prediction. However, in NLP, transformers have become the de-facto standard for representation learning thanks to their strong downstream task transfer. In parallel, the software ecosystem around transformers is maturing rapidly, with libraries like HuggingFace and BertViz enabling streamlined training and introspection. In this work, we make one of the first attempts to systematically evaluate transformers on molecular property prediction tasks via our ChemBERTa model. ChemBERTa scales well with pretraining dataset size, offering competitive downstream performance on MoleculeNet and useful attention-based visualization modalities. Our results suggest that transformers offer a promising avenue of future work for molecular representation learning and property prediction. To facilitate these efforts, we release a curated dataset of 77M SMILES from PubChem suitable for large-scale self-supervised pretraining.
Forward citations
Cited by 43 Pith papers
-
Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent
NMRAgent is an evidential LLM agent for NMR-based molecular structure elucidation that improves accuracy on novel scaffolds and demonstrates utility on real natural products.
-
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
scCycleMol adds a learnable circular cell-cycle head with closed-loop supervision from predicted treated expression, yielding higher r-squared on SciPlex3 gene predictions and improved phase accuracy versus ChemCPA baselines.
-
Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment
LOGICA adds context to pretrained biological LMs via logit-space contrastive alignment with gated adapters, improving AUC on held-out drug-resistance mutation ranking from ~0.55 to ~0.65 while preserving token likelihoods.
-
Augmenting Molecular Language Models with Local $n$-gram Memory
MolGram integrates a conditional n-gram memory module into molecular language models to address locality gaps in SMILES tokenization, improving performance on generation, forward prediction, and retrosynthesis while o...
-
Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction
Chem-GMNet uses sphere-native embeddings, DualSKA attention, and SH-FFN layers to match or beat ChemBERTa-2 on MoleculeNet tasks with fewer parameters and sometimes no pretraining.
-
FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization
FORGE reformulates molecular optimization as context-aware fragment ranking and replacement using mined low-to-high edit pairs, outperforming larger language models and graph methods on standard benchmarks.
-
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
Chirality emerges in SMILES translation models through an abrupt encoder-centered reorganization of representations after a long plateau, identified via checkpoint analysis and ablation.
-
MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model
Querying a fingerprint-conditioned molecule-language model with a calibrated band of spectrum-induced posteriors beats point-threshold fingerprint decoding on NPLIB1 and MassSpecGym.
-
Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
Generic molecular encoders show weak alignment with human olfactory rating geometry and no clear predictive increment over an RDKit-Morgan chemistry baseline, while human rating geometry is reproducible within but onl...
-
OLEDLM: A Unified Language Model for OLED Molecular Design
A GRPO-aligned LLaMA model generates OLED SMILES conditioned on target S1 and f, with the property-alignment evidence largely coming from the very BERT predictor used as reward.
-
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
A magnetic hypergraph neural network with chemistry-inspired directional flow improves ADMET property prediction on several benchmarks.
-
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
Augmenting SLM prompts with a GNN expert's prediction, confidence, and an extracted important subgraph improves zero-shot toxicity/mutagenicity accuracy on MUTAG and Tox21, with gains up to 74% relative to SMILES-only...
-
Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
Pre-training improves CLMs' encoding of molecular substructures especially in upper layers while fine-tuning selectively modifies task-relevant substructures more than others.
-
What Does a Chemical Language Model Know About Molecules?
Sparse autoencoders on MolFormer reveal position-tracking latents in early layers and atom-in-substructure plus pharmacologically relevant features in later layers, with non-canonical SMILES causing greater representa...
-
Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements
Closed-loop LM-agent auto research finds some transferable gains on molecular property prediction benchmarks via external data but shows non-transfer for model and feature edits selected on validation.
-
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
Decentralized AI agent teams self-organize around hypotheses, critique proposals, and share knowledge to outperform single-agent baselines on biomedical ML, language-model optimization, and protein fitness tasks.
-
Training distribution determines the ceiling of drug-blind cancer sensitivity prediction
Drug-blind cancer sensitivity prediction is limited by evaluation metric and training distribution rather than drug representation complexity.
-
Can LLMs Predict Polymer Physics Just by Reading Synthesis and Processing Prose?
PolyLM fine-tunes a 9B-parameter LLM on 185k papers to predict polymer properties from text alone, achieving median R² of 0.74 on 68k held-out samples.
-
Molecules Meet Language: Confound-Aware Representation Learning and Chemical Property Steering in Transformer-VAE Latent Spaces
Chemically meaningful steering for properties like cLogP and TPSA emerges in entangled Transformer-VAE latent spaces only after controlling for SELFIES representation confounds through residualization and decoded traversals.
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
Benchmark across 78 endpoint-split entries finds classical ML winning 47.4% of best performances over pretrained models, GNNs, and LLMs, with performance depending on model-task-split fit rather than scale.
-
NOSE: Neural Olfactory-Semantic Embedding with Tri-Modal Orthogonal Contrastive Learning
NOSE aligns molecular, receptor, and linguistic modalities in a shared embedding space via tri-modal orthogonal contrastive learning and weak positive samples, achieving SOTA performance and zero-shot generalization o...
-
SIGMA: Semantic Identifier Grouping for Molecular Autoregression
A same-suffix contrastive objective makes autoregressive molecular-string models more invariant to how a molecule is written, improving generation fidelity on the reported ZINC benchmark.
-
FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics
FlexMS is a new flexible benchmarking framework that lets researchers dynamically combine deep learning architectures and evaluate their mass spectrum prediction performance on public metabolomics datasets using multi...
-
Foundation Models for Discovery and Exploration in Chemical Space
MIST models up to 10x larger than prior work, fine-tuned on over 400 structure-property tasks, match or exceed SOTA on benchmarks and demonstrate zero-shot olfactory perception mapping consistent with hyperbolic geometry.
-
SmellNet: A Large-scale Dataset for Real-world Smell Recognition
SmellNet supplies 828k gas-sensor time series across 50 substances plus 43 mixtures; ScentFormer reaches 63.3% top-1 accuracy on classification and 50.2% top-1@0.1 on mixture prediction.
-
ChemCrow: Augmenting large-language models with chemistry tools
ChemCrow augments LLMs with 18 expert chemistry tools to autonomously plan and execute syntheses and guide molecular discoveries in organic synthesis, drug discovery, and materials design.
-
Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction
Chem World unifies 17 chemical mixture datasets into 10 property tracks, and Mixture-PINN’s soft physics regularizers beat common encoder–aggregator baselines on most tracks and OOD splits.
-
Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
A reliability-aware evidential fusion model combines four docking engines and produces well-calibrated confidence scores that enable selective prediction with up to 25.7% error reduction.
-
A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It
Marginal conformal prediction under-covers minority classes by tens of points on imbalanced molecular datasets while hitting global coverage; Mondrian calibration recovers the target.
-
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
Closed-loop cell-cycle supervision on predicted treated expression raises SciPlex3 phase accuracy by 0.54–0.62 points while keeping all-gene R² within 0.003 of matched chemCPA baselines.
-
A large-scale foundation model enables simulation-to-real adaptation for nuclear magnetic resonance-based molecular structure analysis
UltraNMR is a large foundation model pre-trained on 158M simulated NMR spectra that transfers to experimental data, achieving SOTA on structure analysis tasks and enabling a 94M-molecule spectral library plus real-wor...
-
MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry
MolE-RAG is a training-free RAG framework that augments LLMs with literature, molecular context, and structural analogs to improve performance on nine molecular property prediction tasks.
-
When Tabular Foundation Models Transfer Across Modalities: A Systematic Evaluation Across 95 Datasets, 7 Modalities, and Two Regimes
A tabular foundation model pipeline with ETF preprocessing transfers across 7 modalities on 95 datasets, matching lightweight tuned baselines on frozen features at much higher speed while providing calibration for deployment.
-
SPADE: Faster Drug Discovery by Learning from Sparse Data
SPADE selects ligands more efficiently than deep learning or Bayesian optimization, needing fewer tests on average to identify high-quality drug candidates for novel proteins.
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
A benchmark across 156 comparisons finds classical ML models win 116 times while larger pretrained and LLM models win far fewer, showing predictive performance depends on model-task fit rather than scale.
-
When Active Learning Falls Short: An Empirical Study on Chemical Reaction Extraction
Active learning for chemical reaction extraction frequently produces non-monotonic learning curves and fails to deliver stable gains over random sampling because of strong pretraining, structured CRF decoding, and lab...
-
Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining
Lit2Vec delivers a documented, reproducible pipeline that extracts and annotates a large licensed chemistry paper corpus from S2ORC with paragraph embeddings and subfield labels.
-
Persistent Manifold Learning of Protein Properties
Persistent manifold learning with Boundary-Induced Graph Laplacians plus language-model features beats prior SOTA Pearson correlation on metalloprotein–ligand and SKEMPI wild-type protein–protein affinity benchmarks.
-
Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings
Tri-modal late fusion of SchNet geometry, ChemBERTa SMILES, and DCN descriptors reaches 0.0207 eV MAE on QM9 U0 atomization energy, a 20.6% gain over a controlled SchNet baseline under 1M parameters.
-
GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction
GLACIER combines graph, SMILES, and descriptor encoders with Finsler fusion and contrastive distillation to produce an efficient multimodal model for molecular property prediction.
-
Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction
Large benchmark shows classical ML and GNNs outperform pretrained large models on most of 22 drug-discovery endpoints under strict cross-validation.
-
Regression with Large Language Models for Materials and Molecular Property Prediction
Fine-tuned LLaMA 3 achieves regression performance on QM9 molecular properties and 28 materials properties from composition strings that rivals random forests but is 5-10x worse than specialized models using atomic co...
-
Machine Learning Based Prediction of Proton Conductivity in Metal-Organic Frameworks
A newly built database of proton-conductive MOFs supports descriptor and transformer machine learning models that predict conductivity with MAE 0.91.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.