Pith. sign in

REVIEW 17 cited by

ChemBERTa-2: Towards Chemical Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.01712 v1 pith:XFETXKGV submitted 2022-09-05 cs.LG cs.AIq-bio.BM

ChemBERTa-2: Towards Chemical Foundation Models

classification cs.LG cs.AIq-bio.BM
keywords pretrainingmoleculartaskschemberta-2chemicaldownstreamfoundationimprovements
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large pretrained models such as GPT-3 have had tremendous impact on modern natural language processing by leveraging self-supervised learning to learn salient representations that can be used to readily finetune on a wide variety of downstream tasks. We investigate the possibility of transferring such advances to molecular machine learning by building a chemical foundation model, ChemBERTa-2, using the language of SMILES. While labeled data for molecular prediction tasks is typically scarce, libraries of SMILES strings are readily available. In this work, we build upon ChemBERTa by optimizing the pretraining process. We compare multi-task and self-supervised pretraining by varying hyperparameters and pretraining dataset size, up to 77M compounds from PubChem. To our knowledge, the 77M set constitutes one of the largest datasets used for molecular pretraining to date. We find that with these pretraining improvements, we are competitive with existing state-of-the-art architectures on the MoleculeNet benchmark suite. We analyze the degree to which improvements in pretraining translate to improvement on downstream tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

    cs.AI 2025-10 unverdicted novelty 8.0

    oMeBench and oMeS provide the first large-scale expert-annotated benchmark and dynamic scoring method for assessing LLM performance on organic mechanism elucidation and multi-step reasoning.

  2. Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction

    cs.LG 2026-05 unverdicted novelty 7.0

    Chem-GMNet uses sphere-native embeddings, DualSKA attention, and SH-FFN layers to match or beat ChemBERTa-2 on MoleculeNet tasks with fewer parameters and sometimes no pretraining.

  3. The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure

    cs.DL 2026-05 unverdicted novelty 7.0

    Screening of 52,000 bioRxiv preprints finds dual-use-adjacent content routinely present in open titles and abstracts, often exceeding risk thresholds.

  4. What Does a Chemical Language Model Know About Molecules?

    cs.LG 2026-06 unverdicted novelty 6.0

    Sparse autoencoders on MolFormer reveal position-tracking latents in early layers and atom-in-substructure plus pharmacologically relevant features in later layers, with non-canonical SMILES causing greater representa...

  5. The Metric Picks the Winner: Evaluation Choice Flips Model Rankings for Drug-Response Prediction in Unseen Chemistry

    cs.LG 2026-06 unverdicted novelty 6.0

    Under Bemis-Murcko scaffold split on THP-1 DRUG-seq data, inverse-variance proxy ranks linear Morgan fingerprint regression highest while contest wMSE ranks deep fusion models highest, with fusion beating linear by -0...

  6. Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

    cs.LG 2026-04 unverdicted novelty 6.0

    Benchmark across 78 endpoint-split entries finds classical ML winning 47.4% of best performances over pretrained models, GNNs, and LLMs, with performance depending on model-task-split fit rather than scale.

  7. One Mask to Rule Them All: On Hidden Facts after Editing and How to Find Them

    cs.LG 2026-04 conditional novelty 6.0

    ROME and MEMIT knowledge edits share a common weight subset isolable by a compact binary mask that reverses ~70–80% of edits and is necessary for editing success.

  8. Foundation Models for Discovery and Exploration in Chemical Space

    physics.chem-ph 2025-10 unverdicted novelty 6.0

    MIST models up to 10x larger than prior work, fine-tuned on over 400 structure-property tasks, match or exceed SOTA on benchmarks and demonstrate zero-shot olfactory perception mapping consistent with hyperbolic geometry.

  9. Thermodynamically consistent machine learning model for excess Gibbs energy

    cs.LG 2025-09 unverdicted novelty 6.0

    HANNA is a thermodynamically consistent ML model for predicting excess Gibbs energy from molecular structures, trained on various binary mixture data and extended to multi-component mixtures using geometric projection.

  10. SciCore-Mol: Augmenting Large Language Models with Pluggable Molecular Cognition Modules

    cs.AI 2026-05 unverdicted novelty 5.0

    SciCore-Mol augments LLMs with three integrated modules for molecular perception, latent diffusion generation, and reaction reasoning, claiming an 8B open model competes with or exceeds proprietary systems on chemical tasks.

  11. MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification

    cs.LG 2026-05 conditional novelty 5.0

    MSAlign aligns frozen DreaMS and ChemBERTa models with MLPs and candidate-based contrastive learning to outperform prior methods on molecule retrieval from MS/MS spectra while quantifying distribution shift in data splits.

  12. Bolek: A Multimodal Language Model for Molecular Reasoning

    cs.LG 2026-05 unverdicted novelty 5.0

    Bolek injects Morgan fingerprint embeddings into an instruction-tuned text model, then fine-tunes on molecular alignment and synthetic chain-of-thought tasks to improve performance and grounding on 15 TDC binary class...

  13. Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

    cs.LG 2026-04 unverdicted novelty 5.0

    A benchmark across 156 comparisons finds classical ML models win 116 times while larger pretrained and LLM models win far fewer, showing predictive performance depends on model-task fit rather than scale.

  14. A Systematic Survey and Benchmark of Deep Learning for Molecular Property Prediction in the Foundation Model Era

    cs.LG 2026-04 accept novelty 5.0

    A systematic survey and benchmark of four deep learning paradigms for molecular property prediction that organizes the field, critiques current data practices, and outlines three future directions.

  15. CVT Archives and Chemical Embedding Measures for Multi-Objective Quality Diversity in Molecular Design

    physics.comp-ph 2026-04 unverdicted novelty 5.0

    CVT archives with learned chemical embeddings improve median global hypervolume and multi-objective quality diversity in NLO molecular design compared to grid-based archives.

  16. Machine learning for smell: Ordinal odor strength prediction of molecular perfumery components

    physics.chem-ph 2025-12 unverdicted novelty 5.0

    The authors compile an ordinal odor strength dataset for over 2,000 molecules from public sources and demonstrate supervised ML prediction of intensity categories, identifying molecular size, polarity, rings, and bran...

  17. Do Larger Models Really Win in Drug Discovery? A Benchmark Assessment of Model Scaling in AI-Driven Molecular Property and Activity Prediction

    cs.LG 2026-04 unverdicted novelty 4.0

    Large benchmark shows classical ML and GNNs outperform pretrained large models on most of 22 drug-discovery endpoints under strict cross-validation.