Pith. sign in

REVIEW 5 cited by

Molecule Attention Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.08264 v1 pith:CYNLGDAD submitted 2020-02-19 cs.LG physics.comp-phstat.ML

Molecule Attention Transformer

classification cs.LG physics.comp-phstat.ML
keywords attentionmoleculetaskstransformercompetitivelymolecularperformsprediction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Designing a single neural network architecture that performs competitively across a range of molecule property prediction tasks remains largely an open challenge, and its solution may unlock a widespread use of deep learning in the drug discovery industry. To move towards this goal, we propose Molecule Attention Transformer (MAT). Our key innovation is to augment the attention mechanism in Transformer using inter-atomic distances and the molecular graph structure. Experiments show that MAT performs competitively on a diverse set of molecular prediction tasks. Most importantly, with a simple self-supervised pretraining, MAT requires tuning of only a few hyperparameter values to achieve state-of-the-art performance on downstream tasks. Finally, we show that attention weights learned by MAT are interpretable from the chemical point of view.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Chem-GMNet: A Sphere-Native Geometric Transformer for Molecular Property Prediction

    cs.LG 2026-05 unverdicted novelty 7.0

    Chem-GMNet uses sphere-native embeddings, DualSKA attention, and SH-FFN layers to match or beat ChemBERTa-2 on MoleculeNet tasks with fewer parameters and sometimes no pretraining.

  2. Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

    cs.LG 2026-07 conditional novelty 6.0

    Aligning patient-specific gene-regulatory graphs with LINCS-pretrained perturbation embeddings via CLIP-style contrastive learning improves clinical drug-response prediction on TCGA and zero-shot I-SPY2.

  3. TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

    cs.AI 2026-05 unverdicted novelty 5.0

    TIGER is a text-informed ML framework that improves bidirectional enzyme-reaction retrieval by distilling semantic knowledge from sequences and aligning representations across tasks and distributions.

  4. A Systematic Survey and Benchmark of Deep Learning for Molecular Property Prediction in the Foundation Model Era

    cs.LG 2026-04 accept novelty 5.0

    A systematic survey and benchmark of four deep learning paradigms for molecular property prediction that organizes the field, critiques current data practices, and outlines three future directions.

  5. Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval

    q-bio.BM 2025-12 conditional novelty 5.0

    FGW-CLIP, a contrastive method that aligns enzymes and reactions while also aligning within-domain EC structure with a Gromov-Wasserstein regularizer, reports state-of-the-art retrieval on EnzymeMap and ReactZyme.