Pith. sign in

REVIEW 3 cited by

BERTology Meets Biology: Interpreting Attention in Protein Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.15222 v3 pith:Y5KBYUFX submitted 2020-06-26 cs.CL cs.LGq-bio.BM

classification cs.CLcs.LGq-bio.BM
keywords proteinattentionstructuretransformerarchitecturesmodelsproteinsrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer architectures have proven to learn useful representations for protein classification and generation tasks. However, these representations present challenges in interpretability. In this work, we demonstrate a set of methods for analyzing protein Transformer models through the lens of attention. We show that attention: (1) captures the folding structure of proteins, connecting amino acids that are far apart in the underlying sequence, but spatially close in the three-dimensional structure, (2) targets binding sites, a key functional component of proteins, and (3) focuses on progressively more complex biophysical properties with increasing layer depth. We find this behavior to be consistent across three Transformer architectures (BERT, ALBERT, XLNet) and two distinct protein datasets. We also present a three-dimensional visualization of the interaction between attention and protein structure. Code for visualization and analysis is available at https://github.com/salesforce/provis.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 28 citations worldwide. Full citation record

  1. Directed Evolution of Proteins via Bayesian Optimization in Embedding Space

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A Gaussian-process Bayesian optimizer running in protein language model embedding space outperforms regression-based directed evolution baselines on two in silico fitness landscapes.

  2. Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A knowledge-graph-guided preference optimization framework that fine-tunes protein language models to generate fewer sequences similar to known harmful proteins.

  3. Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

    cs.LG 2026-07 conditional novelty 5.0 of 10

    AP-REASONER, an affinity-propagation factor-graph sampler with alpha/beta knobs, improves MSA-based protein LM pretraining for contact prediction and conformational sampling, though gains are modest and partly baselin...

Pith tools