Pith. sign in

REVIEW 4 cited by

Towards Interpretable Protein Structure Prediction with Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08764 v1 pith:JW24JOYS submitted 2025-03-11 q-bio.BM cs.AIcs.LG

classification q-bio.BMcs.AIcs.LG
keywords predictionstructuremodelsproteinsaeslanguageautoencodersesm2-3b
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Protein language models have revolutionized structure prediction, but their nonlinear nature obscures how sequence representations inform structure prediction. While sparse autoencoders (SAEs) offer a path to interpretability here by learning linear representations in high-dimensional space, their application has been limited to smaller protein language models unable to perform structure prediction. In this work, we make two key advances: (1) we scale SAEs to ESM2-3B, the base model for ESMFold, enabling mechanistic interpretability of protein structure prediction for the first time, and (2) we adapt Matryoshka SAEs for protein language models, which learn hierarchically organized features by forcing nested groups of latents to reconstruct inputs independently. We demonstrate that our Matryoshka SAEs achieve comparable or better performance than standard architectures. Through comprehensive evaluations, we show that SAEs trained on ESM2-3B significantly outperform those trained on smaller models for both biological concept discovery and contact map prediction. Finally, we present an initial case study demonstrating how our approach enables targeted steering of ESMFold predictions, increasing structure solvent accessibility while fixing the input sequence. To facilitate further investigation by the broader community, we open-source our code, dataset, pretrained models https://github.com/johnyang101/reticular-sae , and visualizer https://sae.reticular.ai .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Protein language models complete repeats by combining induction heads that copy the aligned residue from the other repeat copy with neurons encoding amino-acid similarity; the approximate-repeat circuit contains and g...

  2. Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Causal interventions show ESMFold's folding trunk first transfers sequence chemistry into its pairwise representation (blocks 0–7), then builds pairwise spatial features that control output geometry (blocks 25+).

  3. Design-CP: Context Parallelism for Design of Protein Nanoparticles

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Context-parallel inference for RFdiffusion 3 enables end-to-end all-atom design of large symmetric protein nanoparticles on multi-GPU hardware without retraining.

  4. Mechanistic Interpretability of Antibody Language Models Using SAEs

    cs.LG 2025-12 unverdicted novelty 6.0 of 10

    TopK SAEs uncover biologically meaningful latent features in antibody language models without guaranteeing causal steering, whereas Ordered SAEs provide reliable generative control at the cost of complex activation patterns.

Pith tools