REVIEW 32 references
Towards Understanding the Shape of Representations in Protein Language Models
T0 review · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Protein language models encode 3D structure best at very local residue distances (around 2 to 8 neighbors) and in layers just before the last, not in the final layer.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The authors use two geometric tools. First, they map both real protein structures and PLM representations to curves and compute shape-space statistics like the average spread and effective dimensionality. They find that the effective dimensionality rises in early layers and then shrinks in later layers, with larger models showing a stronger expansion. Second, they build graphs connecting nearby residues, both in the real 3D structure and in the PLM representation, and compare those graphs at different neighbor counts. This lets them ask: at what scale does the model's representation actually mirror the true protein structure?
The main empirical result is that PLM representations match real protein structure best when each residue is connected to only a few neighbors, around 2 to 8. The match is worst at larger context lengths. Also, the layer with the most structure-faithful representation is usually not the last layer, but one or a few layers before it. The authors suggest this means folding models should be attached to those earlier layers rather than the final output.
Extended reading notes
Core claim
The most load-bearing assertion is that 'PLMs preferentially encode immediate as well as local relations between residues, but start to degrade for larger context lengths. The most structurally faithful encoding tends to occur close to, but before the last layer of the models' (abstract). In Section 3.2 this is stated more specifically: 'in all models many of the curves have a bimodal shape throughout the filtration. This implies that PLM representations encode 3d protein structure at both a very local level, at about 2 neighbors, as well as at a slightly less local level, at about 8 neighbors.' If true, these are empirical regularities about how ESM2-style models organize structural information across layers and context scales.
Load-bearing premise
The central claim rests on the assumption that Euclidean k-nearest-neighbor graphs over PLM embedding coordinates are a meaningful proxy for residue context and structural encoding (Section 2.2). The graph filtration moment normalizes by distances to random point clouds R_i in R^m, but the paper never specifies how those random point clouds are sampled or whether they match the length, marginal distribution, or correlation structure of real PLM representations. If the PLM embedding space is not Euclidean in a way that preserves residue-neighbor relations, or if the random baseline is mismatched, the reported optimal context lengths (k=2 and k=8) and the entire 'structure is encoded' conclusion could be artifacts of the distance and normalization choices rather than properties of the model.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (2)
- interpolation_point_count =
1000
- random_point_cloud_baseline
assumptions (4)
- standard math The SRV representation with quadratic spline interpolation and SO(m) quotient provides a valid Riemannian shape space for comparing protein structures and PLM representations.
- domain assumption Euclidean distance between residue embedding vectors in PLM space is a meaningful notion of residue proximity for structural context.
- domain assumption Random point clouds in R^m are the correct null model for 'no structural encoding' in the graph filtration moment.
- domain assumption Quadratic spline interpolation of residue-level point clouds produces curves whose shape faithfully represents the biological or representational structure of a protein.
Cite this review
Pith. "Pith review of Towards Understanding the Shape of Representations in Protein Language Models." pith.science (2026). https://pith.science/paper/J7RNJLMM
@misc{pith2026250924895,
author = {Pith},
title = {Pith review of: Towards Understanding the Shape of Representations in Protein Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/J7RNJLMM}},
note = {Machine review of arXiv:2509.24895}
}
read the original abstract
While protein language models (PLMs) are one of the most promising avenues of research for future de novo protein design, the way in which they transform sequences to hidden representations, as well as the information encoded in such representations is yet to be fully understood. Several works have attempted to propose interpretability tools for PLMs, but they have focused on understanding how individual sequences are transformed by such models. Therefore, the way in which PLMs transform the whole space of sequences along with their relations is still unknown. In this work we attempt to understand this transformed space of sequences by identifying protein structure and representation with square-root velocity (SRV) representations and graph filtrations. Both approaches naturally lead to a metric space in which pairs of proteins or protein representations can be compared with each other. We analyze different types of proteins from the SCOP dataset and show that the Karcher mean and effective dimension of the SRV shape space follow a non-linear pattern as a function of the layers in ESM2 models of different sizes. Furthermore, we use graph filtrations as a tool to study the context lengths at which models encode the structural features of proteins. We find that PLMs preferentially encode immediate as well as local relations between residues, but start to degrade for larger context lengths. The most structurally faithful encoding tends to occur close to, but before the last layer of the models, indicating that training a folding model ontop of these layers might lead to improved folding performance.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
A robust tangent pca via shape restoration for shape variability analysis
Michel Abboud, Abdesslam Benzinou, and Kamal Nasreddine. A robust tangent pca via shape restoration for shape variability analysis. Pattern Analysis and Applications, 23 0 (2): 0 653--671, 2020
2020
-
[3]
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255, 2020
arXiv 2012
-
[4]
The gene ontology knowledgebase in 2023
Suzi A Aleksander, James Balhoff, Seth Carbon, J Michael Cherry, Harold J Drabkin, Dustin Ebert, Marc Feuermann, Pascale Gaudet, Nomi L Harris, et al. The gene ontology knowledgebase in 2023. Genetics, 224 0 (1): 0 iyad031, 2023
2023
-
[5]
Accurate prediction of protein structures and interactions using a three-track neural network
Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373 0 (6557): 0 871--876, 2021
2021
-
[6]
Peptide binder design with inverse folding and protein structure prediction
Patrick Bryant and Arne Elofsson. Peptide binder design with inverse folding and protein structure prediction. Communications Chemistry, 6 0 (1): 0 229, 2023
2023
-
[7]
Scope: improvements to the structural classification of proteins--extended database to facilitate variant interpretation and machine learning
John-Marc Chandonia, Lindsey Guan, Shiangyi Lin, Changhua Yu, Naomi K Fox, and Steven E Brenner. Scope: improvements to the structural classification of proteins--extended database to facilitate variant interpretation and machine learning. Nucleic acids research, 50 0 (D1): 0 D553--D559, 2022
2022
-
[8]
Target sequence-conditioned design of peptide binders using masked language modeling
Leo Tianlai Chen, Zachary Quinn, Madeleine Dumas, Christina Peng, Lauren Hong, Moises Lopez-Gonzalez, Alexander Mestre, Rio Watson, Sophia Vincoff, Lin Zhao, et al. Target sequence-conditioned design of peptide binders using masked language modeling. Nature Biotechnology, pp.\ 1--9, 2025
2025
Show all 32 references
-
[9]
Emergence of a high-dimensional abstraction phase in language transformers
Emily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco, Jade Yu, Alessandro Laio, and Marco Baroni. Emergence of a high-dimensional abstraction phase in language transformers. arXiv preprint arXiv:2405.15471, 2024
2024 arXiv
-
[10]
Computational Topology : An Introduction
Herbert Edelsbrunner and John Harer. Computational Topology : An Introduction . American Mathematical Soc., 2010
2010
-
[11]
Controllable protein design with language models
Noelia Ferruz and Birte H \"o cker. Controllable protein design with language models. Nature Machine Intelligence, 4 0 (6): 0 521--532, 2022
2022
-
[12]
Se (3)-transformers: 3d roto-translation equivariant attention networks
Fabian Fuchs, Daniel Worrall, Volker Fischer, and Max Welling. Se (3)-transformers: 3d roto-translation equivariant attention networks. Advances in neural information processing systems, 33: 0 1970--1981, 2020
1970
-
[13]
Sparse autoencoders uncover biologically interpretable features in protein language model representations
Onkar Gujral, Mihir Bafna, Eric Alm, and Bonnie Berger. Sparse autoencoders uncover biologically interpretable features in protein language model representations. Proceedings of the National Academy of Sciences, 122 0 (34): 0 e2506316122, 2025
2025
-
[14]
Simulating 500 million years of evolution with a language model
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q Tran, Jonathan Deaton, Marius Wiggert, et al. Simulating 500 million years of evolution with a language model. Science, 387 0 (6736): 0 850--858, 2025
2025
-
[15]
Graph filtration learning
Christoph Hofer, Florian Graf, Bastian Rieck, Marc Niethammer, and Roland Kwitt. Graph filtration learning. In International Conference on Machine Learning, pp.\ 4314--4323. PMLR, 2020
2020
-
[16]
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \' dek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596 0 (7873): 0 583--589, 2021
2021
-
[17]
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Measuring the intrinsic dimension of objective landscapes. arXiv preprint arXiv:1804.08838, 2018
2018 arXiv
-
[18]
Fatcat 2.0: towards a better understanding of the structural diversity of proteins
Zhanwen Li, Lukasz Jaroszewski, Mallika Iyer, Mayya Sedova, and Adam Godzik. Fatcat 2.0: towards a better understanding of the structural diversity of proteins. Nucleic acids research, 48 0 (W1): 0 W60--W64, 2020
2020
-
[19]
Evolutionary-scale prediction of atomic-level protein structure with a language model
Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379 0 (6637): 0 1123--1130, 2023
2023
-
[20]
Protein structure alignment using elastic shape analysis
Wei Liu, Anuj Srivastava, and Jinfeng Zhang. Protein structure alignment using elastic shape analysis. In Proceedings of the First ACM International Conference on Bioinformatics and Computational Biology, pp.\ 62--70, 2010
2010
-
[21]
Variational autoencoder for design of synthetic viral vector serotypes
Suyue Lyu, Shahin Sowlati-Hashjin, and Michael Garton. Variational autoencoder for design of synthetic viral vector serotypes. Nature Machine Intelligence, 6 0 (2): 0 147--160, 2024
2024
-
[22]
Large language models generate functional protein sequences across diverse families
Ali Madani, Ben Krause, Eric R Greene, Subu Subramanian, Benjamin P Mohr, James M Holton, Jose Luis Olmos Jr, Caiming Xiong, Zachary Z Sun, Richard Socher, et al. Large language models generate functional protein sequences across diverse families. Nature biotechnology, 41 0 (8...
2023
-
[23]
Exploring the impact of a transformer's latent space geometry on downstream task performance
Anna C Marbut, John W Chandler, and Travis J Wheeler. Exploring the impact of a transformer's latent space geometry on downstream task performance. arXiv preprint arXiv:2406.12159, 2024
2024 arXiv
-
[24]
Geomstats: A python package for riemannian geometry in machine learning
Nina Miolane, Nicolas Guigui, Alice Le Brigant, Johan Mathe, Benjamin Hou, Yann Thanwerdas, Stefan Heyder, Olivier Peltre, Niklas Koep, Hadi Zaatiti, Hatem Hajri, Yann Cabanes, Thomas Gerald, Paul Chauchat, Christian Shewmake, Daniel Brooks, Bernhard Kainz, Claire Donnat, Susa...
2020
-
[25]
Filtration curves for graph representation
Leslie O'Bray, Bastian Rieck, and Karsten Borgwardt. Filtration curves for graph representation. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp.\ 1267--1275, 2021
2021
-
[26]
Less is more: Local intrinsic dimensions of contextual language models
Benjamin Matthias Ruppik, Julius von Rohrscheidt, Carel van Niekerk, Michael Heck, Renato Vukovic, Shutong Feng, Hsien-chin Lin, Nurul Lubis, Bastian Rieck, Marcus Zibrowius, et al. Less is more: Local intrinsic dimensions of contextual language models. arXiv preprint arXiv:25...
2025
-
[27]
Interplm: Discovering interpretable features in protein language models via sparse autoencoders
Elana Simon and James Zou. Interplm: Discovering interpretable features in protein language models via sparse autoencoders. bioRxiv, pp.\ 2024--11, 2024
2024
-
[28]
Shape analysis of elastic curves in euclidean spaces
Anuj Srivastava, Eric Klassen, Shantanu H Joshi, and Ian H Jermyn. Shape analysis of elastic curves in euclidean spaces. IEEE transactions on pattern analysis and machine intelligence, 33 0 (7): 0 1415--1428, 2010
2010
-
[29]
The geometry of hidden representations of large transformer models
Lucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio, Alessio Ansuini, and Alberto Cazzaniga. The geometry of hidden representations of large transformer models. Advances in Neural Information Processing Systems, 36: 0 51234--51252, 2023
2023
-
[30]
Artificial intelligence using a latent diffusion model enables the generation of diverse and potent antimicrobial peptides
Yeji Wang, Minghui Song, Fujing Liu, Zhen Liang, Rui Hong, Yuemei Dong, Huaizu Luan, Xiaojie Fu, Wenchang Yuan, Wenjie Fang, et al. Artificial intelligence using a latent diffusion model enables the generation of diverse and potent antimicrobial peptides. Science Advances, 11 ...
2025
-
[31]
Scoring function for automated assessment of protein structure template quality
Yang Zhang and Jeffrey Skolnick. Scoring function for automated assessment of protein structure template quality. Proteins: Structure, Function, and Bioinformatics, 57 0 (4): 0 702--710, 2004
2004
-
[32]
Protein language models learn evolutionary statistics of interacting sequence motifs
Zhidian Zhang, Hannah K Wayment-Steele, Garyk Brixi, Haobo Wang, Dorothee Kern, and Sergey Ovchinnikov. Protein language models learn evolutionary statistics of interacting sequence motifs. Proceedings of the National Academy of Sciences, 121 0 (45): 0 e2406285121, 2024
2024
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.