Performers approximate full-rank softmax attention in Transformers via FAVOR+ random features for linear complexity, with theoretical guarantees of unbiased estimation and competitive results on pixel, text, and protein tasks.
BERTology Meets Biology: Interpreting Attention in Protein Language Models
3 Pith papers cite this work, alongside 28 external citations. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
ESM2 predicts N-terminal methionine via retrieval of a positional prior from the BOS token through distributed attention circuits rather than direct recognition, revealed by a norm-direction decomposition of rotary attention scores.
A literature review synthesizing generative AI uses in healthcare with emphasis on diffusion and transformer models.
citing papers explorer
-
Rethinking Attention with Performers
Performers approximate full-rank softmax attention in Transformers via FAVOR+ random features for linear complexity, with theoretical guarantees of unbiased estimation and competitive results on pixel, text, and protein tasks.
-
Retrieval and competition: how a protein foundation model starts a protein
ESM2 predicts N-terminal methionine via retrieval of a positional prior from the BOS token through distributed attention circuits rather than direct recognition, revealed by a norm-direction decomposition of rotary attention scores.
-
Recent Advances in Generative AI for Healthcare Applications
A literature review synthesizing generative AI uses in healthcare with emphasis on diffusion and transformer models.