Pith. sign in

Predicting Attention Sparsity in Transformers

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Transformers' quadratic complexity with respect to the input sequence length has motivated a body of work on efficient sparse approximations to softmax. An alternative path, used by entmax transformers, consists of having built-in exact sparse attention; however this approach still requires quadratic computation. In this paper, we propose Sparsefinder, a simple model trained to identify the sparsity pattern of entmax attention before computing it. We experiment with three variants of our method, based on distances, quantization, and clustering, on two tasks: machine translation (attention in the decoder) and masked language modeling (encoder-only). Our work provides a new angle to study model efficiency by doing extensive analysis of the tradeoff between the sparsity and recall of the predicted attention graph. This allows for detailed comparison between different models along their Pareto curves, important to guide future benchmarks for sparse attention models.

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Perspectives on Tsallis Statistics for Artificial Intelligence

cs.AI · 2026-08-02 · conditional · novelty 3.0

Independent AI methods, including sparsemax attention, Tsallis-entropy reinforcement learning, Student-t generative models, and robust losses, are instances of a single 'q-dial' deformation of Boltzmann-Gibbs statistics, with q best treated as a learnable parameter.

citing papers explorer

Showing 1 of 1 citing paper.

  • Perspectives on Tsallis Statistics for Artificial Intelligence cs.AI · 2026-08-02 · conditional · none · ref 35 · internal anchor

    Independent AI methods, including sparsemax attention, Tsallis-entropy reinforcement learning, Student-t generative models, and robust losses, are instances of a single 'q-dial' deformation of Boltzmann-Gibbs statistics, with q best treated as a learnable parameter.