Pith. sign in

The Senses: A Comprehensive Reference

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Transformers Learn Faster with Semantic Focus

cs.LG · 2025-06-17 · conditional · novelty 6.0

Input-dependent top-k sparse attention makes small transformers converge faster and generalize as well as full attention, while input-agnostic sparsity does not, and the effect is tied to reduced dispersion of attention scores.

citing papers explorer

Showing 1 of 1 citing paper.

  • Transformers Learn Faster with Semantic Focus cs.LG · 2025-06-17 · conditional · none · ref 6

    Input-dependent top-k sparse attention makes small transformers converge faster and generalize as well as full attention, while input-agnostic sparsity does not, and the effect is tied to reduced dispersion of attention scores.