← back to paper
arxiv: 2606.27748 · 2 revisions
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel