Flexformer learns attention kernels by treating spectral frequencies as trainable parameters in random Fourier feature-based linear attention, with stationary and nonstationary variants that outperform fixed-kernel baselines.
Long range arena: A benchmark for efficient transformers
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Complex Diffusion Maps with ω-parameterized complex kernels reveal inherent harmonic representations and improve discrimination over real-valued diffusion methods on synthetic data, fMRI, and EEG.
citing papers explorer
-
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Flexformer learns attention kernels by treating spectral frequencies as trainable parameters in random Fourier feature-based linear attention, with stationary and nonstationary variants that outperform fixed-kernel baselines.
-
Complex Diffusion Maps with $\omega$-Parameterized Kernels Revealing Inherent Harmonic Representations
Complex Diffusion Maps with ω-parameterized complex kernels reveal inherent harmonic representations and improve discrimination over real-valued diffusion methods on synthetic data, fMRI, and EEG.