Pith. sign in

REVIEW 3 major objections 6 minor 86 references

Converting Transformers into DGNNs Form

T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read A neural layer that replaces self-attention with a learned digraph Fourier convolution beats all fifteen Transformer baselines on Long-Range Arena, averaging 75.94% accuracy.

desk verdict The LRA gain is unverified and likely an artifact of copied baselines; the Givens-rotation + kernel polynomial method is a real idea worth testing. read the letter →

arxiv 2502.00585 v3 pith:6DIS53JX submitted 2025-02-01 cs.LG cs.CL

classification cs.LGcs.CL MSC 68T07
keywords self-attentionreplacementdigraphconvolutionFouriertransformunitarymatrixparametrizationkernelpolynomialmethodLong-RangeArenaefficienttransformerslongsequencemodeling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the self-attention mechanism inside a Transformer can be replaced outright with a directed-graph convolution, turning the whole architecture into a directed graph neural network form. The replacement, called Synvolution, and its kernel-filtered variant Kernelution, synthesize a learnable unitary digraph shift operator that plays the role of the attention matrix, with linearithmic time complexity and a dense, full-rank mixing matrix. The authors test the resulting model, Converter, on Long-Range Arena, long-document classification, and DNA taxonomy classification, where it reports the best accuracy among all compared models, including a 75.94% Long-Range Arena average that beats the second-best model by about fourteen points. If the results are right, a simple and cheap spectral layer can outperform the quadratic softmax attention that anchors modern Transformers.

What carries the argument

The load-bearing object is Synvolution, a learnable unitary digraph convolution defined as $\mathrm{Synv}(XW_V) = \Phi^{-1}\bigl(\exp(i\Lambda) \odot (\Phi XW_V)\bigr)$, where $\Phi$ is a synthetic unitary eigenvector matrix and $\exp(i\Lambda)$ is a diagonal matrix of synthetic eigenvalues; together they act as the digraph shift operator and frequency response of a digraph Fourier transform. The matrix $\Phi$ is built by the order-1 DHHP parametrization, a unitary diagonal matrix times one lower-unitary Hessenberg factor times one upper-unitary Hessenberg factor assembled from Givens rotations, so the transform runs in $O(N \log N)$ time via parallel scan rather than $O(N^2)$ matrix multiplication. Kernelution replaces the raw eigenvalue multiplier with a kernel-polynomial filter $p_{\mathrm{KP}}(\Lambda)$ whose learnable Chebyshev coefficients are penalized by the kernel polynomial loss, which simulates a data-dependent, Gibbs-damped kernel. A Gated Feed-Forward Network maps the complex-valued output back to real values, and PostScaleNorm stabilizes training.

What would settle it

Re-run all fifteen Long-Range Arena baselines under Converter's own protocol, the same Bayesian tuning budget, seeds, tokenization, and normalization, and check whether the 75.94% average and the fourteen-point gap over Luna-256 survive. Separately, measure the reconstruction error of 1-DHHP against random unitary and DFT matrices of growing size to test whether the Givens-Hessenberg product actually covers the unitary group as Assumption 1 claims.

Watch

Extended reading notes

Core claim

The paper's central claim is that a Transformer whose self-attention is swapped for a synthetic unitary digraph convolution outperforms the vanilla Transformer and fourteen efficient-Transformer variants on every one of the five Long-Range Arena tasks, with scores of 60.38% on ListOps, 86.44% on Text, 83.41% on Retrieval, 61.02% on Image, and 88.43% on Pathfinder, a 75.94% average that exceeds the second-best model, Luna-256 at 61.95%, by roughly fourteen points. The same Converter reports the best accuracy on arXiv-document classification at 16K and 32K tokens (81.77% and 82.34%) and on DNA taxonomy classification (84.59% on the Bos/Sus task and 59.49% on the Mus/Rattus task). The paper reads these results as evidence that the softmax similarity bottleneck is not necessary: a full-rank, dense, learnable mixing operator built from Givens rotations and spectral filtering can carry long-range dependency modeling by itself. Its ablations add that the GRU-based relative position embedding and the kernel polynomial loss each contribute substantially to the final accuracy.

Load-bearing premise

The headline performance claim depends on the fifteen Long-Range Arena baselines being tuned and evaluated under the same protocol as Converter, which the paper does not document (their numbers match published results), and the theoretical claim depends on the unproven Assumption 1 that at most $\lceil N/4 \rceil$ orders of the L-DHHP rotation product can represent any dense unitary matrix.

Editorial extensions

If this is right

  • Self-attention is not required for Transformer-level performance: a linearithmic, full-rank, dense spectral mixing layer suffices on sequences up to 32K tokens.
  • The kernel polynomial method can stand in for the multi-head operation, with the new kernel polynomial loss supplying a principled, order-increasing penalty that mimics adaptive Gibbs damping.
  • Because the mixing matrix is unitary and full-rank by construction, Converter's layer avoids the low-rank bottleneck that limits softmax and kernel attention as depth grows.
  • The same spectral layer transfers across raw text, flattened pixels, long documents, and DNA sequences, indicating the construction is domain-agnostic rather than tuned to one data type.
  • A single-head spectral layer with linearithmic cost is enough to beat models built around sparse, low-rank, or kernelized attention on the tested benchmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the Table 1 baselines were re-run under Converter's exact tuning protocol, the fourteen-point Long-Range Arena margin could shrink; the residual gap on ListOps and Image would then be the honest measure of the mechanism's advantage.
  • The 1-DHHP fast transform is a general structured unitary layer, so it could slot into other quadratic-complexity positions such as cross-attention decoders or state-space mixers, which the paper names as future work for cross-attention.
  • The kernel polynomial loss ties regularization strength to filter order; letting the order $K$ itself be scheduled or learned during training would be a natural extension from coarse to fine spectral resolution.
  • Because the unitary constraint keeps the spectrum on the unit circle, deeper Converter stacks are plausibly more stable than unconstrained attention; a controlled depth-scaling study would test whether the spectral construction delays rank collapse.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes Converter, a Transformer variant in which the self-attention module is replaced by a synthetic unitary digraph convolution called Synvolution, built on the order-L DHHP parametrization of unitary matrices and a learned diagonal of eigenvalues. The authors also introduce Kernelution, which applies the kernel polynomial method to the spectral filter, a kernel polynomial loss, and a gated feed-forward network for complex-valued activations. The paper claims linearithmic time complexity, full-rank dense mixing, and state-of-the-art accuracy on the Long-Range Arena benchmark, long document classification, and DNA taxonomy classification, with an average LRA accuracy of 75.94% versus 61.95% for the best baseline.

Significance. If the empirical results are reproducible under a matched protocol, this is a significant contribution: a simple, full-rank, dense attention alternative with linearithmic complexity that outperforms a wide range of Transformer variants on long-sequence benchmarks would be of broad interest. The theoretical framing via digraph signal processing is novel, and the L-DHHP construction with a fast parallel-scan implementation is a concrete algorithmic contribution. The paper also includes useful ablations showing the contribution of RPE and the kernel polynomial loss. However, the main significance currently rests on an empirical comparison whose fairness is not documented, and the central theoretical proposition is contingent on an unproven assumption.

major comments (3)
  1. [§5.1 and Appendix F.1] The headline claim that Converter surpasses all 14 baselines on all five LRA tasks by a 14-point average margin depends entirely on Table 1. Appendix F.1 states that 'all four datasets and seven baseline models' were tuned with Bayesian optimization under a 16 GB memory constraint, but Table 1 contains five LRA tasks and fifteen models, and Tables 7 and 8 provide no LRA baseline hyperparameters. The reported baseline numbers closely match those in Tay et al. (2021b), which suggests they were transcribed from the literature rather than re-run under the paper's protocol. As presented, the comparison is not validated: the tuning budget, seeds, preprocessing, and evaluation procedure for the baselines are unknown. Please re-run all baselines under the same protocol, report their hyperparameters, and provide the provenance of each value in Table 1, or substantially soften the superiority claim.
  2. [Appendix E, proof of Proposition 2] Proposition 2 is not derived but merely restated: the proof says that since DFT, DWHT, DCT, and DST are unitary matrices, Assumption 1 implies they can be represented by L-DHHP. Because Assumption 1 already asserts that L-DHHP can construct every dense unitary matrix, the proposition carries no independent content. The assumption itself is nontrivial and unproven, and the paper gives no constructive procedure or numerical evidence that ⌈N/4⌉ orders suffice. Please either prove Assumption 1 or a weaker version that covers the specific transforms, or explicitly mark the theoretical claims as conditional.
  3. [§5 and Appendix F.1] No seed information or error bars are reported anywhere in the experiments, and no code repository link is provided. LRA results are known to be sensitive to initialization and tuning, so single-run accuracy values without seeds are insufficient to support a 14-point margin. Please report the number of runs, the variance, and release the code with a clear reproducibility script; this is a load-bearing issue for the empirical claims.
minor comments (6)
  1. [§4.2] The text uses both 'Kernolution' and 'Kernelution' (and later 'Synolution'); please standardize the spelling.
  2. [§5.4] The ablation text says 'Converter achieves the highest performance when using PRE, followed by APE and SPE', but the table shows no 'PRE' row; this should be 'RPE', the recurrent position embedding used elsewhere.
  3. [§5.1] The text describes LRA as containing 'five multi-class classification tasks', but Text and Retrieval are binary classification tasks; please correct the wording.
  4. [Appendix F.1] The sentence 'For all four datasets and seven baseline models' contradicts the actual experimental scope; the appendix should accurately state the number of datasets and baselines.
  5. [References] The reference list contains a typo in 'Elena V oita' (a stray space) and should be corrected.
  6. [Equation (7)] The product notation in Equation (7) is difficult to read and appears malformed in the preprint; please use explicit index bounds or a clearer notation.

Circularity Check

1 steps flagged · score 2.0 of 10

The only reduce-by-construction step is Proposition 2, which is a direct instantiation of unproven Assumption 1; the empirical comparison is external and not circular.

  1. other [Section 4.1 (Assumption 1 / Proposition 2) and Appendix E (Proof of Proposition 2)]
    "Since our method is based on the Givens rotation method, we make Assumption 1. Under this assumption, we can establish the following propositions. Assumption 1. For constructing an arbitrary N × N dense unitary matrix, at most ⌈ N/4 ⌉ orders are sufficient for L-DHHP. Proposition 2. L-DHHP captures the discrete unitary transforms, including discrete Fourier transform (DFT), the discrete Walsh–Hadamard transform (DWHT), the discrete cosine transform (DCT), the discrete sine transform (DST), and their inverses exactly."

    The proof in Appendix E reduces Proposition 2 to Assumption 1: 'By Assumption 1, any N × N unitary matrix can be exactly constructed using L-DHHP... Therefore, as specific cases of unitary matrices, DFT, DWHT, DCT, and DST can all be exactly represented by L-DHHP.' Since Assumption 1 already asserts that every dense unitary matrix is representable by L-DHHP, the proposition is a one-line instantiation of the assumption together with the definitional fact that these transforms are unitary. The claimed expressive-completeness result is assumed rather than independently derived. The paper is transparent that this is conditional, so this is a minor tautological step rather than a hidden circularity, but the proposition adds no evidence beyond the assumption.

full rationale

The central performance claims are tested against external benchmarks (LRA, long-document, and DNA taxonomy), so they are not circular: no baseline accuracy or target value is used as a fitted input, and no parameter fitted to one subset is relabeled as a prediction on a closely related subset. There are no load-bearing self-citations by the authors. The only step that reduces to its own input by construction is Proposition 2, which is a direct corollary of unproven Assumption 1; however, the empirical results do not depend on this proposition, and the paper explicitly labels the premise as an assumption. Protocol concerns about LRA baseline tuning (Appendix F.1) are a correctness and validity risk, not a circularity. Overall score 2 reflects this single auxiliary tautological theoretical step while the main empirical claim is independently benchmarked.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No new physical entities are posited. The novel mathematical objects, Synvolution and Kernelution, are model components rather than entities with independent evidence. The free parameters listed are the hand-chosen constants that directly affect the reported results.

free parameters (3)
  • K, maximum Chebyshev order = 2 (all tasks)
    Chosen by hand; the kernel polynomial expansion order controls the filter approximation.
  • eta, KPL weight = 0.001 for ListOps, Text, Retrieval, Pathfinder; 0.01 for Image; 0.1 for LongDoc and Ensembl
    Tunable hyperparameter connecting cross-entropy and kernel polynomial loss, selected per dataset.
  • DHHP order L = 1
    The paper sets L=1 for all experiments, restricting the unitary family but reducing complexity.
assumptions (4)
  • ad hoc to paper Assumption 1: at most ceil(N/4) orders of L-DHHP suffice for any N x N dense unitary matrix
    Unproven; underpins Proposition 2, which claims L-DHHP captures DFT, DWHT, DCT, DST and their inverses exactly.
  • standard math Chebyshev polynomial interpolation theorems for differentiable and analytic functions
    Used to justify that the Chebyshev filter can approximate arbitrary continuous filters in Kernelution.
  • standard math Kernel polynomial method with Gibbs damping factors
    Basis for replacing the multi-head operation with polynomial filtering and for the kernel polynomial loss, sourced from Weiße et al.
  • standard math LQ decomposition of a square matrix into a lower triangular matrix and Givens rotations
    Foundation of the LHHP and DHHP parameterizations, from Givens (1958).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Converting Transformers into DGNNs Form." pith.science (2026). https://pith.science/paper/6DIS53JX

@misc{pith2026250200585,
  author       = {Pith},
  title        = {Pith review of: Converting Transformers into DGNNs Form},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6DIS53JX}},
  note         = {Machine review of arXiv:2502.00585}
}
read the original abstract

Recent advances in deep learning have established Transformer architectures as the predominant modeling paradigm. Central to the success of Transformers is the self-attention mechanism, which scores the similarity between query and key matrices to modulate a value matrix. This operation bears striking similarities to digraph convolution, prompting an investigation into whether digraph convolution could serve as an alternative to self-attention. In this study, we formalize this concept by introducing a synthetic unitary digraph convolution based on the digraph Fourier transform. The resulting model, which we term Converter, effectively converts a Transformer into a Directed Graph Neural Network (DGNN) form. We have tested Converter on Long-Range Arena benchmark, long document classification, and DNA sequence-based taxonomy classification. Our experimental results demonstrate that Converter achieves superior performance while maintaining computational efficiency and architectural simplicity, which establishes it as a lightweight yet powerful Transformer variant.

Figures

Figures reproduced from arXiv: 2502.00585 by the authors.

Figure 1
Figure 1. Converter architecture. head, resulting in unidirectional symmetric self-attention. Similarly, for digraph convolution, the achievement of graph convolution can be implemented by symmetrizing the adjacency matrix of a digraph. Third, the softmax function in self-attention results in a row-wise normalized digraph adjacency form. Since digraph convolution closely resembles self-attention, we investigated replacing sel… view at source ↗
Figure 2
Figure 2. Illustration of the entire Kernelution process. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. An illustration of Gibbs phenomenon when using the kernel polynomial method with different kernels to [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

86 extracted references · 74 canonical work pages

  1. [1]

    Attention Is All You Need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention Is All You Need . In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30, pages 5998--6008. Curran Associates, I...

  2. [2]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, V...

  3. [3]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale . In International Conference on Learning Representations, 2021

  4. [4]

    Transforming the Language of Life: Transformer Neural Networks for Protein Prediction Tasks

    Ananthan Nambiar, Maeve Heflin, Simon Liu, Sergei Maslov, Mark Hopkins, and Anna Ritz. Transforming the Language of Life: Transformer Neural Networks for Protein Prediction Tasks . In Proceedings of the 11th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics. Association for Computing Machinery, 2020

  5. [5]

    Transformer Dissection: An Unified Understanding for Transformer ' s Attention via the Lens of Kernel

    Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. Transformer Dissection: An Unified Understanding for Transformer ' s Attention via the Lens of Kernel . In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing...

  6. [6]

    Rethinking Attention with Performers

    Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. Rethinking Attention with Performers . In International Conference on Learning Representations, 2021

  7. [7]

    cosFormer: Rethinking Softmax In Attention

    Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. cosFormer: Rethinking Softmax In Attention . In International Conference on Learning Representations, 2022

  8. [8]

    Attention is not all you need: pure attention loses rank doubly exponentially with depth

    Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 2793--2803. PMLR, 07 2021

Show all 86 references
  1. [9]

    Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. Breaking the softmax bottleneck: A high-rank RNN language model. In International Conference on Learning Representations, 2018

  2. [10]

    softmax is not enough (for sharp out-of-distribution)

    Petar Veli c kovi\' c , Christos Perivolaropoulos, Federico Barbero, and Razvan Pascanu. softmax is not enough (for sharp out-of-distribution). arXiv preprint arXiv: 2410.01104, 2024

  3. [11]

    Synthesizer: Rethinking Self-Attention for Transformer Models

    Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. Synthesizer: Rethinking Self-Attention for Transformer Models . In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings ...

  4. [12]

    FNet: Mixing Tokens with Fourier Transforms

    James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. FNet: Mixing Tokens with Fourier Transforms . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4296--43...

  5. [13]

    Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-Attention

    Tong Yu, Ruslan Khalitov, Lei Cheng, and Zhirong Yang. Paramixer: Parameterizing Mixing Links in Sparse Factors Works Better than Dot-Product Self-Attention . In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 681--690, 2022 a

  6. [14]

    Big Bird: Transformers for Longer Sequences

    Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big Bird: Transformers for Longer Sequences . In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Li...

  7. [15]

    Silver and H

    R.N. Silver and H. R \" o der. Densities of States of Mega-Dimensional Hamiltonian Matrices . International Journal of Modern Physics C, 05 0 (04): 0 735--753, 1994

  8. [16]

    Calculating the density of states and optical-absorption spectra of large quantum systems by the plane-wave moments method

    Lin-Wang Wang. Calculating the density of states and optical-absorption spectra of large quantum systems by the plane-wave moments method. Physical Review B, 49: 0 10154--10158, 04 1994

  9. [17]

    Dielectric Constants of Silicon Quantum Dots

    Lin-Wang Wang and Alex Zunger. Dielectric Constants of Silicon Quantum Dots . Physical Review Letters, 73: 0 1039--1042, 08 1994

  10. [18]

    Kouri, and David K

    Amrendra Vijay, Donald J. Kouri, and David K. Hoffman. Scattering and Bound States: A Lorentzian Function-Based Spectral Filter Approach . The Journal of Physical Chemistry A, 108 0 (41): 0 8987--9003, 10 2004

  11. [19]

    The kernel polynomial method

    Alexander Wei e, Gerhard Wellein, Andreas Alvermann, and Holger Fehske. The kernel polynomial method. Reviews of Modern Physics, 78: 0 275--306, 05 2006

  12. [20]

    Chebyshev Expansion Techniques , pages 545--577

    Alexander Wei e and Holger Fehske. Chebyshev Expansion Techniques , pages 545--577. Springer Berlin Heidelberg, 2008

  13. [21]

    Long Range Arena: A Benchmark for Efficient Transformers

    Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. Long Range Arena: A Benchmark for Efficient Transformers . In International Conference on Learning Representations, 2021 b

  14. [22]

    Generating Long Sequences with Sparse Transformers

    Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating Long Sequences with Sparse Transformers . arXiv preprint arXiv: 1904.10509, 2019

  15. [23]

    Reformer: The Efficient Transformer

    Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The Efficient Transformer . In International Conference on Learning Representations, 2020

  16. [24]

    Scatterbrain: Unifying sparse and low-rank attention

    Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra, and Christopher R\' e . Scatterbrain: Unifying sparse and low-rank attention. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, vol...

  17. [25]

    MetaFormer Baselines for Vision

    Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. MetaFormer Baselines for Vision . arXiv preprint arXiv: 2210.13452, 2022 b

  18. [26]

    MetaFormer Is Actually What You Need for Vision

    Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. MetaFormer Is Actually What You Need for Vision . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10819--10829, 06 2022 c

  19. [27]

    Are Sixteen Heads Really Better than One? In H

    Paul Michel, Omer Levy, and Graham Neubig. Are Sixteen Heads Really Better than One? In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32, pages 14014--14024. Curran Asso...

  20. [28]

    Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

    Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797...

  21. [29]

    Multi-Head Attention: Collaborate Instead of Concatenate

    Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. Multi-Head Attention: Collaborate Instead of Concatenate . arXiv preprint arXiv: 2006.16362, 2020

  22. [30]

    Low-Rank Bottleneck in Multi-head Attention Models

    Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. Low-Rank Bottleneck in Multi-head Attention Models . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Procee...

  23. [31]

    Graph filters for signal processing and machine learning on graphs

    Elvin Isufi, Fernando Gama, David I Shuman, and Santiago Segarra. Graph filters for signal processing and machine learning on graphs. IEEE Transactions on Signal Processing, pages 1--32, 2024

  24. [32]

    Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete Signal Processing on Graphs . IEEE Transactions on Signal Processing, 61 0 (7): 0 1644--1656, 2013 a

  25. [33]

    Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete signal processing on graphs: Graph fourier transform . In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6167--6170, 2013 b

  26. [34]

    Aliaksei Sandryhaila and Jos \' e M. F. Moura. Discrete Signal Processing on Graphs: Frequency Analysis . IEEE Transactions on Signal Processing, 62 0 (12): 0 3042--3054, 2014

  27. [35]

    Rahul Singh, Abhishek Chakraborty, and B. S. Manoj. Graph Fourier transform based on directed Laplacian . In 2016 International Conference on Signal Processing and Communications (SPCOM), pages 1--5, 2016

  28. [36]

    Laplacians and the cheeger inequality for directed graphs

    Fan Chung. Laplacians and the cheeger inequality for directed graphs. Annals of Combinatorics, 9 0 (1): 0 1--19, 04 2005

  29. [37]

    Ala \' i z, and Johan A

    Micha \" e l Fanuel, Carlos M. Ala \' i z, and Johan A. K. Suykens. Magnetic eigenmaps for community detection in directed networks. Physical Review E, 95: 0 022302, 02 2017

  30. [38]

    Ala \' i z, \' A ngela Fern \' a ndez, and Johan A.K

    Micha \" e l Fanuel, Carlos M. Ala \' i z, \' A ngela Fern \' a ndez, and Johan A.K. Suykens. Magnetic Eigenmaps for the visualization of directed networks . Applied and Computational Harmonic Analysis, 44 0 (1): 0 189--199, 2018

  31. [39]

    Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform

    Nir Ailon and Bernard Chazelle. Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform . In Proceedings of the Thirty-Eighth Annual ACM Symposium on Theory of Computing, STOC '06, pages 557--563. Association for Computing Machinery, 2006

  32. [40]

    A sparse Johnson: Lindenstrauss transform

    Anirban Dasgupta, Ravi Kumar, and Tam\' a s Sarlos. A sparse Johnson: Lindenstrauss transform . In Proceedings of the Forty-Second ACM Symposium on Theory of Computing, STOC '10, pages 341--350. Association for Computing Machinery, 2010

  33. [41]

    Fastfood — Approximating Kernel Expansions in Loglinear Time

    Quoc Le, Tamas Sarlos, and Alexander Smola. Fastfood — Approximating Kernel Expansions in Loglinear Time . In Sanjoy Dasgupta and David McAllester, editors, Proceedings of the 30th International Conference on Machine Learning, volume 28 of Proceedings of Machine Learning Resea...

  34. [42]

    Orthogonal Random Features

    Felix Xinnan X Yu, Ananda Theertha Suresh, Krzysztof M Choromanski, Daniel N Holtmann-Rice, and Sanjiv Kumar. Orthogonal Random Features . In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29, pages...

  35. [43]

    Deep Fried Convnets

    Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang. Deep Fried Convnets . In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 12 2015

  36. [44]

    ACDC: A Structured Efficient Linear Layer

    Marcin Moczulski, Misha Denil, Jeremy Appleyard, and Nando de Freitas. ACDC: A Structured Efficient Linear Layer . In International Conference on Learning Representations, 2016

  37. [45]

    Hammond, Pierre Vandergheynst, and R \' e mi Gribonval

    David K. Hammond, Pierre Vandergheynst, and R \' e mi Gribonval. Wavelets on graphs via spectral graph theory. Applied and Computational Harmonic Analysis, 30 0 (2): 0 129--150, 2011

  38. [46]

    Implicit Neural Representations with Periodic Activation Functions

    Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit Neural Representations with Periodic Activation Functions . In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing S...

  39. [47]

    Computation of Plain Unitary Rotations Transforming a General Matrix to Triangular Form

    Wallace Givens. Computation of Plain Unitary Rotations Transforming a General Matrix to Triangular Form . Journal of the Society for Industrial and Applied Mathematics, 6 0 (1): 0 26--50, 1958

  40. [48]

    Learning Latent Permutations with Gumbel-Sinkhorn Networks

    Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek. Learning Latent Permutations with Gumbel-Sinkhorn Networks . In International Conference on Learning Representations, 2018

  41. [49]

    Monarch: Expressive Structured Matrices for Efficient and Accurate Training

    Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher R \' e . Monarch: Expressive Structured Matrices for Efficient and Accurate Training . In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csa...

  42. [50]

    Sparse factorization of square matrices with application to neural attention modeling

    Ruslan Khalitov, Tong Yu, Lei Cheng, and Zhirong Yang. Sparse factorization of square matrices with application to neural attention modeling. Neural Networks, 152: 0 160--168, 2022

  43. [51]

    Fast Training of Convolutional Networks through FFTs

    Micha \"e l Mathieu, Mikael Henaff, and Yann LeCun. Fast Training of Convolutional Networks through FFTs . International Conference on Learning Representations, 2013

  44. [52]

    Spectral Graph Theory

    Fan Chung. Spectral Graph Theory . American Mathematical Society, 1997

  45. [53]

    Trefethen

    Lloyd N. Trefethen. Approximation Theory and Approximation Practice, Extended Edition . Society for Industrial and Applied Mathematics, 2019

  46. [54]

    Edwin Hewitt and Robert E. Hewitt. The Gibbs-Wilbraham phenomenon: An episode in fourier analysis . Archive for History of Exact Sciences, 21 0 (2): 0 129--160, 06 1979

  47. [55]

    Wong, and Lidia S

    Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao. Learning Deep Transformer Models for Machine Translation . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1810--1822. Association for ...

  48. [56]

    Nguyen and Julian Salazar

    Toan Q. Nguyen and Julian Salazar. Transformers without Tears: Improving the Normalization of Self-Attention . In Proceedings of the 16th International Conference on Spoken Language Translation. Association for Computational Linguistics, 11 2019

  49. [57]

    PyTorch: An Imperative Style, High-Performance Deep Learning Library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...

  50. [58]

    On the Relation between Position Information and Sentence Length in Neural Machine Translation

    Masato Neishi and Naoki Yoshinaga. On the Relation between Position Information and Sentence Length in Neural Machine Translation . In Mohit Bansal and Aline Villavicencio, editors, Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 32...

  51. [59]

    Decoupled Weight Decay Regularization

    Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization . In International Conference on Learning Representations, 2019

  52. [60]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The Long-Document Transformer . arXiv preprint arXiv: 2004.05150, 2020

  53. [61]

    Li, Madian Khabsa, Han Fang, and Hao Ma

    Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self-Attention with Linear Complexity . arXiv preprint arXiv: 2006.04768, 2020

  54. [62]

    Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran c ois Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, vol...

  55. [63]

    Sparse Sinkhorn Attention

    Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. Sparse Sinkhorn Attention . In Hal Daum \' e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 9438-...

  56. [64]

    o mformer: A Nystr \

    Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. Nystr \" o mformer: A Nystr \" o m-based Algorithm for Approximating Self-Attention . Proceedings of the AAAI Conference on Artificial Intelligence, 35 0 (16): 0 14138--14148...

  57. [65]

    Luna: Linear Unified Nested Attention

    Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer. Luna: Linear Unified Nested Attention . In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021

  58. [66]

    ListOps: A Diagnostic Dataset for Latent Tree Learning

    Nikita Nangia and Samuel Bowman. ListOps: A Diagnostic Dataset for Latent Tree Learning . In Silvio Ricardo Cordeiro, Shereen Oraby, Umashanthi Pavalanathan, and Kyeongmin Rim, editors, Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Co...

  59. [67]

    Maas, Raymond E

    Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning Word Vectors for Sentiment Analysis . In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea, editors, Proceedings of the 49th Annual Meeting of the Association for Computational...

  60. [68]

    Radev, Pradeep Muthukrishnan, and Vahed Qazvinian

    Dragomir R. Radev, Pradeep Muthukrishnan, and Vahed Qazvinian. The ACL Anthology Network Corpus . In Min-Yen Kan and Simone Teufel, editors, Proceedings of the 2009 Workshop on Text and Citation Analysis for Scholarly Digital Libraries ( NLPIR 4 DL ) , pages 54--61. Associatio...

  61. [69]

    Learning Multiple Layers of Features from Tiny Images

    Alex Krizhevsky. Learning Multiple Layers of Features from Tiny Images . Technical report, 2009

  62. [70]

    Learning long-range spatial dependencies with horizontal gated recurrent units

    Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre. Learning long-range spatial dependencies with horizontal gated recurrent units. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural I...

  63. [71]

    Disentangling neural mechanisms for perceptual grouping

    Junkyung Kim, Drew Linsley, Kalpit Thakkar, and Thomas Serre. Disentangling neural mechanisms for perceptual grouping. In International Conference on Learning Representations, 2020

  64. [72]

    Parallel and serial grouping of image elements in visual perception

    Roos Houtkamp and Pieter R Roelfsema. Parallel and serial grouping of image elements in visual perception. J. Exp. Psychol. Hum. Percept. Perform., 36 0 (6): 0 1443--1459, 12 2010

  65. [73]

    Long length document classification by local convolutional feature aggregation

    Liu Liu, Kaile Liu, Zhenghai Cong, Jiali Zhao, Yefei Ji, and Jun He. Long length document classification by local convolutional feature aggregation. Algorithms, 11 0 (8), 2018

  66. [74]

    Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. Convolutional Sequence to Sequence Learning . In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Le...

  67. [75]

    Transformer Language Models without Positional Encodings Still Learn Positional Information

    Adi Haviv, Ori Ram, Ofir Press, Peter Izsak, and Omer Levy. Transformer Language Models without Positional Encodings Still Learn Positional Information . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 1382--1390. Association for Computational L...

  68. [76]

    Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings

    Ta-Chung Chi, Ting-Han Fan, Li-Wei Chen, Alexander Rudnicky, and Peter Ramadge. Latent Positional Information is in the Self-Attention Variance of Transformer Language Models Without Positional Embeddings . In Proceedings of the 61st Annual Meeting of the Association for Compu...

  69. [77]

    The Impact of Positional Encoding on Length Generalization in Transformers

    Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. The Impact of Positional Encoding on Length Generalization in Transformers . In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Info...

  70. [78]

    Choose a Transformer: Fourier or Galerkin

    Shuhao Cao. Choose a Transformer: Fourier or Galerkin . In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021

  71. [79]

    Efficient Attention: Attention With Linear Complexities

    Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. Efficient Attention: Attention With Linear Complexities . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 3531--3539, 01 2021

  72. [80]

    Sparse Attention with Linear Units

    Biao Zhang, Ivan Titov, and Rico Sennrich. Sparse Attention with Linear Units . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6507--6520. Association for Computational Linguistics, 11 2021

  73. [81]

    SimA: Simple Softmax-Free Attention for Vision Transformers

    Soroush Abbasi Koohpayegani and Hamed Pirsiavash. SimA: Simple Softmax-Free Attention for Vision Transformers . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2607--2617, 01 2024

  74. [82]

    Skyformer: Remodel Self-Attention with Gaussian Kernel and Nystr \" o m Method

    Yifan Chen, Qi Zeng, Heng Ji, and Yun Yang. Skyformer: Remodel Self-Attention with Gaussian Kernel and Nystr \" o m Method . In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pa...

  75. [83]

    Scalable Parallel Programming with CUDA: Is CUDA the parallel programming model that application developers have been waiting for? Queue, 6 0 (2): 0 40--53, 03 2008

    John Nickolls, Ian Buck, Michael Garland, and Kevin Skadron. Scalable Parallel Programming with CUDA: Is CUDA the parallel programming model that application developers have been waiting for? Queue, 6 0 (2): 0 40--53, 03 2008

  76. [84]

    Untersuchungen \"u ber Fouriersche Reihen

    Leopold Fej \'e r. Untersuchungen \"u ber Fouriersche Reihen . Mathematische Annalen, 58: 0 51--69, 1904

  77. [85]

    Discourse on Fourier series

    Cornelius Lanczos. Discourse on Fourier series. University mathematical monographs. Oliver and Boyd, 1966

  78. [86]

    Veki \' c and S

    M. Veki \' c and S. R. White. Smooth boundary conditions for quantum lattice systems. Physical Review Letters, 71: 0 4283--4286, 12 1993

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.