Pith. sign in

REVIEW 3 cited by

Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19798 v2 pith:FT7I6FOL submitted 2023-05-31 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords self-attentionksvdasymmetrickerneloptimizationprimal-attentionrepresentationattention
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Recently, a new line of works has emerged to understand and improve self-attention in Transformers by treating it as a kernel machine. However, existing works apply the methods for symmetric kernels to the asymmetric self-attention, resulting in a nontrivial gap between the analytical understanding and numerical implementation. In this paper, we provide a new perspective to represent and optimize self-attention through asymmetric Kernel Singular Value Decomposition (KSVD), which is also motivated by the low-rank property of self-attention normally observed in deep layers. Through asymmetric KSVD, $i$) a primal-dual representation of self-attention is formulated, where the optimization objective is cast to maximize the projection variances in the attention outputs; $ii$) a novel attention mechanism, i.e., Primal-Attention, is proposed via the primal representation of KSVD, avoiding explicit computation of the kernel matrix in the dual; $iii$) with KKT conditions, we prove that the stationary solution to the KSVD optimization in Primal-Attention yields a zero-value objective. In this manner, KSVD optimization can be implemented by simply minimizing a regularization loss, so that low-rank property is promoted without extra decomposition. Numerical experiments show state-of-the-art performance of our Primal-Attention with improved efficiency. Moreover, we demonstrate that the deployed KSVD optimization regularizes Primal-Attention with a sharper singular value decay than that of the canonical self-attention, further verifying the great potential of our method. To the best of our knowledge, this is the first work that provides a primal-dual representation for the asymmetric kernel in self-attention and successfully applies it to modeling and optimization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain

    cs.LG 2025-05 conditional novelty 5.0 of 10

    AGF replaces softmax attention with a learned singular-value-domain graph filter that runs in O(n d^2), and reports moderate accuracy improvements on UEA and LRA benchmarks.

  2. A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper reports that learned value matrices do not align with KPCA quantities, the projection-loss decrease is dominated by output-norm collapse, and the reported Gram eigenvalue statistics cannot be reproduced with...

  3. Mirror Descent on Reproducing Kernel Banach Spaces

    cs.LG 2024-11 conditional novelty 5.0 of 10

    The authors design a functional mirror descent for reproducing kernel Banach spaces and prove conditional linear and O(1/√t) convergence, with a finite-center p-norm RKBS instantiation.

Pith tools