Pith. sign in

REVIEW 2 cited by

Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11467 v1 pith:MKWOVQCZ submitted 2025-02-17 cs.LG math.FA

classification cs.LGmath.FA
keywords transformerspolynomialsapproximationcapabilitynetworksymmetricabilityacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformers are a type of neural network that have demonstrated remarkable performance across various domains, particularly in natural language processing tasks. Motivated by this success, research on the theoretical understanding of transformers has garnered significant attention. A notable example is the mathematical analysis of their approximation power, which validates the empirical expressive capability of transformers. In this study, we investigate the ability of transformers to approximate column-symmetric polynomials, an extension of symmetric polynomials that take matrices as input. Consequently, we establish an explicit relationship between the size of the transformer network and its approximation capability, leveraging the parameter efficiency of transformers and their compatibility with symmetry by focusing on the algebraic properties of symmetric polynomials.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Dimension-Free Approximation of Deep Neural Networks for Symmetric Korobov Functions

    cs.LG 2025-11 conditional novelty 7.0 of 10

    Symmetric squared-ReLU networks approximate symmetric Korobov functions at rate O(m^{-1}) with a dimension-independent prefactor, and gradient-based learning achieves M^{-2/3} generalization error.

  2. Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

    stat.ML 2026-02 conditional novelty 5.0 of 10

    Standard Transformers attain the minimax optimal rate m^{-2γ/(2γ+dn)} (up to logs) for nonparametric regression of Hölder C^{s,λ} targets on [0,1]^{d×n}.

Pith tools