The paper claims that linear self-attention defines a parametric endofunctor whose layered stacking is the free monad, but the construction is mostly restatement and has serious technical flaws.
Metric Space Magnitude and Generalisation in Neural Networks
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Deep learning models have seen significant successes in numerous applications, but their inner workings remain elusive. The purpose of this work is to quantify the learning process of deep neural networks through the lens of a novel topological invariant called magnitude. Magnitude is an isometry invariant; its properties are an active area of research as it encodes many known invariants of a metric space. We use magnitude to study the internal representations of neural networks and propose a new method for determining their generalisation capabilities. Moreover, we theoretically connect magnitude dimension and the generalisation error, and demonstrate experimentally that the proposed framework can be a good indicator of the latter.
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Self-Attention as a Parametric Endofunctor: A Categorical Framework for Transformer Architectures
The paper claims that linear self-attention defines a parametric endofunctor whose layered stacking is the free monad, but the construction is mostly restatement and has serious technical flaws.