REVIEW 3 cited by
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present our position on the elusive quest for a general-purpose framework for specifying and studying deep learning architectures. Our opinion is that the key attempts made so far lack a coherent bridge between specifying constraints which models must satisfy and specifying their implementations. Focusing on building a such a bridge, we propose to apply category theory -- precisely, the universal algebra of monads valued in a 2-category of parametric maps -- as a single theory elegantly subsuming both of these flavours of neural network design. To defend our position, we show how this theory recovers constraints induced by geometric deep learning, as well as implementations of many architectures drawn from the diverse landscape of neural networks, such as RNNs. We also illustrate how the theory naturally encodes many standard constructs in computer science and automata theory.
Forward citations
Cited by 3 Pith papers
-
The Program Hypergraph: Multi-Way Relational Structure for Geometric Algebra, Spatial Compute, and Physics-Aware Compilation
The Program Hypergraph extends binary program semantic graphs to arbitrary-arity hyperedges to faithfully represent multi-way relations in geometric algebra and spatial architectures.
-
Topos Theory for Generative AI and LLMs
The paper claims the category of LLM functions is a topos and uses that to propose new compositional architectures like pullbacks, pushouts, and subobject classifiers, but gives no implementation or complete proofs.
-
Relational inductive biases on attention mechanisms
Attention mechanisms are classified by their relational inductive bias: self-attention assumes a complete graph, masked attention a total order, strided attention p-previous connections, encoder-decoder a bipartite gr...
Discussion (0). Continue with ORCID to comment.