Pith. sign in

Transformer feed-forward layers are key-value memories

11 Pith papers cite this work. Polarity classification is still indexing.

11 Pith papers citing it

citation-role summary

method 1

citation-polarity summary

years

2026 11

verdicts

UNVERDICTED 11

roles

method 1

polarities

use method 1

representative citing papers

Graph Memory Transformer (GMT)

cs.LG · 2026-04-26 · unverdicted · novelty 7.0

Graph Memory Transformer replaces FFN sublayers with a graph memory cell using 128 centroids and transition matrices per block, yielding stable training at 82.2M parameters but higher validation loss than a 103M dense baseline.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

cs.LG · 2026-03-27 · unverdicted · novelty 7.0

Muon achieves higher storage capacity than SGD and matches Newton's method in one-step recovery rates for associative memory under power-law distributions, while saturating at larger critical batch sizes and showing faster initial multi-step dynamics.

Rosetta: Composable Native Multimodal Pretraining

cs.CV · 2026-07-01 · unverdicted · novelty 5.0

Rosetta proposes a composable multimodal pretraining method with MAOP to prevent catastrophic forgetting when expanding modalities beyond standard MoE and MoT approaches.

Cubit: Token Mixer with Kernel Ridge Regression

cs.LG · 2026-05-07 · unverdicted · novelty 5.0 · 2 refs

Cubit replaces Transformer's attention with a closed-form Kernel Ridge Regression token mixer and reports larger gains as training sequence length increases.

citing papers explorer

Showing 11 of 11 citing papers.