REVIEW 2 cited by
Reservoir Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas in machine learning, we explore a variety of non-linear "reservoir" layers interspersed with regular transformer layers, and show improvements in wall-clock compute time until convergence, as well as overall performance, on various machine translation and (masked) language modelling tasks.
Forward citations
Cited by 2 Pith papers
-
Quantum Reservoir Computing: Recent Advances and Future Directions
A comprehensive survey of quantum reservoir computing that proposes a common system model, a memory-architecture taxonomy, and resource-accounting standards, concluding that no broad quantum advantage is currently dem...
-
Perturbative Gradient Training: A novel training paradigm for bridging the gap between deep neural networks and physical reservoir computing
Random-perturbation, forward-pass-only training is demonstrated on simulated and physical reservoir networks, but it underperforms backpropagation on the transformer test and shows no verified pre-reservoir learning.
Discussion (0). Sign in to comment.