Pith. sign in

REVIEW 2 cited by

Reservoir Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15045 v2 pith:KKUJ66P5 submitted 2020-12-30 cs.CL

classification cs.CL
keywords layersmachineperformancereservoirtransformerscomputeconvergencedemonstrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas in machine learning, we explore a variety of non-linear "reservoir" layers interspersed with regular transformer layers, and show improvements in wall-clock compute time until convergence, as well as overall performance, on various machine translation and (masked) language modelling tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantum Reservoir Computing: Recent Advances and Future Directions

    quant-ph 2026-07 accept novelty 4.0 of 10

    A comprehensive survey of quantum reservoir computing that proposes a common system model, a memory-architecture taxonomy, and resource-accounting standards, concluding that no broad quantum advantage is currently dem...

  2. Perturbative Gradient Training: A novel training paradigm for bridging the gap between deep neural networks and physical reservoir computing

    cs.LG 2025-06 reject novelty 4.0 of 10

    Random-perturbation, forward-pass-only training is demonstrated on simulated and physical reservoir networks, but it underperforms backpropagation on the transformer test and shows no verified pre-reservoir learning.

Pith tools