Pith. sign in

REVIEW 1 cited by

Resurrecting Recurrent Neural Networks for Long Sequences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.06349 v1 pith:H4UI5KPA submitted 2023-03-11 cs.LG

classification cs.LG
keywords rnnsdeeplongperformancessmsfastrecurrentwhile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train. Deep state-space models (SSMs) have recently been shown to perform remarkably well on long sequence modeling tasks, and have the added benefits of fast parallelizable training and RNN-like fast inference. However, while SSMs are superficially similar to RNNs, there are important differences that make it unclear where their performance boost over RNNs comes from. In this paper, we show that careful design of deep RNNs using standard signal propagation arguments can recover the impressive performance of deep SSMs on long-range reasoning tasks, while also matching their training speed. To achieve this, we analyze and ablate a series of changes to standard RNNs including linearizing and diagonalizing the recurrence, using better parameterizations and initializations, and ensuring proper normalization of the forward pass. Our results provide new insights on the origins of the impressive performance of deep SSMs, while also introducing an RNN block called the Linear Recurrent Unit that matches both their performance on the Long Range Arena benchmark and their computational efficiency.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 43 citations worldwide. Full citation record

  1. Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    RTU-based exact RTRL enables streaming deep RL under partial observability, sustaining long credit assignment where one-step TBPTT collapses and matching batched PPO on several POPGym tasks.

Pith tools