Pith. sign in

REVIEW 3 cited by

Expressive power of recurrent neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.00811 v2 pith:TOC2PUAI submitted 2017-11-02 cs.LG

classification cs.LG
keywords networksexpressiveshallowdeepnetworkneuralpowerrecurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks are surprisingly efficient at solving practical tasks, but the theory behind this phenomenon is only starting to catch up with the practice. Numerous works show that depth is the key to this efficiency. A certain class of deep convolutional networks -- namely those that correspond to the Hierarchical Tucker (HT) tensor decomposition -- has been proven to have exponentially higher expressive power than shallow networks. I.e. a shallow network of exponential width is required to realize the same score function as computed by the deep architecture. In this paper, we prove the expressive power theorem (an exponential lower bound on the width of the equivalent shallow network) for a class of recurrent neural networks -- ones that correspond to the Tensor Train (TT) decomposition. This means that even processing an image patch by patch with an RNN can be exponentially more efficient than a (shallow) convolutional network with one hidden layer. Using theoretical results on the relation between the tensor decompositions we compare expressive powers of the HT- and TT-Networks. We also implement the recurrent TT-Networks and provide numerical evidence of their expressivity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The impact of allocation strategies in subset learning on the expressive power of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    In a teacher-student setup, maximal expressive power for a fixed learnable-weight budget is characterized by even row or column distribution in linear RNNs and feedforward networks.

  2. Reassessing Muon for Matrix Factorization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Muon's advantage over AdamW is problem-dependent: it loses or ties on plain low-rank factorization and completion but wins on nonnegative matrix factorization.

  3. A Scalable Factorization Approach for High-Order Structured Tensor Recovery

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Gradient descent on the Stiefel manifold recovers Tucker and tensor-train tensors with linear convergence whose initialization requirement and rate scale polynomially with the tensor order N.

Pith tools