REVIEW 5 cited by
Trellis Networks for Sequence Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present trellis networks, a new architecture for sequence modeling. On the one hand, a trellis network is a temporal convolutional network with special structure, characterized by weight tying across depth and direct injection of the input into deep layers. On the other hand, we show that truncated recurrent networks are equivalent to trellis networks with special sparsity structure in their weight matrices. Thus trellis networks with general weight matrices generalize truncated recurrent networks. We leverage these connections to design high-performing trellis networks that absorb structural and algorithmic elements from both recurrent and convolutional models. Experiments demonstrate that trellis networks outperform the current state of the art methods on a variety of challenging benchmarks, including word-level language modeling and character-level language modeling tasks, and stress tests designed to evaluate long-term memory retention. The code is available at https://github.com/locuslab/trellisnet .
Forward citations
Cited by 5 Pith papers
-
DEQuify your force field: More efficient simulations using deep equilibrium models
Recasting EquiformerV2 as a deep equilibrium model with warm-started fixed points gives faster and often more accurate force predictions on standard MD benchmarks, though energy accuracy on OC20 is worse.
-
Compressive Transformers for Long-Range Sequence Modelling
Compressive Transformer sets new records on WikiText-103 (17.1 ppl) and Enwik8 (0.97 bpc) via memory compression and introduces the PG-19 long-range language benchmark.
-
Mogrifier LSTM
Mogrifier LSTM, which applies repeated mutual gating between the input and previous hidden state, outperforms the LSTM on PTB, Wikitext-2, Enwik8, and MWC language modeling.
-
From Local Patterns to Global Understanding: Cross-Stock Trend Integration for Enhanced Predictive Modeling
CSTI, a federated-learning-style scheme that aggregates per-stock models and then fine-tunes on each stock, improves stock prediction in many but not all tested settings.
-
The Good, The Efficient and the Inductive Biases: Exploring Efficiency in Deep Learning Through the Use of Inductive Biases
A dissertation synthesizing the author's papers on continuous kernel convolutions and symmetry-preserving architectures, claiming these inductive biases improve deep learning efficiency.
Discussion (0). Continue with ORCID to comment.