Highway and Residual Networks learn Unrolled Iterative Estimation

Klaus Greff , Rupesh K. Srivastava , J\"urgen Schmidhuber

Authors on Pith no claims yet

classification 💻 cs.NE cs.AIcs.LG

keywords networksarchitectureshighwayresidualestimationfeaturesiterativelayers

read the original abstract

The past year saw the introduction of new architectures such as Highway networks and Residual networks which, for the first time, enabled the training of feedforward networks with dozens to hundreds of layers using simple gradient descent. While depth of representation has been posited as a primary reason for their success, there are indications that these architectures defy a popular view of deep learning as a hierarchical computation of increasingly abstract features at each layer. In this report, we argue that this view is incomplete and does not adequately explain several recent findings. We propose an alternative viewpoint based on unrolled iterative estimation -- a group of successive layers iteratively refine their estimates of the same features instead of computing an entirely new representation. We demonstrate that this viewpoint directly leads to the construction of Highway and Residual networks. Finally we provide preliminary experiments to discuss the similarities and differences between the two architectures.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Eliciting Latent Predictions from Transformers with the Tuned Lens
cs.LG 2023-03 accept novelty 7.0

Training per-layer affine probes on frozen transformers yields more reliable latent predictions than the logit lens and enables detection of malicious inputs from prediction trajectories.
Attention U-Net: Learning Where to Look for the Pancreas
cs.CV 2018-04 unverdicted novelty 6.0

Attention gates added to U-Net automatically focus on target organs in CT images and improve segmentation performance on abdominal datasets.