Pith. sign in

Avoiding pathologies in very deep networks

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Choosing appropriate architectures and regularization strategies for deep networks is crucial to good predictive performance. To shed light on this problem, we analyze the analogous problem of constructing useful priors on compositions of functions. Specifically, we study the deep Gaussian process, a type of infinitely-wide, deep neural network. We show that in standard architectures, the representational capacity of the network tends to capture fewer degrees of freedom as the number of layers increases, retaining only a single degree of freedom in the limit. We propose an alternate network architecture which does not suffer from this pathology. We also examine deep covariance functions, obtained by composing infinitely many feature transforms. Lastly, we characterize the class of models obtained by performing dropout on Gaussian processes.

citation-role summary

background 1

citation-polarity summary

fields

cs.NE 1

years

2019 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Additive function approximation in the brain

cs.NE · 2019-09-05 · conditional · novelty 5.0

Sparse random feature networks with in-degree d are equivalent to order-d additive models, and a distribution of in-degrees yields a mixture of additive kernels.

citing papers explorer

Showing 1 of 1 citing paper.

  • Additive function approximation in the brain cs.NE · 2019-09-05 · conditional · none · ref 47 · internal anchor

    Sparse random feature networks with in-degree d are equivalent to order-d additive models, and a distribution of in-degrees yields a mixture of additive kernels.