Pith. sign in

Neural Redshift: Random Networks are not Random Functions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Our understanding of the generalization capabilities of neural networks (NNs) is still incomplete. Prevailing explanations are based on implicit biases of gradient descent (GD) but they cannot account for the capabilities of models from gradient-free methods nor the simplicity bias recently observed in untrained networks. This paper seeks other sources of generalization in NNs. Findings. To understand the inductive biases provided by architectures independently from GD, we examine untrained, random-weight networks. Even simple MLPs show strong inductive biases: uniform sampling in weight space yields a very biased distribution of functions in terms of complexity. But unlike common wisdom, NNs do not have an inherent "simplicity bias". This property depends on components such as ReLUs, residual connections, and layer normalizations. Alternative architectures can be built with a bias for any level of complexity. Transformers also inherit all these properties from their building blocks. Implications. We provide a fresh explanation for the success of deep learning independent from gradient-based training. It points at promising avenues for controlling the solutions implemented by trained models.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

When Do Neural Networks Learn World Models?

cs.LG · 2025-02-13 · conditional · novelty 7.0

With Boolean variables, a low-degree bias, and a task distribution weighted toward simple functions of the latents, multi-task training provably recovers the latent world model up to permutations and negations.

citing papers explorer

Showing 1 of 1 citing paper.

  • When Do Neural Networks Learn World Models? cs.LG · 2025-02-13 · conditional · none · ref 74 · internal anchor

    With Boolean variables, a low-degree bias, and a task distribution weighted toward simple functions of the latents, multi-task training provably recovers the latent world model up to permutations and negations.