The exact finite-depth output distribution of a width-k depth-n Boolean machine depends only on Hamming weight, biases toward low and high weights up to depth ~2^k, then collapses to the constants true and false.
Deep-layered machines have a built-in occam’s razor.arXiv preprint arXiv:2603.01217, 2026
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Trained ResNets on CIFAR-10 retain measurable dependence on initialization scale under low-LR SGD (26.5 pp test accuracy spread) but not under Adam, indicating that practical inductive bias is shaped by the forgetting time scale of the optimizer and regularizers.
citing papers explorer
-
Exact distribution of the output of a deep-layered machine
The exact finite-depth output distribution of a width-k depth-n Boolean machine depends only on Hamming weight, biases toward low and high weights up to depth ~2^k, then collapses to the constants true and false.
-
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
Trained ResNets on CIFAR-10 retain measurable dependence on initialization scale under low-LR SGD (26.5 pp test accuracy spread) but not under Adam, indicating that practical inductive bias is shaped by the forgetting time scale of the optimizer and regularizers.