REVIEW 3 major objections 5 minor 51 references
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read The simplest Monte Carlo rule—mutate one parameter, keep it only if loss falls—can train deep nets, wide nets, and a small Transformer without any gradients.
desk verdict A clean feasibility demo: one-parameter Monte Carlo mutation trains deep FC nets, extreme pruning, and a small Transformer without gradients—useful, not yet a practical BP replacement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Monte Carlo mutation–optimization selection (MCA): randomly select one or more parameters, apply a finite perturbation (or set the weight to zero for pure pruning), and accept the change only when the empirical loss does not increase and the hard amplitude constraint is respected. The same rule is applied layer-wise with local-field caching so that only the affected forward slice is recomputed.
What would settle it
Train the same 21-layer fully connected net or the same three-layer Transformer under identical data and compute budgets with a modern optimizer; if MCA cannot reach comparable test accuracy or if its wall-clock time grows prohibitively with depth or sequence length, the practicality claim fails.
Extended reading notes
Core claim
A single-parameter Monte Carlo mutation–selection rule (randomly perturb a weight or bias, accept if loss does not increase, optionally subject to a hard magnitude bound) is sufficient to train deep fully connected networks, extremely wide single-hidden-layer networks, and a simple Transformer architecture to nontrivial accuracy on MNIST classification and Tiny Shakespeare language modeling, without backpropagation, residual connections, or (in many experiments) layer normalization.
Load-bearing premise
That full-batch or large-batch accuracy numbers on MNIST and Tiny Shakespeare, obtained with one fixed set of step-size and bound parameters, already show the method is a practical general alternative for deep networks.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that a minimal Monte Carlo mutation–selection rule (MCA)—randomly perturb one or more parameters and accept the change only if the empirical loss does not increase—can train deep neural networks without backpropagation. On a single GPU, the method is shown to train fully connected nets deeper than 20 layers, single-hidden-layer nets up to width 16 384, and a simple multi-layer Transformer on MNIST classification and Tiny Shakespeare character-level language modeling. Additional experiments demonstrate pure pruning training, discrete-weight support, Gaussian activations, and extreme sparsification. A theoretical section argues that MCA recovers gradient descent in the small-perturbation limit while remaining applicable to non-smooth, discrete, and hard-constrained settings where BP fails.
Significance. If the feasibility claim holds under broader conditions, the work supplies a concrete, implementable gradient-free alternative that sidesteps vanishing/exploding gradients and directly supports discrete weights, hard amplitude constraints, and pure pruning. The pure-pruning and extreme-sparsity results (remaining weights ~0.33 % while retaining >95 % test accuracy) are of independent interest for model compression and for studying redundancy. The Transformer language-modeling demonstration, though modest in scale, is a nontrivial proof-of-principle that the same rule can train attention-based sequence models. The theoretical comparison (local-limit recovery of gradient descent, escape from saddles via finite mutations, hard-constraint generalization bounds) is carefully framed and does not overclaim replacement of BP. These elements make the manuscript a useful contribution to gradient-free and physics-inspired training methods, provided the empirical scope is clarified.
major comments (3)
- §2.1 and the abstract claim that MCA “can practically train deep networks.” All deep fully-connected experiments use full-batch (60 k) or large-batch (30 k) training with a single hyper-parameter triple (δ=0.01, Δ=0.1, He-style init) transferred from a 5-layer net. No mini-batch noise, multi-seed statistics, or scaling curves on harder data are reported. Because acceptance of beneficial mutations is evaluated on the empirical loss of the current batch, the acceptance rate may collapse under realistic mini-batch variance; without that evidence the “practical alternative” claim remains under-supported for general deep-network practice.
- §2.5 Transformer language-modeling result: after 48 h the 3-layer character-level model reaches test accuracy 0.47 / perplexity 5.86. The generated sample (Table 1) shows recognizable structure but also pseudo-words and local errors. The manuscript presents this as a “proof-of-principle,” yet the abstract and conclusion still list it among demonstrations that MCA trains “even a simple Transformer.” A clearer quantitative baseline (same architecture trained by BP under identical compute) and an explicit statement of the performance gap are needed so that the claim is not overstated.
- §2.1 BP comparison is limited to a single full-batch Adam run on a 5-layer ReLU net (~98.19 % test). No modern mini-batch BP baselines, residual architectures, or learning-rate schedules appear for the deeper nets or for the Gaussian/Transformer cases. While the paper correctly states that extreme accuracy competition is not its focus, the absence of any competitive BP reference for depths >5 and for the Transformer leaves the relative practicality claim difficult to assess.
minor comments (5)
- §2.1: “10 000 Monte Carlo operations = 1 epoch” is non-standard; a short justification or conversion to wall-clock / FLOPs would help readers compare with BP epochs.
- Figure 2(c) heat-maps are informative but lack a color-bar scale and an explicit statement of how “completely disconnected sites” are thresholded.
- §3.2 Eq. (9): the proportionality constant for E[ΔW_acc] is left unspecified; a one-line derivation or citation would tighten the local-limit argument.
- References [5] and [26] list 2026 publication years; verify that these are not future-dated placeholders.
- CUDA Fortran implementation details (caching strategy, single vs double precision) are mentioned only briefly; a short appendix or code-availability statement would improve reproducibility.
Circularity Check
No load-bearing circularity: MCA is defined by an explicit accept/reject rule on measured loss; deep-net feasibility is shown by new full-batch experiments on MNIST/Tiny Shakespeare, not by definitional identity or fitted-input predictions.
-
self citation load bearing
[§1 Introduction, paragraphs on Boltzmann machines and MCA history]
"Subsequently, the method was extended to the training of general neural networks and shallow fully connected networks [35, 36, 37, 38]: by designing an appropriate algorithm that calculates only the changed parts caused by mutations rather than evolving the entire network, training of shallow fully connected networks became feasible on a CPU. ... However, limited by the computational efficiency of CPUs, the above methods are mainly confined to shallow or small-scale network training"
The citations are to the same author’s prior MCA papers. They establish only that the mutation–selection idea already existed for shallow nets; they do not force or define the new deep-net / Transformer results reported in §2. The load is historical, not definitional, so the circularity is minor and non-central.
full rationale
The paper’s central claims are empirical demonstrations that a simple mutation–selection rule (mutate a parameter or channel, accept if loss does not increase) can train deep fully-connected nets, wide single-hidden-layer nets, and a small Transformer on MNIST and Tiny Shakespeare. The algorithm is stated self-containedly in §2 (Eq. 1 and surrounding text) and evaluated on held-out accuracy and generated text; success is not forced by construction from any fitted constant. Prior self-citations ([32–41], especially [35–38]) supply only the historical precursor that MCA worked for shallow nets on CPU; they are not invoked as uniqueness theorems, ansatzes, or load-bearing premises that make the deep-net results true by definition. Section 3’s theoretical comparisons (local-limit reduction to gradient descent, second-order escape from saddles, Lipschitz bounds under hard constraints) are standard expansions that do not reduce any claimed prediction to an input by identity. No fitted-input-called-prediction, self-definitional loop, or renaming of a known result appears. The single minor self-citation chain is background, not circular, yielding score 1.
Assumptions & free parameters
free parameters (5)
- mutation amplitude δ =
0.01
- hard weight bound Δ =
0.1
- Gaussian transfer coefficient γ =
0.4 / 0.04 / 1.0
- initialization scale =
He / √(0.1/n)
- epoch definition (MC trials per epoch) =
10000
assumptions (3)
- domain assumption A mutation that does not increase empirical cross-entropy on the training set is a useful update for generalization.
- domain assumption Full-batch (or large-batch) evaluation of loss is a valid and sufficient training signal.
- ad hoc to paper Layer-normalization or its absence does not qualitatively change the feasibility claim.
Cite this review
Pith. "Pith review of Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks." pith.science (2026). https://pith.science/paper/CU2S3HGU
@misc{pith2026260708406,
author = {Pith},
title = {Pith review of: Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/CU2S3HGU}},
note = {Machine review of arXiv:2607.08406}
}
read the original abstract
Backpropagation (BP) dominates deep learning training, but its reliance on gradients brings inherent troubles -- vanishing and exploding gradients. The pursuit of gradient-free methods has long been a goal in the field of artificial intelligence. This paper shows that indeed the simplest Monte Carlo algorithm implemented on a single GPU -- randomly mutate a parameter, keep it if the loss decreases, otherwise retry -- can practically train deep networks. This gradient-free method does not even need common techniques such as batch normalization or residual connections to directly train sufficiently deep networks. More remarkably, its flexibility extends to several nontrivial scenarios: it enables pure pruning training, supports discrete weights, accommodates unconventional transfer functions such as Gaussian, and reveals the substantial redundancy of deep networks. We have demonstrated its feasibility on deep networks with more than 20 layers, single-hidden-layer wide networks with up to 16,384 hidden neurons, and even a simple Transformer architecture trained on both image classification (MNIST) and character-level language modeling (Tiny Shakespeare). This simple gradient-free method may offer a complementary perspective for understanding the self-organization and learning mechanisms of neural networks, and also provides an alternative route for building physically inspired deep learning systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. Learning represen- tations by back-propagating errors.Nature,323, 533–536 (1986)
work page 1986
-
[2]
Deep learning.Nature,521, 436– 444 (2015)
LeCun, Y., Bengio, Y., & Hinton, G. Deep learning.Nature,521, 436– 444 (2015)
work page 2015
-
[3]
Ho, T. K. Random decision forests. InProceedings of the Third Inter- national Conference on Document Analysis and Recognition, Vol. 1, pp. 278–282 (1995)
work page 1995
-
[4]
Random forests.Machine Learning,45, 5–32 (2001)
Breiman, L. Random forests.Machine Learning,45, 5–32 (2001)
work page 2001
-
[5]
Theory of random forests.Annual Review of Statistics and Its Application,13, 99–121 (2026)
Scornet, E., & Hooker, G. Theory of random forests.Annual Review of Statistics and Its Application,13, 99–121 (2026)
work page 2026
-
[6]
J.Learning with Kernels: Support Vec- tor Machines, Regularization, Optimization, and Beyond
Sch¨ olkopf, B., & Smola, A. J.Learning with Kernels: Support Vec- tor Machines, Regularization, Optimization, and Beyond. MIT Press (2002). 18
work page 2002
-
[7]
Du, K.-L., Jiang, B., Lu, J., Hua, J., & Swamy, M. N. S. Exploring kernel machines and support vector machines: Principles, techniques, and future directions.Mathematics,12(24), 3935 (2024)
work page 2024
-
[8]
Novikov, A., Podoprikhin, D., Osokin, A., & Vetrov, D. Tensorizing neural networks. InAdvances in Neural Information Processing Systems 28442–450 (2015)
work page 2015
Show all 51 references
-
[9]
Unsupervised Generative Modeling Using Matrix Product States.Physical Review X, 8, 031012 (2018)
Han, Z.-Y., Wang, J., Fan, H., Wang, L., & Zhang, P. Unsupervised Generative Modeling Using Matrix Product States.Physical Review X, 8, 031012 (2018)
2018
-
[10]
Simulation of Quantum Circuits Using the Big- Batch Tensor Network Method.Physical Review Letters,128, 030501 (2022)
Pan, F., & Zhang, P. Simulation of Quantum Circuits Using the Big- Batch Tensor Network Method.Physical Review Letters,128, 030501 (2022)
2022
-
[11]
M., & Schwab, D
Stoudenmire, E. M., & Schwab, D. J. Supervised learning with tensor networks. InAdvances in Neural Information Processing Systems 29 (2016)
2016
-
[12]
Machine learning with tree tensor networks, CP rank constraints, and tensor dropout.IEEE Transactions on Pattern Analysis and Machine Intelligence,46(12), 7825–7832 (2024)
Chen, H., & Barthel, T. Machine learning with tree tensor networks, CP rank constraints, and tensor dropout.IEEE Transactions on Pattern Analysis and Machine Intelligence,46(12), 7825–7832 (2024)
2024
-
[13]
O., Mostafa, H., & Zenke, F
Neftci, E. O., Mostafa, H., & Zenke, F. Surrogate gradient learning in spiking neural networks: Bringing the power of gradient-based opti- mization to spiking neural networks.IEEE Signal Processing Magazine, 36(6), 51–63 (2019)
2019
-
[14]
A solution to the learning dilemma for recurrent networks of spiking neurons.Nature Communications,11, 3625 (2020)
Bellec, G., Scherr, F., Subramoney, A., Hajek, E., Salaj, D., Legenstein, R., & Maass, W. A solution to the learning dilemma for recurrent networks of spiking neurons.Nature Communications,11, 3625 (2020)
2020
-
[15]
High-performance deep spiking neural networks with 0.3 spikes per neuron.Nature Communi- cations,15, 6793 (2024)
Stanojevic, A., Wo´ zniak, S., Bellec, G., et al. High-performance deep spiking neural networks with 0.3 spikes per neuron.Nature Communi- cations,15, 6793 (2024)
2024
-
[16]
The Forward-Forward algorithm: Some preliminary inves- tigations
Hinton, G. The Forward-Forward algorithm: Some preliminary inves- tigations. arXiv:2212.13345 (2022)
2022 arXiv
-
[17]
W., Rosenbluth, M
Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., & Teller, E. Equation of state calculations by fast computing machines. The Journal of Chemical Physics,21(6), 1087–1092 (1953)
1953
-
[18]
P., & Binder, K.A Guide to Monte Carlo Simulations in Statistical Physics, 5th ed
Landau, D. P., & Binder, K.A Guide to Monte Carlo Simulations in Statistical Physics, 5th ed. Cambridge University Press (2021)
2021
-
[19]
Monte Carlo study of quantized SU(2) gauge theory.Phys- ical Review D,21(8), 2308–2315 (1980)
Creutz, M. Monte Carlo study of quantized SU(2) gauge theory.Phys- ical Review D,21(8), 2308–2315 (1980). 19
1980
-
[20]
Geant4—a simulation toolkit.Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment,506(3), 250–303 (2003)
Agostinelli, S., Allison, J., Amako, K., Apostolakis, J., Araujo, H., et al. Geant4—a simulation toolkit.Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment,506(3), 250–303 (2003)
2003
-
[21]
Performance portable Monte Carlo particle transport on Intel, NVIDIA, and AMD GPUs.EPJ Web of Conferences,302, 04010 (2024)
Tramm, J., Romano, P., Shriwise, P., Lund, A., Doerfert, J., Stein- brecher, P., Siegel, A., & Ridley, G. Performance portable Monte Carlo particle transport on Intel, NVIDIA, and AMD GPUs.EPJ Web of Conferences,302, 04010 (2024)
2024
-
[22]
M.Bayesian Learning for Neural Networks
Neal, R. M.Bayesian Learning for Neural Networks. Springer (1996)
1996
-
[23]
Andrieu, C., de Freitas, N., Doucet, A., & Jordan, M. I. An introduction to MCMC for machine learning.Machine Learning,50, 5–43 (2003)
2003
-
[24]
Bayesian neural networks via MCMC: A Python-based tutorial.IEEE Access,12, 70519–70549 (2024)
Chandra, R., & Simmons, J. Bayesian neural networks via MCMC: A Python-based tutorial.IEEE Access,12, 70519–70549 (2024)
2024
-
[25]
Dropout as a Bayesian approximation: Rep- resenting model uncertainty in deep learning
Gal, Y., & Ghahramani, Z. Dropout as a Bayesian approximation: Rep- resenting model uncertainty in deep learning. InProceedings of the 33rd International Conference on Machine Learning, PMLR48, 1050–1059 (2016)
2016
-
[26]
A survey on uncertainty quantification methods for deep learning.ACM Computing Surveys, 58(7), 1–35 (2026)
He, W., Jiang, Z., Xiao, T., Xu, Z., & Li, Y. A survey on uncertainty quantification methods for deep learning.ACM Computing Surveys, 58(7), 1–35 (2026)
2026
-
[27]
P., Kumar, A., Ermon, S., & Poole, B
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. Score-based generative modeling through stochastic differen- tial equations. InInternational Conference on Learning Representations (2021)
2021
-
[28]
J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of Go with deep neural networks and tree search.Nature,529, 484–489 (2016)
2016
-
[29]
H., Hinton, G
Ackley, D. H., Hinton, G. E., & Sejnowski, T. J. A learning algorithm for Boltzmann machines.Cognitive Science,9(1), 147–169 (1985)
1985
-
[30]
E., & Sejnowski, T
Hinton, G. E., & Sejnowski, T. J. Learning and relearning in Boltz- mann machines. In D. E. Rumelhart, J. L. McClelland, & the PDP Research Group,Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1: Foundations, pp. 282–317. MIT Press (1986)
1986
-
[31]
Hinton, G. E. Boltzmann machine.Scholarpedia,2(5), 1668 (2007). 20
2007
-
[32]
Designing asymmetric neural networks with associative mem- ory.Physical Review E,70, 066137 (2004)
Zhao, H. Designing asymmetric neural networks with associative mem- ory.Physical Review E,70, 066137 (2004)
2004
-
[33]
Pattern recognition using asymmetric attractor neural networks.Physical Review E,72, 066111 (2005)
Jin, T., & Zhao, H. Pattern recognition using asymmetric attractor neural networks.Physical Review E,72, 066111 (2005)
2005
-
[34]
Storing limit cycles using delayed feedback neural networks.Advances in Complex Systems,11(03), 433–442 (2008)
Wu, G., & Zhao, H. Storing limit cycles using delayed feedback neural networks.Advances in Complex Systems,11(03), 433–442 (2008)
2008
-
[35]
A global algorithm for training multilayer neural networks
Zhao, H., & Jin, T. A global algorithm for training multilayer neural networks. arXiv:physics/0607046 (2006)
2006 arXiv
-
[36]
Time Series Forecasting Using Multilayer Neural Network Constructed by a Monte-Carlo Based Algorithm
Xin, P., & Zhao, H. Time Series Forecasting Using Multilayer Neural Network Constructed by a Monte-Carlo Based Algorithm. InProceed- ings of the 2009 1st IEEE Symposium on Web Society (SWS 2009), 264–267 (2009)
2009
-
[37]
A general theory for training learning machine
Zhao, H. A general theory for training learning machine. arXiv:1704.06885 (2017)
2017 arXiv
-
[38]
black-box
Zhao, H. Inferring the dynamics of “black-box” systems using a learning machine.Science China Physics, Mechanics & Astronomy,64, 270511 (2021)
2021
-
[39]
Feed-back neural networks with discrete weights.Neural Computing and Applications,22(6), 1063–1069 (2013)
Wang, L., Zhou, Q., Jin, T., & Zhao, H. Feed-back neural networks with discrete weights.Neural Computing and Applications,22(6), 1063–1069 (2013)
2013
-
[40]
Correlation between eigenvalue spectra and dynamics of neural networks.Neural Computation,21(10), 2931– 2941 (2009)
Zhou, Q., Jin, T., & Zhao, H. Correlation between eigenvalue spectra and dynamics of neural networks.Neural Computation,21(10), 2931– 2941 (2009)
2009
-
[41]
An eval- uation of the dynamics of diluted neural network.International Journal of Computational Intelligence Systems,9(6), 1191–1199 (2016)
Wang, L., Shen, J., Zhou, Q., Shang, Z., Chen, H., & Zhao, H. An eval- uation of the dynamics of diluted neural network.International Journal of Computational Intelligence Systems,9(6), 1191–1199 (2016)
2016
-
[42]
P., & Zhou, H
Huang, H. P., & Zhou, H. J. Combined local search strategy for learning in networks of binary synapses.EPL (Europhysics Letters),96, 58003 (2011)
2011
-
[43]
Han, S., Pool, J., Tran, J., & Dally, W. J. Learning both weights and connections for efficient neural networks. InAdvances in Neural Infor- mation Processing Systems 28 (NeurIPS 2015)
2015
-
[44]
B., Fourati, R., et al
Tmamna, J., Ayed, E. B., Fourati, R., et al. Pruning deep neural net- works for green energy-efficient models: A survey.Cognitive Computa- tion,16, 2931–2952 (2024). 21
2024
-
[45]
Pierro, A., Abreu, S., Timcheck, J., Stratmann, P., Wild, A., & Shrestha, S. B. Accelerating linear recurrent neural networks for the edge with unstructured sparsity. InProceedings of the 42nd Inter- national Conference on Machine Learning, PMLR267, 49382–49398 (2025)
2025
-
[46]
Energy-information trade-off induces continuous and discontinuous phase transitions in lat- eral predictive coding.Science China Physics, Mechanics & Astronomy, 67, 260511 (2024)
Huang, Z.-Y., Zhou, R., Huang, M., & Zhou, H.-J. Energy-information trade-off induces continuous and discontinuous phase transitions in lat- eral predictive coding.Science China Physics, Mechanics & Astronomy, 67, 260511 (2024)
2024
-
[47]
E., Eltawil, A
Zhang, L., Krestinskaya, O., Fouda, M. E., Eltawil, A. M., & Salama, K. N. Quantized convolutional neural networks: a hardware perspective. Frontiers in Electronics,6, 1469802 (2025)
2025
-
[48]
J., et al
Aguirre, F., Sebastian, A., Le Gallo, M., Song, W., Wang, T., Yang, J. J., et al. Hardware implementation of memristor-based artificial neural networks.Nature Communications,15, Article 1974 (2024)
1974
-
[49]
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J., & Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. InInternational Conference on Learning Representations(2019)
2019
-
[50]
S., Michaud, E., Tegmark, M., & Williams, M
Liu, Z., Kitouni, O., Nolte, N. S., Michaud, E., Tegmark, M., & Williams, M. Towards understanding grokking: An effective theory of representation learning. InAdvances in Neural Information Processing Systems 35, 34651–34663 (2022)
2022
-
[51]
Network dynamics-based framework for understanding deep neural networks.Science China Physics, Mechanics & Astronomy,69, 270511 (2026)
Lin, Y., Feng, S., Zhang, Y., & Zhao, H. Network dynamics-based framework for understanding deep neural networks.Science China Physics, Mechanics & Astronomy,69, 270511 (2026). 22
2026
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.