REVIEW 6 major objections 4 minor 121 references
Compact Bayesian Neural Networks via pruned MCMC sampling
T0 review · 6 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Pruning a Bayesian neural network by each weight's signal-to-noise ratio after MCMC training, then briefly resampling the surviving weights, cuts the network to under a quarter of its size while retaining accuracy and uncertainty estimates.
desk verdict Useful MCMC-BNN pruning recipe with honest reef application, but the uncertainty claim is untested and the convergence evidence is thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the post-pruning resampling stage (Stage 4 of Algorithm 1), applied to weights selected by two pruning ratios computed from MCMC posterior samples: the signal-to-noise ratio $|\mu_i|/\sigma_i$ and the signal-plus-noise ratio $|\mu_i| + \sigma_i$, where $\mu_i$ and $\sigma_i$ are the posterior mean and standard deviation of weight $i$. Weights whose score falls below the user-defined threshold $\lambda$ are set to zero, and the surviving weights are then resampled by Langevin MCMC—the step the paper calls novel for these criteria—so the compact model's posterior can absorb information lost from the pruned weights. Convergence of the resampled chains is checked with the Gelman-Rubin potential scale reduction factor.
What would settle it
Run at least three resampling chains of 10,000 iterations or more on the surviving weights at 75% pruning and compute per-weight Gelman-Rubin values; if many weights exceed an R-hat of 1.1, the short resampling phase has not converged and the uncertainty estimates are not trustworthy. A complementary check is to compare predictive-interval calibration of the compact model on held-out data against the full model.
Extended reading notes
Core claim
Stated on the paper's own terms, the discovery is that a Bayesian neural network trained with Langevin MCMC can be made compact without sacrificing predictive or probabilistic performance: after sampling the posterior for 50,000 iterations, the authors sort weights by either $|\mu_i|/\sigma_i$ (STN) or $|\mu_i|+\sigma_i$ (SPN), zero out those below a threshold $\lambda$, and then resample the surviving weights for 1000 Langevin iterations without burn-in. Across three regression and five classification datasets, including two real-world coral reef lithology sets, the pruned-and-resampled models retained accuracy at 75% pruning, SPN proved best for regression and STN for classification, and resampling improved results for every method. The paper's headline quantitative claim is over 75% network-size reduction with retained generalisation performance.
Load-bearing premise
The method assumes that after pruning, 1000 Langevin MCMC iterations without burn-in are enough for the surviving weights to converge to a faithful posterior, so the compact model's predictions and uncertainties remain valid; the paper itself reports higher R-hat values at high pruning rates.
Editorial extensions
If this is right
- At 75% pruning, structured STN/SPN pruning with resampling preserves benchmark accuracy where random pruning loses it.
- Resampling is an essential post-pruning step: it improved accuracy across all datasets and recovered up to 25% of performance at high pruning levels.
- The criterion choice matters: SPN gives the most precise regression models, while STN gives the most accurate classifiers.
- Compact MCMC-trained BNNs become a viable option for resource-limited deployments that need uncertainty estimates, such as drill-core analysis and marine robotics.
- The pruned-resampled BNNs approach but do not match Random Forest AUC on the reef-core datasets, while adding the uncertainty information Random Forest lacks.
Reading between the lines
- A longer resampling phase with a proper burn-in might reduce the elevated R-hat values the paper reports; the 1000-iteration choice looks like a computational economy rather than a convergence guarantee.
- Because the pruning scores are posterior statistics of the full chain, the method inherits any convergence bias of the initial MCMC run; using tempered or parallel-tempered MCMC could shift which weights are pruned.
- The same STN/SPN scores could be computed from a variational posterior, so a natural testable extension is whether the prune-then-resample recipe transfers to variational Bayesian networks.
- The comparison with Random Forest suggests the practical case for these compact BNNs rests on uncertainty information rather than raw accuracy; a calibration study on the reef-core classes would test that case directly.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a pruning framework for Bayesian neural networks trained with Langevin MCMC. After a full MCMC run, weights and biases are ranked by signal-to-noise (STN) or signal-plus-noise (SPN) criteria, and those below a threshold are zeroed; the surviving parameters are then resampled with a short Langevin MCMC run. The approach is evaluated on three regression and five classification datasets, including two coral reef drill-core lithology datasets, with comparisons to random pruning. The paper claims that the resulting compact BNNs retain generalization performance and uncertainty estimation, and reports experiments at 25%, 50%, and 75% pruning levels.
Significance. If the reported results hold, the paper would provide a practical way to obtain compact MCMC-trained BNNs with uncertainty estimates, potentially useful for real-world applications such as drill-core classification. The empirical comparison of structured versus random pruning across multiple datasets and 30 independent runs is a genuine strength, as is the release of code. However, the paper's most distinctive claim—that post-pruning resampling preserves the posterior's uncertainty estimates—is not directly tested, and the convergence evidence for the resampled chains is weak. The structured-pruning-over-random-pruning pattern is visible in Tables 2 and 3, but several load-bearing presentation choices and missing baselines prevent the central claims from being fully verified.
major comments (6)
- [Section 4.2] The abstract claims that the compact BNN retains its ability to estimate uncertainty via the posterior distribution, but Section 4.2 reports only RMSE, classification accuracy, and AUC. No calibration, coverage, predictive log-likelihood, or predictive-variance metric is measured, so the uncertainty-preservation claim is not directly tested; this is a central claim of the paper and should be verified or removed.
- [Section 4.5 and Algorithm 1 Stage 4] The validity of the post-pruning resampling step is not established. The text reports that the resampled chains show higher R-hat values at high pruning rates, and the rebuttal relies on trace plots of one parameter per dataset in Figure 6. The reported sampling budget is also inconsistent: Section 4.1 says 50,000 samples with an additional 1000 samples without burn-in, while the Figure 6 caption refers to 25,000 post-burn-in samples and 900 post-burnin resampling samples. Please provide proper convergence diagnostics for the resampled posterior (e.g., R-hat for all parameters, effective sample size, or a longer resampling run) and align the reported numbers.
- [Tables 2 and 3] No unpruned baseline is reported in the tables. The columns labeled 'Resampling No' refer to the pruned network without resampling, not to the full BNN, so the reader cannot verify the claim that performance is retained relative to the full network at any pruning level. The paper should report the full-network (0% pruning) RMSE and accuracy alongside the pruned values, so the 'retaining generalisation performance' claim can be assessed.
- [Section 4.4 and Table 3] The statement that 'STN consistently outperforms both RND and SPN across all classification datasets and pruning levels' is not supported by the table. For example, at 25% pruning with resampling, Ionosphere SPN achieves 92.73 versus STN 92.55, and Expedition 310 SPN achieves 37.11 versus STN 36.78; at 50% pruning with resampling on Abalone, SPN achieves 78.37 versus STN 78.27. Please qualify this claim or provide a paired statistical comparison.
- [Algorithm 1, Step 3] The Metropolis-Hastings acceptance probability is incorrectly specified. The expression α = min(1, P(θ')q(θ_i|θ) / P(θ_i)q(θ''|θ_i)) uses an undefined θ'' and does not state the proposal density q for the Langevin proposal in Equation (7). Without the correct proposal-ratio term, the sampler as written is not a valid MH algorithm and cannot be reproduced. Please define q and give the correct acceptance ratio (or state that the implementation uses an alternative valid scheme).
- [Abstract and Sections 4.3-4.4] The abstract claims 'over 75% reduction in network size,' but the experiments evaluate pruning levels of 25%, 50%, and 75% only. The largest tested reduction is exactly 75%; no result supports 'over 75%.' Either test higher pruning rates or revise the abstract to 'up to 75%.'
minor comments (4)
- [Affiliations] The affiliation line contains a typo: 'Asutralia' should be 'Australia.'
- [Section 4.5 and Figure 7] 'German-Rubin' should be 'Gelman-Rubin' in the text, the figure caption, and the conclusion.
- [Section 3.1.3] 'datsets' is a typo for 'datasets.'
- [Algorithm 1, Stage 3] The condition 'Pruning Ratio<λ' is unclear; the manuscript should specify that the pruning criterion is the STN or SPN value for each weight, not a generic ratio.
Circularity Check
No significant circularity: the pruning results are empirical measurements at fixed pruning levels, and the Bayesian setup is standard.
full rationale
The paper's load-bearing claim is that STN/SPN pruning followed by short Langevin resampling compresses MCMC-trained BNNs with limited performance loss. I traced the derivation chain and found no equation or fitted parameter that is renamed as a prediction. The pruning fractions 0.25, 0.50, and 0.75 are user-chosen evaluation levels, not outputs of the method; lambda is a threshold used to set those levels, and the paper explicitly calls it user-defined: 'We select lambda as a constant that determines the threshold for pruning.' The STN and SPN scores are imported from external prior work [54,99] and applied directly, rather than derived from the present paper's results. The likelihoods and priors are written out in full (Equations 8-14) and are standard Gaussian, inverse-Gamma, and multinomial forms; the citations to the authors' own earlier Langevin BNN papers provide background support, not the justification for the empirical pruning comparison. The only internal inconsistency, the high R-hat at high pruning rates contradicted by selected trace plots in Section 4.5, is a convergence and validity concern, not a circularity. No step exhibits the reduction pattern of Equation X equaling Equation Y by construction, or a fitted input being called a prediction.
Assumptions & free parameters
free parameters (4)
- Pruning threshold lambda (per criterion, per dataset) =
Implicitly set to hit fixed pruning fractions (25%, 50%, 75%); numeric values not reported
- Gaussian prior variance sigma-squared on weights =
Not stated in text
- Inverse-Gamma hyperparameters nu1, nu2 for observation noise tau-squared (regression) =
Not stated in text
- Langevin step size epsilon =
Not stated in text
assumptions (4)
- standard math The Langevin proposal (Eq 7) combined with a Metropolis-Hastings acceptance correction yields samples from the correct posterior p(theta | D)
- domain assumption Chains of 50,000 Langevin iterations converge to the posterior despite multimodal BNN posteriors
- ad hoc to paper Zeroing a weight is a valid removal that needs no bias or activation compensation before resampling
- domain assumption Gaussian priors on weights and inverse-Gamma on noise variance are adequate for the data
Cite this review
Pith. "Pith review of Compact Bayesian Neural Networks via pruned MCMC sampling." pith.science (2026). https://pith.science/paper/VZZMNIMN
@misc{pith2026250106962,
author = {Pith},
title = {Pith review of: Compact Bayesian Neural Networks via pruned MCMC sampling},
year = {2026},
howpublished = {\url{https://pith.science/paper/VZZMNIMN}},
note = {Machine review of arXiv:2501.06962}
}
read the original abstract
Bayesian Neural Networks (BNNs) offer robust uncertainty quantification in model predictions, but training them presents a significant computational challenge. This is mainly due to the problem of sampling multimodal posterior distributions using Markov Chain Monte Carlo (MCMC) sampling and variational inference algorithms. Moreover, the number of model parameters scales exponentially with additional hidden layers, neurons, and features in the dataset. Typically, a significant portion of these densely connected parameters are redundant and pruning a neural network not only improves portability but also has the potential for better generalisation capabilities. In this study, we address some of the challenges by leveraging MCMC sampling with network pruning to obtain compact probabilistic models having removed redundant parameters. We sample the posterior distribution of model parameters (weights and biases) and prune weights with low importance, resulting in a compact model. We ensure that the compact BNN retains its ability to estimate uncertainty via the posterior distribution while retaining the model training and generalisation performance accuracy by adapting post-pruning resampling. We evaluate the effectiveness of our MCMC pruning strategy on selected benchmark datasets for regression and classification problems through empirical result analysis. We also consider two coral reef drill-core lithology classification datasets to test the robustness of the pruning model in complex real-world datasets. We further investigate if refining compact BNN can retain any loss of performance. Our results demonstrate the feasibility of training and pruning BNNs using MCMC whilst retaining generalisation performance with over 75% reduction in network size. This paves the way for developing compact BNN models that provide uncertainty estimates for real-world applications.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
M. Abdar, F. Pourpanah, S. Hussain, D. Rezazadegan, L. Liu, M. Ghavamzadeh, P. Fieguth, X. Cao, A. Khosravi, U. R. Acharya, V . Makarenkov, S. Nahavandi, A review of uncertainty quantification in deep learning: Techniques, applications and challenges, Information Fusion 76 (2021) 243–297. doi:10.1016/j.inffus.2021.05.008. URL https://www.sciencedirect.com...
-
[2]
R. Chandra, J. Simmons, Bayesian Neural Networks via MCMC: A Python-Based Tutorial, IEEE Access 12 (2024) 70519–70549. doi: 10.1109/ACCESS.2024.3401234. URL https://ieeexplore.ieee.org/document/10530647/
-
[3]
R. Chandra, K. Jain, R. V . Deo, S. Cripps, Langevin-gradient parallel tempering for Bayesian neural learning, Neurocomputing 359 (2019) 315–326. doi:10.1016/j.neucom.2019.05.082. URL https://www.sciencedirect.com/science/article/ pii/S0925231219308069
-
[4]
Gelman, J
A. Gelman, J. B. Carlin, H. S. Stern, D. B. Rubin, Bayesian Data Analysis, 0th Edition, Chapman and Hall /CRC, 1995. doi:10.1201/ 9780429258411
1995
-
[5]
R. Van De Schoot, S. Depaoli, R. King, B. Kramer, K. M ¨artens, M. G. Tadesse, M. Vannucci, A. Gelman, D. Veen, J. Willemsen, C. Yau, Bayesian statistics and modelling, Nature Reviews Methods Primers 1 (1) (2021) 1. doi:10.1038/s43586-020-00001-2 . URL https://www.nature.com/articles/ s43586-020-00001-2
-
[6]
C. P. Robert, G. Casella, The Metropolis—Hastings Algorithm, in: Monte Carlo Statistical Methods, Springer New York, New York, NY , 2004, pp. 267–320. URL http://link.springer.com/10.1007/ 978-1-4757-4145-2_7
2004
-
[7]
G. O. Roberts, R. L. Tweedie, Exponential convergence of Langevin distributions and their discrete approximations, Bernoulli 2 (4) (1996) 341 – 363, publisher: Bernoulli Society for Mathematical Statistics and Probability
1996
-
[8]
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, L. K. Saul, An introduction to variational methods for graphical models, Machine learning 37 (1999) 183–233, publisher: Springer
1999
Show all 121 references
-
[9]
R. M. Neal, Bayesian learning for neural networks, V ol. 118, Springer Science & Business Media, 2012. 16
2012
- [10]
-
[11]
Louizos, M
C. Louizos, M. Welling, Multiplicative normalizing flows for varia- tional bayesian neural networks, in: International Conference on Ma- chine Learning, PMLR, 2017, pp. 2218–2227
2017
-
[12]
N. M. Nguyen, M.-N. Tran, R. Chandra, Sequential reversible jump MCMC for dynamic Bayesian neural networks, Neurocomputing 564 (2024) 126960. doi:10.1016/j.neucom.2023.126960. URL https://www.sciencedirect.com/science/article/ pii/S0925231223010834
2024
-
[13]
Papamarkou, J
T. Papamarkou, J. Hinkle, M. T. Young, D. Womble, Challenges in Markov Chain Monte Carlo for Bayesian Neural Networks, Statistical Science 37 (3) (Aug. 2022). doi:10.1214/21-STS840
2022 doi
-
[14]
J. Pall, R. Chandra, D. Azam, T. Salles, J. M. Webster, R. Scalzo, S. Cripps, Bayesreef: A Bayesian inference framework for modelling reef growth in response to environmental change and biological dy- namics, Environmental Modelling & Software 125 (2020) 104610. doi:10.1016/j....
2020
- [15]
-
[16]
Liang, J
T. Liang, J. Glossner, L. Wang, S. Shi, X. Zhang, Pruning and quantiza- tion for deep neural network acceleration: A survey, Neurocomputing 461 (2021) 370–403. doi:10.1016/j.neucom.2021.07.045. URL https://linkinghub.elsevier.com/retrieve/pii/ S0925231221010894
2021 doi
-
[17]
Y . He, L. Xiao, Structured Pruning for Deep Convolutional Neural Net- works: A Survey, IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (5) (2024) 2900–2919. doi:10.1109/TPAMI.2023. 3334614. URL https://ieeexplore.ieee.org/document/10330640/
2024
-
[18]
325–333 vol.1
Sietsma, Dow, Neural net pruning-why and how, in: IEEE International Conference on Neural Networks, IEEE, San Diego, CA, USA, 1988, pp. 325–333 vol.1. doi:10.1109/ICNN.1988.23864. URL http://ieeexplore.ieee.org/document/23864/
1988
-
[19]
Rawat, Z
W. Rawat, Z. Wang, Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review, Neural Computation 29 (9) (2017) 2352–2449. doi:10.1162/neco\_a\_00990. URL https://direct.mit.edu/neco/article/29/9/ 2352-2449/8292
2017 doi
-
[20]
Blalock, J
D. Blalock, J. J. Gonzalez Ortiz, J. Frankle, J. Guttag, What is the State of Neural Network Pruning?, Proceedings of Machine Learning and Systems 2 (2020) 129–146. URL https://proceedings.mlsys.org/paper_files/paper/ 2020/hash/6c44dc73014d66ba49b28d483a8f8b0d-Abstract. html
2020
-
[21]
S. Han, J. Pool, J. Tran, W. Dally, Learning both Weights and Connec- tions for Efficient Neural Network, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ ae0eb3eed39d2bcef4622b2...
2015
-
[22]
Sindhwani, T
V . Sindhwani, T. Sainath, S. Kumar, Structured Transforms for Small-Footprint Deep Learning, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ 851300ee84c2b80ed40f51ed26d866fc-Ab...
2015
-
[23]
J. Wu, C. Leng, Y . Wang, Q. Hu, J. Cheng, Quantized Convolutional Neural Networks for Mobile Devices, 2016, pp. 4820–4828. URL https://openaccess.thecvf.com/content_cvpr_2016/ html/Wu_Quantized_Convolutional_Neural_CVPR_2016_ paper.html
2016
-
[24]
Courbariaux, Y
M. Courbariaux, Y . Bengio, J.-P. David, BinaryConnect: Training Deep Neural Networks with binary weights during propagations, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ 3e...
2015
-
[25]
Y . Wang, C. Xu, C. Xu, D. Tao, Packing Convolutional Neural Net- works in the Frequency Domain, IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (10) (2019) 2495–2510. doi:10.1109/ TPAMI.2018.2857824. URL https://ieeexplore.ieee.org/document/8413170/
2019
- [26]
-
[27]
Y . He, X. Zhang, J. Sun, Channel Pruning for Accelerating Very Deep Neural Networks, 2017, pp. 1389–1397. URL https://openaccess.thecvf.com/content_iccv_2017/ html/He_Channel_Pruning_for_ICCV_2017_paper.html
2017
-
[28]
N. Lee, T. Ajanthan, P. Torr, SNIP: SINGLE-SHOT NETWORK PRUN- ING BASED ON CONNECTION SENSITIVITY, 2018. URL https://openreview.net/forum?id=B1VZqjAcYX
2018
- [29]
- [30]
- [31]
-
[32]
A. Onan, S. Koruko ˘glu, H. Bulut, A hybrid ensemble pruning approach based on consensus clustering and multi-objective evolutionary algo- rithm for sentiment classification, Information Processing & Manage- ment 53 (4) (2017) 814–833. doi:10.1016/j.ipm.2017.02.008. URL https:...
2017 doi
-
[33]
F. E. Fernandes Jr, G. G. Yen, Pruning deep convolutional neural net- works architectures with evolution strategy, Information Sciences 552 (2021) 29–47, publisher: Elsevier
2021
-
[34]
J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge Distillation: A Survey, International Journal of Computer Vision 129 (6) (2021) 1789–1819. doi:10.1007/s11263-021-01453-z . URL https://link.springer.com/10.1007/ s11263-021-01453-z
2021 doi
-
[35]
Schmidhuber, Learning Complex, Extended Sequences Using the Principle of History Compression, Neural Computation 4 (2) (1992) 234–242
J. Schmidhuber, Learning Complex, Extended Sequences Using the Principle of History Compression, Neural Computation 4 (2) (1992) 234–242. doi:10.1162/neco.1992.4.2.234. URL https://direct.mit.edu/neco/article/4/2/234-242/ 5634
1992 doi
-
[36]
J. Yim, D. Joo, J. Bae, J. Kim, A Gift From Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning, 2017, pp. 4133–4141. URL https://openaccess.thecvf.com/content_cvpr_2017/ html/Yim_A_Gift_From_CVPR_2017_paper.html
2017
- [37]
-
[38]
G. Chen, W. Choi, X. Yu, T. Han, M. Chandraker, Learning Efficient Ob- ject Detection Models with Knowledge Distillation, in: I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, V ol...
2017
-
[39]
Z. Li, P. Xu, X. Chang, L. Yang, Y . Zhang, L. Yao, X. Chen, When Ob- ject Detection Meets Knowledge Distillation: A Survey, IEEE Transac- tions on Pattern Analysis and Machine Intelligence 45 (8) (2023) 10555– 10579. doi:10.1109/TPAMI.2023.3257546. URL https://ieeexplore.ieee...
2023
-
[40]
Takashima, S
R. Takashima, S. Li, H. Kawai, An Investigation of a Knowledge Dis- tillation Method for CTC Acoustic Models, in: 2018 IEEE International 17 Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Calgary, AB, 2018, pp. 5809–5813. doi:10.1109/ICASSP. 2018.8461995...
2018
-
[41]
Asami, R
T. Asami, R. Masumura, Y . Yamaguchi, H. Masataki, Y . Aono, Domain adaptation of DNN acoustic models using knowledge distillation, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, New Orleans, LA, 2017, pp. 5185–5189. doi:10.11...
2017
-
[42]
H. Fu, S. Zhou, Q. Yang, J. Tang, G. Liu, K. Liu, X. Li, LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding, Proceedings of the AAAI Conference on Ar- tificial Intelligence 35 (14) (2021) 12830–12838.doi:10.1609/aaai. v35i14.1...
2021 doi
- [43]
-
[44]
N, Dropout: A simple way to prevent neural networks from overfit- ting, Journal of Machine Learning Research 15 (1) (2014) 1929
S. N, Dropout: A simple way to prevent neural networks from overfit- ting, Journal of Machine Learning Research 15 (1) (2014) 1929. URL https://cir.nii.ac.jp/crid/1370290617597926172
2014
-
[45]
H. Wu, X. Gu, Towards dropout training for convolutional neural networks, Neural Networks 71 (2015) 1–10. doi: 10.1016/j.neunet.2015.07.007. URL https://www.sciencedirect.com/science/article/ pii/S0893608015001446
2015 doi
-
[46]
S. Park, N. Kwak, Analysis on the Dropout Effect in Convolutional Neu- ral Networks, in: S.-H. Lai, V . Lepetit, K. Nishino, Y . Sato (Eds.), Com- puter Vision – ACCV 2016, Springer International Publishing, Cham, 2017, pp. 189–204. doi:10.1007/978-3-319-54184-6\_12
2016 doi
-
[47]
Y . Gal, Z. Ghahramani, A Theoretically Grounded Application of Dropout in Recurrent Neural Networks, in: Advances in Neural Infor- mation Processing Systems, V ol. 29, Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper/2016/hash/ 076a0c97d09cf1a0ec3e19c7f...
2016
-
[48]
V . Pham, T. Bluche, C. Kermorvant, J. Louradour, Dropout Improves Recurrent Neural Networks for Handwriting Recognition, in: 2014 14th International Conference on Frontiers in Handwriting Recognition, 2014, pp. 285–290, iSSN: 2167-6445. doi:10.1109/ICFHR.2014. 55
2014 doi
-
[49]
C. Lee, K. Cho, W. Kang, Mixout: E ffective Regularization to Finetune Large-scale Pretrained Language Models (Sep. 2019). URL https://arxiv.org/abs/1909.11299v2
2019 arXiv
-
[50]
T. Yang, J. Deng, X. Quan, Q. Wang, S. Nie, AD-DROP: Attribution- Driven Dropout for Robust Language Model Fine-Tuning, Advances in Neural Information Processing Systems 35 (2022) 12310–12324. URL https://proceedings.neurips. cc/paper_files/paper/2022/hash/ 4fdf8d49476a8001c91...
2022
-
[51]
Y . Li, W. Ma, C. Chen, M. Zhang, Y . Liu, S. Ma, Y . Yang, A Survey on Dropout Methods and Experimental Verification in Recommendation, IEEE Transactions on Knowledge and Data Engineering 35 (7) (2023) 6595–6615, conference Name: IEEE Transactions on Knowledge and Data Engine...
2023
-
[52]
Y . Gal, Z. Ghahramani, Dropout as a Bayesian Approximation: Repre- senting Model Uncertainty in Deep Learning, in: Proceedings of The 33rd International Conference on Machine Learning, PMLR, 2016, pp. 1050–1059, iSSN: 1938-7228. URL https://proceedings.mlr.press/v48/gal16.html
2016
-
[53]
J. Hron, A. Matthews, Z. Ghahramani, Variational Bayesian dropout: pitfalls and fixes, in: Proceedings of the 35th International Conference on Machine Learning, PMLR, 2018, pp. 2019–2028, iSSN: 2640-3498. URL https://proceedings.mlr.press/v80/hron18a.html
2018
-
[54]
Graves, Practical Variational Inference for Neural Networks, in: Advances in Neural Information Processing Systems, V ol
A. Graves, Practical Variational Inference for Neural Networks, in: Advances in Neural Information Processing Systems, V ol. 24, Curran Associates, Inc., 2011. URL https://papers.nips.cc/paper_files/paper/2011/ hash/7eb3c8be3d411e8ebfab08eba5f49632-Abstract.html
2011
-
[55]
Sum, C.-s
J. Sum, C.-s. Leung, G. H. Young, L.-w. Chan, W.-k. Kan, An Adaptive Bayesian Pruning for Neural Networks in a Non-Stationary Environ- ment, Neural Computation 11 (4) (1999) 965–976, conference Name: Neural Computation. doi:10.1162/089976699300016539. URL https://ieeexplore.ie...
1999 doi
-
[56]
Sharma, E
H. Sharma, E. Jennings, Bayesian neural networks at scale: a per- formance analysis and pruning study, The Journal of Supercomputing 77 (4) (2021) 3811–3839. doi:10.1007/s11227-020-03401-z . URL https://doi.org/10.1007/s11227-020-03401-z
2021 doi
- [57]
-
[58]
Beckers, B
J. Beckers, B. Van Erp, Z. Zhao, K. Kondrashov, B. De Vries, Principled Pruning of Bayesian Neural Networks through Variational Free Energy Minimization, IEEE Open Journal of Signal Processing (2023) 1–9doi: 10.1109/OJSP.2023.3337718. URL https://ieeexplore.ieee.org/document/10334001/
2023
-
[59]
M. Gu, S. Sun, Neural Langevin Dynamical Sampling, IEEE Ac- cess 8 (2020) 31595–31605, conference Name: IEEE Access. doi:10.1109/ACCESS.2020.2972611. URL https://ieeexplore.ieee.org/abstract/document/ 8988164
2020
-
[60]
Parayil, H
A. Parayil, H. Bai, J. George, P. Gurram, Decentralized Langevin Dynamics for Bayesian Learning, in: Advances in Neural Information Processing Systems, V ol. 33, Curran Associates, Inc., 2020, pp. 15978– 15989. URL https://proceedings.neurips.cc/paper/2020/hash/ b8043b9b976639...
2020
- [61]
-
[62]
G ¨urb¨uzbalaban, X
M. G ¨urb¨uzbalaban, X. Gao, Y . Hu, L. Zhu, Decentralized Stochastic Gradient Langevin Dynamics and Hamiltonian Monte Carlo, Journal of Machine Learning Research 22 (239) (2021) 1–69. URL http://jmlr.org/papers/v22/21-0307.html
2021
-
[63]
Garriga-Alonso, V
A. Garriga-Alonso, V . Fortuin, Exact Langevin Dynamics with Stochas- tic Gradients, 2020. URL https://openreview.net/forum?id=Rprd8aVUYkE
2020
-
[64]
Chandra, L
R. Chandra, L. Azizi, S. Cripps, Bayesian Neural Learning via Langevin Dynamics for Chaotic Time Series Prediction, in: D. Liu, S. Xie, Y . Li, D. Zhao, E.-S. M. El-Alfy (Eds.), Neural Information Process- ing, Springer International Publishing, Cham, 2017, pp. 564–573. doi: 1...
2017 doi
-
[65]
LeCun, J
Y . LeCun, J. Denker, S. Solla, Optimal Brain Damage, in: Advances in Neural Information Processing Systems, V ol. 2, Morgan-Kaufmann, 1989. URL https://proceedings.neurips.cc/paper/1989/hash/ 6c9882bbac1c7093bd25041881277658-Abstract.html
1989
-
[66]
Hassibi, D
B. Hassibi, D. Stork, Second order derivatives for network pruning: Optimal Brain Surgeon, in: Advances in Neural Information Processing Systems, V ol. 5, Morgan-Kaufmann, 1992. URL https://proceedings.neurips.cc/paper/1992/hash/ 303ed4c69846ab36c2904d3ba8573050-Abstract.html
1992
-
[67]
Str ¨om, Phoneme Probability Estimation with Dynamic Sparsely Connected Artificial Neural Networks, The Free Speech Journal (5) (Oct
N. Str ¨om, Phoneme Probability Estimation with Dynamic Sparsely Connected Artificial Neural Networks, The Free Speech Journal (5) (Oct. 1997). URL https://citeseerx.ist.psu.edu/ document?repid=rep1&type=pdf&doi= a9392b9299972452ea6fbc3c605f76bb1e21ae42
1997
- [68]
-
[69]
Mostafa, X
H. Mostafa, X. Wang, Parameter e fficient training of deep convolutional neural networks by dynamic sparse reparameterization, in: Proceed- ings of the 36th International Conference on Machine Learning, PMLR, 2019, pp. 4646–4655, iSSN: 2640-3498. 18 URL https://proceedings.mlr...
2019
- [70]
-
[71]
J. N. Siems, A. Klein, C. Archambeau, M. Mahsereci, Dynamic Pruning of a Neural Network via Gradient Signal-to-Noise Ratio, 2021. URL https://openreview.net/forum?id=34awaeWZgya
2021
-
[72]
Frederick, M
T. Frederick, M. Greg, Deep neural network compression by in-parallel pruning-quantization, IEEE Transactions on Pattern Analysis and Ma- chine Intelligence (2018)
2018
-
[73]
Hoefler, D
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, A. Peste, Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks, Journal of Machine Learning Research 22 (241) (2021) 1–124
2021
-
[74]
S.-K. Yeom, P. Seegerer, S. Lapuschkin, A. Binder, S. Wiedemann, K.- R. M¨uller, W. Samek, Pruning by explaining: A novel criterion for deep neural network pruning, Pattern Recognition 115 (2021) 107899, pub- lisher: Elsevier
2021
-
[75]
Zemouri, N
R. Zemouri, N. Omri, F. Fnaiech, N. Zerhouni, N. Fnaiech, A new growing pruning deep learning neural network algorithm (GP-DLNN), Neural Computing and Applications 32 (24) (2020) 18143–18159, pub- lisher: Springer
2020
-
[76]
Vadera, S
S. Vadera, S. Ameen, Methods for pruning deep neural networks, IEEE Access 10 (2022) 63280–63300, publisher: IEEE
2022
-
[77]
Cheng, M
H. Cheng, M. Zhang, J. Q. Shi, A survey on deep neural network pruning: Taxonomy, comparison, analysis, and recommendations, IEEE Transactions on Pattern Analysis and Machine IntelligencePublisher: IEEE (2024)
2024
- [78]
-
[79]
Anwar, K
S. Anwar, K. Hwang, W. Sung, Structured Pruning of Deep Convolu- tional Neural Networks, J. Emerg. Technol. Comput. Syst. 13 (3) (2017) 32:1–32:18. doi:10.1145/3005348. URL https://doi.org/10.1145/3005348
2017 doi
-
[80]
Y . Tang, S. You, C. Xu, J. Han, C. Qian, B. Shi, C. Xu, C. Zhang, Re- born filters: Pruning convolutional neural networks with limited data, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 34, 2020, pp. 5972–5980, issue: 04
2020
-
[81]
Z. Wang, C. Li, X. Wang, Convolutional neural network pruning with structural redundancy reduction, in: Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, 2021, pp. 14913– 14922
2021
-
[82]
Mantzaris, G
D. Mantzaris, G. Anastassopoulos, A. Adamopoulos, Genetic algorithm pruning of probabilistic neural networks in medi- cal disease estimation, Neural Networks 24 (8) (2011) 831–835. doi:10.1016/j.neunet.2011.06.003. URL https://www.sciencedirect.com/science/article/ pii/S089360...
2011 doi
-
[83]
Poyatos, D
J. Poyatos, D. Molina, A. D. Martinez, J. Del Ser, F. Herrera, Evo- PruneDeepTL: An evolutionary pruning model for transfer learning based deep neural networks, Neural Networks 158 (2023) 59–82. doi:10.1016/j.neunet.2022.10.011. URL https://www.sciencedirect.com/science/articl...
2023 doi
-
[84]
K. O. Stanley, R. Miikkulainen, Evolving Neural Networks through Augmenting Topologies, Evolutionary Computation 10 (2) (2002) 99– 127, conference Name: Evolutionary Computation. doi:10.1162/ 106365602320169811. URL https://ieeexplore.ieee.org/document/6790655
2002
-
[85]
Cant ´u-Paz, Pruning neural networks with distribution estimation algorithms, in: Genetic and Evolutionary Computation Conference, Springer, 2003, pp
E. Cant ´u-Paz, Pruning neural networks with distribution estimation algorithms, in: Genetic and Evolutionary Computation Conference, Springer, 2003, pp. 790–800
2003
-
[86]
Yang, Y .-P
S.-H. Yang, Y .-P. Chen, An evolutionary constructive and pruning algo- rithm for artificial neural networks and its prediction applications, Neu- rocomputing 86 (2012) 140–149, publisher: Elsevier
2012
-
[87]
R. K. Samala, H.-P. Chan, L. M. Hadjiiski, M. A. Helvie, C. Richter, K. Cha, Evolutionary pruning of transfer learned deep convolutional neural network for breast cancer diagnosis in digital breast tomosyn- thesis, Physics in Medicine & Biology 63 (9) (2018) 095005, publisher:...
2018 doi
-
[88]
Y . Zhou, G. G. Yen, Z. Yi, Evolutionary Shallowing Deep Neural Net- works at Block Levels, IEEE Transactions on Neural Networks and Learning Systems 33 (9) (2022) 4635–4647, conference Name: IEEE Transactions on Neural Networks and Learning Systems. doi:10. 1109/TNNLS.2021.3059529
2022
-
[89]
P. M. Williams, Bayesian Regularization and Pruning Using a Laplace Prior, Neural Computation 7 (1) (1995) 117–143. doi:10.1162/neco.1995.7.1.117. URL https://direct.mit.edu/neco/article/7/1/117-143/ 5830
1995 doi
- [90]
-
[91]
van Baalen, C
M. van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y . Wang, T. Blankevoort, M. Welling, Bayesian Bits: Unifying Quantization and Pruning, in: H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, H. Lin (Eds.), Advances in Neural Information Processing Systems, V ol. 33, Curran...
2020
- [92]
-
[93]
S. Chib, E. Greenberg, Understanding the Metropolis-Hastings Algorithm, The American Statistician 49 (4) (1995) 327–335. doi:10.1080/00031305.1995.10476177. URL http://www.tandfonline.com/doi/abs/10.1080/ 00031305.1995.10476177
1995 arXiv
-
[94]
Vladimirova, J
M. Vladimirova, J. Verbeek, P. Mesejo, J. Arbel, Understanding Priors in Bayesian Neural Networks at the Unit Level, International Conference on Machine Learning, 2019, pp. 6458–6467
2019
-
[95]
Izmailov, S
P. Izmailov, S. Vikram, M. D. Ho ffman, A. G. G. Wilson, What Are Bayesian Neural Network Posteriors Really Like?, in: M. Meila, T. Zhang (Eds.), Proceedings of the 38th International Conference on Machine Learning, V ol. 139 of Proceedings of Machine Learning Research, PMLR, ...
2021
-
[96]
Chandra, M
R. Chandra, M. Jain, M. Maharana, P. N. Krivitsky, Revisiting Bayesian autoencoders with MCMC, IEEE Access 10 (2022) 40482–40495, pub- lisher: IEEE
2022
-
[97]
Chandra, A
R. Chandra, A. Bhagat, M. Maharana, P. N. Krivitsky, Bayesian graph convolutional neural networks via tempered MCMC, IEEE Access 9 (2021) 130353–130365, publisher: IEEE
2021
-
[98]
Kapoor, A
A. Kapoor, A. Negi, L. Marshall, R. Chandra, Cyclone trajectory and intensity prediction with uncertainty quantification using variational recurrent neural networks, Environmental Modelling & Software 162 (2023) 105654. doi:10.1016/j.envsoft.2023.105654. URL https://www.scienc...
2023
-
[99]
E. T. Nalisnick, On priors for Bayesian neural networks, University of California, Irvine, 2018
2018
-
[100]
Gelman, D
A. Gelman, D. B. Rubin, Inference from Iterative Simulation Using Mul- tiple Sequences, Statistical Science 7 (4) (Nov. 1992). doi:10.1214/ ss/1177011136
1992
-
[101]
A. S. Weigend, Time series prediction: forecasting the future and under- standing the past, Routledge, 2018
2018
-
[102]
URL https://www.kaggle.com/datasets/robervalt/ sunspots
Solar Influences Data Analysis Center, Sunspots (2018). URL https://www.kaggle.com/datasets/robervalt/ sunspots
2018
-
[103]
Warwick Nash, Tracy Sellers, Simon Talbot, Andrew Cawthorn, Wes Ford, Abalone, published: UCI Machine Learning Repository (1994)
1994
-
[104]
Sigillito, S
V . Sigillito, S. Wing, S. Hutton, K. Baker, Ionosphere, published: UCI Machine Learning Repository (1989). 19
1989
-
[105]
R. A. Fisher, Iris (1936). doi:10.24432/C56C76. URL https://archive.ics.uci.edu/dataset/53
1936 doi
-
[106]
J. M. Webster, Y . Yokoyama, C. Cotterill, the Expedition 325 Scientists, Proceedings of the Integrated Ocean Drilling Program, V ol. 325, Inte- grated Ocean Drilling Program Management International, Inc., Tokyo, 2011. URL https://doi.org/10.2204/iodp.proc.325.2011
2011 doi
-
[107]
G. F. Camoin, Y . Iryu, Expedition 310 Scientists, Proceedings of the Integrated Ocean Drilling Program, V ol. 310, Integrated Ocean Drilling Program Management International, Inc., Washington, DC, 2007. URL https://doi.org/10.2204/iodp.proc.310.2007
2007 doi
-
[108]
C. D. Woodro ffe, J. M. Webster, Coral reefs and sea-level change, Ma- rine Geology 352 (2014) 248–267, publisher: Elsevier
2014
-
[109]
J. M. Webster, J. C. Braga, M. Humblet, D. C. Potts, Y . Iryu, Y . Yokoyama, K. Fujita, R. Bourillot, T. M. Esat, S. Fallon, others, Re- sponse of the Great Barrier Reef to sea-level and environmental changes over the past 30,000 years, Nature Geoscience 11 (6) (2018) 426–432,...
2018
-
[110]
K. L. Sanborn, J. M. Webster, Y . Yokoyama, A. Dutton, J. C. Braga, D. A. Clague, J. B. Paduan, D. Wagner, J. J. Rooney, J. R. Hansen, New evidence of Hawaiian coral reef drowning in response to meltwater pulse-1A, Quaternary Science Reviews 175 (2017) 60–72. URL https://doi.o...
2017 doi
-
[111]
R. Deo, J. M. Webster, T. Salles, R. Chandra, ReefCoreSeg: A Clustering-Based Framework for Multi-Source Data Fusion for Segmen- tation of Reef Drill Cores, IEEE Access 12 (2024) 12164–12180, con- ference Name: IEEE Access. doi:10.1109/ACCESS.2023.3341156. URL https://ieeexplo...
2024
-
[112]
A. P. Bradley, The use of the area under the ROC curve in the evaluation of machine learning algorithms, Pattern recognition 30 (7) (1997) 1145– 1159, publisher: Elsevier
1997
-
[113]
Rocha, S
A. Rocha, S. K. Goldenstein, Multiclass from binary: Expanding one- versus-all, one-versus-one and ecoc-based approaches, IEEE transac- tions on neural networks and learning systems 25 (2) (2013) 289–302, publisher: IEEE
2013
-
[114]
Huang, C
J. Huang, C. Ling, Using AUC and accuracy in evaluating learning algo- rithms, IEEE Transactions on Knowledge and Data Engineering 17 (3) (2005) 299–310. doi:10.1109/TKDE.2005.50
2005 doi
-
[115]
T. K. Ho, Random decision forests, in: Proceedings of 3rd International Conference on Document Analysis and Recognition, V ol. 1, 1995, pp. 278–282 vol.1. doi:10.1109/ICDAR.1995.598994
1995
-
[116]
T. L. Insua, L. Hamel, K. Moran, L. M. Anderson, J. M. Webster, G. F. Camoin, Advanced classification of carbonate sediments based on phys- ical properties, Sedimentology 62 (2) (2015) 590–606. URL https://doi.org/10.1111/sed.12168
2015 doi
-
[117]
Y . Iryu, Y . Takahashi, K. Fujita, G. Camoin, G. Cabioch, H. Mat- suda, T. Sato, K. Sugihara, J. M. Webster, H. Westphal, Sea level his- tory recorded in the Pleistocene carbonate sequence in IODP Hole 310- M0005D, off Tahiti, Island Arc 19 (4) (2010) 690–706. URL https://doi...
2010
-
[118]
Westphal, K
H. Westphal, K. Heindel, M. Brandano, J. Peckmann, Genesis of micro- bialites as contemporaneous framework components of deglacial coral reefs, Tahiti (IODP 310), Facies 56 (3) (2010) 337–352. URL https://doi.org/10.1007/s10347-009-0207-3
2010 doi
-
[119]
Zhang, M
J. Zhang, M. D. Shields, The e ffect of prior probabilities on quantifi- cation and propagation of imprecise probabilities resulting from small datasets, Computer Methods in Applied Mechanics and Engineering 334 (2018) 483–506. doi:10.1016/j.cma.2018.01.045. URL https://www.sc...
2018 doi
-
[120]
Chandra, A
R. Chandra, A. Kapoor, Bayesian neural multi-source transfer learning, Neurocomputing 378 (2020) 54–64, publisher: Elsevier
2020
-
[121]
We show the changes to each model’s posterior prediction performance with 0 to 75 % pruning using each pruning scheme
Appendix 20 Figure 10: Post pruning regression performance over 30 experimental runs on all regression datasets. We show the changes to each model’s posterior prediction performance with 0 to 75 % pruning using each pruning scheme. The bold lines in each plot represent the mea...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.