Pith. sign in

REVIEW 6 major objections 4 minor 121 references

Compact Bayesian Neural Networks via pruned MCMC sampling

T0 review · 6 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Pruning a Bayesian neural network by each weight's signal-to-noise ratio after MCMC training, then briefly resampling the surviving weights, cuts the network to under a quarter of its size while retaining accuracy and uncertainty estimates.

desk verdict Useful MCMC-BNN pruning recipe with honest reef application, but the uncertainty claim is untested and the convergence evidence is thin. read the letter →

arxiv 2501.06962 v1 pith:VZZMNIMN submitted 2025-01-12 cs.LG cs.AI

classification cs.LGcs.AI MSC 62F1568T07
keywords BayesianneuralnetworksMCMCnetworkpruningsignal-to-noiseratioLangevindynamicsuncertaintyquantificationlithologyclassificationmodelcompression
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Bayesian neural networks (BNNs) quantify prediction uncertainty by placing a distribution over every weight, but they are slow to train and costly to deploy. This paper claims that most of those weights are redundant and can be removed after MCMC training, provided the remaining weights are given a short resampling run; with that step, networks pruned to 75% smaller retain their accuracy and their uncertainty estimates on the benchmarks and reef drill-core datasets tested. The pruning is guided by two posterior statistics of each weight—signal-to-noise and signal-plus-noise ratios—and is compared against random pruning, which degrades performance sharply at high pruning rates. The authors argue this makes uncertainty-aware models practical in resource-constrained applications such as drill-core lithology classification, underwater robotics, and remote sensing.

What carries the argument

The load-bearing machinery is the post-pruning resampling stage (Stage 4 of Algorithm 1), applied to weights selected by two pruning ratios computed from MCMC posterior samples: the signal-to-noise ratio $|\mu_i|/\sigma_i$ and the signal-plus-noise ratio $|\mu_i| + \sigma_i$, where $\mu_i$ and $\sigma_i$ are the posterior mean and standard deviation of weight $i$. Weights whose score falls below the user-defined threshold $\lambda$ are set to zero, and the surviving weights are then resampled by Langevin MCMC—the step the paper calls novel for these criteria—so the compact model's posterior can absorb information lost from the pruned weights. Convergence of the resampled chains is checked with the Gelman-Rubin potential scale reduction factor.

What would settle it

Run at least three resampling chains of 10,000 iterations or more on the surviving weights at 75% pruning and compute per-weight Gelman-Rubin values; if many weights exceed an R-hat of 1.1, the short resampling phase has not converged and the uncertainty estimates are not trustworthy. A complementary check is to compare predictive-interval calibration of the compact model on held-out data against the full model.

Watch

Extended reading notes

Core claim

Stated on the paper's own terms, the discovery is that a Bayesian neural network trained with Langevin MCMC can be made compact without sacrificing predictive or probabilistic performance: after sampling the posterior for 50,000 iterations, the authors sort weights by either $|\mu_i|/\sigma_i$ (STN) or $|\mu_i|+\sigma_i$ (SPN), zero out those below a threshold $\lambda$, and then resample the surviving weights for 1000 Langevin iterations without burn-in. Across three regression and five classification datasets, including two real-world coral reef lithology sets, the pruned-and-resampled models retained accuracy at 75% pruning, SPN proved best for regression and STN for classification, and resampling improved results for every method. The paper's headline quantitative claim is over 75% network-size reduction with retained generalisation performance.

Load-bearing premise

The method assumes that after pruning, 1000 Langevin MCMC iterations without burn-in are enough for the surviving weights to converge to a faithful posterior, so the compact model's predictions and uncertainties remain valid; the paper itself reports higher R-hat values at high pruning rates.

Editorial extensions

If this is right

  • At 75% pruning, structured STN/SPN pruning with resampling preserves benchmark accuracy where random pruning loses it.
  • Resampling is an essential post-pruning step: it improved accuracy across all datasets and recovered up to 25% of performance at high pruning levels.
  • The criterion choice matters: SPN gives the most precise regression models, while STN gives the most accurate classifiers.
  • Compact MCMC-trained BNNs become a viable option for resource-limited deployments that need uncertainty estimates, such as drill-core analysis and marine robotics.
  • The pruned-resampled BNNs approach but do not match Random Forest AUC on the reef-core datasets, while adding the uncertainty information Random Forest lacks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A longer resampling phase with a proper burn-in might reduce the elevated R-hat values the paper reports; the 1000-iteration choice looks like a computational economy rather than a convergence guarantee.
  • Because the pruning scores are posterior statistics of the full chain, the method inherits any convergence bias of the initial MCMC run; using tempered or parallel-tempered MCMC could shift which weights are pruned.
  • The same STN/SPN scores could be computed from a variational posterior, so a natural testable extension is whether the prune-then-resample recipe transfers to variational Bayesian networks.
  • The comparison with Random Forest suggests the practical case for these compact BNNs rests on uncertainty information rather than raw accuracy; a calibration study on the reef-core classes would test that case directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 4 minor

Summary. The paper proposes a pruning framework for Bayesian neural networks trained with Langevin MCMC. After a full MCMC run, weights and biases are ranked by signal-to-noise (STN) or signal-plus-noise (SPN) criteria, and those below a threshold are zeroed; the surviving parameters are then resampled with a short Langevin MCMC run. The approach is evaluated on three regression and five classification datasets, including two coral reef drill-core lithology datasets, with comparisons to random pruning. The paper claims that the resulting compact BNNs retain generalization performance and uncertainty estimation, and reports experiments at 25%, 50%, and 75% pruning levels.

Significance. If the reported results hold, the paper would provide a practical way to obtain compact MCMC-trained BNNs with uncertainty estimates, potentially useful for real-world applications such as drill-core classification. The empirical comparison of structured versus random pruning across multiple datasets and 30 independent runs is a genuine strength, as is the release of code. However, the paper's most distinctive claim—that post-pruning resampling preserves the posterior's uncertainty estimates—is not directly tested, and the convergence evidence for the resampled chains is weak. The structured-pruning-over-random-pruning pattern is visible in Tables 2 and 3, but several load-bearing presentation choices and missing baselines prevent the central claims from being fully verified.

major comments (6)
  1. [Section 4.2] The abstract claims that the compact BNN retains its ability to estimate uncertainty via the posterior distribution, but Section 4.2 reports only RMSE, classification accuracy, and AUC. No calibration, coverage, predictive log-likelihood, or predictive-variance metric is measured, so the uncertainty-preservation claim is not directly tested; this is a central claim of the paper and should be verified or removed.
  2. [Section 4.5 and Algorithm 1 Stage 4] The validity of the post-pruning resampling step is not established. The text reports that the resampled chains show higher R-hat values at high pruning rates, and the rebuttal relies on trace plots of one parameter per dataset in Figure 6. The reported sampling budget is also inconsistent: Section 4.1 says 50,000 samples with an additional 1000 samples without burn-in, while the Figure 6 caption refers to 25,000 post-burn-in samples and 900 post-burnin resampling samples. Please provide proper convergence diagnostics for the resampled posterior (e.g., R-hat for all parameters, effective sample size, or a longer resampling run) and align the reported numbers.
  3. [Tables 2 and 3] No unpruned baseline is reported in the tables. The columns labeled 'Resampling No' refer to the pruned network without resampling, not to the full BNN, so the reader cannot verify the claim that performance is retained relative to the full network at any pruning level. The paper should report the full-network (0% pruning) RMSE and accuracy alongside the pruned values, so the 'retaining generalisation performance' claim can be assessed.
  4. [Section 4.4 and Table 3] The statement that 'STN consistently outperforms both RND and SPN across all classification datasets and pruning levels' is not supported by the table. For example, at 25% pruning with resampling, Ionosphere SPN achieves 92.73 versus STN 92.55, and Expedition 310 SPN achieves 37.11 versus STN 36.78; at 50% pruning with resampling on Abalone, SPN achieves 78.37 versus STN 78.27. Please qualify this claim or provide a paired statistical comparison.
  5. [Algorithm 1, Step 3] The Metropolis-Hastings acceptance probability is incorrectly specified. The expression α = min(1, P(θ')q(θ_i|θ) / P(θ_i)q(θ''|θ_i)) uses an undefined θ'' and does not state the proposal density q for the Langevin proposal in Equation (7). Without the correct proposal-ratio term, the sampler as written is not a valid MH algorithm and cannot be reproduced. Please define q and give the correct acceptance ratio (or state that the implementation uses an alternative valid scheme).
  6. [Abstract and Sections 4.3-4.4] The abstract claims 'over 75% reduction in network size,' but the experiments evaluate pruning levels of 25%, 50%, and 75% only. The largest tested reduction is exactly 75%; no result supports 'over 75%.' Either test higher pruning rates or revise the abstract to 'up to 75%.'
minor comments (4)
  1. [Affiliations] The affiliation line contains a typo: 'Asutralia' should be 'Australia.'
  2. [Section 4.5 and Figure 7] 'German-Rubin' should be 'Gelman-Rubin' in the text, the figure caption, and the conclusion.
  3. [Section 3.1.3] 'datsets' is a typo for 'datasets.'
  4. [Algorithm 1, Stage 3] The condition 'Pruning Ratio<λ' is unclear; the manuscript should specify that the pruning criterion is the STN or SPN value for each weight, not a generic ratio.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the pruning results are empirical measurements at fixed pruning levels, and the Bayesian setup is standard.

full rationale

The paper's load-bearing claim is that STN/SPN pruning followed by short Langevin resampling compresses MCMC-trained BNNs with limited performance loss. I traced the derivation chain and found no equation or fitted parameter that is renamed as a prediction. The pruning fractions 0.25, 0.50, and 0.75 are user-chosen evaluation levels, not outputs of the method; lambda is a threshold used to set those levels, and the paper explicitly calls it user-defined: 'We select lambda as a constant that determines the threshold for pruning.' The STN and SPN scores are imported from external prior work [54,99] and applied directly, rather than derived from the present paper's results. The likelihoods and priors are written out in full (Equations 8-14) and are standard Gaussian, inverse-Gamma, and multinomial forms; the citations to the authors' own earlier Langevin BNN papers provide background support, not the justification for the empirical pruning comparison. The only internal inconsistency, the high R-hat at high pruning rates contradicted by selected trace plots in Section 4.5, is a convergence and validity concern, not a circularity. No step exhibits the reduction pattern of Equation X equaling Equation Y by construction, or a fitted input being called a prediction.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The load-bearing ingredients the reader does not pay for upstream: (i) the pruning threshold lambda, calibrated by the authors to hit the reported 25/50/75% sparsity levels rather than set by a principled rule; (ii) numeric hyperparameters (prior variance sigma-squared, inverse-Gamma nu1 and nu2, Langevin step size epsilon) that are not reported in the text and are inherited from related work [2, 3, 64]; (iii) the assumption that 1000 post-pruning MCMC iterations with no burn-in restore a valid posterior, which the paper's own Gelman-Rubin numbers put in doubt. These are tuning and convergence assumptions, not new constants or entities; the criteria themselves (STN, SPN) and the sampler (Langevin MCMC) come from the cited literature.

free parameters (4)
  • Pruning threshold lambda (per criterion, per dataset) = Implicitly set to hit fixed pruning fractions (25%, 50%, 75%); numeric values not reported
    Eqs 15 and 16 define pruning by lambda; Section 3.1.2 says 'For each model, we prune the parameters with the lowest SPN and STN ratios', so lambda is calibrated to reach the chosen sparsity levels. The headlined '75% reduction' is therefore an evaluation point chosen by the authors, not an emergent quantity.
  • Gaussian prior variance sigma-squared on weights = Not stated in text
    Used in log-priors Eqs 8 and 9 and posterior Eqs 13 and 14. The prior is a Gaussian with zero mean and variance sigma-squared; Section 2.2 says priors follow [2, 3], but no numeric value is given in this paper, so the reader cannot tell whether sigma-squared was tuned.
  • Inverse-Gamma hyperparameters nu1, nu2 for observation noise tau-squared (regression) = Not stated in text
    Appear in Eq 8; values are not reported, and the posterior, and hence every weight's posterior mean and standard deviation used for pruning, depends on them.
  • Langevin step size epsilon = Not stated in text
    Eq 7 defines the Langevin proposal with step size epsilon; the quality of mixing, and thus the posterior mean/std used in pruning scores, depends on this choice. No value or adaptation schedule is reported.
assumptions (4)
  • standard math The Langevin proposal (Eq 7) combined with a Metropolis-Hastings acceptance correction yields samples from the correct posterior p(theta | D)
    Section 2.3 and Algorithm 1 Stage 2. This is standard if the acceptance ratio includes the reverse proposal density; the paper does not display that term, so correctness of the sampler is assumed rather than demonstrated.
  • domain assumption Chains of 50,000 Langevin iterations converge to the posterior despite multimodal BNN posteriors
    Section 4.1 sets 50,000 samples; the introduction notes 'the problem of sampling multimodal posterior distributions'. Convergence is claimed via Gelman-Rubin (Figure 7), which can miss lack of convergence on multimodal targets.
  • ad hoc to paper Zeroing a weight is a valid removal that needs no bias or activation compensation before resampling
    Stage 3 of Algorithm 1 sets pruned weights to zero; the only compensation is the 1000-iteration resampling of Stage 4, whose convergence is the weakest assumption.
  • domain assumption Gaussian priors on weights and inverse-Gamma on noise variance are adequate for the data
    Eqs 8-9 and 13-14. The paper itself notes in Section 5 that 'improper or overly simplistic priors may result in biased outcomes, especially when dealing with multimodal distributions in small ecological datasets'. These are the standard BNN priors used in [2, 3, 64].

how reviews work

0 comments
Cite this review

Pith. "Pith review of Compact Bayesian Neural Networks via pruned MCMC sampling." pith.science (2026). https://pith.science/paper/VZZMNIMN

@misc{pith2026250106962,
  author       = {Pith},
  title        = {Pith review of: Compact Bayesian Neural Networks via pruned MCMC sampling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VZZMNIMN}},
  note         = {Machine review of arXiv:2501.06962}
}
read the original abstract

Bayesian Neural Networks (BNNs) offer robust uncertainty quantification in model predictions, but training them presents a significant computational challenge. This is mainly due to the problem of sampling multimodal posterior distributions using Markov Chain Monte Carlo (MCMC) sampling and variational inference algorithms. Moreover, the number of model parameters scales exponentially with additional hidden layers, neurons, and features in the dataset. Typically, a significant portion of these densely connected parameters are redundant and pruning a neural network not only improves portability but also has the potential for better generalisation capabilities. In this study, we address some of the challenges by leveraging MCMC sampling with network pruning to obtain compact probabilistic models having removed redundant parameters. We sample the posterior distribution of model parameters (weights and biases) and prune weights with low importance, resulting in a compact model. We ensure that the compact BNN retains its ability to estimate uncertainty via the posterior distribution while retaining the model training and generalisation performance accuracy by adapting post-pruning resampling. We evaluate the effectiveness of our MCMC pruning strategy on selected benchmark datasets for regression and classification problems through empirical result analysis. We also consider two coral reef drill-core lithology classification datasets to test the robustness of the pruning model in complex real-world datasets. We further investigate if refining compact BNN can retain any loss of performance. Our results demonstrate the feasibility of training and pruning BNNs using MCMC whilst retaining generalisation performance with over 75% reduction in network size. This paves the way for developing compact BNN models that provide uncertainty estimates for real-world applications.

Figures

Figures reproduced from arXiv: 2501.06962 by the authors.

Figure 1
Figure 1. Framework for compact BNNs with network pruning post-sampling (training) where the weights [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Lithology classification of drill core through visual analysis for a [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Class distribution for Expedition 325 and 310 Datasets [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: evaluates the selected pruning methods for BNNs us￾ing three regression datasets: Lazer, Sunspot, and Abalone. We observe that the accuracy (RMSE) generally deteriorates as the pruning level rises, particularly for random pruning. This trend is evident across all datas…
Figure 5
Figure 5. Figure 5: Classification accuracy of Bayesian neural networks across di [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Posterior distribution and parameter trace plots for BNNs models across di [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Gelman-Rubin diagnostic values (Rˆ) for assessing MCMC sampling convergence across different the selected datasets. We computed the diagnostic score for the initial posterior sampling, and after pruning/resampling at different levels (0.25, 0.50, and 0.75). performance…
Figure 9
Figure 9. Figure 9: ROC curve and AUC for the six classes in the Expedition 325 dataset, [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: Post pruning regression performance over 30 experimental runs on all regression datasets. We show the changes to each model’s posterior prediction [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: Post pruning classification performance over 30 experimental runs on all classification datasets. We show the changes to each model’s posterior prediction [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 43 canonical work pages

  1. [1]

    Abdar, F

    M. Abdar, F. Pourpanah, S. Hussain, D. Rezazadegan, L. Liu, M. Ghavamzadeh, P. Fieguth, X. Cao, A. Khosravi, U. R. Acharya, V . Makarenkov, S. Nahavandi, A review of uncertainty quantification in deep learning: Techniques, applications and challenges, Information Fusion 76 (2021) 243–297. doi:10.1016/j.inffus.2021.05.008. URL https://www.sciencedirect.com...

  2. [2]

    Chandra, J

    R. Chandra, J. Simmons, Bayesian Neural Networks via MCMC: A Python-Based Tutorial, IEEE Access 12 (2024) 70519–70549. doi: 10.1109/ACCESS.2024.3401234. URL https://ieeexplore.ieee.org/document/10530647/

  3. [3]

    Chandra, K

    R. Chandra, K. Jain, R. V . Deo, S. Cripps, Langevin-gradient parallel tempering for Bayesian neural learning, Neurocomputing 359 (2019) 315–326. doi:10.1016/j.neucom.2019.05.082. URL https://www.sciencedirect.com/science/article/ pii/S0925231219308069

  4. [4]

    Gelman, J

    A. Gelman, J. B. Carlin, H. S. Stern, D. B. Rubin, Bayesian Data Analysis, 0th Edition, Chapman and Hall /CRC, 1995. doi:10.1201/ 9780429258411

  5. [5]

    Van De Schoot, S

    R. Van De Schoot, S. Depaoli, R. King, B. Kramer, K. M ¨artens, M. G. Tadesse, M. Vannucci, A. Gelman, D. Veen, J. Willemsen, C. Yau, Bayesian statistics and modelling, Nature Reviews Methods Primers 1 (1) (2021) 1. doi:10.1038/s43586-020-00001-2 . URL https://www.nature.com/articles/ s43586-020-00001-2

  6. [6]

    C. P. Robert, G. Casella, The Metropolis—Hastings Algorithm, in: Monte Carlo Statistical Methods, Springer New York, New York, NY , 2004, pp. 267–320. URL http://link.springer.com/10.1007/ 978-1-4757-4145-2_7

  7. [7]

    G. O. Roberts, R. L. Tweedie, Exponential convergence of Langevin distributions and their discrete approximations, Bernoulli 2 (4) (1996) 341 – 363, publisher: Bernoulli Society for Mathematical Statistics and Probability

  8. [8]

    M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, L. K. Saul, An introduction to variational methods for graphical models, Machine learning 37 (1999) 183–233, publisher: Springer

Show all 121 references
  1. [9]

    R. M. Neal, Bayesian learning for neural networks, V ol. 118, Springer Science & Business Media, 2012. 16

  2. [10]

    Y . Gal, Z. Ghahramani, Bayesian Convolutional Neural Networks with Bernoulli Approximate Variational Inference, arXiv:1506.02158 (Jan. 2016). doi:10.48550/arXiv.1506.02158. URL http://arxiv.org/abs/1506.02158

  3. [11]

    Louizos, M

    C. Louizos, M. Welling, Multiplicative normalizing flows for varia- tional bayesian neural networks, in: International Conference on Ma- chine Learning, PMLR, 2017, pp. 2218–2227

  4. [12]

    N. M. Nguyen, M.-N. Tran, R. Chandra, Sequential reversible jump MCMC for dynamic Bayesian neural networks, Neurocomputing 564 (2024) 126960. doi:10.1016/j.neucom.2023.126960. URL https://www.sciencedirect.com/science/article/ pii/S0925231223010834

  5. [13]

    Papamarkou, J

    T. Papamarkou, J. Hinkle, M. T. Young, D. Womble, Challenges in Markov Chain Monte Carlo for Bayesian Neural Networks, Statistical Science 37 (3) (Aug. 2022). doi:10.1214/21-STS840

  6. [14]

    J. Pall, R. Chandra, D. Azam, T. Salles, J. M. Webster, R. Scalzo, S. Cripps, Bayesreef: A Bayesian inference framework for modelling reef growth in response to environmental change and biological dy- namics, Environmental Modelling & Software 125 (2020) 104610. doi:10.1016/j....

  7. [15]

    Trippe, R

    B. Trippe, R. Turner, Overpruning in Variational Bayesian Neural Net- works, arXiv:1801.06230 [stat] (Jan. 2018). doi:10.48550/arXiv. 1801.06230. URL http://arxiv.org/abs/1801.06230

  8. [16]

    Liang, J

    T. Liang, J. Glossner, L. Wang, S. Shi, X. Zhang, Pruning and quantiza- tion for deep neural network acceleration: A survey, Neurocomputing 461 (2021) 370–403. doi:10.1016/j.neucom.2021.07.045. URL https://linkinghub.elsevier.com/retrieve/pii/ S0925231221010894

  9. [17]

    Y . He, L. Xiao, Structured Pruning for Deep Convolutional Neural Net- works: A Survey, IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (5) (2024) 2900–2919. doi:10.1109/TPAMI.2023. 3334614. URL https://ieeexplore.ieee.org/document/10330640/

  10. [18]

    325–333 vol.1

    Sietsma, Dow, Neural net pruning-why and how, in: IEEE International Conference on Neural Networks, IEEE, San Diego, CA, USA, 1988, pp. 325–333 vol.1. doi:10.1109/ICNN.1988.23864. URL http://ieeexplore.ieee.org/document/23864/

  11. [19]

    Rawat, Z

    W. Rawat, Z. Wang, Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review, Neural Computation 29 (9) (2017) 2352–2449. doi:10.1162/neco\_a\_00990. URL https://direct.mit.edu/neco/article/29/9/ 2352-2449/8292

  12. [20]

    Blalock, J

    D. Blalock, J. J. Gonzalez Ortiz, J. Frankle, J. Guttag, What is the State of Neural Network Pruning?, Proceedings of Machine Learning and Systems 2 (2020) 129–146. URL https://proceedings.mlsys.org/paper_files/paper/ 2020/hash/6c44dc73014d66ba49b28d483a8f8b0d-Abstract. html

  13. [21]

    S. Han, J. Pool, J. Tran, W. Dally, Learning both Weights and Connec- tions for Efficient Neural Network, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ ae0eb3eed39d2bcef4622b2...

  14. [22]

    Sindhwani, T

    V . Sindhwani, T. Sainath, S. Kumar, Structured Transforms for Small-Footprint Deep Learning, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ 851300ee84c2b80ed40f51ed26d866fc-Ab...

  15. [23]

    J. Wu, C. Leng, Y . Wang, Q. Hu, J. Cheng, Quantized Convolutional Neural Networks for Mobile Devices, 2016, pp. 4820–4828. URL https://openaccess.thecvf.com/content_cvpr_2016/ html/Wu_Quantized_Convolutional_Neural_CVPR_2016_ paper.html

  16. [24]

    Courbariaux, Y

    M. Courbariaux, Y . Bengio, J.-P. David, BinaryConnect: Training Deep Neural Networks with binary weights during propagations, in: Advances in Neural Information Processing Systems, V ol. 28, Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/hash/ 3e...

  17. [25]

    Y . Wang, C. Xu, C. Xu, D. Tao, Packing Convolutional Neural Net- works in the Frequency Domain, IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (10) (2019) 2495–2510. doi:10.1109/ TPAMI.2018.2857824. URL https://ieeexplore.ieee.org/document/8413170/

  18. [26]

    H. Li, A. Kadav, I. Durdanovic, H. Samet, H. P. Graf, Pruning Filters for Efficient ConvNets, arXiv:1608.08710 [cs] (Mar. 2017). doi:10. 48550/arXiv.1608.08710. URL http://arxiv.org/abs/1608.08710

  19. [27]

    Y . He, X. Zhang, J. Sun, Channel Pruning for Accelerating Very Deep Neural Networks, 2017, pp. 1389–1397. URL https://openaccess.thecvf.com/content_iccv_2017/ html/He_Channel_Pruning_for_ICCV_2017_paper.html

  20. [28]

    N. Lee, T. Ajanthan, P. Torr, SNIP: SINGLE-SHOT NETWORK PRUN- ING BASED ON CONNECTION SENSITIVITY, 2018. URL https://openreview.net/forum?id=B1VZqjAcYX

  21. [29]

    Frankle, G

    J. Frankle, G. K. Dziugaite, D. M. Roy, M. Carbin, Stabilizing the Lot- tery Ticket Hypothesis, arXiv:1903.01611 [cs, stat] (Jul. 2020). doi: 10.48550/arXiv.1903.01611. URL http://arxiv.org/abs/1903.01611

  22. [30]

    Z. Liu, M. Sun, T. Zhou, G. Huang, T. Darrell, Rethinking the Value of Network Pruning, arXiv:1810.05270 [cs, stat] (Mar. 2019). doi: 10.48550/arXiv.1810.05270. URL http://arxiv.org/abs/1810.05270

  23. [31]

    T. Gale, E. Elsen, S. Hooker, The State of Sparsity in Deep Neural Networks, arXiv:1902.09574 [cs, stat] (Feb. 2019). doi:10.48550/ arXiv.1902.09574. URL http://arxiv.org/abs/1902.09574

  24. [32]

    A. Onan, S. Koruko ˘glu, H. Bulut, A hybrid ensemble pruning approach based on consensus clustering and multi-objective evolutionary algo- rithm for sentiment classification, Information Processing & Manage- ment 53 (4) (2017) 814–833. doi:10.1016/j.ipm.2017.02.008. URL https:...

  25. [33]

    F. E. Fernandes Jr, G. G. Yen, Pruning deep convolutional neural net- works architectures with evolution strategy, Information Sciences 552 (2021) 29–47, publisher: Elsevier

  26. [34]

    J. Gou, B. Yu, S. J. Maybank, D. Tao, Knowledge Distillation: A Survey, International Journal of Computer Vision 129 (6) (2021) 1789–1819. doi:10.1007/s11263-021-01453-z . URL https://link.springer.com/10.1007/ s11263-021-01453-z

  27. [35]

    Schmidhuber, Learning Complex, Extended Sequences Using the Principle of History Compression, Neural Computation 4 (2) (1992) 234–242

    J. Schmidhuber, Learning Complex, Extended Sequences Using the Principle of History Compression, Neural Computation 4 (2) (1992) 234–242. doi:10.1162/neco.1992.4.2.234. URL https://direct.mit.edu/neco/article/4/2/234-242/ 5634

  28. [36]

    J. Yim, D. Joo, J. Bae, J. Kim, A Gift From Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning, 2017, pp. 4133–4141. URL https://openaccess.thecvf.com/content_cvpr_2017/ html/Yim_A_Gift_From_CVPR_2017_paper.html

  29. [37]

    Hinton, O

    G. Hinton, O. Vinyals, J. Dean, Distilling the Knowledge in a Neural Network, version Number: 1 (2015). doi:10.48550/ARXIV.1503. 02531. URL https://arxiv.org/abs/1503.02531

  30. [38]

    G. Chen, W. Choi, X. Yu, T. Han, M. Chandraker, Learning Efficient Ob- ject Detection Models with Knowledge Distillation, in: I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, R. Garnett (Eds.), Advances in Neural Information Processing Systems, V ol...

  31. [39]

    Z. Li, P. Xu, X. Chang, L. Yang, Y . Zhang, L. Yao, X. Chen, When Ob- ject Detection Meets Knowledge Distillation: A Survey, IEEE Transac- tions on Pattern Analysis and Machine Intelligence 45 (8) (2023) 10555– 10579. doi:10.1109/TPAMI.2023.3257546. URL https://ieeexplore.ieee...

  32. [40]

    Takashima, S

    R. Takashima, S. Li, H. Kawai, An Investigation of a Knowledge Dis- tillation Method for CTC Acoustic Models, in: 2018 IEEE International 17 Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Calgary, AB, 2018, pp. 5809–5813. doi:10.1109/ICASSP. 2018.8461995...

  33. [41]

    Asami, R

    T. Asami, R. Masumura, Y . Yamaguchi, H. Masataki, Y . Aono, Domain adaptation of DNN acoustic models using knowledge distillation, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, New Orleans, LA, 2017, pp. 5185–5189. doi:10.11...

  34. [42]

    H. Fu, S. Zhou, Q. Yang, J. Tang, G. Liu, K. Liu, X. Li, LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding, Proceedings of the AAAI Conference on Ar- tificial Intelligence 35 (14) (2021) 12830–12838.doi:10.1609/aaai. v35i14.1...

  35. [43]

    X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, Q. Liu, TinyBERT: Distilling BERT for Natural Language Understanding, ver- sion Number: 5 (2019). doi:10.48550/ARXIV.1909.10351. URL https://arxiv.org/abs/1909.10351

  36. [44]

    N, Dropout: A simple way to prevent neural networks from overfit- ting, Journal of Machine Learning Research 15 (1) (2014) 1929

    S. N, Dropout: A simple way to prevent neural networks from overfit- ting, Journal of Machine Learning Research 15 (1) (2014) 1929. URL https://cir.nii.ac.jp/crid/1370290617597926172

  37. [45]

    H. Wu, X. Gu, Towards dropout training for convolutional neural networks, Neural Networks 71 (2015) 1–10. doi: 10.1016/j.neunet.2015.07.007. URL https://www.sciencedirect.com/science/article/ pii/S0893608015001446

  38. [46]

    S. Park, N. Kwak, Analysis on the Dropout Effect in Convolutional Neu- ral Networks, in: S.-H. Lai, V . Lepetit, K. Nishino, Y . Sato (Eds.), Com- puter Vision – ACCV 2016, Springer International Publishing, Cham, 2017, pp. 189–204. doi:10.1007/978-3-319-54184-6\_12

  39. [47]

    Y . Gal, Z. Ghahramani, A Theoretically Grounded Application of Dropout in Recurrent Neural Networks, in: Advances in Neural Infor- mation Processing Systems, V ol. 29, Curran Associates, Inc., 2016. URL https://proceedings.neurips.cc/paper/2016/hash/ 076a0c97d09cf1a0ec3e19c7f...

  40. [48]

    V . Pham, T. Bluche, C. Kermorvant, J. Louradour, Dropout Improves Recurrent Neural Networks for Handwriting Recognition, in: 2014 14th International Conference on Frontiers in Handwriting Recognition, 2014, pp. 285–290, iSSN: 2167-6445. doi:10.1109/ICFHR.2014. 55

  41. [49]

    C. Lee, K. Cho, W. Kang, Mixout: E ffective Regularization to Finetune Large-scale Pretrained Language Models (Sep. 2019). URL https://arxiv.org/abs/1909.11299v2

  42. [50]

    T. Yang, J. Deng, X. Quan, Q. Wang, S. Nie, AD-DROP: Attribution- Driven Dropout for Robust Language Model Fine-Tuning, Advances in Neural Information Processing Systems 35 (2022) 12310–12324. URL https://proceedings.neurips. cc/paper_files/paper/2022/hash/ 4fdf8d49476a8001c91...

  43. [51]

    Y . Li, W. Ma, C. Chen, M. Zhang, Y . Liu, S. Ma, Y . Yang, A Survey on Dropout Methods and Experimental Verification in Recommendation, IEEE Transactions on Knowledge and Data Engineering 35 (7) (2023) 6595–6615, conference Name: IEEE Transactions on Knowledge and Data Engine...

  44. [52]

    Y . Gal, Z. Ghahramani, Dropout as a Bayesian Approximation: Repre- senting Model Uncertainty in Deep Learning, in: Proceedings of The 33rd International Conference on Machine Learning, PMLR, 2016, pp. 1050–1059, iSSN: 1938-7228. URL https://proceedings.mlr.press/v48/gal16.html

  45. [53]

    J. Hron, A. Matthews, Z. Ghahramani, Variational Bayesian dropout: pitfalls and fixes, in: Proceedings of the 35th International Conference on Machine Learning, PMLR, 2018, pp. 2019–2028, iSSN: 2640-3498. URL https://proceedings.mlr.press/v80/hron18a.html

  46. [54]

    Graves, Practical Variational Inference for Neural Networks, in: Advances in Neural Information Processing Systems, V ol

    A. Graves, Practical Variational Inference for Neural Networks, in: Advances in Neural Information Processing Systems, V ol. 24, Curran Associates, Inc., 2011. URL https://papers.nips.cc/paper_files/paper/2011/ hash/7eb3c8be3d411e8ebfab08eba5f49632-Abstract.html

  47. [55]

    Sum, C.-s

    J. Sum, C.-s. Leung, G. H. Young, L.-w. Chan, W.-k. Kan, An Adaptive Bayesian Pruning for Neural Networks in a Non-Stationary Environ- ment, Neural Computation 11 (4) (1999) 965–976, conference Name: Neural Computation. doi:10.1162/089976699300016539. URL https://ieeexplore.ie...

  48. [56]

    Sharma, E

    H. Sharma, E. Jennings, Bayesian neural networks at scale: a per- formance analysis and pruning study, The Journal of Supercomputing 77 (4) (2021) 3811–3839. doi:10.1007/s11227-020-03401-z . URL https://doi.org/10.1007/s11227-020-03401-z

  49. [57]

    Blundell, J

    C. Blundell, J. Cornebise, K. Kavukcuoglu, D. Wierstra, Weight Un- certainty in Neural Networks, arXiv:1505.05424 [cs, stat] (May 2015). doi:10.48550/arXiv.1505.05424. URL http://arxiv.org/abs/1505.05424

  50. [58]

    Beckers, B

    J. Beckers, B. Van Erp, Z. Zhao, K. Kondrashov, B. De Vries, Principled Pruning of Bayesian Neural Networks through Variational Free Energy Minimization, IEEE Open Journal of Signal Processing (2023) 1–9doi: 10.1109/OJSP.2023.3337718. URL https://ieeexplore.ieee.org/document/10334001/

  51. [59]

    M. Gu, S. Sun, Neural Langevin Dynamical Sampling, IEEE Ac- cess 8 (2020) 31595–31605, conference Name: IEEE Access. doi:10.1109/ACCESS.2020.2972611. URL https://ieeexplore.ieee.org/abstract/document/ 8988164

  52. [60]

    Parayil, H

    A. Parayil, H. Bai, J. George, P. Gurram, Decentralized Langevin Dynamics for Bayesian Learning, in: Advances in Neural Information Processing Systems, V ol. 33, Curran Associates, Inc., 2020, pp. 15978– 15989. URL https://proceedings.neurips.cc/paper/2020/hash/ b8043b9b976639...

  53. [61]

    A. Look, M. Kandemir, Di fferential Bayesian Neural Nets, arXiv:1912.00796 [cs, stat] (Feb. 2020). doi:10.48550/arXiv. 1912.00796. URL http://arxiv.org/abs/1912.00796

  54. [62]

    G ¨urb¨uzbalaban, X

    M. G ¨urb¨uzbalaban, X. Gao, Y . Hu, L. Zhu, Decentralized Stochastic Gradient Langevin Dynamics and Hamiltonian Monte Carlo, Journal of Machine Learning Research 22 (239) (2021) 1–69. URL http://jmlr.org/papers/v22/21-0307.html

  55. [63]

    Garriga-Alonso, V

    A. Garriga-Alonso, V . Fortuin, Exact Langevin Dynamics with Stochas- tic Gradients, 2020. URL https://openreview.net/forum?id=Rprd8aVUYkE

  56. [64]

    Chandra, L

    R. Chandra, L. Azizi, S. Cripps, Bayesian Neural Learning via Langevin Dynamics for Chaotic Time Series Prediction, in: D. Liu, S. Xie, Y . Li, D. Zhao, E.-S. M. El-Alfy (Eds.), Neural Information Process- ing, Springer International Publishing, Cham, 2017, pp. 564–573. doi: 1...

  57. [65]

    LeCun, J

    Y . LeCun, J. Denker, S. Solla, Optimal Brain Damage, in: Advances in Neural Information Processing Systems, V ol. 2, Morgan-Kaufmann, 1989. URL https://proceedings.neurips.cc/paper/1989/hash/ 6c9882bbac1c7093bd25041881277658-Abstract.html

  58. [66]

    Hassibi, D

    B. Hassibi, D. Stork, Second order derivatives for network pruning: Optimal Brain Surgeon, in: Advances in Neural Information Processing Systems, V ol. 5, Morgan-Kaufmann, 1992. URL https://proceedings.neurips.cc/paper/1992/hash/ 303ed4c69846ab36c2904d3ba8573050-Abstract.html

  59. [67]

    Str ¨om, Phoneme Probability Estimation with Dynamic Sparsely Connected Artificial Neural Networks, The Free Speech Journal (5) (Oct

    N. Str ¨om, Phoneme Probability Estimation with Dynamic Sparsely Connected Artificial Neural Networks, The Free Speech Journal (5) (Oct. 1997). URL https://citeseerx.ist.psu.edu/ document?repid=rep1&type=pdf&doi= a9392b9299972452ea6fbc3c605f76bb1e21ae42

  60. [68]

    M. Zhu, S. Gupta, To prune, or not to prune: exploring the e fficacy of pruning for model compression, arXiv:1710.01878 [cs, stat] (Nov. 2017). doi:10.48550/arXiv.1710.01878. URL http://arxiv.org/abs/1710.01878

  61. [69]

    Mostafa, X

    H. Mostafa, X. Wang, Parameter e fficient training of deep convolutional neural networks by dynamic sparse reparameterization, in: Proceed- ings of the 36th International Conference on Machine Learning, PMLR, 2019, pp. 4646–4655, iSSN: 2640-3498. 18 URL https://proceedings.mlr...

  62. [70]

    Dettmers, L

    T. Dettmers, L. Zettlemoyer, Sparse Networks from Scratch: Faster Training without Losing Performance, arXiv:1907.04840 [cs, stat] (Aug. 2019). doi:10.48550/arXiv.1907.04840. URL http://arxiv.org/abs/1907.04840

  63. [71]

    J. N. Siems, A. Klein, C. Archambeau, M. Mahsereci, Dynamic Pruning of a Neural Network via Gradient Signal-to-Noise Ratio, 2021. URL https://openreview.net/forum?id=34awaeWZgya

  64. [72]

    Frederick, M

    T. Frederick, M. Greg, Deep neural network compression by in-parallel pruning-quantization, IEEE Transactions on Pattern Analysis and Ma- chine Intelligence (2018)

  65. [73]

    Hoefler, D

    T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, A. Peste, Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks, Journal of Machine Learning Research 22 (241) (2021) 1–124

  66. [74]

    S.-K. Yeom, P. Seegerer, S. Lapuschkin, A. Binder, S. Wiedemann, K.- R. M¨uller, W. Samek, Pruning by explaining: A novel criterion for deep neural network pruning, Pattern Recognition 115 (2021) 107899, pub- lisher: Elsevier

  67. [75]

    Zemouri, N

    R. Zemouri, N. Omri, F. Fnaiech, N. Zerhouni, N. Fnaiech, A new growing pruning deep learning neural network algorithm (GP-DLNN), Neural Computing and Applications 32 (24) (2020) 18143–18159, pub- lisher: Springer

  68. [76]

    Vadera, S

    S. Vadera, S. Ameen, Methods for pruning deep neural networks, IEEE Access 10 (2022) 63280–63300, publisher: IEEE

  69. [77]

    Cheng, M

    H. Cheng, M. Zhang, J. Q. Shi, A survey on deep neural network pruning: Taxonomy, comparison, analysis, and recommendations, IEEE Transactions on Pattern Analysis and Machine IntelligencePublisher: IEEE (2024)

  70. [78]

    Molchanov, S

    P. Molchanov, S. Tyree, T. Karras, T. Aila, J. Kautz, Pruning Convolutional Neural Networks for Resource E fficient Inference, arXiv:1611.06440 [cs] (Jun. 2017). doi:10.48550/arXiv.1611. 06440. URL http://arxiv.org/abs/1611.06440

  71. [79]

    Anwar, K

    S. Anwar, K. Hwang, W. Sung, Structured Pruning of Deep Convolu- tional Neural Networks, J. Emerg. Technol. Comput. Syst. 13 (3) (2017) 32:1–32:18. doi:10.1145/3005348. URL https://doi.org/10.1145/3005348

  72. [80]

    Y . Tang, S. You, C. Xu, J. Han, C. Qian, B. Shi, C. Xu, C. Zhang, Re- born filters: Pruning convolutional neural networks with limited data, in: Proceedings of the AAAI Conference on Artificial Intelligence, V ol. 34, 2020, pp. 5972–5980, issue: 04

  73. [81]

    Z. Wang, C. Li, X. Wang, Convolutional neural network pruning with structural redundancy reduction, in: Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, 2021, pp. 14913– 14922

  74. [82]

    Mantzaris, G

    D. Mantzaris, G. Anastassopoulos, A. Adamopoulos, Genetic algorithm pruning of probabilistic neural networks in medi- cal disease estimation, Neural Networks 24 (8) (2011) 831–835. doi:10.1016/j.neunet.2011.06.003. URL https://www.sciencedirect.com/science/article/ pii/S089360...

  75. [83]

    Poyatos, D

    J. Poyatos, D. Molina, A. D. Martinez, J. Del Ser, F. Herrera, Evo- PruneDeepTL: An evolutionary pruning model for transfer learning based deep neural networks, Neural Networks 158 (2023) 59–82. doi:10.1016/j.neunet.2022.10.011. URL https://www.sciencedirect.com/science/articl...

  76. [84]

    K. O. Stanley, R. Miikkulainen, Evolving Neural Networks through Augmenting Topologies, Evolutionary Computation 10 (2) (2002) 99– 127, conference Name: Evolutionary Computation. doi:10.1162/ 106365602320169811. URL https://ieeexplore.ieee.org/document/6790655

  77. [85]

    Cant ´u-Paz, Pruning neural networks with distribution estimation algorithms, in: Genetic and Evolutionary Computation Conference, Springer, 2003, pp

    E. Cant ´u-Paz, Pruning neural networks with distribution estimation algorithms, in: Genetic and Evolutionary Computation Conference, Springer, 2003, pp. 790–800

  78. [86]

    Yang, Y .-P

    S.-H. Yang, Y .-P. Chen, An evolutionary constructive and pruning algo- rithm for artificial neural networks and its prediction applications, Neu- rocomputing 86 (2012) 140–149, publisher: Elsevier

  79. [87]

    R. K. Samala, H.-P. Chan, L. M. Hadjiiski, M. A. Helvie, C. Richter, K. Cha, Evolutionary pruning of transfer learned deep convolutional neural network for breast cancer diagnosis in digital breast tomosyn- thesis, Physics in Medicine & Biology 63 (9) (2018) 095005, publisher:...

  80. [88]

    Y . Zhou, G. G. Yen, Z. Yi, Evolutionary Shallowing Deep Neural Net- works at Block Levels, IEEE Transactions on Neural Networks and Learning Systems 33 (9) (2022) 4635–4647, conference Name: IEEE Transactions on Neural Networks and Learning Systems. doi:10. 1109/TNNLS.2021.3059529

  81. [89]

    P. M. Williams, Bayesian Regularization and Pruning Using a Laplace Prior, Neural Computation 7 (1) (1995) 117–143. doi:10.1162/neco.1995.7.1.117. URL https://direct.mit.edu/neco/article/7/1/117-143/ 5830

  82. [90]

    Neklyudov, D

    K. Neklyudov, D. Molchanov, A. Ashukha, D. Vetrov, Struc- tured Bayesian Pruning via Log-Normal Multiplicative Noise, arXiv:1705.07283 [stat] (Nov. 2017). doi:10.48550/arXiv.1705. 07283. URL http://arxiv.org/abs/1705.07283

  83. [91]

    van Baalen, C

    M. van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y . Wang, T. Blankevoort, M. Welling, Bayesian Bits: Unifying Quantization and Pruning, in: H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, H. Lin (Eds.), Advances in Neural Information Processing Systems, V ol. 33, Curran...

  84. [92]

    Mathew, D

    S. Mathew, D. B. Rowe, Pruning a neural network using Bayesian infer- ence, arXiv:2308.02451 [cs, stat] (Aug. 2023).doi:10.48550/arXiv. 2308.02451. URL http://arxiv.org/abs/2308.02451

  85. [93]

    S. Chib, E. Greenberg, Understanding the Metropolis-Hastings Algorithm, The American Statistician 49 (4) (1995) 327–335. doi:10.1080/00031305.1995.10476177. URL http://www.tandfonline.com/doi/abs/10.1080/ 00031305.1995.10476177

  86. [94]

    Vladimirova, J

    M. Vladimirova, J. Verbeek, P. Mesejo, J. Arbel, Understanding Priors in Bayesian Neural Networks at the Unit Level, International Conference on Machine Learning, 2019, pp. 6458–6467

  87. [95]

    Izmailov, S

    P. Izmailov, S. Vikram, M. D. Ho ffman, A. G. G. Wilson, What Are Bayesian Neural Network Posteriors Really Like?, in: M. Meila, T. Zhang (Eds.), Proceedings of the 38th International Conference on Machine Learning, V ol. 139 of Proceedings of Machine Learning Research, PMLR, ...

  88. [96]

    Chandra, M

    R. Chandra, M. Jain, M. Maharana, P. N. Krivitsky, Revisiting Bayesian autoencoders with MCMC, IEEE Access 10 (2022) 40482–40495, pub- lisher: IEEE

  89. [97]

    Chandra, A

    R. Chandra, A. Bhagat, M. Maharana, P. N. Krivitsky, Bayesian graph convolutional neural networks via tempered MCMC, IEEE Access 9 (2021) 130353–130365, publisher: IEEE

  90. [98]

    Kapoor, A

    A. Kapoor, A. Negi, L. Marshall, R. Chandra, Cyclone trajectory and intensity prediction with uncertainty quantification using variational recurrent neural networks, Environmental Modelling & Software 162 (2023) 105654. doi:10.1016/j.envsoft.2023.105654. URL https://www.scienc...

  91. [99]

    E. T. Nalisnick, On priors for Bayesian neural networks, University of California, Irvine, 2018

  92. [100]

    Gelman, D

    A. Gelman, D. B. Rubin, Inference from Iterative Simulation Using Mul- tiple Sequences, Statistical Science 7 (4) (Nov. 1992). doi:10.1214/ ss/1177011136

  93. [101]

    A. S. Weigend, Time series prediction: forecasting the future and under- standing the past, Routledge, 2018

  94. [102]

    URL https://www.kaggle.com/datasets/robervalt/ sunspots

    Solar Influences Data Analysis Center, Sunspots (2018). URL https://www.kaggle.com/datasets/robervalt/ sunspots

  95. [103]

    Warwick Nash, Tracy Sellers, Simon Talbot, Andrew Cawthorn, Wes Ford, Abalone, published: UCI Machine Learning Repository (1994)

  96. [104]

    Sigillito, S

    V . Sigillito, S. Wing, S. Hutton, K. Baker, Ionosphere, published: UCI Machine Learning Repository (1989). 19

  97. [105]

    R. A. Fisher, Iris (1936). doi:10.24432/C56C76. URL https://archive.ics.uci.edu/dataset/53

  98. [106]

    J. M. Webster, Y . Yokoyama, C. Cotterill, the Expedition 325 Scientists, Proceedings of the Integrated Ocean Drilling Program, V ol. 325, Inte- grated Ocean Drilling Program Management International, Inc., Tokyo, 2011. URL https://doi.org/10.2204/iodp.proc.325.2011

  99. [107]

    G. F. Camoin, Y . Iryu, Expedition 310 Scientists, Proceedings of the Integrated Ocean Drilling Program, V ol. 310, Integrated Ocean Drilling Program Management International, Inc., Washington, DC, 2007. URL https://doi.org/10.2204/iodp.proc.310.2007

  100. [108]

    C. D. Woodro ffe, J. M. Webster, Coral reefs and sea-level change, Ma- rine Geology 352 (2014) 248–267, publisher: Elsevier

  101. [109]

    J. M. Webster, J. C. Braga, M. Humblet, D. C. Potts, Y . Iryu, Y . Yokoyama, K. Fujita, R. Bourillot, T. M. Esat, S. Fallon, others, Re- sponse of the Great Barrier Reef to sea-level and environmental changes over the past 30,000 years, Nature Geoscience 11 (6) (2018) 426–432,...

  102. [110]

    K. L. Sanborn, J. M. Webster, Y . Yokoyama, A. Dutton, J. C. Braga, D. A. Clague, J. B. Paduan, D. Wagner, J. J. Rooney, J. R. Hansen, New evidence of Hawaiian coral reef drowning in response to meltwater pulse-1A, Quaternary Science Reviews 175 (2017) 60–72. URL https://doi.o...

  103. [111]

    R. Deo, J. M. Webster, T. Salles, R. Chandra, ReefCoreSeg: A Clustering-Based Framework for Multi-Source Data Fusion for Segmen- tation of Reef Drill Cores, IEEE Access 12 (2024) 12164–12180, con- ference Name: IEEE Access. doi:10.1109/ACCESS.2023.3341156. URL https://ieeexplo...

  104. [112]

    A. P. Bradley, The use of the area under the ROC curve in the evaluation of machine learning algorithms, Pattern recognition 30 (7) (1997) 1145– 1159, publisher: Elsevier

  105. [113]

    Rocha, S

    A. Rocha, S. K. Goldenstein, Multiclass from binary: Expanding one- versus-all, one-versus-one and ecoc-based approaches, IEEE transac- tions on neural networks and learning systems 25 (2) (2013) 289–302, publisher: IEEE

  106. [114]

    Huang, C

    J. Huang, C. Ling, Using AUC and accuracy in evaluating learning algo- rithms, IEEE Transactions on Knowledge and Data Engineering 17 (3) (2005) 299–310. doi:10.1109/TKDE.2005.50

  107. [115]

    T. K. Ho, Random decision forests, in: Proceedings of 3rd International Conference on Document Analysis and Recognition, V ol. 1, 1995, pp. 278–282 vol.1. doi:10.1109/ICDAR.1995.598994

  108. [116]

    T. L. Insua, L. Hamel, K. Moran, L. M. Anderson, J. M. Webster, G. F. Camoin, Advanced classification of carbonate sediments based on phys- ical properties, Sedimentology 62 (2) (2015) 590–606. URL https://doi.org/10.1111/sed.12168

  109. [117]

    Y . Iryu, Y . Takahashi, K. Fujita, G. Camoin, G. Cabioch, H. Mat- suda, T. Sato, K. Sugihara, J. M. Webster, H. Westphal, Sea level his- tory recorded in the Pleistocene carbonate sequence in IODP Hole 310- M0005D, off Tahiti, Island Arc 19 (4) (2010) 690–706. URL https://doi...

  110. [118]

    Westphal, K

    H. Westphal, K. Heindel, M. Brandano, J. Peckmann, Genesis of micro- bialites as contemporaneous framework components of deglacial coral reefs, Tahiti (IODP 310), Facies 56 (3) (2010) 337–352. URL https://doi.org/10.1007/s10347-009-0207-3

  111. [119]

    Zhang, M

    J. Zhang, M. D. Shields, The e ffect of prior probabilities on quantifi- cation and propagation of imprecise probabilities resulting from small datasets, Computer Methods in Applied Mechanics and Engineering 334 (2018) 483–506. doi:10.1016/j.cma.2018.01.045. URL https://www.sc...

  112. [120]

    Chandra, A

    R. Chandra, A. Kapoor, Bayesian neural multi-source transfer learning, Neurocomputing 378 (2020) 54–64, publisher: Elsevier

  113. [121]

    We show the changes to each model’s posterior prediction performance with 0 to 75 % pruning using each pruning scheme

    Appendix 20 Figure 10: Post pruning regression performance over 30 experimental runs on all regression datasets. We show the changes to each model’s posterior prediction performance with 0 to 75 % pruning using each pruning scheme. The bold lines in each plot represent the mea...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.