REVIEW 4 cited by
Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The posteriors over neural network weights are high dimensional and multimodal. Each mode typically characterizes a meaningfully different representation of the data. We develop Cyclical Stochastic Gradient MCMC (SG-MCMC) to automatically explore such distributions. In particular, we propose a cyclical stepsize schedule, where larger steps discover new modes, and smaller steps characterize each mode. We also prove non-asymptotic convergence of our proposed algorithm. Moreover, we provide extensive experimental results, including ImageNet, to demonstrate the scalability and effectiveness of cyclical SG-MCMC in learning complex multimodal distributions, especially for fully Bayesian inference with modern deep neural networks.
Forward citations
Cited by 4 Pith papers
-
Stochastic Weight Sharing for Bayesian Neural Networks
2DGBNN compresses Bayesian neural networks by clustering weight means and variances into shared 2D Gaussians, reducing parameter counts by up to 99% on ImageNet-scale models with small accuracy losses.
-
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
A training-dynamics abstention method matches deep ensembles at a fraction of the training cost, and a five-term error budget explains why selective classifiers still fall short of the oracle.
-
ConBatch-BAL: Batch Bayesian Active Learning under Budget Constraints
Two budget-aware batch active learning heuristics, greedy and dynamic thresholding, reduce the number of labeling rounds needed to reach target accuracy on new building image datasets compared with random selection.
-
A Novel Active Learning Approach to Label One Million Unknown Malware Variants
Proposes ViT-BNN for malware active learning with a tuned Gaussian scaling parameter, but results are not reproducible and the central claims are unsupported.
Discussion (0). Continue with ORCID to comment.