Pith. sign in

REVIEW 1 cited by

URSABench: Comprehensive Benchmarking of Approximate Bayesian Inference Methods for Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.04466 v1 pith:NUKXMQFH submitted 2020-07-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords methodsapproximatebayesiandeepinferencecomprehensiverobustnessscalability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While deep learning methods continue to improve in predictive accuracy on a wide range of application domains, significant issues remain with other aspects of their performance including their ability to quantify uncertainty and their robustness. Recent advances in approximate Bayesian inference hold significant promise for addressing these concerns, but the computational scalability of these methods can be problematic when applied to large-scale models. In this paper, we describe initial work on the development ofURSABench(the Uncertainty, Robustness, Scalability, and Accu-racy Benchmark), an open-source suite of bench-marking tools for comprehensive assessment of approximate Bayesian inference methods with a focus on deep learning-based classification tasks

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neurosymbolic Artificial Intelligence for Robust Network Intrusion Detection: From Scratch to Transfer Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Reusing a frozen pretrained autoencoder, retrained clustering, and fine-tuned XGBoost on ACI-IoT-2023 outperforms FcNN and 1D-CNN with about half the training data, and metamodel-based UQ outperforms score-based UQ on...

Pith tools