REVIEW 7 cited by
Bayesian Recurrent Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this work we explore a straightforward variational Bayes scheme for Recurrent Neural Networks. Firstly, we show that a simple adaptation of truncated backpropagation through time can yield good quality uncertainty estimates and superior regularisation at only a small extra computational cost during training, also reducing the amount of parameters by 80\%. Secondly, we demonstrate how a novel kind of posterior approximation yields further improvements to the performance of Bayesian RNNs. We incorporate local gradient information into the approximate posterior to sharpen it around the current batch statistics. We show how this technique is not exclusive to recurrent neural networks and can be applied more widely to train Bayesian neural networks. We also empirically demonstrate how Bayesian RNNs are superior to traditional RNNs on a language modelling benchmark and an image captioning task, as well as showing how each of these methods improve our model over a variety of other schemes for training them. We also introduce a new benchmark for studying uncertainty for language models so future methods can be easily compared.
Forward citations
Cited by 7 Pith papers
-
Modeling continuous-time stochastic processes using $\mathcal{N}$-Curve mixtures
A mixture of Bezier curves with Gaussian control points, trained with a mixture density network, generates smooth multi-modal sequences in one inference step and outperforms comparable baselines on two tasks.
-
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
GRAD fuses spatial, wavelet, and Fourier visual features with kinematic robot data through graph attention and adversarial alignment, and reports top accuracy plus improved corruption tolerance on two surgical gesture...
-
Error-quantified Conformal Inference for Time Series
ECI adds a smoothed error-quantification term to the online conformal update rule and proves distribution-free long-run coverage bounds, yielding tighter prediction sets on real time-series benchmarks.
-
U-CAM: Visual Explanation using Uncertainty based Class Activation Maps
U-CAM uses gradients of aleatoric and predictive uncertainty losses to sharpen visual attention maps and improve VQA accuracy over standard baselines.
-
Bayesian Learning in Structural Dynamics: A Comprehensive Review and Emerging Trends
A comprehensive review that organizes Bayesian inference in structural dynamics into physical model learning and data-centric statistical model learning, with applications and open challenges.
-
Alternators With Noise Models
Alternator++ adds trainable noise-prediction networks and a noise-matching loss to Alternators, but the proposed training target is ill-defined and the reported improvements are mixed.
-
A Statistical Framework for Model Selection in LSTM Networks
A statistical model selection framework for LSTMs is proposed, but it largely recombines existing methods and provides no convincing evidence of improvement.
Discussion (0). Continue with ORCID to comment.