Pith. sign in

REVIEW 3 cited by

BayesFormer: Transformer with Uncertainty Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.00826 v1 pith:PHCRSYEA submitted 2022-06-02 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords transformeruncertaintyacquisitionactivearchitecturesbayesformerestimatesfunction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer has become ubiquitous due to its dominant performance in various NLP and image processing tasks. However, it lacks understanding of how to generate mathematically grounded uncertainty estimates for transformer architectures. Models equipped with such uncertainty estimates can typically improve predictive performance, make networks robust, avoid over-fitting and used as acquisition function in active learning. In this paper, we introduce BayesFormer, a Transformer model with dropouts designed by Bayesian theory. We proposed a new theoretical framework to extend the approximate variational inference-based dropout to Transformer-based architectures. Through extensive experiments, we validate the proposed architecture in four paradigms and show improvements across the board: language modeling and classification, long-sequence understanding, machine translation and acquisition function for active learning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

    cs.AI 2026-01 conditional novelty 7.0 of 10

    Treating LoRA adapters as sparse-GP random variables with a normalizing flow yields up to 84% ECE and 76% NLL reduction at ~1.2x training cost, reverting to standard LoRA in a point-mass limit.

  2. Sound Probabilistic Safety Bounds for Large Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Guided expansion of a few generation-tree branches yields provably valid but extremely small lower bounds on LLM harm probability; in several reported runs the baseline Monte Carlo estimate is orders of magnitude larger.

  3. Energy-Based Transformers are Scalable Learners and Thinkers

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Energy-Based Transformers learn to predict by gradient-descent minimization of a learned energy function, and the paper reports faster pretraining scaling and inference-time thinking gains over Transformer++ and Diffu...

Pith tools