Pith. sign in

REVIEW 1 cited by

Bias-Variance Tradeoffs in Single-Sample Binary Gradient Estimators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.03549 v2 pith:LQJ7KMTA submitted 2021-10-07 cs.LG cs.NE

classification cs.LGcs.NE
keywords binaryestimatorestimatorsgradientlearningmethodsmodelsnetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Discrete and especially binary random variables occur in many machine learning models, notably in variational autoencoders with binary latent states and in stochastic binary networks. When learning such models, a key tool is an estimator of the gradient of the expected loss with respect to the probabilities of binary variables. The straight-through (ST) estimator gained popularity due to its simplicity and efficiency, in particular in deep networks where unbiased estimators are impractical. Several techniques were proposed to improve over ST while keeping the same low computational complexity: Gumbel-Softmax, ST-Gumbel-Softmax, BayesBiNN, FouST. We conduct a theoretical analysis of bias and variance of these methods in order to understand tradeoffs and verify the originally claimed properties. The presented theoretical results allow for better understanding of these methods and in some cases reveal serious issues.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Principled Bayesian Framework for Training Binary and Spiking Neural Networks

    cs.LG 2025-05 conditional novelty 5.0 of 10

    A variational Bayesian framework with importance-weighted straight-through estimators trains binary and spiking networks without normalization layers, matching surrogate-gradient baselines on CIFAR-10, DVS Gesture, and SHD.

Pith tools