REVIEW 2 cited by
Rao-Blackwellizing the Straight-Through Gumbel-Softmax Gradient Estimator
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Gradient estimation in models with discrete latent variables is a challenging problem, because the simplest unbiased estimators tend to have high variance. To counteract this, modern estimators either introduce bias, rely on multiple function evaluations, or use learned, input-dependent baselines. Thus, there is a need for estimators that require minimal tuning, are computationally cheap, and have low mean squared error. In this paper, we show that the variance of the straight-through variant of the popular Gumbel-Softmax estimator can be reduced through Rao-Blackwellization without increasing the number of function evaluations. This provably reduces the mean squared error. We empirically demonstrate that this leads to variance reduction, faster convergence, and generally improved performance in two unsupervised latent variable models.
Forward citations
Cited by 2 Pith papers
-
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
For a two-layer binary network with Gaussian inputs, O(n^2) samples guarantee ergodic convergence of STE training and O(n^4) guarantee that iterates revisit the optimal weights, even under label noise.
-
A Principled Bayesian Framework for Training Binary and Spiking Neural Networks
A variational Bayesian framework with importance-weighted straight-through estimators trains binary and spiking networks without normalization layers, matching surrogate-gradient baselines on CIFAR-10, DVS Gesture, and SHD.
Discussion (0). Sign in to comment.