Pith. sign in

REVIEW 17 cited by

Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.05662 v4 pith:K62EWTFO submitted 2019-03-13 cs.LG math.OCstat.ML

Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets

classification cs.LG math.OCstat.ML
keywords gradienttraininglossactivationchaincoarsepopulationrule
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Training activation quantized neural networks involves minimizing a piecewise constant function whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengio et al., 2013) in the backward pass only, so that the "gradient" through the modified chain rule becomes non-trivial. Since this unusual "gradient" is certainly not the gradient of loss function, the following question arises: why searching in its negative direction minimizes the training loss? In this paper, we provide the theoretical justification of the concept of STE by answering this question. We consider the problem of learning a two-linear-layer network with binarized ReLU activation and Gaussian input data. We shall refer to the unusual "gradient" given by the STE-modifed chain rule as coarse gradient. The choice of STE is not unique. We prove that if the STE is properly chosen, the expected coarse gradient correlates positively with the population gradient (not available for the training), and its negation is a descent direction for minimizing the population loss. We further show the associated coarse gradient descent algorithm converges to a critical point of the population loss minimization problem. Moreover, we show that a poor choice of STE leads to instability of the training algorithm near certain local minima, which is verified with CIFAR-10 experiments.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Network Quantization by Learning Low-Loss Subspaces

    cs.CV 2026-06 unverdicted novelty 7.0

    Learning quantization-aware linear paths in weight space yields a midpoint whose direct quantization matches quantization-aware training performance without using straight-through estimators.

  2. Training single-electron and single-photon stochastic physical neural networks

    quant-ph 2026-04 unverdicted novelty 7.0

    Single-electron and single-photon stochastic physical neural networks achieve over 97% MNIST test accuracy when trained with empirical outputs in the backward pass using few trials per layer.

  3. MusicMark: A Robust Generative Watermarking Framework for Music Generation

    cs.SD 2026-07 conditional novelty 6.5

    Embedding watermark bits into diffusion semantic latents via a frozen-backbone adapter yields far more robust music provenance than post-hoc watermarking under codecs and cover-song attacks.

  4. Energy Efficiency Maximization for Hybrid RIS-Aided Communications via Deep Unfolding

    eess.SP 2026-07 conditional novelty 6.0

    Deep-unfolded alternating optimization for hybrid RIS mode selection and binary phases yields ~30% higher energy efficiency than plain projected gradient and most of the gain with only ~10% of elements active.

  5. Neural-Network-Assisted Binary Template Construction for Matrix-Based Pattern Matching in the STCF MDC

    hep-ex 2026-07 conditional novelty 6.0

    Offline multi-objective neural optimization yields compact binary trigger–recovery template pairs that retain ~98% of true MDC hits at 95% detector efficiency under multi-background conditions without changing the onl...

  6. LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization

    cs.CL 2026-06 unverdicted novelty 6.0

    LC-QAT achieves data-efficient 2-bit weight-only QAT for LLMs by representing quantized weights as a learned affine transform over discrete vectors, supporting end-to-end optimization from a high-quality PTQ start.

  7. Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations

    cs.AR 2026-05 unverdicted novelty 6.0

    BMRUs enable a direct one-to-one mapping from learned parameters to current-mode analog circuit elements, with discrete hysteretic outputs suppressing noise by at least 20x and supporting sub-microwatt RNN inference i...

  8. LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation

    cs.LG 2026-04 unverdicted novelty 6.0

    LBLLM achieves better accuracy than prior binarization methods for LLMs by decoupling weight and activation quantization through initialization, layer-wise distillation, and learnable activation scaling.

  9. LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization

    cs.CL 2026-06 unverdicted novelty 5.0

    LC-QAT is a 2-bit weight-only vector quantization aware training framework for LLMs that uses linear-constrained affine mappings to achieve data-efficient optimization and outperform prior QAT methods.

  10. Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications

    cs.CV 2026-05 unverdicted novelty 5.0

    A task-aware flow-based generative framework optimizes subsampling masks in compressed sensing, reporting SOTA PSNR of 25.17 dB at 5% rate on CelebA and 29.24 dB for 8x MRI on fastMRI.

  11. Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations

    cs.AR 2026-05 unverdicted novelty 5.0

    BMRUs enable analog recurrent neural network hardware via discrete outputs that suppress noise 20-fold, with one-to-one parameter-to-circuit mapping and linear power scaling for recurrence.

  12. A Composite Activation Function for Learning Stable Binary Representations

    cs.LG 2026-05 unverdicted novelty 5.0

    HTAF is a sigmoid-tanh composite that approximates the Heaviside function to allow stable gradient training of binary activation networks, yielding ICBMs with stable discretization and competitive performance on image tasks.

  13. A Controlled Diagnostic Study of Hardware-Induced Distortions in Hardware-Aware Training

    cs.LG 2026-05 unverdicted novelty 5.0

    Hardware-aware training compensates some distortions like read noise but fails on others like stuck-at faults and IR-drop, separated by three gradient diagnostics.

  14. Improving Generalization by Permutation Routing Across Model Copies

    cs.LG 2026-05 unverdicted novelty 5.0

    Replicating models and routing their local losses via permutations from a mixing kernel Q enables structured message sharing that improves generalization.

  15. JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency

    cs.CL 2026-04 unverdicted novelty 5.0

    JoyAI-LLM Flash delivers a 48B MoE LLM with 2.7B active parameters per token via FiberPO RL and dense multi-token prediction, released with checkpoints on Hugging Face.

  16. Molecular Design beyond Training Data with Novel Extended Objective Functionals of Generative AI Models Driven by Quantum Annealing Computer

    q-bio.QM 2026-02 unverdicted novelty 5.0

    Quantum annealing combined with a Neural Hash Function lets generative models create molecules that are more drug-like than classical versions or the training set itself.

  17. $\text{Log}_\text{b}$Quant: Quantizing Language Models in Logarithmic Space

    cs.CL 2026-07 unverdicted novelty 4.0

    Log_b Quant is an adjustable-base logarithmic quantization technique that outperforms tensor-wise asymmetric linear quantization at 4-bit precision on language model benchmarks while providing memory savings.