Pith. sign in

super hub Tool reference

Categorical Reparameterization with Gumbel-Softmax

Tool reference. 75% of classified Pith citations use this work as a method, library, or software dependency, not as a substantive claim.

113 Pith papers citing it
3,241 external citations · Pith
Method reference 75% of classified citations
abstract

Categorical variables are a natural choice for representing discrete structure in the world. However, stochastic neural networks rarely use categorical latent variables due to the inability to backpropagate through samples. In this work, we present an efficient gradient estimator that replaces the non-differentiable sample from a categorical distribution with a differentiable sample from a novel Gumbel-Softmax distribution. This distribution has the essential property that it can be smoothly annealed into a categorical distribution. We show that our Gumbel-Softmax estimator outperforms state-of-the-art gradient estimators on structured output prediction and unsupervised generative modeling tasks with categorical latent variables, and enables large speedups on semi-supervised classification.

hub tools

citation-role summary

method 6 background 2

citation-polarity summary

claims ledger

  • abstract Categorical variables are a natural choice for representing discrete structure in the world. However, stochastic neural networks rarely use categorical latent variables due to the inability to backpropagate through samples. In this work, we present an efficient gradient estimator that replaces the non-differentiable sample from a categorical distribution with a differentiable sample from a novel Gumbel-Softmax distribution. This distribution has the essential property that it can be smoothly annealed into a categorical distribution. We show that our Gumbel-Softmax estimator outperforms state-o

authors

co-cited works

representative citing papers

Selective Disclosure Watermarking for Large Language Models

cs.CR · 2026-07-06 · accept · novelty 7.0

HeRo recursively partitions the LLM vocabulary into a hierarchy, embedding multi-bit payloads across layers so that verifiers with different keys recover only their authorized portion while preserving the original sampling distribution.

Adversarial Trust Poisoning in Vehicular Collaborative Perception

cs.CR · 2026-05-21 · unverdicted · novelty 7.0

TrustFlip weaponizes consistency-based trust defenses in vehicular collaborative perception by using physical adversarial objects to induce inconsistencies that are misattributed to benign vehicles, leading to their exclusion and reduced system performance.

Dynamic Chunking for Diffusion Language Models

cs.CL · 2026-05-15 · unverdicted · novelty 7.0

DCDM replaces positional blocks with learnable semantic chunks via differentiable Chunking Attention, yielding consistent gains over block and unstructured diffusion baselines up to 1.5B parameters.

Test-time Sparsity for Extreme Fast Action Diffusion

cs.CV · 2026-05-13 · unverdicted · novelty 7.0

Test-time sparsity with a parallel pipeline and omnidirectional feature reuse accelerates action diffusion by 5x to 47.5 Hz while cutting FLOPs 92% with no performance loss.

Marginal multi-object multi-frame blind deconvolution

astro-ph.IM · 2026-05-12 · unverdicted · novelty 7.0

A marginal estimator for blind deconvolution in solar imaging that integrates out object uncertainty to improve regularization and allow plug-and-play hyperparameter optimization.

Approximation-Free Differentiable Oblique Decision Trees

cs.LG · 2026-05-08 · unverdicted · novelty 7.0

DTSemNet gives an exact, invertible neural-network encoding of hard oblique decision trees that supports direct gradient training for both classification and regression without probabilistic softening or quantized estimators.

PhySPRING: Structure-Preserving Reduction of Physics-Informed Twins via GNN

cs.RO · 2026-05-08 · unverdicted · novelty 7.0

PhySPRING uses differentiable GNNs to learn hierarchical coarsened spring-mass topologies and parameters from observations, delivering up to 2.3x speedup on PhysTwin benchmarks and comparable robot policy success rates in zero-shot Real2Sim substitution.

Learning to Theorize the World from Observation

cs.LG · 2026-05-05 · unverdicted · novelty 7.0

NEO is a probabilistic neural model that induces compositional programs as a learned Language of Thought from non-textual observations and executes them via a shared transition model to enable explanation-driven generalization.

citing papers explorer

Showing 50 of 113 citing papers.