REVIEW 14 cited by
Discrete Variational Autoencoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Probabilistic models with discrete latent variables naturally capture datasets composed of discrete classes. However, they are difficult to train efficiently, since backpropagation through discrete variables is generally not possible. We present a novel method to train a class of probabilistic models with discrete latent variables using the variational autoencoder framework, including backpropagation through the discrete latent variables. The associated class of probabilistic models comprises an undirected discrete component and a directed hierarchical continuous component. The discrete component captures the distribution over the disconnected smooth manifolds induced by the continuous component. As a result, this class of models efficiently learns both the class of objects in an image, and their specific realization in pixels, from unsupervised data, and outperforms state-of-the-art methods on the permutation-invariant MNIST, Omniglot, and Caltech-101 Silhouettes datasets.
Forward citations
Cited by 14 Pith papers
-
GradInf: Gradient Estimation as Probabilistic Inference
Gradient estimation of probabilistic programs reduces soundly to probabilistic inference after programmable coupling and factorization, enabling new low-variance estimators that beat baselines.
-
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
HPO enables unbiased policy optimization in hybrid action spaces by mixing differentiable simulation gradients with score-function estimates, outperforming PPO as continuous dimensions increase.
-
Multi-Mode Quantum Annealing for Generative Representation Learning with Boltzmann Priors
A multi-mode quantum annealing approach enables VAEs with Boltzmann priors, showing faster training and better generation than Gaussian-prior VAEs on MNIST, Fashion-MNIST, and CelebA plus improved out-of-distribution ...
-
DSA: Dynamic Step Allocation for Fast Autoregressive Video Generation
DSA adds a jointly trained confidence head to autoregressive video diffusion models that dynamically allocates fewer or more denoising steps per frame, achieving 22.63 FPS real-time generation on H100 while matching V...
-
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.
-
Advancing ALS Applications with Large-Scale Pre-training: Dataset Development and Downstream Assessment
A new large-scale ALS point cloud pre-training dataset, sampled by land cover and slope diversity, improves downstream task performance when used to pre-train BEV-MAE.
-
Sample- and Parameter-Efficient Auto-Regressive Image Models
XTRA shows that block-wise causal autoregressive pre-training improves sample and parameter efficiency over patch-level autoregressive image models such as AIM.
-
PixelVAE++: Improved PixelVAE with Discrete Prior
PixelVAE++ replaces Gaussian latents with discrete RBM-prior latents in a PixelCNN++ decoder, achieving small log-likelihood gains on MNIST, Omniglot, and CIFAR-10, but with no code release and weak evidence that the ...
-
Molecular Design beyond Training Data with Novel Extended Objective Functionals of Generative AI Models Driven by Quantum Annealing Computer
Quantum annealing combined with a Neural Hash Function lets generative models create molecules that are more drug-like than classical versions or the training set itself.
-
Variational Sparse Paired Autoencoders (vsPAIR) for Inverse Problems and Uncertainty Quantification
vsPAIR couples a Gaussian VAE over observations with a spike-and-slab sparse VAE over the quantity of interest via a learned latent mapping, yielding fast inverse reconstructions whose active latent dimensions can be ...
-
Symmetry Understanding of 3D Shapes via Chirality Disentanglement
Abstract-level claim: decorate 3D shape vertices with chirality features drawn from 2D foundation models via Diff3F, enabling left-right disentanglement; the supplied full text is a different paper, so the claim is un...
-
TractoEmbed: Modular Multi-level Embedding framework for white matter tract segmentation
TractoEmbed fuses streamline, cluster, and patch embeddings to improve white matter tract segmentation, reaching 93.04% accuracy with hyperlocal point clouds.
-
A Survey on Deep Learning Architectures for Point Cloud Classification and Segmentation
A systematic literature survey that categorizes deep learning architectures for point cloud classification, part segmentation, and semantic segmentation, evaluates them on benchmarks, and discusses innovations, limita...
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.