REVIEW 65 cited by
Tutorial on Variational Autoencoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In just three years, Variational Autoencoders (VAEs) have emerged as one of the most popular approaches to unsupervised learning of complicated distributions. VAEs are appealing because they are built on top of standard function approximators (neural networks), and can be trained with stochastic gradient descent. VAEs have already shown promise in generating many kinds of complicated data, including handwritten digits, faces, house numbers, CIFAR images, physical models of scenes, segmentation, and predicting the future from static images. This tutorial introduces the intuitions behind VAEs, explains the mathematics behind them, and describes some empirical behavior. No prior knowledge of variational Bayesian methods is assumed.
Forward citations
Showing 60 of 65 Pith papers that cite this
-
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
CT-ScanGaze, the first public eye-tracking dataset on CT volumes, contains 909 scans with radiologist gaze, reports, and findings, and CT-Searcher, a 3D scanpath model, beats adapted 2D baselines on it.
-
Integrating Random Effects in Variational Autoencoders for Dimensionality Reduction of Correlated Data
LMMVAE extends variational autoencoders with linear-mixed-model-style random effects, improving dimensionality reduction for correlated data.
-
Uncertainty-aware damage identification in short-span bridges via physics-informed variational autoencoder
A PI-GCVAE with a differentiable eigenvalue decoder and Gaussian-copula latents recovers true stiffness posteriors on noisy synthetic short-span bridge data at ~79% 95%-coverage.
-
One-shot Conditional Sampling: MMD meets Nearest Neighbors
Conditional distributions can be sampled in one forward pass by training a generator to minimize a nearest-neighbor estimate of expected conditional MMD, with convergence guarantees.
-
Plug-and-Play Latent Diffusion for Electromagnetic Inverse Scattering with Application to Brain Imaging
Latent-space plug-and-play diffusion posterior sampling yields state-of-the-art synthetic EM brain image reconstructions without paired measurement-label training data.
-
DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.
-
RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online
Pretraining a multitask VAE on online ASMR audio and finetuning only the activity head on vibration data improves cross-user activity recognition accuracy over baselines in a 4-participant study.
-
From Gallery to Wrist: Realistic 3D Bracelet Insertion in Videos
A hybrid pipeline that renders bracelets with 3D Gaussian Splatting and then refines each frame with a diffusion model yields realistic, temporally consistent bracelet insertion into dynamic wrist videos.
-
ProDiff: Prototype-Guided Diffusion for Minimal Information Trajectory Imputation
ProDiff reconstructs intermediate trajectory points from just two endpoints by combining prototype learning with a denoising diffusion model, outperforming existing imputation methods on two mobility datasets.
-
Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
Align-DA uses direct preference optimization to align a score-based data assimilation prior with assimilation accuracy, forecast skill, and physical adherence rewards, improving analysis quality over unaligned diffusi...
-
Variational Autoencoder Framework for Hyperspectral Retrievals (Hyper-VAE) of Phytoplankton Absorption and Chlorophyll a in Coastal Waters for NASA's EMIT and PACE Missions
A VAE trained on coastal bio-optical data retrieves phytoplankton absorption spectra and chlorophyll a from EMIT/PACE hyperspectral reflectance with more stable performance than an MDN baseline.
-
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
A training-free pipeline makes diffusion text-to-video generation run on an iPhone 15 Pro with quality close to GPU output, at the cost of slower generation.
-
Layer Separation: Adjustable Joint Space Width Images Synthesis in Conventional Radiography
A layer separation network generates adjustable joint space width synthetic finger X-rays from real radiographs, improving downstream rheumatoid arthritis analysis models.
-
FL-CLEANER: byzantine and backdoor defense by CLustering Errors of Activation maps in Non-iid fedErated leaRning
FL-CLEANER filters Byzantine and backdoor model updates in federated learning under non-IID data by scoring clients with a conditional variational autoencoder on activation-map reconstruction errors and clustering the scores.
-
On the Statistical Capacity of Deep Generative Models
Push-forwards of Gaussian or log-concave latent variables through Lipschitz neural networks are always sub-Gaussian or sub-exponential, so common deep generative models cannot generate heavy-tailed distributions.
-
Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and Convergence
PI-FT, a policy-iteration algorithm for KL-regularized diffusion fine-tuning, converges linearly to the globally optimal control under Lipschitz smoothness assumptions.
-
GAS: Generative Auto-bidding with Post-training Search
A single base bidding policy, refined by small transformer critics and a search-and-vote procedure, lifts auto-bidding performance on benchmark and live Kuaishou traffic.
-
Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis
PRDL learns a promptable Gaussian distribution over patch representations during DINO-style pretraining and uses it to augment WSI classifiers, improving AUC on lung EGFR and cancer subtyping benchmarks.
-
Scalable Modeling of Spatiotemporal Data using the Variational Autoencoder: an Application in Glaucoma
A two-stage variational autoencoder with per-latent-dimension linear regression predicts future glaucoma visual fields more accurately than a classical patient-level spatiotemporal model, especially with few baseline visits.
-
Causal Transfer in Medical Image Analysis
Causal Transfer Learning unifies structural causal models, invariant risk minimisation and counterfactuals with transfer learning to produce domain-robust medical image models.
-
Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner
Diffusion-completer training with a trajectory aligner makes diffusion-based auto-bidding work at scale, improving conversion value by 29.9% on a sparse public benchmark and by 2.0% in production at Kuaishou.
-
From Data to Decision: A Multi-Stage Framework for Class Imbalance Mitigation in Optical Network Failure Analysis
On experimental optical-network data, threshold adjustment improves failure-detection F1 by up to 15.3%, while CTGAN data augmentation improves failure-identification F1 by up to 24.2%.
-
Physical Layer Authentication Based on Hierarchical Variational Auto-Encoder for Industrial Internet of Things
A hierarchical autoencoder-plus-variational-autoencoder scheme authenticates industrial IoT transmitters from channel impulse responses, claiming higher F1 than three baselines without attacker channel priors.
-
Deciphering the Small-Angle Scattering of Polydisperse Hard Spheres using Deep Learning
A VAE trained on simulated small-angle scattering of polydisperse hard spheres generates scattering curves more accurately than Percus-Yevick theory and recovers volume fraction and polydispersity from test curves to ...
-
Bridging the Last Mile of Prediction: Enhancing Time Series Forecasting with Conditional Guided Flow Matching
CGFM uses an auxiliary model's predictions as the source for flow matching to learn forecast residuals and improve time series forecasts.
-
Black Hole Spectroscopy with Conditional Variational Autoencoder
A CVAE trained on simulated ringdown waveforms produces posterior estimates of remnant black hole parameters that match Bayesian inference, including overtones and a braneworld tidal charge parameter.
-
DiffPattern-Flex: Efficient Layout Pattern Generation via Discrete Diffusion
DiffPattern-Flex generates DRC-clean VLSI layout patterns with a discrete diffusion topology model plus a white-box legalization solver, reporting diversity 11.713 and 100% legality on the ICCAD 2014 benchmark.
-
Variational OOD State Correction for Offline Reinforcement Learning
DASP adds a variational density-aware term to offline RL policy optimization and reports improved average scores on MuJoCo and AntMaze benchmarks.
-
Exploring Representation-Aligned Latent Space for Better Generation
Aligning VAE latents with DINOv2 semantic features improves latent diffusion model image generation on ImageNet by about 15% FID.
-
InDeed: Interpretable image deep decomposition with guaranteed generalizability
InDeed decomposes images into low-rank, sparse, and noise parts using a network whose modules mirror a hierarchical Bayesian model, and proposes test-time adaptation for out-of-distribution data.
-
Towards Unraveling and Improving Generalization in World Models
Small zero-drift latent errors can regularize world models, and Jacobian regularization helps when latent errors have nonzero drift.
-
Reward driven workflows for unsupervised explainable analysis of phases and ferroic variants from atomically resolved imaging data
Reward-driven hyperparameter selection guided by domain-wall straightness and continuity steers unsupervised clustering and variational autoencoders toward physically meaningful phase and ferroic-variant maps in Sm-do...
-
Structure learning with Temporal Gaussian Mixture for model-based Reinforcement Learning
A variational Gaussian mixture with an online component-pruning and matching mechanism is combined with a Dirichlet-categorical transition model and belief-based Q-learning to solve small mazes from continuous observations.
-
DeepClean -- self-supervised artefact rejection for intensive care waveform data using deep generative learning
A convolutional variational autoencoder trained only on clean ICU arterial blood pressure data detects waveform artefacts at about 90% sensitivity and specificity and outperforms PCA reconstruction.
-
Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms
A status map of 2015-2025 learning-based motion planning in dynamic environments, organized by four roles learning can play: direct policy, classical-planner augmentation, hybrid coupling, and training support.
-
Variational meta-learning inference for low dimensional neural system identification
A variational (VAE-style) extension of manifold meta-learning that adds Laplace-approximation uncertainty bounds to low-data nonlinear system identification.
-
Deep Learning Option Pricing with Market Implied Volatility Surfaces
A VAE-compressed volatility surface plus a small neural network can approximate QuantLib prices for American puts and arithmetic Asian options in a single forward pass.
-
Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.
-
Comparing Normalizing Flows with Kernel Density Estimation in Estimating Risk of Automated Driving Systems
Normalizing flows fit scenario parameter densities better than KDE on held-out data, but produce roughly 40 times lower collision risk estimates, with no ground truth to decide which is correct.
-
Deep-Learning Investigation of Vibrational Raman Spectra for Plant-Stress Analysis
DIVA uses a variational autoencoder on first-derivative Raman spectra to cluster plant stress states and identify significant peaks without manual preprocessing.
-
Towards Foundation Auto-Encoders for Time-Series Anomaly Detection
A univariate VAE with dilated convolutions is proposed as a simple 'foundation' model for time-series anomaly detection, with preliminary zero-shot experiments on two datasets.
-
Generative Machine Learning in Adaptive Control of Dynamic Manufacturing Processes: A Review
A review proposes a four-part functional taxonomy of ML-enhanced adaptive manufacturing control and analyzes where generative models fit, identifying gaps and future research directions.
-
Unsupervised outlier detection to improve bird audio dataset labels
An unsupervised pipeline using autoencoders plus hierarchical clustering flags out-of-species sounds in Xeno-Canto bird recordings, but accuracy varies strongly across species.
-
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
A conditional variational autoencoder with inverse autoregressive flow generates diverse intonations in voice conversion by sampling a latent style vector conditioned on phoneme posteriorgrams.
-
Image Watermarking of Generative Diffusion Models
A new watermarking scheme for diffusion models trains an autoencoder to embed and recover image watermarks through the generation process, but the reported robustness is undermined by flawed evaluation and an unjustif...
-
Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training
Weight-matrix probability densities in a toy autoencoder are evolved with the Fokker-Planck equation driven by the ADAM update, and the resulting output distributions roughly match training at epoch 5.
-
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
Under SETH, the paper claims VAR models cannot be approximated faster than O(n^4) when attention entries are Theta(sqrt(log n)), but can be approximated in O(n^{2+o(1)}) when entries are o(sqrt(log n)).
-
Circuit Complexity Bounds for Visual Autoregressive Model
The authors show that a simplified formalization of the VAR image generation model lies in DLOGTIME-uniform TC0, meaning it can be simulated by constant-depth threshold circuits with polynomial size and precision.
-
Leveraging Self-Training and Variational Autoencoder for Agitation Detection in People with Dementia Using Wearable Sensors
A variational autoencoder plus self-training pipeline is reported to detect agitation in dementia patients from wristband sensor data, reaching 90.18% balanced accuracy with XGBoost.
-
Dimensionality Reduction Techniques for Global Bayesian Optimisation
A VAE-based latent-space Bayesian optimisation framework with Matérn-5/2 kernels and Sequential Domain Reduction solves more 100D benchmark problems than BO-SDR and REMBO in small numerical experiments.
-
Variational Encoder-Decoders for Learning Latent Representations of Physical Systems
A variational encoder-decoder with KL and covariance regularization reconstructs Hanford groundwater pressure from 1475 inputs through 50 latent variables, below the linear CCA dimension estimate of about 147.
-
Cluster Specific Representation Learning
A meta-algorithm that learns cluster-specific embedding functions jointly with cluster assignments improves clustering and denoising over standard representation learning baselines.
-
Deep Generative Model Driven Protein Folding Simulation
A variational autoencoder can guide adaptive molecular dynamics to fold Fs-peptide to 1.6 Å RMSD, and to approach the native state of a designed beta-beta-alpha protein at 4.4 Å.
-
End to End Autoencoder MLP Framework for Sepsis Prediction
An autoencoder-MLP sepsis predictor achieves moderate accuracy on three ICU cohorts, but the autoencoder is trained only with classification loss, and reported gains over baselines lack error bars and may use test dat...
-
GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective
A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.
-
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.
-
AI/ML-Based Automatic Modulation Recognition: Recent Trends and Future Possibilities
A controlled replication benchmark of nine AMR models on RadioML-2016A, plus experiments showing that moderate SNR training ranges and added recurrent layers can improve accuracy.
-
PXGen: A Post-hoc Explainable Method for Generative Models
PXGen is a post-hoc, training-free explanation framework that scores anchor samples with intrinsic and extrinsic criteria, groups them by thresholds, and selects representative examples via k-dispersion or k-center.
-
Iterative Encoding-Decoding VAEs Anomaly Detection in NOAA's DART Time Series: A Machine Learning Approach for Enhancing Data Integrity for NASA's GRACE-FO Verification and Validation
An iterative VAE is proposed for cleaning NOAA DART time series, but the claimed improvement over classical methods is supported only by visual inspection of a single station.
-
InferPy: Probabilistic Modeling with Deep Neural Networks Made Easy
InferPy provides a compact, high-level Python API for hierarchical probabilistic models with deep neural networks, built on TensorFlow Probability and Keras.
Discussion (0). Continue with ORCID to comment.