Pith. sign in

REVIEW 65 cited by

Tutorial on Variational Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1606.05908 v3 pith:IHIWY2US submitted 2016-06-19 stat.ML cs.LG

classification stat.MLcs.LG
keywords vaesvariationalautoencodersbehindcomplicatedimagestutorialalready
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In just three years, Variational Autoencoders (VAEs) have emerged as one of the most popular approaches to unsupervised learning of complicated distributions. VAEs are appealing because they are built on top of standard function approximators (neural networks), and can be trained with stochastic gradient descent. VAEs have already shown promise in generating many kinds of complicated data, including handwritten digits, faces, house numbers, CIFAR images, physical models of scenes, segmentation, and predicting the future from static images. This tutorial introduces the intuitions behind VAEs, explains the mathematics behind them, and describes some empirical behavior. No prior knowledge of variational Bayesian methods is assumed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 65 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,379 citations worldwide. See all 65 Pith citations

  1. CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling

    cs.CV 2025-07 conditional novelty 7.0 of 10

    CT-ScanGaze, the first public eye-tracking dataset on CT volumes, contains 909 scans with radiologist gaze, reports, and findings, and CT-Searcher, a 3D scanpath model, beats adapted 2D baselines on it.

  2. Integrating Random Effects in Variational Autoencoders for Dimensionality Reduction of Correlated Data

    stat.ML 2024-12 conditional novelty 7.0 of 10

    LMMVAE extends variational autoencoders with linear-mixed-model-style random effects, improving dimensionality reduction for correlated data.

  3. Uncertainty-aware damage identification in short-span bridges via physics-informed variational autoencoder

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A PI-GCVAE with a differentiable eigenvalue decoder and Gaussian-copula latents recovers true stiffness posteriors on noisy synthetic short-span bridge data at ~79% 95%-coverage.

  4. One-shot Conditional Sampling: MMD meets Nearest Neighbors

    stat.ML 2025-09 conditional novelty 6.0 of 10

    Conditional distributions can be sampled in one forward pass by training a generator to minimize a nearest-neighbor estimate of expected conditional MMD, with convergence guarantees.

  5. Plug-and-Play Latent Diffusion for Electromagnetic Inverse Scattering with Application to Brain Imaging

    eess.SP 2025-09 conditional novelty 6.0 of 10

    Latent-space plug-and-play diffusion posterior sampling yields state-of-the-art synthetic EM brain image reconstructions without paired measurement-label training data.

  6. DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.

  7. RARR : Robust Real-World Activity Recognition with Vibration by Scavenging Near-Surface Audio Online

    cs.SD 2025-08 conditional novelty 6.0 of 10

    Pretraining a multitask VAE on online ASMR audio and finetuning only the activity head on vibration data improves cross-user activity recognition accuracy over baselines in a 4-participant study.

  8. From Gallery to Wrist: Realistic 3D Bracelet Insertion in Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A hybrid pipeline that renders bracelets with 3D Gaussian Splatting and then refines each frame with a diffusion model yields realistic, temporally consistent bracelet insertion into dynamic wrist videos.

  9. ProDiff: Prototype-Guided Diffusion for Minimal Information Trajectory Imputation

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ProDiff reconstructs intermediate trajectory points from just two endpoints by combining prototype learning with a denoising diffusion model, outperforming existing imputation methods on two mobility datasets.

  10. Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences

    physics.ao-ph 2025-05 conditional novelty 6.0 of 10

    Align-DA uses direct preference optimization to align a score-based data assimilation prior with assimilation accuracy, forecast skill, and physical adherence rewards, improving analysis quality over unaligned diffusi...

  11. Variational Autoencoder Framework for Hyperspectral Retrievals (Hyper-VAE) of Phytoplankton Absorption and Chlorophyll a in Coastal Waters for NASA's EMIT and PACE Missions

    cs.LG 2025-04 conditional novelty 6.0 of 10

    A VAE trained on coastal bio-optical data retrieves phytoplankton absorption spectra and chlorophyll a from EMIT/PACE hyperspectral reflectance with more stable performance than an MDN baseline.

  12. On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A training-free pipeline makes diffusion text-to-video generation run on an iPhone 15 Pro with quality close to GPU output, at the cost of slower generation.

  13. Layer Separation: Adjustable Joint Space Width Images Synthesis in Conventional Radiography

    eess.IV 2025-02 conditional novelty 6.0 of 10

    A layer separation network generates adjustable joint space width synthetic finger X-rays from real radiographs, improving downstream rheumatoid arthritis analysis models.

  14. FL-CLEANER: byzantine and backdoor defense by CLustering Errors of Activation maps in Non-iid fedErated leaRning

    cs.CR 2025-01 conditional novelty 6.0 of 10

    FL-CLEANER filters Byzantine and backdoor model updates in federated learning under non-IID data by scoring clients with a conditional variational autoencoder on activation-map reconstruction errors and clustering the scores.

  15. On the Statistical Capacity of Deep Generative Models

    stat.ML 2025-01 conditional novelty 6.0 of 10

    Push-forwards of Gaussian or log-concave latent variables through Lipschitz neural networks are always sub-Gaussian or sub-exponential, so common deep generative models cannot generate heavy-tailed distributions.

  16. Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and Convergence

    cs.LG 2024-12 conditional novelty 6.0 of 10

    PI-FT, a policy-iteration algorithm for KL-regularized diffusion fine-tuning, converges linearly to the globally optimal control under Lipschitz smoothness assumptions.

  17. GAS: Generative Auto-bidding with Post-training Search

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A single base bidding policy, refined by small transformer critics and a search-and-vote procedure, lifts auto-bidding performance on benchmark and live Kuaishou traffic.

  18. Promptable Representation Distribution Learning and Data Augmentation for Gigapixel Histopathology WSI Analysis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PRDL learns a promptable Gaussian distribution over patch representations during DINO-style pretraining and uses it to augment WSI classifiers, improving AUC on lung EGFR and cancer subtyping benchmarks.

  19. Scalable Modeling of Spatiotemporal Data using the Variational Autoencoder: an Application in Glaucoma

    stat.AP 2019-08 conditional novelty 6.0 of 10

    A two-stage variational autoencoder with per-latent-dimension linear regression predicts future glaucoma visual fields more accurately than a classical patient-level spatiotemporal model, especially with few baseline visits.

  20. Causal Transfer in Medical Image Analysis

    cs.CV 2026-03 accept novelty 5.0 of 10

    Causal Transfer Learning unifies structural causal models, invariant risk minimisation and counterfactuals with transfer learning to produce domain-robust medical image models.

  21. Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner

    cs.GT 2025-09 conditional novelty 5.0 of 10

    Diffusion-completer training with a trajectory aligner makes diffusion-based auto-bidding work at scale, improving conversion value by 29.9% on a sparse public benchmark and by 2.0% in production at Kuaishou.

  22. From Data to Decision: A Multi-Stage Framework for Class Imbalance Mitigation in Optical Network Failure Analysis

    cs.LG 2025-08 conditional novelty 5.0 of 10

    On experimental optical-network data, threshold adjustment improves failure-detection F1 by up to 15.3%, while CTGAN data augmentation improves failure-identification F1 by up to 24.2%.

  23. Physical Layer Authentication Based on Hierarchical Variational Auto-Encoder for Industrial Internet of Things

    eess.SP 2025-08 conditional novelty 5.0 of 10

    A hierarchical autoencoder-plus-variational-autoencoder scheme authenticates industrial IoT transmitters from channel impulse responses, claiming higher F1 than three baselines without attacker channel priors.

  24. Deciphering the Small-Angle Scattering of Polydisperse Hard Spheres using Deep Learning

    cond-mat.soft 2025-07 conditional novelty 5.0 of 10

    A VAE trained on simulated small-angle scattering of polydisperse hard spheres generates scattering curves more accurately than Percus-Yevick theory and recovers volume fraction and polydispersity from test curves to ...

  25. Bridging the Last Mile of Prediction: Enhancing Time Series Forecasting with Conditional Guided Flow Matching

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CGFM uses an auxiliary model's predictions as the source for flow matching to learn forecast residuals and improve time series forecasts.

  26. Black Hole Spectroscopy with Conditional Variational Autoencoder

    gr-qc 2025-06 conditional novelty 5.0 of 10

    A CVAE trained on simulated ringdown waveforms produces posterior estimates of remnant black hole parameters that match Bayesian inference, including overtones and a braneworld tidal charge parameter.

  27. DiffPattern-Flex: Efficient Layout Pattern Generation via Discrete Diffusion

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DiffPattern-Flex generates DRC-clean VLSI layout patterns with a discrete diffusion topology model plus a white-box legalization solver, reporting diversity 11.713 and 100% legality on the ICCAD 2014 benchmark.

  28. Variational OOD State Correction for Offline Reinforcement Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DASP adds a variational density-aware term to offline RL policy optimization and reports improved average scores on MuJoCo and AntMaze benchmarks.

  29. Exploring Representation-Aligned Latent Space for Better Generation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Aligning VAE latents with DINOv2 semantic features improves latent diffusion model image generation on ImageNet by about 15% FID.

  30. InDeed: Interpretable image deep decomposition with guaranteed generalizability

    cs.CV 2025-01 reject novelty 5.0 of 10

    InDeed decomposes images into low-rank, sparse, and noise parts using a network whose modules mirror a hierarchical Bayesian model, and proposes test-time adaptation for out-of-distribution data.

  31. Towards Unraveling and Improving Generalization in World Models

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Small zero-drift latent errors can regularize world models, and Jacobian regularization helps when latent errors have nonzero drift.

  32. Reward driven workflows for unsupervised explainable analysis of phases and ferroic variants from atomically resolved imaging data

    cond-mat.mtrl-sci 2024-11 conditional novelty 5.0 of 10

    Reward-driven hyperparameter selection guided by domain-wall straightness and continuity steers unsupervised clustering and variational autoencoders toward physically meaningful phase and ferroic-variant maps in Sm-do...

  33. Structure learning with Temporal Gaussian Mixture for model-based Reinforcement Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    A variational Gaussian mixture with an online component-pruning and matching mechanism is combined with a Dirichlet-categorical transition model and belief-based Q-learning to solve small mazes from continuous observations.

  34. DeepClean -- self-supervised artefact rejection for intensive care waveform data using deep generative learning

    stat.ML 2019-08 conditional novelty 5.0 of 10

    A convolutional variational autoencoder trained only on clean ICU arterial blood pressure data detects waveform artefacts at about 90% sensitivity and specificity and outperforms PCA reconstruction.

  35. Learning-Based Motion Planning for Dynamic Environments: From Foundational Algorithms to Emerging Paradigms

    cs.RO 2026-08 conditional novelty 4.0 of 10

    A status map of 2015-2025 learning-based motion planning in dynamic environments, organized by four roles learning can play: direct policy, classical-planner augmentation, hybrid coupling, and training support.

  36. Variational meta-learning inference for low dimensional neural system identification

    cs.LG 2026-07 conditional novelty 4.0 of 10

    A variational (VAE-style) extension of manifold meta-learning that adds Laplace-approximation uncertainty bounds to low-data nonlinear system identification.

  37. Deep Learning Option Pricing with Market Implied Volatility Surfaces

    q-fin.CP 2025-09 conditional novelty 4.0 of 10

    A VAE-compressed volatility surface plus a small neural network can approximate QuantLib prices for American puts and arithmetic Asian options in a single forward pass.

  38. Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.

  39. Comparing Normalizing Flows with Kernel Density Estimation in Estimating Risk of Automated Driving Systems

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Normalizing flows fit scenario parameter densities better than KDE on held-out data, but produce roughly 40 times lower collision risk estimates, with no ground truth to decide which is correct.

  40. Deep-Learning Investigation of Vibrational Raman Spectra for Plant-Stress Analysis

    cs.LG 2025-07 reject novelty 4.0 of 10

    DIVA uses a variational autoencoder on first-derivative Raman spectra to cluster plant stress states and identify significant peaks without manual preprocessing.

  41. Towards Foundation Auto-Encoders for Time-Series Anomaly Detection

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A univariate VAE with dilated convolutions is proposed as a simple 'foundation' model for time-series anomaly detection, with preliminary zero-shot experiments on two datasets.

  42. Generative Machine Learning in Adaptive Control of Dynamic Manufacturing Processes: A Review

    cs.LG 2025-04 conditional novelty 4.0 of 10

    A review proposes a four-part functional taxonomy of ML-enhanced adaptive manufacturing control and analyzes where generative models fit, identifying gaps and future research directions.

  43. Unsupervised outlier detection to improve bird audio dataset labels

    cs.LG 2025-04 conditional novelty 4.0 of 10

    An unsupervised pipeline using autoencoders plus hierarchical clustering flags out-of-species sounds in Xeno-Canto bird recordings, but accuracy varies strongly across species.

  44. Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder

    cs.SD 2025-04 conditional novelty 4.0 of 10

    A conditional variational autoencoder with inverse autoregressive flow generates diverse intonations in voice conversion by sampling a latent style vector conditioned on phoneme posteriorgrams.

  45. Image Watermarking of Generative Diffusion Models

    eess.IV 2025-02 reject novelty 4.0 of 10

    A new watermarking scheme for diffusion models trains an autoencoder to embed and recover image watermarks through the generation process, but the reported robustness is undermined by flawed evaluation and an unjustif...

  46. Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training

    cs.LG 2025-01 conditional novelty 4.0 of 10

    Weight-matrix probability densities in a toy autoencoder are evolved with the Fokker-Planck equation driven by the ADAM update, and the resulting output distributions roughly match training at epoch 5.

  47. On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis

    cs.LG 2025-01 reject novelty 4.0 of 10

    Under SETH, the paper claims VAR models cannot be approximated faster than O(n^4) when attention entries are Theta(sqrt(log n)), but can be approximated in O(n^{2+o(1)}) when entries are o(sqrt(log n)).

  48. Circuit Complexity Bounds for Visual Autoregressive Model

    stat.ML 2025-01 reject novelty 4.0 of 10

    The authors show that a simplified formalization of the VAR image generation model lies in DLOGTIME-uniform TC0, meaning it can be simulated by constant-depth threshold circuits with polynomial size and precision.

  49. Leveraging Self-Training and Variational Autoencoder for Agitation Detection in People with Dementia Using Wearable Sensors

    cs.AI 2024-12 reject novelty 4.0 of 10

    A variational autoencoder plus self-training pipeline is reported to detect agitation in dementia patients from wristband sensor data, reaching 90.18% balanced accuracy with XGBoost.

  50. Dimensionality Reduction Techniques for Global Bayesian Optimisation

    math.OC 2024-12 conditional novelty 4.0 of 10

    A VAE-based latent-space Bayesian optimisation framework with Matérn-5/2 kernels and Sequential Domain Reduction solves more 100D benchmark problems than BO-SDR and REMBO in small numerical experiments.

  51. Variational Encoder-Decoders for Learning Latent Representations of Physical Systems

    cs.LG 2024-12 reject novelty 4.0 of 10

    A variational encoder-decoder with KL and covariance regularization reconstructs Hanford groundwater pressure from 1475 inputs through 50 latent variables, below the linear CCA dimension estimate of about 147.

  52. Cluster Specific Representation Learning

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A meta-algorithm that learns cluster-specific embedding functions jointly with cluster assignments improves clustering and denoising over standard representation learning baselines.

  53. Deep Generative Model Driven Protein Folding Simulation

    q-bio.BM 2019-08 conditional novelty 4.0 of 10

    A variational autoencoder can guide adaptive molecular dynamics to fold Fs-peptide to 1.6 Å RMSD, and to approach the native state of a designed beta-beta-alpha protein at 4.4 Å.

  54. End to End Autoencoder MLP Framework for Sepsis Prediction

    cs.LG 2025-08 reject novelty 3.0 of 10

    An autoencoder-MLP sepsis predictor achieves moderate accuracy on three ICU cohorts, but the autoencoder is trained only with classification loss, and reported gains over baselines lack error bars and may use test dat...

  55. GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective

    cs.AI 2025-07 unverdicted novelty 3.0 of 10

    A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.

  56. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

  57. AI/ML-Based Automatic Modulation Recognition: Recent Trends and Future Possibilities

    cs.LG 2025-02 conditional novelty 3.0 of 10

    A controlled replication benchmark of nine AMR models on RadioML-2016A, plus experiments showing that moderate SNR training ranges and added recurrent layers can improve accuracy.

  58. PXGen: A Post-hoc Explainable Method for Generative Models

    cs.LG 2025-01 reject novelty 3.0 of 10

    PXGen is a post-hoc, training-free explanation framework that scores anchor samples with intrinsic and extrinsic criteria, groups them by thresholds, and selects representative examples via k-dispersion or k-center.

  59. Iterative Encoding-Decoding VAEs Anomaly Detection in NOAA's DART Time Series: A Machine Learning Approach for Enhancing Data Integrity for NASA's GRACE-FO Verification and Validation

    cs.LG 2024-12 reject novelty 3.0 of 10

    An iterative VAE is proposed for cleaning NOAA DART time series, but the claimed improvement over classical methods is supported only by visual inspection of a single station.

  60. InferPy: Probabilistic Modeling with Deep Neural Networks Made Easy

    cs.LG 2019-08 conditional novelty 3.0 of 10

    InferPy provides a compact, high-level Python API for hierarchical probabilistic models with deep neural networks, built on TensorFlow Probability and Keras.

See all 65 Pith citations

Pith tools