Pith. sign in

REVIEW 2 cited by

Amortized Variational Inference: When and Why?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11018 v4 pith:MWRKGHNN submitted 2023-07-20 stat.ML cs.LG

classification stat.MLcs.LG
keywords inferencea-vilatentmodelsvariablevariationalf-viused
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In a probabilistic latent variable model, factorized (or mean-field) variational inference (F-VI) fits a separate parametric distribution for each latent variable. Amortized variational inference (A-VI) instead learns a common inference function, which maps each observation to its corresponding latent variable's approximate posterior. Typically, A-VI is used as a step in the training of variational autoencoders, however it stands to reason that A-VI could also be used as a general alternative to F-VI. In this paper we study when and why A-VI can be used for approximate Bayesian inference. We derive conditions on a latent variable model which are necessary, sufficient, and verifiable under which A-VI can attain F-VI's optimal solution, thereby closing the amortization gap. We prove these conditions are uniquely verified by simple hierarchical models, a broad class that encompasses many models in machine learning. We then show, on a broader class of models, how to expand the domain of AVI's inference function to improve its solution, and we provide examples, e.g. hidden Markov models, where the amortization gap cannot be closed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Active Inference Agents for Delayed and Long-Horizon Environments

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A policy-conditional world model trained under active inference enables single-lookahead planning over hundreds of steps and beats a DQN baseline on energy-efficient control of parallel machines.

  2. Density-Informed Pseudo-Counts for Calibrated Evidential Deep Learning

    stat.ML 2026-02 conditional novelty 5.0 of 10

    DIP-EDL sets Dirichlet pseudo-counts to the product of marginal input density and learned class probabilities, concentrating on the true label distribution while sending OOD inputs to the prior.

Pith tools