Pith. sign in

REVIEW 4 cited by

Amortized In-Context Bayesian Posterior Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.06601 v1 pith:GX2EQCVK submitted 2025-02-10 cs.LG cs.AIstat.ML

Amortized In-Context Bayesian Posterior Estimation

classification cs.LG cs.AIstat.ML
keywords inferenceposterioramortizedbayesianestimationin-contextmethodscontext
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Bayesian inference provides a natural way of incorporating prior beliefs and assigning a probability measure to the space of hypotheses. Current solutions rely on iterative routines like Markov Chain Monte Carlo (MCMC) sampling and Variational Inference (VI), which need to be re-run whenever new observations are available. Amortization, through conditional estimation, is a viable strategy to alleviate such difficulties and has been the guiding principle behind simulation-based inference, neural processes and in-context methods using pre-trained models. In this work, we conduct a thorough comparative analysis of amortized in-context Bayesian posterior estimation methods from the lens of different optimization objectives and architectural choices. Such methods train an amortized estimator to perform posterior parameter inference by conditioning on a set of data examples passed as context to a sequence model such as a transformer. In contrast to language models, we leverage permutation invariant architectures as the true posterior is invariant to the ordering of context examples. Our empirical study includes generalization to out-of-distribution tasks, cases where the assumed underlying model is misspecified, and transfer from simulated to real problems. Subsequently, it highlights the superiority of the reverse KL estimator for predictive problems, especially when combined with the transformer architecture and normalizing flows.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Universal priors: solving empirical Bayes via Bayesian inference and pretraining

    stat.ML 2026-02 conditional novelty 8.0

    A simple random prior-on-prior lets pretrained transformers achieve near-optimal empirical Bayes regret uniformly over all test priors, and length generalization matches α-posterior inference.

  2. Efficient Autoregressive Inference for Transformer Probabilistic Models

    stat.ML 2025-10 conditional novelty 7.0

    A causal autoregressive buffer enables efficient batched autoregressive sampling and joint density evaluation in set-based transformer models by caching context and attending to prior predictions.

  3. Reinforced sequential Monte Carlo for amortised sampling

    cs.LG 2025-10 conditional novelty 6.0

    A method that trains neural samplers using SMC-collected off-policy samples and an importance-weighted replay buffer improves mode coverage on multi-modal targets.

  4. The Milky Way - Large Magellanic Cloud Interaction with Simulation Based Inference

    astro-ph.GA 2025-10 conditional novelty 5.0

    Simulation-based inference on outer-halo star velocities gives a Milky Way reflex speed of 26.4 km/s and an LMC enclosed mass of 9.2×10^10 solar masses within 50 kpc.