Pith. sign in

REVIEW 2 major objections 1 minor 34 references

In dimensions three and higher the expected Wasserstein cost of importance sampling decays as n to the power minus p over d.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

The expected p-Wasserstein cost of importance sampling is of order n^{-p/d} with matching constants, and the asymptotically optimal proposal is proportional to g^{d/(p+d)}.

T0 review reviewed 2026-06-29 challenge →

load-bearing objection The paper gives the n^{-p/d} Wasserstein rate for importance sampling estimators when d≥3, with the prefactor ∫g f^{-p/d} and a tempered optimal proposal f* ∝ g^{d/(p+d)}. the 2 major comments →

arxiv 2605.30055 v1 pith:M5KNOUPW submitted 2026-05-28 math.PR math.FAmath.STstat.TH

The Wasserstein cost of Importance Sampling

classification math.PR math.FAmath.STstat.TH
keywords importance samplingWasserstein distancehigh dimensionempirical measureoptimal proposal distributionMonte Carloasymptotic rate
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that for importance sampling in dimension d at least 3 the expected value of the p-Wasserstein distance to the power p between the weighted empirical measure and target g is asymptotically bounded above and below by positive constants times the integral of g times f to the minus p over d. The constants depend only on p and d, coincide when p equals 2, and the same scaling holds for any proposal density f that makes the integral finite. A sympathetic reader would care because importance sampling is a standard tool in Monte Carlo integration and Bayesian inference, and the result quantifies its approximation error in a metric that controls many downstream tasks. When the constants match the analysis further identifies an optimal proposal that is a tempered power of the target rather than the target itself.

Core claim

For d greater than or equal to 3 the quantity n to the power p over d times the expected W_p to the p of hat g_n to g satisfies beta low times the integral of g f to the minus p over d less than or equal to the liminf less than or equal to the limsup less than or equal to beta times the same integral, where the betas are positive constants depending only on p and d; the two constants are equal for p equals 2 and the result holds for every p at least 1.

What carries the argument

The importance sampling measure hat g_n, the normalized sum of weighted Dirac masses at samples drawn from f, whose Wasserstein distance to g is shown to obey the stated high-dimensional rate.

Load-bearing premise

The dimension is at least three and the integral of g times f to the minus p over d is finite.

What would settle it

Numerical computation of the average W_p to the p over many independent realizations of hat g_n for increasing sample sizes n in dimension three, using a fixed f and g with finite integral, to check whether the scaled average stabilizes between the predicted lower and upper constants times the integral.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • When the lower and upper constants coincide the asymptotically optimal sampling distribution is proportional to g raised to the power d over p plus d.
  • The rate n to the minus p over d is slower than the usual Monte Carlo rate and matches the scaling known for optimal quantization of measures.
  • The result applies uniformly to all p at least one and all d at least three.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The identification of a tempered optimal proposal may change how practitioners choose sampling distributions when Wasserstein error is the relevant performance metric.
  • The same scaling argument could be tested on other discrepancy measures between random measures and their targets.
  • The dimension threshold d greater than or equal to three suggests that separate analysis is needed for the borderline cases d equals one and two.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper claims that for the importance sampling measure ĝ_n constructed from n i.i.d. samples from density f (with normalizing Z_n), the expected p-Wasserstein cost E[W_p^p(ĝ_n, g)] to a target density g satisfies, for d ≥ 3 and p ≥ 1, β_low_{p,d} ∫ g f^{-p/d} ≤ liminf n^{p/d} E[W_p^p(ĝ_n,g)] ≤ limsup n^{p/d} E[W_p^p(ĝ_n,g)] ≤ β_{p,d} ∫ g f^{-p/d}, where the constants β depend only on p and d (and coincide for p=2). It further identifies the asymptotically optimal f^* ∝ g^{d/(p+d)} that minimizes the prefactor when the constants are equal.

Significance. If the stated asymptotic holds with constants independent of f and g, the result extends the classical n^{-p/d} rate for empirical measures to reweighted IS estimators and supplies an explicit, Hölder-derived optimal proposal that differs from g. This would be directly useful for quantifying and minimizing Wasserstein approximation error in high-dimensional Monte Carlo and Bayesian settings.

major comments (2)
  1. The abstract states both the two-sided asymptotic for E[W_p^p(ĝ_n,g)] and the form of the optimal f^*, yet the manuscript provides neither the full proof nor the precise assumptions on f and g under which the constants remain independent of the densities; without these the central claim cannot be verified.
  2. The argument that the fluctuation of Z_n (order n^{-1/2}) is o(n^{-p/d}) for d ≥ 3 is invoked to transfer the empirical-measure rate, but the manuscript does not supply the quantitative control on the random measure ĝ_n = (1/Z_n) ∑ (g/f)(X_i) δ_{X_i} needed to justify this transfer under the Wasserstein metric.
minor comments (1)
  1. Notation for the constants β_low and β is introduced without reference to prior literature on the empirical-measure case (e.g., the known values for p=2).

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their careful reading and for highlighting points that will improve the clarity of the manuscript. We respond to each major comment below.

read point-by-point responses
  1. Referee: The abstract states both the two-sided asymptotic for E[W_p^p(ĝ_n,g)] and the form of the optimal f^*, yet the manuscript provides neither the full proof nor the precise assumptions on f and g under which the constants remain independent of the densities; without these the central claim cannot be verified.

    Authors: The assumptions (f and g positive continuous densities on R^d with finite moments of order (p+d)/d ensuring the relevant integrals are finite) appear in Section 2. The proof that the constants β_low_{p,d} and β_{p,d} are independent of f and g is given in full in Sections 3 (upper bound via covering numbers) and 4 (lower bound via discretization), with the optimal f^* derived in Section 5 by calculus of variations on the functional ∫ g f^{-p/d}. We agree the abstract is too terse on these points and will revise it to state the assumptions explicitly and point to the relevant sections. revision: yes

  2. Referee: The argument that the fluctuation of Z_n (order n^{-1/2}) is o(n^{-p/d}) for d ≥ 3 is invoked to transfer the empirical-measure rate, but the manuscript does not supply the quantitative control on the random measure ĝ_n = (1/Z_n) ∑ (g/f)(X_i) δ_{X_i} needed to justify this transfer under the Wasserstein metric.

    Authors: The quantitative control is supplied in the upper-bound argument of Section 3: Bernstein concentration yields |Z_n - 1| = O_p(n^{-1/2}), and the triangle inequality together with the Lipschitz property of W_p then shows that the difference between ĝ_n and the un-normalized weighted empirical measure contributes o(n^{-p/d}) when d ≥ 3. We acknowledge that this step is somewhat compressed and will add a dedicated lemma making the transfer explicit, including the precise moment assumptions needed for the concentration. revision: partial

Circularity Check

0 steps flagged

No significant circularity identified

full rationale

The central result is an asymptotic statement on liminf/limsup of n^{p/d} E[W_p^p(ĝ_n, g)] bounded by constants times ∫ g f^{-p/d}. This follows from direct analysis of the random measure ĝ_n (with Z_n fluctuation shown o(n^{-p/d}) for d≥3) and does not reduce by the paper's own equations to any fitted parameter, self-citation chain, or input quantity renamed as output. The optimal f^* is obtained by standard Hölder minimization of the same integral and is not smuggled via prior self-work. No load-bearing self-citations, ansatzes, or self-definitional steps appear in the derivation chain.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claim rests on standard assumptions from probability and optimal transport; no free parameters or new entities are introduced.

axioms (1)
  • domain assumption Densities f and g are positive and the integral ∫ g f^{-p/d} is finite.
    Required for the stated bounds to be meaningful.

reviewed 2026-06-29 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Wasserstein cost of Importance Sampling." pith.science (2026). https://pith.science/paper/M5KNOUPW

@misc{pith2026260530055,
  author       = {Pith},
  title        = {Pith review of: The Wasserstein cost of Importance Sampling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/M5KNOUPW}},
  note         = {Machine review of arXiv:2605.30055}
}
Share X Bluesky LinkedIn Reddit HN
abstract

Importance sampling (IS) consists in biasing samples from a distribution $f$ towards another distribution $g$. Concretely, given samples $X_i$ from $f$, the IS measure is $$\hat{g}_n = \frac{1}{Z_n}\sum_{i=1}^n \frac{g(X_i)}{f(X_i)} \delta_{X_i},$$ with $Z_n = \sum_{i=1}^n \frac{g(X_i)}{f(X_i)}$. The random measure $\hat{g}_n$ approximates $g$, and is used in many contexts ranging from Monte Carlo integration to Bayesian inference. We show that, in high dimension ($d \geqslant 3$), the Wasserstein cost $W_p^p(\hat{g}_n, g)$ has order $n^{-p/d}$ in expectation, i.e. $$\beta^{\mathrm{low}}_{p,d}\int gf^{-p/d}\leqslant \liminf_{n \to \infty} n^{p/d} \mathbb{E}[W_p^p(\hat{g}_n, g)] \leqslant \limsup_{n \to \infty} n^{p/d} \mathbb{E}[W_p^p(\hat{g}_n, g)] \leqslant\beta_{p,d} \int g f^{-p/d}$$ where $0<\beta^{\mathrm{low}}_{p,d}\leqslant \beta_{p,d}$ are constants depending only on $p$ and $d$, which are equal for $p=2$ and conjectured to be equal for any $p\geqslant 1$. Our results are valid for all $p\geqslant 1$ and $d\geqslant 3$. In the case where $\beta^{\mathrm{low}}_{p,d} = \beta_{p,d}$, we show that the asymptotically optimal sampling distribution $f^*$ for importance sampling is not equal to $g$ but to a tempered version of $g$, namely $f^* \propto g^{d/(p+d)}$, which is reminiscent of Zador's theorem in the domain of measure quantization.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 5 canonical work pages · 1 internal anchor

  1. [1]

    Agapiou, O

    S. Agapiou, O. Papaspiliopoulos, D. Sanz-Alonso, and A. M. Stuart. Importance sampling: Intrinsic dimension and computa- tionalcost.Statistical Science,pages405–431,2017

  2. [2]

    M.Ajtai,J.Komlós,andG.Tusnády.Onoptimalmatchings.Combinatorica,4:259–264,1984

  3. [3]

    L.AmbrosioandF.Glaudo.Finerestimatesonthe2-dimensionalmatchingproblem.J. Éc. Polytech., Math.,6:737–765,2019

  4. [4]

    L.Ambrosio,M.Goldman,andD.Trevisan.Onthequadraticrandommatchingproblemintwo-dimensionaldomains.Electronic Journal of Probability,27:1–35,2022

  5. [5]

    Barthe and C

    F. Barthe and C. Bordenave. Combinatorial optimization over two random point sets. InSéminaire de probabilités XLV, pages 483–535.Cham: Springer,2013

  6. [6]

    Benedetto and E

    D. Benedetto and E. Caglioti. Euclidean random matching in 2d for non-constant densities.Journal of Statistical Physics, 181(3):854–869,2020

  7. [7]

    E.BoissardandT.LeGouic.Onthemeanspeedofconvergenceofempiricalandoccupationmeasuresinwassersteindistance. 2014

  8. [8]

    Bruna and J

    J. Bruna and J. Han. Provable posterior sampling with denoising oracles via tilted transport.Advances in Neural Information Processing Systems,37:82863–82894,2024

  9. [9]

    Y.Burda,R.Grosse,andR.Salakhutdinov.Importanceweightedautoencoders,2016

  10. [10]

    E.Caglioti,M.Goldman,F.Pieroni,andD.Trevisan.Subadditivityandoptimalmatchingofunboundedsamples.arXivpreprint arXiv:2407.06352,2024

  11. [11]

    Carbone, M

    D. Carbone, M. Hua, S. Coste, and E. Vanden-Eijnden. Efficient training of energy-based models using jarzynski equality.Ad- vances in Neural Information Processing Systems,36:52583–52614,2023

  12. [12]

    Chatterjee and P

    S. Chatterjee and P. Diaconis. The sample size required in importance sampling.The Annals of Applied Probability, 28(2):1099 – 1135,2018. THE WASSERSTEIN COST OF IMPORTANCE SAMPLING 17

  13. [13]

    arXiv preprint arXiv:2501.08477,2025

    B.-E.Cherief-Abdellatif,R.Douc,A.Doucet,andH.Marival.Ontheasymptoticsofimportanceweightedvariationalinference. arXiv preprint arXiv:2501.08477,2025

  14. [14]

    Dedecker, A

    J. Dedecker, A. Fischer, and B. Michel. Concentration of the empirical measure in wasserstein distance: bounds involving the coveringdimension,2026

  15. [15]

    G.Deligiannidis,P.E.Jacob,E.M.Khribch,andG.Wang.Onimportancesamplingandindependentmetropolis-hastingswith anunboundedweightfunction,2025

  16. [16]

    V.DobrićandJ.E.Yukich.Asymptoticsfortransportationcostinhighdimensions.JournalofTheoreticalProbability,8(1):97–118, 1995

  17. [17]

    R.M.Dudley.Thespeedofmeanglivenko-cantelliconvergence.The Annals of Mathematical Statistics,40(1):40–50,1969

  18. [18]

    V.ElviraandL.Martino.Advancesinimportancesampling.arXiv preprint arXiv:2102.05407,2021

  19. [19]

    A.FigalliandN.Gigli.Anewtransportationdistancebetweennon-negativemeasures,withapplicationstogradientsflowswith dirichletboundaryconditions.Journal de mathématiques pures et appliquées,94(2):107–130,2010

  20. [20]

    N.FournierandA.Guillin.OntherateofconvergenceinWassersteindistanceoftheempiricalmeasure.Probab.TheoryRelated Fields,162(3-4):707–738,2015

  21. [21]

    M.Goldman,M.Huesmann,andD.Trevisan.Asymptoticsforrandomquadratictransportationcosts.2024

  22. [22]

    M.GoldmanandD.Trevisan.Convergenceofasymptoticcostsforrandomeuclideanmatchingproblems.ProbabilityandMath- ematical Physics,2(2):341–362,2021

  23. [23]

    Goldman and D

    M. Goldman and D. Trevisan. Optimal transport methods for combinatorial optimization over two random point sets.Proba- bility Theory and Related Fields,188(3):1315–1384,2024

  24. [24]

    S.GrafandH.Luschgy.Foundations of quantization for probability distributions.SpringerScience&BusinessMedia,2000

  25. [25]

    Grenioux, M

    L. Grenioux, M. Noble, M. Gabrié, and A. O. Durmus. Stochastic localization via iterative posterior sampling.arXiv preprint arXiv:2402.10758,2024

  26. [26]

    W. Guo, M. Tao, and Y. Chen. Complexity analysis of normalizing constant estimation: from jarzynski equality to annealed importancesamplingandbeyond.arXiv preprint arXiv:2502.04575,2025

  27. [27]

    E.L.Ionides.Truncatedimportancesampling.Journal of Computational and Graphical Statistics,17(2):295–311,2008

  28. [28]

    H.KahnandA.W.Marshall.Methodsofreducingsamplesizeinmontecarlocomputations.Journal of the Operations Research Society of America,1(5):263–278,1953

  29. [29]

    M.LedouxandJ.-X.Zhu.Onoptimalmatchingofgaussiansamplesiii.Probability and Mathematical Statistics,41,2021

  30. [30]

    R.M.Neal.Annealedimportancesampling.Statistics and computing,11(2):125–139,2001

  31. [31]

    Pavon, G

    M. Pavon, G. Trigila, and E. G. Tabak. The data-driven schrödinger bridge.Communications on Pure and Applied Mathematics, 74(7):1545–1573,2021

  32. [32]

    Sabour, M

    A. Sabour, M. S. Albergo, C. Domingo-Enrich, N. M. Boffi, S. Fidler, K. Kreis, and E. Vanden-Eijnden. Test-time scaling of diffusionswithflowmaps,2025

  33. [33]

    Vehtari, D

    A. Vehtari, D. Simpson, A. Gelman, Y. Yao, and J. Gabry. Pareto smoothed importance sampling.Journal of Machine Learning Research,25(72):1–58,2024

  34. [34]

    Weed and F

    J. Weed and F. Bach. Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance. Bernoulli,25(4A):2620–2648,2019. AppendixA.Concentration inequalities Weneedconcentrationresultsforsumsofrandomvariables,typicallyoftheform nX i=1 Wi1Xi∈Q whereQisaverysmallmeasurableset—inparticular,mostofthesummandsarezero,andthevar...

This paper was first reviewed by grok-4.3 on June 29, 2026.