Pith. sign in

REVIEW 3 major objections 4 minor 48 references

A Revisit to Rate-distortion Theory via Optimal Weak Transport

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims to find a new parametric form of the rate-distortion function using optimal weak transport, tying it to Schrödinger bridge equations and rederiving when the Shannon lower bound is achieved.

desk verdict Worth a careful referee: the OWT reformulation is genuinely useful and the main theorems look right, but Theorem 7's proof has an unjustified minimax swap and Theorem 8's proof is too terse; the stress-test's falsity claim does not hold up. read the letter →

arxiv 2501.09362 v3 pith:SQQ52R47 submitted 2025-01-16 cs.IT math.IT

classification cs.ITmath.IT MSC 94A3449Q22
keywords rate-distortiontheoryoptimalweaktransportSchrödingerbridgeproblemShannonlowerboundparametricrepresentationabstractalphabetsentropic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that the rate-distortion function $R(D)$ — the minimum coding rate needed to keep expected distortion at or below $D$ — can be derived from optimal weak transport, a recently introduced generalization of optimal transport in which moving a point $x$ costs a function of the whole conditional distribution of the reconstruction, not just of a single destination. Working on abstract alphabets modeled as Polish spaces (complete separable metric spaces), it derives a parametric representation of $R(D)$ that ties the function to the Schrödinger bridge problem: for any $\beta$ in the subdifferential of $R$ at $D$ (the set of its supporting slopes), the function equals an infimum over reconstruction measures $\nu$ of an explicit functional built from $e^{-\beta\rho(x,y)}$ plus a correction term $L(\nu,\beta)$. When an optimal reconstruction measure exists, the correction vanishes and the optimal joint distribution takes the product-like form $d\pi^* \propto e^{-\beta\rho(x,y)} d\mu d\nu^*$, yielding a Shannon-lower-bound-style formula. As a byproduct, the paper rederives the known achievability conditions for the Shannon lower bound under squared-error distortion without invoking variational calculus. A sympathetic reader would care because the representation offers a new, more geometric handle on $R(D)$ and points numerical methods developed for Schrödinger bridges at rate-distortion computation.

What carries the argument

The carrying object is the rate-distortion function rewritten as a weak-transport problem: the paper decomposes $R(D) = \inf_{\nu} \inf_{\pi \in \Pi(\mu,\nu), E_\pi \rho \le D} I(X;Y)$ and then relaxes the distortion constraint to $J(\nu,\beta) = \inf_{\pi \in \Pi(\mu,\nu)} [ I(X;Y) + \beta (E_\pi \rho - D) ]$. The workhorse identity expresses $J(\nu,\beta)$ as relative entropy against the reference measure $\gamma = K e^{-\beta\rho(x,y)} d\mu d\nu$, so the inner minimization becomes an entropic optimal transport problem whose optimizer has the multiplicative density $d\pi^*/d\gamma = f(x)g(y)$ by the Schrödinger bridge characterization (Lemma 1). That multiplicative structure is what produces the correction term $L(\nu,\beta)$, assembled from the $g(y)$ solving the Schrödinger equations. Existence of the inner minimizer is supplied by the refined weak-transport existence theorem under the paper's moment and lower-semicontinuity assumptions, and the final representation is assembled by restricting $\beta$ to the subdifferential $\partial R(D)$ of the convex rate-distortion function.

What would settle it

Take a binary source with Hamming distortion at a distortion level where a subgradient $\beta$ is known, compute the claimed infimum-over-$\nu$ expression in Theorem 7, and compare the result with the exactly known $R(D)$ from the classical alternating-minimization algorithm for rate-distortion; any gap would expose the $\inf_\nu \sup_\beta$ swap as the failing step. Alternatively, find any source obeying the paper's assumptions for which an optimal reconstruction $\nu^*$ exists but the optimal joint distribution is not proportional to $e^{-\beta\rho(x,y)} d\mu d\nu^*$, which would refute Theorem 8.

Watch

Extended reading notes

Core claim

The central claim is Theorem 7: under the paper's assumptions on the source, the loss function, and finite moments, the rate-distortion function admits the parametric representation $R(D) = \inf_{\nu} \{ -\int \log(\int e^{-\beta\rho(x,y)} d\nu) d\mu - \beta D + L(\nu,\beta) \}$ for every $\beta$ in the subdifferential $\partial R(D)$, where $L(\nu,\beta)$ is a correction term defined through the function $g(y)$ that solves the Schrödinger bridge equations for the reference measure $\gamma = K e^{-\beta\rho} d\mu d\nu$. The route is to rewrite $R(D)$ as an infimum over reconstruction measures $\nu$ of a weak-transport problem, relax the distortion constraint with a Lagrange multiplier $\beta$, and then use the known structure of Schrödinger-bridge optimizers — $d\pi^*/d\gamma = f(x)g(y)$ — to evaluate the inner minimization. Theorem 8 adds that if an optimal reconstruction $\nu^*$ exists, the optimal joint distribution is $d\pi^* = e^{-\beta\rho(x,y)} (\int e^{-\beta\rho(x,y)} d\nu^*)^{-1} d\mu d\nu^*$, so that $L(\nu^*,\beta) = 0$ and $R(D) = -\int_X \log(\int_Y e^{-\beta\rho(x,y)} d\nu^*) d\mu - \beta D$, the form anticipated by the Shannon lower bound. As a stated payoff, Corollary 2 reproduces earlier achievability conclusions for the Shannon lower bound: for quadratic distortion on $\mathbb{R}^n$, the RD function coincides with the bound when the support of the optimal reproduction has an accumulation point; when the bound is not achieved, that support consists of isolated singularities, and a bounded such support forces the reconstruction alphabet to be finite and discrete.

Load-bearing premise

The load-bearing step is the interchange of the infimum over reconstruction measures and the supremum over the Lagrange multiplier — the $\inf_\nu \sup_\beta = \sup_\beta \inf_\nu$ step inside the chain (31) — for which the paper cites no minimax theorem and proves no convex-concave structure; if that swap fails, the parametric representation of $R(D)$ does not follow from the surrounding estimates.

Editorial extensions

If this is right

  • For any abstract source satisfying the paper's assumptions, $R(D)$ can be written as an explicit infimum over reconstruction measures of a functional of $e^{-\beta\rho(x,y)}$, for each $\beta$ in the subdifferential of $R$ at $D$.
  • Where an optimal reconstruction exists, the optimal test channel has the form $d\pi^* \propto e^{-\beta\rho(x,y)} d\mu d\nu^*$, tying rate-distortion-optimal channels directly to Schrödinger bridges.
  • The Shannon lower bound is achieved for squared-error distortion when the optimal reproduction's support has an accumulation point; failing that, the support is made of isolated singularities, and a bounded such support forces a finite discrete reproduction alphabet.
  • The connection suggests that algorithms developed for Schrödinger bridge problems can be brought to bear on computing rate-distortion functions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the representation in Theorem 7 survives numerical checks, the correction term $L(\nu,\beta)$ can be read as the price of using a non-optimal reconstruction measure; that suggests an alternating scheme that solves Schrödinger bridge equations at fixed $\nu$ and then updates $\nu$, with $L(\nu,\beta)$ as a convergence certificate.
  • A natural test case is finite alphabets, where the classical alternating-minimization algorithm gives $R(D)$ exactly: formula (24) should collapse to the familiar fixed-point equations of that algorithm, and checking that would anchor the whole framework.
  • The same weak-transport perspective could in principle be applied to other information-theoretic quantities that are infima over couplings with one free marginal, such as the capacity-cost function or the information bottleneck, giving each a Schrödinger-bridge-style representation.
  • The scope of Theorem 8 is left open because it assumes an optimal reconstruction exists; identifying source classes where that existence is provable, such as Gaussian sources or compact alphabets, would settle how widely the Shannon-lower-bound-style formula holds.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper attempts to reformulate classical rate-distortion theory in the language of optimal weak transport. The authors introduce an OWT formulation of the rate-distortion Lagrangian, invoke existence theorems for weak transport and the static Schrödinger problem, and propose a parametric representation of R(D) (Theorem 7) involving an infimum over reconstruction measures and a correction term L(ν,β) defined through Schrödinger equations. They then claim (Theorem 8) that when an optimal reconstruction exists, the optimal joint distribution takes the Schrödinger form dπ* = e^{-βρ}/(∫e^{-βρ}dν*) dµ×dν*, yielding the Shannon lower bound, and they use this to reproduce K. Rose's theorem on SLB achievability (Corollary 2).

Significance. The topic is of potential interest: a rigorous connection between rate-distortion theory and Schrödinger bridges could yield new structural and computational insights. The paper correctly applies several external results (the weak-transport existence theorems of Backhoff-Veraguas et al. and Csiszár's parametric representation), and the manipulation of the Lagrangian via weak transport is instructive. However, the central new claims are not reliable as stated: Theorem 8 is false, Theorem 7 includes an unjustified minimax interchange, and Corollary 2 is asserted without proof. The advertised connection to Schrödinger bridges is therefore not established.

major comments (3)
  1. [Appendix F, Eq. (39), Theorem 8] The proof of Theorem 8 uses Csiszár's Lemma 1.4 to write dπ* = α(x)e^{-βρ(x,y)} dµ dν* and then identifies α(x) with (∫e^{-βρ(x,y)}dν*)^{-1}. This identification is equivalent to the balance equation ∫ e^{-βρ(x,y)} / ∫ e^{-βρ(x,y')} dν*(y') dµ(x) = 1 for ν*-a.e. y, which is precisely the tightness condition for the Shannon lower bound. Nothing in the existence of an optimal reconstruction for the rate-distortion problem forces this balance. A concrete counterexample is any non-uniform finite-alphabet source with at least three symbols under Hamming distortion: it satisfies Assumptions 1–6, admits an optimal reconstruction for every D, yet R(D) strictly exceeds the expression in (25) for small D. Hence Theorem 8 and the Schrödinger-bridge representation (24)–(25) are false as stated.
  2. [Appendix E, Eq. (31), Theorem 7] The chain of equalities in Eq. (31) interchanges the infimum over ν and the supremum/maximum over β — including 'inf_ν sup_β' to 'sup_β inf_ν', and 'inf_ν max_{β∈∂R(D)}' to 'max_{β∈∂R(D)} inf_ν' — without any minimax theorem or verification of a convex-concave saddle-point structure. The displayed equalities are therefore unsupported. Although the final statement 'R(D)=inf_ν J(ν,β) for β∈∂R(D)' can be obtained directly from Csiszár's subgradient inequality and the definition of R(D), the theorem as stated includes the stronger inf-max equality, which is not established by the given argument.
  3. [Section IV and Corollary 2] Corollary 2, which is advertised as a main byproduct reproducing K. Rose's results without variational calculus, is not proved in the manuscript. The text merely says 'By (24) along with K. Rose's methods of using the completeness of Hermite polynomials, we can reproduce the following conclusion,' with no derivation supplied. Rose's argument is highly nontrivial, and the premise (24) is furnished by the false Theorem 8, so the claimed reproduction is invalid as it stands.
minor comments (4)
  1. [Section II-B and throughout] The spelling 'Schödinger' should be 'Schrödinger' in several places, and there are numerous typographical artifacts (e.g., '/greaterorequalslant' and 'heorem 2') that should be corrected in a revised manuscript.
  2. [Theorems 3 and 4] Both theorems are cited as 'Theorem 2.1 in [7]', which appears inconsistent; the numbering in Csiszár's paper should be checked and the two results distinguished.
  3. [Theorem 7, definition of L(ν,β)] The definition of L(ν,β) is elliptical: it refers to 'g(y) satisfying the Schrödinger equations (7)', but g is determined only up to a multiplicative constant and the domain of the infimum over ν is not made explicit. A self-contained definition would improve readability.
  4. [Appendix D, Proposition 1] The proof of Proposition 1 applies the Wasserstein triangle inequality and bounds W_t(µ,ν) by c^{-1/t} D^{1/t}; the argument assumes D is finite and should state this explicitly.

Circularity Check

1 steps flagged · score 6.0 of 10

Theorem 8's Schrödinger-bridge/SLB conclusion is inserted, not derived: Appendix F silently identifies Csiszár's α(x) with the Gibbs normalization.

  1. other [Theorem 8, Appendix F, displayed equation (39); cf. main-text assertion before Theorem 8]
    "Furthermore, by Lemma 1.4 in [7], we have dπ⋆ = α(x)e^{−βρ(x,y)}dµ × dν⋆ = e^{−βρ(x,y)} / ∫_Y e^{−βρ(x,y)}dν⋆ dµ × dν⋆. (39)"

    The quoted step reduces the theorem's conclusion to a rewriting. Lemma 1.4 in [7], as invoked, provides the factorization dπ⋆ = α(x)e^{−βρ}dµdν⋆ together with an integral constraint on α, but it does not identify α(x) pointwise with the Gibbs normalization (∫ e^{−βρ(x,y)}dν⋆(y))^{-1}. That identification is exactly the balance equation ∫ e^{−βρ(x,y)}(∫ e^{−βρ(x,y′)}dν⋆(y′))^{-1}dµ(x)=1 for ν⋆-a.e. y, i.e., the condition that the Shannon lower bound is tight. This balance is the content of (24)–(25), not a consequence of the existence of an optimal reconstruction ν⋆ nor of Lemma 1.4. Thus the proof assumes the target identity in the final equality of (39).

full rationale

The paper is mostly a genuine re-derivation using external mathematics: Theorem 7 builds on Gozlan et al.'s weak-transport formulation and on the Backhoff-Veraguas–Pammer existence/semicontinuity theorems, and Corollary 2 reproduces Rose's own Hermite-polynomial method; reference [42] is a self-citation but is not load-bearing. The unchecked inf/sup interchange in eq. (31) is a correctness/rigor gap rather than a circular reduction, so it does not raise the circularity score by itself. The genuinely circular step is localized to Theorem 8 and Appendix F. Eq. (39) claims that Csiszár's Lemma 1.4 gives dπ⋆ = α(x)e^{−βρ}dµdν⋆ and then immediately rewrites α(x) as (∫ e^{−βρ}dν⋆)^{-1}. The lemma, as used, supplies only the factorization plus the ν⋆-a.e. integral condition; the pointwise Gibbs form is equivalent to the balance equation and therefore to SLB tightness, which is exactly (24)–(25). The proof therefore inserts the central conclusion rather than deriving it. Because this clean Schrödinger-bridge/SLB formula is the paper's advertised payoff, the circularity is substantive, although the surrounding OWT existence material and the classical RD lemmas remain independent and non-circular.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The paper introduces no fitted constants and no new physical or mathematical entities. It relies on the external optimal weak transport existence theory and on classical rate-distortion properties from Csiszár. The specific assumptions A4-A6 are ad hoc to this paper and are stronger than the classical Csiszár conditions, which limits the claimed generality.

assumptions (8)
  • domain assumption Assumption 1: inf_y ρ(x,y) = 0 for all x.
    Standard normalization of the loss function, taken from Csiszár's setup.
  • domain assumption Assumption 2: existence of random variables ξ, η with I(ξ;η) < ∞ and Eρ(ξ,η) < ∞.
    Ensures the rate-distortion function is not identically infinite.
  • domain assumption Assumption 3: there exists a finite set B ⊂ Y with ∫ ρ(x,B) dµ < ∞.
    Used to guarantee finiteness of R(D) above D_min, following Csiszár.
  • ad hoc to paper Assumption 4: (x,p) → ∫ ρ(x,y) dp is jointly lower semicontinuous on X × P(Y).
    Needed for the first weak transport existence theorem; fails for squared Euclidean distance, as the paper notes.
  • ad hoc to paper Assumption 5: (x,p) → ∫ ρ(x,y) dp is jointly lower semicontinuous on X × P_t(Y).
    Weaker than Assumption 4 and allows ρ(x,y) = ||x-y||^t; supports the refined existence theorem.
  • ad hoc to paper Assumption 6: X = Y are Polish, µ ∈ P_t(X), and ρ(x,y) ≥ c · d(x,y)^t for some c > 0.
    Stronger than Csiszár's assumptions; used to prove the reconstruction distribution lies in P_t(Y) (Proposition 1).
  • standard math Existence and semicontinuity theorems for weak transport (Theorem 1 from [20], Theorem 2 from [21]).
    External results invoked to guarantee minimizers of J(ν,β).
  • standard math Schrödinger bridge structure of entropic optimal transport minimizers (Lemma 1 from [21]).
    External result used to derive the product form of the optimal coupling density.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Revisit to Rate-distortion Theory via Optimal Weak Transport." pith.science (2026). https://pith.science/paper/SQQ52R47

@misc{pith2026250109362,
  author       = {Pith},
  title        = {Pith review of: A Revisit to Rate-distortion Theory via Optimal Weak Transport},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SQQ52R47}},
  note         = {Machine review of arXiv:2501.09362}
}
read the original abstract

This paper revisits the rate-distortion theory from the perspective of optimal weak transport, as recently introduced by Gozlan et al. While the conditions for optimality and the existence of solutions are well-understood in the case of discrete alphabets, the extension to abstract alphabets requires more intricate analysis. Within the framework of weak transport problems, we derive a parametric representation of the rate-distortion function, thereby connecting the rate-distortion function with the Schr\"odinger bridge problem, and establish necessary conditions for its optimality. As a byproduct of our analysis, we reproduce K. Rose's conclusions regarding the achievability of Shannon lower bound concisely, without reliance on variational calculus.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 43 canonical work pages

  1. [1]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication, ” The Bell System Technical Journal , vol. 27, no. 3, pp. 379–423, 1948

  2. [2]

    Coding theorems for a discrete source with a fidelity criterion,

    ——, “Coding theorems for a discrete source with a fidelity criterion,” International Convention Record , vol. 7, pp. 325–350, 1959

  3. [3]

    Berger, Rate-distortion theory: A mathematical basis for data com- pression

    T. Berger, Rate-distortion theory: A mathematical basis for data com- pression. Englewood Cliffs: Prentice-Hall, 1971

  4. [4]

    Computation of channel capacity and rate-di stortion func- tions,

    R. Blahut, “Computation of channel capacity and rate-di stortion func- tions,” IEEE Trans. Inf. Theory , vol. 18, no. 4, pp. 460–473, 1972

  5. [5]

    R. E. Blahut, Principles and practice of information theory . Addison- Wesley Longman Publishing Co., Inc., 1987

  6. [6]

    T. M. Cover and J. A. Thomas, Elements of Information Theory , 2nd ed. New Y ork, NY , USA: Wiley, 2006

  7. [7]

    On an extremum problem of information theor y,

    I. Csiszár, “On an extremum problem of information theor y,” Studia Scientiarum Mathematicarum Hungarica , vol. 9, no. 1, pp. 57–71, 1974

  8. [8]

    An algorithm for computing the capacity of a rbitrary discrete memoryless channels,

    S. Arimoto, “An algorithm for computing the capacity of a rbitrary discrete memoryless channels,” IEEE Trans. Inf. Theory , vol. 18, no. 1, pp. 14–20, 1972

Show all 48 references
  1. [9]

    Rate distor tion theory for general sources with potential application to image com pression,

    F. Rezaei, N. Ahmed, and C. D. Charalambous, “Rate distor tion theory for general sources with potential application to image com pression,” International Journal of Applied Mathematical Sciences , vol. 3, no. 2, pp. 141–165, 2006

  2. [10]

    Rate-disto rtion theory for general sets and measures,

    E. Riegler, H. Bölcskei, and G. Koliander, “Rate-disto rtion theory for general sets and measures,” in Proc. lEEE Int. Symp. Inf. Theory (ISIT) , V ail, CO, USA, 2018, pp. 101–105

  3. [11]

    Lossy compr ession of general random variables,

    E. Riegler, G. Koliander, and H. Bölcskei, “Lossy compr ession of general random variables,” Information and Inference: A Journal of the IMA, vol. 12, no. 3, pp. 1759–1829, 2023

  4. [12]

    The rate-distortion dimensi on of sets and measures,

    T. Kawabata and A. Dembo, “The rate-distortion dimensi on of sets and measures,” IEEE Trans. Inf. Theory, vol. 40, no. 5, pp. 1564–1572, 1994

  5. [13]

    Successive refinement of abst ract sources,

    V . Kostina and E. Tuncel, “Successive refinement of abst ract sources,” IEEE Trans. Inf. Theory , vol. 65, no. 10, pp. 6385–6398, 2019

  6. [14]

    Asymptotic evaluatio n of certain Markov process expectations for large time, I,

    M. D. Donsker and S. S. V aradhan, “Asymptotic evaluatio n of certain Markov process expectations for large time, I,” Commun. Pure Appl. Math., vol. 28, no. 1, pp. 1–47, 1975

  7. [15]

    Successive refinement of in formation,

    W. H. Equitz and T. M. Cover, “Successive refinement of in formation,” IEEE Trans. Inf. Theory , vol. 37, no. 2, pp. 269–275, 1991

  8. [16]

    Transportation distance, shannon informa tion, and source coding,

    R. M. Gray, “Transportation distance, shannon informa tion, and source coding,” in GRETSI Symposium on Signal and Image Processing , 2013, https://ee.stanford.edu/ gray/gretsi.pdf

  9. [17]

    Estimating the rate- distortion function by Wasserstein gradient descent,

    Y . Y ang, S. Eckstein, M. Nutz, and S. Mandt, “Estimating the rate- distortion function by Wasserstein gradient descent,” Advances in Neural Information Processing Systems , vol. 36, 2024

  10. [18]

    Kan torovich duality for general transport costs and applications,

    N. Gozlan, C. Roberto, P .-M. Samson, and P . Tetali, “Kan torovich duality for general transport costs and applications,” J. Funct. Anal. , vol. 273, no. 11, pp. 3327–3405, 2017

  11. [19]

    A survey of the Schrödinger problem and som e of its connections with optimal transport,

    C. Léonard, “A survey of the Schrödinger problem and som e of its connections with optimal transport,” arXiv:1308.0215, 2013

  12. [20]

    Ex istence, du- ality, and cyclical monotonicity for weak transport costs,

    J. Backhoff-V eraguas, M. Beiglböck, and G. Pammer, “Ex istence, du- ality, and cyclical monotonicity for weak transport costs, ” Calculus of V ariations and Partial Differential Equations , vol. 58, no. 203, 2019

  13. [21]

    Applications of w eak transport theory,

    J. Backhoff-V eraguas and G. Pammer, “Applications of w eak transport theory,” Bernoulli, vol. 28, no. 1, pp. 370–394, 2022

  14. [22]

    A mapping approach to rate-distortion comput ation and analysis,

    K. Rose, “A mapping approach to rate-distortion comput ation and analysis,” IEEE Trans. Inf. Theory , vol. 40, no. 6, pp. 1939–1952, 1994

  15. [23]

    O ptimal transport for domain adaptation,

    N. Courty, R. Flamary, D. Tuia, and A. Rakotomamonjy, “O ptimal transport for domain adaptation,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 39, no. 9, pp. 1853–1865, 2016

  16. [24]

    Regularized optimal transp ort for dynamic semi-supervised learning,

    M. E. Hamri and Y . Bennani, “Regularized optimal transp ort for dynamic semi-supervised learning,” arXiv:2103.11937, 2021

  17. [25]

    Wasserstein g enerative ad- versarial networks,

    M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein g enerative ad- versarial networks,” in International Conference on Machine Learning . PMLR, 2017, pp. 214–223

  18. [26]

    Multi-prototype space learning for commonsense-ba sed scene graph generation,

    L. Chen, Y . Song, Y . Cai, J. Lu, Y . Li, Y . Xie, C. Wang, and G. He, “Multi-prototype space learning for commonsense-ba sed scene graph generation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 2, 2024, pp. 1129–1137

  19. [27]

    Transportation cost for Gaussian and ot her product measures,

    M. Talagrand, “Transportation cost for Gaussian and ot her product measures,” Geometric & Functional Analysis (GAF A) , vol. 6, no. 3, pp. 587–600, 1996

  20. [28]

    Connecting GANs, mean- field games, and optimal transport,

    H. Cao, X. Guo, and M. Laurière, “Connecting GANs, mean- field games, and optimal transport,” SIAM Journal on Applied Mathematics , vol. 84, no. 4, pp. 1255–1287, 2024

  21. [29]

    Matching for causal effects via multimarginal unbalanced optimal transport,

    F. Gunsilius and Y . Xu, “Matching for causal effects via multimarginal unbalanced optimal transport,” arXiv:2112.04398, 2021

  22. [30]

    Convexity of mutual informatio n along the Ornstein-Uhlenbeck flow,

    A. Wibisono and V . Jog, “Convexity of mutual informatio n along the Ornstein-Uhlenbeck flow,” in International Symposium on Information Theory and Its Applications (ISITA) , Singapore, 2018, pp. 55–59

  23. [31]

    Transportation proof of an inequality by Anantharam, Jog and Nair,

    T. A. Courtade, “Transportation proof of an inequality by Anantharam, Jog and Nair,” arXiv:1901.10893, 2019

  24. [32]

    Optimal transport meets information science: from measure concentration, to information theory, to machine learning ,

    Y . Bai, “Optimal transport meets information science: from measure concentration, to information theory, to machine learning ,” PhD Thesis, University of Delaware, 2022

  25. [33]

    Information constrained op timal trans- port: From Talagrand, to Marton, to Cover,

    Y . Bai, X. Wu, and A. Özgür, “Information constrained op timal trans- port: From Talagrand, to Marton, to Cover,” IEEE Trans. Inf. Theory , vol. 69, no. 4, pp. 2059–2073, 2023

  26. [34]

    Villani, Optimal transport: old and new , ser

    C. Villani, Optimal transport: old and new , ser. Grundlehren der mathematischen Wissenschaften. Springer, 2009

  27. [35]

    Charac- terization of a class of weak transport-entropy inequaliti es on the line,

    N. Gozlan, C. Roberto, P .-M. Samson, Y . Shu, and P . Tetal i, “Charac- terization of a class of weak transport-entropy inequaliti es on the line,” Annales de l’Institut Henri Poincaré, Probabilités et Stat istiques, vol. 54, no. 3, pp. 1667 – 1693, 2018

  28. [36]

    On a mixture of Brenier and Str assen theorems,

    N. Gozlan and N. Juillet, “On a mixture of Brenier and Str assen theorems,” Proc. Lond. Math. Soc. , vol. 120, no. 3, pp. 434–463, 2020

  29. [37]

    We ak monotone rearrangement on the line,

    J. Backhoff-V eraguas, M. Beiglböck, and G. Pammer, “We ak monotone rearrangement on the line,” Electronic Communications in Probability , vol. 25, pp. 1–16, 2020

  30. [38]

    Stability of mart ingale optimal transport and weak optimal transport,

    J. Backhoff-V eraguas and G. Pammer, “Stability of mart ingale optimal transport and weak optimal transport,” Ann. Appl. Probab., vol. 32, no. 1, pp. 721–752, 2022

  31. [39]

    Weak transpor t for non- convex costs and model-independence in a fixed-income marke t,

    B. Acciaio, M. Beiglböck, and G. Pammer, “Weak transpor t for non- convex costs and model-independence in a fixed-income marke t,” Math. Finance, vol. 31, no. 4, pp. 1423–1453, 2021

  32. [40]

    Über die umkehrung der naturgesetze,

    E. Schrödinger, “Über die umkehrung der naturgesetze, ” Sitzungs- berichte der Preussischen Akademie der Wissenschaften. Ph ysikalisch- Mathematische Klasse , vol. 144, pp. 144–153, 1931

  33. [41]

    Introduction to entropic optimal transport,

    M. Nutz, “Introduction to entropic optimal transport, ” Lecture notes, Columbia University, 2021

  34. [42]

    A revisit to rate-dis tortion theory via optimal weak transport,

    J. Zou, L. Fan, J. Gao, and J. Wang, “A revisit to rate-dis tortion theory via optimal weak transport,” arXiv:2501.09362, 2025

  35. [43]

    Dupuis and R

    P . Dupuis and R. S. Ellis, A weak convergence approach to the theory of large deviations . John Wiley & Sons, 2011

  36. [44]

    Plug-in estimatio n of Schrödinger bridges,

    A.-A. Pooladian and J. Niles-Weed, “Plug-in estimatio n of Schrödinger bridges,” arXiv:2408.11686, 2024

  37. [45]

    Wasserstein proximal algorit hms for the Schrödinger bridge problem: Density control with nonlinea r drift,

    K. Caluya and A. Halder, “Wasserstein proximal algorit hms for the Schrödinger bridge problem: Density control with nonlinea r drift,” IEEE Trans. Automatic Control , vol. 67, no. 3, pp. 1163–1178, 2021

  38. [46]

    An optimal transport approac h for the Schrödinger bridge problem and convergence of sinkhorn alg orithm,

    S. Marino and A. Gerolin, “An optimal transport approac h for the Schrödinger bridge problem and convergence of sinkhorn alg orithm,” J. Sci. Comput. , vol. 85(2), no. 27, 2020

  39. [47]

    Random coding strategies for minimum entro py,

    E. Posner, “Random coding strategies for minimum entro py,” IEEE Trans. Inf. Theory , vol. 21, no. 4, pp. 388–391, 1975

  40. [48]

    Polyanskiy and Y

    Y . Polyanskiy and Y . Wu, Information Theory: From Coding to Learn- ing. Cambridge University Press, 2025. APPENDIX A PROOF OF LEMMA 3 For any fixed πx ∈ P (Y) with Eπρ ≤ D, one can always find the corresponding ν ∈ P (Y) and π ∈ Π( µ, ν), where for any Borel measurable functio...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.