Pith. sign in

REVIEW 6 minor 13 references

Pinsker's inequality for adapted total variation

T0 review · 0 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read An adapted version of Pinsker's inequality holds: $ATV(\mu,\nu)\le \sqrt{n}\sqrt{2H(\mu|\nu)}$, and the constant $\sqrt{n}$ is sharp.

desk verdict A clean, tight adapted Pinsker inequality with a sound proof; worth refereeing as a short note. read the letter →

arxiv 2506.22106 v1 pith:FC3XSIPQ submitted 2025-06-27 math.PR cs.ITmath.IT

classification math.PRcs.ITmath.IT MSC 60E1549Q22
keywords PinskerinequalityadaptedtotalvariationrelativeentropybicausalcouplingcausaloptimaltransportWassersteindistancetightnessstochasticprocesses
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This note establishes an adapted Pinsker inequality: for probability measures $\mu,\nu$ on a product space $X_1\times\cdots\times X_n$, the adapted total variation distance satisfies $ATV(\mu,\nu)\le \sqrt{n}\sqrt{2H(\mu|\nu)}$. The extra factor $\sqrt{n}$ is the price of requiring the comparison to respect the causal filtration structure of the processes, and it is shown to be unavoidable: for $n$-fold products of slightly biased Bernoulli measures the ratio $ATV(\mu,\nu)^2/(nH(\mu|\nu))$ tends to $2$ as the bias goes to zero, so no smaller constant works in general. The statement matters because adapted Wasserstein distances are the natural way to compare laws of stochastic processes in applications from stochastic control to distributionally robust optimization and machine learning, where transports must not use future information.

What carries the argument

The load-bearing object is Lemma 2.1, a recursive decomposition of $ATV$. It says that $ATV$ equals the total variation of the first marginal plus the integral, against the pointwise minimum of the two marginals, of the total variation between the successive conditional kernels, continued stage by stage; an optimal bicausal coupling is exactly one whose kernels place maximal mass on the diagonal at every layer. The decomposition turns a single constrained infimum over bicausal couplings into a sequence of ordinary transport problems, using the characterization that a coupling is bicausal precisely when its successive disintegrations are themselves ordinary couplings of the corresponding conditional laws, chosen measurably. Once the decomposition is available, the proof needs only the classical Pinsker inequality, Jensen's inequality, and the chain rule of relative entropy.

What would settle it

Compute $ATV(\mu,\nu)$ and $H(\mu|\nu)$ for any pair of $n$-step laws: a single pair with $H(\mu|\nu)<\infty$ and $ATV(\mu,\nu)>\sqrt{n}\sqrt{2H(\mu|\nu)}$ would refute the theorem. For the sharpness claim, evaluate the product Bernoulli family of Corollary 2.3 with fixed $n$ and $0<\varepsilon\ll 1$: here $ATV(\mu,\nu)=2-2(1-\varepsilon)^n$ and $H(\mu|\nu)=n(2\varepsilon^2+o(\varepsilon^2))$, so the normalized ratio $ATV^2/(nH)$ tends to $2$; if a numerical evaluation of this family ever produced a ratio above $2$, the bound as stated would be false.

Watch

Extended reading notes

Core claim

The central claim is that causal, filtration-respecting comparisons of process laws are quantitatively controlled by relative entropy. Theorem 1.1 asserts $ATV(\mu,\nu)\le \sqrt{n}\sqrt{2H(\mu|\nu)}$ for all probabilities on $X_1\times\cdots\times X_n$, where $ATV$ is the adapted total variation distance obtained by restricting couplings to bicausal transport plans. This is the direct generalization of the classical bound $TV(\mu,\nu)\le \sqrt{2H(\mu|\nu)}$, recovered when $n=1$. The proof decomposes $ATV$ into the total variation of the first marginals plus the expected total variation of the successive conditional kernels, applies the ordinary Pinsker bound to each layer, and reassembles the layers through Jensen's inequality and the chain rule of relative entropy. Corollary 2.3 proves the bound is tight: in the product Bernoulli example, $ATV(\mu,\nu)^2/(nH(\mu|\nu))\to 2$ as the bias tends to zero, so the constant $\sqrt{n}$ cannot be improved.

Load-bearing premise

The load-bearing premise is the imported structural fact that any transport plan respecting the time order of both processes can be assembled step-by-step from ordinary couplings of their successive conditional laws, with the kernels chosen measurably; if that fact failed, the recursive decomposition of $ATV$ that carries the whole proof would not be available.

Editorial extensions

If this is right

  • If two $n$-step process laws have relative entropy at most $\delta$, their adapted total variation is at most $\sqrt{n}\sqrt{2\delta}$; in particular, entropy convergence of process laws forces causal transport convergence at rate $\sqrt{n\delta}$.
  • The bound is tight even among product measures: flat product Bernoulli laws already drive the ratio $ATV^2/(nH)$ to its maximal value $2$, so the $\sqrt{n}$ constant cannot be improved by restricting to exchangeable or memoryless processes.
  • Since $ATV\ge TV$, the inequality subsumes the classical Pinsker bound and gives a quantitative link between KL divergence and process-level causal distances, which is useful whenever a model is trained by minimizing relative entropy against a target process law.
  • Combined with the earlier linear comparison $ATV\le (2n-1)TV$, the new inequality gives a dimension-dependent bound $ATV\le \sqrt{n}\sqrt{2H}$, which is sharper than the linear comparison precisely when entropy is small relative to TV.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The recursive decomposition suggests a template: any stagewise cost built from a bounded metric by taking the maximum over coordinates should admit an entropy bound with the same $\sqrt{n}$ horizon factor; the discrete metric here is the simplest instance.
  • Sharpness among product measures indicates that the $\sqrt{n}$ factor is a genuine cost of the horizon, not of path dependence or long memory; even independent repeated experiments pay the full factor in the worst case.
  • A natural testable extension is whether the constant can be improved when the conditional kernels are Lipschitz or otherwise regular, since the paper's examples are purely atomic and do not exercise regularity.
  • The discrete-time proof is built on summing $n$ layers, so a continuous-time analogue would need a different argument; the paper's own mention of continuous-time transport inequalities suggests such an analogue is not immediate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. This short note proves an adapted Pinsker inequality: for probability measures μ and ν on a product Polish space X1 × ... × Xn, the adapted total variation satisfies ATV(μ, ν) ≤ √n √(2H(μ|ν)). The proof decomposes ATV recursively into a sum of conditional total variation terms (Lemma 2.1), applies classical Pinsker termwise, and then uses Jensen's inequality together with the entropy chain rule to obtain the √n factor. A product Bernoulli example (Corollary 2.3) shows that the constant √n is asymptotically sharp as the marginals approach each other.

Significance. The result gives a clean, dimension-dependent Pinsker inequality for adapted transport, complementing the known equivalence ATV ≤ (2n−1)TV and improving the dependence on n. The proof is transparent, parameter-free, and the tightness example is explicit. The recursive decomposition lemma for ATV is independently useful. The main limitation is that this is a short observation rather than a deep structural theorem, and the proof imports the standard bicausal-coupling characterization from [2] without proof.

minor comments (6)
  1. [Lemma 2.1 proof, general induction] In the displayed induction hypothesis, the differential 'dπ_{x1,y1}(x1,y1)' should be 'dπ_{x1,y1}(x2,y2)', and the terms 'ν_{x1:2}' and 'µ_{x1}∧ν_{x1}' in the off-diagonal part should read 'ν_{y1:2}' and 'µ_{x1}∧ν_{y1}'; although these coincide on the diagonal and the intended formula is clear, the current notation makes the induction step unnecessarily confusing.
  2. [Theorem 1.1 proof] In the definition of the auxiliary measure m, the atom should be written 'δ0' rather than 'δt'.
  3. [Corollary 2.3] The entropy expression 'H(µ, |ν)' contains a misplaced comma and bar; it should be 'H(µ|ν)'.
  4. [Corollary 2.3] The displayed expansion of '(1−ε)^{2n}' has a typographical garble '42n(2n−1)/2'; the intended coefficient '4·(2n)(2n−1)/2' is correct.
  5. [References] Reference [2] spells 'Proposition' as 'Propositon'.
  6. [Lemma 2.2] The notation '({Ω = {0,1}, 2Ω})' appears with a mismatched brace; this is a minor formatting issue.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the adapted Pinsker inequality is derived from independent standard results plus classical inequalities, and the tightness example is computed from the definition.

full rationale

The derivation chain is self-contained in the relevant sense. Theorem 1.1 is proved from Lemma 2.1, which is a representation of ATV obtained from the defining bicausal-coupling infimum using the characterization [2, Prop. 5.1] and measurable selection [2, Prop. 5.2]; this is an imported standard theorem, not the target inequality, and it is parameter-free with stated assumptions that do not include ATV <= sqrt(n) sqrt(2H). The rest of the proof applies classical Pinsker, Jensen, and the entropy chain rule; no fitted parameter or 'predicted' quantity is involved, and the target bound is never assumed. Corollary 2.3 computes ATV for product Bernoulli measures directly from Lemma 2.1 and the elementary entropy expansion, giving an independent tightness example. The only overlap with the authors' prior work is the cited bicausal-coupling characterization, which is external mathematical support rather than a reduction of the conclusion to its own input. Hence no circular step is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard external results: classical Pinsker, the chain rule for entropy, the bicausal-coupling characterization, and Jensen's inequality. No free parameters are fitted and no new entities are introduced.

assumptions (4)
  • standard math Classical Pinsker inequality TV(μ,ν) ≤ sqrt(2H(μ|ν))
    Used for the base case and to bound each conditional TV term in Lemma 2.1. Standard result, applied without proof.
  • standard math Chain rule for relative entropy via successive disintegrations
    Used in the final step of Theorem 1.1 to turn sums of conditional entropies into H(μ|ν). Standard for probability measures on product spaces.
  • standard math Characterization of bicausal couplings via successive disintegrations, including measurable selection ([2, Prop. 5.1 and 5.2])
    Basis for Lemma 2.1's recursive formula for ATV. Taken from Backhoff-Veraguas, Beiglböck, Lin, and Zalashko (2017).
  • standard math Jensen's inequality applied to a measure m of total mass at most n
    Used in the proof of Theorem 1.1 to bound the sum of square roots by sqrt(n) times the square root of the summed conditional entropies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Pinsker's inequality for adapted total variation." pith.science (2026). https://pith.science/paper/FC3XSIPQ

@misc{pith2026250622106,
  author       = {Pith},
  title        = {Pith review of: Pinsker's inequality for adapted total variation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FC3XSIPQ}},
  note         = {Machine review of arXiv:2506.22106}
}
abstract

Pinsker's classical inequality asserts that the total variation $TV(\mu, \nu)$ between two probability measures is bounded by $\sqrt{ 2H(\mu|\nu)}$ where $H$ denotes the relative entropy (or Kullback-Leibler divergence). Considering the discrete metric, $TV$ can be seen as a Wasserstein distance and as such possesses an adapted variant $ATV$. Adapted Wasserstein distances have distinct advantages over their classical counterparts when $\mu, \nu$ are the laws of stochastic processes $(X_k)_{k=1}^n, (Y_k)_{k=1}^n$ and exhibit numerous applications from stochastic control to machine learning. In this note we observe that the adapted total variation distance $ATV$ satisfies the Pinsker-type inequality $$ ATV(\mu, \nu)\leq \sqrt{n} \sqrt{2 H(\mu|\nu)}.$$

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 12 canonical work pages

  1. [2]

    Backhoff-Veraguas, M

    J. Backhoff-Veraguas, M. Beiglb¨ ock, Y. Lin, and A. Zalashko. Causal transport in discrete time and applications.SIAM J. Optim. , 27(4):2528–2562, 2017

  2. [1]

    Acciaio, A

    B. Acciaio, A. Kratsios, and G. Pammer. Designing universal causal deep learning models: The geometric (hyper) transformer. Mathematical Finance, 34(2):671–735, 2024

  3. [3]

    Bartl, M

    D. Bartl, M. Beiglb¨ ock, and G. Pammer. The Wasserstein space of stochastic processes. J. Eur. Math. Soc. , 2024. To appear. arXiv:2104.14245

  4. [4]

    Bartl and J

    D. Bartl and J. Wiesel. Sensitivity of multiperiod optimization problems with respect to the adapted wasserstein distance. SIAM J. Financ. Math. , 14(2):704–720, 2023

  5. [5]

    Blanchet, M

    J. Blanchet, M. Larsson, J. Park, and J. Wiesel. Bounding adapted wasserstein metrics. 2024

  6. [6]

    Cont and F

    R. Cont and F. R. Lim. Causal transport on path space. 2024

  7. [7]

    Eckstein and G

    S. Eckstein and G. Pammer. Computational methods for adapted optimal transport. Ann. Appl. Probab. , 34(1A):675 – 713, 2024

  8. [8]

    F¨ ollmer

    H. F¨ ollmer. Optimal couplings on Wiener space and an extension of Talagrand’s transport inequality. In Stochastic analysis, filtering, and stochastic optimization , pages 147–175. Springer, Cham, [2022] ©2022

Show all 13 references
  1. [9]

    Jiang and J

    Y. Jiang and J. Obloj. Sensitivity of causal distributionally robust optimization. 2024

  2. [10]

    Lassalle

    R. Lassalle. Causal transference plans and their Monge–Kantorovich problems. Stoch. Anal. Appl., 36(3):452–484, 2018

  3. [11]

    G. C. Pflug and A. Pichler. Multistage Stochastic Optimization . Springer Series in Operations Research and Financial Engineering. Springer, Cham, 2014

  4. [12]

    C. Villani. Optimal Transport, Old and New, volume 338 ofGrundlehren der mathematischen Wissenschaften. Springer, 2009

  5. [13]

    Xu and B

    T. Xu and B. Acciaio. Conditional COT-GAN for video prediction with kernel smoothing. In NeurIPS 2022 Workshop on Robustness in Sequence Modeling , 2022

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.