REVIEW 6 minor 13 references
Pinsker's inequality for adapted total variation
T0 review · 0 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read An adapted version of Pinsker's inequality holds: $ATV(\mu,\nu)\le \sqrt{n}\sqrt{2H(\mu|\nu)}$, and the constant $\sqrt{n}$ is sharp.
desk verdict A clean, tight adapted Pinsker inequality with a sound proof; worth refereeing as a short note. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Lemma 2.1, a recursive decomposition of $ATV$. It says that $ATV$ equals the total variation of the first marginal plus the integral, against the pointwise minimum of the two marginals, of the total variation between the successive conditional kernels, continued stage by stage; an optimal bicausal coupling is exactly one whose kernels place maximal mass on the diagonal at every layer. The decomposition turns a single constrained infimum over bicausal couplings into a sequence of ordinary transport problems, using the characterization that a coupling is bicausal precisely when its successive disintegrations are themselves ordinary couplings of the corresponding conditional laws, chosen measurably. Once the decomposition is available, the proof needs only the classical Pinsker inequality, Jensen's inequality, and the chain rule of relative entropy.
What would settle it
Compute $ATV(\mu,\nu)$ and $H(\mu|\nu)$ for any pair of $n$-step laws: a single pair with $H(\mu|\nu)<\infty$ and $ATV(\mu,\nu)>\sqrt{n}\sqrt{2H(\mu|\nu)}$ would refute the theorem. For the sharpness claim, evaluate the product Bernoulli family of Corollary 2.3 with fixed $n$ and $0<\varepsilon\ll 1$: here $ATV(\mu,\nu)=2-2(1-\varepsilon)^n$ and $H(\mu|\nu)=n(2\varepsilon^2+o(\varepsilon^2))$, so the normalized ratio $ATV^2/(nH)$ tends to $2$; if a numerical evaluation of this family ever produced a ratio above $2$, the bound as stated would be false.
Extended reading notes
Core claim
The central claim is that causal, filtration-respecting comparisons of process laws are quantitatively controlled by relative entropy. Theorem 1.1 asserts $ATV(\mu,\nu)\le \sqrt{n}\sqrt{2H(\mu|\nu)}$ for all probabilities on $X_1\times\cdots\times X_n$, where $ATV$ is the adapted total variation distance obtained by restricting couplings to bicausal transport plans. This is the direct generalization of the classical bound $TV(\mu,\nu)\le \sqrt{2H(\mu|\nu)}$, recovered when $n=1$. The proof decomposes $ATV$ into the total variation of the first marginals plus the expected total variation of the successive conditional kernels, applies the ordinary Pinsker bound to each layer, and reassembles the layers through Jensen's inequality and the chain rule of relative entropy. Corollary 2.3 proves the bound is tight: in the product Bernoulli example, $ATV(\mu,\nu)^2/(nH(\mu|\nu))\to 2$ as the bias tends to zero, so the constant $\sqrt{n}$ cannot be improved.
Load-bearing premise
The load-bearing premise is the imported structural fact that any transport plan respecting the time order of both processes can be assembled step-by-step from ordinary couplings of their successive conditional laws, with the kernels chosen measurably; if that fact failed, the recursive decomposition of $ATV$ that carries the whole proof would not be available.
Editorial extensions
If this is right
- If two $n$-step process laws have relative entropy at most $\delta$, their adapted total variation is at most $\sqrt{n}\sqrt{2\delta}$; in particular, entropy convergence of process laws forces causal transport convergence at rate $\sqrt{n\delta}$.
- The bound is tight even among product measures: flat product Bernoulli laws already drive the ratio $ATV^2/(nH)$ to its maximal value $2$, so the $\sqrt{n}$ constant cannot be improved by restricting to exchangeable or memoryless processes.
- Since $ATV\ge TV$, the inequality subsumes the classical Pinsker bound and gives a quantitative link between KL divergence and process-level causal distances, which is useful whenever a model is trained by minimizing relative entropy against a target process law.
- Combined with the earlier linear comparison $ATV\le (2n-1)TV$, the new inequality gives a dimension-dependent bound $ATV\le \sqrt{n}\sqrt{2H}$, which is sharper than the linear comparison precisely when entropy is small relative to TV.
Reading between the lines
- The recursive decomposition suggests a template: any stagewise cost built from a bounded metric by taking the maximum over coordinates should admit an entropy bound with the same $\sqrt{n}$ horizon factor; the discrete metric here is the simplest instance.
- Sharpness among product measures indicates that the $\sqrt{n}$ factor is a genuine cost of the horizon, not of path dependence or long memory; even independent repeated experiments pay the full factor in the worst case.
- A natural testable extension is whether the constant can be improved when the conditional kernels are Lipschitz or otherwise regular, since the paper's examples are purely atomic and do not exercise regularity.
- The discrete-time proof is built on summing $n$ layers, so a continuous-time analogue would need a different argument; the paper's own mention of continuous-time transport inequalities suggests such an analogue is not immediate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This short note proves an adapted Pinsker inequality: for probability measures μ and ν on a product Polish space X1 × ... × Xn, the adapted total variation satisfies ATV(μ, ν) ≤ √n √(2H(μ|ν)). The proof decomposes ATV recursively into a sum of conditional total variation terms (Lemma 2.1), applies classical Pinsker termwise, and then uses Jensen's inequality together with the entropy chain rule to obtain the √n factor. A product Bernoulli example (Corollary 2.3) shows that the constant √n is asymptotically sharp as the marginals approach each other.
Significance. The result gives a clean, dimension-dependent Pinsker inequality for adapted transport, complementing the known equivalence ATV ≤ (2n−1)TV and improving the dependence on n. The proof is transparent, parameter-free, and the tightness example is explicit. The recursive decomposition lemma for ATV is independently useful. The main limitation is that this is a short observation rather than a deep structural theorem, and the proof imports the standard bicausal-coupling characterization from [2] without proof.
minor comments (6)
- [Lemma 2.1 proof, general induction] In the displayed induction hypothesis, the differential 'dπ_{x1,y1}(x1,y1)' should be 'dπ_{x1,y1}(x2,y2)', and the terms 'ν_{x1:2}' and 'µ_{x1}∧ν_{x1}' in the off-diagonal part should read 'ν_{y1:2}' and 'µ_{x1}∧ν_{y1}'; although these coincide on the diagonal and the intended formula is clear, the current notation makes the induction step unnecessarily confusing.
- [Theorem 1.1 proof] In the definition of the auxiliary measure m, the atom should be written 'δ0' rather than 'δt'.
- [Corollary 2.3] The entropy expression 'H(µ, |ν)' contains a misplaced comma and bar; it should be 'H(µ|ν)'.
- [Corollary 2.3] The displayed expansion of '(1−ε)^{2n}' has a typographical garble '42n(2n−1)/2'; the intended coefficient '4·(2n)(2n−1)/2' is correct.
- [References] Reference [2] spells 'Proposition' as 'Propositon'.
- [Lemma 2.2] The notation '({Ω = {0,1}, 2Ω})' appears with a mismatched brace; this is a minor formatting issue.
Circularity Check
No circularity: the adapted Pinsker inequality is derived from independent standard results plus classical inequalities, and the tightness example is computed from the definition.
full rationale
The derivation chain is self-contained in the relevant sense. Theorem 1.1 is proved from Lemma 2.1, which is a representation of ATV obtained from the defining bicausal-coupling infimum using the characterization [2, Prop. 5.1] and measurable selection [2, Prop. 5.2]; this is an imported standard theorem, not the target inequality, and it is parameter-free with stated assumptions that do not include ATV <= sqrt(n) sqrt(2H). The rest of the proof applies classical Pinsker, Jensen, and the entropy chain rule; no fitted parameter or 'predicted' quantity is involved, and the target bound is never assumed. Corollary 2.3 computes ATV for product Bernoulli measures directly from Lemma 2.1 and the elementary entropy expansion, giving an independent tightness example. The only overlap with the authors' prior work is the cited bicausal-coupling characterization, which is external mathematical support rather than a reduction of the conclusion to its own input. Hence no circular step is present.
Assumptions & free parameters
assumptions (4)
- standard math Classical Pinsker inequality TV(μ,ν) ≤ sqrt(2H(μ|ν))
- standard math Chain rule for relative entropy via successive disintegrations
- standard math Characterization of bicausal couplings via successive disintegrations, including measurable selection ([2, Prop. 5.1 and 5.2])
- standard math Jensen's inequality applied to a measure m of total mass at most n
Cite this review
Pith. "Pith review of Pinsker's inequality for adapted total variation." pith.science (2026). https://pith.science/paper/FC3XSIPQ
@misc{pith2026250622106,
author = {Pith},
title = {Pith review of: Pinsker's inequality for adapted total variation},
year = {2026},
howpublished = {\url{https://pith.science/paper/FC3XSIPQ}},
note = {Machine review of arXiv:2506.22106}
}
abstract
Pinsker's classical inequality asserts that the total variation $TV(\mu, \nu)$ between two probability measures is bounded by $\sqrt{ 2H(\mu|\nu)}$ where $H$ denotes the relative entropy (or Kullback-Leibler divergence). Considering the discrete metric, $TV$ can be seen as a Wasserstein distance and as such possesses an adapted variant $ATV$. Adapted Wasserstein distances have distinct advantages over their classical counterparts when $\mu, \nu$ are the laws of stochastic processes $(X_k)_{k=1}^n, (Y_k)_{k=1}^n$ and exhibit numerous applications from stochastic control to machine learning. In this note we observe that the adapted total variation distance $ATV$ satisfies the Pinsker-type inequality $$ ATV(\mu, \nu)\leq \sqrt{n} \sqrt{2 H(\mu|\nu)}.$$
Reference graph
Works this paper leans on
-
[2]
J. Backhoff-Veraguas, M. Beiglb¨ ock, Y. Lin, and A. Zalashko. Causal transport in discrete time and applications.SIAM J. Optim. , 27(4):2528–2562, 2017
work page 2017
-
[1]
B. Acciaio, A. Kratsios, and G. Pammer. Designing universal causal deep learning models: The geometric (hyper) transformer. Mathematical Finance, 34(2):671–735, 2024
work page 2024
- [3]
-
[4]
D. Bartl and J. Wiesel. Sensitivity of multiperiod optimization problems with respect to the adapted wasserstein distance. SIAM J. Financ. Math. , 14(2):704–720, 2023
work page 2023
-
[5]
J. Blanchet, M. Larsson, J. Park, and J. Wiesel. Bounding adapted wasserstein metrics. 2024
work page 2024
- [6]
-
[7]
S. Eckstein and G. Pammer. Computational methods for adapted optimal transport. Ann. Appl. Probab. , 34(1A):675 – 713, 2024
work page 2024
- [8]
Show all 13 references
-
[9]
Jiang and J
Y. Jiang and J. Obloj. Sensitivity of causal distributionally robust optimization. 2024
2024
-
[10]
Lassalle
R. Lassalle. Causal transference plans and their Monge–Kantorovich problems. Stoch. Anal. Appl., 36(3):452–484, 2018
2018
-
[11]
G. C. Pflug and A. Pichler. Multistage Stochastic Optimization . Springer Series in Operations Research and Financial Engineering. Springer, Cham, 2014
2014
-
[12]
C. Villani. Optimal Transport, Old and New, volume 338 ofGrundlehren der mathematischen Wissenschaften. Springer, 2009
2009
-
[13]
Xu and B
T. Xu and B. Acciaio. Conditional COT-GAN for video prediction with kernel smoothing. In NeurIPS 2022 Workshop on Robustness in Sequence Modeling , 2022
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.