REVIEW 2 major objections 4 minor 1 cited by
Convex Split Lemma without Inequalities
T0 review · 2 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read The convex split lemma, previously an inequality bounding how well a uniform mixture approximates a product state, becomes an exact equality when max mutual information is replaced by collision mutual information.
desk verdict The paper's central equality-based convex split lemma is false as stated; the independent universal bound may survive but has a proof gap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the collision relative entropy $Q_2(\rho\|\sigma)=\operatorname{Tr}[(\sigma^{-1/4}\rho\sigma^{-1/4})^2] = \operatorname{Tr}[\rho\,\sigma^{-1/2}\rho\,\sigma^{-1/2}]$, the $\alpha=2$ sandwiched Rényi divergence. Its quadratic dependence on $\rho$ makes the cross-terms in the convex split mixture factor exactly: off-diagonal pairs $x\ne x'$ contribute $Q_2(\rho_R\|\omega_R)$, diagonal pairs contribute $Q_2(\rho_{RA}\|\omega_R\otimes\sigma_A)$, producing Eq. (18). The proof of Theorem 2 additionally relies on a unitary-covariance reduction (Lemma 2 plus Corollary 2) that lets smoothing of spectral functions be done on probability vectors, and on a constructed local effect $\Lambda$ that connects the smallest non-zero eigenvalue to the Rényi entropy $H_\alpha$.
What would settle it
Calculate both sides of Eq. (18) for a random two-qubit state $\rho_{RA}$ and a random $\sigma_A$; any mismatch would falsify the equality. Separately, for a single-qubit state $\omega$ and a low-rank effect $\Lambda$ with $\operatorname{Tr}[\Lambda^2\omega]=1-2\delta$ at $\delta=0.01$, measure the trace distance between $\omega$ and $\Lambda\omega\Lambda/\operatorname{Tr}[\Lambda^2\omega]$; if it exceeds $\sqrt{2\delta}$, the constant $c=2-\sqrt{3}$ in Theorem 2 does not follow from the stated proof.
Extended reading notes
Core claim
On the paper's own terms, the discovery is Lemma 1: for $\tau^{RAn}$ the uniform mixture over where a correlated copy sits, $Q_2(\tau^{RAn}\|\omega_R\otimes\sigma_A^{\otimes n}) = \frac{n-1}{n}Q_2(\rho_R\|\omega_R)+\frac{1}{n}Q_2(\rho_{RA}\|\omega_R\otimes\sigma_A)$. Since $Q_2$ is the collision relative entropy (the sandwiched Rényi relative entropy of order 2), this downgrades the earlier inequality to an identity and, with $\omega_R=\rho_R$, yields $D_2(\tau\|\rho_R\otimes\sigma^{\otimes n}) = \log(1+\mu/n)$ and $P^2 \le \mu/(\mu+n)$. The paper further claims Theorem 2: $I_\varepsilon^{\max}(A:B)_\rho \le H_\alpha(A)_\rho - \tilde H_\beta^{\uparrow}(A|B)_\rho + \left(\frac{2}{\beta-1}+\frac{1}{1-\alpha}\right)\log\frac{1}{c\varepsilon^2}$ with $c=2-\sqrt{3}$, for all $\alpha\in(0,1)$ and $\beta>1$. From this, Theorem 3 gives $\limsup_{n\to\infty} \frac{1}{n}\operatorname{Cost}_\varepsilon(\mathcal N^{\otimes n}) \le I(A:B)_{\mathcal N}$, proving the achievability half of the reverse quantum Shannon theorem as a corollary.
Load-bearing premise
The exact constant in the universal bound rests on the assumption that a successful post-selection with probability at least $1-2\delta$ changes the state by at most $\sqrt{2\delta}$; standard versions of this assertion give a larger constant, so the stated bound is not fully supported as written.
Editorial extensions
If this is right
- Quantum state splitting cost improves to $\operatorname{Cost}_\varepsilon(\rho^{AA'}) \le \frac12 I_2^{\varepsilon-\delta}(R:A')_\rho + \log(1/\delta)$, with the collision mutual information no larger than the max mutual information.
- The smoothed max mutual information is controlled dimension-independently by additive Rényi entropies, so one-shot capacities no longer need dimension-dependent constants.
- For channel simulation under LOSE, $\limsup_{n\to\infty} \frac1n \operatorname{Cost}_\varepsilon(\mathcal N^{\otimes n}) \le I(A:B)_{\mathcal N}$, recovering the reverse quantum Shannon theorem's achievability direction from the universal bound.
- The purified-distance bound $P^2(\tau,\rho_R\otimes\sigma^{\otimes n}) \le \mu/(\mu+n)$ follows directly from Corollary 1 and improves the earlier $\sqrt{\mu_{\max}/2n}$ trace-distance scaling.
- The equality extends to weighted mixtures, with $Q_2 = (1-t)Q_2(\rho_R\|\omega_R)+tQ_2(\rho_{RA}\|\omega_R\otimes\sigma_A)$ for $t=\sum_x p_x^2$, showing that the uniform mixture is the optimal choice.
Reading between the lines
- Because the equality is exact, I expect the convex split method to sharpen other single-shot protocols where max information was the bottleneck, such as state redistribution and channel coding with finite blocklengths.
- The universal bound likely admits a cleaner form with an optimized constant if the gentle-measurement step is corrected; the qualitative dimension-free statement should survive.
- One could test numerically whether $\frac12 I_2^{\varepsilon-\delta}+\log(1/\delta)$ beats $\frac12 I_\varepsilon^{\max}+\text{const}$ on random bipartite states; typical gaps would show how much of the improvement comes from using collision rather than max information.
- The equality suggests a direct operational meaning for collision mutual information as the precise one-shot cost measure in convex-split-mediated protocols, not merely an upper bound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two main contributions. The first is an equality-based convex split lemma (Lemma 1, Eq. (18)) that replaces the max mutual information in the standard convex split lemma with the collision mutual information, together with applications to quantum state splitting (Theorem 1) and comparisons with other convex-split variants. The second is a universal, dimension-independent upper bound on the smoothed max mutual information (Theorem 2), expressed in terms of Rényi entropies and an explicit function f_{\alpha,\beta}(\varepsilon), with an application to the reverse quantum Shannon theorem (Theorem 3). The proofs are analytic and rely on a small number of external results from the literature, notably [11], [22], and [32].
Significance. If the stated results hold, the paper gives a genuinely useful refinement of a standard tool: an exact convex-split relation controlled by collision mutual information, with a directly applicable quantum-state-splitting bound. The universal bound on smoothed max mutual information is also significant because dimension-independent bounds of this type are scarce and have direct asymptotic consequences. I checked the main suspected weakness, the cross-term step in the proof of Lemma 1, and found it correct: for x\neq x\prime, the expression in Eq. (54) reduces to Q_2(\rho_R\|\omega_R) because (\omega\otimes\sigma)^{-1/2}(\rho_R\otimes\sigma)(\omega\otimes\sigma)^{-1/2}=(\omega^{-1/2}\rho_R\omega^{-1/2})\otimes I_A, and the partial trace removes any dependence on correlations in \rho_{RA}. The numerical counterexample in the review note evaluates Q_2(\rho_R\|\omega_R) as 4, but with the definition in Eq. (9) it is 2, and then both sides of Eq. (18) equal 3. The central problems I found are instead in the proof of Theorem 2, where the gentle-measurement step uses a stronger trace-distance factor than the standard lemma provides, and in the prefactor of Eq.
major comments (2)
- [Proof of Theorem 2, after Eq. (131)] The proof asserts that Tr[\Lambda^A \omega^A \Lambda^A] \geq 1-2\delta implies, by the gentle measurement lemma, that \omega_\Lambda is \sqrt{2\delta}-close to \omega. The standard gentle measurement lemma applied to the effect E=\Lambda^2 gives \|\omega_\Lambda-\omega\|_1 \leq 2\sqrt{1-\mathrm{Tr}[E\omega]} \leq 2\sqrt{2\delta}, not \sqrt{2\delta}. Consequently the condition "2\delta \leq \varepsilon_1^2" should be "8\delta \leq \varepsilon_1^2". With the paper's choice \delta=(2-\sqrt{3})\varepsilon^2 and \varepsilon_1=(\sqrt{3}-1)\varepsilon, the claimed \varepsilon_1-closeness is not guaranteed, so the constant c=2-\sqrt{3} in Eq. (102), and hence the explicit form of f_{\alpha,\beta}(\varepsilon) in Theorem 2 and Eq. (139), is not established as written. The theorem can likely be repaired by taking \delta=c\varepsilon^2/4 or another rescaling, but the displayed bound and its proof must be revised.
- [Remark after Lemma 1, Eq. (20)] The claimed prefactor 1/4 in the trace-distance bound is inconsistent with Eq. (10). Since D_2=\log(1+\mu/n) and Eq. (10) gives D_2\geq\log(1+\|\rho-\sigma\|_1^2), one obtains (1/2)\|\tau-\rho\otimes\sigma^{\otimes n}\|_1 \leq (1/2)\sqrt{\mu/n}, not (1/4)\sqrt{\mu/n}. A simple classical counterexample is p=(1,0), q=(1/2,1/2), for which D_2(p\|q)=\log 2 and the trace distance is 1/2, violating the claimed 1/4 bound. This error does not affect Theorem 1, which uses Corollary 1 rather than Eq. (20), but the displayed improvement and the comparison in item (ii) of the Remark must be corrected or removed.
minor comments (4)
- [Proof of Theorem 2, after Eq. (137)] The sentence "the relation \varepsilon_0+\varepsilon_1\leq\varepsilon holds with equality" is not correct for the chosen values \varepsilon_1=(\sqrt{3}-1)\varepsilon and \varepsilon_0=(2-\sqrt{3})\varepsilon^2: equality holds only at \varepsilon=1. Since only the inequality is needed, this is a wording issue, but the statement should be fixed.
- [Lemma 3, proof] The proof refers to "Lemma ??" for the Löwner-monotonicity of D_{\max} in its second argument; a proper reference or proof should be supplied.
- [Application: Reverse Quantum Shannon Theorem, Eq. (32) vs Eq. (145) and Eq. (159)] The definition of I_{\alpha,\beta}(A:B)_\mathcal{N} is inconsistent across the paper: Eq. (32) uses H_\alpha(A)-\tilde H_\beta^\uparrow(A|B), Eq. (145) uses H_\alpha(B)-\tilde H_\beta^\uparrow(B|R), and the proof of Theorem 3 uses H_\alpha(A)-\tilde H_\beta^\uparrow(A|B). The systems should be matched consistently after the exchange of A and \tilde A in the post-selection argument.
- [Theorem 3, proof, notation near Eq. (149)] The notation for the dimension in the post-selection parameter \varepsilon_n should be clarified: Eq. (144) writes d^2-1, while the proof writes |A|^2-1 and later swaps A with \tilde A. Please ensure a single convention is used throughout.
Circularity Check
No circularity: the central claims are derived from definitions and from external results by other authors; the apparent mathematical issues are correctness defects, not self-referential reasoning.
full rationale
The derivation chain is not circular. Lemma 1 is a direct algebraic expansion of Q2 (Eqs. 50-56); whether Eq. (54) is valid is a mathematical-correctness question, not a reduction of the conclusion to its input. Theorem 2 is proved using external results [22], [11], and inequalities (91)-(95) that originate outside the paper; no parameter is fitted and then renamed a prediction. The self-citations that appear ([5], [26], [27], [28]) are background framework or technical facts, and none is the sole justification for a central claim: the trace-norm-to-classical channel fact from [5] is elementary and independently verifiable, the min/max interchange of [27] is used only in motivation, and [28] is cited for background bounds not used in the final proof of Theorem 2. The applications re-use the existing QSS protocol of [8]/[11] and the post-selection technique of [32], which is legitimate reuse rather than circularity. The reader's noted defects (invalid cross-term in Lemma 1; loose constant in the gentle-measurement application around Eq. (131)) would, if confirmed, make claims unsupported, but they are not instances of circularity. Hence an honest non-finding is appropriate.
Assumptions & free parameters
assumptions (6)
- standard math Finite-dimensional Hilbert spaces and standard properties of sandwiched Rényi divergences, including data processing, monotonicity in α, and the direct-sum property (38).
- domain assumption The gentle measurement lemma, in the concrete form that ΛωΛ is √(2δ)-close to ω if Tr[Λ²ω] ≥ 1-2δ.
- domain assumption The smoothed max-relative entropy bound Dε_max(ρ||σ) ≤ D_β(ρ||σ) + (1/(β-1)) log(1/ε²) from [22].
- domain assumption The post-selection technique and de Finetti state approximation from [32].
- domain assumption Quasi-convexity and additivity properties of smooth max-information from [11].
- standard math Uhlmann's theorem relating purified distance of purifications and data processing inequalities for trace and purified distance.
Cite this review
Pith. "Pith review of Convex Split Lemma without Inequalities." pith.science (2026). https://pith.science/paper/TFBK4T4S
@misc{pith2026250206526,
author = {Pith},
title = {Pith review of: Convex Split Lemma without Inequalities},
year = {2026},
howpublished = {\url{https://pith.science/paper/TFBK4T4S}},
note = {Machine review of arXiv:2502.06526}
}
read the original abstract
We introduce a refinement to the convex split lemma by replacing the max mutual information with the collision mutual information, transforming the inequality into an equality. This refinement yields tighter achievability bounds for quantum source coding tasks, including state merging and state splitting. Furthermore, we derive a universal upper bound on the smoothed max mutual information, where "universal" signifies that the bound depends exclusively on R\'enyi entropies and is independent of the system's dimensions. This result has significant implications for quantum information processing, particularly in applications such as the reverse quantum Shannon theorem.
Figures
Forward citations
Cited by 1 Pith paper
-
Quantum Information Decoupling Beyond Finite Dimensions
Under finite entropy of the manipulated system, infinite-dimensional IID decoupling and quantum state merging achieve the same optimal rates as in finite dimensions (H(A) and 1/2 I(A:R)).
Reference graph
Works this paper leans on
- [25]
- [11]
- [22]
-
[32]
K. Fang, G. Gour, and X. Wang, Towards the ulti- mate limits of quantum channel discrimination (2022), arXiv:2110.14842 [quant-ph]
work page Pith review arXiv 2022
-
[1]
It eliminates the need to invoke the law of large numbers in proving fundamental results in quan- tum Shannon theory, often leading to more in- tuitive and transparent formulations, even in the asymptotic limit
-
[2]
It provides a robust approach for scenarios where the i.i.d. approximation fails, including finite- resource systems, distributed networks with a lim- ited number of nodes, and near-term quantum de- vices operating in noisy or low-repetition regimes. Despite its success, obtaining exact analytical expres- sions for optimal single-shot rates in QIP tasks r...
-
[3]
An equality-based convex split lemma
-
[4]
A universal upper bound on the smoothed max mu- tual information . The convex split lemma [8], a fundamental tool in- spired by classical rejection sampling techniques, plays a central role in proving achievability results for quan- tum source coding. Here, we present an equality-based formulation that replaces the conventional max mutual information with...
work page Pith review arXiv 2025
Show all 47 references
-
[5]
Renner (2005), phD Thesis
R. Renner (2005), phD Thesis
2005
-
[6]
Hayashi, Quantum Information: An Introduction (Springer Berlin Heidelberg, 2006)
M. Hayashi, Quantum Information: An Introduction (Springer Berlin Heidelberg, 2006)
2006
-
[7]
Tomamichel, Quantum Information Processing with Finite Resources: Mathematical Foundations, Springer- Briefs in Mathematical Physics (Springer International Publishing, 2015)
M. Tomamichel, Quantum Information Processing with Finite Resources: Mathematical Foundations, Springer- Briefs in Mathematical Physics (Springer International Publishing, 2015)
2015
-
[8]
Khatri and M
S. Khatri and M. M. Wilde, Principles of quantum communication theory: A modern approach (2024), arXiv:2011.04672 [quant-ph]
2024 arXiv
-
[9]
Gour, Resources of the quantum world (2024), arXiv:2402.05474 [quant-ph]
G. Gour, Resources of the quantum world (2024), arXiv:2402.05474 [quant-ph]
2024
-
[10]
Berta, Single-shot quantum state merging (2009), arXiv:0912.4495 [quant-ph]
M. Berta, Single-shot quantum state merging (2009), arXiv:0912.4495 [quant-ph]
2009 arXiv
-
[12]
Anshu, V
A. Anshu, V. K. Devabathini, and R. Jain, Phys. Rev. Lett. 119, 120506 (2017)
2017
-
[13]
Horodecki, J
M. Horodecki, J. Oppenheim, and A. Winter, Nature 436, 673 (2005)
2005
-
[14]
Abeyesinghe, I
A. Abeyesinghe, I. Devetak, P. Hayden, and A. Winter, Proceedings of the Royal Society A: Mathematical, Phys- ical and Engineering Sciences 465, 2537 (2009)
2009
-
[15]
Berta, M
M. Berta, M. Christandl, and R. Renner, Communica- tions in Mathematical Physics 306, 579 (2011)
2011
-
[16]
C. H. Bennett, I. Devetak, A. W. Harrow, P. W. Shor, and A. Winter, IEEE Transactions on Information The- ory 60, 2926 (2014)
2014
-
[17]
M¨ uller-Lennert, F
M. M¨ uller-Lennert, F. Dupuis, O. Szehr, S. Fehr, and M. Tomamichel, Journal of Mathematical Physics 54, 122203 (2013)
2013
-
[18]
M. M. Wilde, A. Winter, and D. Yang, Communications in Mathematical Physics 331, 593 (2014)
2014
- [19]
-
[20]
Gour and M
G. Gour and M. Tomamichel, Phys. Rev. A 102, 062401 (2020)
2020
-
[21]
Beigi and A
S. Beigi and A. Gohari, IEEE Transactions on Informa- tion Theory 60, 7980 (2014)
2014
-
[23]
Anshu, R
A. Anshu, R. Jain, and N. A. Warsi, IEEE Transactions on Information Theory 65, 1287 (2019)
2019
-
[24]
Anshu and R
A. Anshu and R. Jain, npj Quantum Information 8, 97 (2022)
2022
-
[26]
Regula, L
B. Regula, L. Lami, and N. Datta, Tight relations and equivalences between smooth relative entropies (2025), arXiv:2501.12447 [quant-ph]
2025 arXiv
-
[27]
Sason, IEEE Transactions on Information Theory 62, 23 (2016)
I. Sason, IEEE Transactions on Information Theory 62, 23 (2016)
2016
-
[28]
M. D. Reid and R. C. Williamson, Journal of Machine Learning Research 12, 731 (2011)
2011
-
[29]
Li and Y
K. Li and Y. Yao, Communications in Mathematical Physics 405, 160 (2024)
2024
-
[30]
Chitambar and G
E. Chitambar and G. Gour, Rev. Mod. Phys. 91, 025001 (2019)
2019
-
[31]
Gour and A
G. Gour and A. Winter, Phys. Rev. Lett. 123, 150401 (2019)
2019
-
[33]
H. Qi, Q. Wang, and M. M. Wilde, Journal of Physics A: Mathematical and Theoretical 51, 444002 (2018)
2018
-
[34]
Cooney, M
T. Cooney, M. Mosonyi, and M. M. Wilde, Communica- tions in Mathematical Physics 344, 797 (2016)
2016
-
[35]
Ciganovi´ c, N
N. Ciganovi´ c, N. J. Beaudry, and R. Renner, IEEE Trans- 6 actions on Information Theory 60, 1573 (2014). [32] M. Christandl, R. K¨ onig, and R. Renner, Phys. Rev. Lett. 102, 020504 (2009). Appendix Preliminaries Notations: The letters A, B, C, and R will be used to denote bo...
2014
-
[36]
The initial state is therefore ρRAA′ ⊗ ϕ⊗n
Alice and Bob borrow n copies of the (entangled) state ϕAB. The initial state is therefore ρRAA′ ⊗ ϕ⊗n
-
[37]
Alice apply the isometry channel V ∈CPTP(AA′An → LAn) to her systems, resulting in the state V(ρRAA′ ⊗ ϕ⊗n)
-
[38]
Alice applies a basis measurement on system L in the basis {|x⟩L}x∈[n], and communicate the outcome x of the measurement to Bob
-
[39]
Alice swap the system (register) Ax with A1 ≡ A, and Bob swap the system Bx with B1 ≡ B Since on pure states, the trace distance equals the purified distance, we get from (88) that the state V(ρRAA′ ⊗ ϕ⊗n) at the second step is δn-close to τ R(LAn)Bn . By the data-processing i...
-
[40]
The case α = 1: Here, I1(A : B)ρ, often referred to simply as I(A : B)ρ, is expressed as I(A : B)ρ = D ρAB ρA ⊗ ρB . (97)
-
[41]
The case α = 2: This case is given by I2(A : B)ρ := min σ∈D(B) logQ2 ρAB ρA ⊗ σB . (98)
-
[42]
(99) We also consider the smoothed version of the max mutual information, defined for all ε ∈ (0, 1) and ρ ∈ D(AB) as: I ε max(A : B)ρ := min ρ′∈Bε(ρ) Imax(A : B)ρ′
The case α = ∞: This is expressed in terms of the max-relative entropy as Imax(A : B)ρ := min σ∈D(B) Dmax ρAB ρA ⊗ σB . (99) We also consider the smoothed version of the max mutual information, defined for all ε ∈ (0, 1) and ρ ∈ D(AB) as: I ε max(A : B)ρ := min ρ′∈Bε(ρ) Imax(A...
-
[43]
Minimization over all states ωAB that are ε0-close to ρAB for some ε0 ∈ (0, 1)
-
[44]
golden unit
Minimization is over all effects Λ ∈ eff(A) such that ωAB Λ := ΛAωABΛA Tr h (ΛA)2 ωA i (120) is ε1-close to ωAB, for another ε1 ∈ (0, 1). Note that from the triangle inequality we get that ωAB Λ is (ε0 + ε1)-close to ρAB. Hence, we will choose ε0, ε1 ∈ (0, 1) that satisfies ε0...
-
[45]
Alice simulates V ⊗n in her lab
-
[46]
Alice applies the channel Qn,ℓ = Θ′ n [idℓ] ∈ CPTP(EnA′n → EnBn), where the superchannel Θ ′ n is an εn-error QSS protocol that is used to simulate the action of the channel idEnA′n→EnBn on the pure state ρCnAnEnA′ n := V ⊗n ξCnAn ˜An n . (151)
-
[47]
We now discuss the technical details of the protocol
Alice discard system En. We now discuss the technical details of the protocol. By construction, the three steps of the protocol result with the channel P ˜An→Bn n,ℓ = TrEn ◦ QEnA′n→EnBn n,ℓ ◦ V ˜A→EA′ ⊗n . (152) Since Θ′ n is εn-QSS with respect to the state ρAnCnEnA′ n . Thus...
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.