{"id":"1028806f-8567-4a31-aabf-16f0cd714413","arxiv_id":"2608.11031","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For any finite set T and law mu on T, the largest expected Bernoulli-process value under coupling with mu is, up to universal constants, the rate-distortion integral of RD_mu(t) and the ell-1-plus-Gaussian decomposition delta_T(mu).","lead":"This paper proves a new, information-theoretic version of the Bernoulli theorem, characterizing the expected supremum of random sign processes through a rate-distortion integral and a distributional decomposition. A smart generalist should care because it offers a simpler proof of a celebrated theorem and a stronger distributional result that connects probability to information theory.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lower bound of Theorem 2.1 rests on unproved type-lifting identity Eq. (51); if that identity or its proof depends on the Bednorz–Latała theorem, the claimed independent proof collapses.","rationale":"Read in good faith, the analytic core of Sections 3–5 and Appendix A is coherent: the area theorem in Lemma 3.3, the leave-one-coordinate-out estimator in Lemma 3.6, the cMMSE-to-RD step in Proposition 4.3, and the multiscale decomposition in Lemma 5.3 all appear internally consistent at the level of the manuscript. I found no internal contradiction in those parts. The decisive weak point is Lemma 6.1's Eq. (51), which the reader also identified. It is used once, but the fixed-law lower bound ∫ RD_μ ≲ B(μ), and hence the full distributional theorem, depends on it. Because the acknowledgments explicitly describe prior notes that derive rate–distortion statements 'with the Bernoulli theorem as input', one cannot rule out circularity without seeing the proof of the Bernoulli type-lifting identity. This is not a refutation: if the identity has an independent proof, the concern disappears. But it does make acceptance conditional on a self-contained derivation of Eq. (51) that does not invoke the theorem being proved. Since the reader already returned CONDITIONAL for the same reason, my read does not change the verdict.","tokens_in":23333,"tokens_out":18234,"duration_ms":210104,"concrete_test":"Isolate Eq. (51) as a standalone lemma and supply a self-contained proof from the definition B(μ)=sup_coupling E⟨ε,X⟩ and type-class combinatorics, checking every step for any invocation of Theorem 1.2, Theorem 2.3, or [BL14]. If the only available proof of [Liu25, Lemma 5] or [vH25, Prop. 3.1] passes through the Bernoulli theorem or van Handel's notes in a way that uses it, then Lemma 6.1 is circular and the proof of Theorem 2.1 must be reworked before the distributional claim can be accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 6.1, the left-to-right direction of Theorem 2.1, hinges on Eq. (51): B(μ) = lim_N N^(-1) b(T_N(μ)). The manuscript states this as a consequence of [Liu25, Lemma 5] and [vH25, Prop. 3.1] but gives no derivation in the present text. This is not a cosmetic omission: Eq. (51) is exactly the step that converts a set-level Bernoulli bound, Theorem 2.2 applied to T_N(μ), into the fixed-law bound ∫ RD_μ ≲ B(μ). If the identity is false, or if its proof in the cited works uses the Bernoulli theorem as an input, then Lemma 6.1 collapses, and with it the distributional lower bound and the claim of an independent proof of the Bernoulli theorem. The acknowledgment that van Handel's notes yield rate–distortion statements 'together with the Bernoulli theorem as input' makes the circularity risk concrete, and the manuscript does not show that the Bernoulli type-lifting identity avoids that input.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims a new information-theoretic proof of the Bernoulli theorem of Bednorz and Latała. For a finite set T and a law μ on T, it defines a rate–distortion functional RD_μ(t), a Cauchy-channel Bayes risk cMMSE_μ(t), and a decomposition functional δ_T(μ). Theorem 2.1 asserts that B(μ), the largest expected value of the Bernoulli process over couplings of the index with law μ, is, up to universal constants, equal to the integrated rate–distortion value ∫ RD_μ(t) dt and to δ_T(μ); at the set level, this gives b(T) ≍ R(T) ≍ Λ(T). The proof is organized as a cycle: Section 3 shows Bernoulli width dominates the integrated Cauchy cMMSE; Section 4, using a Cauchy information–estimation inequality proved in Appendix A, shows the cMMSE integral dominates the rate–distortion integral; Section 5 proves that the rate–distortion integral dominates δ_T(μ) and, via a minimax argument, that Λ(T) is dominated by R(T); Section 6 uses a type-lifting identity to obtain the missing fixed-law lower bound ∫ RD_μ ≲ B(μ) and a Gaussianization argument for the reverse bound. The paper also claims that this gives the first fixed-law, distributional characterization of Bernoulli-process suprema.","tokens_in":23527,"tokens_out":8606,"duration_ms":87850,"significance":"If the proof is correct, this is a significant result. It would provide a new proof of the Bednorz–Latała theorem that avoids their chaining machinery, introduce an information-theoretic functional for Bernoulli processes analogous to the Gaussian majorizing-measure functional, and yield a genuine distributional strengthening that existing Bernoulli proofs do not appear to give. The paper is detailed and internally consistent in its main comparisons: the Cauchy-area argument in §3, the information–estimation inequality of Appendix A, the rate–distortion comparison in §4, and the minimax step in §5 are all carefully developed with explicit constants. The central weakness is the unproved type-lifting identity (Eq. 51), which is load-bearing for the fixed-law lower bound.","major_comments":[{"comment":"The type-lifting identity B(μ) = lim_{N→∞} (1/N) b(T_N(μ)) is stated without proof and is load-bearing for the left-to-right direction of Theorem 2.1. Lemma 6.1 first derives, for compatible N, the bound ∫_0^∞ RD_μ(t) dt ≤ 80π b(T_N(μ))/N + o(1) using Theorem 2.2 and an entropy-loss estimate, and then converts this into the required fixed-law lower bound ∫ RD_μ ≲ B(μ) precisely through Eq. (51). If Eq. (51) were false, or if its proof in the cited works used the Bernoulli theorem as an input, the claimed independent proof would collapse. The manuscript cites [Liu25, Lemma 5] and [vH25, Proposition 3.1] but gives no derivation. The concern is made concrete by the acknowledgment that van Handel's notes yield the rate–distortion characterization only 'together with the Bernoulli theorem as input.' The authors state that the present proof instead uses Eq. (18) to obtain the rate–distortion characterization, but Eq. (51) is a separate identity and the text does not show that it follows from Eq. (18) or that its proof in the cited references avoids the Bernoulli theorem. The authors must either prove Eq. (51) in the present paper or provide a reference with a complete proof that does not invoke the Bernoulli theorem; without this, the fixed-law lower bound, and with it the distributional strengthening and the independent-proof claim, is not established.","section":"§6, Lemma 6.1, Eq. (51)"},{"comment":"The proof structure indicated in §2.4 says that type lifting gives ∫_0^∞ RD_μ(t) dt ≲ B(μ), but the actual proof in §6 imports the identity (51) from prior work rather than deriving it. This is not a cosmetic omission: the entropy-loss correction in Lemma 6.1 is quantitatively small (O(log N)), but the passage from b(T_N(μ)) to B(μ) is an exact asymptotic identity that must be justified. If Eq. (51) is unavailable or depends on the very theorem being proved, then the lower bound in Theorem 2.1 is unsupported and the set-level lower bound b(T) ≳ Λ(T) also loses its noncircular foundation through the chain R(T) ≲ b(T) in Theorem 2.2.","section":"§2.4 and §6, proof of Lemma 6.1"}],"minor_comments":[{"comment":"The typeset text contains repeated OCR-like artifacts, most notably the string '/suppress la' inside 'Bednorz and Lata/suppress la' and in several later occurrences of the name; these should be cleaned up.","section":"Throughout"},{"comment":"The Lipschitz bound for RD with respect to total-variation distance is stated without proof and is plausible, but a short derivation would help the reader: the factor 2d min{n, diam_2(T)^2/t^2} follows from a maximal coupling and the boundedness of the capped quadratic distortion.","section":"§6, Lemma 6.1, rational approximation"},{"comment":"In the comparison of the dyadic sum with the integral, the monotonicity of RD_μ(t) is used; a sentence explicitly noting that the infimum in (9) is nonincreasing in t would make the step easier to verify.","section":"§5.3, Lemma 5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a strong manuscript, and the main technical comparisons appear sound. However, the type-lifting identity Eq. (51) is a load-bearing imported result whose independence from the Bernoulli theorem is not demonstrated. The editor should require the authors either to supply a self-contained proof of Eq. (51) or to identify a reference that proves it without using the Bernoulli theorem. If they cannot do so, the manuscript should not be accepted in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read of arXiv:2608.11031. It deserves a serious referee. The paper does two real things: it gives a new route to the Bednorz–Latała Bernoulli theorem using a Cauchy additive channel and a new Cauchy information–estimation inequality, and it proves a distributional strengthening (Theorem 2.1) that was previously available only for Gaussian processes. For a field that has lived with BL's heavy machinery, a cleaner proof plus a fixed-law characterization is a meaningful step, not a repackaging.\n\nThe skeleton is coherent. Sections 3–4 connect Bernoulli width, cMMSE, and rate–distortion; Section 5 performs the ℓ1-plus-Gaussian decomposition; Section 6 closes the loop. The Cauchy channel is the right analogue: the Poisson kernel is harmonic, and the scalar information–estimation inequality in Appendix A is a genuine calculation. The differentiation-under-the-integral step is compressed but correct in substance. The posterior-replica coupling and the kernel argument in Lemma A.3 are solid. I checked the chaining decomposition in Lemma 5.3 and the telescoping argument; it works. The proof of Theorem 2.3 avoids the BL route and uses only the Gaussian MMT as input.\n\nSoft spots, in proportion. The main one is Eq. (51) in Lemma 6.1: the identity B(μ)=lim_N (1/N)b(T_N(μ)) is imported from [Liu25] and unpublished notes of van Handel, with no derivation in the text. It is load-bearing for the distributional lower bound ∫RD_μ ≲ B(μ). A referee should verify that this identity is proved in the cited sources without using the Bernoulli theorem. But the stress-test worry that this kills the independent proof of the Bernoulli theorem is wrong: the set-level proof uses Theorem 2.2 and Theorem 2.3 and never invokes Eq. (51). The identity is a fixed-law issue, not a set-level one.\n\nTwo smaller points. Lemma 5.2 invokes Sion's minimax theorem on the noncompact domain (R^n)^T; coercivity likely allows truncation, but as written it is a gap. The rational-to-general approximation in Lemma 6.1 is routine, and the stated continuity bounds are fine.\n\nBottom line: the paper is for people working on empirical processes, generic chaining, and minimax estimation. It is not a desk reject. Send it to a competent referee with instructions to check Eq. (51) and its provenance. If that identity is clean, Theorem 2.1 is a major result. I would bring it to reading group and would cite the distributional theorem if I worked in that area.","headline":"A genuinely new and largely correct information-theoretic route to the Bernoulli theorem; the imported type-lifting identity is the main gap, but it does not threaten the independent set-level proof.","tokens_in":24034,"tokens_out":14887,"would_cite":true,"duration_ms":141338,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60G15","46B09","62F15","94A34","94A15"],"pacs":[],"model":"deepseek-v4-flash","headline":"For every prescribed law of the index, Bernoulli-process suprema are characterized, up to universal constants, by an information-theoretic rate–distortion integral.","keywords":["Bernoulli processes","Rademacher suprema","rate-distortion theory","Cauchy channel","Bayesian estimation","majorizing measures","information-estimation inequality","fixed-law characterization"],"falsifier":"A concrete check on the riskiest step: take a finite $T$ and a non-rational law $\\mu$ (for example, weights proportional to $\\sqrt{2}$ on three points), compute $B(\\mu)$ by direct coupling optimization, compute the limit $\\lim_{N\\to\\infty} b(\\mathcal{T}_N(\\mu))/N$ for the exact type classes $\\mathcal{T}_N(\\mu)$, and compare; a mismatch would contradict Eq. (51). For the theorem itself, one could compute the ratio of $B(\\mu)$ to $\\int_0^\\infty RD_\\mu(t)\\,dt$ on a sequence of sets with growing dimension, such as the vertices of a cube or the $\\ell^1$ ball; if that ratio tends to $0$ or $\\infty$, the universal-constant claim fails.","tokens_in":23130,"feed_emoji":"🎲","tokens_out":10034,"duration_ms":87302,"temperature":0.7,"pith_summary":"The paper sets out to prove that the maximum of a Bernoulli process over a finite set is captured by an information-theoretic quantity. For any prescribed law $\\mu$ of the index, the largest expected value $B(\\mu)$ attainable by coupling that index with the independent $\\pm 1$ variables is, up to universal constants, the area under a rate–distortion curve $\\int_0^\\infty RD_\\mu(t)\\,dt$, and also equals a certain $\\ell^1$-plus-Gaussian decomposition cost $\\delta_T(\\mu)$. This is a distributional, fixed-law strengthening of the Bernoulli theorem, which previously was known only in a worst-case form. The proof proceeds by translating lower bounds into Bayesian-estimation limits in a Cauchy additive channel, then comparing the resulting capped Bayes risk with the rate–distortion integral. A sympathetic reader would care because it offers a new proof of a long-standing theorem and, for the first time for a non-Gaussian process, a characterization that holds for each individual index law.","feed_headline":"Rate–distortion integral pins Bernoulli-process suprema","feed_subtitle":"New Bayesian proof via a Cauchy channel fixes the largest Rademacher expectation for every index law.","key_machinery":"The central mechanism is the Cauchy additive channel $Y_t=X+tZ$, with $Z$ having independent standard Cauchy coordinates, paired with the capped quadratic distortion $\\varphi_t(x,x')=\\sum_{i=1}^n (1\\wedge |x_i-x'_i|^2/t^2)$. The load-bearing analytic input is a Cauchy information–estimation inequality: the mutual information of the input and the Cauchy observation is controlled by an integral of the posterior-replica capped risk, playing the role that the Gaussian I–MMSE identity plays for Gaussian channels. An interpolation potential whose endpoint is the Bernoulli width converts the width into an area under a Bayes-risk curve; a posterior-replica coupling turns that risk into the rate–distortion functional; and a multiscale distributional decomposition converts the rate–distortion integral into the $\\ell^1$-plus-Gaussian form. Type lifting, comparing a law with the uniform laws on exact empirical-distribution fibers, supplies the return comparison from the set-level bound back to each prescribed-law $B(\\mu)$.","core_discovery":"The central discovery is Theorem 2.1: for every finite $T\\subset\\mathbb{R}^n$ and every probability measure $\\mu$ on $T$, one has $B(\\mu)\\asymp \\int_0^\\infty RD_\\mu(t)\\,dt\\asymp \\delta_T(\\mu)$, with universal constants. Here $B(\\mu)$ is the supremum of $\\mathbb{E}\\langle\\varepsilon,X\\rangle$ over couplings of $X\\sim\\mu$ with a Rademacher vector $\\varepsilon$; $RD_\\mu(t)$ is the infimum, over couplings of $\\mu$ with itself, of mutual information plus the capped quadratic distortion $\\sum_i (1\\wedge |x_i-x'_i|^2/t^2)$; and $\\delta_T(\\mu)$ is the infimum, over decompositions of the identity map $a:T\\to\\mathbb{R}^n$, of the expected $\\ell^1$ cost of $a$ plus the Gaussian functional of the residual law $(\\mathrm{id}_T-a)_\\#\\mu$. Supremizing over $\\mu$ recovers the classical Bernoulli theorem $b(T)\\asymp R(T)\\asymp \\Lambda(T)$, where $R$ is the set-level rate–distortion area and $\\Lambda$ is the $\\ell^1$-plus-Gaussian width. On the paper's own terms, this is the first prescribed-law characterization of suprema of a Bernoulli process.","pith_inferences":["A natural extension the authors leave implicit is that the same Bayesian scheme should work for other symmetric stable noises whose kernels have similar harmonic and stability properties, yielding fixed-law characterizations for other subexponential processes.","The unproved type-lifting identity cited from prior work could in principle be derived within this framework; if it can be, the distributional theorem becomes self-contained and the assumption identified below can be dropped.","The rate–distortion formulation suggests an algorithmic route: for structured sets $T$, estimating $RD_\\mu(t)$ by alternating optimization over couplings may be easier than computing Bernoulli width directly, and comparing the two for moderate dimension would provide a testable validation of the constant-scale equivalence.","Because $B(\\mu)$ gives per-law information, the theorem may transfer the fixed-law Gaussian results to sums of independent signs, potentially feeding back into symmetrization bounds for empirical processes with prescribed weight distributions."],"forward_implications":["If correct, the distributional theorem gives a previously unavailable fixed-law characterization: for any prescribed index law, the largest attainable Bernoulli-process expectation is, up to constants, an explicit rate–distortion integral.","The proof yields a new, information-theoretic route to the Bernoulli theorem, bypassing the original coordinate-dropping and adaptive-decomposition machinery and replacing it with Bayesian estimation in a Cauchy channel.","The set-level equivalence $b(T)\\asymp R(T)\\asymp\\Lambda(T)$ means Bernoulli width can be bounded from below by computable rate–distortion areas, giving a new sufficient condition for lower bounds in empirical-process theory.","The Cauchy information–estimation inequality fills the role of the Gaussian I–MMSE identity, so the Gaussian/Bernoulli dictionary used here becomes a template for treating other additive-noise settings.","Because the two-sided comparison holds per law, any upper bound on the rate–distortion integral for a given $\\mu$ immediately gives an upper bound on $B(\\mu)$, and any lower bound gives a lower bound on $b(T)$ after supremizing."],"supporting_citations":[{"why":"The Bernoulli theorem whose lower bound is reproved; defines the target functional $\\Lambda(T)$ and the classical equivalence $b(T)\\asymp\\Lambda(T)$.","marker":"[BL14]"},{"why":"Formulated the Bernoulli conjecture and introduced the capped quadratic distortion that the proof uses as its scale-dependent loss.","marker":"[Tal94]"},{"why":"Supplies the upper half of the Gaussian majorizing-measure theorem used to convert Gaussian width into $\\gamma_2$ control.","marker":"[Fer75]"},{"why":"Supplies the Gaussian majorizing-measure theorem and the fixed-law Gaussian characterization used in the proof of Theorem 2.3.","marker":"[Tal87]"},{"why":"Supplies the type-lifting identity invoked in Lemma 6.1 (Eq. 51) and the Gaussian rate–distortion formulation being adapted to Bernoulli processes.","marker":"[Liu25]"},{"why":"Provides the other source of the type-lifting identity, specifically its Proposition 3.1 for exact type classes, used for the lower-bound direction.","marker":"[vH25]"},{"why":"The Bayesian proof of the Gaussian majorizing-measure theorem whose proof structure the present paper follows.","marker":"[Zad26]"},{"why":"The Gaussian I–MMSE identity whose role is replaced, in the Bernoulli setting, by the new Cauchy information–estimation inequality.","marker":"[GSV05]"},{"why":"Raises the question answered by the Cauchy information–estimation inequality, motivating the derivation in Appendix A.","marker":"[Ver23]"}],"fun_headline_variants":["Rate–distortion functional pins Bernoulli supremum for every law","Bayesian proof ties Bernoulli supremum to rate–distortion integral","New characterization: Bernoulli supremum via Bayesian estimation","Cauchy channel reveals Bernoulli supremum law"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise that is not proved in this manuscript is the type-lifting identity: for a given index law $\\mu$, the prescribed-law value $B(\\mu)$ is assumed equal to the limit, as $N\\to\\infty$, of the normalized Bernoulli width of the set of $N$-tuples whose empirical distribution is exactly $\\mu$; the paper cites this from existing work (Lemma 6.1, Eq. 51), and if that identity failed, the lower-bound direction of the main theorem would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Rate–distortion functional pins Bernoulli supremum for every law","Bayesian proof ties Bernoulli supremum to rate–distortion integral","New characterization: Bernoulli supremum via Bayesian estimation","Cauchy channel reveals Bernoulli supremum law"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000899,"raw_usage":{"total_tokens":3876,"prompt_tokens":954,"completion_tokens":2922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":2857}},"tokens_in":570,"tokens_out":2922,"duration_ms":20557,"temperature":1.0,"reasoning_tokens":2857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:15:29.647815+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check on the riskiest step: take a finite $T$ and a non-rational law $\\mu$ (for example, weights proportional to $\\sqrt{2}$ on three points), compute $B(\\mu)$ by direct coupling optimization, compute the limit $\\lim_{N\\to\\infty} b(\\mathcal{T}_N(\\mu))/N$ for the exact type classes $\\mathcal{T}_N(\\mu)$, and compare; a mismatch would contradict Eq. (51). For the theorem itself, one could compute the ratio of $B(\\mu)$ to $\\int_0^\\infty RD_\\mu(t)\\,dt$ on a sequence of sets with growing dimension, such as the vertices of a cube or the $\\ell^1$ ball; if that ratio tends to $0$ or $\\infty$, the universal-constant claim fails.","supporting_citations":[],"review_version":2}