{"id":"ffa6fac2-561b-45b9-b8cf-721eea514e80","arxiv_id":"1908.05713","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For Gaussian multiterminal coding at high resolution, the rate gap to centralized coding is conjectured to equal half the sum of squared inverse-covariance off-diagonal terms over pairs never co-encoded, times d squared, and this is verified for up to three sources.","lead":"The paper proposes a simple formula for the extra rate a distributed Gaussian source coding system needs compared to a centralized one when distortion is small. It proves the formula for systems with up to three sources, leaving the general case as a conjecture.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Displayed rate formulas invert the 1+sqrt factor, yielding a negative gap for independent sources and contradicting the paper's own expansions; the proof as written needs correction.","rationale":"The reader's weakest_assumption concerns the asymptotic structure of the SDP optimizer D* in Lemmas 2 and 3. That structure is asserted tersely but appears derivable from the stated inequalities (10) and the Neumann series (11); I do not find a concrete flaw there. The load-bearing problem I found is more elementary and more visible: the exact rate formulas in Lemma 1 and Lemma 4, as printed, have the (1 + sqrt(...)) factor in the denominator, which makes the independent-source case violate the basic inequality r_S ≥ r_C and makes Eq. (23) algebraically inconsistent with the displayed formula. This is not a matter of an unproved asymptotic; it is a direct internal contradiction. The theorem's final expressions match the conjecture, and the intended formulas are clear from the expansions, so the result is likely correct after fixing the typographical inversion. But as submitted, a reader cannot derive the proof's stated inequalities from the stated lemmas. The verdict should remain CONDITIONAL, requiring correction of the displayed formulas and a re-check of the derived asymptotics. The reader's stated concern about D* is legitimate but secondary; the manuscript's own inconsistent formulas are the first barrier to verification. My proposed test directly exposes the contradiction and would be a quick check for the authors or referees.","tokens_in":14440,"tokens_out":36070,"duration_ms":264152,"concrete_test":"Set θ = 0 and d1 = d2 = d in Lemma 1's displayed formula and compute r - r_C. The formula yields -1/2 log 4, contradicting Eq. (3)'s requirement r_S ≥ r_C. Independently derive the sum rate for two independent Gaussian sources from the cited result [2]: it must equal r_C, confirming that the displayed formula should have (1 + sqrt(1 + 4θ^2 d^2)) in the numerator, not the denominator. Recompute Eq. (23) from Lemma 4's displayed r(d1,d2,d3) with that corrected placement and verify that the expansion + (1/2) θ_{2,3}^2 d^2 + o(d^2) is reproduced exactly.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Lemma 1 and Lemma 4 display exact rate formulas of the form r = (1/2) log( det(Γ) / (2 d1 d2 ... (1 + sqrt(1 + 4θ^2 d1 d2 ...))) ), placing the factor (1 + sqrt(...)) in the denominator. For independent sources (θ = 0) and d1 = d2 = d, Lemma 1 therefore gives r = (1/2) log( det(Γ) / (4 d^2) ), which is (1/2) log 4 below the centralized rate r_C = (1/2) log( det(Γ) / d^2 ). This violates the fundamental inequality r_S(d) ≥ r_C(d) stated in Eq. (3). The subsequent algebra in Section III-A and in Lemma 4, e.g., Eq. (23), treats the factor as if it were in the numerator: Eq. (23) expands to + (1/2) θ^2 d^2 + o(d^2), which only follows if r(d1,d2,d2) = (1/2) log( det(Γ)·(1+sqrt(...)) / (2 d1 d2^2) ). The same inverted-factor pattern appears in the auxiliary expression \tilde r in Appendix B near Eq. (33). Thus the displayed exact characterizations are mutually inconsistent with the proof's own asymptotics and with known lower bounds. Since Theorem 1's verification relies on these formulas, the proof as written is not sound; it would need a systematic correction of the factor placement, after which the intended argument appears salvageable.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Conjecture 1: for any generalized Gaussian multiterminal source coding system in the high-resolution regime, the gap between the distributed sum rate r_S(d) and the centralized rate r_C(d) is asymptotically (1/2) times the sum over pairs (i,j) that never appear together in any encoder of theta_{i,j}^2 d^2, plus o(d^2). The main result, Theorem 1, verifies this conjecture for L <= 3. The proof reduces to five non-redundant covers up to relabeling, and uses exact semidefinite-programming characterizations from prior work (Wang-Chen-Wu, Wang-Chen, Oohama) together with high-resolution determinant expansions. The upper-bound directions are tied to published exact characterizations; the lower-bound tightness arguments in Lemmas 2 and 3 rely on asserted asymptotic structure of the optimal distortion covariance matrix D*.","tokens_in":14731,"tokens_out":6865,"duration_ms":64202,"significance":"If the result is correct, it gives a remarkably clean topology-dependent formula for the high-resolution behavior of a problem whose exact solution is generally not available in closed form, and it unifies the known two-terminal and three-terminal cases. The conjecture is falsifiable and the consistency checks with cover domination and equivalence are valuable. The proof strategy is sound in outline: it reduces to a small list of covers and leverages existing exact characterizations rather than assuming the conjecture. The paper is honest that the full conjecture is open for L > 3. However, the manuscript contains a load-bearing notational/sign error in the displayed exact rate formulas, and the lower-bound proofs in Lemmas 2 and 3 omit the key asymptotic derivation of the optimizer D*. Both issues are fixable, but they must be addressed before the proof can be considered complete.","major_comments":[{"comment":"The displayed exact rate formulas place the factor (1 + sqrt(1 + 4 theta^2 d1 d2)) in the denominator, i.e., r = (1/2) log( det(Gamma) / (2 d1 d2 (1 + sqrt(...))) ). For independent sources (theta=0) and d1=d2=d, this gives r = (1/2) log( det(Gamma) / (4 d^2) ), which is (1/2) log 4 below r_C(d) = (1/2) log( det(Gamma) / d^2 ) and violates the fundamental inequality r_S(d) >= r_C(d) in Eq. (3). The subsequent algebra in Section III-A and Eq. (23) treats the factor as if it were in the numerator, and only that version yields the claimed positive O(d^2) gap. The same inverted-factor pattern appears in the auxiliary expression r-tilde in Appendix B near Eq. (33). Please correct the factor placement systematically and re-derive the expansions; the intended asymptotic result is evidently recoverable.","section":"Section III-A, Lemma 1; Section III-D, Lemma 4; Appendix B near Eq. (33)"},{"comment":"The lower-bound tightness arguments are incomplete at a load-bearing point. The sentences 'It can be shown by leveraging (10) and (11) that xi*_{ell,ell}=d+o(d), ..., and d*_{i,j}=-theta_{i,j} d^2 + o(d^2)' assert exactly the asymptotic structure of the optimizer needed to obtain det(D*) = d^3 - (sum theta^2) d^5 + o(d^5). No derivation is displayed, and this asymptotic structure is essential for the lower bound. Please supply the missing argument or cite a specific lemma from a prior paper that proves it.","section":"Section III-B, after Eq. (14), and Section III-C, after Eq. (21)"}],"minor_comments":[{"comment":"The caption contains a typo: 'wi th with L sources' should be 'with L sources'.","section":"Fig. 1 caption"},{"comment":"The word 'wich' appears twice ('wich is contradictory'); it should be 'which'.","section":"Section III-B and III-C"},{"comment":"In the branch theta_{1,2} > 0 of the definition of U_{2,3}, the noise variable is written N_{1,2}; it should be N_{2,3}.","section":"Appendix C, definition of U_{2,3}"},{"comment":"The displayed closed-form optimizer for the convex problem is garbled: the variables d^2 and d1 appear to be misplaced in the fraction. Please rewrite the expression and verify that it leads to the stated asymptotics (24) and (25).","section":"Section III-D, after Eq. (23)"},{"comment":"The statements 'For any d sufficiently close to 0, we can choose alpha_ell such that ...' in the proofs of Lemmas 4 and 5 would benefit from a brief continuity or monotonicity argument showing that the desired distortion values are simultaneously attainable.","section":"Section III-D and III-E"}],"recommendation":"major_revision","confidential_remarks":"The central result is likely correct and the proof strategy is plausible, but the displayed-rate sign error is a serious inconsistency that must be corrected, and the lower-bound asymptotic structure of D* in Lemmas 2 and 3 must be proved rather than asserted. The paper leans heavily on the authors' own prior exact characterizations; this is legitimate but means the incremental novelty is moderate. I would not reject on that basis if the technical gaps are fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Xie, Tu, Zhou, and Chen propose a clean conjectural expression for the high-resolution gap between the distributed and centralized rate-distortion functions in generalized Gaussian multiterminal source coding, and verify it for up to three sources. The conjecture — the gap is half the sum of squared off-diagonal precision entries for pairs never co-observed — is simple and organizes the L=2 and L=3 cases under one formula. That is genuinely new, even if the proof is a case-by-case application of existing exact characterizations.\n\nThe stress-test note about inverted factors does not hold up on reading the paper. The displayed formulas in Lemmas 1 and 4 put the (1+sqrt(...)) factor in the numerator, not the denominator. The paper's own algebra in Section III-A and around Eq. (23) is consistent with that placement: for θ=0 the formulas reduce exactly to the centralized rate, and for small θ the expansion gives + (1/2)θ^2 d^2, as the conjecture requires. So there is no violation of r_S ≥ r_C here.\n\nThe real soft spot is the lower-bound tightness arguments in Lemmas 2 and 3. For the distributed three-source case and the {12},{3} case, the paper asserts that the optimizer D* of the exact SDP characterization has diagonal entries d+o(d) and off-diagonal entries -θ d^2+o(d^2), then uses that to derive the lower bound. The assertion is plausible and probably follows from (10) and (11), but the derivation is not shown. A referee should ask for it. This is an exposition gap, not an obvious error, but it is the step that carries the lower-bound proof.\n\nThe paper leans on earlier exact characterizations from the authors' own group (Wang and Chen, 2013–2014). That is legitimate — those are published in IEEE Trans. IT — and the paper does not try to rederive them. The reliance is worth noting but is not a flaw.\n\nOverall: a solid, honestly scoped contribution. The conjecture is interesting and the L≤3 verification is a useful data point. I'd send it to peer review with the request that the authors expand the lower-bound derivations. If those hold up, it should be accepted as a regular subfield paper.","headline":"A clean conjecture for the high-resolution gap in Gaussian multiterminal source coding, verified for L≤3, with a proof sound in outline but with a sketched lower-bound step; the stress-test factor-inversion concern is a misreading.","tokens_in":15253,"tokens_out":6067,"would_cite":true,"duration_ms":48513,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A34","94A15","90C22"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single conjectured formula characterizes the high-resolution rate gap in generalized Gaussian multiterminal source coding, and the paper proves it for up to three sources.","keywords":["generalized Gaussian multiterminal source coding","high-resolution regime","rate-distortion function","precision matrix","asymptotic gap","centralized coding","distributed coding","semidefinite programming"],"falsifier":"Take the four-source fully distributed system $\\mathcal{S}=\\{\\{1\\},\\{2\\},\\{3\\},\\{4\\}\\}$ with a covariance matrix whose precision matrix has nonzero off-diagonal entries. Solve the sum-rate SDP from [12] at decreasing distortion $d$ and compare $r_{\\mathcal{S}}(d)-r_{\\mathcal{C}}(d)$ with $\\frac{1}{2}\\sum_{1\\le i<j\\le 4}\\theta_{i,j}^2d^2$; a coefficient mismatch would refute Conjecture 1. For a direct check of the three-source proof, compute the SDP minimizer $D^*$ in Lemma 2 for a concrete $\\Gamma$ with large $\\theta_{1,2}$ and verify whether $d^*_{1,2}=-\\theta_{1,2}d^2+o(d^2)$ and $d^*_{\\ell,\\ell}=d+o(d)$.","tokens_in":14232,"feed_emoji":"📡","tokens_out":12862,"duration_ms":99813,"temperature":0.7,"pith_summary":"This paper asks how much more rate is needed when $L$ jointly Gaussian sources are compressed by separate encoders that each see only a subset of the sources, compared with a single encoder that sees everything. In the high-resolution regime, as allowed distortion $d$ tends to zero, it conjectures a sharp closed-form answer: the rate gap equals $\\frac{1}{2}d^2$ times the sum of squared off-diagonal entries $\\theta_{i,j}^2$ of the precision matrix $\\Theta$, summed over pairs $(i,j)$ that no encoder observes together, up to $o(d^2)$. The paper proves this formula for every system with at most three sources, by checking the five essentially different encoder topologies. If the conjecture holds generally, it turns a problem normally expressed only through semidefinite programs into a simple calculation depending on the encoder graph and the precision matrix.","feed_headline":"Multiterminal coding's rate gap is a sum of squared precision entries","feed_subtitle":"Verified for all topologies up to three sources, the gap grows as d squared times half the squared precision entries.","key_machinery":"The argument rests on exact semidefinite-programming characterizations of the sum rate for Gaussian multiterminal coding, obtained in earlier work for the relevant topologies. The load-bearing asymptotic identity is the expansion of the error-covariance matrix $D=(\\Theta+\\Xi^{-1})^{-1}$ when the auxiliary matrix $\\Xi$ is small: $D=\\Xi-\\Xi\\Theta\\Xi+o(d^2)$ in the relevant entries. Setting the diagonal of $D$ to $d+o(d)$ forces the off-diagonal entries to $-\\theta_{i,j}d^2+o(d^2)$, giving $\\det D=d^L-\\bigl(\\sum_{(i,j)\\in E(\\mathcal{S})}\\theta_{i,j}^2\\bigr)d^{L+2}+o(d^{L+2})$. Since the centralized rate is $r_{\\mathcal{C}}(d)=\\frac{1}{2}\\log(\\det\\Gamma/d^L)$, the logarithm of this determinant produces exactly the conjectured half-times-squares gap.","core_discovery":"The central claim is Conjecture 1: for any cover $\\mathcal{S}$ of $\\{1,\\ldots,L\\}$, $r_{\\mathcal{S}}(d)-r_{\\mathcal{C}}(d)=\\frac{1}{2}\\sum_{(i,j)\\in E(\\mathcal{S})}\\theta_{i,j}^2d^2+o(d^2)$, where $E(\\mathcal{S})$ is the set of source pairs never contained together in any encoder's subset. Theorem 1 establishes this for $L\\le 3$. The proof covers all five non-redundant covers up to relabeling: the two-source distributed system, the three-source fully distributed system, and the three-source systems with encoder subsets $\\{\\{1,2\\},\\{3\\}\\}$, $\\{\\{1,2\\},\\{1,3\\}\\}$, and $\\{\\{1,2\\},\\{1,3\\},\\{2,3\\}\\}$. In the last case the gap is actually $o(d^2)$, so the fully paired three-encoder system matches centralized coding through second order.","pith_inferences":["If the conjecture holds for all $L$, the second-order penalty is governed only by conditional dependencies between sources that no encoder sees together: pairs with $\\theta_{i,j}=0$, which are conditionally independent given the other sources, contribute nothing, tying the formula to Gaussian graphical models.","The same $d^2$ scaling suggests that the next-order $o(d^2)$ term may admit a systematic expansion in terms of triples of sources or cycles in the encoder hypergraph; the exact $L=3$ formulas in the paper could calibrate such an expansion before tackling $L=4$.","A numerical test for $L=4$ is immediately available: the SDP characterizations extend to arbitrary covers, so one can compute $r_{\\mathcal{S}}(d)$ at small $d$ for randomly chosen precision matrices and compare the fitted quadratic coefficient with the conjectured expression."],"forward_implications":["For $L\\le 3$, the high-resolution rate-distortion function of every generalized Gaussian multiterminal system is explicitly $r_{\\mathcal{C}}(d)+\\frac{1}{2}\\sum_{(i,j)\\in E(\\mathcal{S})}\\theta_{i,j}^2d^2+o(d^2)$, with no hidden dependence on the topology beyond the set $E(\\mathcal{S})$.","Systems whose encoder subsets cover every pair of sources pay no second-order penalty: the gap is $o(d^2)$, and for the three-source fully paired system the paper shows the rates are equal for all small $d$.","The gap depends only on the squared entries of the precision matrix, not on the individual variances $\\gamma_{i,i}$ or on the signs of the correlations.","Because the sum of $\\theta_{i,j}^2$ over $E(\\mathcal{S})$ decreases when a cover is refined, the formula reproduces the known ordering $r_{\\mathcal{C}}\\le r_{\\mathcal{S}}\\le r_{\\mathcal{S}'}$ for dominating covers $\\mathcal{S}'$.","A full proof for arbitrary $L$ would reduce high-resolution multiterminal coding to computing a graph-theoretic quantity from the encoder hypergraph and the precision matrix."],"supporting_citations":[{"why":"Supplies the reverse water-filling formula for the centralized rate $r_{\\mathcal{C}}(d)$, the baseline in the gap.","marker":"[1]"},{"why":"Supplies the exact rate region for the two-encoder quadratic Gaussian problem used for the $L=2$ case and for the converse bound in the $\\{\\{1,2\\},\\{1,3\\}\\}$ case.","marker":"[2]"},{"why":"Provides the semidefinite-programming characterization of the sum rate from which Lemma 2 is deduced.","marker":"[12]"},{"why":"Underlies the SDP characterization used for the $\\{\\{1,2\\},\\{3\\}\\}$ system in Lemma 3.","marker":"[15]"},{"why":"Supplies the vector multiterminal SDP results (Theorems 8 and 9) used to construct the upper-bound matrices in Lemmas 2 and 3.","marker":"[16]"},{"why":"Establishes that the first-order gap vanishes, motivating the second-order expansion pursued here.","marker":"[21]"},{"why":"The Berger scheme is one of the two sources of the achievability bound used in Appendices B and C.","marker":"[23]"},{"why":"The Tung scheme completes the Berger-Tung achievability argument used in Appendices B and C.","marker":"[24]"}],"fun_headline_variants":["Gap = half sum of squared precisions, proven for three sources","Fully paired three encoders match centralized rate to second order","Three-source proof confirms gap formula: sum of squared precisions","Multiterminal gap: sum of squared precisions for unpaired source pairs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"In the lower-bound halves of the three-source proofs, the optimizer of the semidefinite program is assumed to have error-covariance diagonal entries $d+o(d)$ and off-diagonal entries $-\\theta_{i,j}d^2+o(d^2)$; if some positive-definite covariance produced an optimizer with a different asymptotic shape, the claimed second-order gap would not follow from these arguments.","fun_headline_variants_meta":{"raw":{"variants":["Gap = half sum of squared precisions, proven for three sources","Fully paired three encoders match centralized rate to second order","Three-source proof confirms gap formula: sum of squared precisions","Multiterminal gap: sum of squared precisions for unpaired source pairs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0018,"raw_usage":{"total_tokens":7017,"prompt_tokens":799,"completion_tokens":6218,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":415,"completion_tokens_details":{"reasoning_tokens":6141}},"tokens_in":415,"tokens_out":6218,"duration_ms":46693,"temperature":1.0,"reasoning_tokens":6141,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:05:45.212285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the four-source fully distributed system $\\mathcal{S}=\\{\\{1\\},\\{2\\},\\{3\\},\\{4\\}\\}$ with a covariance matrix whose precision matrix has nonzero off-diagonal entries. Solve the sum-rate SDP from [12] at decreasing distortion $d$ and compare $r_{\\mathcal{S}}(d)-r_{\\mathcal{C}}(d)$ with $\\frac{1}{2}\\sum_{1\\le i<j\\le 4}\\theta_{i,j}^2d^2$; a coefficient mismatch would refute Conjecture 1. For a direct check of the three-source proof, compute the SDP minimizer $D^*$ in Lemma 2 for a concrete $\\Gamma$ with large $\\theta_{1,2}$ and verify whether $d^*_{1,2}=-\\theta_{1,2}d^2+o(d^2)$ and $d^*_{\\ell,\\ell}=d+o(d)$.","supporting_citations":[{"cited_title":"Cover and J","cited_arxiv_id":null,"evidence_quote":"Supplies the reverse water-filling formula for the centralized rate $r_{\\mathcal{C}}(d)$, the baseline in the gap."},{"cited_title":"Rate region of the quadratic Gaussian two-encoder source-coding problem,","cited_arxiv_id":null,"evidence_quote":"Supplies the exact rate region for the two-encoder quadratic Gaussian problem used for the $L=2$ case and for the converse bound in the $\\{\\{1,2\\},\\{1,3\\}\\}$ case."},{"cited_title":"On the sum rate of Gaussian mu ltiter- minal source coding: New proofs and results,","cited_arxiv_id":null,"evidence_quote":"Provides the semidefinite-programming characterization of the sum rate from which Lemma 2 is deduced."},{"cited_title":"V ector Gaussian two-terminal sour ce coding,","cited_arxiv_id":null,"evidence_quote":"Underlies the SDP characterization used for the $\\{\\{1,2\\},\\{3\\}\\}$ system in Lemma 3."},{"cited_title":"V ector Gaussian multiterminal sou rce coding,","cited_arxiv_id":null,"evidence_quote":"Supplies the vector multiterminal SDP results (Theorems 8 and 9) used to construct the upper-bound matrices in Lemmas 2 and 3."},{"cited_title":"Multiterminal source coding wi th high reso- lution,","cited_arxiv_id":null,"evidence_quote":"Establishes that the first-order gap vanishes, motivating the second-order expansion pursued here."},{"cited_title":"Multiterminal source coding,","cited_arxiv_id":null,"evidence_quote":"The Berger scheme is one of the two sources of the achievability bound used in Appendices B and C."},{"cited_title":"Multiterminal source coding,","cited_arxiv_id":null,"evidence_quote":"The Tung scheme completes the Berger-Tung achievability argument used in Appendices B and C."}],"review_version":1}