{"id":"c517bc3d-5586-431f-ad4a-7ab549a33dcc","arxiv_id":"2501.13784","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"For sources that are conditionally independent given the receiver's side information, the paper claims an exact rate-distortion region for reconstructing a latent variable, but the proofs are deferred and the numerical method is validated only in a single-encoder case.","lead":"Multiple devices each compress what they observe and send it to a central receiver that also has side information and needs to reconstruct a hidden variable. This paper states a compact formula for the optimal compression trade-off in a special case, but leaves the key proof steps and most numerical validation to a later version.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed equality Ra(D)=Ro(D)=R*(D) rests on an unproved transfer of Wagner's direct outer bound and Gastpar's achievable scheme to a latent distortion variable T; the paper omits the proofs of Lemma 1, Lemma 2, Theorem 3, and Corollary 4.","rationale":"The central claim of the paper is an exact single-letter rate-distortion region. Its proof is entirely delegated to known results for direct distributed source coding. The reader's weakest_assumption identifies exactly this: the transfer from direct reconstruction of X_1,...,X_M to indirect reconstruction of a latent T with distortion d(T,\\hat T) is asserted, not proved. I agree that this is the most load-bearing concern. The theorem may be true, and the formulas are plausible, but an exact characterization cannot be accepted on the strength of 'easily established based on Wagner's proof' when the problem has a different distortion variable and a different Markov structure. Corollary 4's simplification from subset-sum constraints to individual constraints is a second unproved step in the same chain: it requires the auxiliaries to be conditionally independent given Y, which is plausible for the Gastpar inner bound but not shown for the outer-bound auxiliaries. The numerical algorithm issues (non-convexity admitted in the convergence analysis, equal-weight Lagrangian rather than a general weighted sum) are real but secondary; even a perfect algorithm would not supply the missing coding theorem. Therefore the reader's REJECT verdict is appropriate: the manuscript's main contribution is not established. I would keep the verdict unchanged.","tokens_in":8182,"tokens_out":29654,"duration_ms":298379,"concrete_test":"Independently re-derive Lemma 2 from Wagner [15] for the indirect setting: write out the full single-letterization for M=2 with W_i=(M_i,J,...) and a decoder g(W_1,W_2,Y), verifying that (i) R_i ≥ I(X_i;W_i|Y) and (ii) E[d(T,g(W_1,W_2,Y))]≤D hold simultaneously. Then prove Corollary 4 by showing the W_i from the outer bound can be replaced by auxiliaries satisfying conditional independence given Y without increasing rates. If either step fails, Theorem 3 is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 1's proof is omitted ('we omit the detailed proof here due to space limitations'); Lemma 2 and Theorem 3 are delegated to Wagner [15] ('can be easily established'); Corollary 4 is deferred to a longer version. This is the load-bearing point because [15] and [18] are direct-coding theorems: their rate inequalities are derived for reconstructing the observed sources X_i under per-component distortion constraints. The present problem replaces those constraints by a single constraint E[d(T,g(W,Y))]≤D on a latent variable T that no encoder observes. It is not automatic that the same auxiliary variables W_i that satisfy the rate lower bounds also satisfy the single-letter distortion constraint, nor that the outer-bound auxiliaries can be chosen conditionally independent given Y so that the subset-sum conditions in (2) collapse to the individual conditions (7). If either step fails, the exact region could contain additional terms or require structured (Körner–Marton type) coding, and the claimed separation into M independent Wyner-Ziv constraints would be false. The numerical section cannot substitute for this: its only external validation is the M=1 Wyner-Ziv case, which does not exercise the coupling among encoders.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper considers a distributed source coding problem in which M encoders observe correlated sources X_1,...,X_M, a decoder has side information Y, and the goal is to reconstruct a latent variable T under a distortion constraint. The main theoretical claim (Theorem 3) is that, when the sources are conditionally independent given Y, the rate-distortion region is given by individual Wyner-Ziv inequalities R_i ≥ I(X_i;W_i) − I(W_i;Y). The paper also proposes an iterative Blahut-Arimoto-style algorithm to compute the region and presents numerical examples for M=2 binary sources. The abstract and introduction frame the problem in the context of task-oriented semantic communication and distributed learning.","tokens_in":8404,"tokens_out":5153,"duration_ms":45003,"significance":"An exact single-letter characterization for this indirect multiterminal source coding problem would be a useful contribution, particularly for task-oriented communication. The numerical algorithm, if properly justified, could provide a practical tool for evaluating such regions. However, as submitted, the central theorem is asserted without proof: the proofs of Lemma 1, Lemma 2, Theorem 3, and Corollary 4 are either omitted, deferred, or dismissed as 'easily established' from prior works. The convergence analysis of the algorithm is internally contradictory, and the numerical validation only covers the M=1 case. Hence the paper currently offers a conjecture and a heuristic algorithm rather than a verifiable information-theoretic result.","major_comments":[{"comment":"The central theorem is not proved. Lemma 1's proof is omitted with 'we omit the detailed proof here due to space limitations'; Lemma 2 and Theorem 3 are dismissed with 'can be easily established based on Wagner's proof'; Corollary 4's proof is deferred to a longer version. This is load-bearing because the cited results [15] and [18] are for direct reconstruction of the observed sources, whereas the present problem imposes a single distortion constraint on a latent variable T that no encoder observes. It is not automatic that the auxiliary variables that satisfy the rate lower bounds also satisfy the distortion constraint, nor that the outer-bound auxiliaries can be chosen conditionally independent given Y so that the subset-sum conditions in (2) collapse to the individual conditions (7). Without these steps, the claimed exact region Ra(D)=Ro(D)=R*(D) is unsupported.","section":"Section II, Lemma 1, Lemma 2, Theorem 3, Corollary 4"},{"comment":"The convergence argument is internally contradictory. The text first claims that the Lagrangian is convex and therefore the proposed iterative optimization framework achieves the global minimum, then immediately concedes that the expected distortion term includes a product of variables and 'the Lagrangian may exhibit non-convex behavior'. This contradiction means the algorithm's convergence to the true rate-distortion function is not established; monotone non-increasing Lagrangian values alone do not imply global optimality in non-convex problems, and the appeal to [21] only supports 'highly effective inner bounds', not exact computation.","section":"Section III, Convergence analysis"},{"comment":"The numerical validation is insufficient. The only external comparison is the M=1 case (Fig. 4), which reduces to a point-to-point problem and does not exercise the coupling among encoders. The M=2 example with T={X1,X2} and sum-Hamming distortion is not compared against any known region, outer bound, or independent computation. Therefore the numerical section cannot support the claimed exact region for M>1.","section":"Section IV, numerical example"},{"comment":"The Lagrangian in (8) minimizes the unweighted sum of rates ∑_{i∈M}(I(W_i;X_i)−I(W_i;Y)) plus a distortion penalty. Since the rate-distortion region is a set of M-dimensional rate vectors, minimizing only the sum rate with a single Lagrange multiplier λ traces at best one boundary point per λ; to characterize the full region, a weighted sum with distinct per-encoder weights is generally required. The paper does not justify that equal weights suffice, and in asymmetric problems they generally do not, so the proposed algorithm does not compute the entire rate-distortion region as claimed.","section":"Section III, Eq. (8)"}],"minor_comments":[{"comment":"The statement says 'for all l ∈ {1,...,M}', but the variable is i; this should be corrected to 'for all i ∈ {1,...,M}'.","section":"Corollary 4"},{"comment":"There is a typo: 'considerd' should be 'considered'.","section":"Section I, last paragraph"},{"comment":"The definition of A^c as 'all elements in the set A ⊆ {1,...,M} that are not in A' is circular; it should be the complement of A in {1,...,M}.","section":"Equation (2)"},{"comment":"The notation in (13) is confusing: the Lagrangian is written as Lλ(Q, q*_{\\m}, q'), but the right-hand side sums over all i∈M, mixing subscripts. The expression should be clarified.","section":"Lemma 6, Eq. (13)"},{"comment":"The paper states that when M=1 the problem reduces to the traditional point-to-point Wyner-Ziv problem. This is only true if T equals the observed source X1; in general it is the remote Wyner-Ziv problem, so the comparison in Fig. 4 needs the condition T=X1 to be stated.","section":"Section IV, M=1 claim"},{"comment":"Reference [20] is malformed: 'C. Q., T. M. Cover, and J. A. Thomas' should be 'T. M. Cover and J. A. Thomas'.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript is at a very preliminary stage: the core proofs are explicitly omitted or deferred, and the convergence analysis contains an acknowledged contradiction. The claimed exact region may be correct, but as submitted the paper does not meet the standard for a journal publication. The authors should be encouraged to provide complete proofs of Lemma 1, Lemma 2, Theorem 3, and Corollary 4, and to clarify the convergence guarantee of the algorithm, before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this paper states an exact rate-distortion region for distributed indirect source coding with decoder side information under conditional independence, but the central theorem is never proved in the manuscript. The proofs of Lemma 1, Lemma 2, Theorem 3, and Corollary 4 are either omitted, deferred, or dismissed as “easily established” from Wagner’s outer bound and Gastpar’s achievable scheme. Those prior results are for direct reconstruction of the observed sources, and the transfer to a latent variable T that no encoder observes is not automatic. That is the load-bearing gap.\n\nWhat is genuinely new: the problem formulation itself — M encoders observing correlated sources, decoder wanting a latent T under distortion, with side information Y — is a useful generalization, and the simplification to per-encoder Wyner-Ziv rates in Corollary 4 is a clean and plausible benchmark. The Lagrangian updates in Lemmas 5–7 are carefully derived, and the M=1 Wyner-Ziv check in Fig. 4 is a good sanity test.\n\nThe soft spots, in proportion: (1) The missing proof is not a matter of space; it is the whole result. Wagner’s and Gastpar’s inequalities are derived for per-component distortion on the sources, and it is not shown that the same auxiliary variables satisfy the distortion constraint on T, nor that the outer-bound auxiliaries can be chosen conditionally independent given Y so the subset sums collapse. If those steps fail, the exact region could require structured (Körner–Marton type) coding. (2) The convergence analysis first claims convexity and a global minimum, then admits the Lagrangian may be non-convex and falls back to “highly effective inner bounds.” That is a self-contradiction, and the two-source example does not resolve it. (3) The numerical example is not really indirect: T = {X1, X2} with additive Hamming distortion is direct reconstruction of the pair. The only external validation is M=1, which does not exercise encoder coupling. (4) Minor: the algorithm appears to minimize sum rate rather than a weighted sum, so it may not trace the whole rate-distortion boundary.\n\nBottom line: the result is plausible, and the paper is worth a serious referee, but it should not be accepted as is. A referee should demand the complete proof of Theorem 3 and its corollaries, plus a numerical example that actually exercises the latent-variable setting. I would not cite it in its current form, but I would bring it to a reading group to debate whether the direct-to-indirect transfer goes through.\n\nRecommendation: send it to peer review, expecting major revision.","headline":"Plausible but unproved: the exact rate-distortion region for distributed indirect source coding rests on omitted proofs, and the numerical example never exercises the indirect setting.","tokens_in":8938,"tokens_out":2763,"would_cite":false,"duration_ms":23597,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A15","94A34","94A29"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims an exact rate-distortion region for distributed indirect source coding when the sources are conditionally independent given decoder side information.","keywords":["distributed source coding","indirect source coding","rate-distortion region","decoder side information","conditional independence","alternating minimization","semantic communication","remote reconstruction"],"falsifier":"Take a small joint distribution with X_1 and X_2 conditionally independent given Y, set T to a nonlinear function such as XOR, and compute the true minimal sum-rate by exhaustive search over all encoder and decoder maps for a distortion target; if any achievable point lies below the sum of the two individual claimed bounds, or if the rectangular claimed region contains points that exhaustive search shows are not achievable, the claimed exact region fails.","tokens_in":7918,"feed_emoji":"📡","tokens_out":7086,"duration_ms":61022,"temperature":0.7,"pith_summary":"The paper claims an exact characterization of the rate-distortion region for a distributed source coding problem in which M encoders each observe one correlated source, a central decoder also has side information Y, and the goal is to reconstruct a latent variable T within a distortion constraint rather than the sources themselves. For the case where the sources are conditionally independent given Y, the paper shows the achievable region and an outer bound coincide, giving per-encoder constraints R_i ≥ I(X_i;W_i) − I(W_i;Y) for auxiliary variables W_i, with a decoder that meets the distortion target. This matters because task-oriented semantic communication and distributed learning systems often need to send only enough information about local observations to recover a remote task variable. The paper also gives an iterative alternating-minimization algorithm to compute the region numerically, which it validates by matching the classical single-encoder side-information curve in the M=1 case.","feed_headline":"Exact rates found for distributed indirect source coding","feed_subtitle":"Rates separate into independent per-encoder constraints when observations share side information.","key_machinery":"The load-bearing structure is the conditional independence X_1,...,X_M ⊥ given Y, which forces the mutual-information cost to decouple across encoders. The argument runs through auxiliary random variables W_i introduced by test channels q_i(w_i|x_i); the rate of encoder i is measured by I(X_i;W_i) − I(W_i;Y), the classical single-encoder side-information expression, and the decoder is a mapping q'(t|y,w) chosen as a Bayes detector to minimize expected distortion. The optimization is carried out on the Lagrangian L_λ = Σ_i [I(W_i;X_i) − I(W_i;Y)] + λ(E[d(T, ˆT)] − D), and the algorithm alternates updates of the back-channel Q_i(w_i|y), the encoder test channel q_i(w_i|x_i), and the decoder q'(t|y,w), producing a monotonically non-increasing Lagrangian. The same Lagrangian is the object used to prove the region.","core_discovery":"The central claim is that when X_1,...,X_M are conditionally independent given the side information Y, the true rate-distortion region R*(D) for reconstructing an indirect source T is exactly the set of rate tuples satisfying R_i ≥ I(X_i;W_i) − I(W_i;Y) for every i, for some auxiliary variables W_i such that a decoding function g(W,Y) achieves average distortion at most D. The paper derives this by taking an achievable region built from independent per-encoder test channels and an outer bound adapted from a general multiterminal bound, and showing the two coincide under the conditional-independence assumption. A corollary reduces the region to M separate single-encoder side-information constraints, so the encoders do not need to coordinate beyond sharing the side-information statistics. The claimed region is stated for discrete memoryless sources with a single-letter distortion measure; the general case without conditional independence is left open.","pith_inferences":["If the proof transfer holds, a natural extension is to Gaussian sources with squared-error distortion, where the per-encoder bounds would likely become closed-form water-filling expressions; the paper does not derive this.","A concrete test of the claim would be to compute the region for a small binary example where the latent variable $T$ is a non-separable function of the sources and compare with exhaustive search over quantizers; a mismatch would indicate that the indirect distortion needs extra terms in the outer bound.","The alternating-minimization algorithm is only guaranteed to converge to the global minimum when the Lagrangian is convex; for non-convex cases the authors resort to random restarts, so the numerical region should be treated as an inner bound unless independently certified.","If confirmed, the result supplies a design rule for distributed learning systems: each client's compression overhead is exactly its individual side-information cost, independent of how many other clients are compressing correlated observations."],"forward_implications":["A rate tuple is achievable exactly when each encoder meets its individual bound $R_i \\ge I(X_i;W_i) - I(W_i;Y)$, so no coordination among encoders is needed beyond knowing the joint statistics.","The single-letter expression gives a computable benchmark for task-oriented compression: for any discrete joint distribution satisfying the conditional-independence condition, the boundary of the rate-distortion region can be traced by sweeping the Lagrange multiplier.","In the $M=1$ limit the region reduces to the classical remote side-information rate-distortion function, and the paper's numerical algorithm reproduces the analytic curve.","For general correlated sources (not conditionally independent given $Y$), the inner and outer regions need not coincide, so the exact region for that case remains open."],"supporting_citations":[{"why":"Defines the remote rate-distortion problem that this setup extends to multiple encoders.","marker":"[10]"},{"why":"Supplies the point-to-point noisy-observation reconstruction result used as a building block.","marker":"[11]"},{"why":"Gives the single-encoder side-information rate-distortion function that the M=1 case must reduce to.","marker":"[13]"},{"why":"Provides the general multiterminal outer bound from which Lemma 2 and Theorem 3 are claimed to follow.","marker":"[15]"},{"why":"Gives the CEO-problem computational framework and inner-bound approach that the alternating-minimization algorithm builds on.","marker":"[17]"},{"why":"Supplies the achievable region for multiple sources with decoder side information that Lemma 1 is stated to extend.","marker":"[18]"},{"why":"Provides the analytic single-encoder side-information curve used to validate the numerical algorithm in the M=1 comparison.","marker":"[22]"}],"fun_headline_variants":["Exact rate region for distributed indirect compression","When side information unlocks exact distributed coding rates","Distributed indirect source coding: exact region characterized","Conditionally independent sources: exact rate-distortion region","Decoder side info yields exact rates in indirect coding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the existing outer and inner bounds for reconstructing the sources directly can be transferred unchanged to reconstructing a latent variable T, even though the detailed proofs of that transfer are omitted here.","fun_headline_variants_meta":{"raw":{"variants":["Exact rate region for distributed indirect compression","When side information unlocks exact distributed coding rates","Distributed indirect source coding: exact region characterized","Conditionally independent sources: exact rate-distortion region","Decoder side info yields exact rates in indirect coding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000518,"raw_usage":{"total_tokens":2457,"prompt_tokens":838,"completion_tokens":1619,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":1549}},"tokens_in":454,"tokens_out":1619,"duration_ms":13667,"temperature":1.0,"reasoning_tokens":1549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:36:52.090893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a small joint distribution with X_1 and X_2 conditionally independent given Y, set T to a nonlinear function such as XOR, and compute the true minimal sum-rate by exhaustive search over all encoder and decoder maps for a distortion target; if any achievable point lies below the sum of the two individual claimed bounds, or if the rectangular claimed region contains points that exhaustive search shows are not achievable, the claimed exact region fails.","supporting_citations":[{"cited_title":"Information transmission with addi- tional noise,","cited_arxiv_id":null,"evidence_quote":"Defines the remote rate-distortion problem that this setup extends to multiple encoders."},{"cited_title":"An improved outer bound for multiterminal source coding,","cited_arxiv_id":null,"evidence_quote":"Provides the general multiterminal outer bound from which Lemma 2 and Theorem 3 are claimed to follow."},{"cited_title":"The Wyner-Ziv problem with multiple sources,","cited_arxiv_id":null,"evidence_quote":"Supplies the achievable region for multiple sources with decoder side information that Lemma 1 is stated to extend."},{"cited_title":"The rate-distortion function for source coding with side information at the decoder,","cited_arxiv_id":null,"evidence_quote":"Provides the analytic single-encoder side-information curve used to validate the numerical algorithm in the M=1 comparison."}],"review_version":1}