{"id":"7496298e-6e63-4cf9-a6c9-d9f6b19a32ff","arxiv_id":"2505.22438","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A synonym-set formulation of variational inference re-derives the rate-distortion-perception tradeoff and is demonstrated with a single progressive image codec.","lead":"The paper recasts perceptual image compression as a problem of encoding a whole set of visually similar 'synonym' images rather than an exact pixel copy, and derives a bitrate-distortion-perception tradeoff from that view. It then builds a single progressive codec that can produce many quality levels and reports quality comparable to existing learned codecs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SVI derivation assumes the very invariance it is supposed to create: Eq. (20) sets p_{tilde x_j|x_i}=p_{tilde x_j|x} for all xi in the ideal synset, and the Jensen steps in Eq. (21) are upper bounds treated as equivalences; without both, Theorem 3.3's triple tradeoff is not derived.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing flaw: Eq. (20) assumes the invariance that the optimization is supposed to produce. My independent reading confirms that this is not a minor technicality; it is the step that lets the proof replace a distribution over synset members with a single reconstruction distribution, and the subsequent Jensen inequalities are upper bounds, not equivalences. The progressive SIC codec and its experiments are a real engineering contribution, and the paper honestly reports its limitations, including the suboptimal detail sampling and the remaining gap to GAN-based methods. However, the experiments do not test the theorem's key identity; they only demonstrate that a particular loss of the postulated form can be optimized. Since the principal claimed contribution is the proof that perceptual image compression must follow the triple tradeoff and the first theoretical explanation of the divergence term's existence, and that proof does not go through as written, the reader's REJECT verdict is appropriate. No verdict change is needed.","tokens_in":30392,"tokens_out":6079,"duration_ms":71190,"concrete_test":"Implement the exact computation of Eq. (21) for a scalar source with a two-element ideal synset X={x1,x2} and a fixed non-ideal encoder: compute the left side of Eq. (21)(a) directly and compare it with the final weighted E-MSE plus E-KLD expression. Unless the encoder already satisfies Eq. (20) and the Jensen inequalities hold with equality, the two sides will differ, showing that Theorem 3.2's equivalence is false for general non-ideal encoders and that Theorem 3.3 is unproven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a minimum-rate theorem, but its proof rests on Lemma 3.2, whose derivation uses the optimization target as a premise. In Appendix A.1, Eq. (20) asserts p_{tilde x_j|x_i}(tilde x_j|x_i)=p_{tilde x_j|tilde y_s}(tilde x_j|tilde y_s)=p_{tilde x_j|x}(tilde x_j|x) for every xi in the ideal synset, justified by an ideal encoder phi*_g that maps all synset members to the same synonymous representation. Yet the derivation is meant to produce the training objective for a non-ideal encoder; before convergence, p_{tilde x_j|x_i} and p_{tilde x_j|x} are generically different, so the equality fails exactly in the regime where the proof is applied. If Eq. (20) is instead granted as a definition of the ideal limit, the argument is circular: this invariance is precisely what the distortion/E-KLD minimization is supposed to create (Figure 1). Separately, steps (d) and (e) of Eq. (21) apply Jensen's inequality to obtain upper bounds and then treat the minimized upper bound as equivalent to the original negative log-likelihood; no equality condition is established for the scale factors alpha_d and alpha_p. Thus Eq. (17)'s equivalence between the synonymous likelihood and the weighted E-MSE plus E-KLD objective is not proven, and Theorem 3.3's constrained rate minimization lacks its main ingredient.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a semantic-information-theoretic analysis of perceptual image compression. It defines an 'ideal synset' of perceptually synonymous images, introduces synonymous variational inference (SVI), and states Lemma 3.2 and Theorem 3.3, claiming that the minimum achievable rate of perceptual compression follows a 'synonymous rate-distortion-perception tradeoff' that contains the existing rate-distortion-perception tradeoff as a special case. It then presents synonymous image compression (SIC), a progressive codec that encodes only the 'synonymous' latent part and samples the remaining detail latents, with experiments comparing DISTS, PSNR, LPIPS, and FID against HiFiC and MS-ILLM. The appendix contains the detailed proofs of the two central results.","tokens_in":30792,"tokens_out":9611,"duration_ms":99954,"significance":"The paper is unusually detailed in its implementation documentation, providing architecture choices, hyperparameter tables, training schedules, and additional GAN fine-tuning experiments, and it is transparent about the preliminary nature of the experimental results. If the theoretical claims were correct, the work would offer a unified semantic-information explanation for the appearance of distribution-divergence terms in perceptual compression and would introduce a practical single-model variable-rate codec. However, the central proof contains load-bearing gaps: the key invariance assumption is circular, Jensen's inequality steps are treated as equivalences, the E-KLD term is identified with a constant rather than derived, and the rate identity is not justified. The main theorem is therefore not established, and the claimed theoretical significance is not currently realized.","major_comments":[{"comment":"The proof of Lemma 3.2 assumes p_{\\tilde x_j|x_i}(\\tilde x_j|x_i)=p_{\\tilde x_j|x}(\\tilde x_j|x) for every x_i in the ideal synset, justified by an ideal encoder \\phi^*_g. This invariance is precisely what the training objective is supposed to learn (Figure 1); before convergence the equality is generically false. Using it as a premise to derive the training objective is therefore circular. Without Eq. (20), the subsequent Bayes manipulation leading to Eq. (21) cannot be performed, so Lemma 3.2 is unproved.","section":"A.1, Eq. (20)"},{"comment":"The derivation applies Jensen's inequality twice to obtain upper bounds on the negative log expected likelihood, then absorbs the slack into scalar weights \\lambda_d and \\lambda_p (Eqs. (23) and (29)). The slack is a function of the distributions p_{x|\\tilde x_j}, p_{\\tilde x_j}, and p_{x_i}, not merely of the synset sizes |X| and |\\tilde X|, so it cannot be absorbed into a constant hyperparameter. Minimizing an upper bound is not equivalent to minimizing the original objective; consequently the claimed equivalence in Eq. (17) is not established.","section":"A.1, Eqs. (21)(d)-(e)"},{"comment":"Eq. (25) shows that the sum of the second and third terms in Eq. (21) equals f(x,X), which the paper identifies as a constant for given x and X. The subsequent split introduces the E-KLD term and the \\delta_p term, and Eq. (28) states E-KLD - \\delta_p = f(x,X). This means the E-KLD is not derived as part of the objective; it is an additional quantity that the authors propose to minimize to 'approximate' a constant. This is not an equivalence and cannot support Lemma 3.2.","section":"A.1, Eqs. (25)-(28)"},{"comment":"The identification of the rate term with I(X;\\hat{\\tilde X}) relies on H_s(\\hat{\\tilde X}|X)=0, justified by log q(\\tilde y_s|x)=0 and a 'determined encoder.' However, reconstruction is not determined by x because the decoder samples \\hat y_{\\epsilon,j}; conditional semantic entropy is not generally zero. Moreover, Theorem 3.3 is stated as a 'minimum achievable rate' without an operational coding theorem (achievability, converse, block-length definitions); the proof only manipulates the variational objective. The central rate-distortion-perception theorem is therefore not proven.","section":"A.2, Eq. (35)"},{"comment":"Because the ideal synset is defined by a perceptual-similarity criterion, the appearance of a distribution-divergence perception term in the objective is an input of the model rather than an independent prediction. The paper's repeated claim to be the first work to 'theoretically explain the fundamental reason for the divergence measure's existence' is stronger than what the derivation supports; at best it shows that a synset defined by perceptual similarity leads to a perceptual penalty.","section":"3.1 and A.2"}],"minor_comments":[{"comment":"The statement that the first term 'equals 0 under the assumption of a uniform density on the unit interval' is inaccurate; a uniform density has a constant log-density value, and the KL term is not itself zero. The text should say the term is constant and can be absorbed.","section":"2.1, Eq. (2); 3.2, Eq. (10)"},{"comment":"The notation for the reconstructed synset is inconsistent: Lemma 3.2 uses \\tilde X, Theorem 3.3 uses \\hat X, and the mutual information I(X;\\hat{\\tilde X}) is not formally defined. Please unify the notation.","section":"3.2 and A.2"},{"comment":"The reference to 'Eq. (3.2)(d)' should be to Eq. (21)(d); there are also several typos, including 'appplied' in Section 3.1 and 'Kodat' in Figure 17.","section":"A.1, after Eq. (28)"},{"comment":"The authors describe the reported experimental results as 'preliminary results' and note insufficient hyperparameter exploration; this caveat should be stated in the main text, as it tempers the claim that the experiments verify the analysis.","section":"C.2"}],"recommendation":"reject","confidential_remarks":"To the editor: the fundamental problems are in the proof of Lemma 3.2 (Eqs. (20)-(21) and (25)-(28)) and Theorem 3.3 (Eq. (35)); the derivation of the E-KLD term is not an equivalence, so the main theorem is unproven. The paper depends heavily on the authors' own synonymity framework (Niu and Zhang, 2024), and the novelty of the 'first work to explain the divergence measure' claim should be carefully scrutinized. The experimental section is transparently preliminary and does not compensate for the proof gaps. I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the theoretical centerpiece doesn't survive close reading, but the progressive SIC codec is a legitimate piece of engineering. I'd send it to reviewers mainly to have them document why the proof fails.\n\nThe new thing here is the progressive channel-sliced codec: one model, 16 rate points, solid DISTS/LPIPS numbers that are competitive with No-GAN MS-ILLM over much of the range. The authors are also unusually candid about limitations, including LPIPS as a stand-in for KL, the FID gap, and the detail-sampling compromise. That part is worth reading.\n\nThe soft spot is Theorem 3.2, and it's load-bearing. Eq. (20) asserts p_{tilde_x_j|x_i}=p_{tilde_x_j|x} for every x_i in the ideal synset, justified by an ideal encoder. But that invariance is precisely what the training is supposed to create. Before convergence it's false; after convergence the derivation is circular. The Jensen steps in (21)(d)-(e) give upper bounds, and the paper then treats minimizing the bound as equivalent to the original objective. No equality condition is given for the alpha_d/alpha_p scaling factors, and the text even admits these absorb the inequality gap. So Lemma 3.2 is not proven as stated. Theorem 3.3 inherits the problem and adds a rate identity I(X; hat dot X) that is asserted rather than derived from a coding theorem. The paper's own Eq. (37) shows the result collapses to Blau-Michaeli when the reconstructed synset has one sample, which undercuts the novelty claim; the \"first explanation of the divergence term\" is not supported.\n\nDon't get me wrong: the synonymity framing is a reasonable way to think about perceptual compression, and the SIC architecture follows from it naturally. But the paper presents the derivation as a proof and it isn't one. The engineering is real; the theory needs major repair. Who gets value from this? Anyone working on semantic-information theory applied to compression, and anyone building multi-rate generative codecs, but they should treat the theoretical section as an open problem rather than a proven result. My recommendation: send it to peer review because the claim is substantive and the flaw is instructive, but be prepared for a reject or a major revision.","headline":"Solid progressive codec engineering, but the SVI proof rests on a circular invariance assumption and Jensen bounds treated as equivalences, so the central theorem does not stand as written.","tokens_in":31326,"tokens_out":2769,"would_cite":false,"duration_ms":33700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A34","94A15","94A17"],"pacs":[],"model":"deepseek-v4-flash","headline":"Viewing each image as a member of a set of perceptually synonymous images, this paper proves that the minimum rate for perceptual compression is a triple tradeoff among rate, distortion, and expected KL divergence, and that this bound…","keywords":["perceptual image compression","rate-distortion-perception tradeoff","synonymous variational inference","semantic information theory","synset","partial semantic KL divergence","progressive image codec","variational inference"],"falsifier":"Test equation (20) directly on a source whose synonym structure is known: train a SIC-style codec, generate several reconstructions from one code by sampling details, re-encode those reconstructions, and measure the dispersion of the recovered synonymous representations. If different synset members yield materially different reconstruction distributions, or if re-encoding different reconstructions gives significantly different $\\hat{y}_s$, the equivalence in Lemma 3.2 fails and the rate bound of Theorem 3.3 is not established. A quantitative companion check on the same synthetic source is to compute the claimed minimum rate and verify that a codec that ignores the synset violates it.","tokens_in":30122,"feed_emoji":"🖼️","tokens_out":13057,"duration_ms":122319,"temperature":0.7,"pith_summary":"This paper re-derives perceptual image compression from a synonymity-based semantic information viewpoint: the 'meaning' of an image is carried not by its pixels but by the set of images that look similar to it under a perceptual criterion, which the authors call the ideal Synset. Its central result, Theorem 3.3, states that the minimum achievable coding rate is set by a triple tradeoff — semantic mutual information between the source and the reconstructed synset, constrained by an expected distortion and an expected KL divergence between source and reconstruction distributions. Because the existing rate-distortion-perception tradeoff is exactly the single-reconstruction special case and classical rate-distortion is the single-image-synset special case, the theorem unifies the optimization goals of traditional and learned codecs. The divergence term that perceptual-compression losses always add is not an ad hoc choice: the analysis shows it arises because the reconstruction target is a set of synonymous images rather than a single image. The accompanying Synonymous Image Compression scheme, which codes only the shared latent part and samples the remaining details, delivers 16 rate points from one progressive model with rate-distortion-perception performance comparable to single-rate generative baselines.","feed_headline":"Perceptual image compression is a three-way tradeoff, theorem shows","feed_subtitle":"A new theorem traces perceptual losses to a 'synset' of similar images; one codec spans 16 bitrates.","key_machinery":"The load-bearing object is the ideal Synset $\\mathcal{X}$, the set of images that share meaning with the source image under a perceptual criterion, together with the partial semantic KL divergence $D_{\\mathrm{KL},s}[q\\|p_s]$, which measures the distance between a syntactic parametric density and a semantic distribution. The latent representation is split into a synonymous part $\\tilde{y}_s$, which must be coded and transmitted, and a detailed part $\\tilde{y}_\\epsilon$, which the decoder samples rather than receives, so only the synonymous part contributes to the rate. Lemma 3.2 converts the synonymous likelihood term into an expected-distortion plus expected-KL objective, and Theorem 3.3 packages that equivalence into the rate bound whose rate term is semantic mutual information $I(X;\\hat{\\mathring{X}}) = H_s(\\hat{\\mathring{X}}) - H_s(\\hat{\\mathring{X}}|X)$. The unification claim is carried by two degenerations: a single-member reconstructed synset reproduces the rate-distortion-perception tradeoff, and a single-member ideal synset reproduces classical rate-distortion.","core_discovery":"On its own terms, the paper's discovery is that perceptual image compression should be optimized against a set of synonymous reconstructions rather than a single reconstruction, and that this shift turns the optimization objective into a synonymous rate-distortion-perception tradeoff. Theorem 3.3 states that the minimum achievable rate is $$R(X) = \\min_{p(\\hat{X}|x)} I(X;\\hat{\\mathring{X}}) \\quad \\text{s.t.} \\quad \\mathbb{E}[d(x,\\hat{x}_i)] \\le D,\\; \\mathbb{E}[D_{\\mathrm{KL}}[p_x \\| p_{\\hat{x}_i}]] \\le P,$$ where the rate is the semantic mutual information between the source and the reconstructed synset. The load-bearing step is Lemma 3.2: minimizing the expected negative log synonymous-likelihood term is equivalent to minimizing a weighted expected distortion plus a weighted expected KL divergence, because the likelihood of the whole synset integrates over its members. The paper claims this is the first theoretical explanation of why a divergence measure must appear in perceptual image compression: the E-KLD term measures the gap between the reconstructed distribution and the ideal synset, and it disappears only when the ideal synset collapses to the original image. Both existing bounds then fall out as degenerations.","pith_inferences":["Editorial inference: the framework predicts that the benefit of synonymity-aware coding grows as the perceptual criterion becomes more resampling-tolerant (DISTS over LPIPS over MSE), a ranking that could be tested by training the same codec under each criterion and comparing the measured rate gap against the bound.","Editorial inference: the limited FID gains the paper reports point to its own next bottleneck — the detail sampler draws from a scalar uniform distribution instead of the conditional vector prior assumed by the derivation, so replacing the sampler is a direct and testable upgrade path.","Editorial inference: the synonymous idempotence constraint, which re-encodes reconstructions and penalizes $\\|\\hat{y}'_s - \\hat{y}_s\\|^2$, gives a practical, human-free measure of how close a codec is to overlapping the ideal synset, and could serve as a convergence criterion in future training.","Editorial inference: the same set-level reasoning should apply to other image restoration tasks whose success is judged by perceptual similarity, predicting that their optimal objectives also contain a divergence term of exactly this form."],"forward_implications":["If Theorem 3.3 is right, the classical rate-distortion bound and the existing rate-distortion-perception bound are the two degenerate edges of a single synonymous tradeoff, obtained by collapsing the ideal synset or the reconstructed synset to one sample.","Because only the synonymous representation is coded, the achievable rate satisfies $I(X;\\hat{\\mathring{X}}) \\le I(X;\\hat{X})$, so a synonymity-aware codec can beat a symbol-level codec at the same rate whenever the decoder is free to sample details rather than receive them.","A single progressive codec with $L$ split levels yields $L$ rate points from one generator; the implemented $L=16$ codec covers the full rate range on Kodak, CLIC2020, and DIV2K, with each rate sharing one analysis and synthesis transform.","The bound identifies the role of the perceptual loss term: it is the expected-KL constraint that drives the reconstructed synset toward the ideal synset, so divergence-based perceptual measures such as LPIPS, DISTS, and adversarial losses are all proxies for the same underlying term."],"supporting_citations":[{"why":"Supplies the synonymity-based semantic information theory the paper builds on: semantic variables, synsets, semantic entropy, and the partial semantic KL divergence that SVI minimizes.","marker":"Niu & Zhang, 2024"},{"why":"Defines the rate-distortion-perception function that Theorem 3.3 generalizes and to which the SIC objective reduces when the reconstructed synset holds a single sample.","marker":"Blau & Michaeli, 2019"},{"why":"Provides the variational autoencoder inference template that SVI adapts by replacing the pixel-level posterior with a synset-level posterior.","marker":"Kingma & Welling, 2013"},{"why":"Establishes the variational image-compression objective and the scale-hyperprior entropy model used for rate estimation in the SIC codec.","marker":"Ballé et al., 2018"},{"why":"Supplies the joint autoregressive and hierarchical prior architecture whose masked convolutions estimate the rate term in the progressive SIC model.","marker":"Minnen et al., 2018"},{"why":"Provides LPIPS, the perceptual measure the implementation substitutes for the exact expected-KL term in the training loss.","marker":"Zhang et al., 2018"},{"why":"HiFiC is the adversarial-loss generative codec that benchmarks the DISTS rate-perception comparisons and the GAN fine-tuning.","marker":"Mentzer et al., 2020"},{"why":"MS-ILLM is the main perceptual-compression baseline compared across DISTS, LPIPS, and FID, including its no-GAN variant trained with LPIPS.","marker":"Muckley et al., 2023"}],"fun_headline_variants":["Perceptual compression is a triple tradeoff, new theorem shows","Synonym sets yield a single codec for 16 bitrates","Semantic synonymity redefines rate-distortion-perception balance","One progressive model handles all bitrates via synset optimization","Image compression's true cost is a three-way tradeoff, theorem says"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The derivation assumes, in equation (20), that every image in the ideal Synset already maps to the same reconstruction distribution under the current encoder — the very invariance the training is meant to create — so the proof leans on its target property before that property exists.","fun_headline_variants_meta":{"raw":{"variants":["Perceptual compression is a triple tradeoff, new theorem shows","Synonym sets yield a single codec for 16 bitrates","Semantic synonymity redefines rate-distortion-perception balance","One progressive model handles all bitrates via synset optimization","Image compression's true cost is a three-way tradeoff, theorem says"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1309,"prompt_tokens":976,"completion_tokens":333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":244}},"tokens_in":592,"tokens_out":333,"duration_ms":4697,"temperature":1.0,"reasoning_tokens":244,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:08:12.945275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Test equation (20) directly on a source whose synonym structure is known: train a SIC-style codec, generate several reconstructions from one code by sampling details, re-encode those reconstructions, and measure the dispersion of the recovered synonymous representations. If different synset members yield materially different reconstruction distributions, or if re-encoding different reconstructions gives significantly different $\\hat{y}_s$, the equivalence in Lemma 3.2 fails and the rate bound of Theorem 3.3 is not established. A quantitative companion check on the same synthetic source is to compute the claimed minimum rate and verify that a codec that ignores the synset violates it.","supporting_citations":[{"cited_title":"and Michaeli, T","cited_arxiv_id":null,"evidence_quote":"Defines the rate-distortion-perception function that Theorem 3.3 generalizes and to which the SIC objective reduces when the reconstructed synset holds a single sample."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the joint autoregressive and hierarchical prior architecture whose masked convolutions estimate the rate term in the progressive SIC model."},{"cited_title":"A., Shechtman, E., and Wang, O","cited_arxiv_id":null,"evidence_quote":"Provides LPIPS, the perceptual measure the implementation substitutes for the exact expected-KL term in the training loss."},{"cited_title":"D., Tschannen, M., and Agustsson, E","cited_arxiv_id":null,"evidence_quote":"HiFiC is the adversarial-loss generative codec that benchmarks the DISTS rate-perception comparisons and the GAN fine-tuning."},{"cited_title":"J., El-Nouby, A., Ullrich, K., Jegou, H., and Verbeek, J","cited_arxiv_id":null,"evidence_quote":"MS-ILLM is the main perceptual-compression baseline compared across DISTS, LPIPS, and FID, including its no-GAN variant trained with LPIPS."}],"review_version":1}