{"id":"04d484e3-a9d6-462e-8fe5-b482e5f5460b","arxiv_id":"2607.05381","paper_version":1,"verdict":"ACCEPT","confidence":"UNKNOWN","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The discrete diffusion NELBO equals data entropy plus an exact path KL to the oracle reverse process, and the denoiser, cavity, and score parameterizations are three interconvertible coordinates of the unique optimal reverse jump rate.","lead":"This paper proves that the negative ELBO of a discrete diffusion model is exactly the data entropy plus the path-space KL divergence between the true and learned reverse processes — not merely a bound. It unifies the denoiser, cavity, and score parameterizations as three coordinates of one object, with closed-form conversions, and shows why denoiser parameterizations diverge for uniform diffusion while cavity parameterizations stay finite.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified. The Oracle Distance identity and its supporting lemmas are rigorously proved via two independent routes (Pythagorean/Bregman assembly and KL chain rule), with standard assumptions that hold for all standard discrete diffusion processes on interior windows.","rationale":"The paper's central theoretical claims are correct and rigorously proved. The Oracle Distance identity is an exact mathematical result, not an approximation, and holds under standard conditions that are satisfied by all common discrete diffusion processes. The two independent proof routes (Pythagorean assembly and KL chain rule) provide strong internal validation. The numerical verification on a toy model confirms the identities without approximation.\n\nThe reader's identified concern (Assumption 2) is technically necessary but practically mild on finite state spaces with standard parameterizations. The more interesting limitation — the teacher-forced nature of the analysis — is explicitly acknowledged by the authors and does not constitute a flaw in the theoretical results.\n\nThe paper makes genuine contributions: (1) the exact NELBO = H(q_0) + path KL identity with boundary terms, (2) the information-theoretic interpretation of the oracle cost, (3) the three-coordinate dictionary with closed-form conversions, and (4) the denoiser divergence result for uniform diffusion. These are substantive advances in the theoretical understanding of discrete diffusion.\n\nThe main gap between theory and practice (scale of validation) is acknowledged and partially addressed by concurrent work [GJS+26]. The verdict of ACCEPT is appropriate.","tokens_in":58196,"tokens_out":9015,"duration_ms":400440,"concrete_test":"Verify the convert-or-pay penalty (Figure 5, Table 6) on a larger-scale model (e.g., a transformer with ≥10M parameters on a real text dataset). Train separate denoiser, cavity, and score heads to convergence, then evaluate each head in all three sampling coordinates — with and without the analytic conversions of Proposition 3. If the off-diagonal penalties grow or fail to vanish with correct conversion at scale, the practical relevance of the coordinate dictionary would be in question. The concurrent results of [GJS+26] on uniform diffusion at scale provide partial evidence, but the full three-coordinate comparison across masked/uniform/GIDD has not been tested beyond the toy model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After careful review, I cannot identify a load-bearing concern that would undermine the central claim. The Oracle Distance theorem (Theorem 3) is proved twice: once via the Pythagorean decomposition (Proposition 2 + Theorem 4), requiring Assumption 3 (fixed support, piecewise continuity), and once via the KL chain rule on path laws (§5.5), requiring only Assumption 2 (rate-support) and the Markov property. Both routes are correct.\n\nThe reader flags Assumption 2 as load-bearing. It is necessary for the Girsanov formula (Theorem 2), but on finite state spaces it reduces to support inclusion: supp bQ̂_t(z_t,·|z_0) ⊆ supp bQ^θ_t(z_t,·). For standard parameterizations (softmax heads, cavity-induced rates), this is automatically satisfied because the forward kernel q_{t|0} has full support on interior windows for masked/uniform/GIDD, and the model rates inherit this support. A neural network with softmax outputs cannot assign exactly zero rate to any jump. The paper also notes (Remark 3) that cavity parameterizations are automatically confined to the correct support S_i, while denoiser parameterizations must learn it — a practical difficulty, not a theoretical gap.\n\nThe information-loss identity J*_t = d/dt H(Z_0|Z_t) (Theorem 4) requires Assumption 3, which holds for all three standard processes on compact interior windows. The proof differentiates H(Z_0|Z_t) using the Kolmogorov forward equation and matches it to the oracle cost via an algebraic identity involving posterior ratios q_{0|t}(z_0|y)/q_{0|t}(z_0|z). The manipulation is correct.\n\nThe telescoping of conditional entropies in the assembly proof is exact: H(Z_0|Z_{t_1}) + [H(Z_0|Z_{t_2}) - H(Z_0|Z_{t_1})] + [H(Z_0) - H(Z_0|Z_{t_2})] = H(q_0). The boundary terms are handled correctly.\n\nThe closest thing to a substantive limitation is that the entire framework characterizes the teacher-forced objective: the path KL is evaluated under the oracle process P⋆ (states z_t ~ q_t), not under the model's own trajectory. Th","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper provides a rigorous, self-contained theoretical account of continuous-time discrete diffusion models. The central result is the Oracle Distance theorem (Theorem 3), which establishes that the negative ELBO equals the data entropy plus an exact path-law KL divergence from the oracle reverse process to the learned one—not merely a bound. The paper identifies the population optimizer as the marginal reverse jump rate (the posterior average of the clean-conditioned rate), shows the irreducible oracle cost equals the information-destruction rate d/dt H(Z₀|Z_t), and proves a universal NELBO floor at the data entropy. For token-factorizable processes, the paper derives three exact coordinate representations of the optimal reverse rate (denoiser, cavity, score) with closed-form conversions, recovers MDM/UDM/SEDD/GIDD as special cases, and proves that a denoiser parameterization makes the uniform-diffusion ELBO diverge at initialization while the cavity parameterization stays finite. Two independent proof routes are provided for both the CTMC ELBO (infinitesimal KL and Girsanov) and the Oracle Distance theorem (Pythagorean assembly and direct path-law factorization). All identities are verified numerically on an exactly solvable toy model.","tokens_in":58442,"tokens_out":3201,"duration_ms":210847,"significance":"The paper makes several strong contributions. It ships parameter-free derivations from standard tools (Girsanov's formula, KL chain rule, Bregman divergence identities, Kolmogorov forward equations) with no fitted constants. The Oracle Distance identity is an exact result, not a bound, and is proved twice via independent routes. The information-loss identity J*_t = d/dt H(Z₀|Z_t) connects the ELBO to the I-MMSE relation in the CTMC setting. The three-coordinate dictionary (denoiser/cavity/score) with exact conversion formulas is a practically useful unification that cleanly separates what each literature loss optimizes. The divergence result for the uniform-diffusion denoiser parameterization (Proposition 6) gives a theoretical account of a known empirical failure mode. The numerical verification, while on a toy model, is exact and tests every identity without approximation. The concurrent work of Gourevitch et al. [GJS+26] independently discovered the denoiser/cavity distinction for uniform diffusion; this paper generalizes that insight to all token-factorizable processes from a single projection principle, which is a meaningful extension.","major_comments":[{"comment":"The paper is theoretically sound. I examined both proof routes for the Oracle Distance theorem (§5.1 assembly via Proposition 2 + Theorem 4, and §5.5 direct path-law factorization), the Girsanov formula (Theorem 2), the Bregman-Pythagorean identity (Lemma 2), the reverse-rate projection (Theorem 5), and the information-loss identity (Theorem 4). The key algebraic steps are correct: the Pythagorean decomposition follows from the standard Bregman conditional-mean identity, the information-loss proof correctly uses the Kolmogorov forward equation and the generator's zero-row-sum property to insert the vanishing term in Eq. (34), and the alternative proof in §5.5 correctly applies the KL chain rule using the Markov property factorization (38). Assumption 2 (rate-support) is load-bearing for the Girsanov formula but reduces to support inclusion on finite state spaces and is automatically满足ed由","section":null},{"comment":"The paper is theoretically sound. I examined both proof routes for the Oracle Distance theorem (§5.1 assembly via Proposition 2 + Theorem 4, and §5.5 direct path-law factorization), the Girsanov formula (Theorem 2), the Bregman-Pythagorean identity (Lemma 2), the reverse-rate projection (Theorem 5), and the information-loss identity (Theorem 4). The key algebraic steps are correct: the Pythagorean decomposition follows from the standard Bregman conditional-mean identity, the information-loss proof correctly uses the Kolmogorov forward equation and the generator's zero-row-sum property, and the alternative proof in §5.5 correctly applies the KL chain rule. Assumption 2 (rate-support) is load-bearing for Girsanov but reduces to support inclusion on finite spaces and is automatically satisfied for standard softmax/cavity parameterizations. Assumption 3 (fixed support, piecewise continuity) ","section":null},{"comment":"The paper is theoretically sound. I examined both proof routes for the Oracle Distance theorem, the Girsanov formula, the Bregman-Pythagorean identity, the reverse-rate projection, and the information-loss identity. The key algebraic steps are correct. Assumption 2 (rate-support) is load-bearing for Girsanov but reduces to support inclusion on finite spaces and is automatically satisfied for standard parameterizations. Assumption 3 is only needed for Theorem 4, not for Theorem 3 itself (as the §5.5 alternative proof shows). I do not identify any load-bearing error that would undermine the central claims.","section":null}],"minor_comments":[{"comment":"§4.2, Definition 1: The footnote defining the Skorokhod space D([t₁,t₂],X) is helpful, but the σ-field is described as 'generated by the coordinate maps' without specifying whether this is the raw or predictable σ-field. For finite state spaces this distinction is immaterial, but a brief clarification would improve rigor.","section":null},{"comment":"§6.2, Remark 3: The convention 0/0 := 0 for terms where both numerator and denominator vanish is introduced here but used implicitly earlier (e.g., in the denoiser rate formula (52)). Stating this convention at first use (or in §4.2) would prevent confusion.","section":null},{"comment":"§7.1, Proposition 5: The Itakura–Saito divergence D_IS(p∥q) = p/q - log(p/q) - 1 is defined inline but its non-negativity is not explicitly stated. A one-line note that D_IS ≥ 0 (by log x ≤ x - 1) would help readers verify that the cavity integrand (62) is non-negative.","section":null},{"comment":"Table 5: The '1/V denoiser NELBO' entry for GIDD references Remark 4 rather than a numbered proposition. Consider elevating Remark 4 to a corollary of Proposition 6 for formal parity, since the divergence result is load-bearing for the practical recommendation to avoid denoiser parameterizations.","section":null},{"comment":"§8: The toy model uses V=8, L=8 with a bigram chain. While appropriate for exact verification, a brief note on why this model is sufficient to test all identities (e.g., it has non-trivial position correlations so denoiser ≠ cavity for uniform/GIDD) would help readers appreciate the design choice.","section":null},{"comment":"Figure 5: The panel labels (a)–(f) are referenced in the text but the figure caption could more explicitly state which theoretical result each panel verifies (e.g., 'Panel (b): Theorem 7 convert-or-pay penalty').","section":null},{"comment":"§3 (Related Work): The paper's relationship to Generator Matching [HHY+25] is discussed at the objective level but the distinction could be sharper. Specifically, a sentence noting that Generator Matching's Propositions 1–2 already contain the conditional-mean optimality and Bregman gradient-sharing at the matching-objective level, while this paper's contribution is the likelihood-side identity (ELBO = entropy + path KL), would make the novelty boundary clearer.","section":null},{"comment":"Several references are dated 2026 (e.g., [GJS+26], [SLY+26], [RNB+26], [ST25]). If these are accepted/published papers, please verify the citation metadata; if they are preprints, consider noting 'preprint' or 'to appear' consistently.","section":null},{"comment":"Notation: The paper uses both Q̂_t (with hat) and bQ_t (with backslash-b) for reverse rates. While defined in §4.2, the two notations appear in different sections and a unified choice would improve readability.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a strong theoretical contribution that should be published. The central results are rigorously proved via two independent routes, the practical insights are concrete and falsifiable, and the relationship to concurrent work [GJS+26] is properly disclosed with the generalization clearly articulated. The only reason for minor revision rather than accept is the collection of presentation issues listed above—none of which affect correctness. The toy-model-only validation is a limitation but is appropriate for a theory paper whose main results are exact identities; the authors acknowledge this and note that [GJS+26] provides scale-level evidence for the uniform special case. I would not request large-scale experiments as a condition of acceptance."},"author_rebuttal":{"model":"glm-5.2","summary":"The referee report is a positive review recommending minor revision. The three 'major comments' appear to be repeated instances of the same assessment: the paper is theoretically sound, the proofs are correct, and no load-bearing errors undermine the central claims. The referee makes two specific observations about assumptions that we can address with minor clarifications.","responses":[{"response":"We thank the referee for the careful and thorough verification of our proofs. We are gratified that the referee independently checked both routes for the Oracle Distance theorem (§5.1 assembly and §5.5 direct factorization), the Girsanov formula (Theorem 2), the Bregman-Pythagorean identity (Lemma 2), the projection result (Theorem 5), and the information-loss identity (Theorem 4), including the key algebraic steps such as the Kolmogorov forward equation insertion in Eq. (34) and the KL chain rule factorization in Eq. (38).","revision_made":"no","referee_comment":"The paper is theoretically sound. Both proof routes for the Oracle Distance theorem, the Girsanov formula, the Bregman-Pythagorean identity, the reverse-rate projection, and the information-loss identity were examined and found correct."},{"response":"We agree with the referee's characterization. Assumption 2 is indeed the absolute-continuity condition required for the Girsanov change-of-measure (Theorem 2) and hence for the CTMC ELBO (Theorem 1). On finite state spaces it is simply support inclusion: supp bQt(zt,·|z0) ⊆ supp bQθ_t(zt,·). For the standard parameterizations studied in §§6–7, this holds automatically: the cavity rate (55) never divides by per-token likelihood and is confined to the forward support Si by construction, while the score rate (57) inherits support from the score head. The denoiser rate (52) requires the additional convention that 0/0 := 0 and that supp πθ_i ⊆ Si(zi_t, t), which we discuss in Remark 3. We will add a brief sentence after Assumption 2 explicitly noting that it reduces to support inclusion on finite spaces and is automatically satisfied for the cavity and score parameterizations, and that the denoiser parameterization requires the support restriction of Remark 3.","revision_made":"yes","referee_comment":"Assumption 2 (rate-support) is load-bearing for the Girsanov formula but reduces to support inclusion on finite state spaces and is automatically satisfied for standard softmax/cavity parameterizations."},{"response":"The referee is exactly right. Assumption 3 (regularity and fixed support) is used only in the proof of Theorem 4 (the information-loss identity J*_t = d/dt H(Z0|Zt)), where it ensures that the conditional entropy H(Z0|Zt) is absolutely continuous and that the differentiation under the sum, the Kolmogorov forward equation substitution, and the zero-row-sum insertion in Eq. (34) are all valid. The Oracle Distance theorem (Theorem 3) itself does not require Assumption 3: the §5.5 alternative proof uses only the reverse-time Markov factorization (38), the KL chain rule, and Girsanov's formula (Theorem 2), none of which needs fixed support or piecewise continuity beyond boundedness and measurability. We already state this in §5.1 ('The regularity assumption is needed only for this second ingredient, not for Theorem 3 itself; §5.5 gives an alternative proof of the latter that avoids it'), but we will make the point more prominent by adding a remark after Theorem 3 explicitly cross-referencing the §5.5 proof's weaker hypotheses.","revision_made":"yes","referee_comment":"Assumption 3 (fixed support, piecewise continuity) is only needed for Theorem 4, not for Theorem 3 itself, as the §5.5 alternative proof shows."}],"tokens_in":58145,"tokens_out":1223,"duration_ms":68922,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"This paper proves something genuinely useful: the negative ELBO of a discrete diffusion model is exactly the data entropy plus the path-space KL from the oracle reverse process to the learned one — not a bound, an identity, with boundary terms made explicit. That is the Oracle Distance theorem (Theorem 3), and it is the main reason to read this paper. The authors give two independent proofs (a Bregman/Pythagorean assembly and a direct KL-chain-rule factorization on path laws), which is good practice and the proofs check out. The information-loss identity J*_t = d/dt H(Z_0|Z_t) connecting the oracle cost to mutual information decay is a clean result, extending the I-MDSE line to arbitrary model rates and finite windows with boundary terms intact. The three-coordinate dictionary (denoiser, cavity, score) with closed-form conversions for token-factorizable processes is the other real contribution. It unifies MDM, UDM, SEDD, and GIDD under one object, explains why denoiser and cavity coincide for masked but not uniform diffusion, and the denoiser divergence result (Proposition 6) gives a concrete theoretical account of why uniform-diffusion denoiser training is ill-conditioned at initialization. The numerical verification on an exactly solvable bigram model is appropriate — every identity is checked without approximation, which is the right way to validate exact theoretical claims. The rate-support assumption (Assumption 2) is load-bearing for the Girsanov formula but is standard: on finite state spaces it reduces to support inclusion, which holds automatically for softmax heads and cavity parameterizations. I agree with the reader and the stress-test that this is not a real concern. The main limitation is that validation is only on a toy model. The theoretical results don't depend on scale, but the practical claim that coordinate conversion matters at model scale is supported here only indirectly, via the concurrent Gourevitch et al. results for uniform diffusion. The authors are honest about this. The paper is also long — §2 reprises §§4–7 by design, which helps readability but adds bulk. This is a well-executed theory paper that clarifies what discrete diffusion training actually optimizes. The central identity is new, the proofs are rigorous, and the dictionary is a genuine unification. It deserves a serious referee.","headline":"Solid theory paper: exact NELBO = entropy + path KL identity, plus a clean denoiser/cavity/score dictionary. Deserves a serious referee.","tokens_in":59128,"tokens_out":1101,"would_cite":true,"duration_ms":96133,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Discrete diffusion ELBO is an exact path distance, not a bound","keywords":[],"falsifier":"If the Oracle Distance identity holds, then for any noising process and any model rate, the data-averaged NELBO minus the data entropy must exactly equal the sum of the reconstruction KL, the terminal-prior KL, and the integrated local rate divergence evaluated at marginal (not clean-conditioned) rates. A single numerical counterexample on a finite-state model would refute it. Additionally, the Pythagorean split predicts that training a denoiser head and reading it as a cavity head without conversion yields a strictly suboptimal reverse process whose generative perplexity exceeds the oracle, a","tokens_in":58495,"feed_emoji":"🔀","tokens_out":1022,"duration_ms":166426,"temperature":0.7,"pith_summary":"This paper proves that the negative ELBO of a discrete diffusion model is not merely a loose likelihood bound but an exact identity: it equals the data entropy plus the KL divergence between the entire trajectory of the true reverse process (the oracle) and the learned one. The unique optimal model is the marginal reverse jump rate, the posterior average of the clean-conditioned rate given only the noisy state. The irreducible cost of training, shared by every noising process, is exactly the rate at which the forward process destroys information about the clean data, which integrates to the data entropy. For sequence models, the paper shows that the denoiser, cavity (bridge plug-in), and concrete score parameterizations are three exact coordinate representations of this single optimal reverse rate, with closed-form conversions, and explains why a denoiser parameterization diverges at initialization under uniform diffusion while the cavity stays finite.","feed_headline":"Discrete diffusion ELBO is an exact path distance, not a bound","feed_subtitle":"The negative ELBO equals data entropy plus the KL between the true and learned reverse trajectories, making all noising processes share one","key_machinery":"The local rate divergence Phi(a,b) = a log(a/b) - a + b, a Bregman divergence of the negative-entropy potential, serves as the per-jump KL between two CTMC intensities. A Bregman-Pythagoras identity splits any model's cost into the oracle cost plus model mismatch with no cross term. For product noising processes, the reverse rate factors into per-token jumps expressible in three coordinates: the denoiser (clean-token posterior given the full noisy state), the cavity law (clean-token posterior given the noisy context but not the local noisy token), and the concrete score (ratio of noisy conditional laws), with closed-form conversions via local Bayes inversion of the forward kernel.","core_discovery":"The central object is the Oracle Distance theorem: for any discrete diffusion noising process and any learned reverse rates, the data-averaged negative ELBO minus the data entropy equals a reconstruction KL at the lower endpoint plus the path-space KL from the oracle reverse process to the learned one. The path KL decomposes into a terminal-prior KL plus an integral of local rate divergences between the true marginal reverse rate and the model rate. This is an exact equality, not an inequality. Its unique optimizer is the marginal reverse rate, obtained by projecting the clean-conditioned bridge rate onto the information available at the noisy state, and its irreducible per-time cost is d/dt","pith_inferences":[],"forward_implications":["ELBO values reported across different noising processes (masked, uniform, GIDD) become directly comparable once boundary terms are retained, since they all share the same entropy floor and the excess measures path divergence to the oracle.","A uniform cavity head initialized to 1/V yields per-token NELBO of log V regardless of process or training window, providing an exact calibration tool for debugging ELBO implementations.","Using a denoiser parameterization for uniform or GIDD diffusion causes the ELBO to diverge at initialization, scaling as (V-1)/V * log(1/beta), while the cavity parameterization remains finite, explaining a known training instability.","The framework cleanly separates reverse-rate error (what the ELBO measures) from sampling factorization error (invisible to the ELBO), showing that even an oracle model with zero rate error degrades at few-step sampling due to the forced product approximation over positions.","The information-uniform clock, defined by the cumulative conditional entropy, drains information linearly and straightens process-dependent cost curves onto a single diagonal, offering a principled alternative to variance-optimal importance sampling for time discretization."],"fun_headline_variants":["Negative discrete diffusion ELBO equals an exact path KL, not a bound","Oracle Distance theorem: discrete diffusion ELBO is an exact path divergence","Discrete diffusion ELBO exactly equals data entropy plus oracle reverse-path KL","All discrete diffusion noising processes share the same best achievable ELBO","Discrete diffusion loss coordinate choice changes what law is actually optimized"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The learned reverse process must assign positive jump rate to every transition the true reverse process uses with positive probability. If the model ever assigns zero rate to a jump the oracle needs, the path-KL identity breaks because the relevant density ratio becomes undefined. This holds for standard parameterizations of masked, uniform, and GIDD diffusion but is an active constraint on the model class: a neural network that collapses its output distribution too narrowly,","fun_headline_variants_meta":{"raw":{"variants":["Negative discrete diffusion ELBO equals an exact path KL, not a bound","Oracle Distance theorem: discrete diffusion ELBO is an exact path divergence","Discrete diffusion ELBO exactly equals data entropy plus oracle reverse-path KL","All discrete diffusion noising processes share the same best achievable ELBO","Discrete diffusion loss coordinate choice changes what law is actually optimized"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":826,"prompt_tokens":735,"completion_tokens":91,"prompt_tokens_details":null},"tokens_in":735,"tokens_out":91,"duration_ms":68484,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T13:36:40.445373+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the Oracle Distance identity holds, then for any noising process and any model rate, the data-averaged NELBO minus the data entropy must exactly equal the sum of the reconstruction KL, the terminal-prior KL, and the integrated local rate divergence evaluated at marginal (not clean-conditioned) rates. A single numerical counterexample on a finite-state model would refute it. Additionally, the Pythagorean split predicts that training a denoiser head and reading it as a cavity head without conversion yields a strictly suboptimal reverse process whose generative perplexity exceeds the oracle, a","supporting_citations":[],"review_version":1}