{"id":"faae4b1f-0db5-4886-a5d0-84639092cb05","arxiv_id":"1908.09429","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MALA-within-Gibbs samplers can achieve dimension-independent acceptance and convergence rates for high-dimensional targets with sparse conditional structure, under block-wise log-concavity.","lead":"This paper proves that MALA-within-Gibbs, a block-wise MCMC algorithm, can keep its acceptance rate and convergence speed independent of the dimension of the target distribution when the distribution has sparse conditional structure and the sampler uses it. This matters for high-dimensional Bayesian inverse problems, where each parameter only interacts with a few neighbors.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof of Theorem 3.6 implicitly assumes a Lipschitz condition on all mixed Hessian blocks; Assumption 3.2 supplies it only for diagonal blocks, so Lemma A.4 is not justified as stated.","rationale":"The reader identified Assumption 3.2 as the weakest point, but focused on the boundedness requirement that excludes Gaussian targets. My stress-test found a different, more load-bearing defect in the same assumption: the stated hypotheses do not provide the Lipschitz control of mixed Hessian blocks that Lemma A.4 and the proof of Theorem 3.6 require. This affects the validity of the main theorem itself, not only its applicability. The issue is concrete and technical: boundedness of ∇_{x_i}v_j is not a Lipschitz bound, and no bound on third derivatives is assumed. Because the contraction argument relies on Lemma A.4 to control differences of averaged second derivatives, the theorem is not proved as written. This is not an ad hominem or a disagreement with current consensus; it is an internal gap between an assumption and its use. The fix is straightforward in principle: add a uniform Lipschitz condition on all mixed second derivatives (equivalently, a uniform bound on third derivatives of log π) to Assumption 3.2, or find another way to prove Lemma A.4. Since the result is likely salvageable with this strengthening and the paper's numerical demonstrations are suggestive, a conditional verdict is appropriate: the authors should be asked to repair Assumption 3.2 and re-verify the proof before the dimension-independence claim is accepted.","tokens_in":33238,"tokens_out":15387,"duration_ms":156771,"concrete_test":"Independently re-derive Lemma A.4 from the printed Assumption 3.2. Concretely, for i ≠ j, attempt to prove ||∇_{x_i}v_j(x) − ∇_{z_i}v_j(z)|| ≤ H_v Σ_{l ∈ I_{i,j}} ||x_l − z_l|| using only ||∇_{x_i}v_j(x)|| ≤ H_v and the diagonal Lipschitz inequality ||∇_{x_j}v_j(x) − ∇_{z_j}v_j(z)|| ≤ H_v ||x−z||. If this derivation cannot be completed, the proof of Theorem 3.6 is incomplete under the stated assumptions; then add to Assumption 3.2 the stronger condition ||∇_{x_i}v_j(x) − ∇_{z_i}v_j(z)|| ≤ H_v||x−z|| for all i,j, restate Theorem 3.6 under this strengthened assumption, and re-check the coupling estimates (1.19)–(1.28).","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Assumption 3.2 states the Lipschitz bound ||∇_{x_j}v_j(x) − ∇_{z_j}v_j(z)|| ≤ H_v ||x−z|| only for the diagonal block of the Hessian. Lemma A.4, however, claims a corresponding Lipschitz bound for every mixed block ∇_{x_i}v_j, and the proof of Theorem 3.6 uses this repeatedly, for example when replacing [x^{k,j}, z^{k,j}] by [x^k, z^k] around Eq. (1.19) and in the bounds on D1 and D2. The boundedness condition ||∇_{x_i}v_j|| ≤ H_v does not imply Lipschitz continuity of the function x ↦ ∇_{x_i}v_j(x): that would require a uniform bound on third derivatives ∇_x∇_{x_i}v_j, which is not assumed (only C^2 is assumed). The symmetry ∇_{x_i}v_j = (∇_{x_j}v_i)^T converts the missing condition into a Lipschitz bound on off-diagonal blocks of v_i, which is likewise not present in Assumption 3.2. Thus Lemma A.4 is not derivable from the stated hypotheses, and the contraction estimates in the core theorem lack support. This is a genuine gap in the central argument, distinct from the restrictiveness of the bounded-gradient assumption. The likely remedy is to strengthen Assumption 3.2 with a uniform Lipschitz condition on all mixed second derivatives, or to replace the offending estimates with a separate argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies MALA-within-Gibbs samplers for target distributions with sparse conditional structure. Under Assumption 3.1 (sparse conditional structure), Assumption 3.2 (bounded vector fields), and Assumption 3.5 (block-wise log-concavity), the authors prove Proposition 3.3, a dimension-independent lower bound on block acceptance probabilities, and Theorem 3.6, a dimension-independent geometric contraction rate for two coupled chains. The paper also discusses practical issues of finding a suitable block partition, illustrates the method on a log-Gaussian Cox point process and an elliptic PDE inverse problem, and compares with pCN and MALA/MMALA.","tokens_in":33479,"tokens_out":6762,"duration_ms":66252,"significance":"If the proofs were fully supported, the paper would make a useful contribution: it identifies a natural structural condition under which partial updating can remove the usual dimension dependence of MALA step sizes and convergence rates, complementing earlier Gaussian localization results in [33]. The paper is honestly written, states its assumptions explicitly, and the numerical experiments are substantive and clearly described. The central theorem, however, relies on a lemma that is not justified by the stated hypotheses, so the main contraction result is currently unsupported as written.","major_comments":[{"comment":"Lemma A.4 is not derivable from Assumption 3.2. Assumption 3.2 contains a Lipschitz bound only for the diagonal Hessian block, ||∇_{x_j} v_j(x) − ∇_{z_j} v_j(z)|| ≤ H_v ||x − z||; for i ≠ j it only asserts the pointwise bound ||∇_{x_i} v_j(x)|| ≤ H_v. Lemma A.4 claims a Lipschitz bound for every mixed block ∇_{x_i} v_j, and its proof applies the diagonal Lipschitz bound to the off-diagonal object ∇_{y_j} v_i(y) − ∇_{z_j} v_i(z). This step is not justified. The gap is load-bearing: Eq. (1.19), used to replace [x^{k,j}, z^{k,j}] by [x^k, z^k], and the estimates on C1 around Eq. (1.25) both rely on Lemma A.4, and these feed directly into the contraction estimate in Theorem 3.6. The authors should either strengthen Assumption 3.2 to include a uniform Lipschitz condition on all mixed Hessian blocks (or a comparable third-derivative condition) and propagate it through the proof, or supply a separate proof of Lemma A.4.","section":"Appendix A.4, Lemma A.4"},{"comment":"The advertised sufficient conditions omit Assumption 3.2. The abstract states that acceptance and step size are dimension-independent when the target has sparse conditional structure and the sampler reflects it; the convergence-rate statement adds only block-wise log-concavity. In fact Proposition 3.3 and Theorem 3.6 also require the bounded-vector-field assumption. This is not a purely cosmetic mismatch: Assumption 3.2 excludes Gaussian targets with sparse precision matrices, the prototypical motivating example discussed in Section 3.1, and it also fails for the log-Gaussian Cox posterior of Section 5.3, whose gradient contains the linear term −[B^{-1}](x − μ1). The authors do acknowledge the restrictiveness of Assumption 3.2, but the abstract and introduction should be revised so that the stated sufficient conditions match the theorem hypotheses, and ideally the paper should discuss what can be proved under weaker growth conditions.","section":"Abstract and Section 3.1"}],"minor_comments":[{"comment":"In the display after Eq. (1.1), the symbol U^n_j appears in the last case; it should presumably be U^k_j.","section":"Appendix A.1, Eq. (1.1)"},{"comment":"The phrase 'uniformly bounded and negative' is slightly misleading because a symmetric negative definite matrix can have positive off-diagonal entries. The definition is clear from the display, but consider saying 'negative definite' and explicitly stating that the off-diagonal entries H_{j,i} may be positive.","section":"Definition 3.4"},{"comment":"The discussion of Table 2 states that acceptance ratios are 'independent of the overall problem dimension' based on comparable tuned values. Since step sizes were tuned separately for each block size and problem, it would help to state explicitly that the observed acceptance ratios are comparable after tuning, rather than fixed a priori.","section":"Section 5.3.3"}],"recommendation":"major_revision","confidential_remarks":"The main correctness issue is the unsupported Lemma A.4; it appears fixable by strengthening Assumption 3.2 or by an additional argument, so I do not recommend rejection. The scope mismatch between the abstract and the actual assumptions is also worth addressing. I saw no evidence of circularity or fitted claims; the numerical experiments are presented honestly, including explicit discussion of violated assumptions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nShort version: this is a worthwhile paper with a real gap in the central convergence proof. The core idea—that MALA-within-Gibbs can obtain dimension-independent acceptance and, under blockwise log-concavity, dimension-independent convergence when the target has sparse conditional structure—is new and worth engaging with. It generalizes the Gaussian localization results of earlier work and handles non-Gaussian targets via a coupling argument. Proposition 3.3 (dimension-independent acceptance) looks right; the proof via Lemma A.2 is coherent and the intuition is clear.\n\nThe problem is Lemma A.4. Assumption 3.2 states a Lipschitz bound only for the diagonal Hessian block of each block gradient v_j. Lemma A.4 needs the same Lipschitz property for every mixed block ∇_{x_i} v_j, and the proof tries to obtain it by symmetry. Symmetry converts ∇_{x_i} v_j into ∇_{x_j} v_i, but Assumption 3.2 controls ∇_{x_i} v_i, not off-diagonal blocks of v_i. So the bound does not follow from the stated hypotheses. Theorem 3.6 uses Lemma A.4 repeatedly, for example when replacing [x^{k,j}, z^{k,j}] by [x^k, z^k] and in the bounds on D1 and D2. The contraction argument therefore lacks support as written. This is not just a restrictive assumption; it is a missing hypothesis in the main proof. The likely fix is to strengthen Assumption 3.2 with uniform Lipschitz continuity of all mixed second derivatives, or to supply a separate argument for the off-diagonal blocks.\n\nOther soft spots, in proportion: Assumption 3.2 excludes Gaussian targets, the most natural class with sparse precision matrices; the authors acknowledge this. The numerical experiments use MMALA-within-Gibbs rather than the analyzed MALA-within-Gibbs, and in the first example blockwise log-concavity is explicitly not satisfied. That limits the empirical support to heuristics plus a nearby algorithm, though the observed dimension-independence of IACT and acceptance is suggestive. No code or data are provided, and IACT estimates lack error bars.\n\nNet: the acceptance result is solid and the paper frames the problem well. The convergence theorem is not proven as stated; it needs repair. A serious referee should engage with this—the gap looks closable with a strengthened assumption—but the paper should not be accepted in its current form.","headline":"A genuinely useful paper on sparse MALA-within-Gibbs, but the main convergence theorem rests on a missing Lipschitz hypothesis and needs a fix before it is citable as proven.","tokens_in":34068,"tokens_out":2868,"would_cite":true,"duration_ms":27972,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","65C05","60J22"],"pacs":[],"model":"deepseek-v4-flash","headline":"MALA-within-Gibbs samplers can keep acceptance, step size, and convergence rate independent of the overall dimension when the target has sparse conditional structure and block-wise log-concavity.","keywords":["MALA-within-Gibbs","sparse conditional structure","dimension independence","block-wise log-concavity","Bayesian inverse problems","Markov chain Monte Carlo","convergence rate","partial updating"],"falsifier":"Take a family of block-wise log-concave targets with sparse conditional structure and bounded log-gradients (for instance, densities with support on a fixed bounded cube), keep the block size fixed, and increase the number of blocks $m$ while holding everything else fixed. If the largest step size that keeps the average block acceptance above a chosen threshold decreases with $m$, or if the integrated autocorrelation time per block grows with $m$, then the claimed dimension-independent $\\tau_0$ and convergence rate are contradicted. A more targeted check: compute the coupled-contraction factor at a fixed $\\tau<\\tau_0$ for increasing $m$ and see whether the measured rate approaches $1-(1-\\delta)\\lambda_H\\tau$ uniformly.","tokens_in":32987,"feed_emoji":"🎲","tokens_out":11748,"duration_ms":113179,"temperature":0.7,"pith_summary":"Markov chain Monte Carlo samplers normally slow down as the dimension of the target distribution grows, and standard scaling results say the step size must shrink. This paper argues that for the Metropolis-adjusted Langevin algorithm used block by block within a Gibbs scan (MALA-within-Gibbs), that slowdown disappears when the target density has sparse conditional structure and the sampler updates the state in blocks matched to that structure. The central theorem shows that under this structure, together with bounded gradients of the log-density and block-wise log-concavity, the per-block acceptance probability is at least $1-M\\sqrt{\\tau}$ and the sampler converges geometrically with a rate independent of the number of blocks. If correct, this gives a concrete route to sampling high-dimensional Bayesian posteriors whose conditional dependence is local, such as spatial fields with short correlation lengths, without re-tuning the algorithm as the domain grows. The paper's two numerical experiments, on a log-Gaussian Cox point process and an elliptic PDE inverse problem, exhibit the predicted dimension independence in practice.","feed_headline":"Sparse local structure keeps this sampler's step size dimension-free","feed_subtitle":"MALA-within-Gibbs also keeps its mixing rate independent of the number of blocks, under stated assumptions.","key_machinery":"The engine of the argument is sparse conditional structure, Assumption 3.1: the Hessian of the log-target satisfies $\\nabla^2_{x_k,x_j}\\log\\pi(x)=0$ for $k\\notin I_j$, with $|I_j|\\le S$ and block sizes bounded by $q$, all independent of $m$. This is what makes a block update low-dimensional. On top of it, block-wise log-concavity (Assumption 3.5) gives a uniformly negative matrix $H(x)$ dominating the block Hessian, which provides the contraction constant $\\lambda_H$ in the convergence rate. The proof technique is a maximal coupling: two chains share proposal noises $\\xi^k_j$, and their accept/reject decisions are coupled through a common uniform variable, which lets the analysis separate the four accept/reject scenarios. The block-distance vector $D^k=(\\|x^k_1-z^k_1\\|,\\ldots,\\|x^k_m-z^k_m\\|)$ is then shown to satisfy a sparse matrix inequality whose operator norm is bounded by $1-(1-\\delta)\\lambda_H\\tau$ for small $\\tau$, giving the dimension-free contraction.","core_discovery":"The central discovery is that the dimension dependence of MALA can be moved out of the sampler entirely when the target's log-density has a sparse Hessian. Under Assumption 3.1, each block gradient $v_j(x)=\\nabla_{x_j}\\log\\pi(x)$ depends only on at most $S$ blocks, so a single block update is a genuinely low-dimensional move no matter how large the full vector is. Assumption 3.2 keeps the gradient and its derivatives bounded by dimension-free constants, which makes the acceptance probability at each block close to one in expectation, $E[\\alpha_j]\\ge 1-M\\sqrt{\\tau}$, with $M$ independent of $m$. With block-wise log-concavity, Theorem 3.6 gives the convergence rate: for any $\\delta>0$ there is $\\tau_0>0$ independent of $m$ such that for $\\tau<\\tau_0$, two coupled MALA-within-Gibbs chains satisfy $$\\sum_{i=1}^m \\left(E\\|x_i^k-z_i^k\\|\\right)^2 \\le \\left(1-(1-\\delta)\\lambda_H\\tau\\right)^{2k}\\sum_{i=1}^m \\left(E\\|$x_i^{0}$-$z_i^{0}$\\|\\right)^2.$$ Starting one chain at the target distribution shows the other converges to it at this dimension-free geometric rate. The theorem is an extension, in the paper's reading, of dimension-independent Gibbs convergence for Gaussian targets with sparse precision matrices to a broader class of block-wise log-concave non-Gaussian targets.","pith_inferences":["Editorial inference: the bounded-gradient assumption is probably not necessary in full strength; since the proof bounds acceptance and contraction locally on the active block set, one could extend the argument to gradients that grow sublinearly in the block size and test numerically whether Gaussian sparse-precision targets, excluded by Assumption 3.2, still show dimension-free acceptance.","Editorial inference: the result elevates coordinate choice to the main design task; the practical recipe suggested by the paper is to search for coordinates in which the posterior Hessian is approximately sparse (as in the Karhunen-Loève parameterization of the elliptic example), which links the sampler to localization strategies used in data assimilation.","Editorial inference: the cost formula (integrated autocorrelation time times number of blocks) makes a testable prediction: for a fixed correlation length, the block size minimizing cost per effective sample should be nearly independent of domain size, so the optimum found at $L=16$ should persist at $L=64$ and beyond; this can be checked by experiment."],"forward_implications":["For any target satisfying Assumptions 3.1, 3.2, and 3.5, the same step size and per-block acceptance behavior work at dimension $n=mq$ as at small $m$, so tuning does not need to be redone as the number of blocks grows.","The sampler's mixing time is dimension-free in the sense of the contraction bound: after $k$ Gibbs cycles the worst-case distance to stationarity shrinks by $(1-(1-\\delta)\\lambda_H\\tau)^{2k}$, independent of $m$.","In Bayesian inverse problems with local prior correlations and local observations, the result says the sampler's speed is governed by the conditional neighborhood size, not by the total state dimension.","Computational cost per effective sample still scales with the number of blocks if each block update requires a full forward model evaluation; the paper identifies the block-size trade-off between cheaper per-step cost (fewer blocks) and lower autocorrelation (more blocks).","The numerical experiments show the predicted dimension independence of integrated autocorrelation time and step size in practice, even in a case where block-wise log-concavity is not verified."],"supporting_citations":[{"why":"Supplies the lemma that zero blocks in the log-density Hessian are equivalent to conditional independence, which is the content of Assumption 3.1.","marker":"[46]"},{"why":"Proves dimension-independent convergence of Gibbs samplers for Gaussian distributions with sparse conditional structure, the result Theorem 3.6 generalizes; the appendix also uses its matrix-norm bound.","marker":"[33]"},{"why":"Provides the optimal-scaling baselines showing that the step size of standard MALA must decrease with dimension, the obstruction that sparse conditional structure is claimed to remove.","marker":"[6, 7, 42, 43]"},{"why":"Analyzes optimal scaling for partially updating MCMC algorithms and gives the general recommendation that block updates should be high-dimensional; the paper positions its sparse-structure result against this baseline.","marker":"[36]"}],"fun_headline_variants":["Sparse Hessian unlocks dimension-free MALA-within-Gibbs","Dimension-free MCMC sampling via sparse conditional structure","MALA-within-Gibbs scales to high-dim with sparse blocks","Sparse structure removes dimension dependence from sampler","Block-wise log-concavity yields dimension-free mixing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Assumption 3.2: the gradient of the log-density and its first derivatives are bounded by constants independent of the overall dimension. The paper itself notes this excludes Gaussian targets, the most natural examples of sparse conditional structure, because their gradients are unbounded; if that boundedness fails, the dimension-independence proof does not apply.","fun_headline_variants_meta":{"raw":{"variants":["Sparse Hessian unlocks dimension-free MALA-within-Gibbs","Dimension-free MCMC sampling via sparse conditional structure","MALA-within-Gibbs scales to high-dim with sparse blocks","Sparse structure removes dimension dependence from sampler","Block-wise log-concavity yields dimension-free mixing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000376,"raw_usage":{"total_tokens":2065,"prompt_tokens":1068,"completion_tokens":997,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":684,"completion_tokens_details":{"reasoning_tokens":915}},"tokens_in":684,"tokens_out":997,"duration_ms":10126,"temperature":1.0,"reasoning_tokens":915,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:11:25.751169+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a family of block-wise log-concave targets with sparse conditional structure and bounded log-gradients (for instance, densities with support on a fixed bounded cube), keep the block size fixed, and increase the number of blocks $m$ while holding everything else fixed. If the largest step size that keeps the average block acceptance above a chosen threshold decreases with $m$, or if the integrated autocorrelation time per block grows with $m$, then the claimed dimension-independent $\\tau_0$ and convergence rate are contradicted. A more targeted check: compute the coupled-contraction factor at a fixed $\\tau<\\tau_0$ for increasing $m$ and see whether the measured rate approaches $1-(1-\\delta)\\lambda_H\\tau$ uniformly.","supporting_citations":[{"cited_title":"Spantini, D","cited_arxiv_id":null,"evidence_quote":"Supplies the lemma that zero blocks in the log-density Hessian are equivalent to conditional independence, which is the content of Assumption 3.1."},{"cited_title":"Morzfeld, X","cited_arxiv_id":null,"evidence_quote":"Proves dimension-independent convergence of Gibbs samplers for Gaussian distributions with sparse conditional structure, the result Theorem 3.6 generalizes; the appendix also uses its matrix-norm bound."},{"cited_title":"Neal and G","cited_arxiv_id":null,"evidence_quote":"Analyzes optimal scaling for partially updating MCMC algorithms and gives the general recommendation that block updates should be high-dimensional; the paper positions its sparse-structure result against this baseline."}],"review_version":1}