{"id":"9e6c9fdb-9890-4442-aca9-1374bceb370c","arxiv_id":"2507.09559","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Variational inference with alpha-divergence matches Hamiltonian Monte Carlo on in-sample fit for spatial lifetime models at roughly half the computation time, but with poorly calibrated uncertainty for some parameters.","lead":"This paper tests whether variational inference, an approximation method, can substitute for slow gold-standard sampling in Bayesian models of time-to-event data with spatial correlations. On two real datasets, the alpha-divergence version gives similar model fit in about half the computing time, though its uncertainty estimates are not always trustworthy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mean-field VI uncertainty is not calibrated against HMC in the pine case, so 'comparable inference performance' is established only for in-sample fit, not for posterior inference.","rationale":"The reader's verdict (CONDITIONAL) is appropriately calibrated. My stress test identifies the same load-bearing weakness: the mean-field variational family is assumed rich enough for reliable posterior inference, but no calibration evidence supports that assumption. The concrete point-estimate/SE comparisons in Tables 4 and 6 make this more than a theoretical worry: discrepancies of many VI standard errors for the baseline hazard and physical-region coefficients indicate that VI posterior widths are too narrow under the independence assumption. In-sample NLL and residual plots are real evidence for point predictive fit, and I credit the paper for reporting them, but they do not test uncertainty calibration or random-effect accuracy. The proposed SBC test is a single, standard check that would settle whether the overconfidence is systematic. The cited supplementary sections are absent from this preprint, so the claimed h-selection and alpha-stability results cannot currently be verified. Thus the verdict should remain CONDITIONAL: the computational speed advantage is plausible and the point estimates are often close, but practitioners should not rely on VI credible intervals until a calibration or coverage study is reported. I do not see a basis for outright rejection: the paper is transparent about the limitation and frames the contribution as a case study.","tokens_in":19212,"tokens_out":6749,"duration_ms":77706,"concrete_test":"Run simulation-based calibration on the spatial PH model: treating fitted values (or prior draws) as truth, simulate R=50 datasets with m=60 sites and similar censoring; rerun the same alpha=0.8 mean-field VI; compute rank of true a, b, beta, s2_gamma, nu, and a subsample of gamma_i among VI posterior draws. If 95% VI interval coverage falls below ~80% or the rank histogram is non-uniform, the Tables 4/6 uncertainties are unreliable. Secondary check: report HMC chain length, warmup, tree depth, acceptance, and whether reported times include warmup.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that alpha-divergence VI has inference performance comparable to HMC rests partly on the assertion that VI posterior means and uncertainties can be used for parameter inference. That assertion is weakest at the mean-field assumption. The paper itself concedes in Section 5 that independent variational distributions can bias spatial random-effect estimation, but it never displays a direct comparison of the gamma random effects between alpha-VI and HMC. The numerical evidence in Tables 4 and 6 already suggests severe undercoverage relative to the HMC reference: for the pine data, baseline hazard a is 0.0444 (VI SE 0.0023) vs 0.0243 (HMC SE 0.0240), a discrepancy of about 8.7 VI SEs; the Coastal Plain physical-region coefficient is 0.0273 (VI SE 0.0501) vs 0.5873 (HMC SE 0.7354), about 11 VI SEs. Similar patterns appear in the GPU results, e.g., Constant and sigma differ by roughly 5-6 VI SEs. Close in-sample NLL values (4481.04 vs 4470.13; 11149.40 vs 11147.80) and residual plots establish that the optimized variational point predictor fits about as well as HMC in sample, but they do not validate posterior intervals or random-effect estimates. Absent a calibration check, the 'uncertainty comparisons' claimed in Section 5 are not supported, and the central claim is overbroad.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the use of variational inference (VI) for Bayesian inference in spatially correlated lifetime data. Two model families are considered: a cumulative-exposure/accelerated-failure-time model for the Titan GPU data and a proportional-hazards model with Weibull baseline for the pine-tree data. The authors implement black-box VI with KL divergence and Rényi α-divergence under a mean-field variational family (diagonal Gaussian components plus independent inverse-gamma components for variance and length-scale parameters), and compare against Hamiltonian Monte Carlo (NUTS in Stan). On the two case studies, α-divergence VI with α=0.8 achieves negative log-likelihood values close to HMC (4481.04 vs 4470.13 for GPU; 11149.40 vs 11147.80 for pine) with roughly half the reported computing time, whereas KL-divergence VI performs markedly worse. The paper concludes that α-divergence VI has inference performance comparable to HMC at lower computational cost.","tokens_in":19401,"tokens_out":7362,"duration_ms":79070,"significance":"The practical question addressed—whether VI can substitute for expensive MCMC in spatial survival models—is relevant, and the case studies use real, publicly available data. The manuscript is transparent about the mean-field independence assumption and its potential bias, and it provides detailed tables and residual plots. If the uncertainty calibration issue can be resolved or the claims appropriately narrowed, the demonstration of a roughly 2x speedup with similar in-sample fit would be useful for practitioners. However, in its current form the central claim is stronger than the evidence: the reported standard errors suggest systematic overconfidence in several VI estimates, and no direct comparison of spatial random effects is presented. These gaps are load-bearing for 'comparable inference performance,' not merely presentation issues.","major_comments":[{"comment":"The conclusion in Section 5 that α-divergence VI has 'comparable inference performance as HMC' is not supported for posterior uncertainty. Table 6 reports pine baseline a = 0.0444 (SE 0.0023) for VI versus 0.0243 (SE 0.0240) for HMC, a difference of roughly 8.7 VI standard errors, and Physical Region (Coastal Plain) = 0.0273 (SE 0.0501) versus 0.5873 (SE 0.7354); Table 4 shows analogous discrepancies, e.g., Constant 2.0235 (0.0166) versus 1.9306 (0.0278) and σ 0.2289 (0.0056) versus 0.2038 (0.0053). These patterns indicate that the mean-field VI credible intervals are narrower than the HMC posterior, i.e., overconfident. The NLL comparisons and residual plots establish in-sample predictive fit, not calibrated posterior intervals. Since Section 5 itself acknowledges that the independence assumption can bias random-effect estimation, the central claim needs to be either accompanied by a calibration comparison (e.g., coverage of HMC posterior means under VI intervals, or interval-overlap diagnostics) or explicitly restricted to point prediction and in-sample fit.","section":"Section 5, Tables 4 and 6"},{"comment":"No direct comparison of spatial random-effect estimates between VI and HMC is given; Figures 5 and 9 display only the α-divergence VI random-effect maps, and Figures 3 and 7 compare residuals rather than random effects. The random effects are the mechanism for spatial correlation, and the Section 5 discussion specifically identifies them as the component most likely to be biased by the mean-field assumption. To support the claim of comparable inference, the authors should report a direct comparison (e.g., a scatterplot of VI versus HMC posterior means, summary of differences, and interval coverage) for the random effects in both case studies.","section":"Section 4.3, Figures 5 and 9"},{"comment":"The computational-efficiency claim is central, but the reported times lack implementation context. The manuscript does not state the software versions, hardware, number of HMC chains, warm-up and post-warm-up draws, VI optimizer settings, or whether all methods were run on the same machine. Table 3 ('Time (minutes) 42.7/55.23/86.53') and Table 5 ('Time (hours) 3.19/5.25/7.79') therefore cannot be independently assessed, and the 2x speedup may not be robust to tuning choices. Please provide a reproducibility appendix with these details and, if possible, timing variability across repeated runs.","section":"Section 4.1, Tables 3 and 5"}],"minor_comments":[{"comment":"Equation (2) uses φ and Φ as a generic location-scale density and cdf, but Section 4.2.1 assumes a standard smallest extreme value error. Please state explicitly that φ and Φ are the SEV density and cdf (or write the Weibull likelihood directly) to avoid ambiguity with the normal-based notation.","section":"Section 3.1.1, Equation (2)"},{"comment":"The crown-class rows list counts whose sum is 20,271 (144.4% of 14,044); please clarify whether these are per-follow-up-tree records rather than per-tree counts, and adjust the N (%) column accordingly.","section":"Table 2"},{"comment":"The data availability statement contains truncated URLs; please provide the complete links and, if available, the analysis code used to produce Tables 3–6 and Figures 3–9.","section":"Data Availability Statement"},{"comment":"The manuscript repeatedly references 'Supplementary Section 1/2/3' but the supplementary material is not included; please make it available for review and cite the relevant results explicitly.","section":"Sections 4.2 and 4.3"}],"recommendation":"major_revision","confidential_remarks":"I see no issue of novelty disclosure; the variational machinery is imported from prior work and the contribution is case-study evidence. The main concern is scope: the authors' claim 'comparable inference performance' is broader than what is shown. The paper fits a case-study-oriented journal; with the requested calibration analyses and narrowed claims it could be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid case-study paper that demonstrates something real. Alpha-divergence VI with a mean-field family gets in-sample predictive fit close to HMC on two spatial lifetime models at about half the compute time, and the speed advantage grows with the number of spatial locations (Figure 6). That is genuinely useful for practitioners. The method is not new — the VR bound, black-box gradient estimator, and mean-field machinery come from Hernandez-Lobato et al. and Li and Turner — but the application to spatially correlated lifetime models is new as far as the cited literature shows, and the two real data analyses give concrete evidence. The KL divergence comparison is a nice control, and the authors are transparent that alpha and MC sample size were chosen empirically and that mean-field independence can bias spatial random-effect estimates.\n\nThe soft spot is real and central. What the numbers support is comparable in-sample fit, not comparable posterior inference. The NLL values are close (4481 vs 4470; 11149 vs 11148) and the residual plots line up, but the VI standard errors in Tables 4 and 6 are much smaller than HMC's on several parameters, with point estimates differing by many VI SEs. Pine baseline hazard a is 0.0444 (SE 0.0023) vs 0.0243 (SE 0.0240); Coastal Plain is 0.0273 (SE 0.0501) vs 0.5873 (SE 0.7354). That is the mean-field undercoverage pattern. The paper never directly compares the spatial random effects between VI and HMC, which is exactly where the authors say the bias should appear. The preprint also references Supplementary Sections 1-3 but does not include them, and the HMC configuration (chain length, tree depth, convergence diagnostics, hardware) is not reported. So the speed comparison is hard to reproduce and the calibration claim cannot be checked.\n\nI would send this out. The empirical comparison warrants referee time. But the authors should either soften the headline claim to 'comparable in-sample fit and point estimates' or add a calibration or coverage study — simulated data with known parameters, or posterior predictive checks — before this is publishable as stated.","headline":"A useful case-study demonstration that alpha-divergence VI roughly matches HMC on in-sample fit for spatial lifetime models at about 2x speed, but the 'comparable inference performance' claim outruns the evidence because the VI credible intervals look overconfident.","tokens_in":20081,"tokens_out":2445,"would_cite":true,"duration_ms":28517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","62H11","62N05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Variational inference using Rényi $\\alpha$-divergence approximates spatially correlated lifetime-data posteriors nearly as accurately as Hamiltonian Monte Carlo while taking about half the computing time.","keywords":["variational inference","alpha-divergence","spatial random effects","lifetime data","Bayesian inference","Hamiltonian Monte Carlo","accelerated failure time model","proportional hazards model"],"falsifier":"Re-estimate the pine tree model with a variational family that allows the spatial random effects to be correlated instead of independent. If the VI point estimate of the baseline hazard parameter $a$ moves from 0.0444 toward the HMC value 0.0243 by more than the reported standard errors, or if VI credible intervals cover the HMC posterior at far below the nominal rate, the claim of comparable inference would be falsified.","tokens_in":18843,"feed_emoji":"⏱️","tokens_out":9362,"duration_ms":98153,"temperature":0.7,"pith_summary":"Bayesian inference for lifetime data with spatial correlations gets slow as the number of locations grows, because sampling the posterior with Hamiltonian Monte Carlo is expensive. The paper asks whether variational inference, which turns inference into optimization, can substitute for sampling while preserving accuracy. It claims yes for the two spatial lifetime models it studies—an accelerated failure time model for Titan GPU failures and a proportional hazards model for pine tree survival—provided the divergence being minimized is Rényi $\\alpha$-divergence rather than KL divergence. With $\\alpha = 0.8$, the variational posterior gives negative log-likelihood values almost identical to HMC and runs about two times faster on both datasets. A sympathetic reader would take the paper as evidence that optimization-based approximate Bayesian inference is a viable fast route for spatial time-to-event models.","feed_headline":"Faster VI matches MCMC fit on spatial lifetime data","feed_subtitle":"Alpha-divergence variational inference hits HMC-level NLL on GPU and pine survival data, at about twice the speed.","key_machinery":"The engine is black-box variational inference built on the variational Rényi bound, $L_{\\alpha} = \\frac{1}{1-\\alpha}\\log E_{q(\\theta|\\eta)}\\left[(p(\\theta,D)/q(\\theta|\\eta))^{1-\\alpha}\\right]$, with the expectation approximated by Monte Carlo samples and the bound optimized with a stochastic gradient optimizer. The variational family is mean-field: independent Gaussian distributions for regression coefficients, log-scale parameters, and spatial random effects, plus inverse-gamma distributions for the spatial variance $s_{\\gamma}^2$ and length scale $\\nu$. Setting $\\alpha < 1$ is the crucial choice because it encourages a mass-covering approximation, whereas KL divergence (the $\\alpha\\to 1$ limit) narrows the approximation around modes; the paper attributes KL's worse fits and slower convergence to this zero-forcing behavior. Sampling from the fitted variational distribution then yields point estimates and credible intervals in the same way sampling from an MCMC posterior does.","core_discovery":"The paper's central claim, in its own terms, is that VI with $\\alpha$-divergence at $\\alpha = 0.8$ has inference performance comparable to HMC for spatially correlated lifetime data while cutting computing time by roughly half. The evidence is two case studies: for Titan GPU lifetimes the negative log-likelihood is 4481.04 for VI versus 4470.13 for HMC, with 42.7 minutes versus 86.53 minutes; for pine tree survival the NLL is 11149.40 versus 11147.80, with 3.19 hours versus 7.79 hours. Point estimates of most coefficients from $\\alpha$-divergence lie closer to HMC than those from KL-divergence VI do, and KL divergence produces visibly worse residual fits and NLL values. In the GPU subset experiments, the time gap between VI and HMC widens as the number of spatial locations increases.","pith_inferences":["Editorial extension: if the roughly 2x speedup extrapolates to thousands of spatial locations, the practical bottleneck for these models shifts from posterior sampling to the manual tuning of the Monte Carlo sample size and the value of $\\alpha$, both of which the paper leaves to empirical selection.","Editorial extension: the reported gaps in baseline-hazard and physical-region estimates between VI and HMC, with much smaller VI standard errors, suggest that a variational family with correlated spatial random effects is the natural next test; the paper names this trade-off but does not quantify it.","Editorial extension: a spatial-aware cross-validation scheme that splits sites by distance rather than randomly could make $\\alpha$ selection and out-of-sample prediction comparisons possible; the paper identifies cross-validation as open but does not implement it."],"forward_implications":["On the two datasets studied, practitioners can get near-HMC posterior summaries for spatial accelerated failure time and proportional hazards models in about half the wall-clock time by using VI at $\\alpha = 0.8$.","KL-divergence VI should not be treated as a safe default for these models, because it produced worse negative log-likelihood, worse residual plots, and slower convergence than $\\alpha$-divergence in both case studies.","The computational advantage of VI over HMC grows with the number of spatial locations, so the method has the most to offer exactly where MCMC is most expensive.","The reported estimates of GPU cage and node effects, pine thinning and region effects, spatial correlation decay curves, and random-effect maps are all obtainable from the variational fit at roughly half the computing cost of the HMC benchmark."],"supporting_citations":[{"why":"Supplies the Titan GPU lifetime dataset and the row/column cabinet spatial structure used in the GPU case study.","marker":"Ostrouchov et al. (2020)"},{"why":"Provides the modified OTB-focused Titan dataset and the Hamiltonian Monte Carlo modeling setup that serves as the benchmark for comparison.","marker":"Min et al. (2025)"},{"why":"Supplies the pine tree survival dataset and the spatial proportional hazards modeling context.","marker":"Li et al. (2015)"},{"why":"Provides the black-box $\\alpha$-divergence minimization approach that the variational inference algorithm is built on.","marker":"Hernandez-Lobato et al. (2016)"},{"why":"Introduces the variational Rényi bound and the monotonicity and mass-covering/zero-forcing properties used to justify choosing $\\alpha < 1$.","marker":"Li and Turner (2016)"},{"why":"Supplies the black-box variational inference gradient estimation technique used for the variational bound.","marker":"Ranganath, Gerrish, and Blei (2014)"},{"why":"Provides the stochastic optimizer used to maximize the variational objective.","marker":"Kingma and Ba (2014)"},{"why":"Provides the Hamiltonian Monte Carlo implementation used to produce the benchmark posterior samples.","marker":"Carpenter et al. (2017)"}],"fun_headline_variants":["Alpha-divergence VI rivals HMC on spatial survival fits, twice as fast","VI cuts runtime in half while matching HMC on spatial lifetime data","Faster spatial survival inference: VI matches HMC accuracy at half cost","Alpha-VI gives HMC-level fit on spatial survival at half compute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mean-field assumption, which forces all parameters and spatial random effects to be independent in the approximate posterior, is rich enough that the optimized variational posterior keeps both the point estimates and the uncertainty close to the true posterior.","fun_headline_variants_meta":{"raw":{"variants":["Alpha-divergence VI rivals HMC on spatial survival fits, twice as fast","VI cuts runtime in half while matching HMC on spatial lifetime data","Faster spatial survival inference: VI matches HMC accuracy at half cost","Alpha-VI gives HMC-level fit on spatial survival at half compute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000919,"raw_usage":{"total_tokens":3920,"prompt_tokens":896,"completion_tokens":3024,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":2944}},"tokens_in":512,"tokens_out":3024,"duration_ms":24118,"temperature":1.0,"reasoning_tokens":2944,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:54:41.112857+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the pine tree model with a variational family that allows the spatial random effects to be correlated instead of independent. If the VI point estimate of the baseline hazard parameter $a$ moves from 0.0444 toward the HMC value 0.0243 by more than the reported standard errors, or if VI credible intervals cover the HMC posterior at far below the nominal rate, the claim of comparable inference would be falsified.","supporting_citations":[{"cited_title":"Maxwell, R","cited_arxiv_id":null,"evidence_quote":"Supplies the Titan GPU lifetime dataset and the row/column cabinet spatial structure used in the GPU case study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the modified OTB-focused Titan dataset and the Hamiltonian Monte Carlo modeling setup that serves as the benchmark for comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the black-box $\\alpha$-divergence minimization approach that the variational inference algorithm is built on."},{"cited_title":"Gerrish, and D","cited_arxiv_id":null,"evidence_quote":"Supplies the black-box variational inference gradient estimation technique used for the variational bound."},{"cited_title":"Gelman, M","cited_arxiv_id":null,"evidence_quote":"Provides the Hamiltonian Monte Carlo implementation used to produce the benchmark posterior samples."}],"review_version":1}