{"id":"8f9a82f9-665f-4058-820c-2e0f31e59da4","arxiv_id":"2607.23788","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A dimension-scheduled multi-chain controller automatically selects diagonal vs low-rank HMC preconditioners from warmup evidence and outperforms fixed Fisher and diagonal baselines on ESS per gradient.","lead":"An automatic multi-chain HMC warmup controller starts diagonal and promotes to low-rank-plus-diagonal mass matrices only when within/between-chain evidence supports it. It beat fixed Fisher low-rank and Welford diagonal warmups on ESS per gradient in headline NUTS benchmarks and turns disagreement or nonlinearity into concrete next-step advice.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline ESS/gradient ratios (2.451–22.572) are reported as bare geometric means with no across-seed uncertainty, and the paper's own paired ablations in Appendix C show the low-rank advantage collapsing to ≈1.0 once a confidence-interval go/no-go criterion is applied.","rationale":"The reader identified the uncalibrated threshold-to-theorem gap as the weakest assumption. That gap is real, but the author flags it explicitly and scopes the theory as conditional; it does not bear on the empirical benchmark claim, which is what the strongest_claim asserts. The more load-bearing issue for that claim is the absence of uncertainty quantification on the headline ratios combined with the paper's own Appendix C null results under its stated go/no-go criterion, plus the warmup-budget confound in the pooled ESS/gradient metric and the favorable-target selection. This is a sharpening, not a reversal: the reader's CONDITIONAL verdict already rests on the empirical claim holding as stated, and my concern says the support for that claim is thinner than the point estimates suggest but not contradicted. The remedy (seed-cluster intervals from the existing frozen corpus) is cheap and decisive, which is exactly what a CONDITIONAL verdict should be conditioned on. Hence UNCHANGED rather than REJECT: nothing indicates the numbers are wrong, only that their robustness is unestablished. I recommend the condition be amended to include per-seed uncertainty reporting on Table 1 alongside the reader's threshold-calibration condition.","tokens_in":17844,"tokens_out":2735,"duration_ms":102755,"concrete_test":"Recompute Table 1 from the frozen per-seed corpus exactly as Appendix C does for the ablations: per-seed paired ratios, geometric mean, and seed-cluster t 95% interval for each of the four headline cells, applying the paper's own go/no-go criterion (interval must exclude 1). If the interval for German-credit vs Fisher low-rank (1.951) includes 1, the headline efficiency claim is not distinguishable from seed noise; if all four exclude 1, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is empirical: on two NUTS targets the controller beat prespecified Fisher low-rank and Welford diagonal warmups in pooled ESS/gradient. Three soft spots compound. (1) Table 1 reports only geometric-mean ratios — no seed count, no per-seed spread, no interval. Yet Appendix C shows the authors know how to do this rigorously: the k=3 matched ablation reports a geometric mean of 1.037 with seed-cluster t-interval [0.998, 1.077] and explicitly declares a null because the interval includes 1; the k=2 matched-diagonal ablation gives ratios 0.869–1.170, i.e., mostly no gain from the low-rank structure the controller selects. So in the controlled settings where uncertainty is quantified, the mechanism implied by the headline (low-rank selection → large efficiency gain) does not reproduce. (2) The two headline targets are both favorable regimes: a synthetic ill-conditioned Gaussian is the canonical case where diagonal fails (making the 22.572 vs Welford ratio close to a foregone conclusion) and German credit is a mild, well-studied posterior. 12/12 low-rank selection on two low-rank-friendly targets does not test selectivity. (3) The pooled ESS/gradient metric charges warmup gradients, and the controller — unlike the 312-transition baselines — also chooses warmup length via its dimension-derived schedule. The ratio therefore conflates estimator selection with budget allocation; a controller that simply stops warmup earlier wins the metric with an identical final preconditioner. None of this contradicts the claim as stated, but the claim's evidentiary base is two targets, no uncertainty quantification, and an internal null result on the implied mechanism.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper proposes an automatic multi-chain warmup controller for Euclidean HMC/NUTS that starts with a diagonal inverse mass matrix and, at dimension-derived window endpoints fixed before sampling, decides whether to promote to a low-rank-plus-diagonal matrix, gather another window, retain a within-region matrix with a population/tempering advisory (persistent within/between disagreement), or advise reparameterization (poor held-out score–position linearity). The formal development gives a route-indexed attractor theorem with an explicit five-term finite-window error budget, operator-confidence consequences for rank/subspace decisions, an exact within/between covariance decomposition, and a finite-horizon Le Cam-style local-transcript lower bound. Empirically, on two NUTS benchmarks the controller selected low rank in 12/12 runs and reports geometric-mean pooled ESS/gradient ratios of 2.451–22.572 against prespecified Fisher low-rank and Welford diagonal baselines, with all 36 runs passing quality checks; a controlled GMM sweep holds the marginal spectrum fixed while the within/between evidence and selection change.","tokens_in":18297,"tokens_out":3358,"duration_ms":63190,"significance":"If the empirical claims hold up under uncertainty quantification, this is a useful contribution: an automatic, budget-driven warmup controller that removes a real manual choice (preconditioner structure) from HMC practice, merged into BlackJAX with pinned revisions, frozen manifests, and checksummed outputs. The preregistered baseline roles, float64/explicit-seed discipline, uniform quality gates, and — unusually — the authors' own reported null results in Appendix C are genuine strengths. The theoretical sections are conditional and honestly scoped, and Theorem 6's local-transcript bound is a clean formulation of a real limitation. However, the theory does not currently certify the deployed controller (by the authors' own statement), and the headline efficiency numbers rest on a confounded metric reported without uncertainty, so the practical significance is not yet established at the level the abstract asserts.","major_comments":[{"comment":"Table 1 reports only bare geometric-mean ESS/gradient ratios: no seed count, no per-seed spread, no interval. Appendix C shows the authors can and do apply a stricter standard: the k=3 matched ablation reports geomean 1.037 with seed-cluster t-interval [0.998, 1.077] and declares a null because the interval covers 1. The headline numbers should meet the same standard: report seeds, per-seed ratios, and intervals for all four Table 1 cells, and state a predeclared go/no-go criterion.","section":"§7, Table 1"},{"comment":"The pooled ESS/gradient metric charges all warmup gradients, but the controller chooses its warmup length via the dimension-derived schedule while both baselines run a fixed 312 transitions per chain. A controller that simply stops warmup earlier wins this metric with an identical final preconditioner. The ratio therefore conflates estimator selection with budget allocation. A matched-budget arm (e.g., the controller's frozen metric plugged into the 312-transition schedule, and vice versa, or sampling-only ESS/gradient) is needed to attribute the gain.","section":"§7 / §3"},{"comment":"The matched ablations undermine the mechanism implied by the headline: where low rank is selected under controlled conditions, low-rank/matched-diagonal projection ESS/gradient ratios are 0.869–1.170 (k=2) and 1.037 [0.998, 1.077] (k=3, declared a null). The Discussion concedes 'no general efficiency gain' from the low-rank structure, but the abstract's framing (12/12 low-rank selection alongside 2–22× gains) implies the selection drives the efficiency. These need reconciling: either the headline gains are schedule/allocation effects, or evidence is needed that the low-rank benefit transfers outside the GMM testbed.","section":"Appendix C vs §7"},{"comment":"§3 defines the Fisher-HMC reference prescription as the 30/55/15 schedule of Seyboldt et al. with L=10/L=80 refresh periods, memory wipes, and step-size reinitialization. Yet Table 1's 'Fisher low-rank warmup' baseline used 312 transitions under the same proportional growing-window schedule as the Welford baseline — i.e., the Fisher estimator transplanted into a Stan-like schedule, not the prespecified Fisher-HMC warmup. The comparison should be relabeled accordingly, or an arm running the actual Fisher prescription should be added; as written the baseline name overstates what was compared.","section":"§3 / §7"},{"comment":"The advertised 12/12 low-rank selection rate is computed on two targets where low rank is expected a priori (an ill-conditioned Gaussian, where the 22.572 ratio vs diagonal is close to foregone, and a mild logistic posterior). This demonstrates promotion but not discrimination: no headline target where diagonal is the right answer is included. The GMM sweep partially tests selectivity but at d=5 with the marginal spectrum pinned. At least one benchmark where diagonal retention is correct would make the selection claim evidentiary rather than assumed.","section":"§7 / §5.4"}],"minor_comments":[{"comment":"4,996 warmup divergences are reported without comment. Even with zero post-warmup divergences, this count deserves discussion: where in the schedule they occur, whether the 20 gradients/transition allowance or window transitions drive them, and whether they affect the evidence tests.","section":"§7"},{"comment":"The SR=10 row shows 1/2 low-rank/diagonal after 0/3 at SR=9 and 9.5 — non-monotone selection at the extreme of the sweep. A sentence on whether this is threshold noise or a real boundary effect would help.","section":"Appendix C, Table 3"},{"comment":"'Universal' in the title is explicitly scoped in §1 to the evaluated HMC-family kernels, which is appropriate, but the abstract does not carry that qualification. Consider scoping the abstract claim as the introduction does.","section":"Title/abstract"},{"comment":"The one-output estimand's comparator is a historical single-chain 2,500-warmup policy, explicitly 'descriptive and not budget-identical.' Since it is not budget-identical, consider moving it out of Table 2 or visually separating it to avoid readers averaging across the two columns.","section":"Appendix B, Table 2"},{"comment":"Figure 2 references a 'population/regional companion, Paper 3,' an unpublished companion paper. The manuscript should stand alone; replace with a generic pointer to population/tempering methods.","section":"Figure 2"},{"comment":"The illustrative Wald radius in Eq. (7) is well-hedged, but the support-recovery assumption ((I−QQᵀ)u = o_p(n^{-1/2})) is strong for a singular Γ_a; a remark on when this is verifiable from the transcript would be useful.","section":"§5.2, Eq. (7)"},{"comment":"Rendering artifacts: 'bR' appears throughout where R̂ is intended (an ArviZ/Unicode issue); 'LR V' is split. Please fix for readability.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"Single-author manuscript with an unusually detailed generative-AI use disclosure; the author accepts responsibility and the code is pinned and checksummed, which mitigates provenance concerns. The text contains forward references to an unpublished companion (\"Paper 3\" in Figure 2); the editor may wish to ask that the manuscript stand alone. Novelty over Bales et al. (2019) and Seyboldt et al. (2026) is real but incremental on the estimator side; the contribution is the controller and evidence gating. The evidentiary asymmetry between the headline table and Appendix C is the central editorial question: the authors clearly know how to quantify uncertainty, and should be held to that standard throughout."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: Lao ships a working multi-chain HMC warmup controller (BlackJAX, commit-locked) that starts diagonal, promotes to low-rank-plus-diagonal from within/between evidence at fixed dimension-derived windows, and turns disagreement or score–position nonlinearity into concrete handoffs. That is useful tooling, not a foundational theorem.\n\nWhat is actually new is the evidence-gated route selection plus abstention/handoff logic, not low-rank mass matrices themselves (Stan windows, Bales, Seyboldt Fisher are all cited cleanly). The paper does several things well. It keeps attraction, structural ID, representativeness, and utility separate. The controlled GMM holds the marginal spike fixed while the within/between signal and selected estimator move—that is the right selection test. Theory (route latch, operator consequences, local-transcript Le Cam bound) is scoped and the author flags that deployed fixed thresholds are not yet calibrated to the five-term error budget. Code, manifests, and preregistered baseline roles are real.\n\nSoft spots, in proportion. Table 1’s geometric-mean ESS/gradient ratios (up to ~22× vs diagonal, ~2× vs Fisher) have no seed spread. The author’s own Appendix C ablations, where intervals are reported, show the low-rank vs matched-diagonal advantage collapsing near 1.0 and a declared null. So the mechanism implied by the headline is weaker than the headline numbers. Only two NUTS targets, both friendly to low-rank structure; 12/12 selection does not prove selectivity. Pooled ESS/gradient also mixes estimator choice with the controller’s own schedule length versus fixed 312-transition baselines. None of that falsifies the stated claim, but it means the efficiency story needs broader targets and uncertainty before you trust the multipliers.\n\nWho it is for: people who build or tune HMC stacks and care about mass-matrix automation and failure modes. Serious referee material—compositional novelty, reproducible artifacts, honest limits. I would engage, cite the controller design if I were writing on adaptation, and push for calibration plus harder posteriors in revision.","headline":"Real shipped auto-warmup controller with honest theory–practice gap; headline ESS ratios look softer once you read the author’s own ablations.","tokens_in":19335,"tokens_out":528,"would_cite":true,"duration_ms":18786,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C05","62F15","65C40"],"pacs":[],"model":"grok-4.5","headline":"An automatic multi-chain HMC warmup controller starts diagonal, promotes to low-rank when evidence supports it, and turns inconclusive or conflicting signals into concrete next actions.","keywords":["Hamiltonian Monte Carlo","warmup adaptation","mass matrix","preconditioner selection","NUTS","low-rank inverse mass","multi-chain diagnostics","ESS per gradient"],"falsifier":"Re-run the preregistered headline NUTS configurations (ill-conditioned Gaussian and German-credit logistic regression, M=8, same gradient budgets and seeds): if the automatic controller fails to select low rank in most runs, or its pooled ESS-per-gradient geometric-mean ratios fall to or below 1 against the Fisher low-rank and Welford diagonal baselines while quality checks still pass, the central efficiency-after-selection claim fails.","tokens_in":19013,"feed_emoji":"🎯","tokens_out":1099,"duration_ms":25613,"temperature":0.7,"pith_summary":"Euclidean HMC needs a step size and a constant preconditioner chosen from short, nonstationary warmup draws, yet standard practice still forces the user to pick the mass-matrix structure in advance and follow a fixed schedule. This paper offers one multi-chain controller that begins with a diagonal inverse mass matrix, checks evidence only at dimension-derived window endpoints, and either promotes to a low-rank-plus-diagonal matrix (choosing the retained rank under sample-support caps) or stays diagonal. When the evidence is inconclusive it simply gathers the next scheduled window; when chains persistently disagree it keeps a within-region matrix and advises a population or tempering method; when score–position linearity is poor it advises reparameterization. On the headline NUTS benchmarks the controller selected low rank in every run and delivered higher pooled ESS per gradient than both a prespecified Fisher low-rank warmup and a Welford diagonal warmup, with all runs passing the post-warmup quality check. A sympathetic reader cares because the method collapses several common warmup heuristics into one evidence-driven path and converts familiar warning signs into actionable guidance instead of silent failure.","feed_headline":"HMC warmup that picks its own preconditioner—and beats fixed baselines","feed_subtitle":"A multi-chain controller starts diagonal, promotes to low rank on evidence, and turns warnings into next steps","key_machinery":"The route-indexed attractor: after a confidence-aware latch promotes an estimator (a “route”), finite-window iterates of that estimator’s population update contract toward its fixed target, with an explicit additive budget for starting-law, step-size, whitening, sampling, and regularization error; eligibility is gated by within-chain and between-means structural tests plus a held-out score–position linearity check.","core_discovery":"A single automatic controller can choose, from multi-chain warmup evidence alone, whether a constant diagonal or low-rank-plus-diagonal inverse mass matrix is supported, which rank to keep, and when to stop adapting the mass matrix—while preserving step-size adaptation and turning persistent within/between disagreement or poor score–position linearity into explicit advice for population methods or reparameterization—and on the evaluated NUTS benchmarks this controller selected low rank in 12/12 runs and improved pooled ESS per gradient relative to both prespecified Fisher low-rank and Welford diagonal baselines.","pith_inferences":["The same within/between evidence split could gate adaptation in other multi-chain MCMC kernels that currently freeze a single metric by hand.","Calibrating the five error terms in the attractor bound would turn the advisory flags into statistically timed stopping rules rather than fixed-threshold heuristics.","The controlled GMM sweep suggests a practical diagnostic for when constant Euclidean geometry is simply the wrong global description, independent of any particular sampler.","Budget-identical comparisons against single-chain historical policies would clarify how much of the reported gain is multi-chain information versus schedule redesign."],"forward_implications":["Users can drop the manual choice of diagonal versus dense/low-rank mass matrix for ordinary Euclidean HMC/NUTS warmup.","Persistent within-/between-chain disagreement becomes an explicit handoff signal to population or tempering methods rather than a silent failure of a single constant preconditioner.","Poor held-out score–position linearity becomes an explicit reparameterization signal while still leaving a usable within-region matrix.","Schedule and rank caps become dimension- and budget-derived defaults instead of per-model tuning knobs.","Once thresholds are calibrated to the finite-window error budget, the same controller can be composed with held-out metric ranking or a regional-exploration companion."],"fun_headline_variants":["HMC controller auto-picks diagonal or low-rank mass matrix from warmup evidence","Multi-chain warmup selects preconditioner rank and structure without manual schedules","Evidence-driven HMC warmup chose low rank in 12/12 NUTS benchmarks","Controller turns HMC preconditioner disagreement into population-method advice","Automatic HMC mass-matrix selection beats Fisher low-rank and Welford baselines"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The controller’s fixed diagnostic thresholds are treated as adequate stand-ins for the theory’s confidence sets and five-term finite-window error budget, even though the paper states that connecting the two still requires calibration.","fun_headline_variants_meta":{"raw":{"variants":["HMC controller auto-picks diagonal or low-rank mass matrix from warmup evidence","Multi-chain warmup selects preconditioner rank and structure without manual schedules","Evidence-driven HMC warmup chose low rank in 12/12 NUTS benchmarks","Controller turns HMC preconditioner disagreement into population-method advice","Automatic HMC mass-matrix selection beats Fisher low-rank and Welford baselines"]},"model":"grok-4.5","effort":"low","cost_usd":0.003799,"raw_usage":{"total_tokens":1277,"prompt_tokens":857,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":37988000,"prompt_tokens_details":{"text_tokens":857,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":336,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":857,"tokens_out":84,"duration_ms":7335,"temperature":1.0,"reasoning_tokens":336,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T12:26:49.103434+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the preregistered headline NUTS configurations (ill-conditioned Gaussian and German-credit logistic regression, M=8, same gradient budgets and seeds): if the automatic controller fails to select low rank in most runs, or its pooled ESS-per-gradient geometric-mean ratios fall to or below 1 against the Fisher low-rank and Welford diagonal baselines while quality checks still pass, the central efficiency-after-selection claim fails.","supporting_citations":[],"review_version":1}