{"id":"cd563f95-1f65-49a4-a88a-8b7db952d9d9","arxiv_id":"1908.04655","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Treating the power-prior parameter beta as an inferred hyperparameter lets one nested sampling run automatically cope with unrepresentative priors and yield corrected evidence estimates.","lead":"This paper presents a Bayesian method for nested sampling that treats the prior-flattening parameter beta as a hyperparameter inferred from the data, then marginalizes over it. The approach aims to make nested sampling robust to unrepresentative priors with little extra cost and uses the inferred beta distribution to diagnose bad priors.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (15)'s factorization of the effective posterior is contradicted by the paper's own Fig. 13: in the asymmetric 4-mode example, beta means differ strongly across modes, yet Eq. (13)/(15) require beta independent of theta.","rationale":"The mathematical construction in Eqs. (9)-(14) is correct in the ideal NS limit: BPR augments the target posterior with beta and the theta-marginal is exactly P(theta). The practical question is what happens when finite-N NS fails on unrepresentative priors. The paper's answer is Eq. (15), a product form that makes the theta-marginal recovery and the evidence rescaling almost trivial. My reading of Section 4.4 is that the paper itself supplies a counterexample to Eq. (15): in the asymmetric four-mode likelihood, the inferred beta values are different in different modes. A product density cannot produce posterior mode means that differ by orders of magnitude, so either the samples are not from the target (10) and the 'effective posterior' model is wrong, or the per-mode beta values are miscomputed. In either case the paper has not established that BPR's output marginalizes to P(theta) in this regime, nor that the histogram-scaled evidence in (16) is unbiased. This is a deeper issue than the evidence-correction being heuristic, although it subsumes it: even if a bin has the largest count, that bin's calibration to pi(beta) presupposes the product form. I therefore agree with the reader that the paper is promising but needs substantial revision, and I do not change the conditional verdict. A targeted independence test on the multimodal beta samples would settle whether Eq. (15) is tenable or whether the paper must weaken its claims to unimodal or globally unrepresentative cases.","tokens_in":19426,"tokens_out":19360,"duration_ms":215016,"concrete_test":"Re-run the asymmetric 4-mode example of Section 4.4 for a case in which all four modes are detected (e.g. furthest mode centre at theta1=theta2=15), with Nlive=100, 300 and 1000 and at least 20 independent realisations each. For each run, compute the per-mode beta samples and perform a permutation test (or Kruskal-Wallis) for equality of the beta distributions across modes. Under Eq. (15) the distributions must be equal; if the p-value remains small or the per-mode mean-beta differences persist as Nlive grows, the effective posterior does not factorize. Then compare the recovered per-mode theta marginals and mode weights with the analytic mixture posterior; if those are biased while the mode-dependent beta persists, BPR does not recover P(theta) in this regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 3, after deriving the exact joint posterior (10), the paper introduces an 'effective' posterior (15), P~_eff(theta,beta) proportional to P(theta) P~(beta), to model finite-N nested-sampling failure and then uses it to conclude that marginalizing beta recovers P(theta) and that the evidence correction in (16) is valid. This factorization is the load-bearing assumption for both claims in the unrepresentative-prior regime. The paper's own Section 4.4 asymmetric 4-mode example contradicts it: Fig. 13(a)-(b) reports estimated log-beta values that differ across modes by orders of magnitude, with modes 1 and 4 (far from the prior centre) having much smaller beta than modes 2 and 3. If (15) held, the conditional distribution of beta would be independent of theta, so the mean beta in every posterior mode would coincide up to Monte Carlo noise. The paper presents this mode-dependent beta as a feature, but it is direct evidence that the NS-BPR output is not distributed as (15). Consequently, the argument that marginalizing beta recovers P(theta) and the calibration of the evidence histogram to pi(beta) are not valid in precisely the multi-modal, spatially asymmetric cases advertised as advantages.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Bayesian posterior repartitioning (BPR), a method for making nested sampling (NS) more robust against unrepresentative priors. The key idea is to introduce an auxiliary hyperparameter β that controls a power transformation of the prior, and to treat β as an unknown to be sampled jointly with the original parameters θ in a single NS run. The authors show that, by construction, the joint posterior on (θ, β) is proportional to P(θ)π(β) (Eq. 10), so that marginalizing over β returns the original posterior P(θ) (Eq. 13) and the evidence is unchanged (Eq. 12). Recognizing that finite-Nlive NS may fail at extreme β values, they posit an 'effective posterior' factorized as P(θ)P~(β) (Eq. 15) and propose a histogram-based correction for the evidence (Eq. 16). Numerical experiments with MultiNest are presented for Gaussian, Laplace, and multimodal likelihoods from 1D to 10D, showing accurate parameter and evidence recovery relative to quadrature, and a negligible computational overhead for representative priors. The paper concludes that BPR should be used as the default in all NS analyses.","tokens_in":19729,"tokens_out":3979,"duration_ms":40669,"significance":"If the method holds up, BPR is a practically valuable contribution to the nested-sampling literature: it automates the choice of the power-prior exponent β, removes the annealing-schedule dependence of the earlier PR method, and offers a diagnostic (β+) for detecting unrepresentative priors. The exact construction in Eqs. (9)-(13) is correct and elegant, and the numerical validation is extensive for smooth unimodal targets. However, the central additional claims—that the effective posterior factorizes as in Eq. (15) and that the evidence correction in Eq. (16) is unbiased—are not established from first principles and are contradicted by the paper's own asymmetric multimodal example. The paper also overstates the robustness of BPR as a default method. Thus the contribution is promising but currently incomplete.","major_comments":[{"comment":"The effective posterior factorization P~_eff(θ,β) ∝ P(θ)P~(β) in Eq. (15) implies that β and θ are independent under the sampling distribution. Yet Fig. 13(a)-(b) shows that, in the asymmetric 4-mode example, the mean log β differs strongly between modes 1/4 and modes 2/3 (by orders of magnitude in Fig. 13(b)). This is direct empirical evidence that the output of NS-BPR is not distributed as Eq. (15) in exactly the multi-modal, spatially asymmetric scenario advertised as a key advantage of BPR. Because the argument that marginalizing β over the effective posterior recovers P(θ) (the basis for the parameter-recovery claim) and the evidence correction in Eq. (16) both rely on Eq. (15), the paper's central practical claims are unsupported for this regime.","section":"Section 3, Eq. (15); Section 4.4, Fig. 13"},{"comment":"The evidence correction in Eq. (16) is heuristic and is not derived from any stated principle. The procedure assumes that the histogram bin with the largest number of β samples corresponds to a region where the NS sampler is unbiased (P~(β)=π(β)), and then scales the histogram so that this bin has volume ∫β_l^β_u π(β)dβ. This is circular in the sense that the same samples used to construct the histogram are used to determine which bin is 'unbiased'. No convergence argument or error bound is given, and no test of the correction is presented for a case in which Eq. (15) is violated (e.g., the asymmetric multimodal example of Sec. 4.4). The evidence estimates in Tables 1, 3, and 4 are all consistent with quadrature, but those cases also exhibit the product form of Eq. (15); the correction's validity in the non-factorized regime is therefore untested.","section":"Section 3, Eq. (16) and surrounding discussion"},{"comment":"In the asymmetric arrangement, the default Nlive=100 yields large and volatile RMSE for modes 1 and 4 (Fig. 13(c)), and for the furthest mode centers the NS run produces no samples at all in those modes (Fig. 13(e)). The paper presents this as defining the limit of applicability, but it directly undermines the abstract's and Conclusion's claim that BPR 'allows' BPR to accommodate spatially asymmetric modes in a robust, hands-off manner. Since the negligible-overhead default-use claim is premised on Nlive=100, this is a load-bearing limitation rather than a minor caveat. The paper should either restrict the robustness claims to the symmetric case or provide a principled way to set Nlive for asymmetric problems.","section":"Section 4.4, Fig. 13(c)-(f)"}],"minor_comments":[{"comment":"Typo: 'settng β = 1' should read 'setting β = 1'.","section":"Section 1, para. 2"},{"comment":"Typo: 'movitation' should read 'motivation'.","section":"Section 3, para. after Eq. (9)"},{"comment":"The caption states the contours correspond to the '2σ (68%) and 3σ (95%)' iso-probability levels. For a Gaussian, 68% corresponds to 1σ and 95% to 2σ, so the percentages are misassigned.","section":"Figure 6 caption"},{"comment":"The captions of Fig. 13 say 'As for Fig 12' while Fig. 12's caption says 'Performance of the BPR method applied to multimodal likelihoods...'; the wording should be made consistent and the differences (per-mode vs averaged quantities) clearly stated.","section":"Section 4.4, Figs. 12 and 13"},{"comment":"The sentence 'one has neither gained nor lost anything by introducing the hyperparameter β' is true only if NS samples the exact joint posterior; this proviso should be restated directly before Eq. (14) to avoid confusion with the effective-posterior discussion that follows.","section":"Section 3, Eq. (14)"}],"recommendation":"major_revision","confidential_remarks":"The reader's report and the skeptic's concern are both on point: the paper's own asymmetric multimodal example contradicts the factorized effective posterior on which the practical claims rest. The authors should be asked to either provide a corrected theory for the non-factorized regime or remove the strong default-use/robustness claims for asymmetric multimodal problems. Also, the evidence correction needs at least a heuristic derivation with explicit assumptions and a test on a case where Eq. (15) is violated. The paper is likely of interest to the Bayesian computation community, but the current claims outrun the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about this paper. First, the core idea is as simple as advertised: run nested sampling on the joint posterior of (θ, β) with the power-posterior repartitioning, and you get samples from P(θ) as the marginal, by construction. The diagnostic use of the β-marginal to detect unrepresentative priors is genuinely new and likely useful. Second, the paper's practical justification for the failure regime rests on a factorization assumption that its own Fig. 13 contradicts, so the method is a promising heuristic, not a theoretically guaranteed default.\n\nWhat is actually new: treating β as a hyperparameter within a single NS run, rather than the annealing schedule of the earlier PR method. The numerical study is broad—1D to 10D, Gaussian, Laplace, and multimodal targets—and shows accurate parameter recovery and evidence estimates in most cases, with small overhead for representative priors. The writing is clear, and the authors are admirably honest that the construction is conceptually almost trivial.\n\nThe soft spots are real. Equation (15) assumes the effective joint posterior factorizes as P(θ) P~(β). If that held, the conditional distribution of β given θ would be independent of θ. But in the asymmetric four-mode example, Fig. 13 shows the mean log β differs by orders of magnitude between modes that are close to versus far from the prior center. That is direct evidence that the sampled distribution is not of the product form. The paper presents this as a feature, but it undermines the central argument that marginalizing β recovers P(θ) in the unrepresentative-prior regime. The evidence correction in Section 3 is likewise heuristic—scaling the β histogram so the largest bin integrates to the prior mass of that bin—and is not derived. It seems to work in the examples where the factorization appears to hold, but the multimodal cases report no evidence estimates at all. The blanket recommendation that BPR should be the default in all NS analyses is extrapolated from MultiNest-only experiments and a limited set of targets.\n\nThese are load-bearing issues for the theory, but not for the practical value of the diagnostic and the simple implementation. The paper deserves a serious referee: the idea is important for a widely used algorithm, the experiments are extensive, and the flaws are identifiable and potentially fixable. I would recommend peer review with the explicit request that the authors either prove or properly bound the factorization assumption, and validate the evidence correction in cases where it fails.\n\nBest,\n[Your name]","headline":"A conceptually simple and potentially useful extension of posterior repartitioning, but the paper's central factorization assumption is contradicted by its own asymmetric multimodal example, so treat the default-use claim with caution.","tokens_in":20176,"tokens_out":4471,"would_cite":false,"duration_ms":45600,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","65C05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Treating the prior-repartitioning power β as a hyperparameter makes nested sampling self-adapting and correct even when the prior is unrepresentative.","keywords":["Bayesian inference","nested sampling","posterior repartitioning","unrepresentative prior","hyperparameter beta","evidence estimation","marginal likelihood","multimodal posterior"],"falsifier":"Run BPR on a constructed problem where the effective β marginal is not flat but strongly bimodal or concentrated near β+, and compare the histogram-corrected log-evidence with exact quadrature. If the correction uses only the largest bin, a visibly non-flat P̃(β) should produce a bias that grows with the departure from flatness; observing that bias would falsify the claim that the correction always recovers the evidence. Alternatively, a case where the sampler fails even at low β, so that no bin of P̃(β) reliably equals π(β), would violate the assumption.","tokens_in":19206,"feed_emoji":"🎲","tokens_out":5785,"duration_ms":53190,"temperature":0.7,"pith_summary":"This paper claims that the auxiliary exponent β in posterior repartitioning can be treated as a hyperparameter and inferred from the data in a single nested-sampling run, rather than tuned through an annealing schedule. The joint posterior then factorizes as P(θ)π(β), so marginalizing over β returns the original posterior and the evidence is unchanged in the ideal limit. When the prior is unrepresentative and nested sampling would fail, the method still yields correct parameter estimates, and a histogram-based correction recovers the evidence. Numerical examples from one to ten dimensions show accuracy and stable computational cost, suggesting this Bayesian posterior repartitioning should be the default mode for nested-sampling analyses.","feed_headline":"Make nested sampling self-tuning: infer the prior power beta","feed_subtitle":"Drawing beta from its own prior fixes unrepresentative-prior failures with negligible overhead.","key_machinery":"The load-bearing identity is the prior–likelihood repartitioning L(θ)π(θ) = L̃(θ,β)π̃(θ|β)π(β) with π̃(θ|β) ∝ π(θ)^β and β ∈ [0,1], which keeps the evidence fixed while broadening the effective prior. The method's second element is the effective posterior factorization: when the sampler fails for extreme β, the returned joint distribution is P(θ)P̃(β) rather than P(θ)π(β), and the marginal P̃(β) is a top-hat-like distribution over [β−,β+]. The third element is the histogram-scaling correction that converts the observed β marginal into an estimate of the missing mass, making the evidence estimate unbiased. Together these pieces let one extra hyperparameter absorb the failure mode of nested sampling with unrepresentative priors.","core_discovery":"By construction, replacing the prior with π̃(θ|β) ∝ π(θ)^β and the likelihood with L̃(θ,β) = L(θ)π(θ)^(1−β)Z_π(β) keeps the product L(θ)π(θ) unchanged. Treating β as uniform on [0,1] and sampling the joint distribution therefore yields samples from P(θ)π(β); the marginal on θ is the original posterior exactly. In realistic finite-live-point runs with an unrepresentative prior, the sampler returns an effective posterior P(θ)P̃(β), where P̃(β) is nonzero only over [β−,β+]; marginalizing over θ still gives P(θ), so parameter inference is correct, and β+ diagnoses how far into the prior wings the likelihood lies. The paper's evidence correction scales the histogram of β samples so that its largest bin has the same integrated mass as the prior π(β) over that bin, then sums the bin volumes to estimate the missing-mass factor ∫P̃(β)dβ; this restores unbiased evidence estimates in all tested examples.","pith_inferences":["β+ could be repurposed as a cheap outlier score in survey-scale pipelines: compute it per dataset and re-analyse only those with small values, rather than re-running every dataset.","If β were given an informative prior instead of uniform, BPR could encode a preference for how aggressively to repartition, turning the method into a tunable shrinkage toward standard NS.","The same joint-reparameterisation idea may transfer to any sampler that uses the prior and likelihood separately, not only to the specific nested-sampling implementation tested here.","A natural further test is to verify that the histogram-scaling evidence correction remains unbiased when the effective β marginal has several modes, since the paper demonstrates multi-modal θ posteriors but only shows unimodal β marginals."],"forward_implications":["A single nested-sampling run with β added as a hyperparameter replaces the multiple runs and annealing schedule of the original PR method, cutting computational cost by orders of magnitude.","Analyses on many datasets can use one standard prior without pre-screening: representative datasets pay negligible overhead, while unrepresentative ones are automatically handled and flagged.","Multi-modal likelihoods whose modes lie at different distances from the prior centre can be sampled correctly in one run, because different β ranges are inferred for different regions.","The upper edge β+ of the β marginal becomes a quantitative, self-calibrating measure of how unrepresentative a prior is for a given dataset.","Evidence estimates remain accurate in the failure regime after the histogram correction, so model comparison is not corrupted by unrepresentative priors."],"supporting_citations":[{"why":"Introduces posterior repartitioning and its annealing-schedule tuning, the method BPR automates and compares against.","marker":"Chen et al. (2018)"},{"why":"Defines nested sampling and the evidence integral that BPR targets.","marker":"Skilling (2006)"},{"why":"Supplies the nested-sampling implementation used for all numerical demonstrations.","marker":"Feroz et al. (2009)"},{"why":"Provides an alternative nested-sampling implementation that BPR is expected to generalize to.","marker":"Handley et al. (2015)"},{"why":"Shows the effective prior in repartitioning can be arbitrary, underpinning the flexibility of the power-prior choice.","marker":"Alsing and Handley (2021)"}],"fun_headline_variants":["Infer beta as a hyperparameter for nested sampling","Bayesian posterior repartitioning: self-adapting beta","Hands-off nested sampling with Bayesian beta inference","Replace annealing with Bayesian beta marginalization","Self-tuning nested sampling by learning the prior power"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evidence correction assumes that, in the bin of the β histogram containing the most samples, the sampler is unbiased and the effective marginal P̃(β) equals the prior π(β), so scaling that bin to the prior's integrated mass correctly estimates the missing mass elsewhere; this is assumed rather than derived.","fun_headline_variants_meta":{"raw":{"variants":["Infer beta as a hyperparameter for nested sampling","Bayesian posterior repartitioning: self-adapting beta","Hands-off nested sampling with Bayesian beta inference","Replace annealing with Bayesian beta marginalization","Self-tuning nested sampling by learning the prior power"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000361,"raw_usage":{"total_tokens":2012,"prompt_tokens":1072,"completion_tokens":940,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":688,"completion_tokens_details":{"reasoning_tokens":866}},"tokens_in":688,"tokens_out":940,"duration_ms":10176,"temperature":1.0,"reasoning_tokens":866,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:35:53.402533+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run BPR on a constructed problem where the effective β marginal is not flat but strongly bimodal or concentrated near β+, and compare the histogram-corrected log-evidence with exact quadrature. If the correction uses only the largest bin, a visibly non-flat P̃(β) should produce a bias that grows with the departure from flatness; observing that bias would falsify the claim that the correction always recovers the evidence. Alternatively, a case where the sampler fails even at low β, so that no bin of P̃(β) reliably equals π(β), would violate the assumption.","supporting_citations":[{"cited_title":"Statistics and Computing pp 1--16","cited_arxiv_id":null,"evidence_quote":"Introduces posterior repartitioning and its annealing-schedule tuning, the method BPR automates and compares against."},{"cited_title":"Bayesian Analysis 1(4):833--860","cited_arxiv_id":null,"evidence_quote":"Defines nested sampling and the evidence integral that BPR targets."},{"cited_title":"Monthly Notices of the Royal Astronomical Society 398(4):1601--1614","cited_arxiv_id":null,"evidence_quote":"Supplies the nested-sampling implementation used for all numerical demonstrations."},{"cited_title":"Monthly Notices of the Royal Astronomical Society 453(4):4384--4398","cited_arxiv_id":null,"evidence_quote":"Provides an alternative nested-sampling implementation that BPR is expected to generalize to."}],"review_version":1}