{"id":"bd4e8f0b-5c73-4afb-a144-812f0e93c750","arxiv_id":"2501.16432","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A RealNVP normalizing flow trained inside nested sampling accelerates Bayesian scans of the Type-II seesaw parameter space and yields posterior constraints on scalar masses and couplings.","lead":"This paper combines a machine-learning classifier and a normalizing-flow generator with nested sampling to map the allowed parameter space of the Type-II seesaw model. The authors report faster convergence than plain nested sampling and give posterior ranges for the model's scalar-sector parameters.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The RNVP-SF correctness claim rests on an unvalidated unbiased-proposal assumption: after the switch to RNVP-F, generated live-point candidates are accepted without density correction, and the truncated RNVP-S baseline cannot certify that the constrained prior is sampled uniformly.","rationale":"The reader's weakest assumption correctly identifies the load-bearing premise: after the switch from RNVP-S to RNVP-F, proposals must remain iid draws from the NS constrained prior within the shrinking hypercube, so that the NS importance weights stay unbiased. My reading of Secs. 3.2.3 and 3.3.3 supports this concern and makes it more concrete: the generator's density is not used for importance correction, and RNVP-F's training set is deliberately drawn from an iso-likelihood shell rather than from the full constrained prior. Filtering such proposals by L > L* does not produce uniform samples from the constrained prior, so the NS bookkeeping is biased unless the generator happens to match the target distribution perfectly, which is not demonstrated. The paper has real strengths: the external toolchain is standard, the SNN classifier accuracy is high and validated, the authors transparently show that RNVP-F alone gives a wrong posterior, and the phenomenological conclusions are plausible. However, the hybrid method's central correctness claim is not yet established. The truncated RNVP-S baseline is insufficient because its tolerance of ~4.09 leaves a meaningful fraction of the evidence uncomputed, and the paper's own note about a possible undiscovered positive-λ4 high-likelihood region is direct evidence of incomplete exploration. These gaps are addressable with an independent sampler comparison, which is why I agree with the CONDITIONAL verdict and do not recommend changing it.","tokens_in":20280,"tokens_out":6140,"duration_ms":63979,"concrete_test":"Run an independent nested-sampling analysis (e.g., MultiNest or PolyChord) on the same Type-II seesaw likelihood and priors, converged to tolerance 0.001, and compare its log-evidence and posterior contours against the RNVP-SF result, with particular attention to the positive-λ4 region. If the independent run finds substantial posterior mass at positive λ4, or if its log-evidence differs from the RNVP-SF result by more than the combined sampling error, the central correctness claim fails. Agreement within tolerance would validate the unbiasedness of the hybrid proposal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that RNVP-SF 'optimizes the runtime while generating the correct posterior' (Sec. 5). For NS to be unbiased, each new live point must be a draw from the uniform constrained prior in the current hypercube. The paper does not establish this for the hybrid run. Sec. 3.3.3 states that RNVP-F selects its training dataset from 'iso-likelihood contours with L slightly lower than the lowest of the live points' rather than from the full constrained prior. The RealNVP is used only as a generator and its learned density is not used to reweight accepted proposals (Sec. 3.2.3). Thus, after the switch to RNVP-F, points that pass L > L* are distributed according to the truncated generator density, not uniformly over the constrained prior. The validation offered is a comparison with an RNVP-S run truncated at tolerance ~4.09 after >700 h (Sec. 3.3.3, Fig. 4) plus a visual statement that the posterior did not change 'perceptibly'. A tolerance of 4.09 versus the target 0.001 means a substantial fraction of the evidence integral remains unaccounted for, so matching its contours cannot certify correctness. Moreover, the authors themselves admit in Sec. 4.1 that there may be 'another high likelihood region with positive λ4, which was not detected by our algorithm'. This is exactly the failure mode expected from a biased proposal that concentrates on the currently learned iso-likelihood shell: an unexplored mode of the constrained prior is never proposed. The unbiasedness premise is therefore load-bearing and unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an ML-assisted nested sampling (NS) framework for BSM parameter spaces, combining an ensemble self-normalizing neural-network classifier with an iteratively retrained RealNVP normalizing flow used as a proposal generator inside NS. Three dataset-selection strategies are compared: RNVP-S (representative prior samples from the bounding hypercube), RNVP-F (training data drawn from iso-likelihood contours just below the current live threshold), and the hybrid RNVP-SF, which switches from RNVP-S to RNVP-F once the sample rate falls below 50. The method is applied to the Type-II seesaw model with seven free parameters, using SPheno and HiggsTools to evaluate HiggsSignals, rho-parameter, and oblique-parameter likelihoods. The authors present posterior constraints on the scalar sector, predictions for charged and neutral scalar masses, doubly charged Higgs collider limits, and estimates of the neutrino Yukawa matrix. The central methodological claim is that RNVP-SF \"optimizes the runtime while generating the correct posterior\" (Sec. 5).","tokens_in":20548,"tokens_out":4501,"duration_ms":43458,"significance":"If the central correctness claim were established, the paper would be a useful practical contribution: it demonstrates a nontrivial integration of a classifier and a normalizing flow into nested sampling, with a physically interesting application and reproducible artifacts (data, figures, and trained models are provided on GitHub). The physics results are plausible and the pipeline is standard in its use of existing spectrum generators and HiggsTools. However, the methodological validation is currently insufficient to support the headline claim, and the authors themselves identify a possibly undetected high-likelihood region with positive lambda4. Because the manuscript is transparent about these gaps and the approach is promising, the correct recommendation is major revision rather than rejection.","major_comments":[{"comment":"The claim that RNVP-SF generates the correct posterior is not supported by the provided validation. In RNVP-F mode, the training set is selected from iso-likelihood contours with L slightly below the lowest live-point likelihood (Sec. 3.3.3), and the RealNVP is used only as a generator with no reweighting by its learned density (Sec. 3.2.3). After the switch, accepted candidates with L > L* are therefore distributed according to the truncated generator density, not uniformly over the constrained prior within the current hypercube, so the NS importance weights are biased. The comparison baseline, RNVP-S, was stopped at tolerance ~4.09 after more than 700 hours, far from the target 0.001, and the agreement is judged visually as not \"perceptibly\" different. A non-converged baseline cannot certify unbiasedness. The authors should provide either a fully converged independent NS run, a density-corrected acceptance scheme for the generator proposals, or diagnostics demonstrating that the live points after the switch are effectively iid draws from the constrained prior.","section":"Sec. 3.3.3, Fig. 4"},{"comment":"The authors state that \"there exists another high likelihood region with positive λ4, which was not detected by our algorithm.\" This admission directly contradicts the claim of a correct posterior: if a high-likelihood mode of the constrained prior is missing, the weighted NS sample under-represents that region and the evidence is underestimated. This is exactly the failure mode expected when the proposal is trained on iso-likelihood shells around the currently known live points. The manuscript should resolve this ambiguity, for example by performing a targeted scan or separate NS run initialized in the positive-λ4 region, or by quantifying that the posterior mass of that region is negligible within the stated credible intervals.","section":"Sec. 4.1, Fig. 5 (bottom-right)"},{"comment":"The initial live points for the NS run are taken from a pre-collected pool of vectors with total χ² ≤ 1335, not from the declared uniform prior over the seven-dimensional parameter box. Nested-sampling theory requires the initial live points to be drawn from the prior; using a low-χ²-truncated pool changes the effective prior and therefore biases the evidence and early posterior weights. The manuscript should quantify the fraction of prior volume excluded by the χ² ≤ 1335 cutoff and demonstrate that this truncation does not affect the final posterior, for example by comparing with an initial pool drawn without that cutoff or by reweighting the initial live points.","section":"Sec. 3.3.1"}],"minor_comments":[{"comment":"The PACS line \"PACS-key discribring text of that key\" is a placeholder and should be removed or filled with actual PACS codes.","section":"Abstract"},{"comment":"The phrase \"large chuck of the parameter space\" should read \"large chunk of the parameter space.\"","section":"Sec. 1"},{"comment":"The legend entries such as \"RN V P− SF\" have inconsistent spacing and should be typeset as RNVP-SF; the vertical line marking the switching point should be defined in the caption.","section":"Fig. 4"},{"comment":"The Yukawa matrix is quoted with means and standard deviations, but the manuscript does not specify the PMNS parametrization, the basis in which Y is given, or the inversion formula used beyond Eq. (12); as written, the result is not reproducible.","section":"Sec. 4.2.2, Eq. (13)"},{"comment":"Equation (10) is typeset ambiguously; the numerator and denominator in the prefactor should be clarified, for example as m_{h±}^2 = ((2√2 μ1 − λ4 v_T) / (4 v_T)) (v_d^2 + 2 v_T^2).","section":"Sec. 2, Eq. (10)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents a promising method but the central correctness claim is not yet established. The authors' candid admission of a possibly missed λ4-positive mode and the unconverged RNVP-S baseline make the validation gap clear. I do not recommend rejection, because the core idea is sound and the gaps appear fixable with additional runs or diagnostics, but the current version should not be accepted as-is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing. The paper is a real engineering contribution: an ensemble SNN classifier plus iteratively retrained RealNVP generator inside nested sampling, with a dataset-selection switch (RNVP-SF) that appears to avoid the mode-collapse failure of the fast RNVP-F variant. The Type-II seesaw posterior itself is new, and the pheno results look plausible. But the central claim—that RNVP-SF “optimizes the runtime while generating the correct posterior”—is not established by the evidence in the manuscript.\n\nWhat it does well: the training loop is iterative and grounded in truth likelihoods from SPheno/HiggsTools; the classifier accuracy is high; the authors explicitly show that RNVP-F alone gives a wrong posterior, which is a useful negative result; and they publish data, figures, and trained models on GitHub. The self-critical note in Sec. 4.1 about a possible undetected positive-λ4 high-likelihood region is creditable and is exactly the right thing to probe.\n\nThe soft spot is load-bearing. After the switch to RNVP-F, the generator is trained on iso-likelihood contours just below L*, not on the full constrained prior, and its learned density is not used to reweight accepted proposals. For NS to stay unbiased, new live points must be draws from the constrained prior in the shrinking hypercube. The paper’s only check is a comparison with an RNVP-S run truncated at tolerance ~4.09 after >700 h, plus a statement that the posterior did not change “perceptibly.” Tolerance 4.09 is far from the target 0.001, so matching its contours is weak evidence. The admitted missed λ4 region is precisely what a biased proposal concentrated on the learned iso-likelihood shell would produce. This is not a refutation of the method, but it is a correctness gap.\n\nMinor concerns: the initial NS pool is pre-filtered by χ2 ≤ 1335, and the manuscript does not state whether runnable code with a commit hash is in the repo, so I could not audit reproducibility. Those are addressable.\n\nWho is it for: people working on ML-accelerated Bayesian scans of BSM parameter spaces, and Type-II seesaw phenomenologists. It deserves a serious referee, not a desk reject. The referee should ask for a converged independent-sampler comparison (or a fully converged RNVP-S run on a reduced problem), a check of the λ4>0 region, and an explicit statement of the proposal-distribution bias and how it is corrected.","headline":"A useful ML-accelerated nested sampling engineering contribution with a new Type-II seesaw posterior, but the headline claim of an unbiased hybrid proposal is not yet backed by a converged baseline or an independent sampler check.","tokens_in":21195,"tokens_out":2178,"would_cite":true,"duration_ms":20152,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Nested sampling of a seven-dimensional BSM parameter space can be accelerated by an iteratively retrained RealNVP generator plus a classifier without biasing the resulting posterior, and the paper demonstrates this on the Type-II seesaw…","keywords":["nested sampling","normalizing flows","RealNVP","self-normalizing neural network","Type-II seesaw","Bayesian posterior","Higgs observables","oblique parameters"],"falsifier":"Run plain nested sampling on the same seven-parameter space to tolerance below 0.001 and compare the marginal posterior in lambda4; if it reveals a positive-lambda4 high-likelihood mode that RNVP-SF missed, the claim of a correct posterior fails. A cheaper diagnostic is a two-sample test between RNVP-F proposals and the exact constrained prior restricted to the live hypercube after the switch, since any systematic mismatch breaks the iid-proposal assumption.","tokens_in":19949,"feed_emoji":"⚛️","tokens_out":8586,"duration_ms":78365,"temperature":0.7,"pith_summary":"This paper is trying to establish that nested sampling, the standard Bayesian evidence and posterior sampler, can be made much faster on particle-physics parameter spaces by inserting machine-learned proposals without corrupting the result. The proposed pipeline inserts two networks into the nested-sampling loop: an ensemble classifier that avoids running the expensive spectrum generator on invalid points, and a RealNVP normalizing-flow generator that is retrained as the run progresses to propose points inside the shrinking live region. The specific schedule named RNVP-SF starts with the faithful hypercube-based selection and switches to fast iso-likelihood selection once the sample rate falls; the authors report that this reaches the termination tolerance with a posterior that is not perceptibly different from the much slower faithful run. Applied to the Type-II seesaw model, the method yields what the authors describe as the first posterior estimates for all seven new parameters under 125 GeV Higgs, rho-parameter, and oblique data, and maps the favoured region for the triplet VEV and charged Higgs masses. A reader should care because the cost bottleneck is generic: any beyond-standard-model scan that must generate a spectrum at every candidate point faces the same slowdown.","feed_headline":"Flow-boosted sampling maps Type-II seesaw in hours","feed_subtitle":"A flow-trained generator keeps nested sampling unbiased while slashing runtime on a seven-parameter particle scan.","key_machinery":"The load-bearing object is the RNVP-SF proposal pipeline. RealNVP is a normalizing flow built from affine coupling layers, an invertible map whose triangular Jacobian makes density evaluation cheap, and here it is trained iteratively on pooled 13-component truth vectors, namely the seven input parameters plus the SM-like Higgs mass, rho shift, three oblique parameters, and the Higgs chi-squared. An ensemble of self-normalizing networks classifies parameter points as spectrum-valid or not before the expensive generator is called. The third piece is the dataset-selection schedule: sample training points from the hypercube bounding the live points, RNVP-S, until the nested-sampling sample rate between two trainings drops below 50, then switch to selecting points near the current iso-likelihood contour, RNVP-F. This schedule is what the paper credits for preserving unbiasedness while recovering speed.","core_discovery":"The paper's central claim is methodological, stated in Section 5 as: the variant called RNVP-SF \"optimizes the runtime while generating the correct posterior.\" Correctness is established by the posterior's high-likelihood region coinciding with the concentration of live points and by comparing against a truncated slow run at tolerance ~4.09; the authors also present the posterior as the first Bayesian mapping of all seven Type-II seesaw parameters. The same section presents the physics outcome of applying the method: the high-vT Type-II seesaw parameter space, constrained by the 125 GeV Higgs data, the rho-parameter, and the oblique parameters, with the favoured scalar masses, mixing angle, and neutrino Yukawa couplings reported in the results.","pith_inferences":["The switch heuristic, namely switching to iso-likelihood dataset selection once the sample rate between two trainings falls below 50, is tuned on this run; a model with widely separated likelihood modes might need a different schedule, and a blended or multi-flow proposal could be more robust.","The authors' own admission that a high-likelihood region at positive lambda4 may be missing is exactly the signature a biased proposal would produce, so a controlled toy model with known evidence is the natural next test of the method.","The same architecture is portable to other beyond-standard-model scans whose bottleneck is an expensive spectrum generator, but unbiasedness is a property of the whole proposal-plus-schedule system and would have to be revalidated for each model rather than assumed."],"forward_implications":["The RNVP-SF scheme reduces the wall time of the Type-II seesaw scan from more than 700 hours at tolerance ~4.09 to convergence at tolerance below 0.001, without, the authors argue, altering the posterior perceptibly.","The resulting posterior maps the large-vT Type-II seesaw parameter space for the first time: the triplet VEV peaks near 1.2-1.4 GeV, and all seven input parameters now have central values, dispersions, and correlations.","The allowed exotic scalar masses are mostly 400-800 GeV, with the charged-Higgs mass splitting at most about 40 GeV and the CP-even mixing angle sin alpha around 10^-3, so the favoured region is not yet excluded by pair-produced doubly charged Higgs searches.","Neutrino Yukawa matrix elements in the favoured region are all near 10^-10, strongly suppressing lepton-flavour- and lepton-number-violating decays.","Because nested sampling produces Bayesian evidence as a by-product, the method also gives the tools to compare Type-II seesaw against other models as data accumulate."],"supporting_citations":[{"why":"Establishes the classifier-based ML-NS pipeline and the amortized-likelihood idea that this paper combines with a flow generator.","marker":"[24]"},{"why":"Defines nested sampling and its weighted-sample posterior, the algorithm whose slowdown motivates the ML assistance.","marker":"[25, 26]"},{"why":"Supplies the RealNVP normalizing-flow architecture used as the proposal generator.","marker":"[130]"},{"why":"Supplies the vacuum-stability and perturbative-unitarity constraints that define the allowed parameter region and the classifier labels.","marker":"[83]"},{"why":"Computes the 125 GeV Higgs likelihood used as the final vector component and as the nested-sampling likelihood.","marker":"[119]"},{"why":"Provides the oblique-parameter fit data and correlations entering the likelihood.","marker":"[91]"},{"why":"Provide the rho-parameter measurements entering the likelihood.","marker":"[92, 93]"},{"why":"Provides the neutrino oscillation data used to derive the posterior distribution of the Yukawa matrix.","marker":"[42]"},{"why":"Generate the particle spectrum whose observables make up the truth vectors used for training and likelihood evaluation.","marker":"[106, 107]"}],"fun_headline_variants":["Flow-boosted nested sampling maps Type-II seesaw","First full Bayesian map of Type-II seesaw via ML","ML-accelerated nested sampling tames seesaw parameter space","Normalizing flow plus nested sampling solves seesaw scan","Seven-parameter seesaw mapped with flow-boosted sampling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The correctness claim rests on the assumption that, after switching to the fast iso-likelihood dataset-selection phase, the generator's proposals are still independent draws from the constrained prior inside the shrinking hypercube, so the nested-sampling importance weights are unbiased; the paper checks this only by comparing against a truncated slow run and by asserting no perceptible change.","fun_headline_variants_meta":{"raw":{"variants":["Flow-boosted nested sampling maps Type-II seesaw","First full Bayesian map of Type-II seesaw via ML","ML-accelerated nested sampling tames seesaw parameter space","Normalizing flow plus nested sampling solves seesaw scan","Seven-parameter seesaw mapped with flow-boosted sampling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1488,"prompt_tokens":814,"completion_tokens":674,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":592}},"tokens_in":430,"tokens_out":674,"duration_ms":6895,"temperature":1.0,"reasoning_tokens":592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:16:49.505592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run plain nested sampling on the same seven-parameter space to tolerance below 0.001 and compare the marginal posterior in lambda4; if it reveals a positive-lambda4 high-likelihood mode that RNVP-SF missed, the claim of a correct posterior fails. A cheaper diagnostic is a two-sample test between RNVP-F proposals and the exact constrained prior restricted to the live hypercube after the switch, since any systematic mismatch breaks the iid-proposal assumption.","supporting_citations":[],"review_version":1}