REVIEW 5 major objections 4 minor 1 cited by
The paper claims that a factorizable normalizing flow can learn, in a single amortized training pass, the full mapping from nuisance parameters to the optimal distribution-transformation fit target, making fully profiled unbinned measuremen
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 23:32 UTC pith:5YSDVYM2
load-bearing objection A genuinely useful amortized profiling idea that currently sells itself short: the abstract promises a Poisson-bootstrap budget and robust coverage that the body never delivers. the 5 major comments →
Profiling systematic uncertainties in Simulation-Based Inference with Factorizable Normalizing Flows
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Central claim: a fully profiled unbinned likelihood fit can be turned into a functional measurement. Factorizable Normalizing Flows (FNF) write the scale and shift of an affine flow as sums of independent linear and quadratic terms per nuisance parameter, with the coefficients learned by neural networks; this keeps the likelihood tractable and lets each systematic be trained from ±1σ samples. A two-step amortized procedure — global best-fit transformation first, then freeze it and train the deformation network while sampling nuisance points — learns a global map from nuisance configuration to the maximum-likelihood DoI transformation in one run, making the profiled transformation available i
What carries the argument
The load-bearing object is the Factorizable Normalizing Flow layer: for each feature dimension j, the scale s_j and shift t_j are written as sums over nuisances of independent linear and quadratic terms, with coefficients output by masked MLPs conditioned on the features. This quadratic-plus-additive ansatz turns systematic variations into smooth, composable deformations that scale linearly with the number of nuisances rather than exponentially. The second key mechanism is the amortized two-step procedure: Step 1 finds the global best-fit DoI transformation; Step 2 freezes it and trains the FNF deformation network by sampling nuisance points from their confidence region, so that the network
Load-bearing premise
The load-bearing premise is that every systematic effect can be written as independent, additive, quadratic functions of the nuisance parameters inside the flow's scale-and-shift space; if real systematics multiply or interact, the profiled uncertainties derived from this transformation class will be biased.
What would settle it
Construct a synthetic experiment with two nuisance parameters whose effects multiply (e.g., the second scales the first's coefficient), train the FNF, and check the profiled 68% interval for the DoI transformation against the true distortion map. If coverage falls well below nominal on repeated trials, or if the amortized profile differs materially from brute-force refits at the same nuisance points, the central claim fails.
If this is right
- Unbinned likelihood fits with many systematic sources become computationally practical: profiling cost moves from per-scan-point retraining to one amortized training pass.
- The fit target can be a full differential distribution via a learnable invertible transformation, enabling unbinned cross-section measurements and simulation calibration without binning.
- Once trained, the profiled transformation is available instantaneously for any nuisance configuration, so systematic uncertainties on the measured distribution can be read off directly from variations along the likelihood contour.
- Adding a new systematic source requires only adding one more factorized network, without retraining the whole model, because of the modular additive structure.
- Each nuisance variation can be trained from its ±1σ samples independently, avoiding exponential growth in simulation data required by a full joint nuisance-space scan.
Where Pith is reading between the lines
- Not demonstrated in the paper: a stress test where two nuisances interact multiplicatively, such as one scaling the other's coefficient, to check whether the additivity assumption biases the profiled DoI. If it does, the framework would need interaction terms or nested factorized layers.
- The same amortized mechanism could be applied to forward modeling, mapping particle-level predictions to detector-level data, so that detector calibration and unfolding are handled by the same learned transformation family; the paper gestures at this but does not implement it.
- The statistical claim could be sharpened by comparing profiled intervals from the amortized network against brute-force per-point refits on low-dimensional problems, quantifying exactly how much coverage error the amortization step introduces.
- Because the learned DoI is invertible, fitted results could be reused as a surrogate for full detector simulation in reinterpretation studies; this follows from the paper's discussion but is not demonstrated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a simulation-based inference framework for unbinned likelihood fits in which the target of the fit is a Distribution of Interest (DoI): an invertible neural-network transformation mapping a reference density to the data density. Systematic uncertainties are modeled through Factorizable Normalizing Flows (FNF): per-nuisance transformations whose affine scale and shift parameters are additive and quadratic in each nuisance parameter (Eqs. 15-16). The central methodological claim is that a two-step training procedure — first a joint global best fit over the DoI and nuisances, then an amortized optimization of a residual systematic-aware transformation over sampled nuisance configurations — learns the profile-likelihood surface in a single run, so that the systematic uncertainty on the DoI can be read off by evaluating the learned map along likelihood contours. The method is validated on a two-dimensional synthetic dataset with two independent nuisance parameters, showing visually good agreement between fitted and generated distributions and giving likelihood scans and principal-component visualizations of the uncertainty.
Significance. If the central claims hold, the framework would be a substantial step toward practical unbinned profile-likelihood fits with many systematic uncertainties and distribution-valued targets. The change-of-variables likelihood machinery (Eqs. 8, 13-17) is standard, and the FNF factorization is an interesting, interpretable ansatz for systematic deformations. The paper also provides a useful synthetic benchmark and detailed architectural details in Appendix B. However, the statistical claims currently outrun the evidence: the advertised Poisson-bootstrap averaged likelihood is absent from the body, the extracted DoI uncertainty is not demonstrated to be a valid confidence set, and the key factorization and amortized-convergence assumptions are not stress-tested. These are load-bearing gaps rather than presentation issues.
major comments (5)
- [Abstract; Sections 2-4] The abstract promises a Poisson-bootstrap ensemble marginalized through an averaged likelihood to deliver a complete statistical-plus-systematic uncertainty budget within a single unbinned likelihood. No section or equation in the body implements this: the only Poisson distribution is the extended-likelihood term in Eq. (6), and there is no bootstrap ensemble, no averaged likelihood, and no treatment of the finite-sample statistical variance of the neural-network DoI. The advertised 'complete' budget is therefore not realized. This must either be implemented and validated or removed/qualified from the claims.
- [Section 3.3.1, Eqs. (24)-(25); Section 4.4.2, Figs. 9-10] The paper extracts the DoI uncertainty by evaluating T_psi along the nuisance-space likelihood contour, but this is not established to be a profile-likelihood interval for a functional parameter. The scans in Fig. 9 fix the transformation T_hat_nu_phi while scanning nu, and Fig. 10 evaluates the learned T_psi without showing that T_psi is the conditional argmax of the likelihood for each nu. Neither is equivalent to a profile likelihood ratio unless the amortized network is proven to track the per-nu optimum. No coverage test is reported, so the claim of 'robust frequentist coverage' is unsupported. I ask for a comparison against brute-force per-nu refits on a grid and a coverage calibration of the resulting DoI bands.
- [Section 3.1, Eqs. (15)-(16)] The additive, quadratic-in-nu factorization in the affine scale/shift space is the load-bearing modeling assumption of the paper. It is asserted rather than derived, and the synthetic benchmark contains only two independent nuisances with simple scaling/shift forms, which cannot validate the factorization for interacting systematics or non-quadratic responses. Since the profiled uncertainties are read off from this model class, a bias in the factorization would directly bias the reported uncertainties. Please add a stress test with correlated/interacting nuisances or provide a quantitative approximation-error study justifying the factorization as sufficient for the target inference.
- [Sections 3.1.2 and 4.4.2] The FNF transformation is trained on the ±1-sigma variation samples, and the systematic uncertainty is then read off from the same fitted model. This is interpolation of training points, not an independent prediction, and accuracy at intermediate or extrapolated nu values is not demonstrated. Additionally, Step 2 samples nuisance configurations only from the 95% confidence interval of the global fit, while Section 3.3.1 claims the model yields the maximum-likelihood transformation for 'any' configuration of nuisances. Please report validation at intermediate and out-of-range nu values, with quantitative error metrics on the learned transformation and density.
- [Section 4.4.3] The paper's own text states that the Gaussian/Hessian approximation used for the principal decomposition is 'not a very accurate description' of the likelihood, yet this approximation is used to define the principal systematic modes and to visualize the uncertainty on the DoI. No quantitative assessment is given of how this inaccuracy propagates into the extracted DoI uncertainty. Since the uncertainty quantification is a central advertised result, this limitation needs to be quantified or the method needs to be replaced by a more accurate likelihood-sensitive decomposition.
minor comments (4)
- [Abstract] The abstract in the top matter promises a Poisson-bootstrap ensemble and averaged likelihood, but the abstract and body of the submitted paper do not mention them. Please harmonize the abstract and body so that advertised contributions match implemented content.
- [Eqs. (18)-(20)] The reparameterization uses Delta_nu/sigma_nu, but the text says training uses '±1-sigma samples.' Please clarify whether sigma_nu is the prior width, the template-variation scale, or the standard deviation of the generated variation samples.
- [Throughout] There are several typos and inconsistent cross-references: 'likelihoqod' in the Introduction, 'procudure' in Section 3.3, 'Sec.15' should be 'Eq. (15)', and 'Figure 4.1' should be 'Figure 5'. A careful proofread is needed.
- [Appendix B] The implementation details are useful, but no code or data release is mentioned. For reproducibility of the synthetic benchmark, please provide access to the code and seed information.
Circularity Check
No significant circularity: the FNF factorization and amortized profiling are fitted surrogates, not reductions to their inputs; missing coverage/bootstrap evidence is a correctness issue, not circularity.
full rationale
The derivation chain does not exhibit a step in which a claimed prediction is equivalent by construction to a fitted input or to a self-citation. Sec. 3.1 (Eqs. 15-16) imposes a factorized quadratic-in-nu ansatz in affine scale/shift space; this is an explicit modeling assumption, not a hidden adoption of the result it later claims. Sec. 3.1.2 says training on +/-1 sigma variation samples 'effectively learns a smooth interpolation'; the later profile evaluation in Sec. 3.3.1 reads uncertainties from the same fitted transformation T_psi along the likelihood contours. That is an amortized surrogate/interpolation rather than an independent prediction, but the paper openly labels it as such and this is standard template-interpolation practice; the uncertainty is not defined as the fit value by an equation. The only self-citation ([30], used in Sec. 3.5 as an example of NF-based morphing) is not load-bearing; the central FNF/profiling argument would survive without it. The stronger claim of 'robust frequentist coverage' (Sec. 3.3.1) is unsupported by any coverage test, and the abstract's Poisson-bootstrap averaged likelihood never appears in Secs. 2-4; the Sec. 4.4.3 Hessian analysis is admitted to be 'not a very accurate description of the likelihood.' These are correctness/completeness deficiencies, not circular reductions. No step reduces Eq. X to Eq. Y by construction or fits a parameter and renames it a prediction.
Axiom & Free-Parameter Ledger
free parameters (5)
- FNF systematic coefficient networks Ψ^(k) (α,β,γ,δ per nuisance per feature) =
not reported (neural-net weights)
- DoI transformation parameters φ (3-layer NSF, 30 bins) =
not reported
- Nominal flow parameters (kinematics flow + conditional y|x flow) =
not reported
- Nuisance sampling range for Step-2 amortized training =
95% CI from the Step-1 likelihood scan; exact bounds not stated
- Prior widths σ_ν used in the reparameterization of Eqs. 18-20 =
not stated explicitly
axioms (6)
- standard math Change-of-variables density formula for invertible transformations with tractable Jacobian determinant
- domain assumption Autoregressive factorization p_f(y,x|ν) = p_f(y|x,ν) p_f(x|ν) is exact, with both sub-flows independently unbiased
- ad hoc to paper Systematic effects decompose as independent additive quadratic functions of each nuisance in the affine flow's latent space (Eqs. 15-16)
- ad hoc to paper Networks trained on ±1σ samples interpolate correctly across the full nuisance range, including near ν=0
- domain assumption Step-2 amortized optimization converges to the conditional argmax for every sampled ν
- domain assumption Mixture fractions η_f(θ,ν) are known or simply parameterized
invented entities (2)
-
Distribution of Interest (DoI)
no independent evidence
-
Factorizable Normalizing Flow (FNF) layer
no independent evidence
read the original abstract
Unbinned likelihood fits maximize the information extracted from experimental data, yet their application in realistic high-dimensional analyses has been fundamentally bottlenecked by the prohibitive computational cost of profiling systematic uncertainties. Furthermore, current machine learning-based inference methods typically estimate scalar parameters, discarding complex high-dimensional correlations. To address this, we propose a general Simulation-Based Inference (SBI) framework that elevates the fit target from scalar parameters to a multivariate Distribution of Interest (DoI), a learnable, invertible transformation of the feature space. We employ Factorizable Normalizing Flows to model systematic variations as parametric deformations, preserving tractability without combinatorial explosion. Crucially, we develop an amortized training strategy that learns the conditional dependence of the DoI on nuisance parameters in a single optimization process, bypassing repetitive training during likelihood scans. To capture the finite-sample statistical variance of the neural network DoI, we introduce a Poisson-bootstrap ensemble, which we marginalize through an averaged likelihood to deliver a complete statistical-plus-systematic uncertainty budget within a single unbinned likelihood. Validated on a synthetic dataset emulating a high-energy physics measurement, our method demonstrates that rigorous, fully profiled unbinned measurements can now be extended to complete differential distributions. By turning the fit into a functional measurement, this approach offers a powerful, unifying framework for a broad range of tasks conventionally treated as distinct problems, from detector calibration and differential cross-sections to unfolding and continuous parameter estimation.
Figures
Forward citations
Cited by 1 Pith paper
-
Proton Structure from Neural Simulation-Based Inference at the LHC
Neural simulation-based inference on unbinned top-quark pair data at 13 TeV yields improved gluon PDF precision over traditional binned analyses while incorporating experimental and theoretical uncertainties.
Reference graph
Works this paper leans on
-
[1]
HistFactory: A tool for creating statistical models for use with RooFit and RooStats
Kyle Cranmer, George Lewis, Lorenzo Moneta, Akira Shibata, and Wouter Verkerke. HistFactory: A tool for creating statistical models for use with RooFit and RooStats. Technical report, New York U., 2012. URL https://cds.cern.ch/record/1456844
arXiv 2012
-
[2]
The CMS statistical analysis and combination tool: COMBINE.Comput
CMS Collaboration. The CMS statistical analysis and combination tool: COMBINE.Comput. Softw. Big Sci., 8 (1):19, 2024. doi:10.1007/s41781-024-00121-4
-
[3]
The frontier of simulation-based inference.Proc
Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based inference.Proc. Nat. Acad. Sci., 117(48):30055–30062, 2020. doi:10.1073/pnas.1912789117
-
[4]
Advancing Tools for Simulation-Based Inference.SciPost Phys
Alexander Held et al. Advancing Tools for Simulation-Based Inference.SciPost Phys. Core, 8:060, 2025. doi:10.21468/SciPostPhysCore.8.2.060
-
[5]
A continuous calibration of the ATLAS flavour-tagging classifiers via optimal transportation maps, 2025
ATLAS Collaboration. A continuous calibration of the ATLAS flavour-tagging classifiers via optimal transportation maps, 2025. Preprint
2025
-
[6]
Mind the Gap: Navigating Inference with Optimal Transport Maps, 2025
Malte Algren, Tobias Golling, Francesco Armando Di Bello, and Christopher Pollard. Mind the Gap: Navigating Inference with Optimal Transport Maps, 2025
2025
-
[7]
An implementation of neural simulation-based inference for parameter estimation in ATLAS.Rept
ATLAS Collaboration. An implementation of neural simulation-based inference for parameter estimation in ATLAS.Rept. Prog. Phys., 2025. doi:10.1088/1361-6633/add370
-
[8]
A Guide to Constraining Effective Field Theories with Machine Learning.Phys
Johann Brehmer, Kyle Cranmer, Gilles Louppe, and Juan Pavez. A Guide to Constraining Effective Field Theories with Machine Learning.Phys. Rev. D, 98(5):052004, 2018. doi:10.1103/PhysRevD.98.052004
-
[9]
Refinable modeling for unbinned SMEFT analyses.Mach
Robert Schöfbeck. Refinable modeling for unbinned SMEFT analyses.Mach. Learn. Sci. Tech., 6:015007, 2025. doi:10.1088/2632-2153/ad9fd1
-
[11]
Ivan Kobyzev, Simon J. D. Prince, and Marcus A. Brubaker. Normalizing Flows: An Introduction and Review of Current Methods.IEEE Trans. Pattern Anal. Mach. Intell., 43(11):3964–3979, 2021. doi:10.1109/TPAMI.2020.2992934
arXiv 2021
-
[12]
Anders Andreassen and Benjamin Nachman. Neural networks for full phase-space reweighting and parameter tuning.Physical Review D, 101(9), May 2020. ISSN 2470-0029. doi:10.1103/physrevd.101.091901. URL http://dx.doi.org/10.1103/PhysRevD.101.091901
-
[13]
Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. Neural spline flows, 2019. URL https://arxiv.org/abs/1906.04032. 26 Profiling systematic uncertainties in SBI with Factorizable Normalizing FlowsPREPRINT
Pith/arXiv arXiv 2019
-
[14]
Masked autoregressive flow for density estimation, 2018
George Papamakarios, Theo Pavlakou, and Iain Murray. Masked autoregressive flow for density estimation, 2018. URLhttps://arxiv.org/abs/1705.07057
Pith/arXiv arXiv 2018
-
[15]
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. URL https://arxiv.org/ abs/1711.05101
Pith/arXiv arXiv 2019
-
[16]
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An Imperative Style, High-Perfo...
2019
-
[17]
Computational Optimal Transport.Found
Gabriel Peyré and Marco Cuturi. Computational Optimal Transport.Found. Trends Mach. Learn., 11(5-6): 355–607, 2019. doi:10.1561/2200000073
-
[18]
Patrick T. Komiske, Eric M. Metodiev, and Jesse Thaler. The Metric Space of Collider Events.Phys. Rev. Lett., 123(4):041801, 2019. doi:10.1103/PhysRevLett.123.041801
-
[19]
OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport
Derek Onken, Samy Wu Fung, Xingjian Li, and Lars Ruthotto. OT-Flow: Fast and Accurate Continuous Normalizing Flows via Optimal Transport. InProceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 9223–9232, 2021. doi:10.1609/aaai.v35i10.17113
-
[20]
Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport.Trans
Alexander Tong, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Kilian Fatras, Guy Wolf, and Yoshua Bengio. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport.Trans. Mach. Learn. Res., 2024, 2024. URL https://openreview.net/forum?id=HgDwiZrpVq. Introduces OT-CFM
2024
-
[21]
Brandon Amos, Lei Xu, and J. Zico Kolter. Input convex neural networks, 2017. URL https://arxiv.org/ abs/1609.07152
Pith/arXiv arXiv 2017
-
[22]
Chin-Wei Huang, Ricky T. Q. Chen, Christos Tsirigotis, and Aaron Courville. Convex potential flows: Universal probability distributions with optimal transport and convex optimization, 2021. URL https://arxiv.org/abs/ 2012.05942
Pith/arXiv arXiv 2021
-
[23]
Springer, Cham, 3rd edition, 2023
Luca Lista.Statistical Methods for Data Analysis: With Applications in Particle Physics, volume 1003 ofLecture Notes in Physics. Springer, Cham, 3rd edition, 2023. ISBN 978-3-031-19933-2. doi:10.1007/978-3-031-19934-9
-
[24]
Generative Unfolding with Distribution Mapping.SciPost Phys., 18: 200, 2025
Anja Butter, Sascha Diefenbacher, Nathan Huetsch, Vinicius Mikuni, Benjamin Nachman, Sofia Pala- cios Schweitzer, and Tilman Plehn. Generative Unfolding with Distribution Mapping.SciPost Phys., 18: 200, 2025. doi:10.21468/SciPostPhys.18.6.200
-
[25]
Data-Driven High-Dimensional Statistical Inference with Generative Models.JHEP, 11:129,
Manuel Szewc et al. Data-Driven High-Dimensional Statistical Inference with Generative Models.JHEP, 11:129,
-
[26]
T2K Collaboration. Machine Learning-Assisted Unfolding for Neutrino Cross-section Measurements with the OmniFold Technique.Phys. Rev. D, 112:012008, 2025. doi:10.1103/PhysRevD.112.012008
-
[27]
Unbinned multivariate observables for global SMEFT analyses from machine learning.JHEP, 03:033, 2023
Raquel Gomez Ambrosio, Jaco ter Hoeve, Maeve Madigan, Juan Rojo, and Veronica Sanz. Unbinned multivariate observables for global SMEFT analyses from machine learning.JHEP, 03:033, 2023. doi:10.1007/JHEP03(2023)033
-
[28]
ATLAS Collaboration. Measurement of off-shell Higgs boson production in the H→ZZ→4ℓ decay channel using a neural simulation-based inference technique in 13 TeV pp collisions with the ATLAS detector.Rept. Prog. Phys., 2025. doi:10.1088/1361-6633/adcd9a
-
[29]
Unbinned inclusive cross-section measurements with machine-learned systematic uncertainties
Lisa Benato, Cristina Giordano, Claudius Krause, Ang Li, Robert Schöfbeck, Dennis Schwarz, Maryam Shooshtari, and Daohan Wang. Unbinned inclusive cross-section measurements with machine-learned systematic uncertainties. Phys. Rev. D, 112:052006, Sep 2025. doi:10.1103/zwzt-1rrw. URL https://link.aps.org/doi/10.1103/ zwzt-1rrw
-
[30]
Caio Daumann, Mauro Donegà, Johannes Erdmann, Massimiliano Galli, Jan Lukas Späh, and Davide Valsecchi. One flow to correct them all: Improving simulations in high-energy physics with a single normalising flow and a switch.Comput. Softw. Big Sci., 8(1):23, 2024. doi:10.1007/s41781-024-00125-0
-
[31]
Analysis-ready generative unfolding, 2025
Anja Butter, Nathan Huetsch, Vinicius Mikuni, Benjamin Nachman, and Sofia Palacios Schweitzer. Analysis-ready generative unfolding, 2025. URLhttps://arxiv.org/abs/2509.02708
Pith/arXiv arXiv 2025
-
[32]
Advancing tools for simulation-based inference.SciPost Physics Core, 8(3), September 2025
Henning Bahl, Víctor Bresó-Pla, Giovanni De Crescenzo, and Tilman Plehn. Advancing tools for simulation-based inference.SciPost Physics Core, 8(3), September 2025. ISSN 2666-9366. doi:10.21468/scipostphyscore.8.3.060. URLhttp://dx.doi.org/10.21468/SciPostPhysCore.8.3.060. 27 Profiling systematic uncertainties in SBI with Factorizable Normalizing FlowsPREPRINT
-
[33]
Simulation-prior independent neural unfolding procedure, 2025
Anja Butter, Theo Heimel, Nathan Huetsch, Michael Kagan, and Tilman Plehn. Simulation-prior independent neural unfolding procedure, 2025. URLhttps://arxiv.org/abs/2507.15084
Pith/arXiv arXiv 2025
-
[34]
Zuko: Normalizing flows in pytorch, 2022
François Rozet et al. Zuko: Normalizing flows in pytorch, 2022. URLhttps://pypi.org/project/zuko. 28
2022
-
[2025]
doi:10.1007/JHEP11(2025)129
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.