REVIEW 4 major objections 5 minor 29 references
Mono-Z Dark Matter Search with Neural Spline Flows Using CMS Run 2015D Open Data
T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Using neural-spline-flow likelihood-ratio scores on public 2015 proton-proton collision data, this analysis sets observed 95% upper limits on the dark matter signal strength for three mediator models, and attributes the 7–12x gap between ob
desk verdict A careful, honest open-data study whose central limits are undercut by the authors' own demonstration that the background model fails in the high-MET tail. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the per-event log-likelihood-ratio score formed from two Neural Spline Flow density estimates: S_h(x) = log p(x|DM_h) − log p(x|SM_channel). A Neural Spline Flow is a normalizing flow whose coupling transforms are monotonic rational-quadratic splines, giving exact tractable densities. Background flows are trained only on control-region events (MET < 50 GeV), one per lepton channel; signal flows are trained on simulated mediator events and evaluated in the same standardized SM feature domain. The score arrays feed a binned profile-likelihood fit over signal and validation regions (MET 50–100 GeV as a sideband) with per-channel normalization nuisances, and asymptotic CL_s formula
What would settle it
Build a background flow that includes events from an independent high-MET sideband (MET between 100 and 200 GeV) and repeat the signal-region fit. If the fitted signal strength drops toward zero and observed limits approach expected limits, the high-MET tail residual was the cause; if the positive signal strength persists with high q0, the residual is intrinsic to the flow/density modeling or to the signal model rather than a CR-to-SR transfer artifact.
Extended reading notes
Core claim
On its own terms, the paper claims that a likelihood-ratio test statistic built from independently trained neural spline flows can serve as the full discriminant for a mono-Z dark matter search, removing the need for a hard upper cut on missing transverse momentum. In the signal region (MET ≥ 50 GeV, |Δφ(MET, Z)| > 2.5, ≤ 1 jet), the per-event score S(x) = log p(x|signal) − log p(x|background), formed from three mediator-specific signal flows and two channel-specific background flows, is fed into a simultaneous signal-region plus validation-region binned profile likelihood. The result is observed 95% CL upper limits on the signal strength μ of <0.0177 (scalar), <0.0362 (vector), and <0.0498
Load-bearing premise
The background score distribution in the signal region—especially the rare events with missing transverse momentum above 100 GeV—is correctly described by a background density trained on MET < 50 GeV events and shifted by one normalization step; the paper's own validation shows simple extrapolations fail in that tail, so if this assumption fails the quoted limits are not valid constraints.
Editorial extensions
If this is right
- The density-ratio score concentrates sensitivity across the full phase space, so the search does not rely on a hard upper MET threshold; events up to the dataset's ~200 GeV MET cap enter the fit.
- A validation-region sideband can supply a data-driven background normalization constraint of roughly 0.85% per channel, even on a small 2.32 fb^-1 dataset.
- The observed-to-expected ratio remains ~7–12 across three mediator hypotheses and three background variants, identifying the high-MET tail as the dominant systematic rather than the signal model.
- The pipeline is transferable to other final states or signal hypotheses by retraining the relevant signal and background flows on the same public data.
Reading between the lines
- Inference: Because early stopping never triggered and validation negative log-likelihood was still decreasing at the final epoch, the five flows are likely undertrained; retraining to convergence could change both signal/background separation and the shape of the high-MET tail.
- Inference: The failure of linear and quadratic extrapolations suggests the background density should be MET-conditioned (e.g., a flow with MET as a conditioning input) to distinguish an extrapolation artifact from a genuine high-MET background population.
- Inference: On a larger dataset, this single-step VR-to-SR shape transfer would likely become the limiting systematic before statistical gains arrive, so a dedicated high-MET control sample or a closure-based reweighting would be needed.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a machine-learning search for dark matter produced in association with a leptonically decaying Z boson, using CMS Run 2015D open data and simplified-model Monte Carlo. Neural Spline Flows are trained on SM control-region data and on three mediator-specific signal MC samples; a per-event log-likelihood-ratio score is built and combined in a simultaneous SR+VR binned profile-likelihood fit. The authors report observed (expected) 95% CL upper limits on the signal strength for scalar, vector, and axial-vector mediators, e.g. mu < 0.0177 (0.0018), and note explicitly that the observed limits are weaker than expected due to a high-MET background-modeling residual. Crucially, the paper's own validation in Section 7.3 shows that the CR-trained NSF density does not extrapolate reliably into the SR tail, and Section 7.2 attributes the large q0 and the 7-12x observed-to-expected ratios to this unresolved residual. Yet Table 2 quotes limits derived from exactly that background model.
Significance. If the limits were valid, this would constitute a novel application of Neural Spline Flow likelihood-ratio scoring to a mono-Z dark matter search using open CMS data, with a plausible claim to methodological novelty and a detailed reproducible pipeline. The paper is unusually transparent: it provides region definitions, fit configurations, numerical JSON outputs, and an explicit, quantified account of its own background-modeling failures. These strengths are real and should be credited. However, the central numerical claim is not supported: the background template used to set the limits is shown by the authors themselves to fail precisely in the signal-like high-MET tail, and the expected-limit/expected-band pairs in the result tables are internally inconsistent. As a physics search, the manuscript cannot stand. The reproducible pipeline might be of interest as a methods-only study, but the physics limit claim as written is not.
major comments (4)
- [Sections 7.2-7.3, Eq. (2), Table 2] The background model used for the limits is the single-step VR-to-SR shape transfer described in Section 7 and Eq. (2). Section 7.3's own validation shows that the CR-trained NSF density does not extrapolate to the high-MET SR tail: linear and quadratic extrapolations of the mean score miss the true SR-tail mean by 90-115 and 300-330 score units, respectively. Section 7.2 states that the 160-181 tail events (MET>=100 GeV) carry mean scores 140-195 units above the VR bulk, where the VR template has negligible support. The profile-likelihood fit then absorbs this shape discrepancy as signal, producing q0 values of 233-327 and observed/expected limit ratios of roughly 7-12. Since the quoted 95% CL limits in Table 2 are derived from exactly this unvalidated background model, their coverage is not established. This is an internal inconsistency between the validation results and the central cl
- [Table 2 and Appendices E.2, E.4, E.5] The reported expected limits and expected bands are mutually inconsistent. In Table 2, the scalar expected limit is mu95_exp = 0.0018, while the 95% expected band is [0.00154, 0.00166]; the vector expected limit 0.0039 is above the quoted 95% band [0.00309, 0.00382]. The same pattern appears in Appendix E.4 (scalar 0.0018 vs [0.00145, 0.00176]) and Appendix E.5. In an asymptotic CLs calculation, the median expected limit must lie inside the central 68% interval and certainly inside the 95% interval. These numbers therefore cannot all be correct, indicating a procedural or computational error in the limit pipeline that directly affects the central results.
- [Section 6 (NSF training and early stopping)] The manuscript states that for all five flows, validation NLL continued to decrease at epoch 200 and early stopping was never triggered; every checkpoint is therefore the final epoch rather than a converged minimum. The authors flag this as 'potential residual undertraining'. This is not a minor caveat: the per-event scores are the entire basis of the likelihood-ratio test statistic, and underfit densities will distort the score distribution, particularly in the sparse high-MET tail where the analysis' main difficulty lies. The paper does not quantify the effect of non-convergence on the reported limits. Without converged flows or a convergence study, the learned densities are not a reliable foundation for the quoted numerical results.
- [Section 3, Table 5, Table 2 (axial-vector signal model)] The axial-vector signal sample combines two physically distinct benchmarks: M_chi=10 GeV, M_V=20 GeV (sigma=1.856 pb) and M_chi=50 GeV, M_V=200 GeV (sigma=0.158 pb). A single Neural Spline Flow is trained on the combined sample, and the same fitted signal strength mu is then converted into separate cross-section limits for each benchmark in Table 2. The signal density used in the likelihood is thus a mixture of two different spectra with different kinematics and different cross sections; it is not the density for either benchmark individually. The resulting cross-section limits are therefore not interpretable as constraints on either benchmark point. This should either be treated as two separate hypotheses or as a mixture with a well-defined composition; as written, the formalism is not well defined.
minor comments (5)
- [Eq. (1)] The notation p(x|SM_ell ell) leaves the channel dependence implicit. Since the two SM flows are trained on different channels, the score should explicitly indicate the channel, e.g. S_h^{(ee)} and S_h^{(mu mu)}.
- [Table 14 vs. Eq. (2)] The fit configuration lists 'Statistical method: chi2 asymptotic CLs approximation,' while Eq. (2) defines a binned Poisson likelihood. The relationship between the Poisson likelihood and the chi2 approximation used for the scans should be clarified.
- [Section 6, footnote after training protocol] The footnote about the significance cap (Z=8.0 due to floating-point underflow) is placed in the middle of the training protocol. It belongs in the results section, and the q0 values themselves should be reported with their numerical precision.
- [Figure 2 caption] The caption says 'Pre-unblinding validation-region score distributions,' but no unblinding procedure or decision rule is described anywhere. Please clarify the blinding protocol or reword the caption.
- [Section 7.1] MET resolution and pileup systematics are explicitly not propagated. Given that the signal region and the dominant background residual are defined by MET, even a rough estimate of these effects is necessary; as written, the systematic budget is incomplete and this gap should be acknowledged more prominently in the conclusions.
Circularity Check
No significant circularity: the CR-trained SM density, independent DM MC, and VR-sideband profile likelihood form a self-contained derivation chain; the documented high-MET residual is a validation failure, not a definitional or self-citation reduction.
full rationale
The derivation chain is not circular. The SM NSF density is trained exclusively on CR events with MET < 50 GeV (Sec. 6); the three DM densities are trained on separate MonoZToLL MC samples (Sec. 3); the per-event score (Eq. 1) is the log-density ratio between these independently trained flows; and the final limits come from a binned profile-likelihood fit (Eq. 2) in which the SR background template is obtained by a single-step VR to SR shape transfer: 'the SM VR score histogram is renormalised to the respective region yield and used as the nominal background prediction for that region' (Sec. 7). This is a sideband transfer, not a fitted parameter renamed as a prediction, and Appendix E.3 explicitly states: 'This approach avoids a pure SR-direct circularity while keeping the SR score shape as the primary discriminant.' The signal-strength mu is not an input to either flow and is not forced by the signal-model choice, since the signal MC is independent of the background data. There are no load-bearing self-citations: the cited CMS/ATLAS results, theory papers, and CERN open-data records are external, and the 'first application' claim is not used to justify the limits. The paper's own validation (Secs. 7.2-7.3 and App. E.4) shows that the CR-trained density does not extrapolate to the high-MET tail: 'Neither functional form reliably extrapolates the CR-trained NSF density into the high-MET SR tail,' and the VR-extrapolated template 'does not close in the high-score tail.' This is a genuine background-modelling and validation failure that undermines the quoted limits as physics constraints, but it is not a circular construction: the mismatch is measured against the same external SR data rather than being generated by the model. Other noted limitations (possible NSF undertraining in Sec. 6; unpropagated MET-resolution and pileup systematics in Sec. 7.1) likewise affect robustness, not circularity. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (5)
- signal strength mu =
scalar 0.0177, vector 0.0362, axial-vector 0.0498 (observed 95% CL upper limits)
- per-channel normalization nuisances theta_mumu, theta_ee =
constrained by VR data, posterior uncertainty ~0.17 sigma
- NSF hyperparameters =
8 transforms, 8 bins, 2x256 hidden units, lr schedule, 200 epochs
- SM feature cleaning caps =
e.g., met pt cap 200 GeV, hadronic recoil cap 564 GeV, derived from data quantiles
- score histogram binning =
30 initial bins merged to min 20 background events
assumptions (6)
- domain assumption Drell-Yan Z+jets events in the control region (MET<50 GeV) are representative of the background shape in the signal region.
- domain assumption The MonoZToLL simulated samples correctly model the kinematics of the simplified-model DM mediators.
- domain assumption Normalizing flows can accurately approximate the 37-dimensional event densities from the available training data.
- standard math Asymptotic formulae for the CLs method apply to the binned likelihood fits.
- domain assumption The single-step VR->SR shape transfer provides a valid nominal background prediction.
- domain assumption The luminosity and lepton efficiency uncertainty values from external CMS measurements apply to this analysis.
Cite this review
Pith. "Pith review of Mono-Z Dark Matter Search with Neural Spline Flows Using CMS Run 2015D Open Data." pith.science (2026). https://pith.science/paper/MLR4TZBU
@misc{pith2026260713771,
author = {Pith},
title = {Pith review of: Mono-Z Dark Matter Search with Neural Spline Flows Using CMS Run 2015D Open Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/MLR4TZBU}},
note = {Machine review of arXiv:2607.13771}
}
abstract
We report a search for dark matter (DM) produced in association with a leptonically decaying \(Z\) boson at \(\sqrt{s}=13\) TeV using CMS Run 2015D open data corresponding to an integrated luminosity of \(2.32\,\mathrm{fb}^{-1}\) together with simplified-model Monte Carlo simulation. Events are selected in the mono-\(Z\rightarrow\ell^+\ell^-\) final state in both the \(\mu\mu\) and \(ee\) channels. Forty kinematic observables are extracted from MINIAOD and MINIAODSIM, cleaned with physics-motivated selections, and reduced to a 37-dimensional feature vector. Five Neural Spline Flows are trained independently to model Standard Model background and mediator-specific DM signal densities. The per-event test statistic is constructed from the log-likelihood ratio between the signal and background density estimates, providing sensitivity across the full kinematic phase space without requiring a hard upper \(\mathrm{MET}\) threshold. A simultaneous profile-likelihood fit combining the two channels yields observed (expected) 95\% confidence level upper limits on the signal-strength parameter of \(\mu<0.0177\) (\(0.0018\)) for the scalar mediator, \(\mu<0.0362\) (\(0.0039\)) for the vector mediator, and \(\mu<0.0498\) (\(0.0069\)) for the axial-vector mediator. The observed limits are weaker than expected because of a residual high-\(\mathrm{MET}\) background-modeling discrepancy rather than evidence for a DM signal. To our knowledge, this is the first application of Neural Spline Flow likelihood-ratio scoring to a mono-\(Z\) dark matter search using CMS Run 2015D open data simultaneously in the \(\mu\mu\) and \(ee\) channels.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Neural spline flows.Advances in Neural Informa- tion Processing Systems, 32, 2019
Conor Durkan, Artur Bekasov, Iain Mur- ray, and George Papamakarios. Neural spline flows.Advances in Neural Informa- tion Processing Systems, 32, 2019
2019
-
[2]
Normaliz- ing flows for probabilistic modeling and inference.J
George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. Normaliz- ing flows for probabilistic modeling and inference.J. Mach. Learn. Res., 22:1–64, 2021
2021
-
[3]
Jessica Goodman, Masahiro Ibe, Arvind Rajaraman, William Shepherd, Tim M. P. Tait, and Hai-Bo Yu. Constraints on dark matter from colliders.Phys. Rev. D, 82:116010, 2010
2010
-
[4]
Bell, Ahmad J
Nicole F. Bell, Ahmad J. Galea, James B. Dent, Thomas D. Jacques, Lawrence M. Krauss, and Thomas J. Weiler. Searching for dark matter at the LHC with a mono-Z. Phys. Rev. D, 86:096011, 2012
2012
-
[5]
Carpenter, Andrew Nelson, Chase Shimmin, Tim M
Linda M. Carpenter, Andrew Nelson, Chase Shimmin, Tim M. P. Tait, and Daniel Whiteson. Collider searches for dark matter in events with a Z boson and missing energy.Phys. Rev. D, 87:074005, 2013
2013
-
[6]
Dark matter benchmark models for early LHC run-2 searches: Report of the AT- LAS/CMS dark matter forum.arXiv preprint, 2015
ATLAS and CMS Collaborations. Dark matter benchmark models for early LHC run-2 searches: Report of the AT- LAS/CMS dark matter forum.arXiv preprint, 2015. Editors: Antonio Boveia and Caterina Doglioni
2015
-
[7]
Searches for dark matter at the LHC: A multivariate analysis in the mono-Z channel.Phys
Alexandre Alves and Kuver Sinha. Searches for dark matter at the LHC: A multivariate analysis in the mono-Z channel.Phys. Rev. D, 92:115005, 2015
2015
-
[8]
Search for new physics in events with a leptonically decaying Z boson and a large transverse momentum imbalance in proton–proton collisions at√s = 13 TeV.Phys
CMS Collaboration. Search for new physics in events with a leptonically decaying Z boson and a large transverse momentum imbalance in proton–proton collisions at√s = 13 TeV.Phys. Rev. D, 97:092005, 2018
2018
Show all 29 references
-
[9]
Search for dark mat- ter produced in association with a leptoni- cally decaying Z boson in proton–proton collisions at √s = 13 TeV.Eur
CMS Collaboration. Search for dark mat- ter produced in association with a leptoni- cally decaying Z boson in proton–proton collisions at √s = 13 TeV.Eur. Phys. J. C, 81:13, 2021
2021
-
[10]
mono-Z ′: searches for dark matter in events with a resonance and missing transverse energy
Marcelo Autran, Kevin Bauer, Tongyan Lin, and Daniel Whiteson. mono-Z ′: searches for dark matter in events with a resonance and missing transverse energy. Phys. Rev. D, 92:115014, 2015
2015
-
[11]
Fox, Roni Harnik, Graham D
Wolfgang Altmannshofer, Patrick J. Fox, Roni Harnik, Graham D. Kribs, and Nir- mal Raj. Dark matter signals in dilepton production at hadron colliders.Phys. Rev. D, 91:015005, 2015
2015
-
[12]
Probing the dark sector through mono-Z boson leptonic decays.Phys
Daneng Yang and Qiang Li. Probing the dark sector through mono-Z boson leptonic decays.Phys. Rev. D, 97:015022, 2018
2018
-
[13]
Elgammal
S. Elgammal. Angular distribution study for high mass dimuon pairs in CMS open 2012 data and for mono-Z ′ model.arXiv preprint, 2024
2012
-
[14]
Testing the boundaries: Normaliz- ing flows for higher dimensional data sets
Humberto Reyes-Gonz´ alez and Riccardo Torre. Testing the boundaries: Normaliz- ing flows for higher dimensional data sets. SciPost Phys., 13:047, 2022
2022
-
[15]
Event generation and den- sity estimation with surjective normalizing flows.SciPost Phys., 13:047, 2022
Rob Verheyen. Event generation and den- sity estimation with surjective normalizing flows.SciPost Phys., 13:047, 2022. See arXiv:2205.01697 for publication details
2022 arXiv
-
[16]
Simulation assisted likelihood-free anomaly detection.Phys
Anders Andreassen, Benjamin Nachman, and David Shih. Simulation assisted likelihood-free anomaly detection.Phys. Rev. D, 101:055004, 2020
2020
-
[17]
DoubleMuon pri- mary dataset in MINIAOD format from RunD of 2015 (/DoubleMuon/Run2015D- 16Dec2015-v1/MINIAOD)
CMS Collaboration. DoubleMuon pri- mary dataset in MINIAOD format from RunD of 2015 (/DoubleMuon/Run2015D- 16Dec2015-v1/MINIAOD). CERN Open Data Portal, Record 24127, 2021. DOI: 10.7483/OPENDATA.CMS.H3TX.ZJZX
2015 doi
-
[19]
Simulated dataset DarkMatter MonoZToLL V Mx-1 Mv- 500 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data
CMS Collaboration. Simulated dataset DarkMatter MonoZToLL V Mx-1 Mv- 500 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16630, 2021. DOI: 10.7483/OPENDATA.CMS.YTLH.E0N7
2015 doi
-
[20]
Simulated dataset DarkMatter MonoZToLL A Mx-10 Mv- 20 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data
CMS Collaboration. Simulated dataset DarkMatter MonoZToLL A Mx-10 Mv- 20 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16575,
2015
-
[21]
Simulated dataset DarkMatter MonoZToLL A Mx-50 Mv- 200 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data
CMS Collaboration. Simulated dataset DarkMatter MonoZToLL A Mx-50 Mv- 200 gDMgQ-1 TuneCUETP8M1 13TeV- madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16597, 2021. DOI: 10.7483/OPENDATA.CMS.34IE.KN6I
2015 doi
-
[22]
Sim- ulated dataset DarkMat- ter MonoZToLL EWK Scalar Mx- 100 Lambda- 3000 TuneCUETP8M1 13TeV-madgraph in MINIAODSIM format for 2015 collision data
CMS Collaboration. Sim- ulated dataset DarkMat- ter MonoZToLL EWK Scalar Mx- 100 Lambda- 3000 TuneCUETP8M1 13TeV-madgraph in MINIAODSIM format for 2015 collision data. CERN Open Data Portal, Record 16601, 2021. DOI: 10.7483/OPEN- DATA.CMS.NW7F.NFGG
2015 doi
-
[23]
Asymptotic formulae for likelihood-based tests of new physics.Eur
Glen Cowan, Kyle Cranmer, Eilam Gross, and Ofer Vitells. Asymptotic formulae for likelihood-based tests of new physics.Eur. Phys. J. C, 71:1554, 2011
2011
-
[24]
Pseudoscalar portal dark matter.Phys
Asher Berlin, Stefania Gori, Tongyan Lin, and Lian-Tao Wang. Pseudoscalar portal dark matter.Phys. Rev. D, 92:015005, 2015
2015
-
[25]
Seng Pei Liew, Michele Papucci, Alessan- dro Vichi, and Kathryn M. Zurek. Mono-X versus direct searches: Simplified models for dark matter at the LHC.JHEP, 03:100, 2017
2017
-
[26]
Search for high- mass new phenomena in the dilepton fi- nal state using proton–proton collisions at √s = 13 TeV with the ATLAS detector
ATLAS Collaboration. Search for high- mass new phenomena in the dilepton fi- nal state using proton–proton collisions at √s = 13 TeV with the ATLAS detector. Phys. Lett. B, 761:372–392, 2016
2016
-
[27]
Search for new high-mass phenomena in the dilepton final state using 36 fb −1 of proton–proton colli- sion data at √s = 13 TeV with the ATLAS detector.Phys
ATLAS Collaboration. Search for new high-mass phenomena in the dilepton final state using 36 fb −1 of proton–proton colli- sion data at √s = 13 TeV with the ATLAS detector.Phys. Lett. B, 776:318–338, 2018
2018
-
[28]
Anomaly detection with density estima- tion.Phys
Benjamin Nachman and David Shih. Anomaly detection with density estima- tion.Phys. Rev. D, 101:075042, 2020
2020
-
[29]
Classifying anomalies through outer density estimation.Phys
Anna Hallin et al. Classifying anomalies through outer density estimation.Phys. Rev. D, 106:055006, 2022
2022
-
[30]
Precision luminosity measurement in proton-proton collisions at √s = 13 TeV in 2015 and 2016 at cms
CMS Collaboration. Precision luminosity measurement in proton-proton collisions at √s = 13 TeV in 2015 and 2016 at cms. Eur. Phys. J. C, 81:800, 2021. 29
2015
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.