{"id":"a8ac7426-8498-49b0-8c65-386a48b63427","arxiv_id":"2607.13771","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A mono-Z dark matter search with neural spline flows on CMS 2015 open data sets limits of mu<0.018–0.050, but these limits are invalid because the background model fails in the high-MET tail.","lead":"This paper searches for dark matter produced with a Z boson in CMS open data, using neural spline flows to score events. It reports 95% CL upper limits that are 7–12 times weaker than expected because of a background-modeling mismatch in the high-MET tail, not because a signal was seen.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quoted 95% CL limits are not supported because the SR background template is the paper's own VR→SR transfer, which Sec. 7.3 shows fails in the high-MET tail; the fit's 7–12× observed/expected ratio is a symptom of that unmodeled shape residual, not a calibrated limit.","rationale":"The paper is commendably transparent about its limitations, and I credit the internal honesty. But a 95% CL upper limit is only a constraint if the background prediction in the signal region is reliable. The paper's own Section 7.3 shows it is not: the SM NSF density trained at MET<50 GeV does not extrapolate to MET≥100 GeV, with extrapolation errors of 90–330 score units. Section 7.2 shows that this unmodeled high-MET tail is what the fit interprets as signal, producing the inflated observed limits. The 5% normalization nuisance is a normalization, not a shape, uncertainty, so it cannot cover the residual. Appendix E.4's VR-extrapolated fit fails in the same tail, ruling out one trivial explanation. I therefore agree with the reader's weakest assumption and with the REJECT verdict; the methodological template may be reusable, but the quoted physics limits are not supported by the analysis as presented.","tokens_in":18361,"tokens_out":9695,"duration_ms":128569,"concrete_test":"Rerun the simultaneous SR+VR profile-likelihood fit and CLs scan after redefining the search region as 50≤MET<100 GeV — the current VR, where the background model closes — keeping all other settings identical to Table 2. If the observed/expected limit ratios collapse from the quoted 7–12 to approximately 1 and μ̂ becomes compatible with 0, the published limits are driven by the unmodeled MET≥100 tail and are not robust constraints. If the ratios persist, the tail residual is not the sole cause and the reader's weakest-assumption would need to be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 7 defines the SR background by a single-step VR→SR shape transfer: the SM VR score histogram is renormalized to the SR yield and used as the nominal background template. The paper's own validation (Sec. 7.3) demonstrates that the CR-trained NSF density does not extrapolate to the high-MET SR tail: linear and quadratic extrapolations of the mean NSF score miss the measured tail by 90–115 and 300–330 score units, respectively. Sec. 7.2 quantifies the consequence: the 160–181 events with MET≥100 GeV (2.4–2.5% of SR) carry mean scores 140–195 units above the VR bulk, in a region where the VR template has negligible support. The profile-likelihood fit then absorbs this shape discrepancy as signal, giving q0≈233–327 (Z capped at 8.0) and observed/expected limit ratios of roughly 7–12. The 5% per-channel normalization nuisance cannot cover a shape shift of this size, and Appendix E.4 confirms that a pure VR-extrapolated template fails closure in the same tail. Since the central numbers in Table 2 are derived from a background model the paper itself shows is wrong in exactly the region where mono-Z DM signal would appear, the 95% coverage of the quoted limits is not established. This is an internal inconsistency between the validation results and the central claim, not a matter of external consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a machine-learning search for dark matter produced in association with a leptonically decaying Z boson, using CMS Run 2015D open data and simplified-model Monte Carlo. Neural Spline Flows are trained on SM control-region data and on three mediator-specific signal MC samples; a per-event log-likelihood-ratio score is built and combined in a simultaneous SR+VR binned profile-likelihood fit. The authors report observed (expected) 95% CL upper limits on the signal strength for scalar, vector, and axial-vector mediators, e.g. mu < 0.0177 (0.0018), and note explicitly that the observed limits are weaker than expected due to a high-MET background-modeling residual. Crucially, the paper's own validation in Section 7.3 shows that the CR-trained NSF density does not extrapolate reliably into the SR tail, and Section 7.2 attributes the large q0 and the 7-12x observed-to-expected ratios to this unresolved residual. Yet Table 2 quotes limits derived from exactly that background model.","tokens_in":18729,"tokens_out":6618,"duration_ms":67934,"significance":"If the limits were valid, this would constitute a novel application of Neural Spline Flow likelihood-ratio scoring to a mono-Z dark matter search using open CMS data, with a plausible claim to methodological novelty and a detailed reproducible pipeline. The paper is unusually transparent: it provides region definitions, fit configurations, numerical JSON outputs, and an explicit, quantified account of its own background-modeling failures. These strengths are real and should be credited. However, the central numerical claim is not supported: the background template used to set the limits is shown by the authors themselves to fail precisely in the signal-like high-MET tail, and the expected-limit/expected-band pairs in the result tables are internally inconsistent. As a physics search, the manuscript cannot stand. The reproducible pipeline might be of interest as a methods-only study, but the physics limit claim as written is not.","major_comments":[{"comment":"The background model used for the limits is the single-step VR-to-SR shape transfer described in Section 7 and Eq. (2). Section 7.3's own validation shows that the CR-trained NSF density does not extrapolate to the high-MET SR tail: linear and quadratic extrapolations of the mean score miss the true SR-tail mean by 90-115 and 300-330 score units, respectively. Section 7.2 states that the 160-181 tail events (MET>=100 GeV) carry mean scores 140-195 units above the VR bulk, where the VR template has negligible support. The profile-likelihood fit then absorbs this shape discrepancy as signal, producing q0 values of 233-327 and observed/expected limit ratios of roughly 7-12. Since the quoted 95% CL limits in Table 2 are derived from exactly this unvalidated background model, their coverage is not established. This is an internal inconsistency between the validation results and the central cl","section":"Sections 7.2-7.3, Eq. (2), Table 2"},{"comment":"The reported expected limits and expected bands are mutually inconsistent. In Table 2, the scalar expected limit is mu95_exp = 0.0018, while the 95% expected band is [0.00154, 0.00166]; the vector expected limit 0.0039 is above the quoted 95% band [0.00309, 0.00382]. The same pattern appears in Appendix E.4 (scalar 0.0018 vs [0.00145, 0.00176]) and Appendix E.5. In an asymptotic CLs calculation, the median expected limit must lie inside the central 68% interval and certainly inside the 95% interval. These numbers therefore cannot all be correct, indicating a procedural or computational error in the limit pipeline that directly affects the central results.","section":"Table 2 and Appendices E.2, E.4, E.5"},{"comment":"The manuscript states that for all five flows, validation NLL continued to decrease at epoch 200 and early stopping was never triggered; every checkpoint is therefore the final epoch rather than a converged minimum. The authors flag this as 'potential residual undertraining'. This is not a minor caveat: the per-event scores are the entire basis of the likelihood-ratio test statistic, and underfit densities will distort the score distribution, particularly in the sparse high-MET tail where the analysis' main difficulty lies. The paper does not quantify the effect of non-convergence on the reported limits. Without converged flows or a convergence study, the learned densities are not a reliable foundation for the quoted numerical results.","section":"Section 6 (NSF training and early stopping)"},{"comment":"The axial-vector signal sample combines two physically distinct benchmarks: M_chi=10 GeV, M_V=20 GeV (sigma=1.856 pb) and M_chi=50 GeV, M_V=200 GeV (sigma=0.158 pb). A single Neural Spline Flow is trained on the combined sample, and the same fitted signal strength mu is then converted into separate cross-section limits for each benchmark in Table 2. The signal density used in the likelihood is thus a mixture of two different spectra with different kinematics and different cross sections; it is not the density for either benchmark individually. The resulting cross-section limits are therefore not interpretable as constraints on either benchmark point. This should either be treated as two separate hypotheses or as a mixture with a well-defined composition; as written, the formalism is not well defined.","section":"Section 3, Table 5, Table 2 (axial-vector signal model)"}],"minor_comments":[{"comment":"The notation p(x|SM_ell ell) leaves the channel dependence implicit. Since the two SM flows are trained on different channels, the score should explicitly indicate the channel, e.g. S_h^{(ee)} and S_h^{(mu mu)}.","section":"Eq. (1)"},{"comment":"The fit configuration lists 'Statistical method: chi2 asymptotic CLs approximation,' while Eq. (2) defines a binned Poisson likelihood. The relationship between the Poisson likelihood and the chi2 approximation used for the scans should be clarified.","section":"Table 14 vs. Eq. (2)"},{"comment":"The footnote about the significance cap (Z=8.0 due to floating-point underflow) is placed in the middle of the training protocol. It belongs in the results section, and the q0 values themselves should be reported with their numerical precision.","section":"Section 6, footnote after training protocol"},{"comment":"The caption says 'Pre-unblinding validation-region score distributions,' but no unblinding procedure or decision rule is described anywhere. Please clarify the blinding protocol or reword the caption.","section":"Figure 2 caption"},{"comment":"MET resolution and pileup systematics are explicitly not propagated. Given that the signal region and the dominant background residual are defined by MET, even a rough estimate of these effects is necessary; as written, the systematic budget is incomplete and this gap should be acknowledged more prominently in the conclusions.","section":"Section 7.1"}],"recommendation":"reject","confidential_remarks":"This paper is unusual in that it documents, in detail, the failure of its own background model and then still presents limits derived from that failed model. The referee report can rely on this internal inconsistency rather than on any external assumption. I recommend rejection. If the authors wish to preserve the methodological contribution, they would need to resubmit as a methods-only study with no physics limit claim, or first redo the analysis with a background model that closes in the SR tail."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is an honest piece of work that nonetheless fails to support its headline result. The authors train neural spline flows to score events in a mono-Z dark-matter search on CMS open data, and they report 95% CL upper limits on signal strength. But their own validation shows the background template breaks exactly where the signal would appear, so the limits are not reliable. The reader's REJECT verdict is right, and the stress-test note lands: Section 7.3 shows the CR-trained flow does not extrapolate to the high-MET tail, with linear and quadratic models missing the tail mean by 90–115 and 300–330 score units. Section 7.2 then explains that this residual drives the 7–12x inflation of observed versus expected limits. The central numbers in Table 2 come from a background model the paper itself proves is wrong in the signal region. That is a load-bearing flaw, not a minor caveat.\n\nCredit where it is due: the paper is unusually transparent. It ships a detailed pipeline, uses public data and MC, defines regions clearly, and explicitly quantifies the failure rather than hiding it. The literature review is fair and cites the relevant density-ratio methods (ANODE, CATHODE, SALAD) and the full Run 2 CMS search. The authors also flag the undertraining of the flows and the missing MET/pileup systematics. As a methodological write-up, it is a solid example of how to document a failed extrapolation in a likelihood-ratio search.\n\nThe soft spots are severe, though. The single-step VR→SR shape transfer is the backbone of the fit, and it does not close in the tail. The scalar cross-section limit is absurdly small (1.76e-9 pb), which signals a normalization issue. The axial-vector sample merges two different benchmark masses, making the physical interpretation murky. And the per-channel normalization nuisance of 5% cannot cover a shape shift of this size. The paper's own diagnosis is that the observed limits are inflated by an unmodeled residual, which is another way of saying the limits are not valid constraints.\n\nFor a reader: this is not a paper whose physics numbers you should quote. It is a cautionary tale about density-ratio methods in sparse tails, and the honest failure analysis could be useful methodologically. I would bring it to a reading group as a case study, but I would not cite the limits myself.\n\nShould it go to peer review? Yes, I think so. It is a serious, reproducible attempt with a clearly articulated problem, and a good referee could either guide a major revision (e.g., using a more robust background model) or provide a definitive rejection with useful feedback. It deserves more than a desk reject.","headline":"A careful, honest open-data study whose central limits are undercut by the authors' own demonstration that the background model fails in the high-MET tail.","tokens_in":19270,"tokens_out":2272,"would_cite":false,"duration_ms":23664,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using neural-spline-flow likelihood-ratio scores on public 2015 proton-proton collision data, this analysis sets observed 95% upper limits on the dark matter signal strength for three mediator models, and attributes the 7–12x gap between ob","keywords":["dark matter","mono-Z","neural spline flows","likelihood ratio","missing transverse momentum","profile likelihood","open data","density estimation"],"falsifier":"Build a background flow that includes events from an independent high-MET sideband (MET between 100 and 200 GeV) and repeat the signal-region fit. If the fitted signal strength drops toward zero and observed limits approach expected limits, the high-MET tail residual was the cause; if the positive signal strength persists with high q0, the residual is intrinsic to the flow/density modeling or to the signal model rather than a CR-to-SR transfer artifact.","tokens_in":18264,"feed_emoji":"🌌","tokens_out":10675,"duration_ms":89331,"temperature":0.7,"pith_summary":"The paper aims to establish that a mono-Z dark matter search can be run end-to-end on public 2015 collision data using a learned density-ratio score instead of hand-crafted variables. Five neural spline flows—two background flows trained on control-region events with missing transverse momentum below 50 GeV for the muon and electron channels, and three signal flows trained on mediator Monte Carlo—produce the per-event score as the log-density difference between signal and background. A simultaneous signal-plus-validation profile-likelihood fit yields observed 95% upper limits on the signal-strength parameter of 0.0177 (scalar), 0.0362 (vector), and 0.0498 (axial-vector), with expected limits roughly 7–12 times smaller. The paper's central interpretive claim is that the gap is driven by a high-MET (≥100 GeV) background-shape residual, not by evidence for dark matter, and it documents that neither linear nor quadratic extrapolations predict that tail.","feed_headline":"Observed mono-Z dark matter limits are 7-12x weaker than expected","feed_subtitle":"Machine-learned search on 2015 open data ties the gap to a high missing-energy tail in the background model.","key_machinery":"The engine is the per-event log-likelihood-ratio score formed from two Neural Spline Flow density estimates: S_h(x) = log p(x|DM_h) − log p(x|SM_channel). A Neural Spline Flow is a normalizing flow whose coupling transforms are monotonic rational-quadratic splines, giving exact tractable densities. Background flows are trained only on control-region events (MET < 50 GeV), one per lepton channel; signal flows are trained on simulated mediator events and evaluated in the same standardized SM feature domain. The score arrays feed a binned profile-likelihood fit over signal and validation regions (MET 50–100 GeV as a sideband) with per-channel normalization nuisances, and asymptotic CL_s formula","core_discovery":"On its own terms, the paper claims that a likelihood-ratio test statistic built from independently trained neural spline flows can serve as the full discriminant for a mono-Z dark matter search, removing the need for a hard upper cut on missing transverse momentum. In the signal region (MET ≥ 50 GeV, |Δφ(MET, Z)| > 2.5, ≤ 1 jet), the per-event score S(x) = log p(x|signal) − log p(x|background), formed from three mediator-specific signal flows and two channel-specific background flows, is fed into a simultaneous signal-region plus validation-region binned profile likelihood. The result is observed 95% CL upper limits on the signal strength μ of <0.0177 (scalar), <0.0362 (vector), and <0.0498","pith_inferences":["Inference: Because early stopping never triggered and validation negative log-likelihood was still decreasing at the final epoch, the five flows are likely undertrained; retraining to convergence could change both signal/background separation and the shape of the high-MET tail.","Inference: The failure of linear and quadratic extrapolations suggests the background density should be MET-conditioned (e.g., a flow with MET as a conditioning input) to distinguish an extrapolation artifact from a genuine high-MET background population.","Inference: On a larger dataset, this single-step VR-to-SR shape transfer would likely become the limiting systematic before statistical gains arrive, so a dedicated high-MET control sample or a closure-based reweighting would be needed."],"forward_implications":["The density-ratio score concentrates sensitivity across the full phase space, so the search does not rely on a hard upper MET threshold; events up to the dataset's ~200 GeV MET cap enter the fit.","A validation-region sideband can supply a data-driven background normalization constraint of roughly 0.85% per channel, even on a small 2.32 fb^-1 dataset.","The observed-to-expected ratio remains ~7–12 across three mediator hypotheses and three background variants, identifying the high-MET tail as the dominant systematic rather than the signal model.","The pipeline is transferable to other final states or signal hypotheses by retraining the relevant signal and background flows on the same public data."],"fun_headline_variants":["Neural spline flow mono-Z search: limits 7-12x weaker than expected","First neural spline flow DM search on open data: limits miss expectation","Open data DM search with neural flows: observed limits off by 7-12x","Mono-Z DM with neural flows: observed limits 7-12x off expectation","Neural spline flows for mono-Z DM: limits 7-12x weaker"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The background score distribution in the signal region—especially the rare events with missing transverse momentum above 100 GeV—is correctly described by a background density trained on MET < 50 GeV events and shifted by one normalization step; the paper's own validation shows simple extrapolations fail in that tail, so if this assumption fails the quoted limits are not valid constraints.","fun_headline_variants_meta":{"raw":{"variants":["Neural spline flow mono-Z search: limits 7-12x weaker than expected","First neural spline flow DM search on open data: limits miss expectation","Open data DM search with neural flows: observed limits off by 7-12x","Mono-Z DM with neural flows: observed limits 7-12x off expectation","Neural spline flows for mono-Z DM: limits 7-12x weaker"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001212,"raw_usage":{"total_tokens":4906,"prompt_tokens":907,"completion_tokens":3999,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":3900}},"tokens_in":651,"tokens_out":3999,"duration_ms":58319,"temperature":1.0,"reasoning_tokens":3900,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:44:55.832629+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a background flow that includes events from an independent high-MET sideband (MET between 100 and 200 GeV) and repeat the signal-region fit. If the fitted signal strength drops toward zero and observed limits approach expected limits, the high-MET tail residual was the cause; if the positive signal strength persists with high q0, the residual is intrinsic to the flow/density modeling or to the signal model rather than a CR-to-SR transfer artifact.","supporting_citations":[],"review_version":1}