{"id":"c0b62b8c-ad00-4f9b-b7a4-5d0c137c89a4","arxiv_id":"2504.14303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A normalizing-flow neural network reconstructs the momenta of the neutrino and dark mediator in single top quark production and reproduces a top-quark spin-correlation variable more accurately than a multilayer perceptron.","lead":"The paper tests machine learning methods to infer the momenta of two invisible particles, the neutrino and a dark matter mediator, in single-top quark events at the LHC. A normalizing flows network reconstructed the angular correlation variable more accurately than a standard neural network, which could sharpen searches for dark matter produced with top quarks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The direct-applicability claim is unsupported: ν-Flows is trained and evaluated only on a 1:1 SM/DM MC mixture with fixed mediator mass, so the learned conditional distribution may not transfer to collision data.","rationale":"The reader identified the core problem: the analysis is purely phenomenological, with no real-data validation, and the underdetermined nature of the inverse problem means the network learns a conditional distribution from simulation. I agree with this assessment. My stress-test pass refines the concern into a concrete, testable failure mode: the training mixture is extremely narrow (one signal mass point and one SM background process, in a 1:1 ratio), so the learned conditional distribution is prior-dependent. Real collider data contains a different mixture and additional backgrounds, so out-of-distribution generalization is required but neither tested nor demonstrated. This is directly relevant to the central claim of direct applicability. The manuscript itself flags the lack of real data in Section 2.1, which is a limitation that the conclusion in Section 4 overstates. I do not see an internal inconsistency or a fatal flaw in the comparative ML result itself: the relative improvement in histogram metrics is large and the code is public, which are genuine strengths. Therefore the appropriate verdict remains CONDITIONAL, unchanged from the reader's assessment. The proposed concrete test — applying the trained network to a realistic mixed-background pseudo-data sample and to alternative mediator mass points — would settle whether the generalisation concern actually lands.","tokens_in":7994,"tokens_out":9629,"duration_ms":99484,"concrete_test":"Generate a Delphes-level pseudo-data sample with a realistic composition after the Section 2.1 selection, for example 80% ttbar, 15% W+jets, and 5% signal, and apply the already-trained ν-Flows without retraining. Compare the reconstructed cosθ_{bar l bar d} histogram with the truth-level histogram of the same mixture. In addition, repeat the evaluation with signal samples generated at mΦ = 200 GeV and mΦ = 600 GeV. If the reconstructed distribution deviates from truth by an amount comparable to or larger than the improvement margin over the MLP (e.g., hist MAE above about 100), the direct-applicability claim is not sustained; if the deviation remains small, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim, that the method 'can be directly applied to collider data' (Section 4), is not supported by the evidence presented. Section 2.1 explicitly states: 'At this stage, the analysis is phenomenological; real collider data is not used.' The network is trained on a 1:1 mixture of SM single-top and signal events with fixed mΦ = 400 GeV and mχ = 1 GeV. Because the inverse problem is underdetermined — six target momentum components are reconstructed from a finite set of visible final-state observables — a normalizing flow learns the conditional distribution p(pν, pΦ | x), which is shaped by the training prior. In real LHC data the selected event mixture will contain other backgrounds (e.g., ttbar, W+jets) and a different signal fraction or mediator mass, so the conditional distribution changes. The paper reports no evaluation on out-of-distribution inputs: no background-only sample, no mixed background-plus-signal pseudo-data, and no variation of model parameters. Consequently, the claim of direct applicability to collider data is an extrapolation beyond the tested regime, not a demonstrated property of the method.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies reconstruction of the momenta of the neutrino and the scalar dark matter mediator in associated single-top-plus-DM production at the LHC, with the goal of reconstructing an angular correlation variable cos(θ_{bar l bar d}) that discriminates signal from SM background. Three machine-learning approaches are compared: a multilayer perceptron (MLP), an autoregressive normalizing flow ('Basic Flows'), and a coupling-layer normalizing flow ('ν-Flows'). The authors report that ν-Flows yields the lowest histogram MAE and χ² scores at parton level and retains an advantage after DELPHES detector simulation, and they conclude that the method can be directly applied to collider data.","tokens_in":8252,"tokens_out":3568,"duration_ms":34834,"significance":"If substantiated, the result is a useful demonstration that conditional normalizing flows outperform point-estimate regressors for an underdetermined kinematic inverse problem in top-quark physics, with a concrete downstream benefit for a DM-search variable. The paper has several strengths: the evaluation uses an external MC-truth benchmark from standard generators (CompHEP/MadGraph, Pythia8, DELPHES), so the comparison is not circular; the analysis is presented at both parton and detector level; and the code is publicly released with a DOI. The main scientific claim, however, is stronger than the evidence: the conclusion that the method 'can be directly applied to collider data' is not supported by the phenomenological setup and the fixed training mixture, and the reported performance metrics lack statistical uncertainties.","major_comments":[{"comment":"The manuscript explicitly states in Section 2.1 that 'At this stage, the analysis is phenomenological; real collider data is not used,' yet Section 4 concludes that the method 'can be directly applied to collider data.' This is an extrapolation beyond the tested regime: the flow is trained and evaluated on a 1:1 SM/DM mixture with fixed mΦ = 400 GeV and mχ = 1 GeV, and the inverse problem is underdetermined, so the learned conditional distribution p(pν,pΦ | x) is shaped by the training prior. No out-of-distribution evaluation is reported: no background-only sample, no mixed background-plus-signal pseudo-data, no variation of the mediator mass or signal fraction, and no data/MC closure test. The direct-applicability claim should either be removed or be replaced by a clearly scoped statement that the method is ready for application to simulated signal-region studies, pending validation on more realistic event mixtures.","section":"Section 2.1 and Section 4"},{"comment":"Table 2 reports MAE, histogram MAE, and χ² score for the three architectures without any statistical uncertainties or multiple-seed statistics. Since neural-network training is stochastic and the χ² values are 335 versus 1557 versus 8985, the reader cannot assess whether the ordering of the methods is stable under retraining or whether the quoted differences are within seed-to-seed variance. The manuscript also does not state the size of the test set or the binning used for the histogram metrics, both of which are needed to interpret the χ² score and to reproduce the comparison. Please add uncertainties over training seeds and specify the test-set size and histogram binning.","section":"Table 2 and Section 3"},{"comment":"The detector-level comparison is presented only as figures, while the central quantitative table (Table 2) is limited to parton level. Since the conclusion in Section 4 explicitly claims superiority 'after simulation of the detector response,' the detector-level histogram MAE and χ² scores should be reported in the same tabular form as the parton-level results. In addition, the aggregation protocol for flow samples is not specified precisely: Section 3 states that the median is used, but the number of samples per event differs in the captions (5 for ν-Flows, 10 for Basic Flows in Figs. 9–10, and 10,000 in Fig. 7), and this choice can affect both the point-wise MAE and the downstream angular distribution.","section":"Section 3, Figures 5 and 6"}],"minor_comments":[{"comment":"The number of flow samples per event is inconsistent across captions: Fig. 7 says 10,000 points are sampled for each Flow-based network, while Figs. 9 and 10 say the median of 10 and 5 points is used. Please reconcile these numbers and state the final sample count used for the reported results.","section":"Figures 7, 9, 10"},{"comment":"The clipping of Basic Flows outputs to the interval [-10, 10] is mentioned in Section 3 but is not included in Table 1 or in the architecture description; please state the clipping range for both flow models and, if the range was tuned, report the tuning range.","section":"Section 2.2.3 and Table 1"},{"comment":"The list of input features is incomplete: the high-level variables are described only as 'various combinations of low-level variables,' without an explicit enumeration or the construction formulas. This hampers reproducibility independently of the public code; please provide the full feature list.","section":"Section 2.1"},{"comment":"The figure captions refer to the 'true_nophi' and 'reconstructed_nophi' distributions, but the term 'nophi' is not defined in the text; clarify which reconstruction uses only the neutrino without the mediator contribution.","section":"Section 2.1 / Figures 5 and 6"},{"comment":"There are several typos and formatting issues, including 'T able 1', 'reconstruiction', and the sentence 'the denominator of such a metric may be near zero' in Section 2.2.1; a careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a fairly standard ML-method comparison applied to an interesting physics case; the novelty is limited but the public code and clear benchmark are positive. The main issue for the journal is the gap between the phenomenological evidence and the 'directly applied to collider data' claim, plus the missing statistical uncertainties in the central comparison table. These are fixable with additional analysis or a substantial softening of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe one-line take: this is a solid, narrow simulation-based benchmark showing that ν-Flows reconstructs the top-quark spin-correlation variable in single-top + DM events far better than an MLP baseline, and it ships code. The caveat is that the final claim \"can be directly applied to collider data\" is not supported by anything in the paper.\n\nWhat's actually new: prior work tried analytic reconstruction [14] and separation studies [23], and the ν-Flows architecture comes from [27] applied to t-tbar. This paper is the first application of normalizing flows to the single-top + DM process with two invisible particles, benchmarked against an MLP. The numbers are striking: histogram MAE drops from 360.6 to 52.4 at parton level, chi-squared from 8985 to 335, and the improvement survives Delphes smearing. The figures show the MLP misses the distribution shape while ν-Flows captures it. Code is on Zenodo, which makes the central comparison reproducible.\n\nThe main soft spot is the direct-applicability sentence. Section 2.1 explicitly says real data is not used; the network is trained on a 1:1 SM/DM mixture with fixed mediator mass 400 GeV and DM mass 1 GeV. Because the problem is underdetermined, the flow learns the conditional distribution implied by that training mixture. Real LHC data will have different background composition, signal fraction, and possibly mediator mass, and the paper reports no evaluation on background-only samples, mixed pseudo-data, or varied model parameters. So the final claim is an extrapolation. It may hold, but the authors need to say it is a promise rather than a demonstrated property.\n\nSecondary issue: Table 2 and Figures 5-6 give single-run metrics with no uncertainties or multiple seeds. For a paper whose main evidence is a metric comparison, that's minor but worth fixing. The Basic Flows convergence handling by clipping outputs to [-10,10] is a bit ad hoc but not a problem.\n\nThe citation pattern is honest and the build on refs [14,23,27] is explicit. No circular derivation. The central comparative result holds up.\n\nWho this is for: people working on ML-based momentum reconstruction in top/DM searches at the LHC. It is an analysis-tool improvement, not new physics. I would send it to peer review; it deserves referee time even if the final claim needs to be reined in. I would not cite it in my own work unless I were in that exact niche.\n\nRecommendation: engage with the work, but ask the authors to soften the direct-applicability claim, add seed statistics, and ideally show one out-of-distribution test.","headline":"A credible ML comparison showing normalizing flows reconstruct a top-spin-correlation variable much better than an MLP, with code released; the direct-applicability claim outruns the phenomenological evidence.","tokens_in":8801,"tokens_out":1768,"would_cite":false,"duration_ms":16382,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a normalizing-flow network, ν-Flows, reconstructs the invisible momenta in associated top-quark plus dark-mediator production well enough to recover the angular correlation variable that separates signal from…","keywords":["dark matter mediator","single top quark production","angular correlations","normalizing flows","neutrino momentum reconstruction","machine learning","LHC phenomenology","top quark spin correlations"],"falsifier":"Train ν-Flows on events from one generator and test it on events from a second generator with the same truth labels; if the histogram MAE of the reconstructed angular variable rises to the level of the MLP baseline (around 360), the claimed performance is generator-specific. A further test on real LHC data in a Standard Model control region would reveal whether the reconstructed distribution matches the expected background shape within uncertainties.","tokens_in":7809,"feed_emoji":"⚛️","tokens_out":7691,"duration_ms":64489,"temperature":0.7,"pith_summary":"The paper tries to establish that a Normalizing Flows architecture called ν-Flows can reconstruct the momenta of the two invisible particles—the neutrino and the scalar dark-matter mediator—in single-top events, and that the reconstructed momenta reproduce the top-quark angular correlation variable $\\cos(\\theta_{\\bar{l}\\bar{d}})$ far more accurately than a multilayer perceptron. This variable is a spin-correlation observable that separates dark-matter production from Standard Model backgrounds in a simplified dark-matter model. At parton level the flow reconstruction gives a histogram mean absolute error of 52.4 and a $\\chi^2$ of 335, versus 360.6 and 8985 for the MLP; after detector simulation the improvement persists. The intended payoff is a practical offline tool that experimental searches could apply directly to LHC data.","feed_headline":"Normalizing flows reconstruct top-quark angular correlations better than MLPs","feed_subtitle":"For dark-matter searches, ν-Flows cuts histogram error from 360 to 52 in the key angular variable.","key_machinery":"The load-bearing object is the ν-Flows architecture, a conditional normalizing flow: an invertible neural transformation that maps a standard normal distribution to the six-dimensional distribution of the neutrino and mediator momenta, conditioned on event observables through an encoding network. The invertible coupling layers split the target variables into two groups in each layer, transform one group with piecewise rational quadratic splines, and use a fully connected network to pass conditioning information between groups. For each event the network is sampled many times and the median is used as the point estimate; the likelihood objective preserves correlations among target variables, which the paper argues is why the reconstructed angular distribution is closer to truth than the MLP's.","core_discovery":"For the process $pp \\to t(\\to \\nu l \\bar{b})q$ plus a scalar dark-matter mediator (mass 400 GeV, dark-matter mass 1 GeV), the paper's central claim is that ν-Flows, a normalizing-flow network with coupling layers and piecewise rational quadratic splines, reconstructs the six target momentum components of the neutrino and mediator from observed final-state objects. When these reconstructed momenta are used to build the angular variable $\\cos(\\theta_{\\bar{l}\\bar{d}})$ in the top-quark rest frame, the resulting histogram matches the true distribution with histogram MAE 52.4 and $\\chi^2$ score 335 at parton level, while the MLP baseline gives 360.6 and 8985. After hadronization and detector smearing, ν-Flows still gives the closest histogram (103.1 and 815.5, compared with 154.6 and 1554.9 for the MLP). The authors therefore conclude that the method 'can be directly applied to collider data' even though the final-state momenta do not uniquely determine the invisible momenta.","pith_inferences":["The same architecture could be tested on other processes with two invisible particles, such as top-quark pair production with additional new-physics particles, where no unique kinematic solution exists.","A decisive check the paper does not report is cross-generator transportability: training on one Monte Carlo generator and testing on another would reveal how much of the success is tied to simulation-specific features.","The per-event sample spread of the flow could be developed into a systematic uncertainty or an event-level weight for the angular distribution, an extension the paper does not exploit."],"forward_implications":["If ν-Flows works on real data, the angular variable $\\cos(\\theta_{\\bar{l}\\bar{d}})$ can be used in experimental searches for dark-matter mediators in the single-top final state, where it separates signal from background.","The method's advantage survives detector smearing, so it is suitable for offline analysis on reconstructed objects rather than only at generator level.","The flow's likelihood-based training does not need analytic solutions for the invisible momenta, removing the obstacle that blocked the analytical approach cited in the paper.","Because the network produces a distribution per event, the reconstructed angular variable can be aggregated by taking the median, while the sample spread gives a measure of reconstruction ambiguity."],"supporting_citations":[{"why":"Introduces the angular variable and provides the MLP baseline the paper compares against.","marker":"[14]"},{"why":"Generates the parton-level events used for training and testing; removing it breaks the dataset.","marker":"[17]"},{"why":"Sets the benchmark model parameters (mχ = 1 GeV, mΦ = 400 GeV, couplings gf = gχ = 1) used in the study.","marker":"[19]"},{"why":"Hadronizes the parton-level events, producing the samples used for the detector-level comparison.","marker":"[21]"},{"why":"Simulates the detector response whose smearing is used in the performance comparison.","marker":"[22]"},{"why":"Supplies the ν-Flows architecture (coupling layers with rational quadratic splines) that the paper adopts.","marker":"[27]"},{"why":"Provides a second generator used to cross-check cross sections and kinematic shapes.","marker":"[15]"}],"fun_headline_variants":["ν-Flows cut angular error 7x in top-DM searches","ν-Flows outperform MLP for top-DM angular reconstruction","Normalizing flows beat MLPs for top-DM angular variable","ν-Flows slash top-DM angular error to 52 from 360","Top-DM angular precision: ν-Flows 7x better than MLP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the Monte Carlo sample used for training is faithful to real LHC collisions, because the observed final state does not determine the neutrino and mediator momenta uniquely, so the network can only learn the conditional distribution encoded in that simulation.","fun_headline_variants_meta":{"raw":{"variants":["ν-Flows cut angular error 7x in top-DM searches","ν-Flows outperform MLP for top-DM angular reconstruction","Normalizing flows beat MLPs for top-DM angular variable","ν-Flows slash top-DM angular error to 52 from 360","Top-DM angular precision: ν-Flows 7x better than MLP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001606,"raw_usage":{"total_tokens":6369,"prompt_tokens":893,"completion_tokens":5476,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":5377}},"tokens_in":509,"tokens_out":5476,"duration_ms":35458,"temperature":1.0,"reasoning_tokens":5377,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:51:52.153504+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train ν-Flows on events from one generator and test it on events from a second generator with the same truth labels; if the histogram MAE of the reconstructed angular variable rises to the level of the MLP baseline (around 360), the claimed performance is generator-specific. A further test on real LHC data in a Standard Model control region would reveal whether the reconstructed distribution matches the expected background shape within uncertainties.","supporting_citations":[{"cited_title":"D’Ambrosio, G","cited_arxiv_id":null,"evidence_quote":"Introduces the angular variable and provides the MLP baseline the paper compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets the benchmark model parameters (mχ = 1 GeV, mΦ = 400 GeV, couplings gf = gχ = 1) used in the study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ν-Flows architecture (coupling layers with rational quadratic splines) that the paper adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a second generator used to cross-check cross sections and kinematic shapes."}],"review_version":1}