{"id":"2dec22fb-f61d-4e2a-b98b-cb772fe71228","arxiv_id":"2505.02739","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A simulated search predicts a 6.8-sigma signal for a 500 GeV dark Higgs in the dimuon plus missing energy channel at the HL-LHC, while heavier masses are not discoverable.","lead":"This paper simulates how the High-Luminosity LHC could spot a hypothetical 'dark Higgs' particle that decays into two muons and missing energy. It uses machine learning to separate the signal from background and finds that only a 500 GeV version would be discoverable with the full dataset.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quoted z=6.77 is the maximum of a BDT-threshold scan evaluated on the same simulated events; the discovery claim is fragile until the threshold-selection bias is checked with a validation split.","rationale":"The paper is a self-contained Monte Carlo sensitivity projection, and its numerical bookkeeping is internally consistent; for example, the BP4 required luminosity for 5sigma scales closely to the quoted z(3000) about 6.7. I agree with the reader's CONDITIONAL verdict and with the caution about detector simulation fidelity. However, I would place the primary load-bearing problem one step earlier in the statistical chain. Table VIII reports z after choosing the BDT response cut that maximizes z on the same events that are then counted. Cross-validation protects the classifier training from overfitting, but it does not account for the optimization over thresholds. This is a known upward bias and it directly inflates the headline 6.77. A split-sample or Asimov test is cheap and would settle it. The paper's final limitation statement about future systematic studies partially mitigates the issue, but the abstract's discovery-potential framing does not. I therefore keep the reader's CONDITIONAL verdict: the analysis is usable as a simplified projection, but the central discovery claim should not be taken at face value until the threshold-selection bias is quantified. The proposed concrete test requires no new event generation and would either confirm or refute the concern.","tokens_in":10196,"tokens_out":13087,"duration_ms":168715,"concrete_test":"Partition the out-of-fold BDT scores for BP4 and the combined background into two independent halves. On half A, scan the BDT cut and record the threshold that maximizes z=S/sqrt(S+B); then apply exactly that threshold to half B and compute z_B. Swap halves and repeat. If either held-out z_B falls below 5, the optimized value 6.77 is inflated by threshold selection and the discovery claim should be re-stated as a projection requiring a pre-specified cut. This check uses only the already generated samples.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claim is Table VIII: BDT gives z=6.77 for BP4 and discovery for lighter benchmarks. The value is obtained by scanning BDT response thresholds and taking the cut where z is maximal for each BP (Section VI, Fig. 8). This is a selection on the same sample used to measure S and B. The k-fold CV used in training prevents overtraining of the classifier response, but it does not protect against the optimization bias of choosing the threshold that maximizes z on the out-of-fold predictions: z_max is the maximum of a noisy function over many correlated thresholds, so it is systematically above the significance that a pre-specified or independently chosen cut would yield. The effect matters because after the cut B=4951 and S=500; Poisson fluctuations at neighboring thresholds are sizeable, and the maximum will preferentially sit on an upward signal fluctuation or a downward background fluctuation. Additionally, z=S/sqrt(S+B) contains no systematic uncertainty in the background; a 10-20% change in background acceptance or normalization moves z from 6.77 to roughly 5 or below. The final section acknowledges systematics only as a future step, while the abstract and headline framing do not carry that caveat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a Monte Carlo study of the dark Higgs model in the mono-Z' portal at the HL-LHC, examining the Z'->mu+mu- plus missing transverse energy signature. Events are generated with MadGraph5_aMC@NLO and Pythia8, passed through Delphes with the CMS pile-up card, and analyzed with TMVA classifiers (BDT, DNN, Likelihood). The BDT is selected as the best classifier, and expected significances are computed for dark Higgs signals with M_Z' between 200 and 1000 GeV at sqrt(s)=14 TeV and L=3000 fb^-1. The central numerical result is a BDT-based significance of z=6.77 for M_Z'=500 GeV (BP4), with lighter benchmark points exceeding 5 sigma and heavier points falling below the discovery threshold.","tokens_in":10455,"tokens_out":5165,"duration_ms":62317,"significance":"If the central numerical claim were robust, the paper would be a useful HL-LHC projection for the dark Higgs dimuon channel, providing a concrete comparison of MVA classifiers and per-benchmark significance predictions. Its strengths include the full MadGraph+Pythia+Delphes simulation chain, five-fold cross-validation in training, an explicit hyperparameter table, and a falsifiable set of benchmark results. However, the quoted significances are computed with a simplified statistical formula that ignores systematic uncertainties, and the BDT cut is chosen to maximize the significance on the same sample used to report it. Both effects bias the headline z values upward, so the paper is best viewed as a phenomenological template that needs additional validation before its discovery claim can be taken at face value.","major_comments":[{"comment":"The 'optimal cut' on the BDT response is selected by scanning thresholds and choosing the one that maximizes z on the same out-of-fold predictions used to quote z. As a result, z=6.77 is the maximum of a noisy function over many correlated thresholds, and is systematically larger than the expected significance for a pre-specified or independently chosen cut. With S=500 and B=4951 at the quoted cut, Poisson fluctuations at neighboring thresholds are sizable, so this is a first-order effect. The authors should report significances for a cut fixed on an independent validation set, or use a three-way split (training/validation/test) and quote the test-sample significance at the validation-selected cut.","section":"Section VI, Eq. (3), Fig. 8, Table VIII"},{"comment":"All quoted significances are purely statistical, z=S/sqrt(S+B), with no systematic uncertainty assigned to signal or background rates, acceptances, or shapes. This is load-bearing: a 10-20% background normalization systematic alone moves the BP4 value from 6.8 toward the 5 sigma boundary (roughly 6.5 at 10% and 5.8 at 20%), and the BP5-BP9 values are already below 3 sigma. The paper should either include a nuisance-parameter treatment with correlated systematic uncertainties or clearly label every number as a statistical-only projection, with a matching caveat in the abstract and summary. The final-section sentence deferring systematics to 'future studies' does not adequately qualify the headline discovery claim.","section":"Eq. (3) and Section VII"},{"comment":"The absolute values of S and B, and therefore the significances, are obtained from Delphes simulation using the CMS Pile-Up configuration card. The paper does not cross-check the resulting muon reconstruction, isolation, and missing transverse energy response against public HL-LHC performance studies or against a full simulation sample, even for one benchmark point. Because the acceptance numbers directly determine the central discovery claim, this fast-simulation assumption should either be validated or assigned a quantitative uncertainty; otherwise the absolute z values inherit an unquantified detector-modeling risk.","section":"Sections IV.A, IV.B and Table VIII"}],"minor_comments":[{"comment":"There are small language slips: 'tabled I' should be 'Table I', and the Table II caption uses 'Pb' where 'pb' is meant.","section":"Section IV.A and Table II caption"},{"comment":"The correlation matrix axis labels in Fig. 4 are nearly unreadable due to overlapping text, and the statement that 'most of them are highly uncorrelated' is difficult to reconcile with entries such as the 77-85% correlations visible in the figure. Please clarify the criterion for 'highly uncorrelated' and improve the figure labels.","section":"Fig. 4 and surrounding text"},{"comment":"The equation for the separation power is typeset confusingly, with the integrand appearing as 'Z ( ˆPs(γ)− ˆPb(γ))2'. Please use a standard definition, e.g., 1/2 * integral((Ps-Pb)^2/(Ps+Pb)) dγ, and define all symbols.","section":"Eq. (2)"},{"comment":"The DNN layout string 'TANH|128, TANH|128, TANH|128, LINEAR' is not self-explanatory; please specify the layer sizes and activation functions explicitly in the table or in a footnote.","section":"Table VI"},{"comment":"The symbol 'NBC' is not defined; please state in the caption that it denotes the number of events after pre-selection but before the BDT cut.","section":"Table VIII"},{"comment":"The sentence 'This indicates good training performance for all three classifiers, as expected due to the use of CV' is imprecise: cross-validation reduces overfitting but does not by itself guarantee that training and test distributions match. Please report the KS test p-values used for the overtraining check.","section":"Section VI, overtraining discussion"}],"recommendation":"major_revision","confidential_remarks":"The paper's central discovery claim is likely inflated by the combination of threshold scanning on the same sample and the absence of systematic uncertainties. The underlying simulation and classifier comparison are standard and coherent, and the identified issues are fixable within the scope of a revision, so I recommend major revision rather than rejection. The authors should also consider de-emphasizing the DNN comparison, which performs worse than the likelihood estimator, or explain the hyperparameter choices that led to that outcome."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a straightforward, well-documented sensitivity projection for the existing dark Higgs / mono-Z' model, and the central number — BDT significance z=6.77 at M_Z'=500 GeV, 3000 fb^-1 — is likely a bit too good because the cut that maximizes z is chosen on the same sample used to quote S and B. The paper is still worth a serious referee, but the discovery framing should be softened pending a fix.\n\nThe genuinely new part is the HL-LHC reach scan: required luminosities for 5-sigma and expected z values for BP4-BP9. The simulation chain is standard and reproducible: MadGraph5_aMC@NLO, Pythia8, Delphes with the CMS pile-up card, TMVA with BDT/DNN/Likelihood, k-fold CV, KS tests, ROC curves. The classifier comparison is competently done, and the background list is sensible. Credit where due: the paper does not oversell the model; it inherits couplings and benchmarks from prior work, which is normal and properly cited.\n\nNow the soft spots, in proportion. The big one is threshold selection bias. The text says (Section VI, Fig. 8, Table VIII) that for each BP they apply the BDT response cut where z is maximum, then report that z. That is maximizing a noisy statistic over many correlated thresholds. The k-fold CV prevents overtraining of the classifier response, but it does not protect against optimizing the cut on the out-of-fold predictions. For BP4, S=500 and B=4951, so Poisson fluctuations are sizeable; the true expected significance at a pre-specified or independently chosen cut will be lower, possibly below 5σ. This is a load-bearing issue for the headline claim.\n\nSecond, z = S/sqrt(S+B) is purely statistical. A 10-20% systematic shift in background acceptance or normalization brings z=6.77 down to roughly 5 or lower. The paper acknowledges systematics only as a future step, which is honest as a limitation, but the abstract and conclusion say “strong discovery potential” without that caveat.\n\nThe minor issues are just that: the DNN with its chosen hyperparameters underperforms BDT and Likelihood, which is plausible but not deeply investigated; the correlation matrices show some correlated variables used anyway, but the importance ranking justifies keeping them.\n\nVerdict: The paper is a solid, clearly written Monte Carlo sensitivity study with one methodological flaw in the central number. It deserves peer review, not a desk rejection. A referee should ask for a validation split or a pre-specified cut to remove the optimization bias, and a systematic uncertainty treatment or at least a sensitivity scan over background normalization. The discovery claim should be rephrased as a preliminary projection until then.","headline":"A clearly documented HL-LHC sensitivity projection for a known dark Higgs model, but the headline z=6.77 is inflated by optimizing the BDT cut on the same events; worth refereeing with a request for a validation split and systematics.","tokens_in":10977,"tokens_out":2357,"would_cite":false,"duration_ms":30825,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A boosted-decision-tree analysis of the dark Higgs Z' dimuon channel finds a statistical significance of 6.77σ at 500 GeV with the full HL-LHC dataset.","keywords":["dark matter","dark Higgs model","Z' boson","dimuon final state","HL-LHC","multivariate analysis","boosted decision trees","missing transverse energy"],"falsifier":"Apply the same BDT cut to the first 3000 $fb^{{-1}}$ of HL-LHC data and fit the dimuon invariant-mass spectrum around 500 GeV in events with large missing transverse energy; if no excess over the Standard Model prediction appears, the claimed $S=500$ events and $z=6.77$ are refuted.","tokens_in":9994,"feed_emoji":"⚛️","tokens_out":13840,"duration_ms":151307,"temperature":0.7,"pith_summary":"The paper asks whether the High-Luminosity LHC, with 3000 $fb^{{-1}}$ at 14 TeV, could discover a dark Higgs model in which a Z' boson decays to a dimuon pair plus missing transverse energy from dark matter. Using a boosted decision tree trained on five kinematic variables, it finds a statistical significance of $z = 6.77$ at a Z' mass of 500 GeV, above the 5σ discovery threshold. It also reports that lighter benchmark points reach 5σ with far less luminosity, while Z' masses from 600 to 1000 GeV stay below 5σ even with the full dataset. A sympathetic reader would take this as evidence that the dimuon-plus-missing-energy channel is a promising early target for the HL-LHC at moderate Z' masses.","feed_headline":"Dark Higgs Z' dimuon search hits 6.77σ at the HL-LHC","feed_subtitle":"A boosted-decision-tree analysis reaches discovery significance at 500 GeV; heavier masses stay below 5σ with the full HL-LHC dataset","key_machinery":"The central mechanism is the trained boosted decision tree (BDT), an ensemble of shallow decision trees that classifies events as signal-like or background-like using five kinematic inputs: the transverse momenta of the two muons, the missing transverse energy, the dimuon invariant mass, and the azimuthal angle between the dimuon system and the missing energy. The analysis scans the BDT response threshold for each benchmark point and takes the cut that maximizes $z = S/\\sqrt{S+B}$, where $S$ and $B$ are the expected signal and background counts after the cut. This classifier, trained with five-fold cross-validation, is shown to achieve the largest area under the ROC curve among the three methods compared, and it is the BDT results that produce the quoted significance values.","core_discovery":"The paper's central claim is that the dark Higgs benchmark point BP4, with $M_{Z'}=500$ GeV and $g_{SM}=0.25$, $g_{DM}=1.0$, would be discovered in the dimuon plus missing transverse energy channel at the HL-LHC: after an optimal cut on the boosted decision tree response, the expected signal is $S = 500$ events against a background of $B = 4951$ events, giving $z = S/\\sqrt{S+B} = 6.77$ with 3000 $fb^{{-1}}$ at 14 TeV. For BP1 through BP3, the paper reports 5σ discovery at integrated luminosities of 4.3, 18.8, and 175 $fb^{{-1}}$ respectively. For BP5 through BP9, masses 600 to 1000 GeV, the expected significance falls from 2.81 to 0.48, below the discovery threshold. The BDT classifier outperforms the deep neural network and the likelihood estimator in the comparison, so the quoted sensitivities are specific to the boosted decision tree.","pith_inferences":["A direct extension would replace the counting significance $z = S/\\sqrt{S+B}$ with a profile-likelihood fit that includes systematic uncertainties on the background shape and acceptance; that would show how much of the 6.77σ survives real-world effects.","The same BDT recipe, retrained on the dielectron or other leptonic channels, could extend the search and cross-check the dark Higgs interpretation without requiring a new model setup.","Because BP1 through BP3 need less than 500 fb^{-1} for discovery, early HL-LHC data could already test this model, and the analysis could be re-optimized on the first year of collisions."],"forward_implications":["At $M_{Z'}=500$ GeV, the dark Higgs benchmark BP4 yields a 6.77σ excess in the dimuon plus missing-energy channel with 3000 fb^{-1}, clearing the discovery threshold.","The lighter benchmarks BP1 through BP3 reach 5σ with only 4.3, 18.8, and 175 fb^{-1} respectively, so the channel is sensitive at modest integrated luminosity for low Z' masses.","For Z' masses between 600 and 1000 GeV, the expected significance stays between 2.81 and 0.48, so the full HL-LHC dataset is not enough to discover the model in this channel.","Since the BDT outperforms the DNN and likelihood classifiers on this final state, the reported sensitivities are tied to using the boosted decision tree; the other classifiers would give weaker results."],"supporting_citations":[{"why":"This reference defines the dark Higgs simplified model and the heavy dark-sector mass assumption that the search targets.","marker":"[26]"},{"why":"This previous search for a leptonically decaying Z' with missing transverse energy motivates the chosen couplings and mass range.","marker":"[28]"},{"why":"This reference supplies the next-to-leading-order event generation and the cross-section values used for signal and background yields.","marker":"[44]"},{"why":"This fast detector simulation converts generated events into reconstructed muons, isolation, and missing transverse energy, setting the acceptances.","marker":"[46]"},{"why":"This reference introduces the boosted decision tree algorithm that the paper selects as its best classifier.","marker":"[38]"},{"why":"This is the multivariate-analysis package used to train, test, and compare the BDT, DNN, and likelihood classifiers.","marker":"[49]"},{"why":"This reference provides the significance formula $z = S/\\sqrt{S+B}$ used to compute the quoted discovery significances.","marker":"[56]"}],"fun_headline_variants":["500 GeV dark Higgs Z' hits 6.77σ at HL-LHC dimuon","BDT spots dark Higgs Z' at 6.77σ in HL-LHC dimuon data","Dark Higgs Z' discovery at 6.77σ with BDT in HL-LHC dimuon","6.77σ dark Higgs Z' signal in HL-LHC dimuon with BDT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result assumes the fast detector simulation used for both signal and background reproduces the real HL-LHC's muon reconstruction, isolation, and missing-energy response; if real detectors are less efficient at high pileup, the predicted significance would shrink.","fun_headline_variants_meta":{"raw":{"variants":["500 GeV dark Higgs Z' hits 6.77σ at HL-LHC dimuon","BDT spots dark Higgs Z' at 6.77σ in HL-LHC dimuon data","Dark Higgs Z' discovery at 6.77σ with BDT in HL-LHC dimuon","6.77σ dark Higgs Z' signal in HL-LHC dimuon with BDT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001059,"raw_usage":{"total_tokens":4454,"prompt_tokens":967,"completion_tokens":3487,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":3384}},"tokens_in":583,"tokens_out":3487,"duration_ms":26783,"temperature":1.0,"reasoning_tokens":3384,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:42:26.063439+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same BDT cut to the first 3000 $fb^{{-1}}$ of HL-LHC data and fit the dimuon invariant-mass spectrum around 500 GeV in events with large missing transverse energy; if no excess over the Standard Model prediction appears, the claimed $S=500$ events and $z=6.77$ are refuted.","supporting_citations":[{"cited_title":"ch/record/2857116","cited_arxiv_id":null,"evidence_quote":"This previous search for a leptonically decaying Z' with missing transverse energy motivates the chosen couplings and mass range."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This reference supplies the next-to-leading-order event generation and the cross-section values used for signal and background yields."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This reference introduces the boosted decision tree algorithm that the paper selects as its best classifier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This reference provides the significance formula $z = S/\\sqrt{S+B}$ used to compute the quoted discovery significances."}],"review_version":1}