{"id":"6a88ebda-9620-4dde-89ce-bc7bbb210ef1","arxiv_id":"2607.24524","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"LDMX Phase II is projected to discover R=2.5 thermal complex-scalar DM at 5σ and to discriminate dark-photon electromagnetic-moment models via 2D Bayes factors.","lead":"A statistical framework projects that LDMX Phase II can discover thermal light dark matter at 5σ for stronger relic-target benchmarks and can tell competing dark-photon interaction models apart from recoil kinematics. It matters because it turns a proposed missing-momentum experiment into a tool that both finds and characterizes a dark sector if a signal appears.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The reader's concern is the right one; I sharpen it: the outer-product background (§2.2) populates kinematically forbidden/poorly motivated regions of the (E, |p_T|) plane, so the background under the signal peak is likely mis-shaped, not just mis-normalized.","rationale":"The reader identified the correct load-bearing assumption and I confirmed no stronger objection exists. I checked the plausible alternatives: (1) The use of asymptotic approximations for the \"≫5σ\" entries (§4.2) despite the paper's own demonstration that asymptotics fail for f(q0|0) — this is not load-bearing because with S = 250–310 signal events against B ≤ 100, discovery is overwhelming under any test-statistic distribution; the exact tail shape of f(q0|0) is irrelevant at that separation. (2) Prior sensitivity of Bayes factors (log-uniform priors over ε ∈ [10^-20, 1], and the ad hoc relic-density prior in §3.2.1/Fig. 13 which even switches to R = 3 benchmarks inconsistent with the rest of the paper) — real, but the paper is transparent about the prior construction and the qualitative 2D-vs-1D gain is prior-robust. (3) The benchmark-4 contour-averaging artifact (§4.3) — acknowledged by the authors themselves. (4) The MadGraph Monte Carlo fluctuation caveat on (M,E)/(C,A) near-zero Bayes factors — also acknowledged. So the background factorization stands as the single concern whose resolution could move the model-selection and borderline-discovery conclusions. It does not, however, undermine the paper's core contribution (the framework and the robust R = 2.5 discovery claim), and the authors disclose it explicitly. Hence UNCHANGED at CONDITIONAL: the condition should be phrased as \"headline model-selection and R = 2.2 results require a correlated 2D Phase-II background model before being quoted as projections,\" which matches the reader's CONDITIONAL with medium correctness risk.","tokens_in":28712,"tokens_out":2484,"duration_ms":93326,"concrete_test":"Replace the outer-product background with a correlated 2D background: either digitize a full 2D (E, pT) photo-nuclear distribution from an LDMX GEANT4 simulation (as in [32]) or, cheaply, reweight the existing outer product by the kinematic/physical correlation (enforce pT/E consistent with the emission-angle distribution of photo-nuclear recoils, after the 40° cut). Renormalize to N_BG = 100 and recompute (a) Z at benchmark 3 and (b) the Fig. 11 heatmap. If benchmark 3 stays >10σ and the ln K ordering/magnitudes shift by <25%, the central claims are robust to the factorization assumption; if any off-diagonal ln K crosses the ln K ≳ 10 decisiveness threshold in either direction, the model-selection claim is background-model dependent and must be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I agree with the reader that §2.2's background construction is the soft spot, and I would sharpen it beyond \"E and |p_T| may be correlated.\" Building the 2D background as the outer product of a |p_T| marginal (Fig. 6 of [32]) and an inclusive E marginal (Fig. 10 of [27]) necessarily places background probability in regions of the (E, |p_T|) plane that a real single recoil electron cannot occupy or occupies with very different weight — most starkly, bins with |p_T| comparable to or exceeding E (the 40° acceptance cut removes only part of this region, and the paper does not state whether the cut is imposed on the background marginals before or after the outer product). For a real photo-nuclear recoil, E and |p_T| are linked by the emission angle; the physical distribution is concentrated along a correlated locus, while the outer product spreads background events across the whole rectangle. The sign of the resulting bias is not obvious a priori: spreading background into unphysical bins dilutes it in the physical low-E/high-|p_T| signal region (flattering discovery significance), but it also mis-shapes the background exactly where the 2D shape discrimination and the Bayes-factor model comparison draw their power. The two headline claims are differently exposed: the 5σ discovery of benchmarks 1 and 3 (S = 310 and 250 vs B ≤ 100 spread over 2500 bins) is almost certainly robust — even a factor-of-several local background underestimate leaves S/B comfortably large, and the \"≫5σ\" values quoted from asymptotics would survive any plausible re-shaping. The genuinely load-bearing exposure is in (a) the Bayes-factor heatmaps (Fig. 11), where model discrimination leans on 2D shape differences between signal hypotheses against a background template whose own 2D shape is assumed, and (b) the R = 2.2 borderline results (Z = 4.4/4.2 at N_BG = 1) where a modest background re-concentration moves the answer across 5σ. The authors flag this honestly in footnote 4, so this is a known-ack","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript builds a binned two-dimensional (recoil energy, transverse momentum) Poisson likelihood framework for LDMX Phase II (10^16 EOT), with log-normal normalization nuisances, and applies it to four benchmark points on the thermal relic target for complex scalar DM with a dark-photon mediator (R = 2.5 and 2.2, m_A' = 0.01 and 0.1 GeV), under two background normalizations (N_BG = 1 and 100). It reports projected 90% C.L. exclusion contours, median discovery significances from full toy pseudo-experiments (arguing correctly that Wilks asymptotics fail at these counts), frequentist and Bayesian parameter estimation for (m_A', epsilon), and Bayes-factor model comparison among kinetic mixing and four dark-electromagnetic-moment interaction models, finding that the R=2.5 benchmarks are discoverable at ≫5σ in both background scenarios and that the 2D analysis gives roughly an order of magnitude larger ln K than an energy-only analysis.","tokens_in":29109,"tokens_out":6227,"duration_ms":216711,"significance":"If the results hold, this is a useful and timely contribution for the light-DM community: it moves beyond counting-only sensitivity estimates to a complete statistical pipeline for LDMX, with concrete, falsifiable projections tied to stated relic-target benchmarks (R=2.5/2.2 at m_A' = 0.01, 0.1 GeV) and to two bracketing background scenarios. Particular strengths worth naming: (i) the explicit demonstration, via large toy ensembles (Appendix B), that the half-chi-square asymptotics fail at LDMX bin counts, with all near-threshold p-values therefore derived from pseudo-experiments; (ii) consistent frequentist (profile likelihood, CLs variant) and Bayesian (dynamic nested sampling, evidences) treatments with log-normal systematics and a hierarchical background-width hyperprior; (iii) the model-selection study showing the (E, |p_T|) analysis gains roughly an order of magnitude in ln K over energy-only, and the instructive misspecification-bias matrix of Fig. 12. The finding that benchmarks 1 and 3 are discoverable at ≫5σ in both background scenarios is likely robust; the R=2.2 threshold results are honestly presented.","major_comments":[{"comment":"The two-dimensional background is constructed as the outer product of a 1D |p_T| photo-nuclear shape (Fig. 6 of [32]) and a 1D inclusive E distribution (Fig. 10 of [27]), assuming E and |p_T| are uncorrelated. This is acknowledged in footnote 4 as an approximation, but two specific problems deserve more than a footnote. (i) The outer product populates regions of the (E, |p_T|) plane that a single recoil electron cannot occupy — most obviously bins with |p_T| approaching or exceeding E, which violate the kinematic bound |p_T| <= E and the stated 40-degree acceptance cut sin^{-1}(|p_T|/E) < 40 deg. The manuscript does not state whether the 40-degree cut is imposed on the marginals before forming the outer product, on the product afterwards, or at all for the background; if applied after, the effective background normalization in the physical region is smaller than the quoted N_BG, and the","section":"§2.2 (Background modelling) and Fig. 4"},{"comment":"The relic-density-prior model discrimination, which the Conclusion highlights ('incorporating a relic density prior ... rendering some previously indistinguishable signal hypotheses statistically distinguishable'), is computed at R = 3 — the mass ratio the Introduction and §2.1 argue is now excluded by DAMIC-M/PandaX-4T electron-recoil constraints, which is the paper's stated motivation for adopting R = 2.5 and 2.2 benchmarks. The justification given (larger signal yields at R = 3) undercuts the result's relevance: the claim of practical interest is whether KM can be distinguished from C/A at the surviving benchmarks. Since at R = 2.5 the yields remain large (Table 4: S = 250 for KM, 5300 for C/A), the exercise should be repeatable at benchmark 3 with the R = 2.5 relic targets; the authors should either present that or clearly motivate why the R = 3 result is the one to quote. A secondar","section":"§5.2, Fig. 13"},{"comment":"Two aspects of the low-yield parameter inference need tightening. (i) The 1-sigma and 2-sigma confidence regions in Fig. 6 and the half-interval widths in Table 2 are constructed under the asymptotic chi-square approximation with k = 2, yet Appendix B demonstrates that the asymptotic distributions fail in precisely this low-count regime, and for benchmarks 2 and 4 (S = 4 and 2) the approximation is not credible. The text concedes the resulting contours can be artifacts of averaging over 25 pseudo-experiments (the 'deceptively tight' BM4 contour), but Table 2 still reports single numbers without the ensemble variance, which the text says is 1-2 orders of magnitude larger than for benchmarks 1 and 3. The variance (or the per-realization interval coverage) should be reported in Table 2, or the BM2/BM4 entries marked as uninformative. (ii) Relatedly, the abstract states that 'both parameters","section":"§4.3, Table 2, and abstract/§6 wording"}],"minor_comments":[{"comment":"§3.2.1 states the median Bayes factor is taken over five pseudo-experiments, while §5.2 and the Fig. 11 caption state ten pseudo-data realizations. Please reconcile.","section":"§3.2.1 vs §5.2"},{"comment":"§3.2.1 requires 'log K ≳ 10' as a stricter criterion than the Jeffreys K > 100 (ln K = 4.6); since Fig. 11 is labelled in ln K, clarify the logarithm base throughout and state the threshold consistently.","section":"§3.2.1"},{"comment":"For benchmarks 1 and 3 the discovery significance is computed with the asymptotic formula of [11] and quoted as '≫5', which is appropriate; however, Appendix B (Fig. 15) displays specific asymptotic values such as Z_as = 58 sigma against toy distributions that the same figure shows are not asymptotic. Either remove the numerical Z_as labels or annotate them as unreliable by the paper's own demonstration.","section":"§4.2 and Appendix B"},{"comment":"The Bayes factors are prior-dependent through the evidence; the log-uniform coupling prior spans epsilon in [10^{-20}, 1], i.e. 21 decades. A brief sensitivity statement (or a justification that the Occam factors cancel in the model pairs considered) would strengthen §5.2.","section":"§3.2 / §5.2"},{"comment":"Fig. 4 illustrates signal and background at m_A' = 1 GeV, epsilon = 10^{-3}, which is not one of the four benchmarks; please note in the caption that the point is illustrative.","section":"Fig. 4"},{"comment":"Typos and wording: 'couple strength' (abstract, should be 'coupling strength'); 'the explore the effect' (§2.2, should be 'to explore'); 'not longer fixed' (§3.1.3, should be 'no longer fixed'); 'Baye's factors' (§5.2); stray spaces in section headings ('F requentist', 'F ormalism', 'T est Statistic').","section":"Various"},{"comment":"Table 2: for BM2 the quoted epsilon bias improves from 0.3 (N_BG=1) to 0.18 (N_BG=100); given the authors' own discussion of averaging artifacts for BM4, it would be worth noting whether this trend is physically meaningful or the same artifact.","section":"§4.3, Table 2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is straightforward: under their Phase-II setup they get robust projected discovery on the still-viable R=2.5 thermal points (benchmarks 1 and 3), recover m_A' and ε when the yield is large, and show that 2D (E,|p_T|) Bayes factors separate KM from M/E from C/A much better than energy-only. That is the right question after DAMIC-M/PandaX pushed R≳3 off the table, and they answer it with a full likelihood rather than a cut-and-count sketch.\n\nWhat is actually new is the combined frequentist+Bayesian 2D pipeline on those post-2025 benchmarks, with large toy ensembles where they correctly refuse Wilks (Appendix B), plus the DEM model-comparison matrices with and without a relic prior. The ingredients are standard—MadGraph dark bremsstrahlung, Cowan-style profile likelihood, dynesty, operators from their earlier work—but the assembly and the numbers are new and usable. They are honest about systematics (log-normal θ_S, θ_B), run both N_BG=1 and 100, and flag the background construction in footnote 4.\n\nThe soft spot is real and is exactly the one the stress-test sharpens. Building the 2D background as the outer product of a published |p_T| shape and an inclusive E shape spreads probability into kinematically awkward corners of the plane. That almost certainly does not kill the ≫5σ claims on benchmarks 1 and 3 (S~250–310 vs B≤100 over 2500 bins). It does matter for the Bayes-factor heatmaps, which lean on 2D shape against that template, and for the R=2.2 points sitting at Z~4. It is a modeling limitation, not a math error, and they own it.\n\nMinor notes: averaging profile-likelihood contours over 25 toys can look tighter than any single realization at low S; CLs is only a ~1.2 factor; no public toys/code, so you cannot re-run. Citation pattern is normal for the subfield.\n\nThis is for people building or reviewing LDMX analyses and for anyone who needs quantitative reach on the surviving thermal strip. It deserves a serious referee. I would engage with it and cite the discovery/exclusion tables and the 2D-vs-1D Bayes comparison when discussing LDMX sensitivity.","headline":"Solid LDMX stats pipeline with real 5σ reach on R=2.5; the outer-product background is the main soft spot, mostly for Bayes factors and the R=2.2 edge cases.","tokens_in":29672,"tokens_out":622,"would_cite":true,"duration_ms":21523,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LDMX Phase II can discover light dark matter on the R=2.5 thermal target and tell dark-photon models apart from the 2D recoil spectrum.","keywords":["light dark matter","LDMX","dark photon","thermal relic target","missing momentum","Bayes factor","parameter estimation","higher electromagnetic moments"],"falsifier":"Apply the identical likelihood pipeline to real LDMX Phase-II data (or a high-fidelity full-detector simulation with correlated backgrounds) at the R=2.5 benchmarks and check whether the measured discovery significance, parameter contours, and Bayes factors match the projections.","tokens_in":29222,"feed_emoji":"⚛️","tokens_out":945,"duration_ms":18708,"temperature":0.7,"pith_summary":"This paper asks whether the planned Light Dark Matter eXperiment can not only find light dark matter but also measure its properties and decide which dark-photon interaction is at work. Using a full likelihood on the two-dimensional recoil-electron energy and transverse-momentum distribution, with signal and background systematics, the authors project that Phase II (10^16 electrons on target) reaches 5σ discovery for complex scalar dark matter on the R=2.5 thermal relic curve under both low and high background assumptions, while the harder R=2.2 curve sits just below that threshold. When a signal is present at those strong benchmarks, both the dark-photon mass and the kinetic mixing are recovered inside their uncertainties by frequentist and Bayesian fits. Bayes factors then separate kinetic mixing from magnetic/electric dipole and charge-radius/anapole hypotheses, and the joint (E, p_T) analysis yields roughly an order of magnitude more discriminating power than energy alone. The same machinery is written so it can be dropped straight onto real LDMX data.","feed_headline":"LDMX can discover light DM and tell dark-photon models apart","feed_subtitle":"Phase II projections show 5σ reach on the R=2.5 relic target and strong 2D model separation","key_machinery":"A binned two-dimensional Poisson likelihood in recoil energy and transverse momentum, with log-normal nuisance parameters on signal and background normalizations, evaluated by full toy pseudo-experiments (frequentist) and dynamic nested sampling (Bayesian evidence and posteriors).","core_discovery":"At four thermal-relic benchmarks for complex scalar dark matter mediated by a dark photon, LDMX Phase II has projected 5σ discovery reach on the R=2.5 targets across N_BG=1 and 100, recovers both m_A' and ε within uncertainties there, and can statistically distinguish kinetic-mixing from higher electromagnetic-moment dark-photon models via Bayes factors, with the two-dimensional recoil analysis far stronger than energy-only.","pith_inferences":["If real photo-nuclear backgrounds prove strongly E–p_T correlated, the quoted Bayes-factor gains from the two-dimensional analysis will shrink and may need re-optimization of the binning.","The hierarchy of model-separation power (2D ≫ energy-only > p_T-only) suggests that any future trigger or analysis cut that discards transverse-momentum information will measurably weaken dark-sector model selection.","Joint likelihoods with direct-detection or other accelerator data sets become straightforward once this LDMX likelihood is public, potentially closing the remaining R=2.2 window faster than either probe alone."],"forward_implications":["A null Phase-II result excludes the R≳2.2 thermal targets over roughly 1 MeV–0.2 GeV in dark-photon mass.","An excess at the strong benchmarks yields ~10% relative uncertainties on both mediator mass and coupling.","The same framework can be reused for other missing-momentum experiments and other electron-coupled dark-sector models.","Adding a relic-density prior on the coupling lifts remaining shape degeneracies between kinetic mixing and charge-radius/anapole models."],"fun_headline_variants":["LDMX Phase II: 5σ light DM reach on R=2.5 relic targets","2D recoils let LDMX separate dark-photon models via Bayes factors","LDMX recovers m_A' and ε on thermal targets, distinguishes DM hypotheses","Projected 5σ discovery and model selection for dark-photon light DM","LDMX 2D analysis outperforms 1D for dark-sector hypothesis tests"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The background shape is built by multiplying independent one-dimensional energy and transverse-momentum distributions and normalizing them to two ad-hoc total counts rather than a full correlated Phase-II background model.","fun_headline_variants_meta":{"raw":{"variants":["LDMX Phase II: 5σ light DM reach on R=2.5 relic targets","2D recoils let LDMX separate dark-photon models via Bayes factors","LDMX recovers m_A' and ε on thermal targets, distinguishes DM hypotheses","Projected 5σ discovery and model selection for dark-photon light DM","LDMX 2D analysis outperforms 1D for dark-sector hypothesis tests"]},"model":"grok-4.5","effort":"low","cost_usd":0.002003,"raw_usage":{"total_tokens":985,"prompt_tokens":868,"num_sources_used":0,"completion_tokens":97,"cost_in_usd_ticks":20028000,"prompt_tokens_details":{"text_tokens":868,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":868,"tokens_out":97,"duration_ms":2662,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T12:25:15.817938+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Apply the identical likelihood pipeline to real LDMX Phase-II data (or a high-fidelity full-detector simulation with correlated backgrounds) at the R=2.5 benchmarks and check whether the measured discovery significance, parameter contours, and Bayes factors match the projections.","supporting_citations":[],"review_version":1}