{"id":"1cbc8f0d-a630-4b46-b81b-8ed8b4e664e5","arxiv_id":"2511.08359","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Simulated future-collider studies project that FCC-ee and FCC-hh could constrain CP-violating Higgs EFT couplings 10-100 times more tightly than the HL-LHC, with the proton-proton machine strongest overall.","lead":"This paper simulates how precisely future particle colliders could detect CP-violating couplings of the Higgs boson, using Monte Carlo events and machine-learned observables. It concludes that FCC-ee and especially FCC-hh would beat the HL-LHC by factors of roughly 10-100 in the most sensitive channels — a direct input to the current European collider roadmap debate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative FCC-hh gains rest on a likelihood with no systematic uncertainties; at FCC-hh statistics are so large that unmodeled shape systematics, not statistics, set the limit, so the advertised 'order of magnitude' factors are not yet supported.","rationale":"I agree with the reader that the most load-bearing weakness is the absence of a systematic-uncertainty budget, but I sharpen it: the order-of-magnitude claim comes from FCC-hh channels where statistical errors are tiny and shape systematics dominate, not from the e+e- channels. The reader's ISR concern is genuine but secondary: even if the FCC-ee ML gains were reduced, the headline 'order of magnitude' is carried by Tab. 5 and Tab. 6. The 100 TeV vs 84 TeV baseline is a real quantitative caveat, but the authors flag it and it would likely degrade factors by at most a few tens of percent, not eliminate the qualitative conclusion. The paper deserves credit for validating LHC yields against ATLAS measurements and for matching existing LHC CP studies; the qualitative ranking (future machines beat HL-LHC; FCC-hh is the strongest) is robust. However, the precise factors of 20 and 23 are statistical extrapolations with no demonstrated systematic floor, and because the central abstract asserts an order-of-magnitude improvement in quantitative terms, this is the single most load-bearing concern. The proposed concrete test directly probes whether a small, plausible systematic uncertainty changes the FCC-hh intervals enough to weaken the headline claim. The reader's CONDITIONAL verdict remains appropriate; no stronger action is warranted without further evidence, but the caveat should be stated prominently.","tokens_in":18751,"tokens_out":6458,"duration_ms":74156,"concrete_test":"Recompute the Tab. 5 FCC-hh H->4l limit after adding to Eq. (4.1) a single profiled nuisance parameter per Phi_4l-m_12 bin representing a 0.5% Gaussian uncertainty on the background-only expectation in that bin (or, alternatively, a 0.5% correlated asymmetry between Phi_4l>0 and Phi_4l<0 bins). If the resulting c_PhiB/Lambda^2 interval widens by more than ~50%, the quoted 20x improvement over HL-LHC is not robust; if it changes by less than ~30%, the no-systematic assumption is adequate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is anchored by Tab. 5 (FCC-hh H->4l: c_PhiB/Lambda^2 ~ +/-0.007, ~20x better than HL-LHC) and Tab. 6 (WBF H->tau tau: c_PhiW/Lambda^2 ~ +/-0.007, ~23x better). Both are produced using the Sec. 4 binned Poisson likelihood, Eq. (4.1), with no nuisance parameters. The paper states in Sec. 4: 'Systematic uncertainties are not explicitly included in the likelihood function... their impact is expected to be small [28].' At FCC-hh, yields are enormous: 364k signal / 89k background in H->4l, and 862k signal / 16M background in WBF. The statistical error in finely binned distributions is therefore sub-percent, so the quoted intervals are set by the detailed shape of Phi_4l vs m_12 or the NN output. Any unmodeled bin-wise distortion of these shapes -- lepton charge/efficiency asymmetries, tau energy scale, background normalization correlated with the CP-odd angle, or detector asymmetries -- is not averaged away by large statistics and will shift the fitted Wilson coefficients. The 'expected small' justification is the same group's ref. [28], not an independent detector-level estimate. The Sec. 5.2 footnote also admits an unquantified factor-of-two interference effect, illustrating that shape-level uncertainties can materially change these projections. The 100 TeV baseline (Sec. 2 footnote) is acknowledged but not propagated into the quoted factors. These issues do not overturn the qualitative ordering, but they directly threaten the quantitative 'order of magnitude' headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents projected sensitivities to CP-violating dimension-six operators in the gauge-Higgs sector at HL-LHC, FCC-ee, LCF, and FCC-hh. It uses CP-odd observables (Delta phi_ll, Phi_4l, Delta phi_jj) and ML-based observables trained on constructive/destructive interference samples, with the binned Poisson likelihood of Eq. (4.1). The central claim is that future colliders, especially FCC-hh, will improve sensitivity by roughly an order of magnitude over the HL-LHC: Tab. 5 gives c_PhiB/Lambda^2 = +/-0.007 TeV^-2 and c_PhiWB/Lambda^2 = +/-0.015 TeV^-2 for H->4l at FCC-hh (~20x better than HL-LHC), and Tab. 6 gives c_PhiW/Lambda^2 = +/-0.007 TeV^-2 for WBF H->tautau (~23x better). For e+e- colliders, ML observables are reported to improve limits by a factor of 4-5, driven by a sign-flip of the interference near the Z pole in the electron channel (Fig. 1(d)).","tokens_in":18966,"tokens_out":5296,"duration_ms":58671,"significance":"If the quantitative projections are reliable, the paper provides a valuable input to the upcoming European Strategy update: it makes a concrete case that FCC-hh, despite a busier background environment, could be the strongest single facility for Higgs CP studies, and it demonstrates the practical benefit of ML-optimised observables. The simulation pipeline is carefully cross-checked against ATLAS yields in Secs. 6-8 (79 H->4l events at Run-II, WBF yields, Z+jets scalings), and the paper is transparent about several limitations (100 TeV baseline, no ISR, factor-of-two interference effects). The methodology is standard and the claimed qualitative ordering of facilities is robust; the main weakness is that the advertised numerical factors are statistical-only projections with no systematic uncertainty budget.","major_comments":[{"comment":"The central 'order of magnitude improvement' claim is produced by a likelihood that contains no systematic uncertainties. At FCC-hh the yields are enormous (Sec. 7: 364k signal / 89k background; Sec. 8: 862k signal / 16M background), so the quoted 95% intervals are set by detailed bin-by-bin shapes of Phi_4l vs m_12 and O_NN, not by counting statistics. The statement that systematics 'are expected to be small [28]' cites the same group's earlier paper and is not an independent detector-level estimate. The authors should either introduce nuisance parameters/binned shape distortions to quantify the effect, or explicitly soften the quantitative factors in the abstract and conclusions.","section":"Sec. 4, Eq. (4.1), Tables 5-6"},{"comment":"The e+e- samples are generated without initial-state radiation and corrected only by a flat k_isr factor. However, the main ML gain (factor 4-5 in Tab. 1) is driven by the sign-flip of the interference as a function of m_ee near the Z pole (Fig. 1(d)). A normalization-only k-factor cannot correct a shape distortion of this kind. In addition, the Sec. 5.2 footnote reports an unquantified factor-of-two change in the interference amplitude when using a dedicated llbb sample. These shape-level uncertainties need to be propagated (e.g., by varying the ISR spectrum or by quoting limits with and without the sign-flip feature) before the FCC-ee/LCF numbers in Tables 1-3 can be taken at face value.","section":"Sec. 5, first paragraph, and Fig. 1(d)"},{"comment":"The FCC-hh projections are computed at sqrt(s)=100 TeV, while the current FCC feasibility-study baseline is 84 TeV. The footnote acknowledges a 'slight reduction in sensitivity' but does not quantify it. Since the headline factors in Tables 5-6 (20x and 23x) are the evidence for the abstract's central claim, the paper should either estimate the 84 TeV cross sections and rerun the key projections, or quote a conservative range for the improvement factors.","section":"Sec. 2, footnote p. 3"}],"minor_comments":[{"comment":"n_k is defined as the expected number of events under the SM-only hypothesis, but in a likelihood it should denote the observed (or Asimov) data; please clarify the notation for expected limits.","section":"Eq. (4.1)"},{"comment":"The caption says 'multiclass NN', but the text in Sec. 6 describes a binary classifier and states that a multiclass network gave no real improvement. The caption should be corrected to match the body.","section":"Sec. 6, Table 4 caption"},{"comment":"'electron-proton collider studies' appears to be a typo; the context indicates electron-positron colliders.","section":"Sec. 5.2, footnote"},{"comment":"The |M_d6|^2 term in Eq. (1.3) is not discussed in the limit-setting procedure. It is presumably negligible for the small Wilson coefficients quoted, but this assumption should be stated explicitly.","section":"Sec. 4 / Eq. (1.3)"},{"comment":"Reference [28] is used to justify the absence of systematic uncertainties, but it is the same group's previous paper. An independent experimental or detector-level estimate would strengthen the argument.","section":"Sec. 4, ref. [28]"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, policy-relevant phenomenological projection and the qualitative ordering (FCC-hh strongest; ML observables help) is likely robust. My recommendation of major_revision is driven solely by the fact that the quantitative 'order of magnitude' claim currently rests on a statistical-only likelihood with no systematic uncertainty budget, and by the unpropagated shape-level concerns in the e+e- analysis. These are fixable within the manuscript's scope; they do not require new physics or a different analysis philosophy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious and mostly solid projection paper. It gives the first consistent, ATLAS-validated comparison of Higgs CP sensitivity across HL-LHC, FCC-ee, LCF, and FCC-hh, with new numerical tables and a couple of genuinely new observations: the electron-channel sign-flip in the interference near the Z pole, and the signal-background interference in llbb. The central claim, that future colliders improve on HL-LHC by roughly an order of magnitude, is supported by their own tables and is probably robust for the qualitative ordering.\n\nThe paper earns credit for reproducibility. Yields are checked against ATLAS in all three channels, the likelihood is clearly specified, and the ML observables are built on published methods rather than invented for this paper. The LCF polarization study is a nice addition. The writing is clear and the internal logic holds up.\n\nThe soft spots are real but not fatal. The biggest is the absence of any systematic-uncertainty budget. At FCC-hh, where signal yields are hundreds of thousands and backgrounds are tens of millions, the quoted limits are shape-limited, not statistics-limited. A flat \"systematics are expected to be small [28],\" citing their own work, does not cover charge/efficiency asymmetries, tau energy scale mismodeling, or background normalization that correlates with the CP-odd angle. The e+e- simulations also correct for missing ISR with a flat k-factor, but the ML gain is driven by a shape-level feature (the m_ll sign-flip), so a normalization-only correction is not obviously adequate. The 100 TeV vs 84 TeV baseline is acknowledged but not propagated; the impact is probably modest, but it should be quantified. The factor-of-two interference effect noted in Sec. 5.2 actually makes the FCC-ee constraints conservative, so that is a minor point.\n\nNone of this overturns the qualitative conclusion: FCC-hh is the strongest overall for Higgs CP, and all future machines beat HL-LHC. But the advertised \"order of magnitude\" factors should be read as idealized statistical projections, not robust predictions.\n\nWho should read this: anyone working on Higgs CP, SMEFT projections, or the ESPPU input process. It deserves a serious referee. The referee should ask for a sensitivity study on shape systematics and a clearer statement on the ISR correction. I would send it to review.","headline":"Solid, ATLAS-validated projection that future colliders give order-of-magnitude Higgs CP improvements; exact FCC-hh factors are statistical-only and will soften with systematics, but the qualitative case holds.","tokens_in":19704,"tokens_out":3031,"would_cite":true,"duration_ms":31382,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Future colliders could tighten Higgs CP-violation limits 20-fold.","keywords":["Higgs CP violation","dimension-six operators","effective field theory","future colliders","CP-odd observables","machine learning","weak boson fusion","H to four leptons"],"falsifier":"Compute the electron-channel e+e- to l+l- H interference at next-to-leading order in the electroweak theory; if the sign-flip near the Z pole persists with the same magnitude, the ML-amplified 4-5x sensitivity gain is real, while if it is washed out or shifted, the projected e+e- limits on c_PhiB and c_PhiWB must be re-evaluated.","tokens_in":1475,"feed_emoji":"⚛️","tokens_out":6586,"duration_ms":98929,"temperature":0.7,"pith_summary":"The paper argues that future electron-positron and proton-proton colliders would improve sensitivity to CP-violating Higgs couplings by roughly an order of magnitude or more compared with the high-luminosity LHC program. It studies three dimension-six operators that mix the Higgs field with weak gauge fields and uses CP-odd angular asymmetries, some machine-learned, to project 95% confidence intervals on the associated Wilson coefficients. Clean electron-positron machines give the best control for two of the three operators, while the 100 TeV proton-proton option is strongest overall, especially in the H to four-lepton and weak-boson-fusion channels. This matters because these couplings are direct probes of new sources of CP violation, which the baryon asymmetry of the universe likely requires.","feed_headline":"Future colliders could sharpen Higgs CP limits 20-fold","feed_subtitle":"Projections across all planned machines say CP violation in the Higgs sector can be measured far more precisely than at the upgraded LHC.","key_machinery":"The analysis is carried by the interference term between the Standard Model amplitude and the dimension-six amplitude, Re(M_SM* M_d6), which is CP-odd and integrates to zero over any CP-even observable. To expose it, the paper uses signed angular variables - Delta phi_ll in Z-Higgs production, Phi_4l in H to four leptons, Delta phi_jj in weak boson fusion - and machine-learned observables O_NN = P^+ - P^- built from classifiers trained to separate positively and negatively weighted interference events. These asymmetries map one-to-one onto the three Wilson coefficients c_PhiB/Lambda^2, c_PhiW/Lambda^2, and c_PhiWB/Lambda^2.","core_discovery":"The paper's central claim is that planned future colliders can probe anomalous CP-violating Higgs interactions roughly an order of magnitude more deeply than the HL-LHC. Using simulated events and a binned profile likelihood, the authors find that the H to four-lepton channel at a 100 TeV proton-proton collider constrains c_PhiB/Lambda^2 to +/-0.007 TeV^-2 and c_PhiWB/Lambda^2 to +/-0.015 TeV^-2, about twenty and 150 times better than HL-LHC projections; weak-boson-fusion H to tau tau with a machine-learned observable constrains c_PhiW/Lambda^2 to +/-0.007 TeV^-2, over twenty times better. The electron-positron machines provide the best constraints on c_PhiB and c_PhiWB in associated Z-Higgs","pith_inferences":["The electron-channel sensitivity gain at e+e- machines rests on a sign-flip in the interference near the Z pole that the muon channel lacks; a direct measurement comparing electron and muon channels would test whether this feature is physical or a generator artifact.","A natural extension is to apply the same interference-sign classification to existing LHC data in similar final states; this paper's framework suggests such gains are possible, but does not itself establish that they survive real systematic uncertainties.","The assumed symmetry of systematic uncertainties is the main place the advertised factors could shrink; a follow-up with full detector simulation and correlated systematics would bound the realistic improvement.","If the sign-flip near the Z pole is confirmed at higher order in electroweak perturbation theory, the ML-amplified asymmetry could become a standard tool for measuring CP violation at future Higgs factories."],"forward_implications":["If the projections hold, a 100 TeV proton-proton collider could measure CP-violating Higgs couplings at the level of roughly 0.01 TeV^-2, about twenty times tighter than the HL-LHC.","The machine-learned observables outperform traditional angular variables in every channel studied, most decisively in weak-boson-fusion H to tau tau, where they roughly double the sensitivity.","Electron-positron and proton-proton facilities are complementary: the e+e- machines dominate for c_PhiB and c_PhiWB in Z-Higgs production, while the pp machine dominates for c_PhiW via weak boson fusion.","Beam polarization at a linear collider improves limits by 1.2 to 1.8, partially compensating for its lower integrated luminosity.","The H to four-lepton channel benefits strongly from finer binning in Phi_4l and the dilepton mass at high-luminosity pp machines, producing gains larger than luminosity scaling alone would suggest."],"fun_headline_variants":["Future colliders could slash Higgs CP uncertainty 20-150x","Planned colliders promise big boost in Higgs CP violation tests","Next colliders could probe Higgs CP violation 150x deeper","Future machines could lift Higgs CP limits by up to 150-fold"],"cache_read_input_tokens":20608,"weakest_assumption_plain":"The projected improvements assume the simulated CP-odd interference shapes - particularly the sign-flip near the Z pole in the electron channel and the sign separation the classifiers learn - are accurate enough that a normalization-only ISR correction and the omission of systematic uncertainties do not materially change the constraints.","fun_headline_variants_meta":{"raw":{"variants":["Future colliders could slash Higgs CP uncertainty 20-150x","Planned colliders promise big boost in Higgs CP violation tests","Next colliders could probe Higgs CP violation 150x deeper","Future machines could lift Higgs CP limits by up to 150-fold"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":1774,"prompt_tokens":690,"completion_tokens":1084,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":434,"completion_tokens_details":{"reasoning_tokens":1011}},"tokens_in":434,"tokens_out":1084,"duration_ms":9885,"temperature":1.0,"reasoning_tokens":1011,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:51:03.898954+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the electron-channel e+e- to l+l- H interference at next-to-leading order in the electroweak theory; if the sign-flip near the Z pole persists with the same magnitude, the ML-amplified 4-5x sensitivity gain is real, while if it is washed out or shifted, the projected e+e- limits on c_PhiB and c_PhiWB must be re-evaluated.","supporting_citations":[],"review_version":1}