{"id":"54542155-4f69-4c4e-a925-ce20ed5f5aec","arxiv_id":"1908.03952","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Lilith-2.0 updates the Higgs signal-strength likelihood tool to ATLAS and CMS Run 2 data, adding variable Gaussian and Poisson likelihoods, arbitrary-dimensional correlations, and new global coupling fits.","lead":"Lilith-2.0 is a public Python tool that turns published Higgs signal-strength measurements into a global likelihood for constraining non-SM Higgs physics. This paper presents the update to LHC Run 2 data (36 fb^-1), adds variable Gaussian and Poisson likelihood options, and gives new global fits for Higgs couplings, 2HDMs, and invisible decays.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The H→ZZ* VH/ttH entries convert one-sided 95% CL limits to two-sided Gaussians with σ = L/1.96, effectively tightening the upper limits by ~16%; the choice was tuned to reproduce the official contour, so the DB's global-fit reliability is not independently validated.","rationale":"The reader's weakest assumption identifies reconstructed likelihoods as the critical input, and the H→ZZ* VH/ttH limit conversion is the clearest concrete instance. I agree with that identification and sharpen it: converting a one-sided 95% CL limit to a two-sided Gaussian via σ = L/1.96 makes the effective one-sided upper limit 0.84L, a 16% tightening that can change exclusion decisions for models with moderately enhanced VH/ttH rates. The paper's choice between one- and two-sided Gaussians based on matching the ATLAS official CF vs. CV contour is a circular validation step; agreement with a 2D projection does not establish that the combined likelihood has correct coverage in the multi-dimensional fits the tool is meant to support. This does not undermine the overall software release: the paper is transparent, the code is public, and the validations are extensive. But it does justify the CONDITIONAL verdict already reached by the reader, so I recommend no change to the verdict. The concrete test would settle the concern by quantifying how much the global-fit results shift under a more faithful treatment of the VH/ttH limits.","tokens_in":23305,"tokens_out":11966,"duration_ms":132856,"concrete_test":"Recompute the Section 6 global fits (CF vs. CV, Cg vs. Cγ, and BR(H→inv)) after replacing the H→ZZ* VH/ttH entries with (a) one-sided Gaussians with σ = L/1.64 using L = 3.70 and 7.51, (b) the published CLs profile if available from HEPData, and (c) omission of these two entries. If the best-fit values or 95% CL interval boundaries move by more than about 10% of the quoted uncertainties relative to DB 19.09, the limit-conversion approximation is material to the claimed constraints and must be documented as a caveat before the database is used for models with enhanced VH/ttH rates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 5.1 (H→ZZ*), ATLAS VH and ttH signal strengths are implemented as µ = 0 ± 1.89 and 0 ± 3.83, converted from 95% CL limits by a two-sided Gaussian (σ = L/1.96). A 95% CL upper limit is one-sided: for a Gaussian centered at zero, the one-sided 95% upper bound is 1.64σ, so the effective limits become 0.84×L (about 3.1 and 6.3) instead of the published L (about 3.7 and 7.5). The database therefore over-constrains positive µ(VH,ZZ*) and µ(ttH,ZZ*) by about 16% relative to the published limits; a model predicting µ(VH,ZZ*) ≈ 3.5 would be allowed by ATLAS at 95% CL but excluded by Lilith. The authors justify this normalization by saying that a one-sided assumption gives a less good match to the official CF vs. CV contour, which means the implementation is tuned against a 2D validation plot rather than tested independently. These entries feed the global fits, so the central claim that the database is 'ready to be used' depends on this approximation not materially shifting the quoted constraints, such as CV = 1.068 ± 0.030 or the 2HDM exclusion regions. No coverage test or sensitivity study is provided for the combined likelihood when individual limit-derived entries are replaced by alternative, equally plausible parametrizations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Lilith-2.0, a public Python library for constraining new physics from Higgs signal-strength measurements, together with an updated XML database (DB 19.09) that incorporates ATLAS and CMS Run 2 results at 36 fb^-1. The main technical novelties are the extension from ordinary Gaussian likelihoods to variable-Gaussian and generalized-Poisson parametrizations, the support of arbitrarily large correlation matrices, and the addition of new production modes (ggZH, tH, bbH). The authors document each experimental input, show validation plots against official ATLAS/CMS contours where available, and give updated global fits for reduced couplings, 2HDM Types I and II, and invisible Higgs decays, finding e.g. C_V = 1.068 +/- 0.030 and BR(H -> inv) < 5% at 95% CL for SM-like couplings. The central claim is that Lilith-2.0 with DB 19.09 is ready for use in constraining a wide class of new physics scenarios.","tokens_in":23756,"tokens_out":7653,"duration_ms":82965,"significance":"If the database is reliable, this is a valuable community resource: it is lightweight, Python-based, publicly available, and it makes efficient use of the best public Higgs measurements, including the full 24x24 CMS correlation matrix. The paper is generally transparent about its approximations, explicitly flagging, for instance, the LO treatment of tHW, the assumed WW/ZZ composition of the ttH VV final state, and the fact that the paper updates rather than replaces the Lilith-1.1 manual. The global-fit results are plausible and consistent with the SM. The chief risk is not the statistical formalism but the fidelity of the reconstructed likelihoods for entries that are digitized from plots or converted from limits; those entries feed the global fits and exclusions, so the 'ready to be used' claim rests on their accuracy.","major_comments":[{"comment":"The conversion of the ATLAS 95% CL upper limits for mu(VH,ZZ*) and mu(ttH,ZZ*) into two-sided Gaussians with sigma = L/1.96 (implemented as 0 +/- 1.89 and 0 +/- 3.83) makes the effective 95% upper bound 1.64 sigma = 0.84 L rather than the published L; if the limits are one-sided, the appropriate Gaussian width would be sigma = L/1.64, so the database is about 16% tighter for positive signal strengths. The authors state that this normalization was chosen because a one-sided assumption gives a less good match to the official C_F vs. C_V contour, which is a tuning against the validation target rather than an independent test. Since these two entries enter all global fits in Section 6, the central claim that DB 19.09 is ready to be used needs a sensitivity study showing how the quoted constraints (C_V = 1.068 +/- 0.030, C_F = 1.045 +/- 0.064, BR(H->inv) < 5%) change when the entries are instead implemented as one-sided Gaussians or with alternative asymmetric widths.","section":"Section 5.1, H->WW and H->tau tau; Fig. 3"},{"comment":"For several ATLAS Run 2 entries the validation is partly circular: the parameters of the approximating likelihood are fitted directly to the experimental 95% CL contour, and the validation plot then compares the Lilith reconstruction with that same contour (e.g., Fig. 3, top-right and bottom panels). This demonstrates internal consistency but not that the parametrization is accurate away from the fitted region or in the tails used for 95% CL exclusions. For these channels no official coupling fit is available, so I ask for an independent cross-check, such as a comparison with the underlying category-level likelihood or a coverage test of the reconstructed likelihood.","section":"Section 5.1, H->mu mu"},{"comment":"The same one-sided/two-sided ambiguity affects the H->mu mu entry: ATLAS reports mu = -0.1 +/- 1.5 with a CLs 95% upper limit of 3.0, and the database implements mu = 0 +/- 1.53. Under a one-sided Gaussian interpretation this corresponds to a 95% upper bound of 1.64 x 1.53 = 2.51, which is more restrictive than the published limit rather than less restrictive. If the intention is to avoid over-constraining, the width should be larger (sigma ~ 1.83 if matching the one-sided limit), or the two-sided interpretation should be explicitly stated and justified.","section":"Section 5.1, H->mu mu"}],"minor_comments":[{"comment":"The text 'This does does not always allow' contains a duplicated 'does' and should read 'This does not always allow'.","section":"Section 2"},{"comment":"In the sentence describing the H->tau tau likelihood, 'mu(VBF,WW) ~ 1.20+0.62-0.56' should read 'mu(VBF,tau tau) ~ 1.20+0.62-0.56'.","section":"Section 5.1, H->tau tau"},{"comment":"The caveat that the HIGG-2017-02 XML file should not be used when C_Z != C_W is important; the code should ideally issue an explicit warning when this file is loaded under assumptions that violate that condition.","section":"Section 5.1, ttH combination"},{"comment":"The statement that 8 TeV cross sections are used for sqrt(s) = 7 TeV because differences are negligible would be more helpful if accompanied by a quantitative estimate or a reference for that estimate.","section":"Section 4"},{"comment":"The text notes that even with the Poisson form the 95% CL contour for H->ZZ* is still 'quite off' before the auxiliary profile likelihoods are used; this reinforces the sensitivity concerns raised in the major comments and should be discussed in the main text rather than only in the appendix.","section":"Appendix A, Fig. 14"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its approximations and the code is likely to be a useful community resource. My main concern is that the validation strategy is partly circular and that the H->ZZ* limit conversion introduces a small but concrete bias; these issues are fixable with sensitivity analyses rather than being fundamental flaws. I would not hold the paper to the standard of exactly reproducing the official likelihoods, but the central 'ready to be used' claim should be backed by robustness tests that are currently missing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is the paper to cite when you use Lilith-2.0 for Run 2 Higgs constraints. It is a software release with validation, not a discovery paper, and the headline numbers — CV at 3–4% precision, BR(H→inv) ≲ 5% — agree with the SM and with other global fits. The reader's conditional verdict is fair, and the stress-test math checks out, but neither concern is fatal.\n\nWhat is actually new is the implementation: variable Gaussian and Poisson likelihoods, an XML format that handles the CMS 24×24 correlation matrix, new production modes (ggZH, tH, bbH), and the curated DB 19.09 database of ATLAS/CMS 36 fb^-1 results. The statistical machinery comes from Barlow and from Berkhout & Plug, and the authors say so. The validation plots are good practice: reconstructed contours drawn against the official ones, with an appendix showing where the parametrization choice matters. Code is on GitHub and interfaced with micrOMEGAs. The citation practice is clean — statistical precedents credited, experimental sources primary. That is real, citable value.\n\nThe soft spots are in the limit conversions, as the stress-test says. For H→ZZ*, the ATLAS 95% CL upper limits on VH and ttH are implemented as two-sided Gaussians with σ = L/1.96. A one-sided 95% bound of L would imply σ = L/1.64, so the effective limits tighten by about 16%: a model predicting μ(VH,ZZ*) near 3.5 would sit within the published ATLAS limit but outside Lilith's likelihood for that channel. The paper's own text admits the two-sided choice was made because the one-sided version matched the official CF-vs-CV contour less well. Those entries are tuned against a validation plot, not independently tested. This is moderate, not fatal: the affected channels are weak, and the headline fit values will not move much. But the 'ready to be used' claim deserves a sensitivity test for these entries, or at least a prominent warning in the XML files.\n\nSame family, smaller: H→μμ is set to 0±1.53 with the stated aim of avoiding over-constraining, yet the one-sided 95% bound from that Gaussian is 2.51, tighter than the quoted 3.0 limit. The justification does not survive contact with the math; the practical impact is negligible.\n\nOther limitations are handled honestly: the ttH VV 95/5 WW/ZZ split carries an explicit warning not to use that file when CZ ≠ CW, and the paper is clear about where digitized plots had to be pressed into service.\n\nBottom line: it deserves a serious referee and will become a standard tool. The revisions I would want are a sensitivity test or explicit caveat for the limit-converted entries, and a corrected sentence on the μμ conversion. For a reading group, it is a good case study in how approximate global likelihoods get built and validated.","headline":"Genuinely useful software update with an honest but imperfect validation: the limit-to-Gaussian conversions in H→ZZ* and H→μμ are the soft spot; the headline global-fit numbers are safe.","tokens_in":24192,"tokens_out":11136,"would_cite":true,"duration_ms":103655,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Lilith-2.0 turns LHC Run 2 Higgs measurements into a reusable global likelihood, with validated fits pinning the Higgs couplings to better than 10 percent.","keywords":["Lilith","Higgs signal strengths","global fits","reduced couplings","kappa framework","two-Higgs-doublet model","invisible Higgs decays","LHC Run 2"],"falsifier":"Extract the true profile likelihoods for the ATLAS H->ZZ* VH and ttH channels from the collaboration's public material and compare them with the database's two-sided Gaussian entries, mu(VH,ZZ*) = 0 +/- 1.89 and mu(ttH,ZZ*) = 0 +/- 3.83; if those true likelihoods are markedly one-sided, the conversion used in DB 19.09 biases the global fits, and the paper's reliability claim would be falsified.","tokens_in":23111,"feed_emoji":"🔬","tokens_out":9351,"duration_ms":98977,"temperature":0.7,"pith_summary":"This paper presents Lilith-2.0, an updated public Python library that converts published LHC Higgs signal-strength measurements into a global likelihood for theorists to use in constraining new physics. The central claim is that this lightweight tool, together with its new database of ATLAS and CMS Run 2 results for 36 $fb^{-1}$, reproduces the experimental collaborations' own coupling fits closely enough to be trusted as a substitute for the full experimental likelihoods. The technical upgrade replaces the old fixed-width Gaussian approximation with variable-width Gaussian and Poisson likelihood forms and allows correlation matrices of arbitrary dimension, which is what makes the Run 2 data usable. On the physics side, the paper reports that the Run 2 data are perfectly compatible with the Standard Model, determine reduced Higgs couplings to better than 10 percent, and bound the invisible Higgs branching fraction at about 5 percent at 95% CL for Standard-Model-like couplings. A sympathetic reader would take the paper's contribution to be demonstrated reliability and ready availability of this approximation machinery rather than any claim of new physics.","feed_headline":"Lilith-2.0 pins Higgs couplings to ~3 percent","feed_subtitle":"Public Python library combines all ATLAS and CMS Run 2 Higgs measurements into one validated global fit.","key_machinery":"The carrying object is the signal-strength likelihood: each measured production-and-decay rate, mu(X,Y), is defined as the ratio of the observed to the Standard-Model rate, and the product of per-measurement likelihoods forms the global likelihood used for fits. The upgrade to version 2.0 lies in the shapes allowed for each piece: a variable Gaussian, whose width grows or shrinks with the parameter value to absorb asymmetric errors; a generalised Poisson, a count-based form with a free shape parameter tuned to the quoted uncertainties; and multi-dimensional Gaussians with arbitrary correlation matrices. The translation from rates to physics is made through reduced couplings CX and CY that scale Standard-Model production and decay amplitudes, which the paper identifies with the kappa framework, with the option to profile over invisible and undetected decay widths. This machinery converts the narrow-width, SM-tensor-structure assumption into concrete intervals on coupling combinations, which is what makes a wide class of new-physics models testable with a single downloaded tool.","core_discovery":"On the paper's own terms, the result is that Lilith-2.0 can reconstruct the likelihoods of the current ATLAS and CMS Run 2 Higgs measurements with enough fidelity that its global fits track the official coupling contours, and that the accompanying database DB 19.09 is ready for general use. Combining the ATLAS and CMS Run 2 data, the global fit of the fermionic and vector reduced couplings gives CF = 1.045+0.064-0.063 and CV = 1.068 +/- 0.030, with the two experiments agreeing at about the 1 sigma level. In the same framework, the visible signal strengths alone already constrain the invisible decay branching fraction to about 5 percent at 95% CL for Standard-Model-like couplings. The slight preference for a vector coupling above one, if taken at face value, pushes two-Higgs-doublet models, where CV is bounded by one, deeper into the alignment limit.","pith_inferences":["Extension: the accuracy of the DB 19.09 parametrisations will get a natural stress test when ATLAS and CMS release their full 139 fb^-1 Run 2 combinations; if the official profile likelihoods for CV and BR(H->inv) shift markedly from Lilith's, the simplified shapes should be revised.","Extension: the paper's own complaints about digitizing plots and hand-typing correlation matrices suggest the next natural step is direct ingestion of STXS-binned data or machine-readable likelihood files, eliminating the lossy reconstruction step.","Extension: the same validation-then-fit pattern used here is transferable to other LHC observables where experiments publish only summary statistics, so the methodology is a template beyond Higgs physics.","Extension: the headline '5 percent bound' on invisible Higgs decays is conditional on Standard-Model-like couplings; theorists with fermion or vector coupling freedom should quote the roughly 15-16 percent bound instead of the headline value."],"forward_implications":["Any new-physics model that only rescales Standard-Model Higgs couplings can be tested immediately against the full Run 2 Higgs dataset by running Lilith-2.0, without re-implementing the experimental likelihoods.","The roughly 3-4 percent determination of CV and sub-10 percent determinations of the fermion and vector couplings put concrete pressure on scenarios predicting percent-level coupling deviations.","The invisible branching fraction bound of about 5 percent at 95% CL for Standard-Model-like couplings directly restricts Higgs-portal dark-matter models, and tightens to about 4 percent when Run 1 data are added.","The preference for CV slightly above one strengthens the case that two-Higgs-doublet models live in the alignment limit, disfavouring large deviations in tan(beta) and cos(beta-alpha).","Because the database format now supports arbitrary correlation matrices, the same machinery can be extended with future full-Run-2 and HL-LHC measurements as they are published."],"supporting_citations":[{"why":"Defines the original Lilith signal-strength likelihood machinery, XML syntax, and ordinary-Gaussian treatment that version 2.0 extends.","marker":"[1]"},{"why":"Provides the variable-Gaussian and generalised-Poisson parametrisations used to handle asymmetric uncertainties.","marker":"[11]"},{"why":"Supplies the CMS combined signal strengths with the 24x24 correlation matrix that is both a database input and a validation target.","marker":"[12]"},{"why":"Provides the ATLAS H->gamma gamma Run 2 results and auxiliary profile likelihoods from which the implemented signal strengths are extracted.","marker":"[16]"},{"why":"Provides the ATLAS H->ZZ* measurements whose profile likelihoods are fitted as Poisson distributions and validated against official contours.","marker":"[21]"},{"why":"Supplies the combined ATLAS and CMS Run 1 measurements used in the Run 2 + Run 1 global fits.","marker":"[10]"},{"why":"Gives the Standard-Model Higgs production cross sections and accuracies used to convert measurements into signal strengths and efficiencies.","marker":"[6]"},{"why":"Defines the kappa-style reduced-coupling framework in which the global fits are expressed.","marker":"[9]"},{"why":"Provides the CMS vector-boson-fusion invisible-decay likelihood grid that contributes to the BR(H->inv) constraint.","marker":"[32]"}],"fun_headline_variants":["Lilith-2.0 tightens Higgs coupling fits to ~3%","New Lilith release sharpens Higgs constraints with Run 2 data","Higgs coupling precision: Lilith-2.0 reaches ~3% from ATLAS+CMS","Lilith-2.0: Higgs signal strengths meet global fits","Run 2 Higgs data: Lilith-2.0 constrains invisible decays to 5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simplified likelihood shapes stored in the database, built from published best-fit values, uncertainties, and correlations, capture the real experimental likelihoods closely enough that approximate entries, such as H->ZZ* limits converted from 95% CL bounds, do not bias the global fit.","fun_headline_variants_meta":{"raw":{"variants":["Lilith-2.0 tightens Higgs coupling fits to ~3%","New Lilith release sharpens Higgs constraints with Run 2 data","Higgs coupling precision: Lilith-2.0 reaches ~3% from ATLAS+CMS","Lilith-2.0: Higgs signal strengths meet global fits","Run 2 Higgs data: Lilith-2.0 constrains invisible decays to 5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000791,"raw_usage":{"total_tokens":3454,"prompt_tokens":885,"completion_tokens":2569,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":2460}},"tokens_in":501,"tokens_out":2569,"duration_ms":17969,"temperature":1.0,"reasoning_tokens":2460,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:56:40.051739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Extract the true profile likelihoods for the ATLAS H->ZZ* VH and ttH channels from the collaboration's public material and compare them with the database's two-sided Gaussian entries, mu(VH,ZZ*) = 0 +/- 1.89 and mu(ttH,ZZ*) = 0 +/- 3.83; if those true likelihoods are markedly one-sided, the conversion used in DB 19.09 biases the global fits, and the paper's reliability claim would be falsified.","supporting_citations":[],"review_version":1}