{"id":"c375044d-e768-406f-b9d4-ae88c7210382","arxiv_id":"1908.05286","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Matrix-element maximisation classifies ttH events with 60-85% of the significance of the traditional integration-based MEM while being up to two orders of magnitude faster.","lead":"This paper tests a faster way to classify particle collisions by guessing the momenta of invisible particles instead of averaging over all possibilities. It reports that the faster method can be up to one hundred times quicker than the standard matrix-element method, with a modest loss in separating signal from background.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Performance is measured on the same 2000-event sample used to choose the optimal cut and algorithm, so the claimed 60-85% of MEM significance is likely optimistic.","rationale":"Good-faith summary: the paper benchmarks matrix-element maximisation against the traditional MEM for ttH with a scan of NLOPT algorithms. The speed advantage is clearly demonstrated in Table II and Figure 4: even the balanced algorithm is roughly an order of magnitude faster than the MadWeight point, and fast local algorithms are much faster still. The performance comparison is the soft spot. Section III B uses one 2000-event sample to perform three coupled tasks: choose the chi0 cut per algorithm, choose the preferred algorithm from Figure 4, and then report the significance of that preferred algorithm. This is an in-sample evaluation of a quantity that has been optimised over cuts and algorithms. With only 1000 signal and 1000 background events and no uncertainty bands, the reported 60%-85% range is an optimistic estimate. The zero-result exclusion rules further move the acceptances in an uncontrolled direction. None of this invalidates the method, but it makes the abstract's quantitative cost/benefit claim contingent on a held-out validation. The reader's conditional verdict already reflects the need for such checks; the present concern sharpens that need, so the verdict remains unchanged.","tokens_in":15569,"tokens_out":9658,"duration_ms":102088,"concrete_test":"Split the 2000 events into two disjoint sets, e.g. 1000 design and 1000 validation, or use 5-fold cross-validation. On the design set only, choose (i) the chi0 cut for each algorithm and (ii) the winning algorithm. Then evaluate the discovery significance on the validation set using the preselected cut and algorithm, and compare with the same protocol for MadWeight. Repeat with bootstrap resampling of the full 2000 events to attach uncertainties. If the validation significance of the selected maximisation algorithm is no longer between 60% and 85% of MadWeight, the reported performance range is inflated by in-sample selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that matrix-element maximisation is far cheaper than the integrated MEM with only a mild performance loss—rests on significance numbers computed in Section III B from a single set of 2000 events (1000 signal, 1000 background). For every algorithm, the optimal cut chi0 is chosen on that same sample (Eq. 9 and Figure 4), and the preferred algorithm GN DIRECT L RAND is selected from the same plot. Thus the quoted '60%-85% of the traditional MEM' is the maximum over a scan of 18 algorithms and the cut choices, not the expected performance of a fixed procedure. With no error bars and 1000 signal and 1000 background events, selection effects are material: the best point in such a scan will tend to look better than a held-out replication. The handling of zero-result events (zero-background events are removed from the background acceptance, and events with both weights zero are ignored) also changes N_b and N_s in an uncontrolled way. The convergence question raised by the reader is real but secondary here: the chosen global algorithm has low zero-result rates, so the dominant threat to the 60-85% figure is in-sample selection rather than poor optima.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, forget the abstract's 'slight reduction': on the authors' own numbers the best maximiser gives about 85% of MadWeight's significance (or 60% for faster algorithms), and the real message is a two-orders-of-magnitude CPU saving with a moderate, tunable loss. That is a useful result for ttH to bb analyses where full MEM is computationally prohibitive.\n\nWhat is new: this is the first systematic application of the matrix-element maximisation idea from Ferreira de Lima, Mattelaer, and Spannowsky to ttH, with a scan over 17 NLOPT derivative-free algorithms, CPU timings, zero-result rates, and significance benchmarks against MadWeight. The comparison is honest in that they use an independent implementation, MadWeight, as baseline. The discussion of reconstructed neutrino momenta is also valuable: it shows frankly that the maximised momenta are biased toward pole masses and should not be interpreted as true momenta, only as seeds for event-deconstruction type tools.\n\nSoft spots. The significance numbers are computed on the same 2000 events used to choose both the cut and the algorithm. Figure 4 is effectively a maximum over 18 algorithms; without error bars or a validation sample, the 60-85% claim is likely optimistic, especially for the best point. The zero-result handling is ad hoc: events with zero background weight are removed from the background acceptance and events with both weights zero are ignored, and the text does not quantify the effect on Ns and Nb. That matters for the significance ranking. There is no code or data release, so the timing and zero-result numbers cannot be reproduced. The 15% transfer-function width is calibrated to ATLAS simulation rather than derived, which is reasonable but adds a systematic layer that the significance comparison ignores. The convergence concern raised in the stress-test is real but secondary: the chosen global algorithm has low zero-result rates, so the bigger issue is in-sample selection.\n\nWho is this for? Practitioners considering MEM for high-multiplicity final states, and tool developers. The paper deserves a serious referee, but the referee should ask for out-of-sample validation, error bars on the significance, and a clearer treatment of zero-result events. I would not cite the 60-85% figure without caveats, but I would cite the benchmark as evidence that maximisation is a practical route to MEM-style classification.","headline":"A useful benchmark showing matrix-element maximisation is far cheaper than integration for ttH at a moderate, likely in-sample-inflated significance loss.","tokens_in":16334,"tokens_out":1994,"would_cite":true,"duration_ms":19710,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that replacing the phase-space integration of the matrix-element method with a maximisation over invisible momenta classifies fully-leptonic ttH events at up to two orders of magnitude lower CPU cost, with discovery…","keywords":["matrix-element method","matrix-element maximisation","ttH production","Higgs to bottom quarks","event classification","invisible particle reconstruction","derivative-free optimisation","LHC searches"],"falsifier":"Run the same 2000-event test set through a high-precision reference maximiser, for example a dense grid or many random restarts over the four free parameters with the same transfer function, and compare each algorithm's reported maximum. If the best algorithm's discovery significance falls outside the paper's 60-85% band relative to that reference, or if the maximised weights frequently miss the reference maximum by more than the stated 1% stopping precision, the central claim that maximisation approximates the MEM classifier is undercut.","tokens_in":15407,"feed_emoji":"⚛️","tokens_out":16202,"duration_ms":140436,"temperature":0.7,"pith_summary":"This paper is trying to establish that the matrix-element method (MEM), the standard way to score LHC events as signal or background by comparing them with theoretical amplitudes, can be run as a maximisation instead of an integration. Applied to top-associated Higgs production with the Higgs decaying to two b-quarks and both tops decaying leptonically, the maximisation version assigns each event a weight by finding the peak of the squared matrix element times a detector-resolution transfer function over the unknown neutrino momenta. The paper reports that this is up to two orders of magnitude faster in CPU time than the traditional integration-based MEM, while the resulting classifier achieves roughly 60-85% of the traditional method's discovery significance, depending on the optimisation algorithm. It also finds that the maximiser returns an estimate of the invisible four-momenta, though with a known bias, and argues such estimates can seed more elaborate reconstruction tools. If true, MEM-style classification becomes practical for high-multiplicity final states and large datasets where integration is prohibitive.","feed_headline":"Matrix-element maximisation cuts ttH scoring cost by ~100x","feed_subtitle":"Replacing integration with optimisation keeps 60-85% of the discovery significance at a fraction of the CPU cost.","key_machinery":"The object that carries the argument is the maximisation weight $w_\\alpha(x) = \\max_{y\\in\\Phi} |M_\\alpha|^2(y)\\, W(x,y)$, where $W(x,y)$ is a transfer function that penalises moving the four b-jet energies away from their measured values (a product of Gaussians with 15% resolution). To make the objective peak, the four free parameters left after momentum conservation are re-expressed as the invariant masses of the top and W propagators -- the 'Main Block D' kinematic template -- and each variable is then mapped to the unit interval through the cumulative distribution function of its Breit-Wigner or Gaussian model. The classifier is the log-ratio $\\chi = \\log w_s / \\log w_b$, and the scan over local and global derivative-free algorithms is what produces the speed-versus-significance tradeoff. The same machinery is generic: it needs only the squared matrix element, a transfer function, and a choice of maximisation variables that peak near the true solution.","core_discovery":"The central claim is that the maximisation weight $w_\\alpha(x) = \\max_{y\\in\\Phi} \\{ |M_\\alpha|^2(y)\\, W(x,y)\\}$ can replace the integrated MEM probability in a signal-vs-background discriminant $\\chi(x)=\\log w_s(x)/\\log w_b(x)$, and that for fully-leptonic $t\\bar{t}h$ with $h\\to b\\bar{b}$ it does so at a fraction of the cost. Scanning sixteen derivative-free optimisation algorithms, the paper finds that the fastest ones reduce the average time per event to a few seconds, while global algorithms such as GN DIRECT L RAND give the best balance of speed and separation power. The cost is a loss in classifier quality: the best maximisation-based discovery significance is about 15% below the traditional MEM, and across algorithms the significance ranges from about 60% to 85% of the MEM value. A second claim is that the maximising phase-space point provides a four-momentum estimate for each neutrino; these are biased (neutrino $p_T$ about 70% too high on average, $p_Z$ close on average but with a huge spread) because the optimum sits at the Breit-Wigner pole masses of the top and W propagators, so the estimates should be most trustworthy when the resonances are narrow.","pith_inferences":["The 60-85% significance gap may be partly an artifact of the stopping criteria and transfer function; using the maximiser's reconstructed momenta as a warm start for a short local integration or a second-stage refinement could recover most of the lost significance while keeping the speed gain.","The bias toward pole masses suggests a calibration route: on signal Monte Carlo, a response map from true to reconstructed neutrino momenta could be built and applied as a correction, turning the method into a quantitative missing-momentum estimator rather than just a classifier.","Because the algorithm scan is process-dependent, other final states will likely need their own scan; the transferable recipe is the peaked-variable reparameterisation and unit-cube mapping, not any single optimiser choice.","The nonzero failure rates (6-13% of events for local algorithms) could be mitigated in practice by restarting the optimiser or falling back to a global algorithm on failure, still likely far cheaper than full phase-space integration."],"forward_implications":["Replacing integration with maximisation makes MEM-style classification feasible for final states with several invisible particles; the cost per event drops to seconds for the fastest algorithms, so large LHC datasets become tractable.","Users can choose an operating point on the speed-significance curve: fast local algorithms for quick scans, global algorithms for maximum separation, and a balanced default between the two.","Every maximised event comes with a concrete four-momentum assignment for the invisible particles, which can be passed to more detailed reconstruction tools or used as a starting point for further MEM calculations.","For processes dominated by narrow-width resonances, the maximised invisible momenta should be close to the true values, because the objective naturally peaks at the pole masses; this offers a cheap way to approximate missing momenta in complicated decay chains.","The method inherits the MEM's advantage over machine-learning classifiers: no training on generated pseudo-data is required, since the weights come from first-principle matrix elements."],"supporting_citations":[{"why":"Introduces the matrix-element maximisation procedure that this paper extends and benchmarks; supplies the core method.","marker":"[26]"},{"why":"Defines the automated matrix-element reweighting method and the Main Block D kinematic template used to make the objective peak; also provides the integration-based reference weight.","marker":"[3]"},{"why":"Demonstrates the traditional integration-based MEM on the ttH process, establishing the physics case and the performance baseline.","marker":"[11]"},{"why":"Provides the automated tree-level matrix elements and the generated event samples used for both signal and background.","marker":"[27]"},{"why":"Supplies the derivative-free optimisation algorithms whose speed and discovery significance are scanned in the performance study.","marker":"[29]"},{"why":"The no-free-lunch theorem for optimisation, cited to justify scanning many algorithms before choosing one.","marker":"[28]"},{"why":"An LHC measurement of Higgs decay to bottom quarks used to calibrate the 15% b-jet energy resolution modelled by the Gaussian transfer function.","marker":"[55]"},{"why":"Provides the signal cross section used to convert acceptance rates into the expected discovery significance.","marker":"[56]"}],"fun_headline_variants":["Optimize, don't integrate: 100x faster ttH scoring","Matrix-element maximization: 100x speed, 85% signal","Swap integration for optimization, 100x faster MEM","Maximization speeds MEM ~100x with minor loss","CPU-efficient MEM: 100x faster, still 85% significant"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the numerical maximum returned by the optimiser is close enough to the true highest value of the matrix-element weight to serve as a reliable classifier score; if optimisation frequently fails or stops early, both the reported significance and the speed ranking could change.","fun_headline_variants_meta":{"raw":{"variants":["Optimize, don't integrate: 100x faster ttH scoring","Matrix-element maximization: 100x speed, 85% signal","Swap integration for optimization, 100x faster MEM","Maximization speeds MEM ~100x with minor loss","CPU-efficient MEM: 100x faster, still 85% significant"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1533,"prompt_tokens":1078,"completion_tokens":455,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":694,"completion_tokens_details":{"reasoning_tokens":366}},"tokens_in":694,"tokens_out":455,"duration_ms":5307,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:18:50.143848+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 2000-event test set through a high-precision reference maximiser, for example a dense grid or many random restarts over the four free parameters with the same transfer function, and compare each algorithm's reported maximum. If the best algorithm's discovery significance falls outside the paper's 60-85% band relative to that reference, or if the maximised weights frequently miss the reference maximum by more than the stated 1% stopping precision, the central claim that maximisation approximates the MEM classifier is undercut.","supporting_citations":[{"cited_title":"Reconstructing the invisible with matrix elements","cited_arxiv_id":"1712.03266","evidence_quote":"Introduces the matrix-element maximisation procedure that this paper extends and benchmarks; supplies the core method."},{"cited_title":"The NLopt nonlinear-optimization package","cited_arxiv_id":null,"evidence_quote":"Supplies the derivative-free optimisation algorithms whose speed and discovery significance are scanned in the performance study."},{"cited_title":"No free lunch theorems for optimization,","cited_arxiv_id":null,"evidence_quote":"The no-free-lunch theorem for optimisation, cited to justify scanning many algorithms before choosing one."}],"review_version":1}