{"id":"fe1443f7-7944-4f36-ba73-d9e3cb3876f4","arxiv_id":"2607.22281","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Adding a moment-decomposition penalty to LHCb's topological b trigger neural networks reduces the efficiency drop at large b-hadron lifetimes with little loss of classification performance.","lead":"LHCb's software trigger uses neural networks to pick out b-hadron decays, but in busy Run 3 collisions it was unintentionally suppressing long-lived candidates. The paper shows two penalty-based fixes and finds that a moment-decomposition penalty flattens the trigger's efficiency at large b lifetimes better than a distance-correlation penalty.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Statistical support for 'MoDe better' is weak: slope differences have overlapping uncertainties and λ values are not matched.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but my load-bearing concern differs from the reader's stated weakest assumption. The reader focused on simulation realism; while that is a valid external-validity concern, the more immediate threat to the central claim is that the evidence internal to simulation is statistically inconclusive. Table 1's slopes have overlapping uncertainties, and no significance test is given for the MoDe-vs-DisCo difference. The unmatched λ values (chosen for classification, not decorrelation) further confound the comparison. These issues are directly testable with a bootstrap and a λ-scan, and they do not rely on assumptions about detector simulation. The paper's main claim is plausible but not established by the presented numbers; therefore the CONDITIONAL verdict stands unchanged. I agree with the reader partially because both concerns are valid, but the statistical weakness is the most load-bearing for the specific headline claim.","tokens_in":10430,"tokens_out":3116,"duration_ms":27930,"concrete_test":"Bootstrap re-analysis of the efficiency curves (Fig. 4) to compute a 95% confidence interval for Δslope ≡ slope_DisCo − slope_MoDe for both 2- and 3-body models, using event-level resampling. If the interval includes zero for either model, the 'MoDe better' claim is not supported. Additionally, repeat the comparison at matched operating points (e.g., equal integrated efficiency or equal mean |slope| across the λ scans in Appendix B) to test whether the ranking is an artifact of the chosen λ values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Sec. 5—that MoDe is better suited—rests on the slopes in Table 1. For 2-body models, MoDe gives -11.23±1.6 ns^-1 and DisCo -12.9±4.7 ns^-1; their uncertainties overlap by over a standard deviation. For 3-body models, MoDe is -6.6±5.5 and DisCo -8.8±1.2; the MoDe interval is very wide and contains the DisCo estimate. No statistical test of the slope difference is provided, and the fits are only for B0→D−π+ (Fig. 4); no slopes are reported for B+→J/ψK+ (Fig. 3). Moreover, the comparison uses λ=0.8 for DisCo and λ=0.2 for MoDe, selected from Fig. 5 based on classification metrics (precision/recall/ROC AUC), not on decorrelation strength. Because the decorrelation strength is not matched between methods, a different λ for either method could plausibly change the ranking—especially for 3-body, where MoDe's slope is consistent with both zero and the DisCo value. The claimed superiority thus rests on point estimates with overlapping uncertainties and an unmatched operating-point comparison, making the conclusion fragile.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses a lifetime-correlation problem in the LHCb topological b trigger, introduced by the higher pile-up conditions of Run 3 (mu = 5.3). The authors identify a background arising from tracks produced in different primary vertices being combined into a candidate, which creates a long-lifetime tail that can bias the trigger efficiency at large decay times. To mitigate this, they train monotonic Lipschitz neural networks for the 2-body and 3-body topological lines using two decorrelation penalties: distance correlation (DisCo) and moment decomposition (MoDe). They compare the resulting trigger efficiencies as a function of true b-hadron lifetime in simulated B+ -> J/psi K+ and B0 -> D- pi+ samples, fit linear slopes above 4 ps for the B0 sample, and conclude that MoDe is better suited for lifetime decorrelation in the topological b trigger.","tokens_in":10776,"tokens_out":2663,"duration_ms":26078,"significance":"If the central claim is established, the paper is practically significant: it identifies a concrete Run 3 trigger limitation and demonstrates two workable mitigation strategies, with potential to preserve large-lifetime acceptance for time-dependent LHCb analyses. The paper is empirically grounded: the authors train and evaluate actual MLNN models in full Pythia/EvtGen/Geant4 simulation, include a no-decorrelation baseline that makes the problem visible, and provide slope fits with uncertainties. The identification of the PV-misassociation background and the demonstration that both penalties reduce the negative large-lifetime slope are useful contributions. However, the headline conclusion that MoDe is superior is currently supported only by point estimates with overlapping uncertainties and an unmatched operating-point comparison; that conclusion needs stronger statistical and methodological support before it is suitable for publication.","major_comments":[{"comment":"The claim that MoDe is \"better suited\" is not statistically established. For the 2-body models, the fitted slopes are -11.23 +/- 1.6 ns^-1 (MoDe) and -12.9 +/- 4.7 ns^-1 (DisCo); for 3-body, -6.6 +/- 5.5 and -8.8 +/- 1.2 ns^-1. In both cases the uncertainties overlap substantially, and for 3-body the MoDe slope is consistent with both zero and the DisCo value. No test of the slope difference is provided, and the fits are independent rather than joint, so the stated ranking rests on point estimates. Please add a quantitative comparison of the slopes (e.g., a difference with propagated uncertainty, a chi-square/confidence-interval statement, or a test that accounts for the correlation between models) before claiming superiority.","section":"Sec. 5, Table 1"},{"comment":"Slopes are reported only for B0 -> D- pi+ (Fig. 4 and Table 1), while the conclusion is stated generally for the topological b trigger. The B+ -> J/psi K+ efficiencies in Fig. 3 are shown only as curves, with no fitted slopes or uncertainties, so the reader cannot judge whether the MoDe-vs-DisCo ranking holds for that decay mode. Either fit and report slopes for both decay modes, or explicitly limit the quantitative claim to B0 -> D- pi+ and provide qualitative evidence for the other sample.","section":"Sec. 5, Figs. 3 and 4"},{"comment":"The comparison uses lambda = 0.8 for DisCo and lambda = 0.2 for MoDe, selected from Fig. 5 based on precision/recall/ROC AUC rather than on decorrelation strength. Since the penalty strength controls the degree of decorrelation, differences in slopes could reflect the unmatched lambda values rather than an intrinsic difference between the methods. To support the ranking, the authors should either match the methods by an achieved decorrelation level (e.g., equal dCorr or equal large-lifetime slope) or show that the conclusion is robust to reasonable choices of lambda on both sides. As written, a different lambda for either method could plausibly change the ordering, especially for the 3-body case.","section":"Sec. 5 and Appendix B"}],"minor_comments":[{"comment":"The phrase \"Appendix Appendix A\" in Sec. 2 should read \"Appendix A\". Similarly, \"Table. 1\" in Sec. 5 should be \"Table 1\". There are also minor grammatical issues, e.g., \"the version of MoDe used not penalizing the turn-on curve\" should be rephrased.","section":"General"},{"comment":"The efficiency points are shown without uncertainty bands or error bars. Since Table 1 gives slope uncertainties, the figure would be clearer if the data points included statistical uncertainties, particularly at large lifetime where the event counts decrease.","section":"Figs. 3 and 4"},{"comment":"The structure of Table 3 is hard to read: the DisCo 3-body row appears to have no penalty strengths listed, and the MoDe 3-body row is not aligned with the 2-body row. Please reformat using separate sub-tables or explicit column entries.","section":"Table 3"},{"comment":"The text says the 3-body MoDe model has a maximum precision at lambda = 0.1 and significant degradation beyond, but then declares lambda = 0.2 optimal. This is a defensible choice, but the reasoning would be clearer if the trade-off were quantified, e.g., by stating how much each metric changes between lambda = 0.1 and 0.2.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering study with a clear empirical demonstration, but the central claim of MoDe superiority is not yet supported by the statistics as presented. The authors should be asked to strengthen the significance analysis and the comparability of the operating points; with those changes the paper could be suitable for EPJC. I do not see grounds for rejection, and I would not want the statistically fragile ranking to be published as it stands."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, the most interesting thing here is the background: the PV-misassociation component at low min(log chi2_IP) that grows with pileup. That's a real observation, well motivated, and the authors show it clearly. The paper also does a clean job of evaluating two published decorrelation penalties in a concrete trigger context. The efficiency curves with the no-decorrelation baseline show both DisCo and MoDe flatten the large-lifetime efficiency relative to the current model. That part is solid.\n\nThe problem is the headline comparison. The paper says MoDe is 'better suited,' but that rests entirely on the slopes in Table 1, and those slopes are statistically indistinguishable. For 2-body: -11.23 +/- 1.6 vs -12.9 +/- 4.7. For 3-body: -6.6 +/- 5.5 vs -8.8 +/- 1.2. The uncertainties overlap. On top of that, the lambda values are not comparable: DisCo is used at lambda=0.8, MoDe at lambda=0.2, each chosen from classification metrics rather than from matching the degree of decorrelation. You can't claim one method beats the other when the comparison is at different operating points and the measured difference is within uncertainty.\n\nAlso, slopes are only reported for B0 -> D-pi+; the B+ -> J/psi K+ efficiencies are shown but no fits are given. That's a missing piece for a quantitative claim.\n\nThe simulation-only nature is a limitation but not a fatal one. This is trigger development; simulation is the standard first step. Still, the PV-misassociation background is the load-bearing new effect, and the paper doesn't discuss validation against Run 3 data or a plan to check it.\n\nBottom line: this is a legitimate engineering study, clearly written, with a solid background identification and a plausible demonstration. The 'MoDe better' conclusion is not supported by the numbers as presented. I'd send it to a referee, but the authors should be asked to either provide matching operating points, a statistical test of the slope difference, or soften the claim. A serious referee would catch this.","headline":"Useful simulation study of lifetime decorrelation for the LHCb topological b trigger, but the claimed MoDe superiority is not statistically supported.","tokens_in":11261,"tokens_out":3007,"would_cite":false,"duration_ms":24895,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding a quadratic moment-decomposition penalty to LHCb's topological b-trigger neural networks flattens their efficiency for long-lived b-hadrons without losing classification power.","keywords":["topological b trigger","lifetime decorrelation","LHCb","HLT2","monotonic Lipschitz neural networks","pile-up background","MoDe","DisCo"],"falsifier":"Take a real Run 3 data sample where the topological b trigger did not fire, e.g., selected by a muon trigger, reconstruct B+ to J/psi K+ candidates, and measure the trigger efficiency as a function of reconstructed B+ decay time for tau>4 ps; if the fitted slope is significantly more negative than the predicted -11 ns^-1, the MoDe-lambda=0.2 claim is falsified.","tokens_in":10377,"feed_emoji":"⏱️","tokens_out":3472,"duration_ms":27978,"temperature":0.7,"pith_summary":"The paper argues that in Run 3 of the LHC, with about five visible proton-proton collisions per bunch crossing, LHCb's topological b trigger suffers from a new background: vertices made from tracks originating in different collisions that mimic long-lived b decays. This background makes the neural-network trigger less efficient for genuinely long-lived b hadrons, which would bias time-dependent measurements. The paper presents two ways to train the trigger networks to decorrelate their score from lifetime and shows in simulation that a moment-decomposition penalty (MoDe) with strength lambda=0.2 flattens the efficiency at large lifetimes better than the distance-correlation (DisCo) penalty. If correct, this keeps the trigger unbiased for particles living longer than about 4-5 ps, preserving sensitivity to mixing and CP-violation measurements.","feed_headline":"MoDe penalty flattens LHCb b-trigger efficiency at large lifetimes","feed_subtitle":"A quadratic moment-decomposition term keeps the trigger unbiased for decay-time analyses, protecting mixing and CP measurements.","key_machinery":"The MoDe penalty expands the cumulative distribution of network scores in bins of lifetime into Legendre polynomials and penalises the difference between the actual distribution and its polynomial fit; with order l=2 and monotonicity constraints it allows the sharp efficiency turn-on at small lifetimes while enforcing flatness at large lifetimes. DisCo uses distance correlation between score and lifetime as a penalty. The comparison is made by fitting linear slopes to the efficiency as a function of true lifetime beyond 4 ps.","core_discovery":"The central claim is that the decay-time bias introduced by pile-up-related track misassociation can be removed by adding a quadratic penalty based on Legendre decomposition of the classifier score distribution as a function of lifetime. Measured on simulated B+ to J/psi K+ and B0 to D- pi+ events passing HLT2, the MoDe-penalised models (lambda=0.2) reduce the negative slope of trigger efficiency versus true lifetime for tau>4 ps from about -25 ns^-1 to about -11 ns^-1 (2-body) and from about -22 ns^-1 to about -7 ns^-1 (3-body, consistent with zero), while preserving precision, recall and ROC AUC. The paper concludes MoDe is better suited than DisCo for lifetime decorrelation in the topolog","pith_inferences":["If the simulation's pile-up model is representative, the same MoDe recipe should transfer to other LHCb inclusive triggers that use the same MLNN architecture, not only the two- and three-body lines.","The method's reliance on true lifetime in training means it is best applied when signal simulation is trustworthy; an analogous data-driven version could use reconstructed decay-time sidebands to estimate the lifetime distribution.","A natural next test is to verify that the efficiency flattening survives when the trigger is run in real Run 3 conditions, where track-association algorithms may behave differently under fluctuating pile-up.","Since MoDe with l=2 explicitly enforces monotonicity at large lifetimes, it could be combined with the monotonic Lipschitz architecture to give a formal guarantee of no lifetime sculpting beyond a chosen scale."],"forward_implications":["Time-dependent CP and mixing analyses using the topological b trigger will no longer lose statistical sensitivity from a sharp drop in acceptance for b hadrons living beyond about 4 ps.","The optimal MoDe working point (lambda=0.2) can be adopted for both 2- and 3-body trigger models without retuning, as it preserves precision and ROC AUC.","The identified PV-misassociation background, prominent at high pile-up, becomes a standard systematic to check in Run 3 trigger and offline selections.","The linear-slope criterion over tau>4 ps provides a simple benchmark for future trigger decorrelation studies."],"fun_headline_variants":["MoDe penalty removes LHCb trigger lifetime bias","MoDe penalty holds LHCb trigger flat at long lifetimes","LHCb b-trigger made lifetime-flat via MoDe penalty","MoDe penalty kills LHCb trigger lifetime dependence"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything is demonstrated in simulated events; if the simulation does not reproduce the Run 3 rate of multiple simultaneous collisions and their track-association properties, the observed flattening may not hold for real data.","fun_headline_variants_meta":{"raw":{"variants":["MoDe penalty removes LHCb trigger lifetime bias","MoDe penalty holds LHCb trigger flat at long lifetimes","LHCb b-trigger made lifetime-flat via MoDe penalty","MoDe penalty kills LHCb trigger lifetime dependence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2489,"prompt_tokens":703,"completion_tokens":1786,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":1716}},"tokens_in":447,"tokens_out":1786,"duration_ms":11070,"temperature":1.0,"reasoning_tokens":1716,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:13:20.623134+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real Run 3 data sample where the topological b trigger did not fire, e.g., selected by a muon trigger, reconstruct B+ to J/psi K+ candidates, and measure the trigger efficiency as a function of reconstructed B+ decay time for tau>4 ps; if the fitted slope is significantly more negative than the predicted -11 ns^-1, the MoDe-lambda=0.2 claim is falsified.","supporting_citations":[],"review_version":1}