{"id":"4390682c-1a31-4c6b-a530-0c6a18be3879","arxiv_id":"2506.08826","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"The new nonclosure loss term, made differentiable with a sigmoid approximation, lets the neural network optimize the ABCD background estimate directly, improving closure and training stability.","lead":"The CMS collaboration presents ABCDisCoTEC, a machine-learning method that trains two neural-network outputs for the ABCD background-estimation technique by directly minimizing the method's prediction error. Applied to a stealth supersymmetry search, it yields more accurate background predictions and more stable training than the previous ABCDisCo approach.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on an unvalidated equivalence between the sigmoid-relaxed nonclosure loss (Eqs. 5-8) and the hard-boundary nonclosure actually used in validation; with no fidelity check or sensitivity study for a=100 and random boundary sampling, the mechanism's generality is unsupported.","rationale":"The reader's weakest assumption is precisely the sigmoid surrogate fidelity, and my read agrees that this is the most load-bearing point. The paper's empirical validation (Figs. 9, 12, 16) is substantial and provides real evidence that the method works on the stealth SUSY case, including validation in observed data; I do not see a reason to reject. However, the method section does not establish a formal or systematic empirical connection between the differentiable loss and the hard nonclosure metric, and the incorrect Eq. (6) suggests the nonclosure-loss mapping warrants closer scrutiny. A simple correlation and sensitivity check would settle whether the concern is real; unless that check fails, the CONDITIONAL verdict should stand.","tokens_in":41313,"tokens_out":14785,"duration_ms":180306,"concrete_test":"Using the trained ABCDisCoTEC model and the simulation validation sample, evaluate on the full boundary grid of Fig. 12 both the smooth loss L_nonclosure (Eq. 5 with Eq. 7, a=100, applying the same per-batch random boundary strategy) and the hard nonclosure C/ (Eq. 2). Compute the rank correlation over grid points, and repeat the comparison with a=10, 30, 100, 300, and 1000 while holding all other hyperparameters fixed. If the rank correlation is weak (e.g., Spearman rho below 0.7), or if the hard C/ at the final analysis boundary (0.44, 0.42) is not minimized near a=100, the sigmoid surrogate is not faithful and the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the assertion that minimizing the smooth L_nonclosure of Eq. (5), computed with the product-sigmoid weights of Eq. (7) as in Eq. (8), directly minimizes the hard-boundary nonclosure of Eq. (2). This equivalence is not established. The smooth loss is a different functional: it weights every event fractionally in all four regions, so it can be reduced by smearing events across randomly drawn boundaries rather than by making S1 and S2 independent for hard regions. The paper provides no bound linking the soft and hard nonclosure, no sensitivity scan over the scale a (the text only states that a=100 \"was found to provide the best closure performance in general\"), and no study of the distribution of random boundaries b1 and b2. Consequently, minimizing Eq. (5) during training could lower the smooth loss without reducing the hard nonclosure at the final analysis boundary, and the low hard nonclosure reported in Fig. 12, while encouraging, does not demonstrate that the surrogate is faithful in other applications. Relatedly, the claimed equivalence of Eq. (6) to Eq. (5) is algebraically wrong when N_B N_C > N_A N_D: the formula in Eq. (6) is non-monotonic for C/ > 2 and even decreases as the nonclosure grows, which further indicates that the loss mapping is not fully characterized. A direct fidelity check is needed before the central claim that the method \"directly minimizes nonclosure\" is accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ABCDisCoTEC, an extension of the ABCDisCo method for background estimation in LHC searches. The key addition is a differentiable nonclosure loss term that replaces hard ABCD region boundaries with a two-dimensional sigmoid relaxation and samples boundaries randomly per training batch. The method is applied to a stealth supersymmetry search using CMS simulation and data, and is compared with variants using only the distance-correlation loss, only the closure loss, both, and with a lambda hyperparameter scan versus the modified differential method of multipliers (MDMM). The validation uses held-out test samples, simulation control regions, and data-driven validation regions constructed inside the ABCD plane.","tokens_in":41646,"tokens_out":7434,"duration_ms":86389,"significance":"If the method works as claimed, it provides a practical way to train classifiers that satisfy the ABCD background-estimation relation across a broad range of boundaries, which is directly useful for many LHC searches. The paper includes strong validation elements: separate test samples, comparisons of loss variants, and a data-based validation-region study in Section 6. The MDMM comparison, if reproducible, is a useful contribution to multiobjective training in this context. The main limitation is that the central claim of 'directly minimizing the nonclosure' relies on a heuristic sigmoid surrogate whose fidelity is not demonstrated, and the reproducibility is limited by missing hyperparameter details. The algebraic issue in Eq. (6) also needs correction.","major_comments":[{"comment":"The statement that Eq. (6) is the nonclosure loss 'in terms of the explicit nonclosure definition' is algebraically incorrect for r = N_B N_C/(N_A N_D) > 1. When r > 1, C/ = r-1, but Eq. (6) gives ((r-1)/(3-r))^2, which differs from Eq. (5) and is non-monotonic; for example at r=2, Eq. (6) equals 1 while Eq. (5) equals 1/9. The equality holds only for r <= 1. Since Eq. (5) is the loss actually used, this does not invalidate the training, but the equivalence claim in the text should be corrected or restricted.","section":"2.2, Eq. (6)"},{"comment":"The central claim that minimizing the sigmoid-relaxed L_nonclosure directly minimizes the hard-boundary nonclosure of Eq. (2) is not established. The soft loss weights every event fractionally in all four regions and is minimized over randomly sampled boundaries, so it is a different functional from the hard-boundary nonclosure used in validation. The paper states a=100 was 'found to provide the best closure performance in general' but shows no fidelity check, no sensitivity scan over a, and no characterization of the boundary sampling distribution. I request a direct comparison: on a fixed test set, plot the soft loss against the hard nonclosure at several training checkpoints, and report the hard nonclosure for a = 10, 30, 100, 300. Without this, the phrase 'directly minimizes the nonclosure' overstates what is demonstrated.","section":"2.2, Eqs. (5)-(8)"},{"comment":"It is not specified whether L_nonclosure and L_DisCo are evaluated on background events only, on signal events only, or on the full batch. This matters because the ABCD relation in Eq. (1) is a statement about background events and does not hold for signal, and Fig. 6 shows signal deliberately concentrated in region A. If the nonclosure loss were computed on the full batch including signal, it would penalize the desired signal topology. The training description should state unambiguously which event classes enter each loss term.","section":"4.1"},{"comment":"The MDMM implementation is not described with enough detail to be reproducible or to support the claimed advantages. The paper does not quote the values or update schedules of the Lagrange multipliers alpha_i, the damping factors c_i, or the constraint targets epsilon_DisCo, nor the number of trainings used in Fig. 14. Since the MDMM comparison is a principal result highlighted in the abstract, these implementation details should be provided in a table or appendix.","section":"5.2, Eqs. (9)-(11)"}],"minor_comments":[{"comment":"The statement 'a=100 was found to provide the best closure performance in general' should be supported by a sensitivity study; as written it is an unexplained empirical choice.","section":"2.2"},{"comment":"Final training hyperparameters, including the lambda values, learning-rate schedule, number of epochs, and early-stopping criterion, are not listed; add a table with the values used for the final models.","section":"4.1"},{"comment":"The sentence 'All other hyperparameters are set to optimal values for each training configuration' is vague; specify the performance criteria used to determine optimality.","section":"4.2"},{"comment":"The quoted 3-15% systematic uncertainty from Ref. [8] is mentioned but no uncertainty band is shown in Fig. 16; indicate the band or refer the reader to the companion paper for the exact procedure.","section":"6, Fig. 16"},{"comment":"The significance formula is a rough approximation; clarify that N_bkg is the background event count and that the nonclosure term is intended to be added in quadrature as a systematic uncertainty.","section":"4.2, Eq. (13)"}],"recommendation":"major_revision","confidential_remarks":"This is a CMS collaboration paper whose physics result lives in a companion paper. The method-focused story is suitable for Machine Learning: Science and Technology if the requested fidelity checks and implementation details are added. The algebraic error in Eq. (6) is a red flag and should be fixed rather than left as a typo. I would not reject the paper, because the empirical validation is strong, but the central 'direct minimization' claim needs to be substantiated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it does what it says: ABCDisCoTEC adds a differentiable nonclosure term to the ABCDisCo loss, so training can directly target background-estimation closure rather than relying only on distance correlation. The empirical validation on the stealth SUSY case is serious: held-out test samples, simulation and data validation regions, and comparisons against DisCo-only and closure-only trainings. The reported low nonclosure across the ABCD plane, and the improved training stability with MDMM (roughly two orders of magnitude fewer iterations), are credible. Second, the main soft spot is the unreviewed link between the smooth sigmoid-relaxed nonclosure loss and the hard-boundary nonclosure used for validation. The paper states a=100 gives the best closure performance but shows no sensitivity scan, no bound, and no study of the random boundary sampling. So the central claim that the method 'directly minimizes nonclosure' is only demonstrated empirically for this single application, not established as a general method.\n\nThe genuinely new bits: the closure loss term itself (the sigmoid relaxation idea is not new, but applying it to nonclosure in this context is), and the use of MDMM to handle the multiobjective training. Both are useful and likely to be adopted by other analyses. The paper is also well-organized and candid about failure modes, which is refreshing.\n\nThe soft spots, in proportion: (1) The surrogate fidelity issue is real but not damning for this paper, because the held-out results do show low hard nonclosure. Still, a referee should ask for an ablation over a and a study of the boundary distribution. (2) Eq. (6) is simply wrong when N_B N_C > N_A N_D; the formula gives (C/(2-C))^2 unconditionally, but it should be (C/(2+C))^2 in that branch. This is a minor algebra slip that should be fixed. (3) No code/data release, which limits reproducibility, but that's common for CMS papers. (4) The MDMM section is a bit underspecified for external use; hyperparameters for the final models are not fully quoted.\n\nOverall: this deserves a serious referee. It is a solid engineering contribution to a widely used background-estimation technique, with credible validation. Recommended with requests for the sensitivity study and the Eq. (6) correction.","headline":"Useful extension of ABCDisCo that adds a differentiable closure loss and MDMM; the main idea works for the demonstrated case, but the surrogate's fidelity is under-validated.","tokens_in":42148,"tokens_out":3726,"would_cite":true,"duration_ms":42309,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ABCDisCoTEC adds a differentiable closure term to the ABCDisCo loss, so neural-network training directly minimizes the ABCD background-estimation error, yielding decorrelated discriminants with improved sensitivity in a stealth…","keywords":["ABCD method","background estimation","distance correlation","nonclosure","differentiable loss","neural network","stealth supersymmetry","LHC"],"falsifier":"Train ABCDisCoTEC on a simulated sample with known signal and background, then measure the hard-boundary nonclosure on an independent test set across a fine grid of ABCD boundary choices and compare it with the smoothed nonclosure loss evaluated at the same boundaries; a weak or reversed correlation between the two would show that the smooth surrogate is not controlling the quantity it claims to control.","tokens_in":1535,"feed_emoji":"🎯","tokens_out":1686,"duration_ms":100769,"temperature":0.7,"pith_summary":"The paper introduces ABCDisCoTEC, a training scheme that makes the ABCD background-estimation method learnable end to end. The central move is to add a differentiable approximation of the nonclosure, the mismatch between the predicted and true background counts in the signal region, to the neural-network loss, so training optimizes the quantity the method actually needs to control. In a stealth supersymmetry search using proton-proton collision data, the two learned discriminants stay strongly signal-sensitive while the nonclosure is small across most of the ABCD plane, which reduces the systematic bias in the background estimate and improves the expected sensitivity to new physics. The paper also shows that the modified differential method of multipliers makes the multi-term training more stable and reaches the same solution in far fewer iterations than manual hyperparameter tuning.","feed_headline":"Neural net trains directly on its background-estimation error","feed_subtitle":"A differentiable closure term replaces manual variable choice, cutting systematic bias in LHC searches.","key_machinery":"The load-bearing piece is the nonclosure loss term, $L_{\\mathrm{nonclosure}} = \\left(\\frac{N_A N_D - N_B N_C}{N_A N_D + N_B N_C}\\right)^2$, which measures how far the ABCD prediction $N_B N_C/N_D$ is from the observed count $N_A$. Because event counts are discrete, the paper replaces hard counting with a two-dimensional sigmoid, $\\sigma(S_1,S_2,b_1,b_2)=1/[(1+e^{-a(S_1-b_1)})(1+e^{-a(S_2-b_2)})]$ with scale $a=100$ and boundaries chosen randomly per batch, so that gradients can flow through the ABCD geometry. The full loss combines binary cross-entropy for classification, distance correlation for independence, and this closure term; the modified differential method of multipliers then treats the decorrelation and closure losses as constraints with learnable multipliers, which stabilizes the training and gives the hyperparameters a physical meaning.","core_discovery":"On the paper's own terms, the central claim is that one can train a neural network to minimize the ABCD nonclosure directly, rather than hoping that minimizing distance correlation between two outputs is enough. The nonclosure loss is $L_{\\mathrm{nonclosure}} = \\left(\\frac{N_A N_D - N_B N_C}{N_A N_D + N_B N_C}\\right)^2$, and it becomes differentiable when hard event counts are replaced by a two-dimensional sigmoid weighting with scale $a=100$ and randomly chosen boundaries per batch. Adding this term to the binary cross-entropy and distance-correlation losses yields two decorrelated discriminants with strong signal-background separation. In the paper's stealth supersymmetry case, the combined loss gives lower average nonclosure and higher normalized significance than either the distance-correlation or the closure term alone, and the resulting background estimates show good agreement between simulation and observed data. The accompanying use of MDMM turns the subordinate losses into constrained objectives with learnable multipliers, which stabilizes training and lets the analyst set physically meaningful targets such as a 10% nonclosure.","pith_inferences":["Beyond the paper: because the smooth nonclosure is only a surrogate, the method's success for a new analysis should be checked by validating the correlation between the smoothed loss and the final hard-boundary nonclosure at the chosen boundaries.","Beyond the paper: random boundary sampling during training effectively averages the closure constraint over many possible ABCD partitions, which may make the learned discriminants more uniformly decorrelated and could be studied as an implicit regularizer.","Beyond the paper: the differentiable-counting trick is not specific to high-energy physics and could be reused wherever a ratio of region counts is optimized with gradient descent, for example in anomaly detection or survey analyses.","Beyond the paper: when the Pareto front is strongly nonconvex, the advantage of MDMM over grid search should be larger than in the convex stealth-supersymmetry example, so a synthetic benchmark with a known nonconvex front would quantify the benefit."],"forward_implications":["Any analysis that relies on the ABCD method can in principle replace hand-selected independent variables with two learned discriminants trained to control nonclosure directly.","Smaller nonclosure translates into smaller systematic uncertainty in the background prediction, which directly improves discovery significance for searches where signal and background look similar.","MDMM converts loss weights into interpretable constraints, so an analyst can target a specific nonclosure value rather than scanning dimensionless hyperparameters.","The sub-ABCD validation procedure (VR I, VR II, VR III) provides a way to test the method in observed data even when no orthogonal validation region exists.","The same sigmoid-relaxation trick can be applied to extended ABCD formulations with additional control regions, as the paper notes."],"supporting_citations":[{"why":"Introduces the ABCD method for data-driven background estimation, the procedure this paper optimizes.","marker":"[1]"},{"why":"Supplies the statistical-model framework in which ABCD region counts enter a likelihood constraint.","marker":"[2]"},{"why":"Documents practical ABCD background estimation; the paper's nonclosure metric follows this prescription.","marker":"[3]"},{"why":"Defines ABCDisCo, the neural-network decorrelation method that ABCDisCoTEC extends with a closure term.","marker":"[4]"},{"why":"Provides distance correlation and its zero-implies-independence theorem, used as the decorrelation loss component.","marker":"[5]"},{"why":"Companion search presenting the stealth supersymmetry analysis details and improved limits that validate the method.","marker":"[8]"},{"why":"Introduces the modified differential method of multipliers used to stabilize multi-objective training.","marker":"[9]"},{"why":"Earlier search for the same signature with simulation-limited background estimation, the baseline this work improves on.","marker":"[15]"}],"fun_headline_variants":["Neural network minimizes background estimation error directly","New loss term trains nets to close the ABCD method gap","ABCDisCoTEC enhances background estimation with closure loss","Directly training on nonclosure reduces LHC background bias","Train on the error you care about: LHC background closure"],"cache_read_input_tokens":44288,"weakest_assumption_plain":"The load-bearing premise is that minimizing the smooth, sigmoid-weighted version of the nonclosure during training actually reduces the true discrete nonclosure at the boundaries used later; if that surrogate is unfaithful, the claimed background accuracy does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Neural network minimizes background estimation error directly","New loss term trains nets to close the ABCD method gap","ABCDisCoTEC enhances background estimation with closure loss","Directly training on nonclosure reduces LHC background bias","Train on the error you care about: LHC background closure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1495,"prompt_tokens":1056,"completion_tokens":439,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":672,"tokens_out":439,"duration_ms":5226,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:01:19.107981+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train ABCDisCoTEC on a simulated sample with known signal and background, then measure the hard-boundary nonclosure on an independent test set across a fine grid of ABCD boundary choices and compare it with the smoothed nonclosure loss evaluated at the same boundaries; a weak or reversed correlation between the two would show that the smooth surrogate is not controlling the quantity it claims to control.","supporting_citations":[{"cited_title":"Search for High Mass Top Quark Production in p anti-p Collisions at S**(1/2) = 1.8 TeV","cited_arxiv_id":"hep-ex/9411001","evidence_quote":"Introduces the ABCD method for data-driven background estimation, the procedure this paper optimizes."},{"cited_title":"Background estimation with the ABCD method featuring the TRooFit toolkit","cited_arxiv_id":null,"evidence_quote":"Documents practical ABCD background estimation; the paper's nonclosure metric follows this prescription."},{"cited_title":"Search for top squarks in final states with many light-flavor jets and 0, 1, or 2 charged leptons in proton-proton collisions at √s=13 TeV","cited_arxiv_id":null,"evidence_quote":"Companion search presenting the stealth supersymmetry analysis details and improved limits that validate the method."},{"cited_title":"Constrained differential optimization","cited_arxiv_id":null,"evidence_quote":"Introduces the modified differential method of multipliers used to stabilize multi-objective training."}],"review_version":1}