{"id":"a6e0104c-2134-4888-bde2-64ad73d314b3","arxiv_id":"2602.17300","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An analytical likelihood based on error-function smearing tensors extracts per-category lepton and photon energy scale and resolution corrections from Z decays without random convolution.","lead":"This paper introduces an analytical likelihood method to extract lepton and photon energy scale and resolution corrections from Z-boson decays, replacing slow random-smearing convolution with exact error-function integrals and automatic differentiation. It is a practical calibration tool for LHC experiments, claiming large speedups and validated on toy and Pythia simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 9 propagates per-lepton Gaussian smearing to m_ll via an unstated first-order delta rule; 'exact analytical' unbiasedness is not established beyond the 1.5% validation, so larger-resolution or pT-cut regimes are unproven.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the weakest assumption is indeed Eq. 9: that equation is the bridge from the per-lepton Gaussian smearing model to the fully analytical likelihood. If its first-order propagation error becomes sizable, both parameter recovery and the Hessian-based covariance are biased. The existing validations are all in the small-sigma regime, so the paper's central claim of unbiased recovery and accurate uncertainties is only established there, not for the larger-resolution or hard-cut cases that the method is likely to encounter. The separate dimension inconsistency in Eq. 15 (it omits the S normalization from Eq. 12) reinforces the need for a corrected derivation before the MC-uncertainty claims are taken as published. However, these are correctable issues rather than a refutation of the method, so no change to the reader's conditional verdict is needed; the proposed test would either corroborate the approximation and justify moving toward acceptance, or force a correction of Eq. 9 and the uncertainty propagation.","tokens_in":13362,"tokens_out":15161,"duration_ms":135319,"concrete_test":"Run the Pythia-based validation with injected per-lepton smearing sigma_b = 5% and 10% (ideally using the same relative-pT categorization) and compare fitted r_b and sigma_b with injected values, including pull distributions. If the fitted bias exceeds statistical uncertainty or the pull widths deviate from unity, Eq. 9's Gaussian approximation is implicated. As a direct analytical check, numerically integrate the exact conditional density of m_ll under Eq. 4 (product of two independent normals) and compare it with Eq. 9's Gaussian for sigma = 2%, 5%, and 10%; report the maximum discrepancy in the resulting alpha tensor. This would define the validity domain of the method in a single quantitative statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central construction uses Eq. 9, which asserts M_ll(m; r_ll, sigma_ll) = N(m; r_ll m_mc, r_ll sigma_ll m_mc), with Eq. 6 defining r_ll and sigma_ll. But from Eq. 1, m_ll is proportional to sqrt(pT1 pT2). Under Eq. 4, each pT is an independent Gaussian; the conditional distribution of m_ll is therefore the distribution of the square root of a product of two independent Gaussians, which is not Gaussian. It has a nonzero second-order mean shift (approximately -sigma^2/4 times m) and skewness of order sigma^3. Eq. 6 and Eq. 9 together amount to a first-order delta-rule approximation, but the paper nowhere states that this is an approximation or quantifies its error. Because the alpha tensor (Eq. 11) and the likelihood (Eq. 12) inherit this assumption, any bias from the approximation propagates directly into the extracted r_b and sigma_b and into the Hessian covariance. The Pythia validation uses sigma_sim = 1.5%, where the induced bias is at or below 0.1%, so it does not probe the failure mode. For categories with larger resolution (e.g., high-|eta| electrons, TeV muons) or with hard pT thresholds (the pT > 25 GeV first bin, photon pT cuts), non-Gaussian tails and category migration are amplified, and the claimed 'unbiased parameter recovery and accurate uncertainty estimates' have no support. The paper's own Sec. 5.4 documents analogous resolution-induced scale biases in the photon case, confirming that this mechanism is real rather than hypothetical.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an analytical likelihood method (IJazZ2.0) for extracting per-category lepton energy scale and resolution corrections from Z→ℓℓ decays, replacing random-smearing convolution with a fully differentiable error-function-based likelihood. The central formulas are the per-lepton Gaussian smearing model (Eq. 4), the combination rules for dilepton scale and resolution (Eq. 6), the Gaussian form of the smeared invariant mass (Eq. 9), and the tensor α_ijc that maps finely binned simulation to data bins (Eq. 11). The method is validated with a toy Breit–Wigner MC and a realistic Pythia Z→ℓℓ sample, showing unbiased parameter recovery at σ_sim=1.5%. A relative-pT categorization is introduced to mitigate migration biases. The approach is extended to photon energy calibration using Z→µµγ decays via a new observable V_DY (Eq. 24), with an iterative fitting procedure and a documented set of biases for large resolutions and pT cuts. The paper also derives an analytical formula for MC-statistical uncertainties (Eqs. 15–19) and reports large CPU gains from automatic differentiation.","tokens_in":13924,"tokens_out":9722,"duration_ms":81427,"significance":"If the derivation were exact as claimed, the method would be a valuable, fast, and numerically stable calibration tool for LHC precision measurements. The paper has concrete strengths: it ships reproducible software on PyPI, demonstrates the CPU gain quantitatively, validates the likelihood against random-smearing, and explicitly documents some limitations in the photon case (Sec. 5.4). The analytical treatment of finite-MC-statistics uncertainties is an ambitious and useful contribution. However, the central Gaussian propagation step (Eq. 9) is only a first-order approximation, and the exactness claim is not supported; the MC-uncertainty derivative (Eq. 15) appears to be missing a normalization factor. These issues affect the core claim of unbiased parameter recovery and accurate uncertainties. With corrections and additional validation, the method could become a solid contribution.","major_comments":[{"comment":"Eq. (9) is presented as a consequence of Eqs. (4) and (6), but it is not exact. From Eq. (1), m_ll² ∝ p_T1 p_T2; under independent Gaussian smearing of each p_T (Eq. 4), the conditional distribution of m_ll is that of a constant times the square root of a product of two Gaussians, which is not Gaussian. To second order in σ, the mean is shifted by ≈ −m_mc sqrt(r1 r2)(σ1²+σ2²)/8 and the distribution is skewed. Eq. (9) is therefore a first-order delta-rule approximation that is never stated or quantified. Since the α tensor (Eq. 11) and the likelihood (Eq. 12) inherit this Gaussian form, the extracted r_b and σ_b and their Hessian uncertainties carry the neglected bias. The Pythia validation uses σ_sim=1.5%, where the bias is ~5×10⁻⁵ relative and invisible; it does not test larger resolutions (high-|η| electrons, high-pT muons). Please either prove the Gaussian form or explicitly introduce","section":"§2.3, Eq. (9)"},{"comment":"The derivative of the NLL with respect to the MC bin content m_jc is missing a normalization factor. With p_ic = A_ic/S_c, A_ic = Σ_j α_ijc m_jc, S_c = Σ_i A_ic, the correct derivative is δnll/δm_jc = (1/S_c)[ n_c Σ_i α_ijc − Σ_i n_ic α_ijc / p_ic ]. Eq. (15) as written lacks the 1/S_c factor; the stated definitions n_c ≡ Σ_i n_ic and p_c ≡ Σ_i p_ic (=1) do not repair the dimensional inconsistency. Because Eq. (19) and the validation in Fig. 4 build on this derivative, please correct Eq. (15), verify against the code, and confirm that Fig. 4 was produced with the corrected formula.","section":"§3, Eq. (15)"},{"comment":"The claim of unbiased parameter recovery is qualified only by the first-pT-bin bias, but the relative-pT recasting procedure itself is heuristic: the absolute-pT bin edges are reconstructed from average relative-pT values, and the two-step fit fixes scale before extracting smearing. The paper asserts this 'closely reproduces' the original categories without a quantitative comparison. Given that the method is intended for differential calibration, please provide a quantitative check of the recast bin accuracy (e.g., bin-by-bin differences in injected parameters or in fitted parameters versus the original absolute-pT categorization).","section":"§4, Fig. 7"}],"minor_comments":[{"comment":"The phrase 'exact analytical treatment' is too strong. Consider wording such as 'analytic treatment within a Gaussian-smearing approximation'.","section":"Abstract and Introduction"},{"comment":"The notation for the resolution combination is garbled: 'σℓℓc =0.5× q σℓ2 b1 +σ ℓ2 b2' should be typeset as σℓℓc = 0.5√(σℓ_b1² + σℓ_b2²).","section":"Eq. (6)"},{"comment":"The 'naive bias prediction' is described only verbally. Please provide the exact formula used to produce the dashed curve so the reader can reproduce it.","section":"Sec. 5.4, Fig. 10"},{"comment":"The statement that the second fit converges to the same smearing parameters as the first 'because no correlations are introduced' in the absolute-pT case could be expanded: a short explanation of why relative-pT introduces correlations would help.","section":"Sec. 4, second fit"},{"comment":"Reference [3] is the software repository; it would be useful to mention the version number and a DOI if available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The validation in Fig. 4 suggests that the authors may have implemented the correct normalization in the code, despite the paper's Eq. (15); I recommend asking them to confirm this during revision. The more serious issue is the first-order nature of Eq. (9), which underlies the 'exact' claim. I do not see circularity or a fundamental flaw in the idea; the paper needs to either prove the Gaussian form or explicitly scope the approximation and validate it at larger resolutions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Useful tool, and the central idea is sound for the regime it targets, but the paper oversells its exactness. The analytical likelihood with the error-function smearing tensor, automatic differentiation, and the relative-pT strategy are real contributions. The V_DY variable for the photon scale is clever. Shipping the PyPI code and validating with toy and Pythia simulations is good practice, and unbiased recovery at 1.5% resolution is demonstrated.\n\nThe main soft spot is Eq. 9. The paper says it follows from Eqs. 4 and 6, but m_ll is proportional to the square root of the product of two pT values. Under independent Gaussian per-lepton smearing, that distribution is not Gaussian; Eq. 9 is a first-order delta-rule approximation. The paper never states that or quantifies the error. At 1.5% resolution it is fine, but for larger resolutions or hard pT thresholds the bias can become visible, and the photon section's own Sec. 5.4 shows the same mechanism producing scale biases. So the 'exact analytical treatment' phrase in the abstract is too strong. The authors should state the approximation explicitly and bound its effect.\n\nEq. 15 as printed is dimensionally wrong. The first-order derivative of the nll with respect to m_jc should carry an overall 1/S_c factor, where S_c is the total expected count in the category. The toy validation in Fig. 4 suggests the code is doing the right thing, so this is likely a typo in the write-up, but as written it will mislead anyone trying to implement or audit the formula. That needs a correction.\n\nThe photon extension is the most candid part: the authors document resolution-induced biases and pT-selection biases and explain what an analyst should do. That is good, but the abstract and conclusion still claim 'unbiased parameter recovery and accurate uncertainty estimates' without these caveats. The residual biases should either be corrected in the software or be carried as explicit systematic uncertainties.\n\nOverall, this is a solid practical contribution: not a new physics result, but a tool that LHC calibration people will likely use. The approximations and the outlier formula are addressable in revision. I would send it to peer review rather than desk reject it.","headline":"Useful calibration tool, but the 'exact' claim is overstated: Eq. 9 is a first-order approximation and Eq. 15 as printed is dimensionally wrong; still worth a serious referee.","tokens_in":14341,"tokens_out":5599,"would_cite":true,"duration_ms":48695,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that an analytical, error-function-based likelihood can replace random smearing in Z->ll calibration, recovering scale and resolution parameters with unbiased estimates and large CPU gains.","keywords":["lepton energy scale","resolution smearing","analytical likelihood","Z->ll calibration","error-function convolution","relative pT categorisation","photon energy scale","automatic differentiation"],"falsifier":"A concrete check: generate Z->ll events with a much larger injected resolution (for example sigma_l = 5-10%) and fit the parameters; if the recovered scale or resolution drifts beyond the reported uncertainty, the Gaussian approximation of Eq. (9) is falsified. Alternatively, compute the exact distribution of m_ll from the product of two independent Gaussian-smearing factors and compare it to the assumed Gaussian shape; any visible deviation at the resolution values used in real experiments would bound the method's validity.","tokens_in":13251,"feed_emoji":"⚛️","tokens_out":4240,"duration_ms":38220,"temperature":0.7,"pith_summary":"The paper aims to show that the standard calibration step of matching simulated Z->ll invariant-mass shapes to data does not need random-number smearing of each simulated event. Instead, it builds an exact analytical expression, based on error functions, for the probability that a smeared event lands in a given mass bin, making the likelihood smooth and differentiable. This lets calibration parameters—per-category energy scale and resolution—be fitted with automatic differentiation, giving a large speed gain and stable minimization. The authors validate that fitted parameters recover injected values in toy and realistic simulations, and they extend the same likelihood to photon energy scale using Z->mumu-gamma decays. If the method holds, it gives experiments a fast, unified calibration tool for leptons and photons.","feed_headline":"One error-function tensor replaces random-smearing in Z->ll fits","feed_subtitle":"A fully differentiable likelihood pulls scale and resolution from Z->ll and Z->mumu-gamma events, cutting fit time from days to minutes.","key_machinery":"The key object is the tensor alpha_ijc (Eq. 11): for each di-lepton category c, the probability that a simulated event in fine bin j migrates to data bin i after applying scale r and resolution sigma, expressed with error functions. It turns the binned likelihood into a smooth, differentiable function of r and sigma, so automatic differentiation can compute gradients exactly. For photons, the observable V_DY (Eq. 24) replaces the invariant mass, making the same machinery applicable to Z->mumu-gamma events. The relative-pT variable r_pT = pT/m_ll is the workaround that removes category-migration biases when calibration categories depend on lepton transverse momentum.","core_discovery":"The central claim is that Eqs. (9) and (11), which model each simulated di-lepton event's smeared mass as a Gaussian with mean r_ll times the Monte Carlo mass and width r_ll sigma_ll times that same mass, and which bin this Gaussian via error functions, yield a likelihood whose fitted parameters are unbiased and whose covariance is accurate. The error-function tensor alpha_ijc turns the binned convolution into a dense but differentiable operation, so gradients can be computed exactly. A relative-pT variable r_pT = pT/m_ll is introduced to avoid biases when categories depend on lepton pT, and a two-step fit (scale first, then resolution) recovers injected parameters in realistic Drell-Yan sim","pith_inferences":["The Gaussian form of Eq. (9) is a first-order approximation; if the lepton pT resolution is large, the exact product-of-Gaussians distribution will deviate, so the claimed unbiasedness likely has a regime bound in resolution fraction.","The V_DY bias for photon resolution smearing above ~2% (Sections 5.4.1-5.4.2) suggests a two-pass strategy: fit resolution from Z->ee events, then smear the simulation and re-fit the photon scale, rather than extracting both simultaneously.","The method's speed opens the possibility of per-event calibration as a function of many variables at once, something that was previously prohibitive with random smearing.","A natural stress test is to run the same likelihood with a non-Gaussian smearing model (e.g., a double-sided Crystal Ball); the error-function integral approach may be generalizable beyond Gaussian shapes."],"forward_implications":["Lepton energy scale and resolution can be simultaneously extracted per category with exact gradients, making large-scale multi-parameter calibrations (100+ parameters, 20M events) run in minutes instead of days.","Finite-MC-statistics uncertainties are propagated to the fitted parameters through an analytic linear-response formula, so the fit reports correct total uncertainties even when simulation size is limiting.","pT-dependent calibrations become usable without bias by rebinning in r_pT rather than absolute pT, and correcting scale before measuring resolution.","Photon energy scale can be measured in Z->mumu-gamma events using the same likelihood, with a small iterative correction for the V_DY approximation.","The same analytical machinery can run on any framework supporting automatic differentiation, including GPUs, making it well suited for the precision-calibration environment of a large collider experiment."],"fun_headline_variants":["Exact Gaussian smearing fit for lepton scale and resolution","Error-function tensor enables fast, unbiased Z->ll calibrations","Relative-pT removes category bias in lepton scale fits","Fully differentiable likelihood cuts Z->ll calibration time","New fit method applies to Z->mumu-gamma photon energies"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole fit rests on approximating the smeared di-lepton invariant mass as Gaussian with mean r_ll times the Monte Carlo mass and width r_ll sigma_ll times that mass (Eq. 9); because m_ll is a product of two Gaussian-smeared pT values, this is only exact in a small-resolution limit, and the paper does not quantify where it breaks.","fun_headline_variants_meta":{"raw":{"variants":["Exact Gaussian smearing fit for lepton scale and resolution","Error-function tensor enables fast, unbiased Z->ll calibrations","Relative-pT removes category bias in lepton scale fits","Fully differentiable likelihood cuts Z->ll calibration time","New fit method applies to Z->mumu-gamma photon energies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000668,"raw_usage":{"total_tokens":2897,"prompt_tokens":774,"completion_tokens":2123,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":2039}},"tokens_in":518,"tokens_out":2123,"duration_ms":15984,"temperature":1.0,"reasoning_tokens":2039,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T22:15:52.326726+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: generate Z->ll events with a much larger injected resolution (for example sigma_l = 5-10%) and fit the parameters; if the recovered scale or resolution drifts beyond the reported uncertainty, the Gaussian approximation of Eq. (9) is falsified. Alternatively, compute the exact distribution of m_ll from the product of two independent Gaussian-smearing factors and compare it to the assumed Gaussian shape; any visible deviation at the resolution values used in real experiments would bound the method's validity.","supporting_citations":[],"review_version":1}