{"id":"4d6a098e-80e3-4ac0-a344-d62f8d08d247","arxiv_id":"2512.07074","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Profile OmniFold jointly unfolds particle-level distributions and profiles detector nuisance parameters via EM, recovering true spectra when the forward model is misspecified.","lead":"This paper extends the OmniFold unfolding algorithm so it can simultaneously unfold particle-level distributions and fit detector nuisance parameters. The result is a fully unbinned method for reducing bias in LHC cross-section measurements when detector simulations are imperfect.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim relies on unvalidated heuristic V for selecting among initializations; V is a detector-level fit statistic not shown to rank particle-level accuracy.","rationale":"The paper's theoretical contributions (Props. 1–3) are sound as far as they go, and the experiments are suggestive. However, the headline claim is specifically that POF 'is able to accurately estimate the true particle-level distribution' in the misspecified forward model case. In the CMS demonstration, this is only true for the run selected by V; other initializations converge to a biased θ and worse fit. Since V is acknowledged to be a heuristic, the positive result depends on an unvalidated selection step. This is the most load-bearing weakness because it directly concerns whether the reported accuracy is a property of the algorithm or of the particular seeds chosen. A concrete simulation-based check can resolve it. This does not invalidate the paper—the method may still work—but it justifies a conditional verdict rather than full acceptance. We therefore agree with the reader's identification of this as the weakest assumption.","tokens_in":24118,"tokens_out":11217,"duration_ms":108379,"concrete_test":"Using the same CMS Open Data setup with known truth (θ_true=1.7), run POF from at least 10 initializations θ^(0) spanning [0.5,2.0]. For each converged solution, compute V from Eq. (25) and a particle-level error metric (e.g., L1 or KS distance between the unfolded and true X distribution). Test whether the solution with the highest V also has the lowest particle-level error and whether V is monotonically related to error across initializations. If V fails to identify the best particle-level solution, the central claim lacks a valid selection mechanism.","verdict_should_be":"UNCHANGED","load_bearing_attack":"POF's central claim—that it accurately estimates the true particle-level distribution under a misspecified forward model—is demonstrated only after selecting among multiple runs using the goodness-of-fit statistic V (Eq. 25). The paper itself states that V 'is a heuristic statistic' and that no principled investigation of its properties has been done (Sec. 3.1). In the CMS study (Sec. 5, Fig. 10), POF converges to θ̂≈1.35 for initializations θ^(0)=1.0 and 1.1, with V<1, while the reported solution uses an initialization that yields V≈1 and θ̂=1.62. This shows the final answer depends on the selection rule. But V measures only detector-level agreement between p(y) and q̃(y); in an ill-posed inverse problem, multiple (ν,θ) pairs can give equally good detector-level fits with different particle-level distributions. No experiment or theory in the paper establishes that the highest-V solution is the most accurate in particle-level space. The fragility is compounded by the estimated-nuisance-parameter sensitivity to w (θ̂=1.42 vs 1.5 in the Gaussian case), which the paper acknowledges. Thus the headline claim is only as strong as the untested V ranking.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Profile OmniFold (POF), an extension of the classifier-based OmniFold algorithm for unbinned unfolding that jointly estimates the particle-level reweighting function ν(x) and a scalar nuisance parameter θ entering the detector response kernel. The authors formulate a population-level likelihood with a prior/penalty on θ, derive EM updates (Propositions 1–3), and implement the algorithm using neural-network density-ratio estimators. They validate POF on a two-dimensional Gaussian example with analytic and estimated w-functions, and on a CMS Open Data simulation for inclusive jet-pT spectra, comparing against the original OmniFold. The paper introduces a heuristic goodness-of-fit statistic V, Eq. (25), to select among runs with different θ initializations. It acknowledges that POF lacks convergence guarantees, that θ estimates are sensitive to the estimated w-function, and that no uncertainty quantification is provided.","tokens_in":24425,"tokens_out":4864,"duration_ms":47022,"significance":"If established, POF would fill a genuine gap: extending simulation-based unbinned unfolding to the practically important case where the forward model depends on nuisance parameters, thereby avoiding the expensive 'repeat the measurement under systematic variations' paradigm. The paper's formal EM derivation, the classifier-based construction of the w-function, and the use of public CMS simulation data are strengths, as are the clear statements of limitations. However, the central empirical claim—that POF accurately recovers the true particle-level distribution under a misspecified forward model—is presently demonstrated only after selecting among multiple initializations using an unvalidated heuristic. Because the inverse problem is ill-posed, detector-level agreement does not by itself imply particle-level accuracy, so the selection rule needs justification. The lack of uncertainty quantification further tempers the 'accurate' claim. The methodology is promising and the derivations appear sound, but the current evidence is conditional.","major_comments":[{"comment":"The central empirical claim relies on the heuristic statistic V to select among multiple initialization runs, but V is not validated as a ranking of particle-level accuracy. The paper states in Sec. 3.1 that V 'is a heuristic statistic' whose statistical properties have not been investigated. In the CMS study (Sec. 5.3, Fig. 10), initializations θ^(0)=1.0 and 1.1 converge to θ̂≈1.35 with V<1, while the reported result is the run with the highest V. Since V is based on the validation accuracy of a detector-level classifier, and since the unfolding problem is ill-posed, there is no guarantee that a high-V solution is more accurate at the particle level. Please provide either a theoretical justification or a controlled experiment (e.g., many simulated truths with varying ν and θ, checking that argmax V selects the solution closest to the truth in a particle-level metric) before the headline","section":"Sec. 3.1, Eq. (25); Sec. 5.3, Fig. 10"},{"comment":"No uncertainty quantification is provided for either the unfolded distribution or the nuisance parameter, as the paper acknowledges in Sec. 6. The observed point estimates with the estimated w-function (Gaussian θ̂=1.42 vs true 1.5; CMS θ̂=1.62 vs true 1.7) and the stated sensitivity of θ̂ to classifier training (Sec. 4.3) make it impossible to judge whether these differences are statistical fluctuations or systematic biases. Even a bootstrap or repeated-simulation variability assessment would materially strengthen the claim that POF 'accurately estimates' the true distribution. Without such quantification, the empirical support for the central claim is incomplete.","section":"Sec. 6; Sec. 4.3; Sec. 5.3"},{"comment":"The lack of convergence guarantees is a practical issue that interacts with the selection rule. The paper notes that the likelihood is not concave and that POF can converge to a local maximum (Sec. 5.3). The proposed remedy—multiple initializations plus the V heuristic—is not accompanied by any guidance on how many initializations are needed, how to choose their range, or how to detect when the V-based selection has failed. This is not a fatal flaw, but it is a load-bearing aspect of the method's applicability and should be addressed, at least empirically, in the revision.","section":"Sec. 3.1; Sec. 5.3"}],"minor_comments":[{"comment":"Please define w_i explicitly: is it the detector-level weight assigned in step 1 of the current iteration? The notation is currently ambiguous.","section":"Eq. (25)"},{"comment":"The dataset description says events are used as both 'simulation' and 'data'; clarify that the 'data' are simulated events from the same detector simulation, and that 'data' is produced by reweighting the MC sample. The distinction between the constructed θ=1.7 'truth' and the nominal θ=1.0 simulation should be stated more prominently.","section":"Sec. 5.1"},{"comment":"There are several typographical issues in the reference list, e.g., 'Physical Reivew Letters' in Andreassen et al. (2020). A careful proofread is needed.","section":"References"},{"comment":"The top/bottom plot labels ('Updated estimates θ̂' and 'Goodness-of-fit statistic') are useful; consider adding them as axis labels in the figures for clarity.","section":"Figs. 4, 7, 10"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations, which is a strength. The EM derivations are coherent and the code/data are publicly available. The main risk is that the reported CMS success depends on selecting a favorable initialization via a heuristic that has not been shown to select the best particle-level solution. I believe this is fixable within the manuscript's scope by adding validation experiments for V and/or uncertainty estimates, but the current version does not fully support the strongest claims in Sec. 6."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the POF paper. Short version: it does what it says—extends OmniFold to profile a nuisance parameter in the forward model, fully unbinned, with coherent EM machinery and honest experiments. I'd send it to peer review, expecting moderate revision.\n\nWhat's new: OmniFold assumes the detector response is exactly known, and Chan & Nachman's profiled unfolding requires binned detector-level data. POF is the first fully unbinned classifier-based EM that updates the reweighting function and the nuisance parameter jointly, using a learned response ratio w(y,x,θ)=p(y|x,θ)/q(y|x). I checked the supplement: the proofs of Props 1–3 are sound, and Prop 3 formalizes a classifier trick Chan & Nachman used without proving. The experiments earn their keep: with the analytic w, the Gaussian example recovers θ̂=1.48 vs true 1.5, and the unfolded density tracks truth while OmniFold is visibly biased. In the CMS jet-resolution study (truth θ=1.7), POF gives θ̂=1.62 and a particle-level result that matches truth. Code and data are posted openly. The paper is also upfront about its limitations—θ̂ sensitivity to w training, no uncertainty quantification, local optima—in Sections 5 and 6.\n\nSoft spots, in proportion. The main one is the selection rule. The algorithm is run from multiple θ(0) and the reported solution is whichever maximizes V, a goodness-of-fit statistic built from the step-1 classifier's validation accuracy. The paper calls V a heuristic and does not study it. In the CMS example, initializations at 1.0 and 1.1 fall into a local optimum θ̂≈1.35 with V<1, and V picks the better run—so the headline demonstration literally depends on V doing its job. The concern that V is detector-level and might not rank particle-level accuracy is fair but not fatal: detector-level fit is the only thing the observed data can constrain, and in the presented cases V selects the solution that is also right at particle level. Still, a referee should ask for a check of the selection rule against known-truth synthetic data, and for reporting of what happens without selection.\n\nSecond, θ̂ is fragile with respect to the estimated w: Gaussian θ̂=1.42 vs 1.5, and the paper notes shifts with w training range and seed. The unfolded density is more robust, which is the primary output, but the nuisance-parameter estimate being this sensitive matters for a method whose selling point is profiling.\n\nThird, there is no uncertainty quantification of any kind. The authors acknowledge this; for a measurement-oriented community it limits immediate deployment.\n\nWho this is for: HEP physicists and statisticians working on simulation-based unfolding. A solid methods paper, not a field-reshaping one. My recommendation: engage with it, send it out, and make the review center on whether V's selection is reliable enough to carry the empirical claims.","headline":"A genuinely useful extension of OmniFold that profiles nuisance parameters in unbinned unfolding; the central claim holds in the demonstrated experiments, but the heuristic initialization-selection rule (V) is the load-bearing soft spot and should be probed in review.","tokens_in":24945,"tokens_out":6030,"would_cite":true,"duration_ms":57371,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces Profile OmniFold, a machine-learning-based unfolding method that simultaneously reweights simulated events and profiles nuisance parameters in the detector response, recovering the true particle-level distribution when","keywords":["unfolding","nuisance parameters","OmniFold","Profile OmniFold","classifier-based density ratio estimation","expectation-maximization","simulation-based inference","cross section measurement"],"falsifier":"Take a simulation with a known true θ* and run POF from a grid of initializations spread across the parameter space; if the solution with the highest V has θ̂ distant from θ* (while a lower-V run is closer to θ*), the selection rule is contradicted. A sharper version: in the Gaussian example with analytic w, if the V-maximizing run ever yields an unfolded density that is further from the true density (by KS distance or L1) than a lower-V run, the central claim that the selected POF solution recovers the truth fails.","tokens_in":23989,"feed_emoji":"⚛️","tokens_out":6588,"duration_ms":58981,"temperature":0.7,"pith_summary":"The paper introduces Profile OmniFold (POF), an extension of the OmniFold algorithm that performs unbinned, simulation-based unfolding of detector-smeared data while simultaneously profiling nuisance parameters in the forward model. POF iteratively reweights simulated particle-level events to match experimental data and updates the nuisance parameters by maximizing the population log-likelihood, using classifiers to estimate all required density ratios. The authors demonstrate, on a Gaussian example and on simulated CMS jet data, that when the detector simulation's response kernel is misspecified, POF recovers the true particle-level distribution, while standard OmniFold yields biased results. The paper also highlights a practical caveat: POF can converge to local maxima depending on the initialization of the nuisance parameter, so the authors propose a classifier-based goodness-of-fit statistic V to select among multiple initializations.","feed_headline":"Profile OmniFold corrects unfolding for detector mismodeling","feed_subtitle":"When the detector simulation is wrong, OmniFold is biased; Profile OmniFold recovers the true spectrum.","key_machinery":"The central object is the conditional density ratio w(y,x,θ)=p(y|x,θ)/q(y|x), which reweights the Monte Carlo response kernel to account for nuisance parameters. POF is an EM algorithm whose Q-function separates into a ν-dependent term and a θ-dependent term, allowing the particle-level reweighting and the nuisance-parameter update to be performed independently each iteration. Density ratios for both the detector-level reweighting and for w itself are estimated with binary classifiers (Proposition 3 factors w into two classifier ratios). To pick among multiple initializations, the method uses the heuristic statistic V, derived from the step-1 classifier's validation accuracy, which is near 1","core_discovery":"Profile OmniFold treats the unfolding inverse problem as a joint maximum-likelihood estimation of the particle-level reweighting function ν(x) and the nuisance parameter θ in the forward model p(y|x,θ). At each EM iteration, it (1) reweights detector-level simulation to match data, (2) pulls the ratio back to particle level, and (3) updates θ by maximizing a Q-function term that depends on the learned conditional density ratio w(y,x,θ)=p(y|x,θ)/q(y|x). The authors prove (Proposition 2) that the ν and θ updates separate, and (Proposition 3) that w can be estimated as a product of two classifier-based density ratios. In experiments, POF with correctly estimated w recovers the true particle-lev","pith_inferences":["If V consistently ranks solutions, it could also serve as a model-misspecification diagnostic for standard OmniFold, flagging runs where the reweighted detector-level distribution never matches data.","The classifier-based factorization of w suggests a block-coordinate extension to multiple nuisance parameters; the scalar demonstration leaves that as a natural next step.","The observed sensitivity of θ̂ to classifier training suggests that ensembles over w, rather than a single fit, may be needed for reliable profiling; the paper notes sensitivity but does not quantify it.","POF could be paired with profile-likelihood or bootstrap techniques to attach uncertainties to both the unfolded density and the nuisance parameter, closing the current gap in uncertainty quantification."],"forward_implications":["With a correctly estimated w, POF recovers the true particle-level distribution in the Gaussian and CMS studies, while OmniFold’s solution is visibly biased.","POF performs unbinned profiling at both detector and particle level, unlike the earlier profiled unfolding approach that required binned detector-level data.","Because POF preserves OmniFold’s classifier-based density-ratio estimation, it can be implemented as a drop-in extension of existing OmniFold software.","The reported sensitivity to initialization implies that users must run multiple starting values and use the V-statistic to select the solution.","The EM separation result (Proposition 2) provides a foundation for adding nuisance-parameter profiling to other simulation-based unfolding methods."],"fun_headline_variants":["Profile OmniFold: unfolding that corrects for detector mismodeling","Unfolding with nuisance parameters: Profile OmniFold","Correcting detector mismodeling with Profile OmniFold","Profile OmniFold: joint unfolding and nuisance parameter estimation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The algorithm's success rests on the untested heuristic that the V-statistic—the step-1 classifier's weighted accuracy—identifies the global maximum of the likelihood; the CMS results show V fails to flag a bad local optimum when initialization is poor, so a scenario with only poor initializations would break the method.","fun_headline_variants_meta":{"raw":{"variants":["Profile OmniFold: unfolding that corrects for detector mismodeling","Unfolding with nuisance parameters: Profile OmniFold","Correcting detector mismodeling with Profile OmniFold","Profile OmniFold: joint unfolding and nuisance parameter estimation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000571,"raw_usage":{"total_tokens":2518,"prompt_tokens":708,"completion_tokens":1810,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":1740}},"tokens_in":452,"tokens_out":1810,"duration_ms":11337,"temperature":1.0,"reasoning_tokens":1740,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:00:08.374998+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a simulation with a known true θ* and run POF from a grid of initializations spread across the parameter space; if the solution with the highest V has θ̂ distant from θ* (while a lower-V run is closer to θ*), the selection rule is contradicted. A sharper version: in the Gaussian example with analytic w, if the V-maximizing run ever yields an unfolded density that is further from the true density (by KS distance or L1) than a lower-V run, the central claim that the selected POF solution recovers the truth fails.","supporting_citations":[],"review_version":1}