{"id":"88a9a45a-ae55-4681-962f-3723547d3929","arxiv_id":"2508.20541","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A transformer-based machine-learning particle-flow algorithm integrated in CMS software gives jet and missing-transverse-momentum performance similar to the standard particle-flow algorithm while running about twice as fast.","lead":"CMS physicists trained a transformer-based machine-learning model to replace the standard particle-flow reconstruction algorithm, and report similar jet and missing-energy performance at about twice the speed. The paper documents the model's integration into CMS software and first checks against proton-proton collision data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'validated on data' claim rests on two observables largely insensitive to a common jet energy scale shift, so MC-to-data transfer is the load-bearing unverified step.","rationale":"The reader's weakest_assumption already identifies the simulation-to-data transfer as the critical risk, and my reading agrees. The simulation-only performance comparisons (Figs. 2 and 3) are internally reasonable and give some support for physics equivalence, but the final claim is about use in real CMS data. The Section 5 data check is too coarse: the chosen observables are not sensitive to a common jet energy scale offset, which is exactly the kind of transfer failure one would worry about after training at a different center-of-mass energy and detector conditions. This is not an accusation of misconduct; it is a standard closure requirement in CMS jet physics. The runtime claim, while comparing CPU PF to GPU MLPF, is at least clearly stated and could be investigated separately; it does not undermine the physics claim as directly. Therefore the conditional verdict should stand, pending a proper data-side JES/JER closure test. I would not reject the paper, because the architecture, integration, and simulation validation are real evidence, and the missing piece is a well-defined validation step that the collaboration can perform.","tokens_in":5977,"tokens_out":5237,"duration_ms":54328,"concrete_test":"Run MLPF and PF-PUPPI jets through the standard CMS JetMET calibration chain on a 2024 run-condition data sample (dijet and Z+jet or gamma+jet) and derive the residual jet energy scale and resolution for MLPF relative to PF-PUPPI as a function of pT and eta. Require the MLPF residual JES to be consistent with the PF JES within the PF systematic uncertainty for pT > 30 GeV and |eta| < 2.5, and compare the MLPF and PF pTmiss response in Z->mumu events. If MLPF requires materially different jet energy corrections or has significantly larger residual JER, the 'comparable physics performance' claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Summary claims MLPF shows 'comparable physics performance, including PU mitigation' and is 'validated on simulation and data.' The weakest load-bearing step is the data validation in Section 5. The model is trained on 14 TeV 2023-condition Monte Carlo (tt, QCD, Z->tautau, PU 55-75), then judged on 13.6 TeV 2024 dijet data using only pTmiss and dijet asymmetry shape agreement (Fig. 4). These observables are weakly sensitive to the failure mode that matters most for replacing PF: an overall jet energy scale shift. Dijet asymmetry is a ratio of the two leading jet pT values, so a common scale factor largely cancels; pTmiss in a back-to-back dijet event is dominated by resolution and soft radiation, not by the absolute jet response. No jet energy scale or resolution closure in data, no per-bin uncertainties, and no quantitative comparison of MLPF jets to a reference object (e.g., Z boson, photon) are presented. If the simulation-to-data transfer fails because of 2024 detector conditions, alignment changes, or the 14 vs 13.6 TeV center-of-mass difference, MLPF could carry an O(1-2%) jet pT offset while the two shown distributions still look compatible. The central claim overstates what has been demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings paper reports on the integration of a machine-learned particle-flow (MLPF) algorithm into the CMS software framework (CMSSW). The model is a transformer-based network with FlashAttention trained on 14 TeV Monte Carlo samples with pileup (tt, QCD, Z->tautau) to reconstruct particle candidates from tracks and calorimeter clusters. The paper describes the training target, loss function, input features, simulation studies comparing MLPF to the standard PF+PUPPI reconstruction, a commissioning study on 13.6 TeV dijet data, and a runtime benchmark showing a ~2x speedup over the CPU-based PF algorithm. The central claim is that MLPF demonstrates comparable physics performance, including pileup mitigation, while reducing runtime by a factor of two, and is validated on both simulation and data.","tokens_in":6212,"tokens_out":3877,"duration_ms":38524,"significance":"The result is significant if the claims hold: replacing the rule-based PF step with an ML model at comparable physics quality and lower runtime would have practical implications for CMS reconstruction, particularly for the HL-LHC computing challenges. The paper makes concrete contributions by describing a CMSSW integration, a realistic training setup, and a benchmark with specific runtime numbers (40 ms/event vs 90 ms/event). It also provides falsifiable performance claims in simulation. However, the physics-performance claim is currently supported only qualitatively, with no quantitative resolution or response numbers and only a weak data validation. Strengths include the explicit runtime benchmark, the architectural details, and the single-particle efficiency/fake-rate comparisons; these are concrete and reproducible elements. The data validation, in contrast, does not yet substantiate the 'validated on data' wording.","major_comments":[{"comment":"The data validation does not support the claim of 'validated on simulation and data' or 'comparable physics performance' in data. The two observables shown, pTmiss and dijet asymmetry, are largely insensitive to a common jet energy scale offset: the dijet asymmetry is a ratio of the two leading jet pT values, so a global scale factor cancels, and pTmiss in a back-to-back dijet event is dominated by resolution and soft radiation rather than by the absolute jet response. No jet energy scale or resolution closure in data, no per-bin uncertainties, and no comparison against a reference object (e.g., Z boson or isolated photon) are provided. If the simulation-to-data transfer fails because of 2024 detector conditions or the 14 TeV vs 13.6 TeV difference, an O(1-2%) jet pT offset would be invisible in these distributions. Please add a quantitative closure test in data, or explicitly revise the claim to state that the data show distribution-level agreement within the statistical precision of this small sample.","section":"Section 5, Fig. 4"},{"comment":"The statement that jet and pTmiss performance are 'comparable to PF' is supported only by qualitative ratio plots; no numerical values for jet response (mean), resolution (width), or pTmiss resolution are given, and the plots contain no uncertainty bands. For a claim that MLPF can replace a core reconstruction algorithm, quantitative metrics with uncertainties are required, at least in the form of a table summarizing the response and resolution for representative pT bins. Additionally, the caption 'We show uncorrected jet pT resolution' is not descriptive of the plotted content, which are response distributions.","section":"Section 4, Fig. 3"},{"comment":"The sqrt(pT) loss weight is an ad hoc tuning parameter that was chosen after observing jet performance ('We have observed that this weight term improves the jet energy scale and resolution performance'). This creates a circularity with the later claim in Section 4 that jet and pTmiss performance are 'not explicitly optimized for during training.' The paper should state how this weight was selected (e.g., on a dedicated validation sample), and acknowledge that the jet-level performance figures incorporate this tuning choice. As written, the choice of the weight is a free parameter that is not specified or justified beyond the observed improvement.","section":"Section 3.2"},{"comment":"The model is trained on 14 TeV simulation with 2023 detector conditions, but the data validation uses 13.6 TeV data from 2024. This domain shift is not addressed. While the center-of-mass difference is small, the change in detector conditions and alignment could affect the input feature distributions. The paper should discuss why this transfer is expected to be valid, or include a simulation-level cross-check at 13.6 TeV with 2024 conditions to demonstrate that the model retains its performance under the changed conditions.","section":"Sections 2 and 5"}],"minor_comments":[{"comment":"The text uses the Unicode ligature 'dĳet' in several places; this should be written as 'dijet'.","section":"Throughout"},{"comment":"The training description states 'eight epochs' and 'approximately 260 hours' on a single A100; please also state the number of training events or steps per epoch to make the training cost reproducible.","section":"Section 3.2"},{"comment":"The single-particle efficiency and fake-rate plot (Fig. 2, right) shows only photons and neutral hadrons, and the 'Reco/Gen' ratio plots have no statistical error bars. Adding uncertainties or bin counts would improve the assessment of the claimed agreement.","section":"Section 4, Fig. 2"},{"comment":"In the pTmiss ratio plots, the right-hand panel appears to use a y-axis offset (-50 to 50) that is not explained in the caption; please clarify what is being plotted.","section":"Section 4, Fig. 3"},{"comment":"The definition of lambda as lambda = pi/2 - theta is non-standard without context; consider stating that this is the complement of the polar angle and why it is used as an input feature.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings-style contribution, and the bar for quantitative support in a journal referee report may be higher than for a typical EPS-HEP proceeding. The central claim of 'comparable physics performance' is plausible but not yet demonstrated with the required rigor: the simulation comparisons lack numerical metrics and uncertainties, and the data validation uses observables that are insensitive to the most relevant failure mode (a common jet energy scale shift). The runtime and integration results are concrete and well described. A major revision that adds quantitative jet response/resolution numbers, statistical uncertainties, and a data closure test, or that appropriately tempers the claims, would bring the paper to a publishable standard."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a credible proceedings status report from the CMS MLPF program. The genuinely new pieces are the FlashAttention transformer variant, the production CMSSW integration via ONNX, the runtime benchmark (40 ms/event on a shared A100 vs 90 ms for PF on CPU), and the first commissioning on 375 pb-1 of 2024 dijet data. Those are real, concrete steps toward making ML-based PF a practical option for HL-LHC.\n\nThe paper is honest about its own limitations. Section 3.1 explicitly describes the unavoidable smearing between Pythia truth and the training target, and the PU masking is clearly explained. The loss function and the sqrt(pT) weighting are presented with the admission that the weight was chosen after observing jet performance—a minor tuning circularity, but disclosed. The runtime comparison uses six parallel jobs per A100 with 8 CPU threads each, which is a realistic production setup.\n\nThe soft spot is the data validation. The Summary says the model is 'validated on simulation and data,' but the data section contains only shape comparisons of pTmiss and dijet asymmetry, with no error bars, no closure test, and no reference-object measurement like Z or photon balance. The stress-test concern is fair: dijet asymmetry is a ratio of the two leading jet pT values, so a common jet energy scale shift cancels; pTmiss in back-to-back dijets is dominated by resolution and soft radiation, not the absolute jet response. Training was done at 14 TeV with 2023 conditions and flat PU 55-75, while the data are 13.6 TeV from 2024 with a different PU profile. If MC-to-data transfer fails, MLPF could carry an O(1-2%) jet pT offset while both shown distributions still look compatible. So the 'comparable physics performance' claim is plausible but not fully pinned down.\n\nFor a proceedings paper this is probably acceptable—it's a status report, not a final physics measurement. But the summary overstates what has been demonstrated. The paper would benefit from either quantitative data validation (JES/JER closure or Z/gamma balance) or a more careful wording like 'no significant discrepancy observed in the commissioning sample.'\n\nI'd send this to a serious referee for a conference proceedings; it's honest, concrete, and the runtime and integration details are worth checking. I'd also cite it if I were working on ML-based reconstruction. It is not a major new idea, but it is a legitimate, useful incremental step.\n\nBest","headline":"A concrete, honest proceedings report on MLPF's transformer variant and CMSSW integration; the physics claim is plausible but the data validation is too thin to fully support 'comparable performance' yet.","tokens_in":6739,"tokens_out":4631,"would_cite":true,"duration_ms":39549,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a machine-learned transformer called MLPF can replace the rule-based particle-flow step in CMS event reconstruction with comparable physics performance—including pileup mitigation—while running about twice as fast.","keywords":["particle flow","machine learning","transformer","CMS","pileup mitigation","jet reconstruction","missing transverse momentum","GPU inference"],"falsifier":"A decisive test would be an independent measurement of the MLPF jet energy response in data, for instance using Z+jet or photon+jet transverse-momentum balance, or a simulation closure test at 2024 conditions with a pileup profile outside the flat 55-75 training range. If the measured response or the dijet asymmetry deviates from the PF-PUPPI result by more than the paper's plots show, the claim of comparable physics performance on data would not be supported.","tokens_in":5764,"feed_emoji":"⚛️","tokens_out":9144,"duration_ms":80718,"temperature":0.7,"pith_summary":"This paper presents a machine-learned particle-flow algorithm, MLPF, as a single learned replacement for the first step of CMS event reconstruction. Instead of hand-written rules that link tracks to calorimeter clusters, a transformer takes all tracks and clusters and directly outputs particle candidates with identity, momentum, and a per-particle pileup score. The authors claim that in simulated top-quark and QCD events with up to 75 pileup interactions, jets and missing transverse momentum match the standard PF+PUPPI chain, and that the same model agrees with PF on a 375 $pb^{-1}$ sample of real 13.6 TeV dijet data. The payoff, if true, is a drop-in module that needs no changes downstream and cuts reconstruction time from about 90 ms to 40 ms per event.","feed_headline":"Transformer-based MLPF matches particle flow in CMS at half the runtime","feed_subtitle":"Single learned pass replaces rule-based linking and pileup subtraction at half the compute.","key_machinery":"The carrying mechanism is the transformer-based MLPF model—a neural network that uses attention over all input elements—trained against a simulation-defined target. Each target particle is assigned to a unique primary input element, charged particles to their originating track and neutral particles to the cluster carrying most of their energy, so the loss factors per input element as four terms: binary existence classification, focal-loss particle-ID classification, binary pileup classification, and four-momentum regression. At inference the model's learned per-particle pileup score replaces PUPPI for neutral and non-tracker charged particles, while tracker charged particles keep the standard treatment; the resulting candidates feed ordinary AK4 jet clustering and missing-transverse-momentum reconstruction, isolating the effect of the learned PF step.","core_discovery":"MLPF is trained end-to-end on Geant4-simulated 14 TeV Run-3-condition events with a flat pileup profile of 55-75 interactions, using a target that aligns with pileup-subtracted Pythia truth particles. The model outputs for each input track or cluster a particle-existence label, particle type, pileup label, and four-momentum, with the loss weighted by sqrt(pT) to protect high-transverse-momentum performance. The paper reports that AK4 jet response and resolution and pTmiss in tt and QCD simulation are comparable to PF with PUPPI pileup mitigation, that MLPF improves neutral-hadron reconstruction efficiency at the same fake rate, and that on dijet data the pTmiss and dijet asymmetry distributions agree with PF. It further shows the model integrated into the CMS offline reconstruction software with stable output, running at about 40 ms/event on GPU versus 90 ms/event for PF on CPU.","pith_inferences":["If the simulation-to-data transfer holds at higher luminosity, the learned per-particle pileup score could be retrained directly on Phase-2 pileup conditions, replacing hand-tuned PUPPI parameters.","Because the model outputs ordinary particle candidates, the training objective could be extended to downstream physics targets such as jet response or missing-transverse-momentum resolution; the authors note their performance was not explicitly optimized for those, so this is a natural next step beyond the paper.","The data validation rests only on shape agreement of two distributions; a dedicated data calibration with an independent reference, such as Z+jet or photon+jet balance, would be a stronger test than what the paper reports."],"forward_implications":["MLPF can replace the PF reconstruction step in the CMS offline reconstruction chain without modifying downstream modules.","Jet momentum response and resolution and missing transverse momentum in tt and QCD simulation with pileup are comparable to PF with PUPPI pileup mitigation.","MLPF improves neutral-hadron reconstruction efficiency relative to PF while keeping the same fake rate.","On a 375 pb^-1 sample of 13.6 TeV dijet data, the model's missing-transverse-momentum and dijet-asymmetry distributions agree with PF.","The GPU implementation runs about twice as fast as PF on CPU, around 40 ms per event versus 90 ms, with stable output across repeated runs."],"supporting_citations":[{"why":"Defines the CMS particle-flow algorithm that MLPF is designed to replace and is compared against.","marker":"[1]"},{"why":"Original MLPF method using graph neural networks whose architecture and loss this work extends with transformers.","marker":"[9]"},{"why":"Prior CMS MLPF progress supplying the particle-level loss and target definition used here.","marker":"[11]"},{"why":"FlashAttention kernels that let the transformer run on GPU at the reported inference speed.","marker":"[12]"},{"why":"CMS pileup mitigation in 13 TeV data, the PF-PUPPI baseline used for jet and pTmiss comparisons.","marker":"[15]"},{"why":"The PUPPI per-particle pileup-weight algorithm whose role MLPF's learned pileup score replaces.","marker":"[16]"},{"why":"Pythia 8.3 provides the simulated truth particles used to define training targets and validation.","marker":"[7]"},{"why":"Geant4 simulation generates the detector response from which tracks and clusters are reconstructed.","marker":"[4]"}],"fun_headline_variants":["MLPF transformer matches CMS particle flow, cuts runtime in half","Single-pass ML particle flow matches PF in CMS at half runtime","CMS MLPF: transformer event reconstruction at half the compute","MLPF transformer cuts CMS particle-flow runtime in half","End-to-end MLPF ties PF accuracy in CMS at half the runtime"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that training on 14 TeV simulated events with 55-75 pileup transfers to real 13.6 TeV collisions; the paper's data check is limited to agreement of the missing-transverse-momentum and dijet-asymmetry shapes, with no quantitative jet energy scale, resolution, or closure validation on data.","fun_headline_variants_meta":{"raw":{"variants":["MLPF transformer matches CMS particle flow, cuts runtime in half","Single-pass ML particle flow matches PF in CMS at half runtime","CMS MLPF: transformer event reconstruction at half the compute","MLPF transformer cuts CMS particle-flow runtime in half","End-to-end MLPF ties PF accuracy in CMS at half the runtime"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000718,"raw_usage":{"total_tokens":3162,"prompt_tokens":816,"completion_tokens":2346,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":2261}},"tokens_in":432,"tokens_out":2346,"duration_ms":14522,"temperature":1.0,"reasoning_tokens":2261,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:43:24.327755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be an independent measurement of the MLPF jet energy response in data, for instance using Z+jet or photon+jet transverse-momentum balance, or a simulation closure test at 2024 conditions with a pileup profile outside the flat 55-75 training range. If the measured response or the dijet asymmetry deviates from the PF-PUPPI result by more than the paper's plots show, the claim of comparable physics performance on data would not be supported.","supporting_citations":[{"cited_title":"Progress towards an improved particle flow algorithm at CMS with machine learning","cited_arxiv_id":"2303.17657","evidence_quote":"Prior CMS MLPF progress supplying the particle-level loss and target definition used here."},{"cited_title":"Geant4 – a simulation toolkit","cited_arxiv_id":null,"evidence_quote":"Geant4 simulation generates the detector response from which tracks and clusters are reconstructed."}],"review_version":2}