{"id":"2b4ed8c5-a86b-414e-9a55-799c4a4223ae","arxiv_id":"2411.19506","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"CMS has deployed a machine-learning anomaly trigger on its level-1 hardware that stably collects events standard triggers miss, marking the first operational use of such a system at a hadron collider's first trigger stage.","lead":"A neural network now runs on the CMS experiment's trigger hardware, scanning every proton collision at 40 million per second and flagging unusual events without being told what to look for. During real data-taking in 2024, it ran stably and selected a different mix of events than the standard trigger system.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The anomaly score's discrimination power is unvalidated: the trigger may be selecting high-multiplicity events rather than new physics, so the central scientific claim lacks direct support.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: an autoencoder's reconstruction error (and here, the related latent-space score) must separate anomalous events from background for the trigger to be scientifically valuable. The paper's evidence is limited to stable rate monitoring and an observed preference for high multiplicity, which is actually a warning sign rather than confirmation of anomaly sensitivity. This concern does not undermine the deployment claim, which is plausible and supported by the monitoring plots, but it does undermine the stronger claim that the collected events are enriched in new physics. The reader's CONDITIONAL verdict already captures this, and my analysis supports that condition: the authors should provide direct signal-efficiency studies or a demonstration that the score is not reducible to trivial event features. I therefore recommend keeping the verdict unchanged. The paper deserves credit for a real hardware deployment, but the scientific payoff remains unestablished in this proceedings write-up.","tokens_in":4573,"tokens_out":2595,"duration_ms":25675,"concrete_test":"Using either the trained VAE (from the paper's authors or a reimplementation) and the L1 object definitions, compute the anomaly score for a sample of simulated ZeroBias events and for simulated benchmark BSM events (e.g., H→4b, SUEP). Measure the separation in anomaly score between signal and background, and then repeat the comparison after conditioning on L1 object multiplicity and pile-up. If the signal-background separation disappears or becomes negligible once multiplicity is matched, the anomaly score is primarily encoding multiplicity rather than anomalous physics. As a complementary check on real data, compute the rank correlation between the AXOL1TL score and the total number of L1 objects in the triggered events; a correlation above 0.9 would strongly support the multiplicity-proxy hypothesis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CMS is now collecting a data stream enriched in anomalous events in real time. That claim depends on the anomaly score—the sum of squared latent means of the AXOL1TL VAE, Section 2.2—being a reliable discriminant for BSM or rare SM physics. The paper provides no direct measurement of signal efficiency for any BSM model, and the only indirect evidence, Fig. 4, shows a preference for high-multiplicity events. Because the training data is ZeroBias at fixed pile-up 62, the VAE may simply learn that high-multiplicity events are rare in the training distribution, making the anomaly score a proxy for L1 object multiplicity. The claimed 46% efficiency gain for H→4b is stated without a supporting study or reference, and the 'orthogonality' to the standard L1 menu may be trivial if it merely reflects the known multiplicity preference. If this is the case, the trigger runs stably but collects events already accessible to standard multi-jet triggers, undermining the novelty and scientific value of the deployed system.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the preparation, deployment, and live testing of two autoencoder-based anomaly detection algorithms, AXOL1TL and CICADA, at the CMS Level-1 trigger. AXOL1TL is a variational autoencoder that takes as input L1 objects (jets, electrons/photons, muons, and MET) in hardware integer precision, computes a latent vector of size 8, and uses the sum of squared latent means as the anomaly score. The network is translated to FPGA firmware via hls4ml and installed in the CMS Global Trigger test crate, which receives the same inputs as the main trigger without affecting data taking. The paper describes the methodology, the five thresholds tested, and presents monitoring data from June 2024 (Fig. 3) showing stable trigger rates, score distributions as a function of L1 object multiplicity (Fig. 4), and invariant mass distributions from scouting data (Fig. 5). The authors claim the AXOL1TL trigger operated stably and selects events orthogonal to the standard L1 menu, with a quoted 46% efficiency gain for H→4b relative to a rule-based trigger at 1 kHz.","tokens_in":4652,"tokens_out":3587,"duration_ms":31373,"significance":"If the deployment claims are correct, this is a noteworthy technical milestone: it demonstrates that unsupervised anomaly detection can be executed in the FPGA-based L1 trigger within the 50 ns latency budget and operated on live proton collision data. The paper also provides evidence of bit-exactness between the HLS emulation and the qkeras model (Fig. 2), and it references a public CMS dataset record [3], which supports reproducibility. However, the broader scientific significance depends on the anomaly score genuinely selecting events that are both new and useful for BSM searches. That property is asserted but not quantitatively demonstrated in this manuscript: no signal efficiency or background rejection for any BSM model is shown, and the only indirect evidence (Fig. 4) shows a preference for high-multiplicity events, which may be a trivial feature of the training distribution. Thus the paper is valuable as an engineering and operational report, but its physics claims are currently unsupported.","major_comments":[{"comment":"The sentence 'In the case of AXOL1TL, it is to have 46% efficiency gain when compared to rest of L1 trigger, when operating at a rate of 1kHz for capturing exotic decay of higgs to four b quarks' states a precise, quantitative physics result without any supporting study, reference, or definition of the efficiency measurement. There is no description of the signal sample, the baseline trigger, or the statistical procedure. Because this number is the only quantitative claim about physics performance in the paper, it is load-bearing for the motivation, but the reader cannot verify or reproduce it. The authors should either provide a reference to a public CMS note or analysis, or remove the claim and replace it with a qualitative statement.","section":"Section 2, paragraph 2"},{"comment":"The claim that 'The dataset triggered by AXOL1TL tends to be orthogonal to events triggered by the regular L1 Trigger menu, as seen in Fig. 4' is not supported by the evidence shown. The right panel of Fig. 4 plots AXOL1TL score as a function of L1 object multiplicity, which demonstrates only a multiplicity preference; it does not quantify the overlap or complementarity between AXOL1TL-triggered events and standard L1 menu events. Orthogonality is a stronger statement that requires, for example, the fraction of AXOL1TL events that pass or fail standard seeds, or a comparison of trigger efficiencies on a common event sample. As written, the 'orthogonality' claim is unsubstantiated, and the subsequent statement that this 'highlights the novelty of the events' overinterprets the figure.","section":"Section 3, Fig. 4 and surrounding text"},{"comment":"The manuscript relies on the assumption that the autoencoder reconstruction error (or its latent-proxy, the sum of squared latent means) is a reliable discriminant for BSM or rare SM physics, but no validation of this assumption is presented. The paper shows no signal efficiency for any BSM model, no background rejection curve, and no closure test demonstrating that anomalous events actually produce high anomaly scores in the chosen L1 input features. In particular, the observed preference for high-multiplicity events (Fig. 4) raises the possibility that the score is a proxy for object multiplicity rather than a physically meaningful anomaly measure. Since the scientific value of the deployed trigger depends on the score's discriminatory power, this is a load-bearing gap. The authors should at least include a simulation-based benchmark (e.g., H→4b or another CMS-endorsed signature) or explicitly state that such validation is deferred to a separate publication.","section":"Section 2, paragraph 2 and Section 2.2"}],"minor_comments":[{"comment":"The phrase 'it is to have 46% efficiency gain' should be corrected to 'it is estimated to have a 46% efficiency gain', and 'exotic decay of higgs to four b quarks' should be 'exotic decay of the Higgs boson to four b quarks'.","section":"Section 2, paragraph 2"},{"comment":"The text states that invariant mass distributions of pairs of 'jets, electrons, and photons' are studied, but the figure caption lists 'jets (left), muons (center), and photons (right)'. This inconsistency should be resolved; if muons were studied, the text should say so, and electrons should be mentioned only if included.","section":"Section 3, Fig. 5 caption"},{"comment":"The rate monitoring plot would benefit from axis labels and units. The caption says 'Global trigger rate monitoring time series' but does not specify the y-axis unit (presumably Hz or kHz), which is important for interpreting stability.","section":"Section 3, Fig. 3 caption"},{"comment":"The choice to use the sum of squared latent means, rather than a full reconstruction-based anomaly score, is stated but not motivated. A brief justification of why this proxy is adequate (e.g., empirical equivalence or latency constraints) would improve the manuscript.","section":"Section 2.2"},{"comment":"There are several typographical issues, including 'multĳet' (should be 'multijet') and minor grammatical errors such as 'The Neural Network makes a prediction for each event within these constraints' where 'each event' is not strictly accurate for all events; these should be corrected in a final proofreading.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a conference proceedings contribution, and the engineering/deployment story is credible and of interest to the HEP community. However, the physics performance claims, particularly the 46% H→4b number and the orthogonality statement, are not supported by the evidence in the paper. The authors may be able to address these by referring to a public CMS note or by explicitly framing the paper as an operational report with physics validation deferred. I recommend major_revision rather than rejection because the central deployment claim appears sound and the missing evidence is local and fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this is a status report, not a validation paper, and the part worth taking seriously is the deployment claim. AXOL1TL, a VAE trained on ZeroBias data, is running in the CMS L1 Global Trigger test crate during 2024 data-taking, and the rate monitoring and score distributions are plausible evidence that it runs stably and selects a different set of events than the standard L1 menu. That alone is a first for a hadron collider experiment, and it matters.\n\nWhat the paper does well: it gives a clear description of the two algorithms (AXOL1TL and CICADA), the input objects, the quantization-aware training and hls4ml FPGA implementation, and the live monitoring. The lack of mass sculpting in the scouting data (Fig. 5) is a useful early check. And it is honest that CICADA is not yet commissioned.\n\nWhere it is soft: the paper makes several quantitative claims that are not backed up here. The 46% efficiency gain for H→4b is stated without a study, reference, or statistical test. The claim that the anomaly triggers are “better” than rule-based triggers is similarly unsubstantiated. More fundamentally, the central scientific assumption—that reconstruction error separates BSM events from SM background—is never directly tested with any BSM sample. The only indirect evidence, Fig. 4, shows a preference for high-multiplicity events, which is consistent with the VAE having learned that high object multiplicity is rare in its fixed-pile-up training sample. If that is what drives the anomaly score, the trigger may be enriching a region that standard multi-jet triggers already cover. That concern is real, and the paper does not address it. No code, model weights, or architecture details are released either, which limits reproducibility.\n\nNone of this kills the deployment claim. The test crate is the right place to try this, the stability evidence looks reasonable for a first run, and the CMS note [3] presumably contains the backup. But as a standalone paper, the reader should treat the physics-sensitivity claims as promises, not results.\n\nWho gets value: anyone working on real-time ML at colliders, and anyone planning an anomaly trigger program. It deserves a serious referee if it were submitted as a full paper, but the referee should insist on seeing the signal-efficiency study and the multiplicity-dependence analysis before the physics claims are repeated.","headline":"A genuine first deployment of an ML anomaly trigger in the CMS L1 test crate, with the physics-sensitivity claims still unsubstantiated.","tokens_in":5321,"tokens_out":2861,"would_cite":true,"duration_ms":23780,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An autoencoder embedded in the Level-1 trigger FPGAs assigns an anomaly score to every 40 MHz LHC collision within the 50 ns latency budget.","keywords":["anomaly detection","variational autoencoder","Level-1 trigger","FPGA","real-time machine learning","CMS experiment","LHC Run 3","trigger-level data scouting"],"falsifier":"Feed simulated or embedded signal events, such as an exotic Higgs boson decaying to four b quarks, through the same L1 reconstruction and the AXOL1TL firmware emulation, and measure the score distribution against ZeroBias background; if the signal does not populate the high-score tails at any of the five thresholds, the claim that reconstruction error discriminates new physics fails. A control check would compare high-scoring and low-scoring events at the same object multiplicity, since the score difference disappearing once multiplicity is matched would indicate the score mostly measures event complexity rather than new physics.","tokens_in":4236,"feed_emoji":"⚛️","tokens_out":8590,"duration_ms":73359,"temperature":0.7,"pith_summary":"This paper reports that the CMS experiment has deployed a variational-autoencoder-based anomaly trigger, AXOL1TL, in the Global Trigger test crate FPGAs and operated it stably on proton-proton collisions during LHC Run 3. The algorithm computes an anomaly score for every collision within the trigger's latency constraints and uses that score to select events that the ordinary rule-based trigger menu would not select. The paper claims the AXOL1TL-triggered events are largely orthogonal to events from the regular L1 trigger and are enriched in high-multiplicity final states, making the dataset a model-independent resource for new-physics searches. If this is right, CMS is now continuously recording an anomalous-event stream in real time for offline analysis.","feed_headline":"An autoencoder now flags anomalous LHC collisions in real time","feed_subtitle":"Events outside the standard trigger menu are now stored for offline analysis, stably throughout June 2024 data-taking.","key_machinery":"The central object is the variational autoencoder: an encoder compresses L1 trigger objects (ten jets, four electron/photon objects, four muons, and missing transverse energy) into an eight-dimensional latent space constrained toward a standard normal distribution, and a decoder tries to reconstruct the input from that compressed representation. Only the encoder runs in real time, and the anomaly score is approximated by the sum of squared latent means, $\\sum_{i=1}^{8}\\mu_i^2$, which avoids running the decoder during inference. To fit inside the latency budget, the network is trained with quantization-aware techniques, translated to high-level synthesis firmware, and bit-exactly matched to the software emulation, so the FPGA score agrees with the trained model.","core_discovery":"On its own terms, the paper establishes that a variational autoencoder, trained on triggerless ZeroBias data from 2023 and compressed to a latent space of size eight, can be implemented in the Global Trigger test crate FPGAs and produce an anomaly score for every 40 MHz collision within the 50 ns latency window. During June 2024 data-taking, five AXOL1TL thresholds ran stably, and the events selected were largely orthogonal to those selected by the standard L1 trigger menu. High-multiplicity events tend to receive higher anomaly scores, and the triggered scouting data show invariant mass distributions of jets, muons, and photons without visible trigger-induced sculpting. The paper's deliverable is therefore a running, model-independent trigger stream of anomalous events rather than a measurement of any particular new-physics signature.","pith_inferences":["The paper does not report a signal-efficiency measurement for any specific BSM model; injecting simulated signals into the L1 object stream and measuring efficiency at each of the five thresholds would be a direct validation, with the 46% Higgs-to-four-b estimate serving as a concrete benchmark.","Because high multiplicity drives the anomaly score, part of the score likely tracks event complexity or pileup; comparing scores at fixed multiplicity or pileup would separate genuine new-physics selection from busy-event selection.","The same deployment pipeline could be applied to other architectures or input feature sets, and a trigger combining AXOL1TL's object-level score with calorimeter-image scores would cover complementary anomaly classes while being testable in the same test crate."],"forward_implications":["Events selected by AXOL1TL are stored for offline use, giving analyses a data stream that was not preselected by any particular BSM model.","The nominal-threshold stream feeds HLT scouting, producing a compact, continuously recorded sample of anomalous events, while the very-tight-threshold stream produces fully reconstructed events for discovery-oriented searches.","The orthogonality of AXOL1TL events to the standard L1 menu adds new event classes, especially high-multiplicity final states, to the CMS collected dataset.","At a 1 kHz rate, the paper estimates AXOL1TL would gain about 46% in efficiency over the rest of the L1 trigger for an exotic Higgs decay to four b quarks, indicating that the anomaly trigger finds events rule-based triggers miss.","Mass distributions in the AXOL1TL scouting data show no obvious selection sculpting, so the collected events may be usable for resonance searches without large trigger-correction uncertainties."],"supporting_citations":[{"why":"Establishes autoencoders on FPGAs as real-time unsupervised new-physics detection at 40 MHz, the approach AXOL1TL adapts to the Global Trigger.","marker":"[7]"},{"why":"Defines the variational autoencoder objective and latent-space construction that AXOL1TL's encoder and score formula use.","marker":"[11]"},{"why":"Documents the CMS dataset collected with AXOL1TL, which is the central deliverable the paper reports.","marker":"[3]"},{"why":"Provides the high-level synthesis workflow that translates the trained network to FPGA firmware for ultra-low-latency inference.","marker":"[1]"},{"why":"Supplies the quantization-aware training approach that makes the network implementable in fixed-point hardware.","marker":"[4]"},{"why":"Describes the data scouting and parking paradigm through which AXOL1TL-selected events are stored with minimal reconstruction.","marker":"[8]"},{"why":"Demonstrates fast FPGA inference of deep networks in particle physics, supporting the feasibility of sub-microsecond trigger inference.","marker":"[5]"}],"fun_headline_variants":["Autoencoder flags anomalous LHC events within 50 ns","CMS Global Trigger runs autoencoder for anomaly search","Real-time anomaly detection on FPGAs for LHC collisions","CMS deploys autoencoder trigger for model-independent physics","Anomaly autoencoder operates at 40 MHz in CMS trigger test crate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything depends on the assumption that events containing new physics will be harder for the autoencoder to reconstruct than ordinary background events, so a high reconstruction-error score reliably marks events worth keeping; if new physics looks just like the background in the features the network sees, the trigger will run perfectly but add nothing.","fun_headline_variants_meta":{"raw":{"variants":["Autoencoder flags anomalous LHC events within 50 ns","CMS Global Trigger runs autoencoder for anomaly search","Real-time anomaly detection on FPGAs for LHC collisions","CMS deploys autoencoder trigger for model-independent physics","Anomaly autoencoder operates at 40 MHz in CMS trigger test crate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000527,"raw_usage":{"total_tokens":2513,"prompt_tokens":882,"completion_tokens":1631,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1548}},"tokens_in":498,"tokens_out":1631,"duration_ms":10760,"temperature":1.0,"reasoning_tokens":1548,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:07:02.369133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed simulated or embedded signal events, such as an exotic Higgs boson decaying to four b quarks, through the same L1 reconstruction and the AXOL1TL firmware emulation, and measure the score distribution against ZeroBias background; if the signal does not populate the high-score tails at any of the five thresholds, the claim that reconstruction error discriminates new physics fails. A control check would compare high-scoring and low-scoring events at the same object multiplicity, since the score difference disappearing once multiplicity is matched would indicate the score mostly measures event complexity rather than new physics.","supporting_citations":[{"cited_title":"Govorkova et al","cited_arxiv_id":null,"evidence_quote":"Establishes autoencoders on FPGAs as real-time unsupervised new-physics detection at 40 MHz, the approach AXOL1TL adapts to the Global Trigger."},{"cited_title":"Hayrapetyan et al","cited_arxiv_id":null,"evidence_quote":"Describes the data scouting and parking paradigm through which AXOL1TL-selected events are stored with minimal reconstruction."}],"review_version":1}