Pith. sign in

REVIEW 5 minor 37 references

Commissioning and Low Latency Operation of the Graph Neural Network Electromagnetic Calorimeter Trigger at the Belle II Experiment

T0 review · 0 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read A graph-neural-network calorimeter trigger now runs at Belle II with 1.053 µs latency, meeting the first-level trigger budget for the first time.

desk verdict Solid commissioning paper: they actually got a GNN calorimeter trigger inside Belle II's hard L1 budget and ran it for months. read the letter →

arxiv 2607.09347 v1 pith:APFA5YLK submitted 2026-07-10 hep-ex eess.SP

classification hep-exeess.SP
keywords ElectromagneticCalorimeterClusteringFPGAsGraphNeuralNetworksMachineLearningParticlePhysicsTriggerBelleII
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports that a graph neural network that treats calorimeter trigger cells as graph nodes can cluster energy deposits and produce trigger bits fast enough for Belle II’s first-level trigger. Through model compression, floorplanning and DSP remapping the authors cut end-to-end latency from 3.168 µs to 1.053 µs, just under the 1.067 µs hard limit. They fully wire the module into the existing trigger chain, add slow-control and online rate monitoring, validate bit-accurate agreement with the offline model, and have operated the system under real beam conditions since December 2025. The result is a working, low-latency GNN trigger module that can now be compared quantitatively with the classical clustering logic already used by the experiment.

What carries the argument

The Sunset CaloClusterNet: a single-GravNet, heterogeneously quantised (≤8-bit power-of-two) object-condensation GNN whose FPGA dataflow, after SLR floorplanning and DSP disabling, realises the 1.053 µs critical path.

What would settle it

A simultaneous beam run in which the same trigger bit is read both via the LEMO cable and via an independent, fully timed optical path whose absolute delay is known to better than 10 ns; if the corrected LEMO latency then exceeds 1.067 µs, the margin claim fails.

Watch

Extended reading notes

Core claim

Hardware–algorithm co-design (single-GravNet Sunset model with ≤8-bit fixed-point layers, manual Super-Logic-Region floorplanning and DSP-to-logic remapping) reduces GNN-ETM end-to-end latency from 3.168 µs to 1.053 µs, satisfying the Belle II L1 ECL budget of 1.067 µs; the module has been commissioned with full slow-control and rate monitoring and operated in the experiment since December 2025.

Load-bearing premise

That the 95-percent arrival-time histogram recorded at the global decision logic, after the programmed frontend offset and simulated latency correction, truly equals the module’s end-to-end latency under all beam conditions.

Editorial extensions

If this is right

  • GNN-ETM can now inject at least one runtime-reconfigurable trigger bit into the active L1 decision via the LEMO path.
  • Online rate counters stored in the EPICS archiver allow direct, quantitative comparison of GNN versus classical clustering under identical beam conditions.
  • Structural reordering of the ECL trigger chain can relax the latency budget to 1.367 µs, opening headroom for richer trigger logic.
  • The same co-design recipe (model compression + floorplanning + DSP remapping) can be reused for other low-latency GNN stages in the Belle II pipeline.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Once the signal-classifier bit is unmasked, the observed rate reduction relative to the classical module may translate into a cleaner physics sample or a higher effective luminosity for rare-event searches.
  • The 14 ns residual margin is small enough that any future increase in detector occupancy or firmware feature set will force either further compression or a move to a larger FPGA generation.
  • Bit-accurate RTL-to-offline agreement demonstrated here becomes a practical template for certifying other machine-learning triggers before they are allowed to veto data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The manuscript reports the commissioning and low-latency operation of the Graph Neural Network Electromagnetic Calorimeter Trigger Module (GNN-ETM) in the Belle II L1 trigger. Building on a prior FPGA prototype, the authors reduce end-to-end latency from 3.168 µs to 1.053 µs via hardware–algorithm co-design (model compression to a single GravNet block with ≤8-bit fixed-point layers, manual SLR floorplanning, and DSP-to-logic remapping), integrate trigger-bit generation and the link to the Global Decision Logic, and add slow-control drivers plus online rate monitoring. RTL simulations achieve bit-accurate agreement with the offline reference; place-and-route timing is closed; live GDL arrival-time measurements under beam conditions confirm a 14 ns margin under the stated configuration; and C2 trigger rates are compared with the existing ICN-ETM for cosmic and beam runs. The module has been operated in Belle II since December 2025.

Significance. This is a concrete, operational demonstration that a GNN-based calorimeter clustering algorithm can meet the hard real-time latency and throughput constraints of a collider L1 trigger and be fully integrated into the experiment’s slow-control and monitoring infrastructure. The combination of cycle-accurate RTL validation, place-and-route resource and timing reports, and live GDL latency and rate measurements provides a reproducible engineering path that other experiments can follow. The work advances the prior prototype from a parallel debug module to a system that is latency-compliant and ready to contribute trigger bits, which is a meaningful step for machine-learning triggers in high-energy physics.

minor comments (5)
  1. Section VI-A / Eq. (1): The construction of ˆt_gnn from the GDL histogram, t_FAM, t_sim and ˆt_0.95 is clear, but a short explicit statement that the 14 ns margin is configuration-dependent (and vanishes without the FAM offset) would help readers who only skim Fig. 9.
  2. Figure 10 and surrounding text: The qualitative discussion of why GNN-ETM w/o sig. rates exceed ICN-ETM rates in beam runs (cluster splitting vs. endcap TC definition) is useful; a sentence noting that a full efficiency/purity comparison is left for future work would set expectations more cleanly.
  3. Section IV-A / Fig. 4: The quantisation scheme (mostly Q1.7/Q2.6, interfaces Q4.12/Q5.11) is stated; a brief remark on whether any layers required non-power-of-two scales or special handling would complete the reproducibility picture.
  4. Throughout: A few typographical inconsistencies remain (e.g., “Organiization”, mixed µs/us notation, and the arXiv date stamp of July 2026 versus “operated since December 2025”). These are easily corrected in production.
  5. Section II: The latency budget of 1.067 µs is stated to differ from the prior paper because the module order is not swapped; a parenthetical cross-reference to the earlier value would avoid confusion for readers of both works.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: latency and rates are independently measured engineering results, not predictions forced by fitted inputs or self-citation chains.

full rationale

This is a hardware commissioning and co-design paper, not a first-principles derivation. The central claim—that end-to-end GNN-ETM latency falls from 3168 ns to 1053 ns and meets the 1067 ns Belle II L1 ECL budget—is established by cycle-accurate RTL simulation (Section V-A, Fig. 6) with bit-accurate agreement to the offline model, plus place-and-route timing closure after model compression, SLR floorplanning, and DSP remapping. Live GDL arrival-time histograms (Eq. 1, Fig. 9) and C2 rate comparisons (Fig. 10) are observational confirmations under cosmic and beam runs, not fitted parameters renamed as predictions. Citation of the prior prototype [12] is ordinary baseline reference for an improved system; it is not used as a uniqueness theorem, ansatz smuggling, or load-bearing external fact that forces the present result. No self-definitional loop, fitted-input-as-prediction, or renaming of a known result appears in the derivation chain. Circularity score is therefore 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard FPGA timing and Belle II system parameters already established in the literature, plus a small set of design choices (quantisation bit-widths, floorplan, DSP disable) that are engineering decisions rather than free parameters fitted to physics data. No new physical entities are postulated.

free parameters (3)
  • layer-wise fixed-point bit-widths (mostly Q1.7 / Q2.6, interfaces Q4.12 / Q5.11)
    Chosen by manual hyperparameter search under a hard 8-bit-per-layer limit; not fitted to physics observables but selected for hardware cost vs. algorithmic fidelity.
  • system clock frequency f_sys = 254.432 MHz for the GNN dataflow accelerator
    Doubled from the previous design by floorplanning; an engineering choice that directly sets the latency budget.
  • programmable Frontend Analog Module delay offset t_FAM = -146 ns
    Runtime-reconfigurable offset used to enlarge the measured latency margin; its value is chosen by operators, not derived.
assumptions (4)
  • domain assumption Belle II L1 ECL trigger latency budget for a module in the ICN-ETM position is 1.067 µs
    Stated as a hard real-time requirement derived from buffer depths and system architecture (Section II); taken as given from prior Belle II documentation.
  • domain assumption Synchronous clock-domain crossings introduce negligible boundary effects relative to the overall latency
    Invoked when the authors introduce user-defined f_sys for the accelerator (Section III).
  • domain assumption Bit-accurate RTL simulation of the full design is a faithful predictor of post-place-and-route behaviour on the XCVU190
    Used to claim functional correctness before hardware commissioning (Section V-A).
  • ad hoc to paper The 95 % quantile of the GDL rising-edge clock-counter histogram, after the stated offsets, equals the true end-to-end GNN-ETM latency
    Eq. (1) and the analysis surrounding Fig. 9; the mapping from observed ˆt to ˆt_gnn is an experimental definition introduced in this work.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Commissioning and Low Latency Operation of the Graph Neural Network Electromagnetic Calorimeter Trigger at the Belle II Experiment." pith.science (2026). https://pith.science/paper/APFA5YLK

@misc{pith2026260709347,
  author       = {Pith},
  title        = {Pith review of: Commissioning and Low Latency Operation of the Graph Neural Network Electromagnetic Calorimeter Trigger at the Belle II Experiment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/APFA5YLK}},
  note         = {Machine review of arXiv:2607.09347}
}
read the original abstract

We present the commissioning and operation of the Graph Neural Network Electromagnetic Calorimeter Trigger Module (GNN-ETM) of the Belle II experiment at the SuperKEKB collider. The GNN-ETM processes calorimeter trigger cells as graph nodes to perform clustering and feature extraction. We fully integrate the system with the successive stages of the first-level trigger, develop slow-control drivers, and add online monitoring capabilities. We optimise the existing FPGA-based architecture through hardware-algorithm co-design, achieving an overall system latency of 1.053 us. Our hardware implementation is validated through register-transfer-level simulations, achieving bit-accurate agreement with the offline reference model. Online monitoring enables the measurement of instantaneous trigger rates, providing a quantitative basis for trigger-level performance studies. In summary, we report on the GNN-ETM as a fully operational, low-latency trigger module with online control and monitoring capabilities, compatible with the latency requirements of the Belle II first-level trigger system.

Figures

Figures reproduced from arXiv: 2607.09347 by the authors.

Figure 1
Figure 1. Simplified overview of the ECL L1 trigger system at the Belle II Experiment. Components of the L1 trigger system that must satisfy hard real-time [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 3
Figure 3. Overview of the GNN-ETM postprocessing stage. Latency-critical submodules are blue; the remaining submodules are green. The external interfaces of this submodule are also shown in [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Most values are quantized to Q1.7 and Q2.6, meaning that most values have to lie in the range [−1, 1] and [−2, 2] respectively. At the interfaces, we keep the Q4.12 and Q5.11 quantisation to retain the full resolution. In comparison to the Armadillo CaloClusterNet, the Sunset CaloClusterNet does not yield a significantly lower algorithmic performance. Q1.7 Quantization Q5.11 Quantization Q2.6 Quantization Q3.5 Quant… view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Floorplan constraints of the GNN-ETM for implementation with AMD Vivado 2024.2. Hierarchical modules are the same as in [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 7
Figure 7. Figure 7: Utilisation of system resources on the AMD Ultrascale XCVU190 [PITH_FULL_IMAGE:figures/full_fig_p005_7.png]
Figure 6
Figure 6. Figure 6: End-to-end latency for the complete inference chain on the UT4 [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 9
Figure 9. Figure 9: Relative occurrence of the GNN-ETM latency tˆgnn of a trigger bit from GNN-ETM, measured at the GDL. GT equals transmission via gigabit transceiver. LEMO equals transmission via twisted-pair cable. Blue and green dashed lines denote the 95 % quantile for the respective…
Figure 10
Figure 10. Figure 10: Comparison of C2 trigger rates between ICN-ETM and GNN-ETM based on monitoring PVs on the GDL. GNN-ETM trigger rates are shown both with the signal classifier (w. sig.) and without the signal classifier (w/o sig.). quantitative statements on the two systems. By making…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 2 linked inside Pith

  1. [1]

    End-to-End Multi-track Reconstruction Using Graph Neural Networks at Belle II,

    L. Reuteret al., “End-to-End Multi-track Reconstruction Using Graph Neural Networks at Belle II,”Computing and Software for Big Science, vol. 9, no. 1, p. 6, Dec. 2025

  2. [2]

    Graph Neural Networks for Charged Particle Tracking on FPGAs,

    A. Elabdet al., “Graph Neural Networks for Charged Particle Tracking on FPGAs,”Frontiers in Big Data, vol. 5, p. 828666, Mar. 2022. IEEE TRANSACTIONS ON NUCLEAR SCIENCE, VOL. XX, NO. XX, XXXX 2026 8

  3. [3]

    Hardware-accelerated GNN-based hit filtering for the Belle II Level-1 trigger,

    G. a. o. Heine, “Hardware-accelerated GNN-based hit filtering for the Belle II Level-1 trigger,”Journal of Instrumentation, vol. 21, no. 02, p. C02007, 2 2026

  4. [4]

    Photon Reconstruction in the Belle II Calorimeter Using Graph Neural Networks,

    F. Wemmeret al., “Photon Reconstruction in the Belle II Calorimeter Using Graph Neural Networks,”Comput. Softw. Big Sci., vol. 7, no. 1, 12 2023

  5. [5]

    Jet tagging via particle clouds,

    H. Qu and L. Gouskos, “Jet tagging via particle clouds,”Physical Review D, vol. 101, no. 5, p. 056019, Mar. 2020

  6. [6]

    Graph neural networks in particle physics,

    J. Shlomi, P. Battaglia, and J.-R. Vlimant, “Graph neural networks in particle physics,”Machine Learning: Science and Technology, vol. 2, no. 2, p. 021001, Jan. 2021

  7. [7]

    JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs,

    Z. Queet al., “JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs,” 8 2025

  8. [8]

    Online track reconstruction with graph neural networks on FPGAs for the ATLAS experiment,

    S. Dittmeier, “Online track reconstruction with graph neural networks on FPGAs for the ATLAS experiment,”EPJ Web Conf., vol. 337, 2025

Show all 37 references
  1. [9]

    Real-Time Graph-based Point Cloud Networks on FP- GAs via Stall-Free Deep Pipelining,

    M. Neuet al., “Real-Time Graph-based Point Cloud Networks on FP- GAs via Stall-Free Deep Pipelining,” in2025 38th SBC/SBMicro/IEEE Symposium on Integrated Circuits and Systems Design (SBCCI), 2025

  2. [10]

    LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics,

    Z. Queet al., “LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics,”ACM Transactions on Embedded Computing Systems, vol. 23, no. 2, pp. 1–28, Mar. 2024

  3. [11]

    Low Latency Edge Classification GNN for Particle Trajectory Tracking on FPGAs,

    S.-Y . Huanget al., “Low Latency Edge Classification GNN for Particle Trajectory Tracking on FPGAs,” in2023 33rd International Conference on Field-Programmable Logic and Applications (FPL). Gothenburg, Sweden: IEEE, Sep. 2023, pp. 294–298

  4. [12]

    Real-Time Graph Neural Networks on FPGAs for the Belle II Electromagnetic Calorimeter,

    I. Haideet al., “Real-Time Graph Neural Networks on FPGAs for the Belle II Electromagnetic Calorimeter,”JINST, 2026

  5. [13]

    SuperKEKB Collider,

    K. Akai, K. Furukawa, and H. Koiso, “SuperKEKB Collider,”Nucl. Instrum. Meth. A, vol. 907, pp. 188–199, 2018

  6. [14]

    Belle II Technical Design Report,

    T. Abeet al., “Belle II Technical Design Report,” 11 2010

  7. [15]

    Data Acquisition System for the Belle II Experiment,

    S. Yamadaet al., “Data Acquisition System for the Belle II Experiment,” IEEE Trans. Nucl. Sci., vol. 62, no. 3, 2015

  8. [16]

    Design of the Global Reconstruction Logic in the Belle II Level-1 Trigger system,

    Y .-T. Laiet al., “Design of the Global Reconstruction Logic in the Belle II Level-1 Trigger system,”Nucl. Instrum. Meth. A, vol. 1078, 2025

  9. [17]

    Status of the electromagnetic calorimeter trigger system at Belle II,

    S. Kimet al., “Status of the electromagnetic calorimeter trigger system at Belle II,”J. Phys. Conf. Ser., vol. 928, 11 2017

  10. [18]

    Electromagnetic calorimeter of the Belle II detector,

    B. Shwartz and BELLE II calorimeter group, “Electromagnetic calorimeter of the Belle II detector,”Journal of Physics: Conference Series, vol. 928, p. 012021, Nov. 2017

  11. [19]

    Electromagnetic calorimeter trigger at Belle,

    B. Cheonet al., “Electromagnetic calorimeter trigger at Belle,”Nucl. Instrum. Meth. A, vol. 494, no. 1, 2002

  12. [20]

    [Online]

    Arm Limited,AMBA AXI-Stream Protocol Specification, 2021, iHI 0051B. [Online]. Available: https://developer.arm.com/documentation/ ihi0051/latest/ [21]IEEE Standard for a Versatile Backplane Bus: VMEbus, Institute of Electrical and Electronics Engineers, New York, NY , USA, 1987

  13. [21]

    Belle2Link: A Global Data Readout and Transmission for Belle II Experiment at KEK,

    D. Sunet al., “Belle2Link: A Global Data Readout and Transmission for Belle II Experiment at KEK,”Physics Procedia, vol. 37, pp. 1933–1939, 2012

  14. [22]

    Chisel: Constructing Hardware in a Scala Embedded Language,

    J. Bachrachet al., “Chisel: Constructing Hardware in a Scala Embedded Language,” inProceedings of the 49th Annual Design Automation Conference. San Francisco, California: ACM, 6 2012

  15. [23]

    The Slow Control and Data Quality Monitoring System for the Belle II Experiment,

    T. Konnoet al., “The Slow Control and Data Quality Monitoring System for the Belle II Experiment,”IEEE Transactions on Nuclear Science, vol. 62, no. 3, pp. 897–902, Jun. 2015

  16. [24]

    Trigger slow control system of the Belle II experi- ment,

    C.-H. Kimet al., “Trigger slow control system of the Belle II experi- ment,”Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, vol. 1014, p. 165748, Oct. 2021

  17. [25]

    Network shared memory framework for the Belle data acquisition control system,

    M. Nakao and S. Suzuki, “Network shared memory framework for the Belle data acquisition control system,” in1999 IEEE Conference on Real-Time Computer Applications in Nuclear Particle and Plasma Physics. 11th IEEE NPSS Real Time Conference. Conference Record (Cat. No.99EX295). ...

  18. [26]

    Learning representations of irregular particle- detector geometry with distance-weighted graph networks,

    S. R. Qasimet al., “Learning representations of irregular particle- detector geometry with distance-weighted graph networks,”Eur. Phys. J. C, vol. 79, no. 7, 1 2019

  19. [27]

    Object condensation: one-stage grid-free multi-object re- construction in physics detectors, graph and image data,

    J. Kieseler, “Object condensation: one-stage grid-free multi-object re- construction in physics detectors, graph and image data,”Eur. Phys. J. C, vol. 80, no. 9, p. 886, 2020

  20. [28]

    Vitis Unified Software Platform,

    AMD, “Vitis Unified Software Platform,” https://www.amd.com/ en/products/software/adaptive-socs-and-fpgas/vitis.html, 2025, version 2024.2, accessed 2025-10-28

  21. [29]

    Vivado Design Suite,

    ——, “Vivado Design Suite,” https://www.amd.com/de/products/ software/adaptive-socs-and-fpgas/vivado.html, 2025, version 2024.2, accessed 2025-10-28

  22. [30]

    Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors,

    C. N. Coelhoet al., “Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors,”Nature Machine Intelligence, vol. 3, no. 8, pp. 675– 686, Aug. 2021. [Online]. Available: https://www.nature.com/articles/ s42256-021-00356-5

  23. [31]

    Code for the GNN-ETM Training and Evaluation,

    I. Haideet al., “Code for the GNN-ETM Training and Evaluation,” https://github.com/ihaide/gnnetm-software, 2026

  24. [32]

    Code for the Quantized GravNet Implementation,

    M. Neuet al., “Code for the Quantized GravNet Implementation,” https://github.com/ihaide/qgravnet, 2026

  25. [33]

    Custom QKeras Fork,

    M. Neu and I. Haide, “Custom QKeras Fork,” https://github.com/ihaide/qkeras, 2026

  26. [34]

    Guptaet al.Deep Learning with Limited Numerical Precision

    S. Guptaet al.Deep Learning with Limited Numerical Precision. [Online]. Available: http://arxiv.org/abs/1502.02551

  27. [35]

    Liu and M

    Z.-G. Liu and M. Mattina. Learning low-precision neural networks without Straight-Through Estimator(STE). [Online]. Available: http: //arxiv.org/abs/1903.01061

  28. [36]

    ModelSim HDL simulator,

    AMD, “ModelSim HDL simulator,” https://eda.sw.siemens.com/en-US/ ic/modelsim/, 2025, version 2023.4, accessed 2025-05-13

  29. [37]

    The EPICS Archiver Appliance,

    M. Shankaret al., “The EPICS Archiver Appliance,”Proceedings of the 15th Int. Conf. on Accelerator and Large Experimental Physics Control Systems, vol. ICALEPCS2015, pp. 4 pages, 0.753 MB, 2015

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.