REVIEW 2 major objections 4 minor 46 references
LHCb's 164-server GPU cluster processes all 32 Tbps of collision data in real time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:16 UTC pith:CCLWV5UC
load-bearing objection LHCb shows a real 32 Tbps trigger-less DAQ+GPU HLT architecture; the integrated claim is credible but the full-system test used a passthrough selection, so treat the headline as slightly ahead of the evidence. the 2 major comments →
A converged architecture for processing 32 Tbps of physics data in real-time at the LHCb experiment
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a converged architecture can handle the full 32 Tbps of LHCb detector output — roughly 150 kB per event at 30 MHz — in real time, reducing the event rate by a factor of 30. The system collapses three traditionally separate roles (readout, event building, and filtering) onto the same servers, connected by a non-blocking two-layer InfiniBand fat-tree. Event building is treated as an all-to-all personalized exchange with strict scheduling, and per-fragment buffers measured in seconds eliminate the need for deep-buffered switches. The assembled events are processed entirely on GPUs by a trigger application that batches events into groups of 600–800 per stream, uses stat
What carries the argument
The load-bearing mechanism is the convergence of readout, event building, and filtering on a single node: FPGA DAQ cards, InfiniBand host adapters, and GPUs share the same server and are kept within NUMA domains to avoid cross-socket traffic. Zero-copy RDMA (kernel bypass) moves event fragments from readout buffers through the fat-tree to builder buffers, and the GPU processes events directly from host memory. The GPU application's multi-event scheduler — which batches events into aligned execution masks and topologically sorts algorithm sequences under two heuristics (spread data providers, tighten masks of expensive algorithms) — is what makes full-event processing at 30 MHz feasible on co
Load-bearing premise
The 32 Tbps end-to-end claim rests on the assumption that the full-system test, which used a data generator and a passthrough selection at 30:1 acceptance, reproduces the I/O, memory, and PCIe behaviour of the real production HLT1 reconstruction closely enough that the isolated 41 MHz GPU peak still holds inside the complete chain.
What would settle it
Run the production physics sequence — not a passthrough selection — at the design 30 MHz input rate on the full 164-node cluster for a sustained period, and check for zero event loss and bounded buffer occupancy. If buffer depths grow without limit or events are dropped, the 32 Tbps claim fails.
If this is right
- The trigger now uses the full detector information instead of a subset from specialized sensors, reducing selection bias and improving physics efficiency.
- The 36% headroom at the GPU stage means additional reconstruction algorithms can be added to HLT1 without hardware changes, expanding the physics programme.
- The architecture demonstrates that custom trigger electronics and deep-buffered switches can be replaced with commodity HPC components, a path available to future experiments.
- The zero-copy all-to-all event-building pattern is a scalable solution for any DAQ facing the incast problem at multi-Tbps rates.
- The cross-architecture (CUDA/HIP/CPU) trigger code and the static-memory design make the software portable to future accelerator hardware.
Where Pith is reading between the lines
- If the passthrough-based full-system test faithfully mimics the memory and PCIe profiles of production reconstruction, the 36% GPU headroom suggests the 32 Tbps claim would survive a full end-to-end production test; that test is not reported here.
- The same converged design points toward the planned LHCb Upgrade II: each node has one free PCIe slot, so adding a third GPU per server could raise throughput further without network changes.
- One could expect other LHC experiments to revisit the idea of a trigger-free readout as GPU performance per watt keeps climbing, since the main bottleneck shown here is memory bandwidth (160 GB/s per node), not compute.
- Because the selection is software-defined, physics changes can be deployed by recompiling the trigger rather than redesigning electronics, shortening the loop between theory and data-taking.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents the Run-3 LHCb data acquisition and first-level trigger (HLT1) architecture: 164 servers equipped with FPGA readout cards, an InfiniBand fat-tree event builder using all-to-all personalized exchange, and a fully GPU-based HLT1 implemented in Allen. The authors claim that this converged architecture processes the full 32 Tbps of detector data in real time at 30 MHz with a 30:1 filter, the highest software data rate in any physics experiment. Measurements include single-server scaling across GPUs, multi-server event-builder scaling, an isolated 41 MHz peak for the production HLT1 sequence, and a full-system integration test at 30 MHz using a data generator and a passthrough selection with 30:1 acceptance.
Significance. If substantiated, this is a major engineering milestone for real-time data processing in high-energy physics: a compact, COTS-based DAQ that handles an order of magnitude more software input data than other LHC experiments, with a production trigger running entirely on GPUs. The paper has clear strengths: a full-scale cluster deployment, explicit instrumentation of throughput and resource utilization, a useful comparison table across experiments, and a detailed description of zero-copy, scheduling, and memory-management techniques. The main weakness is that the end-to-end 32 Tbps claim rests on a full-system test with a passthrough selection rather than the production HLT1 reconstruction, so the isolated 41 MHz GPU headroom has not been shown to survive the complete data path.
major comments (2)
- [§V-B and §V-C] The headline claim (abstract and §VI) that the system processes the full 32 Tbps in real time is not directly demonstrated. The full-system integration test in §V-C uses a 'passthrough selection with an acceptance rate of 30:1' and a data generator, not the production HLT1 reconstruction sequence. The production sequence is measured at 41 MHz 'in isolation' (§V-B), i.e., without event-building traffic, PCIe/DMA contention, or shared-memory backpressure from the full chain. The 36% headroom is necessary but not sufficient: heavier GPU compute could alter kernel durations, memory access patterns, and buffer occupancy, potentially introducing backpressure into the event-builder or contending with DMA transfers. To support the end-to-end claim, the authors should either run the production HLT1 sequence in the full-system test at 30 MHz, or provide quantitative evidence (e.g., latency/backpre
- [Table 1] There is a numerical inconsistency in the central throughput figure. Table 1 lists the LHCb event size as 150 kB at 30 MHz, which corresponds to 4.5 TB/s = 36 Tbps, not 32 Tbps. The abstract and §I meanwhile state 32 Tbps throughout. If the average event size is actually ~133 kB (with 150 kB as a maximum), the table should say so; if the 32 Tbps figure is the design input rate, the event-size/rate row should be consistent. Because 32 Tbps is the paper's primary quantitative claim, this arithmetic discrepancy must be resolved.
minor comments (4)
- [§V-C] The sentence 'The data generator produces additional memory pressure that does not affect the performance of our application' is an unsupported assertion. A comparison run without the generator, or a quantification of the generator's memory footprint, would make the claim verifiable.
- [§IV-C] The manuscript defines peak versus sustained performance but does not report run-to-run variability, error bars, or the number of measurement repetitions. For an engineering performance paper, a statement of measurement uncertainty (even approximate) would strengthen the reported 41 MHz and 30 MHz numbers.
- [§III-C] The acronym MIPS (5,596,800) is used without expansion. Please define it at first use (million instructions per second, presumably).
- [Throughout] Several typos and spacing artifacts appear in the text: 'A TLAS', 'V ertex Locator', 'zeroMQ' (should be ZeroMQ), 'STL thread' (should be std::thread). These should be corrected during copyediting.
Circularity Check
No significant circularity: the 32 Tbps claim rests on in-paper measurements; self-citations are contextual.
full rationale
This is an empirical performance paper rather than a derivation. The central claim—that the converged architecture processes 32 Tbps in real time—is supported by in-paper measurements: the isolated HLT1 peak of 41 MHz with 36% headroom (Sec. V-B), the event-builder scaling test (Fig. 6), and the full integration test at 30 MHz (Sec. V-C). No parameter is fitted and then renamed as a prediction, and no result is defined in terms of another result by construction. The full-system test uses a passthrough selection with a 30:1 acceptance rather than the production HLT1 sequence; this is a validity or representativeness caveat about whether the end-to-end benchmark exercises the same GPU/data-movement behavior as full reconstruction, but it is not a circularity. The self-citations in the paper ([31], [39], [40], [41], [43]) describe the DAQ system, scheduler, Allen framework, and hardware-choice comparison; they are contextual and are not load-bearing for the main throughput claim, which is measured directly in this paper. Accordingly, no circular step is identified.
Axiom & Free-Parameter Ledger
free parameters (3)
- Nominal input rate / event size model =
30 MHz, ~150 kB per event
- HLT1 rejection factor =
30:1
- Event batch size =
600-800 events
axioms (5)
- domain assumption The 'realistic event size model' and nominal 30 MHz non-empty collision rate are representative of actual LHCb Run 3 data load.
- ad hoc to paper The full-system passthrough selection with 30:1 acceptance reproduces the I/O, memory, and PCIe behavior of the production HLT1 physics sequence.
- domain assumption All-to-all personalized exchange over the InfiniBand fat-tree sustains full 32 Tbps without incast with the implemented scheduling.
- domain assumption The GPU HLT1 precision choices (single precision, restricted Kalman filter region, two half-precision pattern-recognition paths) preserve the physics efficiency required by LHCb.
- domain assumption The custom FPGA PCIe receiver cards and commercial NICs deliver the stated per-node rates (150 Gb/s FPGA, 100 Gb/s NIC) under full-system conditions.
read the original abstract
The LHCb detector at the Large Hadron Collider has been upgraded to acquire an unprecedented 32 Tbps of particle-collision data to provide new insights in the High Energy Physics domain. The data produced by the detector is filtered in real-time to select interesting collisions. As part of the upgrade, a pre-filtering stage has been removed leading to a factor 40 increase in data rate. To deal with the high throughput demands of LHCb real-time data processing, we present an off-the-shelf network architecture using zero-copy techniques in conjunction with an efficient, fully-GPU-based filter. Our converged architecture is able to process the full 32 Tbps of particle-collision data in real-time, the highest in any physics experiment to date. Our result extends the reach of the LHCb physics programme and sets a new standard for real-time data processing at particle physics experiments.
Figures
Reference graph
Works this paper leans on
-
[1]
Benedikt, P
M. Benedikt, P . Collier, V . Mertens, J. Poole, and K. Schindl,LHC Design Report, ser. CERN Y ellow Reports: Monographs. Geneva: CERN, 2004. [Online]. Available: https://cds.cern.ch/record/823808
2004
-
[2]
Fruhwirth and M
R. Fruhwirth and M. Regler,Data analysis techniques for high- energy physics. Cambridge University Press, 2000. [Online]. Available: http://inspirehep.net/record/299776?ln=es
2000
-
[3]
A. Strandlie and R. Frühwirth, ‘‘Track and vertex reconstruction: From classical to adaptive methods,’’Reviews of Modern Physics, vol. 82, no. 2, pp. 1419–1458, May 2010, publisher: American Physical Society. [On- line]. Available: https://link.aps.org/doi/10.1103/RevModPhys.82.1419
-
[4]
N. V . Canudas, M. C. Gómez, X. Vilasís-Cardona, and E. G. Ribé, ‘‘Graph Clustering: a graph-based clustering algorithm for the electromagnetic calorimeter in LHCb,’’ Dec. 2022, arXiv:2212.11061 [hep-ex, physics:physics]. [Online]. Available: http://arxiv.org/abs/2212.11061
Pith/arXiv arXiv 2022
-
[5]
LHCb Collaboration, ‘‘The LHCb Detector at the LHC,’’Journal of Instrumentation, vol. 3, no. 08, pp. S08 005–S08 005, Aug. 2008, publisher: IOP Publishing. [Online]. Available: http://stacks.iop.org/1748- 0221/3/i=08/a=S08005?key=crossref.358ac80e1a6b6ba36f68c89dc0c4bed4
2008
-
[6]
——, ‘‘LHCb detector performance,’’International Journal of Modern Physics A, vol. 30, no. 07, p. 1530022, mar 2015. [Online]. Available: https://www.worldscientific.com/doi/abs/10.1142/S0217751X15300227
- [7]
-
[8]
——, ‘‘Upgrade Software and Computing,’’ Tech. Rep., mar 2018. [Online]. Available: https://cds.cern.ch/record/2310827?ln=es
arXiv 2018
-
[9]
P . C. Broekema, J. J. D. Mol, R. Nijboer, A. S. van Amesfoort, M. A. Brentjens, G. M. Loose, W. F. A. Klijn, and J. W. Romein, ‘‘Cobalt: A GPU-based correlator and beamformer for LOFAR,’’Astronomy and Computing, vol. 23, pp. 180–192, Apr. 2018, arXiv:1801.04834 [astro-ph]. [Online]. Available: http://arxiv.org/abs/1801.04834
Pith/arXiv arXiv 2018
-
[10]
Romein, P
J. Romein, P . Broekema, J. Mol, and R. V an Nieuwpoort, ‘‘The LOFAR Correlator: Implementation and Performance Analysis,’’ACM SIGPLAN Notices, vol. 45, pp. 169–178, May 2010
2010
-
[11]
CHIME Collaboration, ‘‘An Overview of CHIME, the Canadian Hydrogen Intensity Mapping Experiment,’’The Astrophysical Journal Supplement Series, vol. 261, no. 2, p. 29, Aug. 2022, arXiv:2201.07869 [astro-ph]. [Online]. Available: http://arxiv.org/abs/2201.07869
Pith/arXiv arXiv 2022
-
[12]
L. Aggarwal, S. Banerjee, S. Bansal, F. Bernlochner, M. Bertemes, V . Bhardwaj, A. Bondar, T. E. Browder, L. Cao, M. Campajola, G. Casarosa, C. Cecchi, R. Cheaib, G. De Pietro, A. Di Canto, M. Dorigo, P . Feichtinger, T. Ferber, B. Fulsom, M. García, G. Gaudino, A. Gaz, A. Glazov, S. Granderath, E. Graziani, D. Greenwald, P . Goldenzweig, I. Heredia, M. H...
Pith/arXiv arXiv 2022
-
[13]
A TLAS Collaboration, ‘‘Operation of the A TLAS trigger system in Run 2,’’Journal of Instrumentation, vol. 15, no. 10, pp. P10 004–P10 004, Oct. 2020, arXiv:2007.12539 [hep-ex, physics:physics]. [Online]. Available: http://arxiv.org/abs/2007.12539
Pith/arXiv arXiv 2020
-
[14]
Fontanesi, ‘‘New trigger strategies for CMS during Run 3,’’ CERN, Geneva, Tech
E. Fontanesi, ‘‘New trigger strategies for CMS during Run 3,’’ CERN, Geneva, Tech. Rep., 2022. [Online]. Available: https://cds.cern.ch/record/2842439
arXiv 2022
-
[15]
D. Rohr, ‘‘Usage of GPUs in ALICE Online and Offline processing during LHC Run 3,’’EPJ Web of Conferences, vol. 251, pp. 04 026–04 026, Aug. 2021. [Online]. Available: https://www.epj- conferences.org/10.1051/epjconf/202125104026
arXiv 2021
-
[16]
2023, arXiv:2302.01238 [hep-ex, physics:nucl-ex, physics:physics]
ALICE Collaboration, ‘‘ALICE upgrades during the LHC Long Shutdown 2,’’ Feb. 2023, arXiv:2302.01238 [hep-ex, physics:nucl-ex, physics:physics]. [Online]. Available: http://arxiv.org/abs/2302.01238
Pith/arXiv arXiv 2023
-
[17]
——, ‘‘The alice experiment at the cern lhc,’’Journal of Instrumentation, vol. 3, no. 08, p. S08002, aug 2008. [Online]. Available: https://dx.doi.org/10.1088/1748-0221/3/08/S08002
-
[18]
A TLAS Collaboration, ‘‘The atlas experiment at the cern large hadron collider,’’Journal of Instrumentation, vol. 3, no. 08, p. S08003, aug 2008. [Online]. Available: https://dx.doi.org/10.1088/1748-0221/3/08/S08003
-
[19]
CMS Collaboration, ‘‘The cms experiment at the cern lhc,’’Journal of Instrumentation, vol. 3, no. 08, p. S08004, aug 2008. [Online]. Available: https://dx.doi.org/10.1088/1748-0221/3/08/S08004
-
[20]
Technical design report
A TLAS Collaboration,ATLAS level-1 trigger: Technical Design Report, ser. Technical design report. A TLAS. Geneva: CERN, 1998. [Online]. Available: http://cds.cern.ch/record/381429 10 VOLUME 11, 2023
1998
-
[21]
CMS Collaboration, ‘‘The CMS trigger system,’’Journal of Instrumentation, vol. 12, no. 01, pp. P01 020–P01 020, Jan. 2017, arXiv:1609.02366 [hep-ex, physics:physics]. [Online]. Available: http://arxiv.org/abs/1609.02366
Pith/arXiv arXiv 2017
-
[22]
Jenni, M
P . Jenni, M. Nessi, M. Nordberg, and K. Smith,ATLAS high-level trigger , data-acquisition and controls: Technical Design Report, ser. Technical design report. A TLAS. Geneva: CERN, 2003. [Online]. Available: https://cds.cern.ch/record/616089
2003
-
[23]
CMS Collaboration, ‘‘Developing gpu-compliant algorithms for cms ecal local reconstruction during lhc run 3 and phase 2,’’Journal of Physics: Conference Series, vol. 2438, no. 1, p. 012027, feb 2023. [Online]. Available: https://dx.doi.org/10.1088/1742-6596/2438/1/012027
-
[24]
F. Costa, A. Kluge, P . V . Vyvre, and for the ALICE Collaboration, ‘‘The detector read-out in alice during run 3 and 4,’’Journal of Physics: Conference Series, vol. 898, no. 3, p. 032011, oct 2017. [Online]. Available: https://dx.doi.org/10.1088/1742-6596/898/3/032011
-
[25]
D. Rohr, S. Gorbunov, A. Szostak, M. Kretz, T. Kollegger, T. Breitner, and T. Alt, ‘‘ALICE HLT TPC Tracking of Pb-Pb Events on GPUs,’’Journal of Physics: Conference Series, vol. 396, no. 1, pp. 012 044–012 044, Dec. 2012
2012
-
[26]
E. Kou, P . Urquijo, B2TiP Theory community, and Belle II Collaboration, The Belle II Physics Book, 2018
2018
-
[27]
Phanishayee, E
A. Phanishayee, E. Krevat, V . V asudevan, D. G. Andersen, G. R. Ganger, G. A. Gibson, and S. Seshan, ‘‘Measurement and analysis of tcp throughput collapse in cluster-based storage systems,’’ inProceedings of the 6th USENIX Conference on File and Storage Technologies, ser. FAST’08. USA: USENIX Association, 2008
2008
-
[28]
Y . Chen, R. Griffit, D. Zats, and R. H. Katz, ‘‘Under- standing tcp incast and its implications for big data work- loads,’’ EECS Department, University of California, Berkeley, Tech. Rep. UCB/EECS-2012-40, Apr 2012. [Online]. Available: http://www2.eecs.berkeley.edu/Pubs/TechRpts/2012/EECS-2012-40.html
2012
-
[29]
J. P . Cachemiche, P . Y . Duval, F. Hachon, R. Le Gac, and F. Réthoré, ‘‘The PCIe-based readout system for the LHCb experiment,’’JINST, vol. 11, no. 02, p. P02013, 2016
2016
-
[30]
E. Zahavi, G. Johnson, D. J. Kerbyson, and M. Lang, ‘‘Optimized infiniband fat-tree routing for shift all-to-all communication patterns,’’Concurrency and Computation: Practice and Experience, vol. 22, no. 2, pp. 217–231, 2010. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1002/cpe.1527
-
[31]
Matev, N
R. Matev, N. Nolte, and A. Pearce, ‘‘Configuration and scheduling of the LHCb trigger application,’’EPJ Web of Conferences, vol. 245, p. 05004, 2020, publisher: EDP Sciences
2020
-
[32]
Mertens, ‘‘The Easiest Hard Problem: Number Partitioning,’’ arXiv:cond-mat/0310317, Oct
S. Mertens, ‘‘The Easiest Hard Problem: Number Partitioning,’’ arXiv:cond-mat/0310317, Oct. 2003, arXiv: cond-mat/0310317. [Online]. Available: http://arxiv.org/abs/cond-mat/0310317
Pith/arXiv arXiv 2003
-
[33]
D. H. Cámpora Pérez, N. Neufeld, and A. Riscos Núñez, ‘‘Search by triplet: An efficient local track reconstruction algorithm for parallel architectures,’’ Journal of Computational Science, vol. 54, p. 101422, sep 2021
2021
-
[34]
P . Fernandez Declara, D. H. Cámpora Pérez, J. Garcia-Blas, D. vom Bruch, J. D. Garcia, and N. Neufeld, ‘‘A parallel-computing algorithm for high-energy physics particle tracking and decoding using GPU architectures,’’IEEE Access, pp. 91 612–91 626, 2019. [Online]. Available: https://ieeexplore.ieee.org/document/8756134/
arXiv 2019
-
[35]
R. E. Kalman, ‘‘A New Approach to Linear Filtering and Prediction Problems,’’Journal of Basic Engineering, vol. 82, no. 1, p. 35, mar 1960
1960
-
[36]
D. H. Cámpora Pérez and O. Awile, ‘‘An efficient low-rank Kalman filter for modern SIMD architectures,’’Concurrency and Computation: Practice and Experience, vol. 30, no. 23, p. e4483, dec 2018. [Online]. Available: http://doi.wiley.com/10.1002/cpe.4483
- [37]
- [38]
- [39]
-
[40]
R. Aaij, J. Albrecht, M. Belous, P . Billoir, T. Boettcher, A. B. Rodríguez, D. vom Bruch, D. H. Cámpora Pérez, A. C. Vidal, D. C. Craik, P . F. Declara, L. Funke, V . V . Gligorov, B. Jashal, N. Kazeev, D. M. Santos, F. Pisani, D. Pliushchenko, S. Popov, R. Quagliani, M. Rangel, F. Reiss, C. S. Mayordomo, R. Schwemmer, M. Sokoloff, H. Stevens, A. Ustyuzh...
Pith/arXiv arXiv 2020
-
[41]
Pisani, T
F. Pisani, T. Colombo, P . Durante, M. Frank, C. Gaspar, L. G. Cardoso, N. Neufeld, and A. Perro, ‘‘Design and commissioning of the first 32 tbit/s event-builder,’’IEEE Transactions on Nuclear Science, pp. 1–1, 2023
2023
-
[42]
C. E. Leiserson, ‘‘Fat-trees: Universal networks for hardware-efficient supercomputing,’’ vol. C-34, no. 10, pp. 892–901
-
[43]
R. Aaij, M. Adinolfi, S. Aiola, S. Akar, J. Albrecht, M. Alexander, S. Amato, Y . Amhis, F. Archilli, M. Bala, G. Bassi, L. Bian, M. P . Blago, T. Boettcher, A. Boldyrev, S. Borghi, A. B. Rodriguez, L. Calefice, M. C. Gomez, D. H. Cámpora Pérez, A. Cardini, M. Cattaneo, V . Chobanova, G. Ciezarek, X. C. Vidal, J. L. Cobbledick, J. A. B. Coelho, T. Colombo...
Pith/arXiv arXiv 2022
-
[44]
Colombo, A
T. Colombo, A. Amihalachioaei, K. Arnaud, F. Alessio, L. Brarda, J.- P . Cachemiche, D. H. Cámpora Pérez, S. Cap, L. Cardoso, F. Cindolo, M. Daoudi, P . Durante, P .-Y . Duval, C. Faerber, P . Fernández, M. Frank, D. Galli, C. Gaspar, F. Hachon, M. Jevaud, B. Jost, R. Le Gac, M. Man- zali, U. Marconi, H. Mohammed, N. Neufeld, F. Pisani, L. Promberger, C. ...
2020
-
[45]
LHCb Collaboration, ‘‘Framework TDR for the LHCb Upgrade II,’’ CERN, Geneva, Tech. Rep., 2021. [Online]. Available: http://cds.cern.ch/record/2776420
arXiv 2021
-
[46]
Rep., 2016, iSBN 978-92-9083-494-6
——, ‘‘Physics case for an LHCb Upgrade II - Opportunities in flavour physics, and beyond, in the HL-LHC era,’’ CERN, Geneva, Tech. Rep., 2016, iSBN 978-92-9083-494-6. [Online]. Available: https://cds.cern.ch/record/2636441 DANIEL HUGO CÁMPORA PÉREZreceived his Ph.D. from the University of Sevilla. He has worked at CERN since 2010, contributing mainly to t...
arXiv 2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.