REVIEW 13 cited by
Fast inference of deep neural networks in FPGAs for particle physics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent results at the Large Hadron Collider (LHC) have pointed to enhanced physics capabilities through the improvement of the real-time event processing techniques. Machine learning methods are ubiquitous and have proven to be very powerful in LHC physics, and particle physics as a whole. However, exploration of the use of such techniques in low-latency, low-power FPGA hardware has only just begun. FPGA-based trigger and data acquisition (DAQ) systems have extremely low, sub-microsecond latency requirements that are unique to particle physics. We present a case study for neural network inference in FPGAs focusing on a classifier for jet substructure which would enable, among many other physics scenarios, searches for new dark sector particles and novel measurements of the Higgs boson. While we focus on a specific example, the lessons are far-reaching. We develop a package based on High-Level Synthesis (HLS) called hls4ml to build machine learning models in FPGAs. The use of HLS increases accessibility across a broad user community and allows for a drastic decrease in firmware development time. We map out FPGA resource usage and latency versus neural network hyperparameters to identify the problems in particle physics that would benefit from performing neural network inference with FPGAs. For our example jet substructure model, we fit well within the available resources of modern FPGAs with a latency on the scale of 100 ns.
Forward citations
Cited by 13 Pith papers
-
Oraqle: An Empirical Analysis of Qubit Readout and Discriminators in Quantum Error Correction
Using real 5-qubit traces, this study shows readout windows can be cut to ~600 ns with negligible QEC penalty and small discriminators match large ones.
-
Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network
A dilated causal CNN quantized to fixed point and synthesized to an FPGA detects charge jumps in superconducting qubits at 6.19 μs latency with 0.843 efficiency, close to the 0.866 of the offline χ2 method on |Δq|∈[0.1,0.5]e.
-
Real-time graph neural networks on FPGAs for the Belle II electromagnetic calorimeter
A GNN-based calorimeter clustering and signal classifier ran on an FPGA inside the Belle II L1 trigger readout path, improving position resolution and photon separation at the cost of exceeding the trigger decision latency.
-
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
A fixed-budget sparse-convolution FPGA framework runs CNNs on <=20 of ~4000 pixels, achieving 0.665 us inference for MicroBooNE with a 73x speedup and ~2% AUC loss.
-
wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
A new open benchmark with 683,176 synthesized hls4ml designs plus GNN/transformer surrogates that predict FPGA resources/latency accurately in-distribution but poorly on out-of-distribution scientific models.
-
Learning Symmetry-Independent Jet Representations via Jet-Based Joint Embedding Predictive Architecture
J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.
-
Lund jet images from generative and cycle-consistent adversarial networks
A least-squares GAN trained on Lund jet plane images reproduces the simulated jet substructure distribution to within a few percent, and a CycleGAN maps between jet categories such as parton-level vs detector-level or...
-
End-to-end workflow for machine learning-based qubit readout with QICK and hls4ml
A quantized neural network deployed on QICK's FPGA performs single-transmon readout at 96% fidelity, 32 ns inference latency, and under 16% LUT overhead, matching classical thresholding methods.
-
Neural Architecture Codesign for Fast Physics Applications
An automated two-stage neural architecture search and compression pipeline discovers FPGA-efficient models for Bragg peak finding and jet classification, beating or matching hand-crafted baselines on accuracy, latency...
-
JEDI-net: a jet identification algorithm based on interaction networks
JEDI-net, an interaction-network jet tagger, outperforms DNN, CNN, and GRU taggers on a five-class simulated LHC jet dataset.
-
Continuous-variable photonic quantum extreme learning machines for fast collider-data selection
A Gaussian photonic QELM with displacement encoding and quadrature/photon-number readout produces polynomial features that, under a linear readout, match or beat small MLPs on top-jet and Higgs classification.
-
Comparative Analysis of FPGA and GPU Performance for Machine Learning-Based Track Reconstruction at LHCb
Using hls4ml, the authors estimate an Alveo U250 FPGA can process the LHCb track-embedding MLP at 1.1 million events per second, exceeding the measured 0.82 million events per second of an RTX 3090 GPU at lower power.
-
Analysis of Hardware Synthesis Strategies for Machine Learning in Collider Trigger and Data Acquisition
For small-to-medium fully connected VAE trigger encoders on an Alveo U200 FPGA, hls4ml achieves lower latency and SNL achieves lower LUT and FF usage at matched latency.
Discussion (0). Continue with ORCID to comment.