Pith. sign in

REVIEW 6 cited by

Fast inference of deep neural networks in FPGAs for particle physics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.06913 v3 pith:ZV75B2XG submitted 2018-04-16 physics.ins-det cs.CVhep-exstat.ML

Fast inference of deep neural networks in FPGAs for particle physics

classification physics.ins-det cs.CVhep-exstat.ML
keywords physicsfpgasneuralparticleinferencelatencynetworkexample
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent results at the Large Hadron Collider (LHC) have pointed to enhanced physics capabilities through the improvement of the real-time event processing techniques. Machine learning methods are ubiquitous and have proven to be very powerful in LHC physics, and particle physics as a whole. However, exploration of the use of such techniques in low-latency, low-power FPGA hardware has only just begun. FPGA-based trigger and data acquisition (DAQ) systems have extremely low, sub-microsecond latency requirements that are unique to particle physics. We present a case study for neural network inference in FPGAs focusing on a classifier for jet substructure which would enable, among many other physics scenarios, searches for new dark sector particles and novel measurements of the Higgs boson. While we focus on a specific example, the lessons are far-reaching. We develop a package based on High-Level Synthesis (HLS) called hls4ml to build machine learning models in FPGAs. The use of HLS increases accessibility across a broad user community and allows for a drastic decrease in firmware development time. We map out FPGA resource usage and latency versus neural network hyperparameters to identify the problems in particle physics that would benefit from performing neural network inference with FPGAs. For our example jet substructure model, we fit well within the available resources of modern FPGAs with a latency on the scale of 100 ns.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

    quant-ph 2026-07 conditional novelty 6.0

    A dilated causal CNN quantized to fixed point and synthesized to an FPGA detects charge jumps in superconducting qubits at 6.19 μs latency with 0.843 efficiency, close to the 0.866 of the offline χ2 method on |Δq|∈[0.1,0.5]e.

  2. On-chip probabilistic inference for charged-particle tracking at the sensor edge

    physics.ins-det 2026-02 unverdicted novelty 6.0

    Neural networks integrated into silicon sensor front-end electronics can regress charged-particle hit positions and angles with calibrated uncertainties from single-layer data while satisfying hardware constraints on ...

  3. Real-time graph neural networks on FPGAs for the Belle II electromagnetic calorimeter

    physics.ins-det 2026-02 conditional novelty 6.0

    A GNN-based calorimeter clustering and signal classifier ran on an FPGA inside the Belle II L1 trigger readout path, improving position resolution and photon separation at the cost of exceeding the trigger decision latency.

  4. SparsePixels: Efficient Convolution for Sparse Data on FPGAs

    cs.AR 2025-12 conditional novelty 6.0

    A fixed-budget sparse-convolution FPGA framework runs CNNs on <=20 of ~4000 pixels, achieving 0.665 us inference for MicroBooNE with a 73x speedup and ~2% AUC loss.

  5. wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

    cs.LG 2025-11 conditional novelty 6.0

    A new open benchmark with 683,176 synthesized hls4ml designs plus GNN/transformer surrogates that predict FPGA resources/latency accurately in-distribution but poorly on out-of-distribution scientific models.

  6. Hybrid neural denoising for resource-efficient near- and sub-threshold radio triggering of extensive air showers

    astro-ph.IM 2026-05 unverdicted novelty 5.0

    Hybrid convolutional denoiser plus classifier improves weak-signal radio triggering for air showers in noisy environments while meeting FPGA hardware constraints.