REVIEW 4 cited by
Fast convolutional neural networks on FPGAs with hls4ml
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on FPGAs. By extending the hls4ml library, we demonstrate an inference latency of $5\,\mu$s using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.
Forward citations
Cited by 4 Pith papers
-
Local Conformal Predictions for Calibrated Surrogates
FALCON is a novel conformal prediction technique that learns locally calibrated confidence intervals for neural network surrogates modeling LHC scattering amplitudes.
-
SparsePixels: Efficient Convolution for Sparse Data on FPGAs
A fixed-budget sparse-convolution FPGA framework runs CNNs on <=20 of ~4000 pixels, achieving 0.665 us inference for MicroBooNE with a 73x speedup and ~2% AUC loss.
-
Hybrid neural denoising for resource-efficient near- and sub-threshold radio triggering of extensive air showers
Hybrid convolutional denoiser plus classifier improves weak-signal radio triggering for air showers in noisy environments while meeting FPGA hardware constraints.
-
Continuous-variable photonic quantum extreme learning machines for fast collider-data selection
A Gaussian photonic QELM with displacement encoding and quadrature/photon-number readout produces polynomial features that, under a linear readout, match or beat small MLPs on top-jet and Higgs classification.
Discussion (0). Sign in to comment.