REVIEW 4 cited by
hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Accessible machine learning algorithms, software, and diagnostic tools for energy-efficient devices and systems are extremely valuable across a broad range of application domains. In scientific domains, real-time near-sensor processing can drastically improve experimental design and accelerate scientific discoveries. To support domain scientists, we have developed hls4ml, an open-source software-hardware codesign workflow to interpret and translate machine learning algorithms for implementation with both FPGA and ASIC technologies. We expand on previous hls4ml work by extending capabilities and techniques towards low-power implementations and increased usability: new Python APIs, quantization-aware pruning, end-to-end FPGA workflows, long pipeline kernels for low power, and new device backends include an ASIC workflow. Taken together, these and continued efforts in hls4ml will arm a new generation of domain scientists with accessible, efficient, and powerful tools for machine-learning-accelerated discovery.
Forward citations
Cited by 4 Pith papers
-
KANEL\'E: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation
Quantized, pruned Kolmogorov-Arnold Networks can be compiled directly into FPGA lookup tables, achieving extreme latency/resource reductions and matching state-of-the-art LUT-based networks on several benchmarks.
-
Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml
A 32-to-2 autoencoder for LHCb PicoCal pulses was synthesized on a Microchip PolarFire FPGA via a new hls4ml backend, achieving 25 ns latency and 3.1% LUT usage per channel in simulation.
-
wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
A new open benchmark with 683,176 synthesized hls4ml designs plus GNN/transformer surrogates that predict FPGA resources/latency accurately in-distribution but poorly on out-of-distribution scientific models.
-
Neural Network Acceleration on MPSoC board: Integrating SLAC's SNL, Rogue Software and Auto-SNL
SNL, aided by the new Auto-SNL converter, achieves lower latency than hls4ml on 3 of 4 benchmark models, at the cost of higher BRAM/FF in some designs.
Discussion (0). Sign in to comment.