Pith. sign in

REVIEW 4 cited by

hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.05579 v3 pith:BGU2FTSL submitted 2021-03-09 cs.LG cs.ARphysics.ins-det

classification cs.LGcs.ARphysics.ins-det
keywords hls4mllearningmachinescientificworkflowaccessiblealgorithmsasic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accessible machine learning algorithms, software, and diagnostic tools for energy-efficient devices and systems are extremely valuable across a broad range of application domains. In scientific domains, real-time near-sensor processing can drastically improve experimental design and accelerate scientific discoveries. To support domain scientists, we have developed hls4ml, an open-source software-hardware codesign workflow to interpret and translate machine learning algorithms for implementation with both FPGA and ASIC technologies. We expand on previous hls4ml work by extending capabilities and techniques towards low-power implementations and increased usability: new Python APIs, quantization-aware pruning, end-to-end FPGA workflows, long pipeline kernels for low power, and new device backends include an ASIC workflow. Taken together, these and continued efforts in hls4ml will arm a new generation of domain scientists with accessible, efficient, and powerful tools for machine-learning-accelerated discovery.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KANEL\'E: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation

    cs.AR 2025-12 conditional novelty 7.0 of 10

    Quantized, pruned Kolmogorov-Arnold Networks can be compiled directly into FPGA lookup tables, achieving extreme latency/resource reductions and matching state-of-the-art LUT-based networks on several benchmarks.

  2. Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

    hep-ex 2026-02 conditional novelty 6.0 of 10

    A 32-to-2 autoencoder for LHCb PicoCal pulses was synthesized on a Microchip PolarFire FPGA via a new hls4ml backend, achieving 25 ns latency and 3.1% LUT usage per channel in simulation.

  3. wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

    cs.LG 2025-11 conditional novelty 6.0 of 10

    A new open benchmark with 683,176 synthesized hls4ml designs plus GNN/transformer surrogates that predict FPGA resources/latency accurately in-distribution but poorly on out-of-distribution scientific models.

  4. Neural Network Acceleration on MPSoC board: Integrating SLAC's SNL, Rogue Software and Auto-SNL

    cs.LG 2025-08 conditional novelty 5.0 of 10

    SNL, aided by the new Auto-SNL converter, achieves lower latency than hls4ml on 3 of 4 benchmark models, at the cost of higher BRAM/FF in some designs.

Pith tools