Pith. sign in

REVIEW 10 cited by

hls4ml: An Open-Source Codesign Workflow to Empower Scientific Low-Power Machine Learning Devices

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.05579 v3 pith:BGU2FTSL submitted 2021-03-09 cs.LG cs.ARphysics.ins-det

classification cs.LGcs.ARphysics.ins-det
keywords hls4mllearningmachinescientificworkflowaccessiblealgorithmsasic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accessible machine learning algorithms, software, and diagnostic tools for energy-efficient devices and systems are extremely valuable across a broad range of application domains. In scientific domains, real-time near-sensor processing can drastically improve experimental design and accelerate scientific discoveries. To support domain scientists, we have developed hls4ml, an open-source software-hardware codesign workflow to interpret and translate machine learning algorithms for implementation with both FPGA and ASIC technologies. We expand on previous hls4ml work by extending capabilities and techniques towards low-power implementations and increased usability: new Python APIs, quantization-aware pruning, end-to-end FPGA workflows, long pipeline kernels for low power, and new device backends include an ASIC workflow. Taken together, these and continued efforts in hls4ml will arm a new generation of domain scientists with accessible, efficient, and powerful tools for machine-learning-accelerated discovery.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Design Rules for Extreme-Edge Scientific Computing on AI Engines

    cs.AR 2026-04 unverdicted novelty 7.0 of 10

    AI Engines enable larger low-latency neural networks for extreme-edge scientific computing on FPGAs than programmable logic, via a new latency-adjusted resource equivalence metric and tailored optimizations.

  2. KANEL\'E: Kolmogorov-Arnold Networks for Efficient LUT-based Evaluation

    cs.AR 2025-12 conditional novelty 7.0 of 10

    Quantized, pruned Kolmogorov-Arnold Networks can be compiled directly into FPGA lookup tables, achieving extreme latency/resource reductions and matching state-of-the-art LUT-based networks on several benchmarks.

  3. HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN Inference

    cs.AR 2026-04 unverdicted novelty 6.0 of 10

    HGQ-LUT delivers a practical LUT-aware training framework with new tensor-based layers, heterogeneous quantization, and a resource surrogate that automates accuracy-efficiency trade-offs for FPGA DNN inference.

  4. On-chip probabilistic inference for charged-particle tracking at the sensor edge

    physics.ins-det 2026-02 unverdicted novelty 6.0 of 10

    Neural networks integrated into silicon sensor front-end electronics can regress charged-particle hit positions and angles with calibrated uncertainties from single-layer data while satisfying hardware constraints on ...

  5. Enabling Low-Latency Machine learning on Radiation-Hard FPGAs with hls4ml

    hep-ex 2026-02 conditional novelty 6.0 of 10

    A 32-to-2 autoencoder for LHCb PicoCal pulses was synthesized on a Microchip PolarFire FPGA via a new hls4ml backend, achieving 25 ns latency and 3.1% LUT usage per channel in simulation.

  6. wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

    cs.LG 2025-11 conditional novelty 6.0 of 10

    A new open benchmark with 683,176 synthesized hls4ml designs plus GNN/transformer surrogates that predict FPGA resources/latency accurately in-distribution but poorly on out-of-distribution scientific models.

  7. da4ml: Distributed Arithmetic for Real-time Neural Networks on FPGAs

    cs.AR 2025-07 unverdicted novelty 6.0 of 10

    A distributed arithmetic algorithm for CMVM operations on FPGAs reduces area by up to one third and latency for quantized neural networks, integrated into hls4ml.

  8. The impact of source and survey modelling on the connection between [O III] emitters and Ly $\alpha$ forest transmission at z ~ 6

    astro-ph.CO 2026-06 unverdicted novelty 5.0 of 10

    Empirical halo-to-[O III] emitter modeling with realistic JWST survey mocks produces cross-correlations consistent with z~6 data within large scatter, but with a ~10 cMpc offset in the 1D peak.

  9. Neural Network Acceleration on MPSoC board: Integrating SLAC's SNL, Rogue Software and Auto-SNL

    cs.LG 2025-08 conditional novelty 5.0 of 10

    SNL, aided by the new Auto-SNL converter, achieves lower latency than hls4ml on 3 of 4 benchmark models, at the cost of higher BRAM/FF in some designs.

  10. WaveDriver: a Laser Guide Star AO System for HWO

    astro-ph.IM 2026-05 unverdicted novelty 3.0 of 10

    WaveDriver is a laser guide star AO concept whose initial simulations indicate it may be required to meet HWO primary mirror segment stability and low-order wavefront stability requirements.

Pith tools