Pith. sign in

REVIEW 1 cited by

The Sparsity Roofline: Understanding the Hardware Limits of Sparse Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.00496 v2 pith:E2EJAMW7 submitted 2023-09-30 cs.CV cs.LG

classification cs.CVcs.LG
keywords sparsityperformancerooflinesparsespeeduphardwaremodelpatterns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Sparsity Roofline, a visual performance model for evaluating sparsity in neural networks. The Sparsity Roofline jointly models network accuracy, sparsity, and theoretical inference speedup. Our approach does not require implementing and benchmarking optimized kernels, and the theoretical speedup becomes equal to the actual speedup when the corresponding dense and sparse kernels are well-optimized. We achieve this through a novel analytical model for predicting sparse network performance, and validate the predicted speedup using several real-world computer vision architectures pruned across a range of sparsity patterns and degrees. We demonstrate the utility and ease-of-use of our model through two case studies: (1) we show how machine learning researchers can predict the performance of unimplemented or unoptimized block-structured sparsity patterns, and (2) we show how hardware designers can predict the performance implications of new sparsity patterns and sparse data formats in hardware. In both scenarios, the Sparsity Roofline helps performance experts identify sparsity regimes with the highest performance potential.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A co-designed 4-bit quantization and temporal sparsity scheme, with a matching dense-sparse accelerator, speeds up diffusion model inference by 6.91x with 51.5% energy reduction at near-FP32 image quality.

Pith tools