REVIEW 3 cited by
Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
DNN accelerators are often developed and evaluated in isolation without considering the cross-stack, system-level effects in real-world environments. This makes it difficult to appreciate the impact of System-on-Chip (SoC) resource contention, OS overheads, and programming-stack inefficiencies on overall performance/energy-efficiency. To address this challenge, we present Gemmini, an open-source*, full-stack DNN accelerator generator. Gemmini generates a wide design-space of efficient ASIC accelerators from a flexible architectural template, together with flexible programming stacks and full SoCs with shared resources that capture system-level effects. Gemmini-generated accelerators have also been fabricated, delivering up to three orders-of-magnitude speedups over high-performance CPUs on various DNN benchmarks. * https://github.com/ucb-bar/gemmini
Forward citations
Cited by 3 Pith papers
-
High-Level Synthesis of Efficient Pipelines with Visibility Control
Visibility control, a unified publish/observe abstraction, lets sequential programs express diverse pipeline hazard-resolution strategies with near-RTL PPA.
-
ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
ReChisel, an LLM agent with reflection and an escape mechanism for non-progress loops, significantly improves Chisel code generation success rates across five LLMs and three benchmarks.
-
Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core
Float16 full training on a single RISC-V core reaches near-float32 accuracy on tested MLPs with ~50% memory savings, plus a +1.15% LUT cost for scalar Zfh support.
Discussion (0). Continue with ORCID to comment.