REVIEW 3 cited by
MIOpen: An Open Source Library For Deep Learning Primitives
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep Learning has established itself to be a common occurrence in the business lexicon. The unprecedented success of deep learning in recent years can be attributed to: abundance of data, availability of gargantuan compute capabilities offered by GPUs, and adoption of open-source philosophy by the researchers and industry. Deep neural networks can be decomposed into a series of different operators. MIOpen, AMD's open-source deep learning primitives library for GPUs, provides highly optimized implementations of such operators, shielding researchers from internal implementation details and hence, accelerating the time to discovery. This paper introduces MIOpen and provides details about the internal workings of the library and supported features. MIOpen innovates on several fronts, such as implementing fusion to optimize for memory bandwidth and GPU launch overheads, providing an auto-tuning infrastructure to overcome the large design space of problem configurations, and implementing different algorithms to optimize convolutions for different filter and input sizes. MIOpen is one of the first libraries to publicly support the bfloat16 data-type for convolutions, allowing efficient training at lower precision without the loss of accuracy.
Forward citations
Cited by 3 Pith papers
-
Adding MFMA Support to gem5
Added and validated MFMA/MCE support in gem5 for AMD MI200 and MI300 GPUs, achieving 1.5% and 1.3% MAPE against real hardware.
-
Optimizing Winograd Convolution on ARMv8 processors
A hand-optimized, fused Winograd convolution for ARMv8 CPUs reports up to 4.7x to 10.6x speedups over NCNN, NNPACK, FastConv, and ACL on Kunpeng 920, Graviton2, and Phytium 2000+.
-
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.
Discussion (0). Continue with ORCID to comment.