Pith. sign in

REVIEW

wav2letter++: The Fastest Open-source Speech Recognition System

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.07625 v1 pith:R4ZD73S2 submitted 2018-12-18 cs.CL

classification cs.CL
keywords wav2letterrecognitionspeechopen-sourcefastestframeworksothersystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces wav2letter++, the fastest open-source deep learning speech recognition framework. wav2letter++ is written entirely in C++, and uses the ArrayFire tensor library for maximum efficiency. Here we explain the architecture and design of the wav2letter++ system and compare it to other major open-source speech recognition systems. In some cases wav2letter++ is more than 2x faster than other optimized frameworks for training end-to-end neural networks for speech recognition. We also show that wav2letter++'s training times scale linearly to 64 GPUs, the highest we tested, for models with 100 million parameters. High-performance frameworks enable fast iteration, which is often a crucial factor in successful research and model tuning on new datasets and tasks.

Discussion (0). Sign in to comment.

Pith tools