Pith. sign in

REVIEW

Scaling Up Online Speech Recognition Using ConvNets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.09727 v1 pith:NQ33DKVQ submitted 2020-01-27 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords latencyaccuracydesignonlinerecognitionspeechsystemthroughput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We design an online end-to-end speech recognition system based on Time-Depth Separable (TDS) convolutions and Connectionist Temporal Classification (CTC). We improve the core TDS architecture in order to limit the future context and hence reduce latency while maintaining accuracy. The system has almost three times the throughput of a well tuned hybrid ASR baseline while also having lower latency and a better word error rate. Also important to the efficiency of the recognizer is our highly optimized beam search decoder. To show the impact of our design choices, we analyze throughput, latency, accuracy, and discuss how these metrics can be tuned based on the user requirements.

Discussion (0). Sign in to comment.

Pith tools