Pith. sign in

REVIEW

Low Latency ASR for Simultaneous Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.09891 v1 pith:66WMOCEN submitted 2020-03-22 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords latencydecodingspeechtranslationwordperformancereducingsimultaneous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

User studies have shown that reducing the latency of our simultaneous lecture translation system should be the most important goal. We therefore have worked on several techniques for reducing the latency for both components, the automatic speech recognition and the speech translation module. Since the commonly used commitment latency is not appropriate in our case of continuous stream decoding, we focused on word latency. We used it to analyze the performance of our current system and to identify opportunities for improvements. In order to minimize the latency we combined run-on decoding with a technique for identifying stable partial hypotheses when stream decoding and a protocol for dynamic output update that allows to revise the most recent parts of the transcription. This combination reduces the latency at word level, where the words are final and will never be updated again in the future, from 18.1s to 1.1s without sacrificing performance in terms of word error rate.

Discussion (0). Continue with ORCID to comment.

Pith tools