Pith. sign in

REVIEW 1 cited by

ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.04596 v3 pith:MSTFRV4G submitted 2023-04-10 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords translationespnet-st-v2languagemodelsspokentoolkitattentionespnet
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community. ESPnet-ST-v2 supports 1) offline speech-to-text translation (ST), 2) simultaneous speech-to-text translation (SST), and 3) offline speech-to-speech translation (S2ST) -- each task is supported with a wide variety of approaches, differentiating ESPnet-ST-v2 from other open source spoken language translation toolkits. This toolkit offers state-of-the-art architectures such as transducers, hybrid CTC/attention, multi-decoders with searchable intermediates, time-synchronous blockwise CTC/attention, Translatotron models, and direct discrete unit models. In this paper, we describe the overall design, example models for each task, and performance benchmarking behind ESPnet-ST-v2, which is publicly available at https://github.com/espnet/espnet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A hierarchical transducer with self-distillation and a tuned blank penalty improves joint speech recognition and translation, matching offline attention models on conversational data.

Pith tools