A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.
Streaming Sequence Transduction through Dynamic Compression
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce STAR (Stream Transduction with Anchor Representations), a novel Transformer-based model designed for efficient sequence-to-sequence transduction over streams. STAR dynamically segments input streams to create compressed anchor representations, achieving nearly lossless compression (12x) in Automatic Speech Recognition (ASR) and outperforming existing methods. Moreover, STAR demonstrates superior segmentation and latency-quality trade-offs in simultaneous speech-to-text tasks, optimizing latency, memory footprint, and quality.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
A survey of 110 SimulST papers shows most systems rely on unrealistic human pre-segmented audio and inconsistent terminology, and it offers a taxonomy and recommendations to fix both.