Pith. sign in

REVIEW 5 cited by

Speech Translation with Large Language Models: An Industrial Practice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.13585 v1 pith:DXWYRYER submitted 2023-12-21 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords llm-stspeechlanguagelargetranslationmodelmodelsaccurate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given the great success of large language models (LLMs) across various tasks, in this paper, we introduce LLM-ST, a novel and effective speech translation model constructed upon a pre-trained LLM. By integrating the large language model (LLM) with a speech encoder and employing multi-task instruction tuning, LLM-ST can produce accurate timestamped transcriptions and translations, even from long audio inputs. Furthermore, our findings indicate that the implementation of Chain-of-Thought (CoT) prompting can yield advantages in the context of LLM-ST. Through rigorous experimentation on English and Chinese datasets, we showcase the exceptional performance of LLM-ST, establishing a new benchmark in the field of speech translation. Demo: https://speechtranslation.github.io/llm-st/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.

  2. StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A unified large speech-language model uses speech chain-of-thought to jointly perform segmentation, generation-policy decisions, and streaming translation.

  3. From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.

  4. Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The IT-IST IWSLT 2025 submission shows that a 1.5B language model with a speech encoder can do reasonable ASR after alignment, but struggles with ST and SQA.

  5. CMU's IWSLT 2025 Simultaneous Speech Translation System

    cs.CL 2025-06 conditional novelty 4.0 of 10

    CMU reports 44.3 BLEU English-to-Chinese and 25.1 BLEU English-to-German on the ACL60/60 dev set with a streaming Wav2Vec2.0-Qwen2.5 system trained on about 3,850 hours of synthesized speech translation data.

Pith tools