Pith. sign in

REVIEW 2 cited by

Task Arithmetic for Language Expansion in Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11274 v3 pith:J5UUUCOP submitted 2024-09-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagemodelstaskarithmetictranslationexistingpairsre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in large language models (LLMs) has gained interest in speech-text multimodal foundation models, achieving strong performance on instruction-tuned speech translation (ST). However, expanding language pairs is costly due to re-training on combined new and previous datasets. To address this, we aim to build a one-to-many ST system from existing one-to-one ST systems using task arithmetic without re-training. Direct application of task arithmetic in ST leads to language confusion; therefore, we introduce an augmented task arithmetic method incorporating a language control model to ensure correct target language generation. Our experiments on MuST-C and CoVoST-2 show BLEU score improvements of up to 4.66 and 4.92, with COMET gains of 8.87 and 11.83. In addition, we demonstrate our framework can extend to language pairs lacking paired ST training data or pre-trained ST models by synthesizing ST models based on existing machine translation (MT) and ST models via task analogies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Pseudo-labelling and prompting with fluent transcriptions improve end-to-end spoken grammatical error correction and feedback for Whisper-based models, but the benefits depend on model size.

  2. A correlation-permutation approach for speech-music encoders model merging

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A layer-wise correlation-permutation alignment lets a merged HuBERT-MERT encoder keep speech performance and improve music scores, beating linear interpolation by 19.5 points on the authors' aggregate score.

Pith tools