Pith. sign in

REVIEW 1 cited by

Assessing Evaluation Metrics for Speech-to-Speech Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.13877 v1 pith:K3E74SCD submitted 2021-10-26 cs.CL cs.SDeess.AS

Assessing Evaluation Metrics for Speech-to-Speech Translation

classification cs.CL cs.SDeess.AS
keywords translationlanguagesspeech-to-speechevaluationstandardizedevaluatemetricspreviously
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Speech-to-speech translation combines machine translation with speech synthesis, introducing evaluation challenges not present in either task alone. How to automatically evaluate speech-to-speech translation is an open question which has not previously been explored. Translating to speech rather than to text is often motivated by unwritten languages or languages without standardized orthographies. However, we show that the previously used automatic metric for this task is best equipped for standardized high-resource languages only. In this work, we first evaluate current metrics for speech-to-speech translation, and second assess how translation to dialectal variants rather than to standardized languages impacts various evaluation methods.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OpenSTBench: Beyond Semantic Evaluation for Speech Translation

    eess.AS 2026-05 unverdicted novelty 6.0

    OpenSTBench supplies a unified multidimensional protocol and code for evaluating S2TT and S2ST systems in both offline and streaming modes across translation, speech, speaker, emotion, temporal, and latency dimensions.