{"id":"25fbdf1e-a872-492e-ace0-53ced171c918","arxiv_id":"2506.02995","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"End-to-end speech translation systems translate idioms worse than text-based systems, frequently producing literal or incorrect outputs, across German and Russian to English.","lead":"This paper tests how well speech-to-text translation systems handle idioms compared to text-based systems, using German and Russian to English data. It finds speech translation models drop sharply on idiomatic phrases, often reverting to literal translations, while text-based models and cascaded systems do better.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"News and idiom test sets are not matched; the reported SLT-specific drop may reflect corpus difficulty rather than idiomaticity.","rationale":"I consider the corpus mismatch the single most load-bearing assumption because the paper's headline contribution is the idiomaticity-specific weakness of SLT. If the news and idiom sets differ in difficulty, genre, or other covariates, the observed COMET gap is not evidence for that claim. The human annotation of literal translations is suggestive, but it is measured only on the idiom set, so it does not provide a control. Additionally, the paper's own Table 2 undercuts the conclusion that the gap is 'more pronounced' for SLT: several MT systems exhibit comparable absolute drops. This does not invalidate the finding that SLT scores lower on idioms than MT/LLMs, but it does mean the quantitative 'pronounced drop' framing is overstated. A matched control experiment would settle whether the effect is idiom-specific. The reader's conditional verdict is therefore appropriate; the concern is addressable without overturning the paper. Agreement with the reader: the weakest assumption identified is the same corpus comparability issue.","tokens_in":15597,"tokens_out":13406,"duration_ms":150611,"concrete_test":"Build a matched control set from the same source corpus: for each of the 250 idiom sentences, take a literal, non-idiomatic sentence from Idioms-InContext-MT (or a literal paraphrase of the idiom's meaning) matched for length and word frequency, and rerun the full evaluation (COMET + human annotation) for the SLT and text-MT systems. If the news-vs-idiom gap in COMET shrinks to near zero or becomes comparable across SLT and MT, the reported 'pronounced SLT drop' is an artifact of corpus mismatch; if the gap persists, the idiomaticity explanation is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's causal claim that end-to-end SLT systems have a specific weakness for idioms rests on the comparability of the two evaluation sets built in §3.3. The news baseline is 250 random sentences from News Commentary, while the idiom set is a manually filtered subset of Idioms-InContext-MT that deliberately keeps only hard, non-literal cases. These sets differ in genre, register, sentence length, vocabulary and topical distribution, and none of these are controlled or matched. The COMET drop from news to idioms could therefore be driven by general difficulty or domain shift rather than idiomaticity. The human annotation in §3.4.2 does not fix this: the 'Literal Translation' category is defined only for idiom items, so it cannot serve as a matched control, and the annotation sample is only 50 items per condition. The conclusion's claim that the gap 'is more pronounced for SLT systems' is also not clearly supported by Table 2: for German, the absolute COMET drop is 0.204 for Whisper vs 0.209 for NLLB; for Russian it is 0.140 vs 0.145. The Limitations section acknowledges synthetic speech and language coverage but never mentions this corpus mismatch, so the missing limitation should be flagged.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a systematic comparison of idiom translation in speech-to-text (SLT) versus text-to-text MT, LLM, and cascaded systems for German→English and Russian→English. It builds a 250-item news set from News Commentary and a 250-item idiom set from Idioms-InContext-MT, synthesizes speech for the text sources, and evaluates the systems with COMET and human annotations, plus DecoderLens layer-wise diagnostics. The central claim is that end-to-end SLT systems (Whisper, SeamlessM4T audio) drop sharply on idiomatic data and often produce literal translations even in later encoder layers, while text-based systems handle idioms better.","tokens_in":15739,"tokens_out":4566,"duration_ms":48026,"significance":"The paper addresses an underexplored problem and provides a useful empirical benchmark: it compares two language pairs, multiple system classes, automatic and human evaluation, and it releases code, evaluation datasets, and annotated subsets. The direct audio-versus-text comparisons within the idiom set are informative and support the existence of an SLT-specific challenge with figurative language. The DecoderLens analysis is an original diagnostic application. The main caveat is that the cross-corpus news-versus-idiom comparison is not controlled, so the strongest causal claim about idiomaticity per se needs reframing or additional evidence.","major_comments":[{"comment":"The news and idiom evaluation sets are constructed independently from different corpora (News Commentary vs. a manually filtered subset of Idioms-InContext-MT) and differ in genre, register, sentence length, and topic; no matching or controlling for these factors is reported. The claim that SLT systems show a 'pronounced performance drop' on idiomatic data therefore confounds idiomaticity with general corpus difficulty. The within-idiom comparisons in §4.2 (audio vs. text on the same idiom items) provide more direct evidence for an SLT-specific weakness, but the cross-corpus drop in Table 2 cannot by itself establish that idioms are the cause. I ask the authors to either add a matched non-idiomatic control condition (e.g., literal uses of the same idiom-containing sentences or genre-matched non-idiomatic sentences) or to restrict the causal claim and state the corpus mismatch explicitly in the Limitations section.","section":"§3.3, Table 2"},{"comment":"The statement that the news-to-idiom gap 'is more pronounced for SLT systems' is not supported by the reported COMET differences. In German, Whisper drops from 0.8437 to 0.6402 (Δ = 0.2035) and NLLB drops from 0.8841 to 0.6749 (Δ = 0.2092); in Russian, Whisper drops from 0.8318 to 0.6916 (Δ = 0.1402) and NLLB drops from 0.8664 to 0.7214 (Δ = 0.1450). No statistical test of an interaction between system type and dataset is provided. The conclusion should be based on the direct audio-vs-text comparisons on the idiom set, where the difference is clear, or on a proper interaction test, not on the unmatched cross-corpus gaps.","section":"§4.1, Table 2, Conclusion point 1"},{"comment":"The human annotation is based on only 50 items per system-domain combination, and no inter-annotator agreement is reported. The DecoderLens analysis in §5 and Figure 3 is therefore built on a small sample and a single adjudicated annotation pass. Reporting agreement, for example Cohen's kappa, and stating the number of unique idiom items versus repeated outputs per model would strengthen the layer-wise claims. This does not invalidate the qualitative pattern, but it limits the strength of the conclusions drawn from the layer-wise distributions.","section":"§3.4.2, §5, Figure 3"}],"minor_comments":[{"comment":"There is a missing space: 'Inthispaper' should be 'In this paper'.","section":"Abstract"},{"comment":"The first sentence of the Limitations section contains a typo: 'Morever' should be 'Moreover'.","section":"Limitations"},{"comment":"In the Russian Idioms table, the standard deviation for Whisper is shown as 'L0.106', which appears to be a typo for '0.106'.","section":"Appendix B, Table 5(b)"},{"comment":"The row labeled 'Seamless (Text MT and LLM)' mixes a system name with a category label; it would be clearer to call it 'SeamlessM4T (text-to-text)' and to list DeepSeek and LLaMA separately, as is already done.","section":"Table 2"},{"comment":"The captions do not state the number of annotated examples per bar; adding this information (for example, 'n = 50 per condition') would help readers assess the reliability of the displayed proportions.","section":"Figure 2 and Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of an NLP/speech-translation venue and the artifact release is a definite strength. The main risk is overclaiming on the cross-corpus comparison; the revision should focus on reframing that claim and adding the requested statistical support. No other concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about. It is the first systematic comparison I know of that puts MT, LLM, end-to-end SLT, and cascaded systems on the same idiomatic test data for two language pairs. The main result—that end-to-end SLT systems score clearly worse on idioms than text-based or cascaded systems, and produce more literal translations—survives scrutiny. Human annotation corroborates the COMET scores, and the DecoderLens layer-wise analysis is a nice diagnostic touch, even if limited to encoder-decoder models. That is genuinely useful for the speech translation community.\n\nThe soft spots are real but fixable. The biggest is the unmatched baseline. The news set comes from News Commentary, while the idiom set is a deliberately filtered subset of Idioms-InContext-MT that keeps only non-literal, hard cases. These sets differ in genre, sentence length, vocabulary, and overall difficulty, so the COMET gap between news and idioms cannot be cleanly attributed to idiomaticity. The paper even acknowledges the filtering but never discusses this confound. The human annotation can't fix it either, because the 'Literal Translation' category only exists for idiom items. This needs to be acknowledged and ideally controlled with a difficulty-matched set or at least softened causal language.\n\nSecond, the abstract says the performance drop is 'more pronounced for SLT systems,' but the absolute drops in Table 2 are nearly identical across SLT and MT: for German, Whisper drops 0.204 while NLLB drops 0.209; for Russian, 0.140 vs 0.145. SLT systems are worse on idioms in absolute terms, and they do produce more literal translations, but the interaction claim is not supported by the paper's own numbers. That is an overstatement that should be corrected.\n\nMinor issues: no confidence intervals or model-vs-model significance tests, only within-model news-vs-idiom tests; synthetic speech with a single voice; the code/data link is promised but not verifiable from the preprint; annotation is only 50 items per condition. None of these overturn the central finding, but they do limit how strongly the quantitative gaps can be interpreted.\n\nBottom line: this deserves a serious referee. The empirical pattern is solid, the contribution is genuinely new, and the weaknesses are addressable in revision. I would recommend a conditional accept: fix the overclaim about a more pronounced drop, discuss the corpus-matching confound, and either add a matched control or temper the causal language.","headline":"Useful first benchmark of idiom translation across MT/SLT/LLM/cascades; central claim holds, but the news-vs-idiom comparison is confounded and one headline claim isn't supported by the paper's own table.","tokens_in":16345,"tokens_out":2114,"would_cite":true,"duration_ms":24654,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Speech-to-text translation systems show a pronounced, statistically significant performance drop on idiomatic sentences, often falling back to literal translations even in their deepest layers, while text-based machine translation and…","keywords":["idiom translation","speech-to-text translation","figurative language","machine translation evaluation","COMET metric","DecoderLens analysis","cascaded speech translation","large language models"],"falsifier":"A reader could take non-idiomatic sentences from the same idiom corpus, or add matched non-idiomatic controls from the same source, and re-run the news-versus-idiom comparison; if the COMET gap largely disappears, the paper's attribution of the drop to idiomaticity would be refuted.","tokens_in":15355,"feed_emoji":"🎙️","tokens_out":5934,"duration_ms":56942,"temperature":0.7,"pith_summary":"This paper sets out to test whether speech-to-text translation (SLT) systems, which take audio directly to translated text, fail on idiomatic expressions more than text-based systems do. The authors compare end-to-end SLT models (Whisper and SeamlessM4T) with text machine translation, large language models, and cascaded ASR-plus-MT pipelines on German-to-English and Russian-to-English idiom and news test sets. The central finding is that SLT systems suffer a much larger quality drop when moving from news to idiomatic data, frequently producing literal word-for-word translations, while text-based systems and large language models preserve figurative meaning more often. The paper interprets this as evidence that end-to-end SLT architectures have a specific weakness with figurative language that cascaded systems partly avoid.","feed_headline":"Speech-to-text translation stumbles hardest on idioms","feed_subtitle":"Direct audio systems lose figurative meaning on German and Russian idioms; text models and cascades cope better.","key_machinery":"The evaluation design pairs two complementary instruments. COMET, a neural semantic-similarity metric, scores whole-sentence translation quality on matched news and idiom sets, making the news-versus-idiom gap the headline quantity. DecoderLens, which swaps the final encoder representation for intermediate layer activations and decodes them into text, lets the authors annotate how translation quality evolves layer by layer and pinpoints where literalization appears. Both are applied to the same encoder-decoder models, so the layer-wise failure modes can be tied directly to the COMET drops.","core_discovery":"The paper claims to establish that end-to-end speech-to-text translation systems have a systematic, measurable weakness on idiomatic input that is not explained by automatic speech recognition errors alone. Across both language pairs, direct SLT systems (Whisper Large v3 and SeamlessM4T speech-to-text) score substantially lower on the COMET metric for idiom sentences than for news sentences (for example, Whisper drops from 0.844 to 0.640 on German-to-English), while text-only MT systems, large language models, and cascaded ASR-to-MT/LLM pipelines retain more of the figurative meaning. Layer-wise DecoderLens analysis shows SLT encoders produce meaningful output only in high layers and tend to settle on literal translations even there, whereas text encoders refine toward correct idiomatic output more gradually. The conclusion is that idiom handling requires idiom-specific training strategies and better internal representations for figurative meaning in SLT architectures.","pith_inferences":["A natural next test the paper does not run: probe whether the audio encoder's compressed representations lose idiom-specific cues before the decoder starts, which would predict that better acoustic-semantic alignment layers would close the gap.","Because the idiom and news sets come from different corpora, part of the measured gap could reflect genre or topic difficulty rather than idiomaticity per se; a matched control set would isolate the idiom effect.","The literalization pattern in higher layers suggests a targeted intervention: adding an idiom-recognition auxiliary task at those layers might reduce SLT literal translations more than adding more data.","Applying the same news-versus-idiom comparison to more language pairs and to spontaneous (non-TTS) speech would show whether the SLT idiom weakness is universal or tied to the synthetic audio used here."],"forward_implications":["If the central claim holds, end-to-end SLT systems need idiom-specific training data or objectives rather than relying on general scaling.","For practical speech translation of content likely to contain idioms, cascaded ASR plus text MT or LLM pipelines are the safer choice.","Literal translation is the dominant failure mode across both SLT and MT, so progress will require models to recognize figurative status before decoding.","The idiom gap is consistently larger for German than for Russian, suggesting language-specific idiom difficulty that future benchmarks should separate.","COMET scores and human annotation align in ranking systems, so automatic semantic metrics can be used to track idiom-specific improvements."],"supporting_citations":[{"why":"Supplies the Idioms-InContext-MT dataset from which the 250 idiom test sentences per language pair are manually selected.","marker":"Stap et al., 2024"},{"why":"Provides the COMET evaluation metric used for all automatic quality scores.","marker":"Rei et al., 2020"},{"why":"Provides the DecoderLens method used for layer-wise translation analysis.","marker":"Langedijk et al., 2024"},{"why":"SeamlessM4T text and speech models evaluated as end-to-end SLT and MT systems.","marker":"Barrault et al., 2023"},{"why":"Updated SeamlessM4T v2 model used in the experiments.","marker":"Barrault et al., 2025"},{"why":"Whisper Large v3 used both as a direct SLT system and as the ASR front-end for cascaded pipelines.","marker":"Radford et al., 2022"},{"why":"NLLB-200 model used as the text-based MT comparison system.","marker":"Team et al., 2022"},{"why":"DeepSeek-V3 LLM used for text translation and cascaded ASR-to-LLM output.","marker":"DeepSeek-AI et al., 2025"},{"why":"Justifies using semantic metrics like COMET over surface-level metrics for figurative language evaluation.","marker":"Song and Xu, 2024"}],"fun_headline_variants":["Direct speech translation loses idiom meaning, text AI wins","Speech-to-text literalizes idioms, MT and LLMs preserve them","End-to-end SLT systems falter on idioms, cascades improve","Idioms trip up speech-to-text AI more than conventional text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes the news and idiom test sets differ mainly in idiomaticity, even though they come from different corpora with different genres and topics; if that assumption fails, the measured performance gap cannot be attributed to idioms alone.","fun_headline_variants_meta":{"raw":{"variants":["Direct speech translation loses idiom meaning, text AI wins","Speech-to-text literalizes idioms, MT and LLMs preserve them","End-to-end SLT systems falter on idioms, cascades improve","Idioms trip up speech-to-text AI more than conventional text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1511,"prompt_tokens":935,"completion_tokens":576,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":551,"tokens_out":576,"duration_ms":6605,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:11:02.782096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could take non-idiomatic sentences from the same idiom corpus, or add matched non-idiomatic controls from the same source, and re-run the news-versus-idiom comparison; if the COMET gap largely disappears, the paper's attribution of the drop to idiomaticity would be refuted.","supporting_citations":[{"cited_title":"The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM Abilities","cited_arxiv_id":"2405.20089","evidence_quote":"Supplies the Idioms-InContext-MT dataset from which the 250 idiom test sentences per language pair are manually selected."},{"cited_title":"DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers","cited_arxiv_id":"2310.03686","evidence_quote":"Provides the DecoderLens method used for layer-wise translation analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies using semantic metrics like COMET over surface-level metrics for figurative language evaluation."}],"review_version":1}