{"id":"d1c5f17f-9dbf-49df-a8cd-a0b246a6ead0","arxiv_id":"2506.04855","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Instruction-example alignment and extreme short demonstrations drive length control in LLM translation; selecting among multiple outputs improves the length-quality tradeoff.","lead":"This paper tests eight open-source language models on translating English into German, French, and Spanish while keeping the output length within ten percent of the source. It finds that showing models extreme short examples, with matching instructions, is the most reliable way to control translation length, which matters for dubbing and subtitling.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed necessity of extreme demonstrations is not established: the match/no-match design never varies prompt type factorially, and zero-shot results in Table 9 show the Short instruction alone already shortens output.","rationale":"The reader's weakest assumption was domain transfer from the MuST-C devset to the IWSLT blind set. That is a real generalization concern, but it targets the final-evaluation claim rather than the central causal claim, which is established on the development split. The stronger problem is internal to the reported design: the match/no-match analysis in Table 3 cannot separate instruction wording from demonstration content because the only mismatch condition uses the uncontrolled Random prompt. The zero-shot rows in Table 9 already show that the Short instruction alone produces short outputs for several models, so the Abstract's 'only when presented with extreme examples' is too strong as written. This concern is concrete, testable, and does not require rejecting the paper: the underlying trends are consistent across models and language pairs, the code and data are released, and a full factorial experiment would either support the interaction or force a more careful restatement. The reader's verdict of CONDITIONAL is therefore appropriate; my read does not change it, but it identifies a different and more direct weakness than the one singled out in the reader's weakest_assumption. I would keep the paper conditional pending the off-diagonal prompt/pool runs.","tokens_in":26464,"tokens_out":8450,"duration_ms":100698,"concrete_test":"Re-run the 5-shot En→De condition for gemma2:27b and llama3:70b with a full factorial prompt×pool design: {Random, Isometric, Same, Short, Tiny} prompts × {Random, Isometric, Same, Short, Tiny} pools, 10 sampling runs per cell, plus zero-shot for each prompt. The critical comparisons are Short/Tiny prompts with Random and Isometric pools. If LR(Short prompt + Random pool) is close to LR(Short prompt + Tiny pool) and below 1.0, extreme demonstrations are not necessary for the shortening effect. If instead LR(Short prompt + Random pool) remains near 1.1 while the matched Tiny cell reaches below 1.0, the alignment claim is supported. Report LC as well, since the claimed advantage may be about precision rather than raw shortening.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract, Section 4.1) is that length control requires well-aligned instructions plus extreme Tiny/Short demonstrations, and that isometric demonstrations alone do not shorten output. The evidence for the alignment and extreme-demonstration components is incomplete. Table 3 (and Table 8) compares matched prompt+pool cells against a single mismatch condition: the uncontrolled Random prompt with the same pools. It never runs the reverse mismatch, e.g., a Short/Tiny prompt with Random demonstrations, or an Isometric prompt with Tiny demonstrations. Consequently, the large match advantages for Short/Tiny could be caused by the instruction wording alone rather than by instruction-example alignment. This is not merely hypothetical: the paper's own zero-shot results in Table 9 show that the explicit Short instruction with no demonstrations already drives the length ratio below 1.0 for Gemma and Llama models (e.g., gemma2:27b En→De: 0.83 vs 1.13 for Random; llama3:70b: 0.92). Section 4.2 even notes the zero-shot LR is below 1.0 for Llama/Gemma, but the Abstract's 'only when presented with extreme examples' and Section 4.1's 'can overcome this bias only when extreme examples are provided' are inconsistent with that observation. A weaker, defensible claim would be that extreme demonstrations, combined with matching instructions, improve the precision of length control (LC) relative to zero-shot extremes; but the necessity of extreme demonstrations for the shortening effect is not supported by the reported design. This is a correctness risk in the causal claim, not a generalization risk, and it is addressable with additional experiments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates prompting strategies for length-controlled (isometric) machine translation with eight open-weight LLMs across En→De, En→Fr, and En→Es. It compares four instruction types (Random, Isometric, Same, Short/Tiny) paired with demonstration pools of matching or non-matching length properties, in 0-, 5-, 10-, and 20-shot settings, and evaluates on the IWSLT 2022 Isometric Shared Task blind set. The authors find that few-shot demonstrations shift output length most when the instruction and pool are aligned, and that selecting among multiple outputs with COMETKIWI after length filtering can rival or beat shared-task baselines on some language pairs.","tokens_in":26681,"tokens_out":5580,"duration_ms":64887,"significance":"The study is a careful and useful empirical contribution: it evaluates eight models on three language pairs with ten runs per setting, re-evaluates all shared-task baselines with the same script, and publicly releases the data. If its central claim were fully supported, the paper would provide actionable guidance for prompt-based length control. However, the headline claim that LLMs 'tend to produce shorter translations only when presented with extreme examples' is contradicted by the paper's own zero-shot results, and the match/no-match evidence for instruction-example alignment is confounded by the choice of the Random prompt as the sole mismatch condition. The work is therefore valuable as a descriptive study, but the main interpretive claim needs substantial revision.","major_comments":[{"comment":"The claim that LLMs 'tend to produce shorter translations only when presented with extreme examples' (Abstract) and that models 'can overcome this bias only when extreme examples are provided' (Section 4.1) is directly contradicted by the paper's own zero-shot results in Table 9. For example, in zero-shot En→De with the Short prompt, gemma2:27b achieves a length ratio of 0.83 and llama3:70b achieves 0.92, both below 1.0 with no demonstrations at all. Section 4.2 even notes that zero-shot Llama and Gemma models produce length ratios below 1.0. The necessity claim should be replaced with a weaker, defensible statement, e.g., that extreme demonstrations combined with matching instructions improve the precision of length control relative to zero-shot prompting, but are not necessary for shortening.","section":"Abstract and Section 4.1"},{"comment":"The match/no-match comparison is confounded: the 'No' condition always uses the uncontrolled Random prompt, while the 'Yes' condition uses the pool-matched prompt. This design never varies the prompt type factorially (e.g., Short/Tiny prompts with Random demonstrations, or Random prompts with Short/Tiny demonstrations are both missing). Given that the zero-shot Short instruction alone already shortens output for several models, the large gaps in Table 3 could be caused entirely by the instruction wording rather than by instruction-example alignment. To support the claimed alignment effect, the authors need a factorial manipulation (or at least the reverse mismatch conditions). This is a load-bearing weakness because the alignment conclusion is the paper's central conceptual contribution.","section":"Section 4.1 and Table 3"},{"comment":"The abstract's claim of 'state-of-the-art performance for some language pairs' rests on a multi-output selection procedure (10 generations, length filtering, COMETKIWI reranking), which is not directly comparable to the single-output shared-task systems. The k=1 rows in Table 4 show that without multi-output selection the same models are below the STRONGBASELINE (e.g., gemma2:27b-k=1 En→De BLEU 19.0 vs. 21.6 for the baseline). The claim should be explicitly qualified as a multi-output selection result, or the paper should report a single-output comparison as the primary one; otherwise the SOTA statement overstates what the setup supports.","section":"Section 5 and Table 4"}],"minor_comments":[{"comment":"The statistical test behind the underlines is not described; please specify the test used and whether any multiple-comparison correction was applied.","section":"Section 4.1 / Table 3"},{"comment":"The row for gemma2:9b-k=1 in En→Fr shows all-zero scores (LR=0.00, LC=0.0, BS=0.00, BLEU=0.0), which appears to be a data error or an unrun cell; please correct or explain this entry.","section":"Table 4"},{"comment":"The sentence 'when demonstrations are also given in few-shot settings for these models, translations are longer, even when the associated demonstrations are short or very short' is confusing because the preceding sentence reports zero-shot length ratios below 1.0; please clarify whether the intended comparison is zero-shot versus few-shot for the same instruction, and how that relates to the paper's main claims.","section":"Section 4.2"},{"comment":"The demonstration pool is drawn from the MuST-C dev set (TED talks), while the final evaluation is on the IWSLT 2022 blind set of YouTube dialogues; the potential domain mismatch and its effect on length-control transfer should be acknowledged in the Limitations section.","section":"Section 2 / Limitations"},{"comment":"The prompt template table is dense but workable; a small example of a fully instantiated prompt would improve readability.","section":"Section 3 / Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical study with careful methodology and useful released data, but the headline claims overreach the evidence that is already present in the paper. The zero-shot contradiction and the confounded match/no-match design are fixable by rewriting the central claims and adding the missing factorial conditions (or at least the reverse mismatches). The 'state-of-the-art' qualification also needs adjustment. With these revisions, the paper could become a well-scoped contribution; I do not see a need for rejection, but the current abstract and Section 4.1 are not defensible as written."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a careful, useful empirical study of length-control prompting for isometric MT, but the paper's headline causal claim is not supported by its own data. The short instruction alone, with no demonstrations, already pulls output length below 1.0 for several models (Table 9). So \"LLMs produce shorter translations only when presented with extreme examples\" is false as stated.\n\nWhat's actually good: the experimental design is solid in many ways. Ten runs per setting, three language pairs, eight open models, and they re-evaluated the shared-task baselines with the same script. The demonstration pools (Random, Isometric, Same, Short, Tiny) are clearly defined, and the observation that isometric demonstrations do not induce shorter output is a real and useful negative result. The multi-output selection with COMETKIWI is a practical recipe for getting a better length-quality tradeoff, and the released data will help others build on it. The paper is honest about many limitations (quantization, small blind set, cost).\n\nThe soft spot is the load-bearing claim. The match/no-match design compares matched prompt+pool cells against a single mismatch condition: the uncontrolled Random prompt with the same pools. It never runs the reverse mismatch (e.g., Short prompt with Random examples), so instruction wording and example type are confounded. The zero-shot results in Table 9 show the Short instruction alone lowers LR for Gemma and Llama (e.g., gemma2:27b En->De: 0.83 vs 1.13 for Random). Section 4.2 even notes the zero-shot LR is below 1.0, but the abstract and Section 4.1 still conclude that extreme examples are necessary. That is inconsistent. A defensible claim would be that extreme demonstrations combined with matching instructions improve the precision of length control relative to zero-shot extremes, and that isometric demonstrations fail to shorten output. That weaker claim still holds. Also, the \"state-of-the-art\" framing is overstated for En->Es, where length compliance trails the strong baseline.\n\nOther issues are minor: no confidence intervals in the appendix, p<0.1 is lenient, and the 200-sentence blind set is small. These do not change the main takeaway.\n\nThis paper is worth a serious referee. The recipe is practical, the negative result about isometric demonstrations is interesting, and the multi-output selection is a useful contribution. But the authors need to reframe the central claim and ideally run the factorial experiments that would support it. I would cite this for the empirical findings, not for the necessity claim.","headline":"Useful empirical recipe for length-controlled MT, but the causal claim that extreme demonstrations are necessary is not supported by the paper's own zero-shot results.","tokens_in":27297,"tokens_out":2854,"would_cite":true,"duration_ms":29543,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Precise length control in few-shot LLM translation requires extreme short demonstrations whose instruction wording matches them; isometric demonstrations alone do not shorten output.","keywords":["isometric machine translation","length control","few-shot prompting","large language models","demonstration selection","prompt alignment","IWSLT 2022 isometric shared task","translation overgeneration"],"falsifier":"Run the authors' best 10-shot Tiny setting on a held-out set of YouTube-dialogue sentences, but replace the Tiny demonstrations with Random ones while keeping the \"shorter than the source\" instruction. If length compliance (within $\\pm 10\\%$) does not visibly exceed the Random-pool baseline, the central claim that extreme demonstrations drive length control is falsified for that domain.","tokens_in":26214,"feed_emoji":"📏","tokens_out":7282,"duration_ms":74346,"temperature":0.7,"pith_summary":"This paper asks what it takes to make open-source LLMs translate with a strict length budget: outputs within ten percent of the source's character count, as dubbing and subtitling require. Across eight LLMs and three language pairs, the authors find that few-shot demonstrations alone do not control length. The decisive factor is pairing a length instruction with demonstrations that are themselves extreme: translations much shorter than their sources (the 'Tiny' and 'Short' pools). Isometric or same-length examples barely shift the output length distribution, because models fall back on the typical translation length they learned in training. The paper then shows that generating ten outputs from different Tiny demonstration sets, filtering to the length-compliant ones, and selecting with a reference-free quality scorer yields competitive isometric MT and state-of-the-art trade-offs for some pairs.","feed_headline":"Tiny examples, not isometric ones, control LLM translation length","feed_subtitle":"Aligning few-shot prompts with ultra-short demonstration pairs lets open LLMs meet the ±10 percent dubbing length bar.","key_machinery":"The load-bearing mechanism is the pairing of a prompt template's length instruction with a demonstration pool built to embody that instruction. The paper defines five pools from the MuST-C devset by target/source length ratio: Random (unfiltered), Isometric (ratio in $[0.9, 1.1]$), Same (the 50 closest to ratio $1.0$), Short (ratio $\\le 1.0$), and Tiny (the 50 smallest ratios, averaging about $0.6$--$0.68$). For each pool the prompt is reworded to match (\"ensure it is shorter than the source\", \"within $\\pm 10\\%$\", etc.), and the matched-pair setting is what produces large length shifts: Tiny/Short pools push En→De ratios down to roughly $0.9$--$1.05$ depending on model, while Isometric and Same pools leave ratios near the Random baseline. A second mechanism is overgeneration control: instructing the model to \"output only the translation\" and cutting generated text at the first newline prevents explanations from inflating length. The final component is multi-output selection: ten independent 10-shot prompts from the Tiny pool, length filtering, and COMETKIWI reranking.","core_discovery":"The paper's central claim is that effective length control in few-shot LLM translation is not a matter of adding length instructions or demonstrations by themselves; it is a matter of alignment. A prompt that asks for a shorter translation accompanied by examples that are drastically shorter (the Tiny pool, the fifty shortest devset translations) reliably compresses output length ratios, while demonstrations that merely satisfy the isometric $\\pm 10\\%$ constraint, or that almost exactly match source length, leave the output ratio close to the unconstrained value. The authors show this across eight open-weight models (Llama 3, Gemma 2, Qwen 2, Mistral, Mixtral) and three directions (En→De, En→Fr, En→Es), and they connect it to a mechanism claim: when instruction and demonstrations agree, the model shifts its implicit length prior; when they disagree, the demonstrations are effectively ignored. On top of this, they show that taking ten outputs from ten different Tiny demonstration sets, discarding all outputs outside the $\\pm 10\\%$ band, and ranking the survivors with COMETKIWI reaches BERTScore/BLEU competitive with or above the 2022 IWSLT isometric-shared-task baselines and several submissions, with the best results for En→De and En→Es and a clear path to synthetic data creation.","pith_inferences":["The Tiny-pool effect suggests that LLMs use few-shot examples as a prior over output length rather than as a constraint; a testable extension is to sweep demonstration ratios continuously (e.g., 0.5, 0.7, 0.9, 1.0) and check whether the output length ratio is a monotone function of the pool's average ratio.","The instruction-alignment finding likely generalizes to other soft constraints that are underrepresented in the training distribution, such as formality or verbosity: when the desired property is rare in typical outputs, extreme aligned demonstrations may be needed, and mismatched prompts will be silently ignored.","In production, the 10-output reranking pipeline could be replaced by an adaptive strategy: start with the uncontrolled prompt, and only for non-compliant outputs retry with Tiny demonstrations, cutting cost roughly in half since about half the samples are already compliant.","Because all models were quantized, the absolute quality numbers are likely understated; the length-control mechanism, being a relative effect between pools, is probably unaffected, and testing on unquantized models would be a cheap way to check."],"forward_implications":["Few-shot length control in LLMs can be engineered without fine-tuning: choosing the shortest available translations as demonstrations and wording the instruction to match them is enough to move output length toward the target.","Isometric demonstrations are not a neutral control; because typical translations are longer than their sources, giving isometric or same-length examples tells the model nothing it does not already assume, so length compliance stays near the unconstrained level.","Generating multiple outputs (e.g., ten) from independent Tiny demonstration sets and selecting a length-compliant one by reference-free quality score gives a practical recipe for isometric MT and a source of synthetic training data.","Beyond five or ten demonstrations, additional shots bring little translation-quality gain, but the demonstration-selection effect on length is robust across model families and language pairs.","The approach reaches state-of-the-art trade-offs for En→De and En→Es against the 2022 shared-task submissions, while En→Fr still trails the strong baseline."],"supporting_citations":[{"why":"Supplies the shared-task setup, the $\\pm 10\\%$ isometric constraint, the length compliance metric, and the strong/weak baselines used for comparison.","marker":"Anastasopoulos et al., 2022"},{"why":"Provides the prompt template and the few-shot demonstration format the paper adapts for in-context samples.","marker":"Zhang et al., 2023a"},{"why":"Documents LLM overgeneration in translation and motivates the 'output only the translation' restriction.","marker":"Bawden and Yvon, 2023"},{"why":"Establishes that prompting strategy strongly affects few-shot MT performance, the starting point of the paper's investigation.","marker":"Vilar et al., 2023"},{"why":"Defines isometric MT as the $\\pm 10\\%$ constraint for dubbing, framing the task.","marker":"Lakew et al., 2022"},{"why":"Supplies COMETKIWI, the reference-free quality estimator used to select among length-compliant outputs.","marker":"Rei et al., 2022"},{"why":"The AppTek constrained submission that the paper's final evaluation is compared against for the state-of-the-art claim.","marker":"Wilken and Matusov, 2022"}],"fun_headline_variants":["Tiny few-shot examples bend LLM translation length more than isometric ones","Align instruction and demo to shrink LLM translation length","Extreme examples beat isometric ones for LLM length control","Few-shot prompt alignment, not isometric demos, drives LLM length control","LLMs only shrink translation length when given extreme short demos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's length-control results on the shared-task blind test set rest on the assumption that demonstration pools drawn from the MuST-C devset transfer their effect to a different genre of YouTube-dialogue test sentences; if the genre mismatch changes how models follow the demonstration-length signal, the final-evaluation gains would not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Tiny few-shot examples bend LLM translation length more than isometric ones","Align instruction and demo to shrink LLM translation length","Extreme examples beat isometric ones for LLM length control","Few-shot prompt alignment, not isometric demos, drives LLM length control","LLMs only shrink translation length when given extreme short demos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001019,"raw_usage":{"total_tokens":4327,"prompt_tokens":1001,"completion_tokens":3326,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":3245}},"tokens_in":617,"tokens_out":3326,"duration_ms":24482,"temperature":1.0,"reasoning_tokens":3245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:31:55.294854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the authors' best 10-shot Tiny setting on a held-out set of YouTube-dialogue sentences, but replace the Tiny demonstrations with Random ones while keeping the \"shorter than the source\" instruction. If length compliance (within $\\pm 10\\%$) does not visibly exceed the Random-pool baseline, the central claim that extreme demonstrations drive length control is falsified for that domain.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines isometric MT as the $\\pm 10\\%$ constraint for dubbing, framing the task."},{"cited_title":"Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, Jos \\'e G","cited_arxiv_id":null,"evidence_quote":"Supplies COMETKIWI, the reference-free quality estimator used to select among length-compliant outputs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The AppTek constrained submission that the paper's final evaluation is compared against for the state-of-the-art claim."}],"review_version":1}