{"id":"56f01bb6-cbbd-4c34-9e7a-01736bc9082f","arxiv_id":"2607.22100","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"MEUSLI, a released family of linear projectors, enables open-source LLM-based ASR across 28 European languages and supports few-hour adaptation to unseen languages and extra speech tasks.","lead":"MEUSLI is a family of small, open-source adapters that connect OpenAI's Whisper speech encoder to multilingual text-only language models, giving them speech-to-text ability in 28 European languages. The paper shows the adapter can be fine-tuned for new languages and for speech translation and topic identification with only a few hours of data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bootstrapping claim rests on false claim that Ukrainian/Albanian are unseen by Whisper; 'new languages' evidence tests only encoder-known languages.","rationale":"The reader's verdict (CONDITIONAL) is appropriate. The paper's primary empirical contribution—an open 28-language projector family with credible WER results and released checkpoints—is supported by the tables and HuggingFace release. However, the bootstrapping claim, highlighted as contribution (3) and in the abstract, depends on the assertion that Ukrainian and Albanian are unseen by the speech encoder. That assertion is false; Whisper explicitly supports both languages. This is a load-bearing concern because it undermines the 'new languages' aspect of the claim, not the overall usefulness of multilingual projector initialization. The recommended concrete test—running the same experiment on a language truly absent from Whisper's training set—would cleanly resolve whether the bootstrapping result extends to genuinely unseen languages. The reader already identified this as the weakest assumption, so agreement is full. The CONDITIONAL verdict stands, conditioned on the authors either correcting the claim or providing evidence from a truly unseen language. No change to the reader's verdict is needed, hence UNCHANGED.","tokens_in":13119,"tokens_out":3643,"duration_ms":40649,"concrete_test":"Check the official Whisper tokenizer/model card for the presence of Ukrainian ('uk') and Albanian ('sq') in the supported language set (e.g., in openai/whisper tokenizer.json or the model's language list). If both are present, as expected, repeat the Section 3.2 bootstrapping protocol on a language provably absent from Whisper's language set (e.g., a regional language like Upper Sorbian or a low-resource language not in Whisper's 99 languages) with the same data budgets (30 hours and 46 minutes). If MEUSLI fine-tuning does not substantially outperform a from-scratch projector on that truly unseen language, the 'languages not seen during training' claim must be retracted or narrowed to 'languages unseen by the LLM.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 explicitly asserts that Ukrainian and Albanian are 'excluded from both the speech encoder and EuroLLM pretraining corpora' to demonstrate bootstrapping of truly unseen languages. This is factually incorrect for the encoder: Whisper-large-v3-turbo was trained on both Ukrainian ('uk') and Albanian ('sq'), as listed in the official Whisper language set. Consequently, the reported gains (16.34% vs. 20.43% WER for Ukrainian, and 75.61% for Albanian after 46 minutes) show that MEUSLI can adapt a multilingual projector to languages the frozen encoder already represents linguistically. They do not demonstrate extension to 'languages not seen during training' at the system level. The abstract's claim 'easily extended to other languages not seen in training' and contribution (3) are therefore overstated. The projector's useful initialization effect remains plausible, but the 'unseen' status—a key differentiator of the method—is unsupported by the presented experiments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MEUSLI, a multilingual linear projector for SLAM-style LLM-based ASR that connects a frozen Whisper-large-v3-turbo encoder to frozen open-source multilingual LLMs (EuroLLM-1.7B-Instruct, EuroLLM-9B, Apertus-8B). The projector is trained on 7,622 hours of speech from Common Voice, FLEURS, and VoxPopuli across 28 European languages, and the authors report WERs on in-domain (CV, FL) and out-of-domain (MLC-SLM) test sets. They further study low-resource fine-tuning for Breton and Maltese, bootstrapping for Ukrainian and Albanian, continual learning via data replay, and transfer to speech translation and topic identification on SIB-Fleurs. The models are released on HuggingFace.","tokens_in":13406,"tokens_out":3173,"duration_ms":33818,"significance":"If the central claims hold, MEUSLI is a valuable open resource: it is the first fully open multilingual SLAM-style projector family covering 28 European languages, with public model weights and a systematic comparison across three LLMs and multiple evaluation sets. The paper also provides practical training details and shows that multilingual pretraining helps low-resource adaptation, which is a useful contribution to the open SpeechLLM ecosystem. The release of the projectors on HuggingFace supports reproducibility and downstream use. However, the extent to which the method 'bootstraps languages not seen in training' is overstated, and the multitask transfer results are partially confounded by training/evaluation data overlap, both of which temper the significance unless corrected.","major_comments":[{"comment":"The bootstrapping experiment claims that Ukrainian and Albanian were 'excluded from both the speech encoder and EuroLLM pretraining corpora.' This is factually incorrect for the encoder: Whisper-large-v3-turbo was trained on both Ukrainian and Albanian. The reported gains (16.34% vs. 20.43% WER for Ukrainian, and 75.61% for Albanian) therefore demonstrate adaptation to languages the frozen encoder already understands, not extension to 'languages not seen in training' as claimed in the abstract and contribution (3). This weakens the 'unseen language' claim substantially. Please either re-run the experiment with truly encoder-unseen languages or reframe the claim as adaptation to languages unseen by the LLM only.","section":"§3.2, Table 4, Abstract, Contribution (3)"},{"comment":"The multitask transfer results are partially circular in an empirical sense: the MEUSLI projector was pretrained on FLEURS audio, and SIB-Fleurs shares the same audio segments, differing only in the added ST and TID annotations. The authors disclose this overlap in Section 5, but the Section 6 conclusion that 'MEUSLI transfers effectively beyond ASR' is stronger than the evidence supports. The current design cannot separate the benefit of multilingual projector pretraining from the benefit of having previously seen the exact audio during ASR pretraining. Please temper the claim or conduct an evaluation on a truly unseen speech corpus for ST/TID.","section":"§4, Table 6, §5"},{"comment":"The text states that 'MEUSLI remains beneficial even when paired with a specialized monolingual encoder.' This is contradicted by the Maltese row: with the adapted encoder, MEUSLI fine-tuning yields 23.7% WER versus 18.0% for the monolingual from-scratch system. The benefit is observed only for Breton (24.5% vs. 27.7%). Please either restrict the claim to Breton, or report variance/error bars and statistical significance before making a general statement.","section":"§3.1, Table 3"}],"minor_comments":[{"comment":"Typo: 'continual leaning' should be 'continual learning'.","section":"Abstract"},{"comment":"'Ukranian' should be 'Ukrainian' in the table header.","section":"Table 4"},{"comment":"The multitask experiments defer detailed methodology to an 'Interspeech 2026 paper' that is not available. For the present manuscript to be self-contained, please include at least the prompt templates, optimization details, and the exact train/eval split sizes.","section":"§4.2"},{"comment":"The Albanian training set is only 46 minutes; the from-scratch baseline fails to converge (389% WER), which may be a degenerate output rather than a meaningful comparison. Report the number of utterances, decoding settings, and any filtering used.","section":"§3.2"},{"comment":"The figure shows the SLAM-ASR pipeline but does not indicate where LoRA is applied to the LLM. Adding a label would improve clarity.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The core 28-language ASR contribution appears sound and the open release is valuable. However, the 'unseen language' claim rests on a factual error about Whisper's training set, and the multitask transfer experiments are partly confounded by data overlap. These are fixable with reframing or additional experiments, but they are load-bearing for two of the paper's four stated contributions. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful empirical paper. The authors release linear projector weights connecting Whisper-large-v3-turbo to three open multilingual LLMs, covering 28 EU languages, and the WER tables make the central claim credible. If you work on speech LLMs or low-resource ASR, this gives you an off-the-shelf baseline and a sensible recipe.\n\nWhat's new: the open release itself, the systematic 28-language evaluation across three LLMs, the fine-tuning results for Breton/Maltese, and the replay-based mitigation of catastrophic forgetting. The multitask extension to ST and TID with small data is a nice bonus, and the authors are upfront in the limitations section that SIB-Fleurs overlaps with FLEURS pretraining data.\n\nWhere I part ways with the paper: Section 3.2 claims Ukrainian and Albanian were 'excluded from both the speech encoder and EuroLLM pretraining corpora.' That is false for the encoder. Whisper-large-v3-turbo was trained on Ukrainian and Albanian, so the bootstrapping experiment does not test truly unseen languages. It tests whether the projector can be adapted to languages the frozen encoder already handles. That is still a useful result — the gains over training a projector from scratch are real — but the 'new language not seen in training' claim in the abstract and contribution (3) is overstated. Some algorithmic details are deferred to a separate paper, which is acceptable for a workshop-style submission but makes full reproduction harder.\n\nNet: the core empirical contribution survives the flaw. The paper deserves a serious referee; the main fix is to re-frame the bootstrapping section and verify the encoder's language coverage before resubmission. I would cite this as the current open multilingual projector baseline, and I'd bring it to a reading group.","headline":"A useful, credible open multilingual projector family for 28 EU languages; the bootstrapping claim is overstated because the test languages are in Whisper's training set, but the core results and release hold up.","tokens_in":13881,"tokens_out":1708,"would_cite":true,"duration_ms":19357,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MEUSLI claims that a single lightweight projector can pair a frozen speech encoder with a frozen multilingual LLM to deliver end-to-end ASR in 28 European languages, and that the same projector can bootstrap new languages and other speech t","keywords":["speech recognition","large language model","multilingual","linear projector","low-resource languages","continual learning","speech translation","topic identification"],"falsifier":"Run the same bootstrapping experiment on a language that is verifiably absent from both the speech encoder's and the LLM's pretraining corpora (a low-resource language with a few hours of publicly available audio). If fine-tuning MEUSLI on that language, with ~30 hours of data, fails to produce a large WER improvement over training a projector from scratch, then the paper's claim that multilingual projector pretraining bootstraps languages unseen by the encoder collapses. A minimal check: verify the encoder's training data list for Ukrainian and Albanian; their presence alone falsifies the pap","tokens_in":13029,"feed_emoji":"🎙️","tokens_out":10773,"duration_ms":97734,"temperature":0.7,"pith_summary":"The paper is trying to establish that one small trainable component — a linear projector that maps speech features into the token-embedding space of a language model — can be trained once on many languages and then reused both for direct multilingual ASR and as a starting point for languages and tasks with very little data. It argues that the projector's multilingual pretraining transfers to low-resource and even supposedly unseen languages, and that this transfer extends beyond transcription to speech translation and topic identification. The concrete evidence is a set of word-error-rate, BLEU, and accuracy tables comparing MEUSLI against both a frozen encoder alone and projectors trained from scratch, plus a public release of the trained projectors. For a reader, the payoff would be a fully open-source, reproducible route to speech-enabled LLMs for the long tail of languages.","feed_headline":"One projector powers open-source ASR in 28 European languages","feed_subtitle":"Train once on 7,600 hours of speech, then fine-tune low-resource languages with a few hours of data.","key_machinery":"The central object is the linear projector: a single hidden layer with ReLU followed by a regression layer (17.31M trainable parameters) that maps temporally downsampled output of a frozen pre-trained speech encoder into the embedding space of a frozen multilingual LLM. A small low-rank adaptation (rank 8, scaling factor 32, 1.38M parameters) is applied on the LLM. Training uses cross-entropy loss on the LLM's token predictions with a text prompt appended to the projected speech embeddings; the same machinery is reused unchanged for fine-tuning, continual learning with data replay, and multitask extensions.","core_discovery":"On the paper's own terms, the central discovery is that a single 17-million-parameter linear projector, trained on about 7,600 hours of open speech data from 28 European languages, links a frozen speech encoder to a frozen multilingual LLM and produces an end-to-end speech recognizer that beats the encoder alone on most of those languages, with especially large gains on low-resource languages. The same projector initialization then improves low-resource fine-tuning, and, using a replay buffer of roughly 1,000 samples per original language, lets new languages be added without catastrophic forgetting. The paper further reports that fine-tuning MEUSLI on under five hours of labeled data per lan","pith_inferences":["The paper's 'unseen language' claim rests on Ukrainian and Albanian being absent from both the speech encoder and the LLM; since the chosen encoder was actually trained on those languages, the experiment demonstrates adaptation to languages the encoder already understands. A truly held-out language test remains an open experiment.","The reported gains over the frozen encoder alone suggest the LLM's linguistic priors are being used, but the paper does not isolate how much of the improvement comes from the encoder's multilingual representation versus the LLM's priors; an ablation removing the LLM (replacing it with a text-only decoder) would help.","A testable extension is to apply the same bootstrapping recipe to a language with no presence in either the encoder or the LLM pretraining (e.g., an endangered language with a few hours of recordings). If the MEUSLI advantage over from-scratch persists, the continual-learning bootstrapping claim survives; if it collapses, the reported gains depend on the encoder's prior knowledge.","Because the projector is only about 17 million parameters while the language model is orders of magnitude larger, the paper implies the alignment bottleneck is small; a natural next experiment is sweeping projector capacity downward to see whether an even simpler linear map suffices at this scale."],"forward_implications":["A fully open-source, reproducible speech-to-text path for 28 European languages becomes available, removing dependence on proprietary or English-only speech-enabled LLMs.","Low-resource languages gain a practical recipe: initialize from the multilingual projector rather than training a projector from scratch, which in the reported experiments roughly halves WER for Breton and Maltese.","New languages can be bootstrapped with tens of hours of data or less, and data replay with about 1,000 samples per original language prevents the catastrophic forgetting seen in naive fine-tuning.","The same projector weights transfer to speech translation and topic identification with under five hours of labeled data per language-task, so the ASR-pretrained projector acts as a general spoken-language-understanding initialization."],"fun_headline_variants":["One projector handles ASR in 28 European languages","Single projector enables multilingual speech recognition","MEUSLI: one projector for ASR across 28 languages","Linear projector beats speech encoder on 28 languages","Train once, extend to new languages with minimal data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that Ukrainian and Albanian were excluded from both the speech encoder's and the LLM's pretraining corpora, so that their improvement proves bootstrapping of truly unseen languages; in fact the speech encoder used was trained on both languages.","fun_headline_variants_meta":{"raw":{"variants":["One projector handles ASR in 28 European languages","Single projector enables multilingual speech recognition","MEUSLI: one projector for ASR across 28 languages","Linear projector beats speech encoder on 28 languages","Train once, extend to new languages with minimal data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000484,"raw_usage":{"total_tokens":2209,"prompt_tokens":710,"completion_tokens":1499,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":1425}},"tokens_in":454,"tokens_out":1499,"duration_ms":10465,"temperature":1.0,"reasoning_tokens":1425,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:45:58.839251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same bootstrapping experiment on a language that is verifiably absent from both the speech encoder's and the LLM's pretraining corpora (a low-resource language with a few hours of publicly available audio). If fine-tuning MEUSLI on that language, with ~30 hours of data, fails to produce a large WER improvement over training a projector from scratch, then the paper's claim that multilingual projector pretraining bootstraps languages unseen by the encoder collapses. A minimal check: verify the encoder's training data list for Ukrainian and Albanian; their presence alone falsifies the pap","supporting_citations":[],"review_version":1}