{"id":"6f22311b-8094-44d5-a7de-2739412ad390","arxiv_id":"2412.09818","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new Singapore-tailored AudioLLM, built by fusing a fine-tuned Whisper encoder with SEA-LION V3, shows strong in-domain speech recognition but mixed task-understanding gains versus a cascaded baseline.","lead":"MERaLiON-AudioLLM is a Singapore-focused speech-text model that combines a fine-tuned Whisper audio encoder with the SEA-LION language model. The paper reports competitive results on local speech benchmarks but its task-understanding gains are weaker than a simple cascaded baseline on many tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The task-specific understanding claim is not currently supported: Section 5.1 specifies no judge model, reports no human agreement or confidence intervals, and Table 1 shows the cascaded baseline beating MERaLiON on most SQA/SDS/SI rows.","rationale":"The reader's weakest assumption identifies the unvalidated LLM-as-a-Judge evaluation as the load-bearing point, and my independent read agrees. The paper is a system report, not a derivation, so its evidence is empirical: a released model, a released dataset, and one benchmark table. The strongest positive evidence is the ASR improvements on NSC-based test sets, but those are in-domain by the paper's own description of training on NSC data. The task-understanding improvements, which are central to the abstract, depend on judge scores whose reliability is never established. My reading of Table 1 adds a sharper problem: the cascaded baseline, which the paper motivates its architecture against, is often better on exactly the task-understanding rows. That does not make the model useless; it makes the overbroad abstract claim unsupported. The reader's CONDITIONAL verdict is appropriate because the model and MNSC release are reproducible, useful contributions, and the gaps in evidence could be closed with a proper judge-validation study and out-of-distribution ASR evaluation. I see no reason to move the verdict to REJECT or ACCEPT; the correct disposition remains conditional on the evaluation being shown to be reliable.","tokens_in":14836,"tokens_out":3154,"duration_ms":34732,"concrete_test":"Run a judge-validation study on the released MERaLiON model and the cascaded baseline: sample 100 items per task from SLUE-SQA-5, Spoken-SQuAD, MNSC-SQA-Part 3/6, MNSC-SDS-Part 5/6, OpenHermes-Audio, Alpaca-GPT4-Audio, MELD-Sentiment, and MELD-Emotion; score outputs with a specified judge (e.g., GPT-4o) and with human raters, reporting pairwise judge-human agreement, per-item variance, and bootstrap confidence intervals. If the rankings between MERaLiON and the cascaded baseline flip under a different judge or do not exceed noise, the task-understanding claim in the abstract is unsupported. Separately, evaluate ASR on a held-out Singaporean English set not derived from NSC to test whether the localization gains generalize beyond the training corpus.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's claim of 'improvements in ... task-specific understanding' rests entirely on the non-ASR/non-ST rows of Table 1. For SQA, SDS, SI, and paralinguistics, the paper says scores come from an LLM-as-a-Judge framework (Section 5.1), but it never names the judge model, gives the prompt template, reports human agreement, or provides confidence intervals. Without these, a 2–10 point margin between systems (e.g., 82.9 vs 80.1 on SLUE-SQA-5; 51.4 vs 42.0 on MNSC-SQA-Part 3) cannot be distinguished from judge noise or prompt sensitivity. The problem is compounded by the paper's own cascaded baseline: MERaLiON is worse than the cascaded Whisper+SEA-LION model on 5 of 8 SQA rows, 3 of 4 SDS rows, both SI rows, and both MELD rows. Thus, if 'improvements' means versus the cascaded pipeline the table contradicts the claim; if it means versus other AudioLLMs, the unvalidated judge cannot support it. The ASR-side localization evidence is also partly in-domain: NSC-derived test sets come from the same corpus used for training, so the large WER gains on MNSC-ASR-Part 2 (0.05 vs 0.19) demonstrate fitting to the training distribution more than generalization. The released model and MNSC corpus are valuable community assets, but the central 'task-specific understanding' claim is not independently established as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report introduces MERaLiON-AudioLLM, a speech-text model that combines a Whisper-large-v2-based encoder fine-tuned on Singaporean and Southeast Asian speech data with the SEA-LION V3 Gemma-2-based decoder, connected by an MLP-100 adaptor. The authors release model weights and a multitask corpus (MNSC), and evaluate the model on ASR, speech translation, spoken QA, dialogue summarization, speech instruction following, and paralinguistics against Qwen2-Audio, WavLLM, SALMONN, and a cascaded Whisper+SEA-LION baseline. The paper claims improvements in both speech recognition and task-specific understanding, and positions the model as the first speech-text model tailored for Singapore's multilingual landscape. The evidence for the task-specific claim rests on LLM-as-a-Judge scores in Table 1, while the ASR gains are mostly on NSC-derived test sets that overlap with the training corpus.","tokens_in":15107,"tokens_out":4643,"duration_ms":45380,"significance":"If the empirical claims survive scrutiny, the contribution would be a useful localized speech-text model and a reusable multitask corpus for Singapore English and regional languages, with publicly released weights that could support downstream work in low-resource and dialect-rich settings. The engineering details (data pipeline, FSDP training, compute infrastructure) are also valuable for practitioners. However, the central claims as written are not yet independently established: the task-specific understanding results are based on an unvalidated judge, the comparison to the cascaded baseline contradicts part of the abstract, and the strongest ASR gains are substantially in-domain. The released assets are a genuine community contribution, but the paper's evaluation must be strengthened and its claims scoped before the contribution can be assessed fairly.","major_comments":[{"comment":"The abstract claims 'improvements in task-specific understanding,' but the only evidence for the SQA, SDS, SI, and paralinguistic rows is the LLM-as-a-Judge framework described in Section 5.1. The manuscript does not name the judge model, give the prompt template, report human agreement or calibration, or provide confidence intervals. On this evidence, differences such as SLUE-SQA-5 82.9 vs 80.1 or MNSC-SQA-Part 3 51.4 vs 42.0 cannot be distinguished from judge noise or prompt sensitivity. This issue is load-bearing for the central claim and needs either validated judge scores with human agreement and intervals, or a re-scoped claim that does not assert general task-specific improvements.","section":"Section 5.1 and Table 1"},{"comment":"The phrase 'improvements in both speech recognition and task-specific understanding' is ambiguous and, on the table as reported, partly contradicted. The cascaded Whisper-large-v2 + SEA-LION baseline beats MERaLiON-AudioLLM on 5 of 8 SQA rows, 3 of 4 SDS rows, both SI rows, both MELD rows, and the two Earnings ASR rows (Earnings21-Test 0.17 vs 0.11; Earnings22-Test 0.20 vs 0.14). If the intended comparison is against other AudioLLMs, the abstract should say so explicitly; if it is against the cascaded pipeline, the table does not support the sentence. The claim needs to be restated with an explicit comparison class and the contradictions addressed.","section":"Table 1 and Abstract"},{"comment":"The NSC-derived test sets come from the same corpus used for training the audio encoder and for multimodal instruction fine-tuning. The large MNSC-ASR-Part-2 gain (0.05 vs 0.19 for Qwen2-Audio) therefore largely demonstrates fit to the training distribution, as the paper itself concedes in Section 5.2 ('given its training on in-domain data'). Because localization is a central contribution, the evaluation needs external held-out local data, or a clear statement that the NSC results are in-domain checks rather than evidence of generalization. The conclusion in Section 8 should be tempered accordingly.","section":"Section 3 and Table 1 (MNSC rows)"},{"comment":"All results are single point estimates with no error bars, significance tests, or multiple-seed variation. For rows where margins are small (e.g., LibriSpeech-Test-Clean 0.03 vs 0.03; CoVoST 2 Zh→En 15.0 vs 16.5), the reported 'competitive' or 'best' rankings are not established. Because the central claim depends on relative performance, the paper should report variability (e.g., confidence intervals or multiple runs) or soften comparative statements.","section":"Table 1 (overall)"}],"minor_comments":[{"comment":"The Hugging Face link is given as 'MERaLiON/AudioLLM' in Section 1 but as 'MERaLiON/MERaLiON-AudioLLM-Whisper-SEA-LION' in the footnote; please unify the URLs.","section":"Section 1 and footnote 3"},{"comment":"In the paragraph on MELD, 'the goal is to identity the sentiment' should be 'identify the sentiment'; in Section 7, 'we intent to explore' should be 'we intend to explore'.","section":"Section 5.2"},{"comment":"For the cascaded baseline, only 'We tuned its hyperparameters and prompt template' is reported; please provide the prompt template and hyperparameter choices so the comparison is reproducible.","section":"Section 5.1"},{"comment":"The MNSC corpus is released, but its processing details (segmentation, deduplication, split assignment, and the synthesis of SQA/SDS/GR data) are deferred to future work; a data card or appendix with these details would strengthen the release's reproducibility.","section":"Section 3"},{"comment":"The underlining and bolding conventions are not consistently visible when the cascaded column contains the best result (e.g., several rows have the best score in the Cascaded Model column but are not underlined in the text version); please clarify the formatting.","section":"Table 1"},{"comment":"The MLP-100 adaptor is claimed to give 'slightly better results' than window-level Qformer and ConvMLP, but no comparison numbers or ablations are shown; please add the supporting experiment or remove the claim.","section":"Section 2.1.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a system/technical report with valuable released assets, but it is written more like an announcement than a rigorous evaluation. The 'first' and 'pioneering' framing should be checked against prior Artel and the evaluation strengthened before publication. The LLM-as-a-Judge issue and the unstated comparison class are the main blockers; the authors should be asked to either add a validated human-evaluation component or substantially narrow the claims in the abstract and conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a system report for a Singapore-tuned speech-text LLM: Whisper-large-v2 encoder fine-tuned on local ASR, an MLP-100 adapter, and SEA-LION (Gemma-2 9B) decoder with LoRA. The architecture is not new — they follow SALMONN/Qwen2-Audio/WavLLM — and they cite those. What is actually new and valuable is the release: the model weights on HuggingFace and a new multitask Singapore English corpus (MNSC). The data pipeline details (30 TB, filtering NSC, de-duplication, synthesised tasks) are useful for anyone building localised audio models.\n\nThe paper is also honest. Section 5.2 explicitly says NSC gains are in-domain because the model was trained on that data, and the limitations section admits catastrophic forgetting of instruction following, a 30-second audio cap, and absent safety alignment. That candor counts.\n\nThe soft spots are in the evaluation. The task-specific understanding claim rests on an LLM-as-a-Judge framework with no judge model named, no prompt template, no human agreement, and no confidence intervals. The table margins are often a few points, so the 2–10 point gaps could be judge noise. Worse, their own cascaded baseline (Whisper + SEA-LION) beats MERaLiON on most SQA, SDS, SI, and both MELD rows. So the abstract's 'improvements in task-specific understanding' is not supported against that baseline; it's only supported against the other AudioLLMs, and then only via an unvalidated judge. The ASR wins on NSC are in-domain, though the Earnings21/22 gains are genuine and the ST results are mixed but include real wins on En→Id and En→Zh.\n\nNone of this makes the paper a waste. The released artifacts are the kind of community contribution that justifies a technical report. But as a scholarly claim, the evaluation needs major work: specify the judge, report human agreement and error bars, and reframe 'improvements' relative to the cascaded pipeline they claim to beat.\n\nWho this is for: people working on speech LLMs for Southeast Asian languages, and anyone comparing fused vs cascaded audio-LLM designs. I'd send it to a serious venue after the eval is repaired, likely as a resource/dataset paper. The model and corpus deserve review; the current abstract does not.","headline":"Useful localized speech-text model and new corpus, but the evaluation table is too weak to back the abstract's task-understanding claim; okay as a resource report.","tokens_in":15727,"tokens_out":2519,"would_cite":true,"duration_ms":26547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MERaLiON-AudioLLM is claimed to be the first speech-text model built for Singapore's multilingual speech, and it reports lower local word-error rates plus gains in spoken question answering and summarization.","keywords":["audio large language model","speech-text model","Singapore English","speech recognition","speech translation","spoken question answering","spoken dialogue summarization","multimodal instruction tuning"],"falsifier":"Select a random sample of responses from the spoken question answering, dialogue summarization, and speech instruction test sets and have human annotators score them with the same rubric; if the judge's rankings do not match human rankings, or if a different judge or prompt reverses the model ordering, the task-understanding claim fails. For the speech recognition part, resample the National Speech Corpus test items to form confidence intervals around the word error rate; the localization claim would be weakened if the 5% versus 19% margin over Qwen2-Audio on MNSC-ASR-Part 2 is within sampling noise.","tokens_in":14585,"feed_emoji":"🗣️","tokens_out":9981,"duration_ms":94401,"temperature":0.7,"pith_summary":"This technical report introduces MERaLiON-AudioLLM, an end-to-end model that takes audio plus a text instruction and produces a text answer, designed for Singapore's multilingual and code-switched speech. The authors claim it is the first speech-text model tailored to this landscape, and that fusing a locally fine-tuned speech encoder with a Southeast Asian language model improves both speech recognition and task-specific understanding compared with general-purpose audio LLMs. If the claim holds, communities with distinctive accents and low-resource languages gain a locally usable alternative to cascaded speech-recognition-plus-LLM pipelines, where recognition errors can propagate into the final answer. The report backs the claim with a curated 260,000-hour training corpus, a public release of weights and a multitask subset of Singapore's National Speech Corpus, and evaluations against several published audio LLMs.","feed_headline":"Singapore-tuned audio LLM beats general models on local speech","feed_subtitle":"Fusing a locally fine-tuned speech encoder with a Southeast Asian LLM improves recognition and spoken-task understanding; weights and…","key_machinery":"The load-bearing mechanism is the MLP-100 adaptor, a two-layer MLP that reshapes the audio encoder's 1,500 frame embeddings, each of dimension 1,280, into 100 tokens of dimension 3,854 to match the text decoder's embedding size. This adaptor, together with the fully fine-tuned MERaLiON-Whisper encoder, is trained end-to-end with the SEA-LION V3 decoder, whose MLP layers receive rank-8 LoRA adapters rather than full fine-tuning. The training objective is standard autoregressive cross-entropy over the text output conditioned on both audio and text instruction tokens, and the data pipeline filters the National Speech Corpus for mislabels, assigns identical transcriptions to the same splits, superimposes two-sided conversations, and segments audio to at most 30 seconds.","core_discovery":"On its own terms, the paper's central discovery is that localization transfers through the whole pipeline: fine-tuning the Whisper-large-v2 encoder on cleaned Singapore English data and connecting it end-to-end to a locally pre-trained LLM yields word error rates of 5% on the Multitask National Speech Corpus prompted readings, versus 19% for Qwen2-Audio and 33% for Whisper-large-v2, while the same model stays competitive on unseen Earnings21 and Earnings22 and improves on spoken question answering, dialogue summarization, speech instruction, and accent and gender recognition relative to the audio-LLM baselines. The improvement is attributed chiefly to data curation and in-domain fine-tuning rather than to new architecture, since the fusion design follows other audio LLMs. The paper also reports that the model underperforms on MELD sentiment and emotion, and flags instruction-following loss and a 30-second audio context limit as known limitations.","pith_inferences":["Beyond the paper: if the automated judge scores are later validated by human agreement, the same recipe of fine-tuning a strong open encoder on a national corpus and fusing it with a locally pre-trained LLM should transfer to other under-resourced languages and dialects, because neither component is Singapore-specific.","A testable extension the paper does not run: measure word error rate on Singaporean conversational or telephone audio recorded after the training corpus was collected, to separate genuine local robustness from in-domain memorization of National Speech Corpus prompts.","An editor's inference: the 30-second limit and reported instruction-following loss suggest the release is a proof of concept; a concrete follow-up is to check whether replaying text-only instruction pairs during multimodal fine-tuning restores instruction following without eroding the speech recognition gains.","With the corpus released, an independent audit could verify that no identical transcription appears in both training and test splits, which would directly test the paper's data-leakage-avoidance claim."],"forward_implications":["A model fine-tuned on local data can beat general-purpose audio LLMs on in-country benchmarks while remaining competitive on standard ones such as LibriSpeech and Common Voice.","End-to-end audio-plus-text fusion gives a single model for speech recognition, translation, spoken question answering, dialogue summarization, and speech instruction, avoiding the error propagation of a separate ASR system feeding an LLM.","Publishing the model weights and the Multitask National Speech Corpus lets other groups reproduce the local-curation recipe and apply it to their own regional languages and accents.","The reported gaps on MELD sentiment and emotion indicate that paralinguistic understanding is not solved by this release and needs additional local data or architectural changes, as the paper itself states."],"supporting_citations":[{"why":"Supplies Whisper-large-v2, the encoder backbone that the paper fine-tunes on local data to create MERaLiON-Whisper.","marker":"[Radford et al., 2023]"},{"why":"Supplies SEA-LION V3, the locally pre-trained LLM decoder whose embedding space the audio tokens are aligned to.","marker":"[AI Singapore, 2024]"},{"why":"Provides the National Speech Corpus, the main Singapore English data source for local accent adaptation and evaluation.","marker":"[Koh et al., 2019]"},{"why":"Defines the AudioBench evaluation set, test tasks, and the LLM-as-a-Judge protocol used for all non-ASR results.","marker":"[Wang et al., 2024]"},{"why":"Qwen2-Audio 7B is a primary end-to-end AudioLLM baseline that the paper compares against on every task.","marker":"[Chu et al., 2024]"},{"why":"WavLLM is a primary baseline, and its dual-encoder curriculum approach is the main methodological contrast.","marker":"[Hu et al., 2024]"},{"why":"SALMONN is a primary baseline and supplies the window-level Qformer alternative that the MLP-100 adaptor is compared to.","marker":"[Tang et al., 2024]"},{"why":"Provides NSC-derived ASR, SQA, and SDS test sets used to measure localized understanding beyond the original corpus.","marker":"[Wang et al., 2025]"},{"why":"Supplies LoRA, the low-rank adaptation method used to partially fine-tune the text decoder while training the full model.","marker":"[Hu et al., 2022]"}],"fun_headline_variants":["Singapore audio LLM cuts local WER from 33% to 5%","Local audio LLM beats Whisper and Qwen on Singapore speech","Singapore-tuned audio LLM achieves 5% WER on local speech","Audio LLM for Singapore: 5% WER, outperforms big baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The task-understanding part of the claim rests on automated judge scores whose reliability is never established, so if the judge is noisy or biased, the reported gains in spoken QA, summarization, instruction, and paralinguistic tasks are unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Singapore audio LLM cuts local WER from 33% to 5%","Local audio LLM beats Whisper and Qwen on Singapore speech","Singapore-tuned audio LLM achieves 5% WER on local speech","Audio LLM for Singapore: 5% WER, outperforms big baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001106,"raw_usage":{"total_tokens":4565,"prompt_tokens":854,"completion_tokens":3711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":3627}},"tokens_in":470,"tokens_out":3711,"duration_ms":29519,"temperature":1.0,"reasoning_tokens":3627,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:41:30.013109+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of responses from the spoken question answering, dialogue summarization, and speech instruction test sets and have human annotators score them with the same rubric; if the judge's rankings do not match human rankings, or if a different judge or prompt reverses the model ordering, the task-understanding claim fails. For the speech recognition part, resample the National Speech Corpus test items to form confidence intervals around the word error rate; the localization claim would be weakened if the 5% versus 19% margin over Qwen2-Audio on MNSC-ASR-Part 2 is within sampling noise.","supporting_citations":[{"cited_title":"SEA-LION (southeast asian languages in one network): A family of large language models for southeast asia","cited_arxiv_id":null,"evidence_quote":"Supplies SEA-LION V3, the locally pre-trained LLM decoder whose embedding space the audio tokens are aligned to."}],"review_version":1}