Pith. sign in

REVIEW 6 cited by

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09818 v3 pith:DTJUJODM submitted 2024-12-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords meralion-audiollmlanguagemodelsaddresslargelinguisticmultilingualsingapore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape. Developed under the National Large Language Models Funding Initiative, Singapore, MERaLiON-AudioLLM integrates advanced speech and text processing to address the diverse linguistic nuances of local accents and dialects, enhancing accessibility and usability in complex, multilingual environments. Our results demonstrate improvements in both speech recognition and task-specific understanding, positioning MERaLiON-AudioLLM as a pioneering solution for region specific AI applications. We envision this release to set a precedent for future models designed to address localised linguistic and cultural contexts in a global framework.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.

  2. AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR

    cs.CL 2025-06 conditional novelty 6.0 of 10

    AsyncSwitch improves code-switched ASR on Whisper by adapting the decoder on text before speech-text alignment and full fine-tuning.

  3. Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Zero-shot audio-language models are not reliable spoof detectors because they over-predict 'spoof', and FP16 quantization keeps this bias while INT8 worsens it.

  4. Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Training a speech-LLM on question-answer pairs generated with both discrete and continuous emotion labels improves its contextual emotion reasoning as scored by an LLM judge.

  5. Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An evaluation of 7 LLMs/LMMs on 3 deception datasets shows fine-tuned text LLMs set benchmarks on review spam while multimodal models lag behind video-based baselines.

  6. Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Whisper-large-v3 encoder with a linear projector and LoRA-tuned Gemma3-12B decoder achieves 16.63% average WER/CER on the MLC-SLM 2025 private test set.

Pith tools