REVIEW 6 cited by
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape. Developed under the National Large Language Models Funding Initiative, Singapore, MERaLiON-AudioLLM integrates advanced speech and text processing to address the diverse linguistic nuances of local accents and dialects, enhancing accessibility and usability in complex, multilingual environments. Our results demonstrate improvements in both speech recognition and task-specific understanding, positioning MERaLiON-AudioLLM as a pioneering solution for region specific AI applications. We envision this release to set a precedent for future models designed to address localised linguistic and cultural contexts in a global framework.
Forward citations
Cited by 6 Pith papers
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
A new spoken math benchmark, Spoken-MQA, shows that current speech-based AI models reason poorly from spoken math input, especially for arithmetic and knowledge-heavy problems.
-
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
AsyncSwitch improves code-switched ASR on Whisper by adapting the decoder on text before speech-text alignment and full fine-tuning.
-
Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
Zero-shot audio-language models are not reliable spoof detectors because they over-predict 'spoof', and FP16 quantization keeps this bias while INT8 worsens it.
-
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
Training a speech-LLM on question-answer pairs generated with both discrete and continuous emotion labels improves its contextual emotion reasoning as scored by an LLM judge.
-
Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings
An evaluation of 7 LLMs/LMMs on 3 deception datasets shows fine-tuned text LLMs set benchmarks on review spam while multimodal models lag behind video-based baselines.
-
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
A Whisper-large-v3 encoder with a linear projector and LoRA-tuned Gemma3-12B decoder achieves 16.63% average WER/CER on the MLC-SLM 2025 private test set.
Discussion (0). Sign in to comment.