REVIEW 8 cited by
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires significant annotation efforts and risks catastrophic forgetting of the original language capabilities. In this work, we present a simple yet effective automatic process for creating speech-text pair data that carefully injects speech paralinguistic understanding abilities into SLMs while preserving the inherent language capabilities of the text-based LLM. Our model demonstrates general capabilities for speech-related tasks without the need for speech instruction-tuning data, achieving impressive performance on Dynamic-SUPERB and AIR-Bench-Chat benchmarks. Furthermore, our model exhibits the ability to follow complex instructions derived from LLMs, such as specific output formatting and chain-of-thought reasoning. Our approach not only enhances the versatility and effectiveness of SLMs but also reduces reliance on extensive annotated datasets, paving the way for more efficient and capable speech understanding systems.
Forward citations
Cited by 8 Pith papers
-
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
SAKURA shows large audio-language models struggle with multi-hop reasoning from speech and audio even when they correctly perceive the needed attribute.
-
Towards Reliable Large Audio Language Model
Training a large audio language model to say 'I don't know' on one audio type (speech, music, or sound) makes it more likely to refuse uncertain questions on the other types.
-
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
Most speech-aware language models follow written output-format instructions far worse than their text-only base LLMs, and Speech-IFEval measures this as catastrophic forgetting.
-
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
A contrastive-style adapter trained on LLM-generated positive and negative audio descriptions improves audio hallucination accuracy to 77.5 percent and audio question answering to 84.3 percent, without changing the fr...
-
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
The authors release MNSC, the largest standardized multitask spoken Singlish corpus, and SingAudioLLM, a multimodal model that sets strong baselines on ASR, spoken QA, dialogue summarization, and paralinguistic QA.
-
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
ORCA predicts the distribution of human correctness ratings for open-ended audio QA answers and matches or beats LLM judges while also estimating annotator disagreement.
-
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
In a three-stage end-to-end spoken language model, experience replay (mixing old data into later training) was the most effective mitigation against catastrophic forgetting, greatly outperforming model merging and LoR...
-
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
AlignFormer, a CTC-guided dynamic-window adapter, lets a frozen LLM trained on ASR data alone achieve high instruction-following rates on zero-shot speech translation and question answering.
Discussion (0). Continue with ORCID to comment.