Pith. sign in

REVIEW 3 cited by

COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.02248 v2 pith:KFT2M7CE submitted 2023-11-03 cs.CL cs.AIeess.AS

classification cs.CLcs.AIeess.AS
keywords speechcosmiccapabilitiesinstruction-followingshotcontextualdatain-context
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a cost-effective method to integrate speech into a large language model (LLM), resulting in a Contextual Speech Model with Instruction-following/in-context-learning Capabilities (COSMIC) multi-modal LLM. Using GPT-3.5, we generate Speech Comprehension Test Question-Answer (SQA) pairs from speech transcriptions for supervised instruction tuning. With under 30 million trainable parameters and only 450 hours of English speech data, COSMIC demonstrates emerging capabilities in instruction-following and in-context learning. Equipped with such capabilities, COSMIC achieves a maximum 33.18 BLEU score in 0-shot EN-to-X speech to text translation (S2TT) and a significant boost in the 1-shot setting. Additionally, there is an average 25.8\% relative Word Error Rate (WER) reduction for 1-shot cross-domain adaptation. COSMIC exhibits a significant automatic speech recognition (ASR) accuracy gain in contextual biasing tasks due to its instruction-following capability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition

    eess.AS 2025-05 conditional novelty 7.0 of 10

    LibriSpeech and Common Voice evaluation sentences leak into the Pile, and controlled LLM pretraining experiments show that contamination biases output probabilities even when error rates barely change.

  2. Speechless: Speech Instruction Training Without Speech for Low Resource Languages

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.

  3. MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

    cs.SD 2026-01 conditional novelty 4.0 of 10

    Post-training an auditory LLM with an episodic in-context objective on high-resource speech data improves low-resource child-speech ASR and audio reasoning beyond direct fine-tuning and vanilla ICL.

Pith tools