REVIEW 2 cited by
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, speech to speech translation. Moreover, large language models (LLMs) pretrained from vast text corpora contain rich linguistic information that can improve accuracy in a variety of tasks. In this paper, we present a decoder-only Discrete Multimodal Language Model (DMLM), which can be flexibly applied to multiple tasks (ASR, T2S, S2TT, etc.) and modalities (text, speech, vision). We explore several critical aspects of discrete multi-modal models, including the loss function, weight initialization, mixed training supervision, and codebook. Our results show that DMLM benefits significantly, across multiple tasks and datasets, from a combination of supervised and unsupervised training. Moreover, for ASR, it benefits from initializing DMLM from a pretrained LLM, and from a codebook derived from Whisper activations.
Forward citations
Cited by 2 Pith papers
-
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
A comprehensive review of end-to-end multi-speaker ASR that contrasts SIMO and SISO architectures and reports that no design wins consistently, with real-world benchmark progress stagnant since 2021.
-
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
In a small multimodal language model for ASR, lip movements give the largest relative benefit at high noise while image/OCR context peaks at moderate noise, but the effect depends on architecture and input format.
Discussion (0). Continue with ORCID to comment.