Pith. sign in

REVIEW 1 cited by

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03668 v2 pith:3UVMOWQR submitted 2023-12-06 eess.AS cs.AIcs.CLcs.LG

classification eess.AScs.AIcs.CLcs.LG
keywords speechmodelmodelspre-trainedlanguageproposedacousticenables
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, most of their applications in ASR involve only one of either a pre-trained speech or a language model. This paper proposes integrating a pre-trained speech representation model and a large language model (LLM) for E2E ASR. The proposed model enables the optimization of the entire ASR process, including acoustic feature extraction and acoustic and language modeling, by combining pre-trained models with a bridge network and also enables the application of remarkable developments in LLM utilization, such as parameter-efficient domain adaptation and inference optimization. Experimental results demonstrate that the proposed model achieves a performance comparable to that of modern E2E ASR models by utilizing powerful pre-training models with the proposed integrated approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligning Pre-trained Models for Spoken Language Translation

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Frozen speech recognition and machine translation models can be aligned by a small connector network to perform end-to-end speech translation, and the connector also serves as a domain adapter.

Pith tools