REVIEW 3 cited by
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have been successfully applied for rescoring automatic speech recognition (ASR) hypotheses. However, their ability to rescore ASR hypotheses of casual conversations has not been sufficiently explored. In this study, we reveal it by performing N-best ASR hypotheses rescoring using Llama2 on the CHiME-7 distant ASR (DASR) task. Llama2 is one of the most representative LLMs, and the CHiME-7 DASR task provides datasets of casual conversations between multiple participants. We investigate the effects of domain adaptation of the LLM and context carry-over when performing N-best rescoring. Experimental results show that, even without domain adaptation, Llama2 outperforms a standard-size domain-adapted Transformer-LM, especially when using a long context. Domain adaptation shortens the context length needed with Llama2 to achieve its best performance, i.e., it reduces the computational cost of Llama2.
Forward citations
Cited by 3 Pith papers
-
Customizing Speech Recognition Model with Large Language Model Feedback
LLM log-probability scores combined with acoustic scores serve as RL rewards to adapt ASR models to new domains without labeled data.
-
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
A geometry-independent multi-talker ASR pipeline combining EEND-VC diarization, GSS and SP-MWF enhancement, and a Whisper/WavLM ROVER ensemble reduces CHiME-8 DASR macro tcpWER by 63% relative.
-
Large Language Models based ASR Error Correction for Child Conversations
LLM-based error correction improves zero-shot Whisper and fine-tuned WavLM child ASR transcriptions, but not fine-tuned Whisper, and conversational context as implemented degrades corrections.
Discussion (0). Continue with ORCID to comment.