A speech LLM fine-tuned with a fair-average loss outperforms BERT and wav2vec2 baselines on holistic L2 oral proficiency scoring, and transfers across test parts and datasets.
Assessment of L2 Oral Proficiency using Speech Large Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The growing population of L2 English speakers has increased the demand for developing automatic graders for spoken language assessment (SLA). Historically, statistical models, text encoders, and self-supervised speech models have been utilised for this task. However, cascaded systems suffer from the loss of information, while E2E graders also have limitations. With the recent advancements of multi-modal large language models (LLMs), we aim to explore their potential as L2 oral proficiency graders and overcome these issues. In this work, we compare various training strategies using regression and classification targets. Our results show that speech LLMs outperform all previous competitive baselines, achieving superior performance on two datasets. Furthermore, the trained grader demonstrates strong generalisation capabilities in the cross-part or cross-task evaluation, facilitated by the audio understanding knowledge acquired during LLM pre-training.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Assessment of L2 Oral Proficiency using Speech Large Language Models
A speech LLM fine-tuned with a fair-average loss outperforms BERT and wav2vec2 baselines on holistic L2 oral proficiency scoring, and transfers across test parts and datasets.