LOPA with SALR achieves RMSE 0.361 on spoken language assessment by enforcing ordinal prototype alignment in latent space from a frozen Whisper encoder.
LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Language Assessment (SLA). While effective, this paradigm often overlooks the intrinsic ordinal structure of language acquisition. This paper works around the necessity of large-scale MLLMs by introducing Latent Ordinal Prototype Alignment (LOPA) for SLA, a prototype-based regularizer that enforces an ordinal geometric prior directly on the latent space. Coupled with Semantic-Anchored Layer Routing (SALR), which adaptively harvests multi-depth representations from a frozen Whisper encoder, our framework achieves an RMSE of 0.361. This performance rivals billion-parameter systems without the need for LLM-based fine-tuning. Further analysis reveals that SALR's synergy with LOPA offers interpretable, criterion-aligned preferences, thereby supporting an efficient and ordinal-aware modeling alternative to current scaling-centric models for SLA.
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment
LOPA with SALR achieves RMSE 0.361 on spoken language assessment by enforcing ordinal prototype alignment in latent space from a frozen Whisper encoder.