A SpeechLLM-based ASR system can be fine-tuned with federated learning and LoRA adapters, reaching word error rates close to centralized training on English and Italian while transmitting only a small fraction of the model.
Federated Self-supervised Speech Representations: Are We There Yet?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The ubiquity of microphone-enabled devices has lead to large amounts of unlabelled audio data being produced at the edge. The integration of self-supervised learning (SSL) and federated learning (FL) into one coherent system can potentially offer data privacy guarantees while also advancing the quality and robustness of speech representations. In this paper, we provide a first-of-its-kind systematic study of the feasibility and complexities for training speech SSL models under FL scenarios from the perspective of algorithms, hardware, and systems limits. Despite the high potential of their combination, we find existing system constraints and algorithmic behaviour make SSL and FL systems nearly impossible to build today. Yet critically, our results indicate specific performance bottlenecks and research opportunities that would allow this situation to be reversed. While our analysis suggests that, given existing trends in hardware, hybrid SSL and FL speech systems will not be viable until 2027. We believe this study can act as a roadmap to accelerate work towards reaching this milestone much earlier.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
A SpeechLLM-based ASR system can be fine-tuned with federated learning and LoRA adapters, reaching word error rates close to centralized training on English and Italian while transmitting only a small fraction of the model.