A speech large language model trained on beamformed multi-channel audio performs directional speech recognition and source localization across 12 discrete angles on simulated smart glasses data.
Room impulse responses (RIRs) from real environments are used to model spatial diver- sity, generating 12 distinct directions at 30° resolution
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
A speech large language model trained on beamformed multi-channel audio performs directional speech recognition and source localization across 12 discrete angles on simulated smart glasses data.