SAKURA shows large audio-language models struggle with multi-hop reasoning from speech and audio even when they correctly perceive the needed attribute.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
SAKURA shows large audio-language models struggle with multi-hop reasoning from speech and audio even when they correctly perceive the needed attribute.