By projecting hidden states to the vocabulary at every layer, the paper shows that failed attribute recognition in three LALMs is marked by mid-network information peaks followed by degradation, and that models rely on direct audio queries rather than consolidated attribute states in text tokens.
Developing instruction-following speech language model without speech instruction-tuning data,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
By projecting hidden states to the vocabulary at every layer, the paper shows that failed attribute recognition in three LALMs is marked by mid-network information peaks followed by degradation, and that models rely on direct audio queries rather than consolidated attribute states in text tokens.