A contrastive-style adapter trained on LLM-generated positive and negative audio descriptions improves audio hallucination accuracy to 77.5 percent and audio question answering to 84.3 percent, without changing the frozen language model.
For example, Identify sounds that are absent as con- trasting examples
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
A contrastive-style adapter trained on LLM-generated positive and negative audio descriptions improves audio hallucination accuracy to 77.5 percent and audio question answering to 84.3 percent, without changing the frozen language model.