Pith. sign in

Paper Citation Record · LEDGER

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2505.14518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14518 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.382288Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.068351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:36:05.992218Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4f1b436-897a-48e5-b3ba-808817bae70e · outbound

This paper cites These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples These models can process audio, speech, and text in- puts at the same time, using text prompts to extract relevant information from audio and speech

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.659290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.049299Z digest=sha256:91444e30e45191a95076c0d5ab44c523c24913d44350269564ab56bb407cfc3f

Observation 75dea77b-fdc2-4e7f-9441-67ad801a1281 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.642485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.054761Z digest=sha256:b85dce386c7989c164cf019aad537fdc6623f7c977e54aa3ec9f27b1e866d58f

Observation e6e12cee-b610-460a-8574-3860481d703c · outbound

This paper cites This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This is achieved by leveraging a backbone-LLM- synthesized dataset, which automatically generates audio- text pairs and contrastive data across general audio scenarios

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.624963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.060552Z digest=sha256:9a6204dec6147db5eadb65ae22116f748df2409586ab9ccc9981760aae9b4f2b

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · outbound

This paper cites Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:72c4d1871ff88add93076cb1e3dc2e56a47e022b8442c3b46f9ef494bf58fe51

Observation 4b49fbe8-c3f8-41aa-aa0b-139433f3aa65 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.608610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.074285Z digest=sha256:1f4133fbf407df5227e3ea9a05860782c0b98b05aa293c2722d84c344a2a623a

Observation fd548852-9e47-4d82-ae5c-7c9c93d96b22 · outbound

This paper cites For example, Replay the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Replay the audio

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.592156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.078959Z digest=sha256:3a6fe80f3175d77b9ddafff27c3dd9feb1dd2c861141555e85c20ce52080efb7

Observation 4eafbe61-5039-4fd9-8492-4e508434e98f · outbound

This paper cites For example, Identify sounds that are absent as con- trasting examples.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples For example, Identify sounds that are absent as con- trasting examples

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.575502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.083888Z digest=sha256:cb5a6be2d48795462bc0593fc59aea2ef4ac2b811e646be9f224624aef3e5e69

Observation 469d3f93-108c-42c4-ae69-6ab4a7775bfe · outbound

This paper cites It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples It aims to generate descriptions of both the sound events that are present and those that are ab- sent in the audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.558975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.088930Z digest=sha256:ce38f5807ce6e4eaa6035cc1634b479bd9e97b3f8b0d479699e159152caa2c26

Observation 3d1bb6d6-4898-459a-93a1-f395612811ea · outbound

This paper cites We utilize the foundation model Whisper 2.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples We utilize the foundation model Whisper 2

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.542602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.093619Z digest=sha256:f62ac3c02683099ffe28e12d84b815302395f9b0bf5b5a21fc477fc6af18d360

Observation 9ee2869b-0654-4519-8eb0-b6f4ff4d55d2 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.217051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.217051Z digest=sha256:2fa3fe77b0da6c8d965490830b2d827d5cfdf95f4ae880e201d18d78e25a0e34

Observation 6c22f2cb-911a-4f9c-b747-96980c7ce5c7 · outbound

This paper cites Birds chirping 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Birds chirping 3

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.510773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.105038Z digest=sha256:e110f2d947d519109faa939e417ece2673bccbf6e8224d2a70ba4105bb374357

Observation cd18291c-a7ab-44f5-ac71-f74799a1a1a7 · outbound

This paper cites Water pouring Contrastive examples of specific sound events not present in the provided audio:.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Water pouring Contrastive examples of specific sound events not present in the provided audio:

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.493382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.109759Z digest=sha256:cd3d151de28c6a9a1d45583650b31a0e186ea469a447f4abe4cdddc418ce015c

Observation 0c68f654-9c13-4206-9363-6336c62a81af · outbound

This paper cites A dog barking 3.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dog barking 3

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.475584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.115197Z digest=sha256:b2176f8e51eeca737b3e6ac4d1b22c9907055583b13f5b80e98f3e55fb3efb9e

Observation 3e5d6622-a47d-4884-8ba0-13685ee69c8d · outbound

This paper cites This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples This study employs the instruction-tuned LLaMA-3.1-8B 3 [26] as the core large language model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.456880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.120346Z digest=sha256:57cdfcc2cdcfe6986920092b0e04d43457a739e648014ca9896dfde130c2d570

Observation 1f355e0a-7368-4b96-8982-c1b4a618a1d9 · outbound

This paper cites The only trainable component is the audio modality adapter, which is randomly initialized.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The only trainable component is the audio modality adapter, which is randomly initialized

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.439017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.125344Z digest=sha256:752eb3cd5559ea4096ebd9b7d2daae9481299667b7fc93a2235b71d24900285e

Observation 9f58033f-dffe-4a94-84cd-1906ac3f5a24 · outbound

This paper cites yes” and “no.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples yes” and “no

Reference 16

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:36:06.421900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.131174Z digest=sha256:229190552576f8e93e9335510d3597d7894f0597555942ad91ce473ba1c0b072

Observation 73d30f03-ffa5-4921-bfc0-ac948d59e3b7 · outbound

This paper cites Table 3: Evaluation results of our proposed models and other baseline models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Table 3: Evaluation results of our proposed models and other baseline models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.398659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.136258Z digest=sha256:e68cbd54574895701a77834e384ac98bc0466b0854d7ebe7fb436b59b58ba85f

Observation 46477ba4-d8cf-4d09-9b06-f83fb282d92a · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.381357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.142059Z digest=sha256:034fa72d4668ca1756e83f6be621ec7374f51f2f7da9cd55849016ff5bc8f215

Observation 5aa1a652-4638-4162-91f6-bfd795b5ca29 · outbound

This paper cites A combined sam- ple includes both sound events that are present and those that are absent within a single sample.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A combined sam- ple includes both sound events that are present and those that are absent within a single sample

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.362735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.147047Z digest=sha256:1b8fa5e534d98b6ee198695e26108cc539016bce41a79ee1cc972e83032cd30f

Observation e98f23b4-17f8-4c5a-bb9f-46ffd8f625fb · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.343924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.151871Z digest=sha256:9a21502a851d7ac6666450e5841fe1af75551629a0874d938c1d985281b9c04f

Observation b5804312-4735-47b5-94cf-2a482074cda7 · outbound

This paper cites Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Additionally, we achieve impressive results on audio un- derstanding and reasoning benchmarks, demonstrating the ro- bustness and versatility of this approach

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.324918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.157172Z digest=sha256:d1c275bc6cd6218371720a2a0ce173a318e75ba5ead2ec718432fd4a885d08e7

Observation e9d9af8f-5a80-480f-a6df-349e62748f55 · outbound

This paper cites Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.304876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.164063Z digest=sha256:614339699c97fa0fc1a0ca1e43931da467afd6fb345984f5d02f2cca5e04dae4

Observation b1951192-a5f2-4a01-960b-1584a69c1371 · outbound

This paper cites an unresolved cited work.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:06.527171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.099029Z digest=sha256:374c0b6012e17f4a0481afdbb6094c2878863eb502ad6fef2ed6c5492da42ced

Observation 94ecf3d1-352d-4326-8166-25c207e8235b · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.287343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.169985Z digest=sha256:96db52e7072fa3120461d400d23cf528dc13fd6aa6980b3d6528bfea56079920

Observation 947c0734-ee66-4116-a7d4-88bbf7f6a839 · outbound

This paper cites A Survey of Hallucination in Large Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey of Hallucination in Large Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.175875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.175875Z digest=sha256:74e6bc7e7b9b8a5da46b7176795d5b77c00761877099889ce4a76a4d634224b1

Observation 0efa8b6b-08f4-4b6f-a1ec-433f9142b586 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.181805Z digest=sha256:9fb9c9e3c776e162c9b228a227f65e85a6dea5f35de6d9d72e12e087dd138e16

Observation 50aa18dc-f008-4cd9-905d-c0258fc47d71 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.187966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.187966Z digest=sha256:bb9467475011d44f27ee482dd679dfe4db73bd3fa4ca3681ff8249e772121b12

Observation 899a92dc-d6d8-43f3-9cf3-ecfca2d1d066 · outbound

This paper cites Chainpoll: A high efficacy method for LLM hallucination detection.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Chainpoll: A high efficacy method for LLM hallucination detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.194841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.194841Z digest=sha256:7fa5a6cf674463ce24701537dd45d19a6c6d918efadaa8c52ec5a614f55979f0

Observation f16b41d7-dbea-452a-b8dc-f9fa2a49215b · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A Survey on Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.200197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.200197Z digest=sha256:ef7b20f41190669a9dca91bb094bf2cb50191f57f35854820bf88fe2652b9572

Observation 63a548d8-ea27-418a-9405-3d6510adc614 · outbound

This paper cites AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.205353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.205353Z digest=sha256:678528975a11da19ee6ce55c1e852a8240b72616ddfa5472d6e45040237987fd

Observation e32b1938-11de-4f94-be17-2b46538f74ee · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.211406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.211406Z digest=sha256:da4d2079425de4ad1fb2d536c5bcd2782df07900d0acc9c291a970eaab722611

Observation e6f3ca03-9bf5-4559-9d97-c60b4a42442b · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.223027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.223027Z digest=sha256:6bbc9c8d6df153e04395c8b590026ccf49eae803280c4bea467c0e6ca9bc2782

Observation c72d2687-f312-4ef5-9dd6-25246c4334ff · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.229747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.229747Z digest=sha256:30af7fc497ccf10acfa5afbc1fa5b3e6cadd133956dac65cd2d889e4ebbc0796

Observation fea80d42-5259-4f28-9d9e-ca33e2f66681 · outbound

This paper cites Qwen2-Audio Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Qwen2-Audio Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.234847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.234847Z digest=sha256:268bb4ac98d2fc656c88831485fe1dd657a931ff1296cc7aa7e7451a7a8bc789

Observation 9da4d6fc-7fda-4c8a-8d55-0d068e4613d0 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.240818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.240818Z digest=sha256:ea9d244f2af70cd938df669632f7d2a30cac6313f76657b5cb3c573bfb2c999f

Observation 5551c408-3c1d-4ff3-93ea-dcaa35d7e734 · outbound

This paper cites Joint audio and speech understanding,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Joint audio and speech understanding,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.246793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.246793Z digest=sha256:fcc99e1ec1c7d8b81e90f73e9c70488302ae36e04bce048fcbe25e8ea8b1a058

Observation 58143505-a894-43eb-afc5-089c1337a98a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Lora: Low-rank adaptation of large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.256738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.251833Z digest=sha256:dfc7d15222acd49466a3ed3eda6de26e33c5cac56b3b2934ccc4fc7b116f730c

Observation 8149759d-1f1a-47ba-9ca2-8662405adcd8 · outbound

This paper cites Minigpt- 4: Enhancing vision-language understanding with advanced large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Minigpt- 4: Enhancing vision-language understanding with advanced large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.238208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.257062Z digest=sha256:149968b9e96a3385e7e88b932fba05b88f12af7cca09ce863defd5aee4829ead

Observation c9704f10-ca59-4fde-a91b-0c38e62b6af4 · outbound

This paper cites DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.263202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.263202Z digest=sha256:75f3699990574592ebd21b438824a1f1ee5a64e2dcdc450961bb15574f4ee2cc

Observation a733ccb4-0c34-4845-87ea-8fbd8657e308 · outbound

This paper cites Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.271017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.271017Z digest=sha256:596bafd53ca8d91e47aef2102ce8dde95308a44348205cec7446a104b3c8b209

Observation aaf93c1a-b881-4f7b-b57f-b6b74eb67bc2 · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.218844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.276624Z digest=sha256:0fc0664a474a776bd6150ad8889f83ac04e8b381c77fcc15e823f81991f4dec4

Observation c1daf752-92ab-424b-9087-ce9f84192261 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.282635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.282635Z digest=sha256:dc19e0f1fc6dcdfbf037fb08b3db5a545281777b16482b54725db09d33cb4730

Observation 1cfbde63-dbe9-409a-abac-3438c11f0782 · outbound

This paper cites GPT-4 Technical Report.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples GPT-4 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.288005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.288005Z digest=sha256:085dac8997a11b275fa318089c3e4bed3d492a2040297781308cf2b7e49fdd58

Observation b6d267d4-164b-4f24-9cf5-0974a5cdf182 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Robust speech recognition via large-scale weak supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.293881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.293881Z digest=sha256:9706767e52cc5fe48e81e33562677d5b7a83b9e1f9a96399c697340b5c1e2530

Observation 0032467e-b94f-4ce2-8acc-2a0cb0182610 · outbound

This paper cites Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.185045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.305383Z digest=sha256:10b513a315b4023731f8f314e96e6d6ad80d0859b0b6ea400cfda94cead3adcd

Observation baa7458a-3db0-4e98-9cc3-2b3f018e6efc · outbound

This paper cites Investigating the Emergent Audio Classification Ability of ASR Foundation Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Investigating the Emergent Audio Classification Ability of ASR Foundation Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.312182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.312182Z digest=sha256:3dc3c13338bb63accad5888993ff3fcbec310a9049554df0cbd521d065514931

Observation 442c1b89-5608-4404-bffc-1aa7a6840d8c · outbound

This paper cites The Llama 3 Herd of Models.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples The Llama 3 Herd of Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.317879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.317879Z digest=sha256:e8a4040a32e07b75db9f3856a64577e2a8316fbd4c714ded9442c74d4633038b

Observation 53a4cc54-84cf-47db-ba68-63991dd43509 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.322965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.322965Z digest=sha256:8eabd813381920164f0057db27e5c02edcb9f40db55b094ea9c63215079a2d85

Observation f2073a1f-f2a2-4ec3-bf6b-424a85df2694 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audio set: An ontology and human-labeled dataset for audio events,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.328829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.328829Z digest=sha256:beec51bfdc02926652b6c71ee111a65f843852443abe9178b6f0082bf880de60

Observation 657373a1-6163-4ee2-8fe8-96ee22e535f5 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Audiocaps: Generat- ing captions for audios in the wild,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.143910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.334531Z digest=sha256:825f07814de2bfa98f658ddcc5b66e64b5d608cdec423a8ddad8952bc85b5ac9

Observation addbd51c-2bf0-4e82-a861-619b15881c24 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Fsd50k: an open dataset of human-labeled sound events,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.124876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.339782Z digest=sha256:f659c5fc3101be361ba35bc246efd5f0fb484809fa8c376c826bcbec1fab7db3

Observation 819f028b-2fba-4549-bdd4-149be57c1ac9 · outbound

This paper cites Sound event envelope estimation in polyphonic mixtures,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Sound event envelope estimation in polyphonic mixtures,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.106117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.344867Z digest=sha256:abae7c4062204fd6d4b12f408f9433f0d5f05b86faa0473cfa865230787111d2

Observation 583f01e1-27e1-4817-ae55-08b307ffa2f1 · outbound

This paper cites ESC: Dataset for Environmental Sound Classi- fication,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples ESC: Dataset for Environmental Sound Classi- fication,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.350164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.350164Z digest=sha256:09a3671e7a02c3dcbd85ae77a360498604d37e796788b7b7a5b328849f737db6

Observation 053ac22b-de92-4431-8d9f-e1cb13272102 · outbound

This paper cites A dataset and taxonomy for urban sound research,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples A dataset and taxonomy for urban sound research,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.087252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.355490Z digest=sha256:13a8296e413e0ec2b8985772c030db7d2acb79615932b49fa4e3c93693c20772

Observation f83ed0ac-928b-4b17-b340-18cfb2f1b5a7 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Clotho- aqa: A crowdsourced dataset for audio question answering,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.066791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.361096Z digest=sha256:66e34a1dd0340fa841bb0c0e0ece4978e992806447959e154c80fe825662ea24

Observation d0dc9d1c-5198-4574-80bf-55f128e4c85b · outbound

This paper cites V ocalsound: A dataset for improv- ing human vocal sounds recognition,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples V ocalsound: A dataset for improv- ing human vocal sounds recognition,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.043758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.368392Z digest=sha256:29b77a267bea3f876b6a98dc6c2464251e2ee966149eaf556ef253f14ce49811

Observation 52dcb272-344a-4460-959a-850a1a95db75 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.375649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.375649Z digest=sha256:2ebb3949b108aec0e4a2c0ce9bc6ee486aae95747759d740643eac4d4eeae091

Observation d96c217d-6b28-4ea2-aeb0-321c26923575 · outbound

This paper cites What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples What do mllms hear? examining the interaction between llm and audio encoder com- ponents in multimodal large language models,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:06.021026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.382288Z digest=sha256:5d19d045aaa43d79f23aa1ada10ceaa47b738ca650576dded13ad97cfe2d1136

Pith citing papers

Observation 0280248b-6ca7-45b5-a5c7-3d5366219030 · inbound

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples cites this paper.

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:36:05.999867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:05.068351Z digest=sha256:72c4d1871ff88add93076cb1e3dc2e56a47e022b8442c3b46f9ef494bf58fe51