Pith. sign in

Paper Citation Record · LEDGER

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2607.13408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13408 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:20:03.868153Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:19:55.932658Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 897b7fd3-7c2c-496e-abb6-b3bc3fcc7be8 · outbound

This paper cites Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:55.932658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:55.932658Z digest=sha256:e4e449a4bb935a8fe495779c222c3df92fb6d9d31bb43a1b6eaa4f15f64ad1ef

Observation 93faeea3-4eac-431b-a8db-3976a2457343 · outbound

This paper cites an unresolved cited work.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.009130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.009130Z digest=sha256:98b935aea73f5f5b99343e60d5cbf39390a441fea2b4bb38cb5a8e63f95449ff

Observation c552ab95-18f0-40c9-bd4d-1a1f569710d2 · outbound

This paper cites Rain is falling con- tinuously.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Rain is falling con- tinuously

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.153317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.153317Z digest=sha256:bf7175efe0724c20384257bd583daa0cd28892927a29d561871502b065f74055

Observation 7bed7c70-e320-44be-81fc-a962e42fcab3 · outbound

This paper cites It starts with the sound ofe 1, shifts toe 2, and ends withe 3.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models It starts with the sound ofe 1, shifts toe 2, and ends withe 3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.317477Z digest=sha256:82b1dedb12cd5869327f6b0520d13c8d52bc3be1905b01662d9bf8f119d10f60

Observation 261244ee-e32a-471f-851a-887ae4373772 · outbound

This paper cites Verifying ALLMs as Judges Table 1 reports ALLM performance on audio understanding benchmarks with ground-truth annotations.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Verifying ALLMs as Judges Table 1 reports ALLM performance on audio understanding benchmarks with ground-truth annotations

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-02T05:19:56.421451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.421451Z digest=sha256:b6bbb8dd09ca1b30003982bf4764ef577e8e3f6f1971912cbd122024b20aac34

Observation 8d3e11d5-03f1-4d28-bbe9-0b3d50843993 · outbound

This paper cites While recent systems achieve strong percep- tual realism, our study shows that they often overlook fine- grained requirements such as sound event completeness and temporal ordering.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models While recent systems achieve strong percep- tual realism, our study shows that they often overlook fine- grained requirements such as sound event completeness and temporal ordering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.578105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.578105Z digest=sha256:d13b9d5ce7489663adcd727db1255e6fa6a4f5a696f69eb8ceee2de07586a258

Observation 73364c19-88fc-4a25-ab64-3abf510cc4e2 · outbound

This paper cites an unresolved cited work.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.685633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.685633Z digest=sha256:9dd049371a2a2febdda15c56f0dbf15e7c9385cf88942c0b614d8785acb90411

Observation 83e62017-cf58-492c-8725-d771dc9ac622 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioGen: Textually Guided Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.888125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.888125Z digest=sha256:7e40e02b072a75bdf11203d3f342ca7d572eb31ad3ab0649ce49dc24612ec64a

Observation 3b8dd948-6f65-466b-a3cc-46a277751328 · outbound

This paper cites MusicLM: Generating Music From Text.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MusicLM: Generating Music From Text

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.911402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.911402Z digest=sha256:8bc853592c44feec0d2ecd9f69f041e790186fb64d8af9ebe3818dbf8bbbfd4f

Observation cd851abf-2a24-4ff8-b554-cb7618658e91 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.963850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.963850Z digest=sha256:ac5b45d87083edaf3f57b77067422cfa3642b741ca506b8ef588e2c84916b351

Observation cd8dac01-cc08-4c03-b8de-40e5cddd78c1 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.089048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.089048Z digest=sha256:11180b7db196513033fc245184e950fa563922f751d3a6a0fc3184ba1523f796

Observation 348e50f9-7e8a-4b55-b7c6-6fe9059b08ca · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound genera- tion,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Diffsound: Discrete diffusion model for text-to-sound genera- tion,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.219463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.219463Z digest=sha256:65cc280f48c977075786220fb60e90c60891235abbda79373ffaa05800dbfcb7

Observation 5c55add6-096b-489f-90d4-2bc60330662f · outbound

This paper cites Towards general-purpose text-instruction-guided voice conversion,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Towards general-purpose text-instruction-guided voice conversion,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.281224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.281224Z digest=sha256:eee9b2356586268f74bf933b365eef971efbb0f47bb825fbb5701e758c2fe9d3

Observation 7c05fc79-d7d6-4cb5-999f-5c2659632ce1 · outbound

This paper cites Simple and controllable music generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Simple and controllable music generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.382379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.382379Z digest=sha256:d54780dffbf52606671a2fc6aefd3615ae95f6b7162a5ca61bef36ff3104ca6f

Observation 7ad275b0-2ef7-460e-b96a-95909d3f937f · outbound

This paper cites Audiolm: a language modeling approach to audio gener- ation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiolm: a language modeling approach to audio gener- ation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.450207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.450207Z digest=sha256:fe7d43ce82dcaff110cc3a959f042f821b38efd4c8e88d1442ea349266e01da3

Observation c3acb87d-c10d-4f9a-8ab0-17fa3de23416 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models SoundStorm: Efficient Parallel Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.547298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.547298Z digest=sha256:8bbc2c990f1f48d3cb14985511f0f71896b59894bc9e85a8d5e3fa32282a74ab

Observation 99729a8c-4d66-420f-867e-6f3bdbb85b88 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.626028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.626028Z digest=sha256:0fcd461acc21c1b5c49d22163331622d53b3227c3ead360b4dc2cc67b1209ed5

Observation f1542adc-c504-493e-b5cc-0d8a9df23b56 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.779448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.779448Z digest=sha256:f1939b4e07918a40c993e72dee9cb2aecbf7571dec9e9f464eaca8bbfe0d870b

Observation af68f29d-04c2-4789-a505-a562ef0ef7b0 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap- ranked preference optimization,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap- ranked preference optimization,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.897476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.897476Z digest=sha256:e8138082e53f67ff7375ee174e8305e32c035ed59f285a7df81c54e45690b76b

Observation c7753fe2-0ae1-45b0-a65f-29f274af05eb · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.993408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.993408Z digest=sha256:6561274161c1d151a7156c71b5c5de22cb5006be6c50b8dd28101bd1409b43d4

Observation 38c0a192-db20-4807-8219-d48127efaa21 · outbound

This paper cites Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.141475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.141475Z digest=sha256:5216dfc351fcb57dbc261287822c834820467c8e3824eb350bfaab11355e8fcb

Observation c69f1c51-9f5b-4103-901d-165ab4231797 · outbound

This paper cites Stable audio open,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Stable audio open,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.321242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.321242Z digest=sha256:1fad68205af7b5c8e012cdfb8b1787a28cdf8c494f0342019748c0fa9958d457

Observation bf2e3ebf-2a9f-4891-995d-e67f26ee13e9 · outbound

This paper cites Etta: Elucidating the design space of text-to-audio models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Etta: Elucidating the design space of text-to-audio models,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.435218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.435218Z digest=sha256:b37105d5f15914d46d937ce9c5f71f30d5eb642785172ac29c012f27541ef454

Observation 137f8c6c-2bdb-45d0-a2e6-47e2e1cae64a · outbound

This paper cites Impact: Iter- ative mask-based parallel decoding for text-to-audio generation with diffusion modeling,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Impact: Iter- ative mask-based parallel decoding for text-to-audio generation with diffusion modeling,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.585617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.585617Z digest=sha256:89cd2b5dcbd07e118a2ed8ec0ec92770cd982b39c046eed0d3876356001b461b

Observation bf6b0368-4ac5-4de3-9de7-09775fd8cde0 · outbound

This paper cites Gen- erative audio language modeling with continuous-valued tokens and masked next-token prediction,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Gen- erative audio language modeling with continuous-valued tokens and masked next-token prediction,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.676033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.676033Z digest=sha256:52bde737ba83dbba48083156157b83b4de1bc2821d72c22aa09f19786150381f

Observation faf29cef-d559-44e2-a290-673d1bece393 · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.829580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.829580Z digest=sha256:f6dc06988f7b83f4b0ae5f37cefe0d0135386022be55371e0d4e07162f7d7e66

Observation f28e2bb0-5c2d-47e1-800d-a3f9a2f950b1 · outbound

This paper cites Improved techniques for training gans,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improved techniques for training gans,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.968651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.968651Z digest=sha256:38a827fa7733692b4eabb7f34110aaa74c0b876ff0d1151335ae81e6f5f38cb3

Observation 68fafd2c-9de4-478b-8e25-965cce8f216a · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Clap learning audio concepts from natural language supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.107979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.107979Z digest=sha256:bf23dc8a5019aa0c1425f1d30714bf14895045c0e3b136d66b87bd438455ff91

Observation 823b4a65-1815-4089-bea7-f33063c84884 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.212491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.212491Z digest=sha256:0ddc690c485104b13d77706d63a6801a623d0cdf93c5f00d283aa1afacee13bc

Observation cebd6c1b-4a32-480c-92df-1ae59215ff7e · outbound

This paper cites Ritta: Model- ing event relations in text-to-audio generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Ritta: Model- ing event relations in text-to-audio generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.335502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.335502Z digest=sha256:21cdb9aba173cdc5121e9ec8c893f482749b281e1baf7e406cc5755e0c6b20e8

Observation e1032269-da52-42aa-bf86-def53b563435 · outbound

This paper cites Aurelius: Relation aware text-to-audio generation at scale,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aurelius: Relation aware text-to-audio generation at scale,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.441087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.441087Z digest=sha256:b592cd00d6adcd2dd061f93785c89daba621643394d91cfd7df1e83b15a0f100

Observation 902de066-9176-4336-8df2-c8a8c476b0ca · outbound

This paper cites Compa: Address- ing the gap in compositional reasoning in audio-language mod- els,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Compa: Address- ing the gap in compositional reasoning in audio-language mod- els,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.620138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.620138Z digest=sha256:d7681aea730c8fec64b63b09084e5fbeeae88c9018c14645d0803028f7fb165e

Observation f9f1773e-935e-4788-a6ef-ea24d36f11a5 · outbound

This paper cites T2a-feedback: Improving basic capabilities of text-to-audio generation via fine-grained ai feed- back,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models T2a-feedback: Improving basic capabilities of text-to-audio generation via fine-grained ai feed- back,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.789891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.789891Z digest=sha256:0e93a9d81646b6c4757ef8bafa44af21b44fa5e198282ae8f5657fcccf063d20

Observation 90a1d0a4-25db-4132-a3a1-62995b35a5b4 · outbound

This paper cites Aqascore: Evaluating semantic alignment in text-to-audio generation via audio question answering,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aqascore: Evaluating semantic alignment in text-to-audio generation via audio question answering,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.939522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.939522Z digest=sha256:7037fc2bbb128b79ca1722a7f53cf99358a8183c27ac37f2b4d3902bc3ac444c

Observation c044ba76-7e09-45c9-997c-dfaf3b24b02c · outbound

This paper cites The Sound of Absence: Audio-Language Embedding Models Struggle with Negation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.076142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.076142Z digest=sha256:fc75ede0a1f1344063046e6972ec900144b783cfcf346a3989a42bb2b9d6afc4

Observation ea01226f-cdac-4df2-a716-7a4c648b1874 · outbound

This paper cites GPT-4 Technical Report.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.190570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.190570Z digest=sha256:50b08912508d2d07ba57913c955e6611e892ac76fa6c31fb10e4eda608f49fd4

Observation 858fae87-adb7-410e-9d45-36588d562764 · outbound

This paper cites Joint audio and speech understanding,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Joint audio and speech understanding,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.385579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.385579Z digest=sha256:0e27f8f19b23e50b85ba85f0ede4ed5a9e692346f2ec52386063242ab92d7be1

Observation 0f35bbe4-350f-4beb-bfd6-a448f097bb72 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.587754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.587754Z digest=sha256:360d1901dee6d491e145d45a6fa42e94dbee779a8d2246a0413fe4e5f33f8bb0

Observation eeb694ec-2354-4fac-a910-630a9669a28b · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.724635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.724635Z digest=sha256:3797e056ff6a579a3c233a95557ef4ec366ae3c68e8d7c2a26ac9ec2b34fd304

Observation 404f282e-2fdc-4d66-a562-0a1f6a605364 · outbound

This paper cites Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.835510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.835510Z digest=sha256:eae3b4088bb6b20f6fa280913375786757962f0e5fff6d855ee60a00784cffad

Observation b6d54ccc-0e7d-4188-864e-5c65a1e75323 · outbound

This paper cites Speechprompt: Prompting speech language models for speech processing tasks,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Speechprompt: Prompting speech language models for speech processing tasks,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.957127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.957127Z digest=sha256:14fecdb7c2bc4730539775e62551d2a91e31335ea4ad9a6654739e69cc98ae90

Observation 2b6a02bc-8860-4fbe-b2d4-40658e22f1b0 · outbound

This paper cites Blsp- emo: Towards empathetic large speech-language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Blsp- emo: Towards empathetic large speech-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.082235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.082235Z digest=sha256:756380970da396f3d3875177cbd346b60da730f4f1170997f477522afff28129

Observation 63214a6e-6535-42e2-aec1-579dd8c11ab5 · outbound

This paper cites Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.281092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.281092Z digest=sha256:dfead6e27408ecc91694aa055877215242a85ec74905100df3b20436b535a8f2

Observation 608bbfa2-db29-48f0-b7a6-465c09ffa510 · outbound

This paper cites Voxtral.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Voxtral

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.370603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.370603Z digest=sha256:0aeb8b222650283a3e93ad5ca52d9f9f5e68b36711091a18e804f757ee43dae1

Observation 24cfd24c-19b1-49d1-98a5-cd9b15b8f084 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Qwen2.5-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.448742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.448742Z digest=sha256:fb550436228bf1d8d1b790d898b21c16c928ffad612992e6e9871f12621f0e7a

Observation 643cea81-7953-4c92-b12c-91f3407746b5 · outbound

This paper cites Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.542825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.542825Z digest=sha256:1286828093756792f1179a6c4f673667d455cad2565158a19146a8454204638e

Observation 4db4dd76-378d-41cc-9a06-1be24eb3645d · outbound

This paper cites From alignment to advancement: Bootstrapping audio- language alignment with synthetic data,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models From alignment to advancement: Bootstrapping audio- language alignment with synthetic data,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.637467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.637467Z digest=sha256:41277a1dc51ff7228b9c9f216d666ff0b044d1db699521d9df1635a410ec1665

Observation c0c038e6-b470-4114-998a-f34b9a784af4 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.741673Z digest=sha256:21425d1c2eb9a5707f575d24b8eed1b08c713f52d552b314d94d5969bf9220ba

Observation 51859531-fef3-4eb6-a497-416d668ceead · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.846967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.846967Z digest=sha256:ea9b1ab33c3ade9271ac5670a79584698d7937f1f91c86a1690a72a9d968fe23

Observation 82966f7a-01b6-4c22-b376-bebc7cb45090 · outbound

This paper cites On the landscape of spoken language models: A comprehensive survey,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models On the landscape of spoken language models: A comprehensive survey,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.913145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.913145Z digest=sha256:0fc1a07d15e61fd7b875cbb84e303c01466de2cb8bc4675226890f0e8112ce64

Observation 1c2430b7-ebd3-4959-89b8-d6be49899dd5 · outbound

This paper cites Audio flamingo 2: An audio- language model with long-audio understanding and expert reason- ing abilities,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio flamingo 2: An audio- language model with long-audio understanding and expert reason- ing abilities,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.010846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.010846Z digest=sha256:8dd94436f2cd5276e2a11aafb64e42d04c95a446d1d6b133518916d181cda434

Observation 437b6dca-2413-4b2e-a766-918f3ce965c9 · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.115586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.115586Z digest=sha256:4adf07591bdde0a6083c2e0464f6cb173e41fe1ed5ff7c152883b98a0ded40da

Observation d21e50ec-7b39-44ad-bc3d-c16394ff7da3 · outbound

This paper cites Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.190850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.190850Z digest=sha256:2ccf0c91b8046fbec897988f0c91923db875e0b2e31ed8971e25da0b8ed1574a

Observation 22d14206-0d01-4eab-9f4a-3efedc4a53f9 · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.256101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.256101Z digest=sha256:654528581692909a0be8c15bdf9fbd865e91885132a4808d86cf21b8d5ec0861

Observation 8f6b06dc-021c-424d-b487-6393d494879e · outbound

This paper cites Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.393294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.393294Z digest=sha256:20d71f0a5992f91e3c7e728053a51ed60a8707a2eb939011c0825ab50c71a766

Observation bc7f16a0-4878-48a4-95f3-1e4acbb3d1cb · outbound

This paper cites Mmau: A mas- sive multi-task audio understanding and reasoning benchmark,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Mmau: A mas- sive multi-task audio understanding and reasoning benchmark,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.526891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.526891Z digest=sha256:0e47feadbde3e9e116d28b5aebf287ed90abcfd01f3aabd28868b4d7d70d747a

Observation 6725413c-ccf2-44a3-92cf-440a8c763c73 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.638568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.638568Z digest=sha256:7afacd8629c5990f952c3c620c35fae6a4cee43923e07c68dd8f51eb1a935ece

Observation 845f6050-4f64-412b-ab73-8ac9afaa3b9a · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.704210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.704210Z digest=sha256:33131359f677e1ee617dca1084ec4f6e09c508460476a50488bc56d2cbe77627

Observation bbe15c22-57ad-49e3-922d-0172e2fe2e30 · outbound

This paper cites Game-time: Evaluating temporal dynamics in spoken language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Game-time: Evaluating temporal dynamics in spoken language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.807832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.807832Z digest=sha256:771c085d3b676386fd1cec1aad338f063ca9d9ee9dfd06c40dfb008b5794a903

Observation 96f2e11a-f2dc-4ef6-a2ca-3096fb59c02c · outbound

This paper cites Aqua-bench: Beyond finding answers to knowing when there are none in audio question answering,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aqua-bench: Beyond finding answers to knowing when there are none in audio question answering,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.895515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.895515Z digest=sha256:bcb70aaf4427988e347c1b93f1dc8490324ae41d805b88e5f445a564910e226a

Observation 4a70f1b7-c7c8-4ec6-ad7f-eb66b453401b · outbound

This paper cites Baton: aligning text-to- audio model using human preference feedback,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Baton: aligning text-to- audio model using human preference feedback,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.015060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.015060Z digest=sha256:7336a9e75935fc55335e9b912a0875d3335dd6cfd9214dc3385678377a37d4ed

Observation 11e27ef6-1daa-4dc8-89f9-9050d62389a4 · outbound

This paper cites Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.080402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.080402Z digest=sha256:79de9e1ac170ae937eb8951de90b347ed0dba8323f70ff48ae0b4165a0af5cf5

Observation bf460991-dbb7-491d-aefe-0c7c596b4ab9 · outbound

This paper cites Audio large language models can be descrip- tive speech quality evaluators,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio large language models can be descrip- tive speech quality evaluators,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.189659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.189659Z digest=sha256:7f7a089dc6e7ba18b24a025a01c0fbb134d1d2b9e17b594a0c4b2eb012662856

Observation b988b51f-4411-49b7-b030-50c369e351cb · outbound

This paper cites AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.280286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.280286Z digest=sha256:a63cb1c49da774839d0f22e0f55a825a2434936ec060f8124eafabf6801accb2

Observation fcfb353f-9297-47b6-a57a-556b13f36469 · outbound

This paper cites Audio-aware large language models as judges for speaking styles,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio-aware large language models as judges for speaking styles,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.382877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.382877Z digest=sha256:801f1feb18c6cedc71398fb89cefab4ab503307843bbd8e052f5d3a0d0ff5faf

Observation 2cc58875-3fa4-4603-9419-b2cabe24bced · outbound

This paper cites InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.452297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.452297Z digest=sha256:46931ab0860c4b73d45da2fadd9ea97b282fac0b384d3dc37f052eff3a8a6de0

Observation ddab289b-5042-4030-a525-02d0dc13abbd · outbound

This paper cites Audioeval: Automatic dual-perspective and multi-dimensional evaluation of text-to-audio-generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audioeval: Automatic dual-perspective and multi-dimensional evaluation of text-to-audio-generation,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.583905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.583905Z digest=sha256:c360165940005d1a1fcb5dc0f64f5a3bb545ab6b2aa5226a8e25da9b48416014

Observation 0a5ecb1b-92f2-4fb1-8830-1452c7d0a9d3 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiocaps: Generat- ing captions for audios in the wild,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.689733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.689733Z digest=sha256:e51ca4d3584dcb01446c39959ea63608cf4e3db85c10628dd377b1d273d7d751

Observation 1c71ab8d-13c3-487b-afad-0214601ac569 · outbound

This paper cites Audiotime: A temporally- aligned audio-text benchmark dataset,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiotime: A temporally- aligned audio-text benchmark dataset,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.790811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.790811Z digest=sha256:ebbf282c5212c4fdd7408c06fbf7f06c539243db0375d5b0eef4ea1f9687054a

Observation fc05b91d-98ce-4a16-b5c4-cc609a2bd884 · outbound

This paper cites ESC: Dataset for Environmental Sound Classifi- cation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models ESC: Dataset for Environmental Sound Classifi- cation,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.868153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.868153Z digest=sha256:ff43cc9fc2410c42c8f678c2e312c57651e3b3057a85dfe43a836cdd0f45878e

Pith citing papers

Observation 897b7fd3-7c2c-496e-abb6-b3bc3fcc7be8 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:55.932658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:55.932658Z digest=sha256:e4e449a4bb935a8fe495779c222c3df92fb6d9d31bb43a1b6eaa4f15f64ad1ef