Pith. sign in

Paper Citation Record · LEDGER

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2607.13408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13408 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:20:03.868153Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:19:55.932658Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 897b7fd3-7c2c-496e-abb6-b3bc3fcc7be8 · outbound

This paper cites Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:55.932658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:55.932658Z digest=sha256:3e4f4b4179e21b81244f328d46ed0d6661f784779f0b25d8fcb056984b82db59

Observation 93faeea3-4eac-431b-a8db-3976a2457343 · outbound

This paper cites an unresolved cited work.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.009130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.009130Z digest=sha256:cd877e23374c7f2ad6e36b20bc7d9f3754d9771d5ca237ed2bc11f994d15ed14

Observation c552ab95-18f0-40c9-bd4d-1a1f569710d2 · outbound

This paper cites Rain is falling con- tinuously.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Rain is falling con- tinuously

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.153317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.153317Z digest=sha256:bab59e5d4b90e5d72788f8d985b14e2443cc78f3fd5dbf4dbe7d3abddaca08a2

Observation 7bed7c70-e320-44be-81fc-a962e42fcab3 · outbound

This paper cites It starts with the sound ofe 1, shifts toe 2, and ends withe 3.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models It starts with the sound ofe 1, shifts toe 2, and ends withe 3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.317477Z digest=sha256:3ef2ea8393d2198560d76b8033e4139b0b1b19a36cf2cdc5bab426e6c50c7897

Observation 261244ee-e32a-471f-851a-887ae4373772 · outbound

This paper cites Verifying ALLMs as Judges Table 1 reports ALLM performance on audio understanding benchmarks with ground-truth annotations.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Verifying ALLMs as Judges Table 1 reports ALLM performance on audio understanding benchmarks with ground-truth annotations

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-08-02T05:19:56.421451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.421451Z digest=sha256:30d5c1f1b2040bcd2f8946ef77d7f5eb5425a3c07b6e3e33116c68cf60083b76

Observation 8d3e11d5-03f1-4d28-bbe9-0b3d50843993 · outbound

This paper cites While recent systems achieve strong percep- tual realism, our study shows that they often overlook fine- grained requirements such as sound event completeness and temporal ordering.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models While recent systems achieve strong percep- tual realism, our study shows that they often overlook fine- grained requirements such as sound event completeness and temporal ordering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.578105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.578105Z digest=sha256:63430e83f7fa001b1d387709d14c0abeffdf17d4629cfe4edc7d564e7732d7b5

Observation 73364c19-88fc-4a25-ab64-3abf510cc4e2 · outbound

This paper cites an unresolved cited work.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.685633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.685633Z digest=sha256:080555bd9ff305d0fbd9a320fc799bdf65aa22236ae40734bff23eb64c89ac1b

Observation 83e62017-cf58-492c-8725-d771dc9ac622 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioGen: Textually Guided Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.888125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.888125Z digest=sha256:e78194e78e598353ba5bf413e844096f861b02538d79f3dbe73d7edda6dd4cd5

Observation 3b8dd948-6f65-466b-a3cc-46a277751328 · outbound

This paper cites MusicLM: Generating Music From Text.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MusicLM: Generating Music From Text

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.911402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.911402Z digest=sha256:25fae6efb52ca245bfa5f66b93e41f904912c2c391f0fedc944e815a097ca75e

Observation cd851abf-2a24-4ff8-b554-cb7618658e91 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:56.963850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:56.963850Z digest=sha256:69f87f920c7901260b56d91baafe76261b6bfd1bc7e38627a8939e0942e6832f

Observation cd8dac01-cc08-4c03-b8de-40e5cddd78c1 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.089048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.089048Z digest=sha256:fbf058b75e38fddd2f5ebaa7d968bd9135a797ccff0f8541ff917654f65cd635

Observation 348e50f9-7e8a-4b55-b7c6-6fe9059b08ca · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound genera- tion,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Diffsound: Discrete diffusion model for text-to-sound genera- tion,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.219463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.219463Z digest=sha256:386a8ec6719f69babea99cda3b5e15fbe5c8b39b6ecd961cded3ae67aa07e6ce

Observation 5c55add6-096b-489f-90d4-2bc60330662f · outbound

This paper cites Towards general-purpose text-instruction-guided voice conversion,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Towards general-purpose text-instruction-guided voice conversion,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.281224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.281224Z digest=sha256:d418d5ab0aa3d71c3c17d02217ef164e1b501e2bafd7da72174794721776e7cb

Observation 7c05fc79-d7d6-4cb5-999f-5c2659632ce1 · outbound

This paper cites Simple and controllable music generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Simple and controllable music generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.382379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.382379Z digest=sha256:db2ad1382bd3e20bc67bdf47def35a065e1ad2c35ddbc6849186f97c465967e2

Observation 7ad275b0-2ef7-460e-b96a-95909d3f937f · outbound

This paper cites Audiolm: a language modeling approach to audio gener- ation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiolm: a language modeling approach to audio gener- ation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.450207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.450207Z digest=sha256:d5074ae8f3eefdee6fbcde7e8982720f63f4d690906260bc805088125132040e

Observation c3acb87d-c10d-4f9a-8ab0-17fa3de23416 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models SoundStorm: Efficient Parallel Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.547298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.547298Z digest=sha256:383dc0af80d07e0dfec9a0561faf6c34dda6692abe77f7deb4bfb2a368982c74

Observation 99729a8c-4d66-420f-867e-6f3bdbb85b88 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.626028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.626028Z digest=sha256:cb727f31947c6669c6c14953a9d21f9e8f5370eb188ca3716b3ac7c2507932ea

Observation f1542adc-c504-493e-b5cc-0d8a9df23b56 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.779448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.779448Z digest=sha256:75bfd0be1b11e0b14eff7790ea52a9d4f4a19c6bf2770135b6fa16d1caddad80

Observation af68f29d-04c2-4789-a505-a562ef0ef7b0 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap- ranked preference optimization,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap- ranked preference optimization,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.897476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.897476Z digest=sha256:d081cb99643b953f4cff8de227cb532c0b9c8ce93e0feb1d7df3de50d4eb1451

Observation c7753fe2-0ae1-45b0-a65f-29f274af05eb · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.993408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.993408Z digest=sha256:264c864023b0ea85ede8f2b66c073373583020da3a2e4350013035aae335d712

Observation 38c0a192-db20-4807-8219-d48127efaa21 · outbound

This paper cites Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.141475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.141475Z digest=sha256:b4aee8ae7776690c1dfec9796f16d64c1f3ac8132a73698cef94895de378acb9

Observation c69f1c51-9f5b-4103-901d-165ab4231797 · outbound

This paper cites Stable audio open,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Stable audio open,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.321242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.321242Z digest=sha256:79ccab8c6039ad51643b319703568db999ef4b6631034d0810abd09961d93bf8

Observation bf2e3ebf-2a9f-4891-995d-e67f26ee13e9 · outbound

This paper cites Etta: Elucidating the design space of text-to-audio models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Etta: Elucidating the design space of text-to-audio models,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.435218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.435218Z digest=sha256:22a7d7dd48e678b69947715f3978583303a129c970f4fc91d2f55b0acfb71f84

Observation 137f8c6c-2bdb-45d0-a2e6-47e2e1cae64a · outbound

This paper cites Impact: Iter- ative mask-based parallel decoding for text-to-audio generation with diffusion modeling,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Impact: Iter- ative mask-based parallel decoding for text-to-audio generation with diffusion modeling,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.585617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.585617Z digest=sha256:da8a5adefe3daa07b935b885fa0b12b6a1f1809be33bad08e14b9ca93c88d38d

Observation bf6b0368-4ac5-4de3-9de7-09775fd8cde0 · outbound

This paper cites Gen- erative audio language modeling with continuous-valued tokens and masked next-token prediction,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Gen- erative audio language modeling with continuous-valued tokens and masked next-token prediction,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.676033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.676033Z digest=sha256:9d23b5f5cfde505d2949c569ea56b7343e87d7159eb06b7fe1d4b8b776eacb9b

Observation faf29cef-d559-44e2-a290-673d1bece393 · outbound

This paper cites Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.829580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.829580Z digest=sha256:c07d9b7f11ae6567b41922db3d81c3d6a6c8637311477e0c901406a3881d8fb1

Observation f28e2bb0-5c2d-47e1-800d-a3f9a2f950b1 · outbound

This paper cites Improved techniques for training gans,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improved techniques for training gans,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:58.968651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:58.968651Z digest=sha256:4cb88e069a787a4c91565c81a6d48a33af614fac7e9698c9245c476188877911

Observation 68fafd2c-9de4-478b-8e25-965cce8f216a · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Clap learning audio concepts from natural language supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.107979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.107979Z digest=sha256:864fcea3f4855751afca15d7f54251c2ba930a98e8bc86fef6282ff61097bdc9

Observation 823b4a65-1815-4089-bea7-f33063c84884 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.212491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.212491Z digest=sha256:735edfcfa4373811c1abb45e6589c280f873b6e4da23a9aae4ab936ac1dec885

Observation cebd6c1b-4a32-480c-92df-1ae59215ff7e · outbound

This paper cites Ritta: Model- ing event relations in text-to-audio generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Ritta: Model- ing event relations in text-to-audio generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.335502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.335502Z digest=sha256:a33974919d89abfa15053a339746b932d5d2c126c93db00aaf7bb43609c9113c

Observation e1032269-da52-42aa-bf86-def53b563435 · outbound

This paper cites Aurelius: Relation aware text-to-audio generation at scale,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aurelius: Relation aware text-to-audio generation at scale,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.441087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.441087Z digest=sha256:e4e68b90e45fc6e617efc1532e55484071d1a1f94c25fc469c56bcf4f14dc9a0

Observation 902de066-9176-4336-8df2-c8a8c476b0ca · outbound

This paper cites Compa: Address- ing the gap in compositional reasoning in audio-language mod- els,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Compa: Address- ing the gap in compositional reasoning in audio-language mod- els,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.620138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.620138Z digest=sha256:e6c1c1cf63c00b7d6ac5968c4d9a7600c7197d0a7a7212ba9a7f7b622c0eb655

Observation f9f1773e-935e-4788-a6ef-ea24d36f11a5 · outbound

This paper cites T2a-feedback: Improving basic capabilities of text-to-audio generation via fine-grained ai feed- back,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models T2a-feedback: Improving basic capabilities of text-to-audio generation via fine-grained ai feed- back,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.789891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.789891Z digest=sha256:2f6150a3e192351e9b445a22c17868a820f8d65e5319eaedaf7e332887b71519

Observation 90a1d0a4-25db-4132-a3a1-62995b35a5b4 · outbound

This paper cites Aqascore: Evaluating semantic alignment in text-to-audio generation via audio question answering,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aqascore: Evaluating semantic alignment in text-to-audio generation via audio question answering,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:59.939522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:59.939522Z digest=sha256:75fb3875b1867c8b25b1956f3ec26ab9eede7286fe003a674448f72b9582c30d

Observation c044ba76-7e09-45c9-997c-dfaf3b24b02c · outbound

This paper cites The Sound of Absence: Audio-Language Embedding Models Struggle with Negation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.076142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.076142Z digest=sha256:4e09ad6c5af76f144b6d65c07504193ecdcf215a853baeb8803a6da887c8cba2

Observation ea01226f-cdac-4df2-a716-7a4c648b1874 · outbound

This paper cites GPT-4 Technical Report.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.190570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.190570Z digest=sha256:c15f1b31db6edf4cce7dbcfb7dcc01221b024243480ee9a98771bb05580c71ce

Observation 858fae87-adb7-410e-9d45-36588d562764 · outbound

This paper cites Joint audio and speech understanding,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Joint audio and speech understanding,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.385579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.385579Z digest=sha256:85210b6f2d683c472490453e2ab726a8d3ccb4566d97193ab634bd82968c1a4e

Observation 0f35bbe4-350f-4beb-bfd6-a448f097bb72 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.587754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.587754Z digest=sha256:9d8ee51e0317d2cdf77f8a94693c1ab7d7a8053d25bd41b1866a07a6f7e5ea4c

Observation eeb694ec-2354-4fac-a910-630a9669a28b · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.724635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.724635Z digest=sha256:c46b36c820c2eb564e4369d931dfe16fec8ab79b680c773d040dff27c201e375

Observation 404f282e-2fdc-4d66-a562-0a1f6a605364 · outbound

This paper cites Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.835510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.835510Z digest=sha256:e3d5822264ddd717d9225cf9b6c8165155aaebd00b8dc69cc3611238ad8c9df5

Observation b6d54ccc-0e7d-4188-864e-5c65a1e75323 · outbound

This paper cites Speechprompt: Prompting speech language models for speech processing tasks,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Speechprompt: Prompting speech language models for speech processing tasks,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:00.957127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:00.957127Z digest=sha256:8e9d28b64376f584535e01d2fe63b32c936a2d28b061a43a8de0686801d81db5

Observation 2b6a02bc-8860-4fbe-b2d4-40658e22f1b0 · outbound

This paper cites Blsp- emo: Towards empathetic large speech-language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Blsp- emo: Towards empathetic large speech-language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.082235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.082235Z digest=sha256:da3d63dad9deb385e60bc173204b0e43358583a34dd6d788bebe2e7ab96e92cb

Observation 63214a6e-6535-42e2-aec1-579dd8c11ab5 · outbound

This paper cites Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.281092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.281092Z digest=sha256:d50e48505bf1b03e6594a93e83a59727fb49b43dfa2814dd31122f92189b0a56

Observation 608bbfa2-db29-48f0-b7a6-465c09ffa510 · outbound

This paper cites Voxtral.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Voxtral

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.370603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.370603Z digest=sha256:044e27e2a54a32bad2802d284512fd6cd2eb68d8b7eb40420a3bc25422d20259

Observation 24cfd24c-19b1-49d1-98a5-cd9b15b8f084 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Qwen2.5-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.448742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.448742Z digest=sha256:68d44328f9593befcbdf33838f9c2f5445dee78b01fb0567cd339dd458c10478

Observation 643cea81-7953-4c92-b12c-91f3407746b5 · outbound

This paper cites Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Teaching audio-aware large language models what does not hear: Mitigating hallucinations through synthesized negative samples,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.542825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.542825Z digest=sha256:037a92913044db0d9fe1bd1581eda25ed6feb1206c226bc3becbcde2a26d24b7

Observation 4db4dd76-378d-41cc-9a06-1be24eb3645d · outbound

This paper cites From alignment to advancement: Bootstrapping audio- language alignment with synthetic data,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models From alignment to advancement: Bootstrapping audio- language alignment with synthetic data,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.637467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.637467Z digest=sha256:4677e31c3b1116eec475e6fff020706347ac9d87db38af7f24987cabd80d2d67

Observation c0c038e6-b470-4114-998a-f34b9a784af4 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.741673Z digest=sha256:1d1a29a8d49680998a9bbacad3e86b8ae8016d14c29035d8036ac05ee134ef73

Observation 51859531-fef3-4eb6-a497-416d668ceead · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.846967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.846967Z digest=sha256:0eb2723baf5cc7bef0af2a88c07afc05c34e01d6df8c5dc98c11d55c753d85b2

Observation 82966f7a-01b6-4c22-b376-bebc7cb45090 · outbound

This paper cites On the landscape of spoken language models: A comprehensive survey,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models On the landscape of spoken language models: A comprehensive survey,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:01.913145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:01.913145Z digest=sha256:8d2c5e64004bb4c4a2e9cd9e79a1a4638a1b59d8bd1c36b604082c4eb3e5f943

Observation 1c2430b7-ebd3-4959-89b8-d6be49899dd5 · outbound

This paper cites Audio flamingo 2: An audio- language model with long-audio understanding and expert reason- ing abilities,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio flamingo 2: An audio- language model with long-audio understanding and expert reason- ing abilities,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.010846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.010846Z digest=sha256:7c9a524a01db71c67453a15b131cb3aac2b882c05b6247bb7afa4be334985ac9

Observation 437b6dca-2413-4b2e-a766-918f3ce965c9 · outbound

This paper cites Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.115586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.115586Z digest=sha256:727bd4a940eb04c4827756aca212f6005067ae330590b44a0aacefe13e89dbf2

Observation d21e50ec-7b39-44ad-bc3d-c16394ff7da3 · outbound

This paper cites Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Can large audio-language models truly hear? tackling hallucinations with multi-task assessment and stepwise audio reasoning,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.190850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.190850Z digest=sha256:b37e9b8477fa8f6d72d763ea4bfd6e7c29a9c08c1789bc6dfd21d3cf5d354653

Observation 22d14206-0d01-4eab-9f4a-3efedc4a53f9 · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.256101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.256101Z digest=sha256:1285e21017eee09cd979287f0ea32e4fa3b5e7206cfd6619b67434ccec9d4855

Observation 8f6b06dc-021c-424d-b487-6393d494879e · outbound

This paper cites Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.393294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.393294Z digest=sha256:41e75f939361b0a7816c25e3d7d4549e22daef60559a29d0dfe45b9d9d5428cd

Observation bc7f16a0-4878-48a4-95f3-1e4acbb3d1cb · outbound

This paper cites Mmau: A mas- sive multi-task audio understanding and reasoning benchmark,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Mmau: A mas- sive multi-task audio understanding and reasoning benchmark,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.526891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.526891Z digest=sha256:e734a566361842a52bf29d532c31fdb543c7fb069917c5fdbe0a9d39c8fe7a86

Observation 6725413c-ccf2-44a3-92cf-440a8c763c73 · outbound

This paper cites MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.638568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.638568Z digest=sha256:d3a0b084740414b19b33563fa0b784fe9c692a69b084646c1f19e2f705ef13af

Observation 845f6050-4f64-412b-ab73-8ac9afaa3b9a · outbound

This paper cites MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.704210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.704210Z digest=sha256:ae4b1e087da14dcc1618ed68256c1354068f86dc69c73c892a0e465f3d2f42b7

Observation bbe15c22-57ad-49e3-922d-0172e2fe2e30 · outbound

This paper cites Game-time: Evaluating temporal dynamics in spoken language models,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Game-time: Evaluating temporal dynamics in spoken language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.807832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.807832Z digest=sha256:81eca2e6e40e00605da7234d14acc06f6bd541f525ec998ae61ac78f1a68c110

Observation 96f2e11a-f2dc-4ef6-a2ca-3096fb59c02c · outbound

This paper cites Aqua-bench: Beyond finding answers to knowing when there are none in audio question answering,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Aqua-bench: Beyond finding answers to knowing when there are none in audio question answering,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:02.895515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:02.895515Z digest=sha256:57f33cd2a5be172d741253713ac2b7ff457e6f4b25cb6d421b0d1bea0a1ad3c1

Observation 4a70f1b7-c7c8-4ec6-ad7f-eb66b453401b · outbound

This paper cites Baton: aligning text-to- audio model using human preference feedback,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Baton: aligning text-to- audio model using human preference feedback,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.015060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.015060Z digest=sha256:1e3dc3418b3982ccebd398e6cfeaaf9ab50ebadee7a53deb3a118e5eebc91c57

Observation 11e27ef6-1daa-4dc8-89f9-9050d62389a4 · outbound

This paper cites Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.080402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.080402Z digest=sha256:70cb47c3e186d6750171c2131c35809053b34c71e5bbe076f96993a991b348b5

Observation bf460991-dbb7-491d-aefe-0c7c596b4ab9 · outbound

This paper cites Audio large language models can be descrip- tive speech quality evaluators,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio large language models can be descrip- tive speech quality evaluators,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.189659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.189659Z digest=sha256:7759887dec1a0c3fc056a8433ddaf1c82fb9ec76b49387d044dff16d5c34f609

Observation b988b51f-4411-49b7-b030-50c369e351cb · outbound

This paper cites AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.280286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.280286Z digest=sha256:fbb83844732be72883bed3c232dd5197e2b4faa41d7a8541f55cfaedb236ff6c

Observation fcfb353f-9297-47b6-a57a-556b13f36469 · outbound

This paper cites Audio-aware large language models as judges for speaking styles,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audio-aware large language models as judges for speaking styles,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.382877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.382877Z digest=sha256:03fbe6157253c8b77d6f2915eacfd272e6c94675c2967cb993c1713e4d319c14

Observation 2cc58875-3fa4-4603-9419-b2cabe24bced · outbound

This paper cites InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.452297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.452297Z digest=sha256:4a9e0b42701b730ec06d871cf09934f66f366811f9422b6c224a8a228e90c85f

Observation ddab289b-5042-4030-a525-02d0dc13abbd · outbound

This paper cites Audioeval: Automatic dual-perspective and multi-dimensional evaluation of text-to-audio-generation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audioeval: Automatic dual-perspective and multi-dimensional evaluation of text-to-audio-generation,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.583905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.583905Z digest=sha256:77514b62a92fcb2f72e2db1aa43b8edd35ca2774ce040665e5529c21f6b14110

Observation 0a5ecb1b-92f2-4fb1-8830-1452c7d0a9d3 · outbound

This paper cites Audiocaps: Generat- ing captions for audios in the wild,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiocaps: Generat- ing captions for audios in the wild,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.689733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.689733Z digest=sha256:e29c8f2732a9ae80109616e7ed6a9c8e42a2d760fece0b86a5c6acd95aa975ac

Observation 1c71ab8d-13c3-487b-afad-0214601ac569 · outbound

This paper cites Audiotime: A temporally- aligned audio-text benchmark dataset,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Audiotime: A temporally- aligned audio-text benchmark dataset,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.790811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.790811Z digest=sha256:985a4f643568e9db7b1a6e04a7e2cf26da5217e50be8aeec3df748fc75c3a901

Observation fc05b91d-98ce-4a16-b5c4-cc609a2bd884 · outbound

This paper cites ESC: Dataset for Environmental Sound Classifi- cation,.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models ESC: Dataset for Environmental Sound Classifi- cation,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T05:20:03.868153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:20:03.868153Z digest=sha256:c10ecc2da9457bbffd58c468c56f71c0bb49a36755a4cb1ecf42a11a63bbd41a

Pith citing papers

Observation 897b7fd3-7c2c-496e-abb6-b3bc3fcc7be8 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:55.932658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:55.932658Z digest=sha256:3e4f4b4179e21b81244f328d46ed0d6661f784779f0b25d8fcb056984b82db59