Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2304.13731.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:27:20.859159Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
15
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 69860cd6-e039-4abb-818e-5f10f765f5be · inbound
Generative Semantic Communication: Diffusion Models Beyond Bit Recovery Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1c1ce906-c66e-457a-9914-2eec6b7bf2f4 · inbound
Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 707a0d2c-634c-4acf-a713-b365b0ec29d6 · inbound
Movie Gen: A Cast of Media Foundation Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca53461d-b78d-4ea0-8b5e-812a18c0efc1 · inbound
MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060332f0-dc0d-4512-9549-4f8f5ab1e690 · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · inbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea5d44d-f34d-4927-b625-ef939d97f6f8 · inbound
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee7105a-96f9-46eb-bcba-338e0d6ac279 · inbound
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa2e99e3-c5b0-481b-8d65-5d64ab32c365 · inbound
Neural Vocoders as Speech Enhancers Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d37e2e-d5cb-4911-88fd-5825de27fe83 · inbound
Overview of the Amphion Toolkit (v0.2) Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 652e3360-2970-470c-80d9-3d11db5d2752 · inbound
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 285
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb759659-6ea3-46ce-a086-8709c5d47140 · inbound
PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca845220-fa6c-49bd-a730-253acc8ed52f · inbound
PerPO: Perceptual Preference Optimization via Discriminative Rewarding Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537e01b1-1e2e-4d28-8e31-4aaebdd6d25d · inbound
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac16a45-5bd5-4f75-a9ae-0834e90d4fc4 · inbound
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67958216-8203-4b5c-bd0c-bd450947469c · inbound
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039e2613-8da5-4e80-a2dd-bbb982f6bfe7 · inbound
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc0151b-43c8-4f68-a72c-50d218cdacc8 · inbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90dc13ef-de70-403c-83d4-b6dcc99ece2f · inbound
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d4c176-df34-4284-a3c6-df2bc1bf91ef · inbound
Ego-centric Predictive Model Conditioned on Hand Trajectories Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73a878b-9120-4d7c-b9d8-0d069b9b32f4 · inbound
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb26b8d8-ce1c-4a41-a413-4f98fe9b0774 · inbound
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8205c891-3462-4416-a8a8-d7edb75a6565 · inbound
AudioMoG: Guiding Audio Generation with Mixture-of-Guidance Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6dc1f721-c254-40e8-bbb1-550d02063719 · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe1b4c2-5e26-453d-bb72-c04d068231c1 · inbound
Omni2Sound: Towards Unified Video-Text-to-Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cc2f9143-e0da-45f8-90c0-5b8af801aee5 · inbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 08905fd2-4446-4e2b-ac4a-c5d44b33ca3a · inbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b57c942-0463-4f7b-86f1-960219ace007 · inbound
Auditing Training Data in Generative Music Models via Black-Box Membership Inference Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3fc835a8-7169-4cb2-8a84-a653542dc1b0 · inbound
AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd8dac01-cc08-4c03-b8de-40e5cddd78c1 · inbound
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a07c5b1f-b16c-444f-af20-39dc2fe7e313 · inbound
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105e8c7f-b948-40ec-9d5a-68954b12baaa · inbound
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.