Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:19.924941Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2412.19351.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:19.924941Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:29:47.889339Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T18:29:48.781610Z
81 of 81 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 38a0f702-0c4e-4914-b2e9-c39311c56b32 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e949966-c858-4c17-bdc1-99c277bce3f6 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models MusicLM: Generating Music From Text
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf6b427-dfef-4d06-bac4-0ccc9c0c2280 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Improving image generation with better captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79dc3aeb-1cd7-4c70-99b8-612df2b5bbb8 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audiolm: a language modeling approach to audio generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b26f449e-3485-4050-89b5-d163165cd51c · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Vggsound: A large-scale audio-visual dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6494ed4b-99c6-42fb-9346-7567e34056d1 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Scaling instruction-finetuned language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81419a77-4f44-471c-a957-a01daf729252 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Simple and controllable music generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5318ae30-fa95-4658-8ecb-75886942d59d · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Look, listen, and learn more: Design choices for deep audio embeddings
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 176eb31a-d2b6-43b9-8be9-8694f0afd531 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31aeb629-08cb-44a1-b4f6-94cf929500b3 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models High fidelity neural audio compression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea64442-a818-405c-a818-802cd5450134 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Diffusion models beat gans on image synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fe67b9-e09e-4f8e-bbc3-5ba7e3f01ba1 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Natural Language Supervision for General-Purpose Audio Representations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2388ac5-ad3e-4afe-92c1-8884cc9c6897 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Scaling rectified flow transformers for high-resolution image synthesis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68c8769-783a-4b13-888c-6a9d48d3cb7d · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Fast timing-conditioned latent audio diffusion
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b4cf4c1b-afda-4ac7-bf37-dbac0f24533b · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Long-form music generation with latent diffusion
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bdeabe8-3bd5-4d4d-8216-902399c5c10b · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Stable Audio Open
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b5a11b-d485-4276-bed9-a1fd84d7200e · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models FLUX that Plays Music
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba8e9df-13c1-4f38-820a-23d50bc64b5e · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audio set: An ontology and human-labeled dataset for audio events
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febe2422-9f4d-443b-8096-6e1f6397f52f · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Text-to-audio generation using instruction guided latent diffusion model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 19c8c6b2-6272-46a0-9a48-d7cf7a87aa43 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252aef73-a3a0-4225-925a-957d08ea0c46 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Taming Data and Transformers for Audio Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3b283e-dbf0-4bbb-bab5-91c692844f7c · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient diffusion training via min-snr weighting strategy
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ccebd0fd-8948-4f10-be94-e676e7ef8baf · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Gaussian Error Linear Units (GELUs)
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f71e10-a9c7-467a-95ff-7ac8edaac624 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Classifier-Free Diffusion Guidance
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9cc778-8f86-43dc-bcdd-3e76505e4626 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Denoising diffusion probabilistic models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce53328-e6b5-416d-8390-c24d099da666 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae4b6d1-0c62-45de-b91b-a1aa2d975d35 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Noise2Music: Text-conditioned Music Generation with Diffusion Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1d2ba8-4f67-4685-a5c4-a13af216205c · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e6614b3-f58f-4196-894f-eba9dae087ce · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Elucidating the design space of diffusion-based generative models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0e4154-1a28-4081-ba3b-4e396ee4a461 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Guiding a Diffusion Model with a Bad Version of Itself
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41df97c-8c19-483b-b8f6-bf7d50b0b82f · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9893b9-a990-452b-b519-6d748d3122ef · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audiocaps: Generating captions for audios in the wild
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c628301e-5e5f-40fe-b64d-e867819ac69a · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e596195-28a8-4d5c-a951-c396965664fb · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Variational diffusion models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751e8003-63d4-44d2-b6d6-8da49c09525a · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Kingma and Max Welling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b717ad-6b0c-4f07-b1c9-d715030bd0c4 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f4cc80e-6c47-4dd4-b606-13ce4387c1ed · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Diffwave: A versatile diffusion model for audio synthesis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06599a77-24f1-49a4-8e98-00979612c6cb · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b58e351c-68d9-4742-a8bd-d02e5e6ce317 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Improving Text-To-Audio Models with Synthetic Captions
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c75c390-6775-4c6a-aa0b-9ae16fe73eed · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient training of audio transformers with patchout
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40aa2056-1b35-4ee2-bf0d-1365d5d79202 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models AudioGen: Textually Guided Audio Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15833d4c-6bf7-41c5-a062-7be1a6f70a18 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434b7642-ce13-487e-9721-c7a5fbee995b · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient neural music generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d3f9a6ff-0611-4c05-bf12-fc1470453737 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ce5c63-f03b-446b-9ab1-a5b6fca078ee · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Bigvgan: A universal neural vocoder with large-scale training
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e89695d-468a-438f-ac72-036c0e3b04b3 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Quality-aware Masked Diffusion Transformer for Enhanced Music Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0309a060-3e82-41b2-b31b-7b2bd48de87e · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Jen-1: Text-guided universal music generation with omnidirectional diffusion models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2631d82b-9bc2-4180-ad83-bca00b1c4835 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Flow Matching for Generative Modeling
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2046cb-c949-4ddb-b41b-20cec4608d34 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Generative Pre-training for Speech with Flow Matching
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a28a237-5297-4e46-b9a9-8cea9619ae13 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audioldm: Text-to-audio generation with latent diffusion models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45117569-e7f0-4991-bf7d-81a43543f2ee · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f419579-be39-41dd-aa00-00ed85b49656 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Decoupled Weight Decay Regularization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d787dc-de0b-4ed0-b48f-810684e6a027 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98e2b40-e347-4d99-a31a-ac6cb5924f74 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models The song describer dataset: a corpus of audio captions for music-and-language evaluation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0ddf404a-7328-4679-9b8f-bcc20eacfdc9 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc78afad-bff9-43ca-be72-4027c2db333b · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Mustango: Toward Controllable Text-to-Music Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9baac49-05ac-4991-b9ee-b7bbb970e9a0 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Mixed Precision Training
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877599ef-0ac9-4456-b80b-34f271828853 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Improving multimodal datasets with image captioning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 15d34930-2554-46a9-b17d-d1ef2d63eac5 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Gpt-4o: A powerful multimodal language model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4241cba8-4b8e-4b62-86f5-7adaaf966ad7 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Scalable diffusion models with transformers
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f5c2ea2-defe-4769-a2ad-5c8fb70ec2be · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Language models are unsupervised multitask learners
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f88baeb1-f7a2-4f7f-baf7-a6b3bdc47293 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb6bce3-9d89-404a-97da-5f7c813628e6 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models The musdb18 corpus for music separation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 28c4716d-3ec0-4b15-b5aa-a296b03921c1 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models High-resolution image synthesis with latent diffusion models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f96203-c3cf-455f-b6a5-cdc7bcb28adc · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Progressive Distillation for Fast Sampling of Diffusion Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b379d49e-be66-494c-a4a4-9a68ceeb91e4 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Mo \^u sai: Efficient text-to-music diffusion models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3010030e-b093-4861-8e03-ce84cb7aea29 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Score-based generative modeling through stochastic differential equations
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846d28ab-e886-49cc-bb39-4d6a2735d0ed · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models auraloss: Audio focused loss functions in pytorch
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0160a295-6494-4ec1-8bb8-2b88f4d68bb1 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Automatic multitrack mixing with a differentiable mixing console of neural audio effects
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7a763d9b-71cb-4685-a43b-a7803e2a0743 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Roformer: Enhanced transformer with rotary position embedding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ded199f-618f-4ce6-b8dd-9ee467885d8a · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Improving and generalizing flow-based generative models with minibatch optimal transport
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28daef5-c711-4649-a51b-a06146b427f9 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Attention is all you need
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c87f90-3104-4ca3-aaaf-cd187fc35c2a · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b1ceba9-f1e3-476b-9189-44c649c5b49d · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models CogVLM: Visual Expert for Pretrained Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace7a19c-5d75-47b2-9b37-e27c2a52372a · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375f9341-8308-4fab-b4da-e3e22a06ffb4 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81f82fa-c6b0-4067-9217-595cf16890a9 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d9a63d-6bc0-4c65-b402-f1a61547b6ae · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38af1539-d2ab-48a6-a29a-3a4df938e4e4 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models @esa (Ref
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe2ca9c-508c-489f-99db-a579f9037f64 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b30830-190f-4d5d-8ae1-efca20564ce1 · outbound
ETTA: Elucidating the Design Space of Text-to-Audio Models LP-MusicCaps: LLM-Based Pseudo Music Captioning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7febed-d781-4472-864a-180d04a118a1 · inbound
A2SB: Audio-to-Audio Schrodinger Bridges ETTA: Elucidating the Design Space of Text-to-Audio Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.