Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:13.413269Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2501.14680.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:59:13.413269Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 566a6805-e497-41c4-86bf-0f81baf37c77 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Denoising diffusion probabilistic models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1aee23-d93d-4c61-b40e-210b61353d5a · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Attention is all you need,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966344f7-3780-459f-b958-da3755165e1f · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning High-resolution image synthesis with latent diffusion models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ffb23d7d-3ea3-4714-b2d0-77e5973346b2 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM: Text-to-audio generation with latent diffusion models,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1ea920f5-f649-4d73-bd31-7294ce13535f · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f61aed68-8741-4be0-bd98-592b7aa181bc · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Fast Timing-Conditioned Latent Audio Diffusion
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47cf95b6-e3a7-4277-abf1-539d7ff7656c · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Grad- TTS: A diffusion probabilistic model for text-to-speech,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 283bd3a9-de8c-4763-b624-c8fe25a07278 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Learning transferable visual models from natural language supervision,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e8729c8-8429-4fc4-99da-92bfcde07589 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3160c0eb-c05e-459a-b0fa-8ec633876133 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MuLan: A joint embedding of music audio and natural language,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 902b2cc3-17b9-4d6d-9cde-b8918855bf36 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0776d894-663a-481a-8e79-878116b01567 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning BERT: Pre-training of deep bidirectional transformers for language understanding,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5aad878-1222-40dc-8dce-9ece449697d7 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MusicLM: Generating Music From Text
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929e3b7e-73e3-4564-82bb-5f0724e54828 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Audio-Text models do not yet leverage natural language,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f197e251-1377-44db-9ba4-6107f6aa9036 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Simple and controllable music generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fcc17df-b228-41c0-95d1-6a3e309169d7 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd9d5b1-9e68-4c48-9885-16227857af87 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f63d4d4-9f04-49f0-a456-e3488eb8d161 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35c20cb4-a0b8-49c4-906f-ee4c6615b9af · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Scaling instruction-finetuned language models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb11ae3a-d7c1-49ff-8c19-59fd83b599cf · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 74a009b1-1e61-42d2-be28-81a8aa34ae0e · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b914bd54-6e1c-478d-b1ce-7810777e15b3 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f89a955-997d-46cd-bcd9-db00163ef125 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Whitening Sentence Representations for Better Semantics and Faster Retrieval
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae17606-0a6b-4b87-96a7-2d04485d714e · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning On the sentence embeddings from pre-trained language models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a221d86c-6487-4f1a-832e-9caeea2f7ffa · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning SimCSE: Simple contrastive learning of sentence embeddings,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 860d36a4-6d5d-4059-815f-5ddf185f3a97 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning FiLM: Visual reasoning with a general conditioning layer,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 758ceddd-7816-45e9-81a8-96cb144ecea3 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Self-attention encoding and pooling for speaker recognition,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3a3d6c11-7b22-4857-afc5-00bc06521ecc · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Progressive Distillation for Fast Sampling of Diffusion Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11db5f13-3cbd-4732-b402-6d7f233250eb · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Variational diffusion models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9af5f36e-2048-4843-9127-d4deb2529023 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Classifier-Free Diffusion Guidance
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d21dee-0e18-4b06-a7ba-4f6e64ec2b5c · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning The MTG-Jamendo dataset for automatic music tagging,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a6eb880-7b8b-4f2d-8ef6-148e5ad6c50c · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning FMA: A dataset for music analysis,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d3f9f4c4-3504-4c0a-93fa-6e0bed2dce2f · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e4b77eb-f779-4c16-bbab-3acac67ab291 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Hybrid transformers for music source separation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae4fa1a8-10e2-45e4-818b-2323d5a6c4e8 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Mustango: Toward controllable text-to-music generation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02fd6909-8ce7-4ff8-9d79-a7315916f4e4 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning AudioLDM training, finetuning, inference and evaluation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0b17c31f-fd4c-4ddc-abc1-1d2f4979e6de · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning LAION-AI/CLAP
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2dc6ea61-dc3e-47dd-9904-901eb5c6d0d8 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Decoupled weight decay regularization,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f46d9230-614a-4a67-887b-75909a96cc60 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Denoising diffusion implicit models,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255e0e09-02aa-4f52-b371-69376eddd4de · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 16834dba-2422-4e04-a209-04fab56c8001 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Stable Audio Open
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2deda01-6e87-4709-9ca8-374ec85eb6cc · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Fr ´echet audio distance: A reference-free metric for evaluating music enhancement algorithms,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce8ed845-9eb0-4e2a-a345-f59323737667 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning Audio generation evaluation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 727c1c7e-0988-4d59-8e07-eedfc6b08e08 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning CNN architectures for large-scale audio classification,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95c438e2-aca3-4bc5-aecc-dbf2e809f8c4 · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e4f311e3-c33e-4515-bc6c-705ae7dd259a · outbound
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning CoLLAT: On adding fine-grained audio understanding to language models using token-level locked-language tuning,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.