Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:18.146028Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2411.12641.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:18.146028Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 990d35bb-ab15-4a86-b710-5824cc4d88dc · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34cade4-29f9-4155-a4e3-af00fe1ce649 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Look, listen, and learn more: Design choices for deep audio embeddings
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c1dbbc6d-6d7c-4958-b17f-ec351a2cf76e · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e601f507-1b25-4895-bf77-6822b18f1367 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LP-MusicCaps: LLM-Based Pseudo Music Captioning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c8d089-5c58-4df5-955c-7d032858bf5d · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SingSong: Generating musical accompaniments from singing
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b17c951-e43c-4d6e-ba22-54cbf1d17685 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint music and language attention models for zero-shot music tagging
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c37ee982-b157-47eb-8672-99a29fa3df47 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Long-form music generation with latent diffusion
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 214f2509-3365-4596-b4c2-900f898a8b6c · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models VampNet: Music Generation via Masked Acoustic Token Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506b6021-d676-4bda-b7e1-a3572e26369f · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio set: An ontology and human-labeled dataset for audio events
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40bc01b0-0948-4894-8ce8-254c2ef195ae · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eca96272-26cb-4398-8d2b-0d765900055d · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de35cce0-8fe5-440a-9b0a-fdbdb3d63b6b · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524d5f25-d188-48a5-8ee9-4bfb8847f6c7 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5579418-771a-4633-a14e-002d6cbfc947 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b0111806-a9f6-4f14-af0e-a808c2b3e499 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Retrieval Augmented Generation of Symbolic Music with LLMs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7959ab72-2e7a-45c1-94b1-1a48bcaa3460 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb143b2-f6e4-474b-86cc-4b67d6bef4cc · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Auto-Encoding Variational Bayes
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c664983b-0f43-43a8-8a62-8aaabf079cfa · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f48f30-1d70-4f13-99b9-16df11f82b7d · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Efficient Training of Audio Transformers with Patchout
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13e77da-de29-4894-a1f3-33bcdf016491 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7e1e4f-e84d-4d2f-83ce-0fc66c659fb9 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Content-based Controls For Music Large Language Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9d8692-1128-4c04-a4e4-833456fff121 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8c9f1b-e2b0-4c2d-9d5a-ed0418f56f1d · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f74cdce-4863-49dc-8d77-68473f76c2d0 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54e9a4e-c430-4d5f-bba0-18693ff386d3 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Novice-AI music co-creation via AI-steering tools for deep generative models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3ae3d660-10ae-42dd-be2d-40be5d851534 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuseCoco: Generating Symbolic Music from Text
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45537720-6e91-4814-9d5e-00371d492646 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f262fcf-3f34-4c01-86fd-2ff4870db5d6 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bec592a-7b1c-4b21-ad97-ddd546f84ca9 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2c73553-2e83-4041-9037-efc495714b3e · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2019.8937170
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7d0675-2c8f-4c2f-9d98-1e41803c0269 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Stem- gen: A music generation model that listens
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94ce5295-f9f1-49ab-9bfe-e085ba0f4556 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Zero-shot image-to-image translation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e186346-0f81-47ff-b771-832e2a9a2092 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Glove: Global vectors for word representation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6505c8e3-76a5-42cf-97d2-7044012f2901 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Deep contextualized word representations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da217417-315c-4b64-89c0-cda45a3910bb · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generalized multi-source inference for text conditioned music diffusion models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 72792de0-2abd-4dfd-ba04-aea455bc04d8 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Paguri: a user experience study of creative interaction with text-to-music models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8511da5-b9d9-409a-846c-139d67be53a5 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio Conditioning for Music Generation via Discrete Bottleneck Features
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3572c32-6a17-4956-85e2-3fb128887240 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b6032a99-19c9-43ad-a077-09b2585925e6 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models URL https://doi.org/10.1109/ICASSP.2019.8683855
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f28cfb2-25c1-4eb9-8f49-b09ede1b1942 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee62ef0-c927-4586-b0fa-f3c5a1a85f37 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae40fbd-b2da-456a-8d2d-f086b3d3373c · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA: Open and Efficient Foundation Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36568fc4-9e0d-42c1-8a3e-651dd21d23c3 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Plug-and-play diffusion features for text-driven image-to-image translation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5f8e728e-061b-4156-9bd6-00b3c5458970 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models WaveNet: A Generative Model for Raw Audio
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e141c705-cb91-474c-a455-2ec8c01301d8 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa745de-7c50-45fe-b3b1-1506c629e7b0 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b21ed84-4989-4f8d-8f77-4858ca1e080e · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deaafc50-5a0a-4bd1-8afd-eccaff0c3403 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Iteratta: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2703ed64-5d4f-4d4a-a51c-7e9cbb6c0e02 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d3a452-f80f-4be2-879c-54b11508117a · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12264b63-e1f6-41f7-bfb4-7b46197d5f43 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7aaae18c-fb86-4487-a150-aefcaa3649a6 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models ChatMusician: Understanding and Generating Music Intrinsically with LLM
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dedc977-28b4-44cd-a5de-2858df3024b6 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d503132-40d5-45a0-8a85-75d23867d85b · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cosmic: A conversational interface for human-ai music co-creation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 28009de3-3aab-4a83-b2f9-07d17531f7d8 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c418d10-b949-4159-add6-a74d84190332 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Survey of Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00db70a-dfd9-42df-9df6-e87647de670c · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Masked Audio Generation using a Single Non-Autoregressive Transformer
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f06bdcae-e553-4106-812e-acbba9ecca99 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models
Reference 1986
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57bb2bc-47fd-4c40-afeb-66e9746f6412 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring XAI for the Arts: Explaining Latent Space in Generative Music
Reference 1996
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bbe475b-5897-4987-86d2-2ab251e89d31 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2003.819861
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370d21ad-03b8-494f-a600-fbf4f2ef17dc · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Improving Text-To-Audio Models with Synthetic Captions
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d803a4f9-532f-4a98-844a-4347d9fe9899 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MoisesDB: A dataset for source separation beyond 4-stems
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b6bd2d2e-5f02-4445-ac86-2279b1482464 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MidiCaps: A large-scale MIDI dataset with text captions
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e7ac71-eccd-4b9e-a718-16e7989b14e1 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5cb6ad-9aa4-430d-ae8a-95a11b661c7a · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Au- diocaps: Generating captions for audios in the wild
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0035bc81-1603-4cf1-8b24-b65ea1e86213 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuLan: A Joint Embedding of Music Audio and Natural Language
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca295f9c-16c9-4a67-9d7e-c3ddf8763784 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a68021-b2fe-4f82-9568-cfe8712a4d44 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Jukebox: A Generative Model for Music
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5decd465-1463-4b20-b74a-8ff02ef74acb · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Musicldm: Enhancing novelty in text- to-music generation using beat-synchronous mixup strategies
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56994326-27b9-4dee-bca5-b81e86bbbaf9 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a50e756-deba-4490-bf58-6fdf0d6ea5b5 · outbound
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.