Pith. sign in

Paper Citation Record · LEDGER

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2411.12641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12641 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:18.146028Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved49
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 990d35bb-ab15-4a86-b710-5824cc4d88dc · outbound

This paper cites MusicLM: Generating Music From Text.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.824039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.824039Z digest=sha256:bb71bfe689d75282d65bd527d9148bfd86b80fcf89b34dec2cc12cd6da19ad7a

Observation d34cade4-29f9-4155-a4e3-af00fe1ce649 · outbound

This paper cites Look, listen, and learn more: Design choices for deep audio embeddings.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Look, listen, and learn more: Design choices for deep audio embeddings

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.319985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.846659Z digest=sha256:0ecdecf11509c0dc3e2bf656611ce65dd97357a79a4498a1afd109a3f6ff9ea8

Observation c1dbbc6d-6d7c-4958-b17f-ec351a2cf76e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.850829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.850829Z digest=sha256:eb18601eda92c821a9639df2dd147eda8de423b8f3fb67f6610fd4f8417d6bd7

Observation e601f507-1b25-4895-bf77-6822b18f1367 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.865512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.865512Z digest=sha256:39e8abe61635db4aee49f42d0af00e973888d65244e597b63f2d438f62db7c17

Observation 86c8d089-5c58-4df5-955c-7d032858bf5d · outbound

This paper cites SingSong: Generating musical accompaniments from singing.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SingSong: Generating musical accompaniments from singing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.869956Z digest=sha256:0162632cdfe1bd22bdcbc9fe8326c383a0271b1f9c0962e8e457d4010a37a88c

Observation 5b17c951-e43c-4d6e-ba22-54cbf1d17685 · outbound

This paper cites Joint music and language attention models for zero-shot music tagging.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint music and language attention models for zero-shot music tagging

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.306287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.874701Z digest=sha256:57cc0e4ed3ced86fd1e803de334aa9c2b7a6887a36ea91742f79da9f1d879d60

Observation c37ee982-b157-47eb-8672-99a29fa3df47 · outbound

This paper cites Long-form music generation with latent diffusion.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Long-form music generation with latent diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.879005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.879005Z digest=sha256:cb224346bd1c1170bcefc7cf7993bbea3a3712dd25396d6d9a3d37bda5ce9715

Observation 214f2509-3365-4596-b4c2-900f898a8b6c · outbound

This paper cites VampNet: Music Generation via Masked Acoustic Token Modeling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.883184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.883184Z digest=sha256:8b3bf87d5732fa0d2c3b495a51c2f188bc658afe2300d421c4024f000d56a514

Observation 506b6021-d676-4bda-b7e1-a3572e26369f · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio set: An ontology and human-labeled dataset for audio events

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.294067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.887501Z digest=sha256:14e073ba6c1de5d322bdd15b4c55ea512e6dd15f186c9c85535751c955d908be

Observation 40bc01b0-0948-4894-8ce8-254c2ef195ae · outbound

This paper cites CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.891727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.891727Z digest=sha256:37f66af7ee2c7a022758f495122ca54d8e90f894a3876bf9d1540d294fdb4b20

Observation eca96272-26cb-4398-8d2b-0d765900055d · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.896554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.896554Z digest=sha256:f94d3ba96aba5863997ca607ad678ed0fb060a4f2cd1ed14766dbffbbb379178

Observation de35cce0-8fe5-440a-9b0a-fdbdb3d63b6b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.905866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.905866Z digest=sha256:85d65d76cf39108d4ea8a8a175af1b5aa79b006f21238371dbf929fa8b5c7bd4

Observation 524d5f25-d188-48a5-8ee9-4bfb8847f6c7 · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.914502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.914502Z digest=sha256:125755fec99a1a91dc4c3b6d6dc37bbbb71054e9d2390939a394e7177387e52f

Observation b5579418-771a-4633-a14e-002d6cbfc947 · outbound

This paper cites an unresolved cited work.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:22:19.280963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.919166Z digest=sha256:4fba7cb1fe15371da237596698a3ba8cff4c5457b59c7de0f03323043b911ce9

Observation b0111806-a9f6-4f14-af0e-a808c2b3e499 · outbound

This paper cites Retrieval Augmented Generation of Symbolic Music with LLMs.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Retrieval Augmented Generation of Symbolic Music with LLMs

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:22:18.874918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.928087Z digest=sha256:de155c2d7febd133a523502bf67000e08f0b8b4c85ffdb8bc152038f5f7764f8

Observation 7959ab72-2e7a-45c1-94b1-1a48bcaa3460 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.932668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.932668Z digest=sha256:d3fd7f63eeeb0357a73345108dc2a119ebfa01cbb486cae789edf7666dbf5801

Observation 3bb143b2-f6e4-474b-86cc-4b67d6bef4cc · outbound

This paper cites Auto-Encoding Variational Bayes.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Auto-Encoding Variational Bayes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.941042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.941042Z digest=sha256:5ab53f57efdffe7e7ceff76a3bf2ce3d2e23f0e24ef766ce244f786513c85a37

Observation c664983b-0f43-43a8-8a62-8aaabf079cfa · outbound

This paper cites SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.949611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.949611Z digest=sha256:853925562bce08196978c4a821c3bb75ea156d5948be384af655e126ad5aeea2

Observation a7f48f30-1d70-4f13-99b9-16df11f82b7d · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Efficient Training of Audio Transformers with Patchout

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.953788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.953788Z digest=sha256:e8ad38acaf73c408a1d5b3cb97f4d43083cce88c79ee05c6f0e9db2cb050daa2

Observation e13e77da-de29-4894-a1f3-33bcdf016491 · outbound

This paper cites PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.957983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.957983Z digest=sha256:9a516468fe88f8fc8fdda20ddd71b39cd94e97e2a5cb58322dcc255f14fc89b2

Observation 1f7e1e4f-e84d-4d2f-83ce-0fc66c659fb9 · outbound

This paper cites Content-based Controls For Music Large Language Modeling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Content-based Controls For Music Large Language Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.962109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.962109Z digest=sha256:9cb6cef13a777fb537c9a49c44453b3342e05cfcdee22c93650f773cd4908b9f

Observation 1d9d8692-1128-4c04-a4e4-833456fff121 · outbound

This paper cites Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.966629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.966629Z digest=sha256:bdb93d38269a93f3841bc139267cca80f16fc641900fe03d3637f9f00346714e

Observation 6e8c9f1b-e2b0-4c2d-9d5a-ed0418f56f1d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.970735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.970735Z digest=sha256:b2b120473e4954227e3bc3801f7599caf3f8b239c38bb464ae4b88bb80cd917b

Observation 7f74cdce-4863-49dc-8d77-68473f76c2d0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.974840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.974840Z digest=sha256:854239ddf5b06e7ea511c847e881ef449d8bbee7ca134eedb1c981241643abad

Observation d54e9a4e-c430-4d5f-bba0-18693ff386d3 · outbound

This paper cites Novice-AI music co-creation via AI-steering tools for deep generative models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Novice-AI music co-creation via AI-steering tools for deep generative models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.252792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.979035Z digest=sha256:61b69f243e90c95ae33ebde1d3c0395d96bf6d149314e0c88bfdc110fe1cd46c

Observation 3ae3d660-10ae-42dd-be2d-40be5d851534 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuseCoco: Generating Symbolic Music from Text

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.982833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.982833Z digest=sha256:fbf5edd7620029cad52aa5d42d3b0e2166f2f15a1fe5a1bda809ed2994b0a75c

Observation 45537720-6e91-4814-9d5e-00371d492646 · outbound

This paper cites Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.986909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.986909Z digest=sha256:2cc97e3b63b056558bd9fd0a3edd6e0ace0a8f34e09fd92dc8f1d41767586fb5

Observation 1f262fcf-3f34-4c01-86fd-2ff4870db5d6 · outbound

This paper cites The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.991759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.991759Z digest=sha256:0d020757655badaaef5d27b3fff5967463bf45d5eace84d86f5519baf4b82b54

Observation 6bec592a-7b1c-4b21-ad97-ddd546f84ca9 · outbound

This paper cites Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.240664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.996053Z digest=sha256:33d8148185f28e8f1e86063cb200325bdd79b0c1317bb96e6fa58d145d072b25

Observation e2c73553-2e83-4041-9037-efc495714b3e · outbound

This paper cites 2019.8937170.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2019.8937170

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-12T17:22:18.000506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.000506Z digest=sha256:c2d99ff2dd2a1f7637a3a5e86ea150d942789a42b4d445a6fe9c8f695842c29a

Observation 2b7d0675-2c8f-4c2f-9d98-1e41803c0269 · outbound

This paper cites Stem- gen: A music generation model that listens.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Stem- gen: A music generation model that listens

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.228074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.012749Z digest=sha256:646d163808846ac5c3b030a2545277cb99200e4f3194a2c1dd3ae751681e2c7b

Observation 94ce5295-f9f1-49ab-9bfe-e085ba0f4556 · outbound

This paper cites Zero-shot image-to-image translation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Zero-shot image-to-image translation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.216388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.016660Z digest=sha256:9d972c3f68e3f41b36e0c7a4ee85ad34e41e39dbcb2629647a59334c90cf1302

Observation 1e186346-0f81-47ff-b771-832e2a9a2092 · outbound

This paper cites Glove: Global vectors for word representation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Glove: Global vectors for word representation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.203529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.020738Z digest=sha256:46962adaab2e40917f695216029ba88c0feee6c95099d70ab9e90c28ecc9a4c5

Observation 6505c8e3-76a5-42cf-97d2-7044012f2901 · outbound

This paper cites Deep contextualized word representations.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Deep contextualized word representations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.033393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.033393Z digest=sha256:2e56ecd30980d1d34b55d7b7b03ffb871d44f2d71cac54215be276e0db1dc209

Observation da217417-315c-4b64-89c0-cda45a3910bb · outbound

This paper cites Generalized multi-source inference for text conditioned music diffusion models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generalized multi-source inference for text conditioned music diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.171499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.037046Z digest=sha256:db7f5cc2dfdc10256813a81722229de53965fbf3154c0e2cf1950f2a8e96299f

Observation 72792de0-2abd-4dfd-ba04-aea455bc04d8 · outbound

This paper cites Paguri: a user experience study of creative interaction with text-to-music models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Paguri: a user experience study of creative interaction with text-to-music models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.041327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.041327Z digest=sha256:80895b0e60b89ecb59cc457ff66085a6dcf9585e4ec5083bd4c082561fc02432

Observation a8511da5-b9d9-409a-846c-139d67be53a5 · outbound

This paper cites Audio Conditioning for Music Generation via Discrete Bottleneck Features.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Audio Conditioning for Music Generation via Discrete Bottleneck Features

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.045976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.045976Z digest=sha256:f17bdf1e1dd71a68a92fe20281b7d496302c8d5239dca706f0c09eca17afdb32

Observation c3572c32-6a17-4956-85e2-3fb128887240 · outbound

This paper cites an unresolved cited work.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:22:19.157232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.050497Z digest=sha256:5b608af0ce4219452608239286ce5000a2688b3342ad0584738424182c2e5400

Observation b6032a99-19c9-43ad-a077-09b2585925e6 · outbound

This paper cites URL https://doi.org/10.1109/ICASSP.2019.8683855.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models URL https://doi.org/10.1109/ICASSP.2019.8683855

Reference 54

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T17:22:18.502761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.062013Z digest=sha256:ea4a0b6371e753dd96ffb4e5e430dea94094cb73b948f6584b81385e8e508e88

Observation 2f28cfb2-25c1-4eb9-8f49-b09ede1b1942 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.066712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.066712Z digest=sha256:53fa27207d431c9429f533fa9c9ea81d63052e6ee6118a4b3ca9ea27dcad2e28

Observation bee62ef0-c927-4586-b0fa-f3c5a1a85f37 · outbound

This paper cites Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.071838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.071838Z digest=sha256:11f258485a79f732dae431f18561d844a47343093e410ed94045491b424af074

Observation 4ae40fbd-b2da-456a-8d2d-f086b3d3373c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.077072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.077072Z digest=sha256:f788625ce59cb22fc514b11861b1c406241e444e6a1e8ebb93c61e6614b18386

Observation 36568fc4-9e0d-42c1-8a3e-651dd21d23c3 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Plug-and-play diffusion features for text-driven image-to-image translation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.143365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.081609Z digest=sha256:a5a5d3b030b8638ba02d071de1ec3b28461245e7c1d6a63012e7945a859e694f

Observation 5f8e728e-061b-4156-9bd6-00b3c5458970 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models WaveNet: A Generative Model for Raw Audio

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.085750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.085750Z digest=sha256:8c4ba5281cd7d457fbd7607e6667e39a32e9d240e6c1e9ffd879747d02a69b6f

Observation e141c705-cb91-474c-a455-2ec8c01301d8 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.094014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.094014Z digest=sha256:3f33fb0008e8d0c40104bb38bc65f5ccdf806cb3090b80291704191007204bd0

Observation 7fa745de-7c50-45fe-b3b1-1506c629e7b0 · outbound

This paper cites Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.098324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.098324Z digest=sha256:e0cdffdd27a6da8562180dd6120a6305213a67c41bb0529decd99975bf6ca62f

Observation 5b21ed84-4989-4f8d-8f77-4858ca1e080e · outbound

This paper cites Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.102619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.102619Z digest=sha256:7641620dd8d8bebed6870500e1b6429f8347d794e399af7408cfa9f731de28dd

Observation deaafc50-5a0a-4bd1-8afd-eccaff0c3403 · outbound

This paper cites Iteratta: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Iteratta: An interface for exploring both text prompts and audio priors in generating music with text-to-audio models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.129595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.106545Z digest=sha256:937f559874ab74eab569bd7fef2362046242ad4d610ffd18ee94b3352741f376

Observation 2703ed64-5d4f-4d4a-a51c-7e9cbb6c0e02 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.110641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.110641Z digest=sha256:92efbfe1d855e2936af122966c21c2d513a5a54716f7b950b5dcfee64f1a3f35

Observation d7d3a452-f80f-4be2-879c-54b11508117a · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.114817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.114817Z digest=sha256:327559007f1d1411e1e79fadf64bd10a096a094ee354a37f7919c06694e02822

Observation 12264b63-e1f6-41f7-bfb4-7b46197d5f43 · outbound

This paper cites JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:22:18.262901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.119324Z digest=sha256:4d525dbe6fd867c7e2b794019e8ff8d9e95bc6900e9adbbb3ad992bb8e73af68

Observation 7aaae18c-fb86-4487-a150-aefcaa3649a6 · outbound

This paper cites ChatMusician: Understanding and Generating Music Intrinsically with LLM.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models ChatMusician: Understanding and Generating Music Intrinsically with LLM

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.123882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.123882Z digest=sha256:8e572cbfe892fcca3b22704bd90fc8d89b335b6ccdf10ecaefd4bd8dc730617a

Observation 3dedc977-28b4-44cd-a5de-2858df3024b6 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.128284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.128284Z digest=sha256:8cc2cfe50a130aac8079382cafe170aa1095106aa4a864b1fb5f5cf9321c078c

Observation 7d503132-40d5-45a0-8a85-75d23867d85b · outbound

This paper cites Cosmic: A conversational interface for human-ai music co-creation.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Cosmic: A conversational interface for human-ai music co-creation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.115111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.132747Z digest=sha256:a8bc04e5ac85e6375dff3677dc7e05b5c8b662aa178aa33d698a20bdef95d167

Observation 28009de3-3aab-4a83-b2f9-07d17531f7d8 · outbound

This paper cites Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.136706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.136706Z digest=sha256:d649278d657d96afb9fe5eab036e97adbbddc4a0bc168e09d77233f6650e7710

Observation 4c418d10-b949-4159-add6-a74d84190332 · outbound

This paper cites A Survey of Large Language Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Survey of Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.141425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.141425Z digest=sha256:df3add6e114e2324956fce6b85fe3dfbeef9a60984d5b32ed39e2e11b381d0b7

Observation b00db70a-dfd9-42df-9df6-e87647de670c · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.146028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.146028Z digest=sha256:5e387725becbb190005a9be85166887056ab1d43b5814e8229d2ff2e5683119c

Observation f06bdcae-e553-4106-812e-acbba9ecca99 · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.008497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.008497Z digest=sha256:b5e81e184c869f57669cb3fbe415ba21ced3290f01fc5a1a6d5c5ff301739d2e

Observation a57bb2bc-47fd-4c40-afeb-66e9746f6412 · outbound

This paper cites Exploring XAI for the Arts: Explaining Latent Space in Generative Music.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Exploring XAI for the Arts: Explaining Latent Space in Generative Music

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.829212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.829212Z digest=sha256:9eac37f5b83574ea343a81a6cdebf4ecbe40fc82bb06c0a045598e85eaa7db57

Observation 6bbe475b-5897-4987-86d2-2ab251e89d31 · outbound

This paper cites 2003.819861.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models 2003.819861

Reference 2004

Resolution
malformed identifier
no resolver link, observed 2026-08-12T17:22:18.090093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.090093Z digest=sha256:754d427d37806528472818c8981e490fb25a5348776aa016619015684c1e092b

Observation 370d21ad-03b8-494f-a600-fbf4f2ef17dc · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Improving Text-To-Audio Models with Synthetic Captions

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.945138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.945138Z digest=sha256:438191b20fc392231430281f0ff5bbb2ed7f750e0a326d5e2b81a662d91c8a2b

Observation d803a4f9-532f-4a98-844a-4347d9fe9899 · outbound

This paper cites MoisesDB: A dataset for source separation beyond 4-stems.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MoisesDB: A dataset for source separation beyond 4-stems

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.189399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:18.024699Z digest=sha256:42013538be8d903152ebdf7947cf99c95d280f65900e67501b04cfbc5bf57daf

Observation b6bd2d2e-5f02-4445-ac86-2279b1482464 · outbound

This paper cites MidiCaps: A large-scale MIDI dataset with text captions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MidiCaps: A large-scale MIDI dataset with text captions

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:18.004318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:18.004318Z digest=sha256:d177291f8927c9fd546e1cf5883ceb6bb4e49b91a1ffb37641817fd374858610

Observation 91e7ac71-eccd-4b9e-a718-16e7989b14e1 · outbound

This paper cites A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models A Comprehensive Survey on Deep Music Generation: Multi-level Representations, Algorithms, Evaluations, and Future Directions

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.923559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.923559Z digest=sha256:854dc023c8b484dd89692febae4908664d8f3c38289661d0ec0c639cbad5519e

Observation bf5cb6ad-9aa4-430d-ae8a-95a11b661c7a · outbound

This paper cites Au- diocaps: Generating captions for audios in the wild.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Au- diocaps: Generating captions for audios in the wild

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.265481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.936966Z digest=sha256:6817094e66e95e93f255ff343e8de530e64b8e8e5f82e551420b2eab755c2d74

Observation 0035bc81-1603-4cf1-8b24-b65ea1e86213 · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.910125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.910125Z digest=sha256:2b52ba29cab736539a82bdb29035385f6476cd6a3157f10a661d614cd88c0092

Observation ca295f9c-16c9-4a67-9d7e-c3ddf8763784 · outbound

This paper cites SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.861124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.861124Z digest=sha256:b009d62bbf652c9dab89e8719a020c4c796b71e2137394aae2d5e86941f1d7ca

Observation 08a68021-b2fe-4f82-9568-cfe8712a4d44 · outbound

This paper cites Jukebox: A Generative Model for Music.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Jukebox: A Generative Model for Music

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.855404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.855404Z digest=sha256:2cab4635c54a2d123742aee645f0084ac5cfe0e56088bc6ea48e6b4c3d6e65a3

Observation 5decd465-1463-4b20-b74a-8ff02ef74acb · outbound

This paper cites Musicldm: Enhancing novelty in text- to-music generation using beat-synchronous mixup strategies.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Musicldm: Enhancing novelty in text- to-music generation using beat-synchronous mixup strategies

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.343974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.838339Z digest=sha256:fc14a29a86a38369a8e50b0b89d2ee3c54132b9209022a33b2f7acb20ffd27c6

Observation 56994326-27b9-4dee-bca5-b81e86bbbaf9 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.355936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.833825Z digest=sha256:82d56db64b438145dae04dcdd6908ea5ae8a52417f798c4fe12093832ac353fd

Observation 0a50e756-deba-4490-bf58-6fdf0d6ea5b5 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:22:19.332606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T17:22:17.842502Z digest=sha256:99c38fa9b64d9012f8c8ed1ea185feca7cb43549a57c16657131b1fd92b76668

Pith citing papers

No inbound Pith citation observations are available.