Pith. sign in

Paper Citation Record · LEDGER

Mustango: Toward Controllable Text-to-Music Generation

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2311.08355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08355 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:28:26.609238Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.335915Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1ac52b9d-41cb-475b-8151-720c9c6fade0 · inbound

Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation cites this paper.

Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:48:52.766523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:48:52.766523Z digest=sha256:ee59264464bfa868f9159c0c01dc15fca516f655bdd32410956671a5ba7af7f8

Observation bc78afad-bff9-43ca-be72-4027c2db333b · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Mustango: Toward Controllable Text-to-Music Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.847605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.847605Z digest=sha256:2f28a61a159b8c98787fc11be07fb4d78296ddf6443dd286750fae4cf65bffa5

Observation f5134e4a-541b-4f83-b567-96336310e21f · inbound

Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning cites this paper.

Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning Mustango: Toward Controllable Text-to-Music Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:23:49.242900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:23:49.242900Z digest=sha256:99ccc993fc0a913ab9fd5f0670a6e033ce17f2a34a6af8b737c1bc8dfdbc0cda

Observation abcdffa6-0467-4f04-94e8-a8f31291579b · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:56.333453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:56.333453Z digest=sha256:3d76fdc823905cfdd558e15a38857c52f7e97eb79c5410e1ce87a5d48fbd7866

Observation 3158d74e-844f-4719-8b6e-4ec31bbdd231 · inbound

Genre Controlled Music Generation via Activation Steering cites this paper.

Genre Controlled Music Generation via Activation Steering Mustango: Toward Controllable Text-to-Music Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:41.140747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:36:41.140747Z digest=sha256:b9f8a6d2e4342d313a1f9587651203f06e69b2f9c5ff4f459125896014a72ef2

Observation 4a2b1cb3-759e-4290-96a2-5d7f6423fc9b · inbound

DanceChat: Large Language Model-Guided Music-to-Dance Generation cites this paper.

DanceChat: Large Language Model-Guided Music-to-Dance Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:26:42.195662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:26:42.195662Z digest=sha256:cdd70786b3f467eb9073261305230e967ea64c1768660c128f256102374679cd

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:a75cbdac1584f1271e63d1f38a1d9268c970db6242538561f6fc707e8f76bdc2

Observation 7b451098-c377-4e48-ac5e-9248f2331f46 · inbound

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners cites this paper.

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners Mustango: Toward Controllable Text-to-Music Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:51.505249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:51.505249Z digest=sha256:efe82a01e793d49fbe371ea23e3364edb7b1dc9822abc5f3d3bbd1c397fb94f2

Observation 8b43321c-ee0f-4bd9-95e1-1eab3c52c66b · inbound

Benchmarking Music Generation Models and Metrics via Human Preference Studies cites this paper.

Benchmarking Music Generation Models and Metrics via Human Preference Studies Mustango: Toward Controllable Text-to-Music Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:41:08.322086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:41:08.322086Z digest=sha256:aa7f45a725591fb7891a2a9793f8fd4e97f4f6f1acd0cb126f7f446926156297

Observation 2570ed29-32d7-4832-a1d2-7ab9448d92ce · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Mustango: Toward Controllable Text-to-Music Generation

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:44.834686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:41854f8fd0f6003ed63b939af5f4e17b87444a127f3d1583fe8b8a4b8c5ef5f2

Observation f2198e82-a9f1-4308-994b-d4b265a5eaf5 · inbound

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization cites this paper.

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Mustango: Toward Controllable Text-to-Music Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:49.274001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:49.274001Z digest=sha256:ccda90a7ba45a83efc83c701ba6a6f0c4667285ed8cadc2dc4af16eeb7401400

Observation 96a2debe-1a19-49b5-b9b0-f38a10d7574e · inbound

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions cites this paper.

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions Mustango: Toward Controllable Text-to-Music Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:31:12.444751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:31:12.444751Z digest=sha256:187e237a4082560b7547c2681a9bc309b798196feef66833547de64432857a2b

Observation 4d39e694-9084-43a8-bad0-ca9069f4858c · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment Mustango: Toward Controllable Text-to-Music Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.561225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.561225Z digest=sha256:5e73a654b072ce6dbfc321c793bfd3ac147d231f2f0d08e2f284e6c822b8f573

Observation c5536b7b-197b-472b-b36a-929ede8c8c6d · inbound

A Survey on Evaluation Metrics for Music Generation cites this paper.

A Survey on Evaluation Metrics for Music Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:27.397749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:27.397749Z digest=sha256:7651eaa4ee8d531109a5f727b6525d54a6298f8e13f7096bfe37c41cf8d0302f

Observation afb0178d-35fc-47ae-b92f-6849f77c9448 · inbound

AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation cites this paper.

AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation Mustango: Toward Controllable Text-to-Music Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:16:08.841052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:16:08.841052Z digest=sha256:0eaabe05c4ae74bdd1e9ba90ba03e2136ece1a8868a79d368e06bc0e22b4b312

Observation 1f5ac5ce-20ee-479a-98d8-f6bcb284dddb · inbound

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization cites this paper.

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization Mustango: Toward Controllable Text-to-Music Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T13:12:14.509681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:12:14.509681Z digest=sha256:825fa35318c08799666dcd48b3e44faa4422471cd4118957e1d4f38f07a9e931

Observation 24ef680d-7125-4f95-ab00-09148a105fa6 · inbound

Segment Transformer: AI-Generated Music Detection via Music Structural Analysis cites this paper.

Segment Transformer: AI-Generated Music Detection via Music Structural Analysis Mustango: Toward Controllable Text-to-Music Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T20:54:14.265474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:54:14.265474Z digest=sha256:1254267a21bbf5d091de26d1cfd52a2639ddd5e22184091249529da42b13fd4e

Observation 394b1895-d2f0-4adf-a7a0-5b1e62f1066c · inbound

Steering Autoregressive Music Generation with Recursive Feature Machines cites this paper.

Steering Autoregressive Music Generation with Recursive Feature Machines Mustango: Toward Controllable Text-to-Music Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:10:54.487022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T05:06:22.119543Z digest=sha256:eded3b6076aa5e3f1f1a43133a7f4f6f9e52c31472123595cefaad924ab88d3c

Observation 8b75afb5-e8bf-4f16-b92d-4f948b9b7f8d · inbound

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions cites this paper.

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions Mustango: Toward Controllable Text-to-Music Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:24.665422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T13:10:29.213917Z digest=sha256:76faf17e0bd9911c4b53a9b08cbd53c0ef17f1c75e364f2a12487e6f02918001

Observation dc665989-1d9d-4ef0-8f1b-81cc125c072d · inbound

FIGMA: Towards FIne-Grained Music retrievAl cites this paper.

FIGMA: Towards FIne-Grained Music retrievAl Mustango: Toward Controllable Text-to-Music Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.213111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T23:32:09.401023Z digest=sha256:7e16599cd85dc1f6acc8f4a9b16f12e7b803a538128f12780a909530db972129

Observation 11dd6db8-79d2-4a3b-bea3-e6693339f5e8 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Mustango: Toward Controllable Text-to-Music Generation

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.337322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:70a5b778f99dad2cf52de1f5d4fc99eaff2e6b60894c99a4eb48ae6f6b833df2

Observation 59ea047e-8644-4b8f-93a2-04a424b2b992 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Mustango: Toward Controllable Text-to-Music Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:b1ed3bb06f5299374b6bdd2cd6a099c303c763302030377df7e984f7b7369734

Observation 8fc956f7-7bdc-42f8-9c5b-bb93f9cf7fbc · inbound

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model cites this paper.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mustango: Toward Controllable Text-to-Music Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.979881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.979881Z digest=sha256:45ffc6e4d2a102e522d25704756b02995515bbb8f46e86a4961e3f380034e389

Observation b301c622-6332-45e6-a3b7-1ad8f5731c5a · inbound

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping cites this paper.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Mustango: Toward Controllable Text-to-Music Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.609238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.609238Z digest=sha256:88a1373187cc5c2a5d1bb2fdf1ec355e33a53f7309d18e056f2ee98904c06f4e