Pith. sign in

Paper Citation Record · LEDGER

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:54:23.104469Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a6515c44-4ef2-4503-b7bd-b62dc5b0f1e6 · outbound

This paper cites MusicLM: Generating Music From Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.857355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.857355Z digest=sha256:170f56a3b196e8abf7bd2faec533919e740a1b55f07957d8728e303724b3c86f

Observation 4f43d365-518a-40f4-84b4-30d389fc9897 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.867787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.867787Z digest=sha256:b1367a92654f67afa793e481948449ce3aec1569ca5e589ac6ba81beb87f3351

Observation 273693a8-f2cf-469e-9228-d62a582ed790 · outbound

This paper cites InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.539916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.873016Z digest=sha256:d3647c325722a89b11a2a3166a6dedb49e130bc55753fa59bcd73c2a087ced4c

Observation e1d1ea82-7ed2-438b-92db-8fe2228288cb · outbound

This paper cites InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.524635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.887371Z digest=sha256:667c262e4cc710d23d974e8a5d11221a8710ab7d554b1546e93c5e69882d218e

Observation ae39f69c-a169-4759-b809-9e29a31f997f · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Fast Timing-Conditioned Latent Audio Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.891858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.891858Z digest=sha256:d2badd440c321c00f80a444c54ca22601f2ede1771e251dccc26d85c0d204d99

Observation 105e8c7f-b948-40ec-9d5a-68954b12baaa · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.896826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.896826Z digest=sha256:439dad02f77c1cf8b6b7cfe75298b44e9be678978ee3d8e573752b2e537d03cf

Observation cc7317ee-6fd8-49ca-b85c-d9dc23da835f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Classifier-Free Diffusion Guidance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.901792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.901792Z digest=sha256:bc6286fc2e577b5302afc514fb9d37910c369a6253c05a36695be0e9d9ffeddb

Observation 4f81140f-afe0-49b6-8345-0031d08cab2b · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.907977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.907977Z digest=sha256:035a53030597ed735cbd2d51dce948fb491e09189df61107cee40875004fbcff

Observation 05672a3b-b744-4038-b8c4-c0a1684da041 · outbound

This paper cites Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T17:54:23.318894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.922897Z digest=sha256:b059949947543d651897892f735adfe9021049f4793b651b3e4189a32c1db35d

Observation c39156a9-052c-4d01-a835-1c5b26ce463e · outbound

This paper cites Symphony Generation with Permutation Invariant Language Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symphony Generation with Permutation Invariant Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.927741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.927741Z digest=sha256:533b3a66da021389ef46db133cbffba8307187c8dad91a52cec2fef0535cccc7

Observation af262a4d-1262-4886-a43d-acd0802b97e5 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuseCoco: Generating Symbolic Music from Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.950607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.950607Z digest=sha256:1c97c0f0ee43066baf9bffab6b615eef6e81f28d78a2dfdcd4b09e25ed7f687c

Observation 8fc956f7-7bdc-42f8-9c5b-bb93f9cf7fbc · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mustango: Toward Controllable Text-to-Music Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.979881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.979881Z digest=sha256:e8910e567ddd85a51ea5c91983b06f5da6b239efc05184003c684f661ecee1e2

Observation 50bc3d6e-8142-479c-bc00-ec77c11e1c30 · outbound

This paper cites Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.032198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.032198Z digest=sha256:a197e8b19544f9a0a6c5ec3f19317ccc8de0e5231433a16ed0d8491d57a3b339

Observation 7ae14d82-eb1b-4d0b-bd1e-92bb6d562f48 · outbound

This paper cites Symbolic Music Generation with Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symbolic Music Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.055772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.055772Z digest=sha256:e58cdd84fa27ebaa091343ae1054781f3aa075d90d901ee715e01426ff7bc389

Observation 3177bb59-396d-4118-8ae3-b8ca441ec781 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.080284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.080284Z digest=sha256:eadae327d937c5402b0170d5506db076b41705e37e4862d14dcdc39cf68deb6b

Observation 81e39820-6994-451f-99b8-a8c7a942db3b · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.085526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.085526Z digest=sha256:58420aabfcdfc211c390354a09f2800e80fc056b46201e1a0e45b25328e5e0c7

Observation 44d78cd7-78f1-4711-9823-c20a555e641b · outbound

This paper cites Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.094990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.094990Z digest=sha256:972987d695ac66abd57ec0024a30ba8f3557d3338e3646e5c590414d0f74a162

Observation 3571aefe-3462-4d1a-8a65-1b5a2c7edae0 · outbound

This paper cites Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.099837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.099837Z digest=sha256:e0c99cc1869ed4b843006831804d384f8d9f701feefe1f9876fb15e610942f48

Observation 663438e4-e295-4e3d-851e-943b5c9e488a · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.104469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.104469Z digest=sha256:7f9cdeef860f234b6881bca0bf2de70950ef283ac23d40f4c0e07ad601431e56

Observation 92bf661a-9c73-41c0-ab1e-0979609781d1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Auto-Encoding Variational Bayes

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.918152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.918152Z digest=sha256:fb0effbde83f1fca6c6453cf034ebd7a90e29825392afad171a88644a7823f1c

Observation 70069825-e6b9-4023-bce2-8e464c80c663 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.882276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.882276Z digest=sha256:177effbb42283cb6ecb237d6c8863ff4bf4d6290f2d399ff9851cd04d3e64ae7

Observation e8e6da20-9c3e-400e-8201-e5f354deb0db · outbound

This paper cites POP909: A Pop-song Dataset for Music Arrangement Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model POP909: A Pop-song Dataset for Music Arrangement Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.090051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.090051Z digest=sha256:9fa56702a897466f413d26b00bad806c55949b94d38e1f02412b5fa411f79ea7

Observation 6fec1bb8-bd79-4231-b913-453677663ad7 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.913422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.913422Z digest=sha256:1a8fd610192857c7da6637810efae31571f20514c622d4b1bfc1c4049ac8b847

Observation 90a167dd-addb-4891-94d6-4db58d9af18d · outbound

This paper cites High Fidelity Neural Audio Compression.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.877436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.877436Z digest=sha256:aae556c41f8b89d7eefda9b7156a767cdd7a4c9af9f63c354bb16269e0115655

Observation 43418b24-ca96-4809-ab7f-55dae6a259ec · outbound

This paper cites GPT-4 Technical Report.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.839229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.839229Z digest=sha256:1e137f5cdaf047d0287840bb6dbbf4c270399921a8c5dc06ce62ee80e5a1944e

Observation 16343409-1236-4d50-895e-c0f0a44e432c · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.862705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.862705Z digest=sha256:b8bc03c95fddf07118f9002bb3d772ae76e5fda4b5cb4c6113f87ec5433cdace

Pith citing papers

No inbound Pith citation observations are available.