Pith. sign in

Paper Citation Record · LEDGER

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:54:23.104469Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a6515c44-4ef2-4503-b7bd-b62dc5b0f1e6 · outbound

This paper cites MusicLM: Generating Music From Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.857355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.857355Z digest=sha256:c2411b8e0b960b65035fd6c475c843ce1ad3e4b01cbf23792f08e8421ac22f5a

Observation 4f43d365-518a-40f4-84b4-30d389fc9897 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.867787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.867787Z digest=sha256:bc57a17ed59bc38da28738ae9a98fe9916231a25493bb4db830d695258638ab6

Observation 273693a8-f2cf-469e-9228-d62a582ed790 · outbound

This paper cites InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.539916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.873016Z digest=sha256:3fa2b0d943a0c3ac78e7177a12eb9b81778dc4d2537ffcb106089211e0c7c602

Observation e1d1ea82-7ed2-438b-92db-8fe2228288cb · outbound

This paper cites InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.524635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.887371Z digest=sha256:f510ca608e1d9a38d64329a76dc00b4e4ac7d1fef4834df9ef24630b75397ecf

Observation ae39f69c-a169-4759-b809-9e29a31f997f · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Fast Timing-Conditioned Latent Audio Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.891858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.891858Z digest=sha256:375d259d3dc4863bf7322fa24d976dba6a70d985fa1d13bd35c222bab4545e5d

Observation 105e8c7f-b948-40ec-9d5a-68954b12baaa · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.896826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.896826Z digest=sha256:22fa54863d39d77316753845a8c7830a476ed9841ed7161fee1279908687eb4a

Observation cc7317ee-6fd8-49ca-b85c-d9dc23da835f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Classifier-Free Diffusion Guidance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.901792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.901792Z digest=sha256:50b89860128498b632c4dd33f492f6e715691f7bc4a33d476e261bd12565931c

Observation 4f81140f-afe0-49b6-8345-0031d08cab2b · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.907977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.907977Z digest=sha256:70b8569f26a6f56f0a2d68be744f8b09dcd1d00d019c4af10ff601b0c91dbbe1

Observation 05672a3b-b744-4038-b8c4-c0a1684da041 · outbound

This paper cites Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T17:54:23.318894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T17:54:22.922897Z digest=sha256:9f2db66d4a1e6b5c885fc485d13772cb1d39957f29a2b0c97d5446279be90d99

Observation c39156a9-052c-4d01-a835-1c5b26ce463e · outbound

This paper cites Symphony Generation with Permutation Invariant Language Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symphony Generation with Permutation Invariant Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.927741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.927741Z digest=sha256:48fa43a9edba3ced50479dfc536cfdd30c36b45deb6bd163f1f9340cdb6313e2

Observation af262a4d-1262-4886-a43d-acd0802b97e5 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuseCoco: Generating Symbolic Music from Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.950607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.950607Z digest=sha256:7c1a2b553ea9705b5ab6c713cef3b1f477e83ffc3185e2d320f543aade4658c3

Observation 8fc956f7-7bdc-42f8-9c5b-bb93f9cf7fbc · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mustango: Toward Controllable Text-to-Music Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.979881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.979881Z digest=sha256:e1d8502d03046ce4ada9973002ff56cd17555c157c674b85f36a916a614b3b04

Observation 50bc3d6e-8142-479c-bc00-ec77c11e1c30 · outbound

This paper cites Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.032198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.032198Z digest=sha256:27042ba83f833cb08e526d3652f725f4f54b9c165516cb1cfb78abb22131a306

Observation 7ae14d82-eb1b-4d0b-bd1e-92bb6d562f48 · outbound

This paper cites Symbolic Music Generation with Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symbolic Music Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.055772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.055772Z digest=sha256:5bd4da3e22ff846e57b95f7bc5e6179d81210f6fd7f3bf57fe833eaf91e4b281

Observation 3177bb59-396d-4118-8ae3-b8ca441ec781 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.080284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.080284Z digest=sha256:e1466e91e8e8ee04492ce620ba42e84e15f94b11b1b82d1ef0ac51b15ab812b7

Observation 81e39820-6994-451f-99b8-a8c7a942db3b · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.085526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.085526Z digest=sha256:de691f15a1d11ad042579d50c5b21dd7031d44dbed78511cd55d9176fe4daa56

Observation 44d78cd7-78f1-4711-9823-c20a555e641b · outbound

This paper cites Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.094990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.094990Z digest=sha256:31f428fa5a4f675bb9583d5605486e85cc607b712fb90d66005d4c93e8a0d7d8

Observation 3571aefe-3462-4d1a-8a65-1b5a2c7edae0 · outbound

This paper cites Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.099837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.099837Z digest=sha256:901357868d4aef0eb9fb7be44b4e068d0d97e7fb35f378ffa78d5ca52589d525

Observation 663438e4-e295-4e3d-851e-943b5c9e488a · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.104469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.104469Z digest=sha256:a0bdafcfe9c1f2d4d6f83719628a8e260e557d62295351f2241d402cf2453cce

Observation 92bf661a-9c73-41c0-ab1e-0979609781d1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Auto-Encoding Variational Bayes

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.918152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.918152Z digest=sha256:a613807ab8b1cff041c5767c68acb27ebceef9c2fa62bea0906bb0c27836f07e

Observation 70069825-e6b9-4023-bce2-8e464c80c663 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.882276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.882276Z digest=sha256:b141c10631b4793c510aa52fba273492961eac27cb3dd51d2df97413ede27251

Observation e8e6da20-9c3e-400e-8201-e5f354deb0db · outbound

This paper cites POP909: A Pop-song Dataset for Music Arrangement Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model POP909: A Pop-song Dataset for Music Arrangement Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.090051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.090051Z digest=sha256:47939acec12198db4b8cdab216057a1e490d16940ca6db6e19aa260f479a8cd0

Observation 6fec1bb8-bd79-4231-b913-453677663ad7 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.913422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.913422Z digest=sha256:84eaa4536fb1e03e3a54e529a8fcbd4d13560d7bd8ed084435023aa6a03cf4a7

Observation 90a167dd-addb-4891-94d6-4db58d9af18d · outbound

This paper cites High Fidelity Neural Audio Compression.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.877436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.877436Z digest=sha256:6b79addf728b10aae4aff6632b3607e71e94a1de444026b6bea9d1a7a35e1930

Observation 43418b24-ca96-4809-ab7f-55dae6a259ec · outbound

This paper cites GPT-4 Technical Report.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.839229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.839229Z digest=sha256:bc45c4d429877802df037408823bd46e3aad38f629f1df3938f2b916670172ff

Observation 16343409-1236-4d50-895e-c0f0a44e432c · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.862705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.862705Z digest=sha256:2ac876aa54c399a97451295bd918fa91b10207bad5aee8ec9970cd8e130581f1

Pith citing papers

No inbound Pith citation observations are available.