Pith. sign in

Paper Citation Record · LEDGER

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04378 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:42:37.133267Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved15
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a8874bf-1f7d-4649-9a84-57fa7402b568 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Self-supervised learning from images with a joint- embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.080911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.912135Z digest=sha256:8162ef74b03c6bba96b918ad9fff4c4fdfb9b41eb157ad8625131bd8ad70a73b

Observation 87ab2416-fe04-4adf-a606-c7d2920dc467 · outbound

This paper cites LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.063959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.918353Z digest=sha256:e6a86d9d9e1c43d147a64e7a456ceab76919420db7cbe9f686c476d4341d8002

Observation f0ce2832-a2bc-4771-a2d8-30745e4a1e7d · outbound

This paper cites Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.048869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.923728Z digest=sha256:32e0cdeafc46aceb47dab7c939053aa6a54093066fdd2714036ec99fcd52a56d

Observation 2cf73a05-3977-49d2-80f3-fc2571c3f015 · outbound

This paper cites Zapata, and Xavier Serra.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Zapata, and Xavier Serra

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.032439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.929545Z digest=sha256:55f27cabd96b4d5d759c85203400fbb47c86df8f7b2c4576cd3552e27a0e89cc

Observation 8ab0eeaf-e0d4-4b3c-90ec-aca0412ce320 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Emerging properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.017269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.934794Z digest=sha256:44d478b94b48a52697ee28a4bbf0e6e0f3754d4148f77dff0a9b458fce075e73

Observation 59777f34-56c3-4d60-b615-a86b8be63621 · outbound

This paper cites Codified audio language modeling learns useful representations for music information retrieval.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Codified audio language modeling learns useful representations for music information retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.001701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.939949Z digest=sha256:0a13887215f52da8313b20c0368b249141d8c6222c42dfb2f8196ae6770a0074

Observation 5c3baffb-2a94-4a1e-8893-96df08d80ae6 · outbound

This paper cites Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.568953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.945515Z digest=sha256:4a4ea3014d18477335e7a29217b64cff1399782a498690c60593563751b6af95

Observation 8b25a6df-49bf-4ad7-9374-fec2b95402df · outbound

This paper cites Pixelflow: Pixel-space generative models with flow, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pixelflow: Pixel-space generative models with flow, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.950691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.950691Z digest=sha256:ad6aa382705cca1e6cc3e6b8bf86135f1534f8c5a2176ada86a1e53e0aefd2d7

Observation b76b3215-d6ea-4ee0-a166-f468f897a65c · outbound

This paper cites Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.955570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.955570Z digest=sha256:ae725b2866592364cc6a69142dcfa2043436b403b7f8883e49c7d7bb506f3786

Observation 0872e412-f945-466a-b80e-e54c31e0fc54 · outbound

This paper cites Diffusion models beat GANs on image synthesis.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion models beat GANs on image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.976481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.960678Z digest=sha256:03fb4b3ae2e1edddfed7c947b3b20c9fa6b4c63fbc79ae4874f307398d7cd4a1

Observation 99673cb7-a190-4952-9601-3eaf6166e80b · outbound

This paper cites Generative modelling in latent space, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Generative modelling in latent space, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.962049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.965421Z digest=sha256:410c332377100489471196ea0c7b11fbe3f9f26858d2b8fded333649b3d3d5ee

Observation ee53d3be-8918-411f-a66a-0973ca62549c · outbound

This paper cites Hawley, and Jordi Pons.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Hawley, and Jordi Pons

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.946586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.970418Z digest=sha256:6772859ef40575bec55615e98663bbbfc010c1385b3c93e794b86b115be40750

Observation ec1dd9e5-264d-4417-b06f-3f8a9553fca9 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Learning and Leveraging World Models in Visual Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.975241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.975241Z digest=sha256:c372b85788b17e5bb50f89af34ebd8be8b4d6790c54aab7b3bde6c3b8662a3a4

Observation 823ffb86-9a19-46ce-8fac-aed9cf3d4be2 · outbound

This paper cites World Models.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.980422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.980422Z digest=sha256:1d7412a6f8120bd44355f5c966dd3d1f3711f2434129344e01d131df2d22499d

Observation 19e4507f-3ca4-4b83-985a-1df8eaaef8b7 · outbound

This paper cites Using a joint-embedding predictive architecture for symbolic music understanding.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Using a joint-embedding predictive architecture for symbolic music understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.929987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.985406Z digest=sha256:89d189a0b759e2945f0b8a5d9ac6c9ba918f441fd74e14c419c6ecf6baa00ea9

Observation 2a4bf12f-9026-413b-a7da-0e8174363d51 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.914256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.990376Z digest=sha256:aaafefe718a68f67ce06c6e44ae4cd2b798b27bb4b6dc0d3ee1332624133fb16

Observation dcf967d7-5d7b-4c45-9c3f-0b29962cccbf · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.898382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.995429Z digest=sha256:d0bd8ca0d8affdae39d69bf9ddd2d7fd287801d278ee6439d1521fb1d79bb600

Observation 58bff056-ca19-4e5f-94e5-a646e70108bd · outbound

This paper cites EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.882781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.000539Z digest=sha256:1c347df8d3656ad50e3f1a95b7a31c09b11daccb00306fbace297cc6b19b0ae8

Observation 776a0ebf-080a-4d0b-bcc9-900b302cffdc · outbound

This paper cites How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.005741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.005741Z digest=sha256:71aa00d8731d5334066eb52bcffabd1e6e4e22cda40fa946ad0229dc1d069aa5

Observation cb009084-0471-46a6-918f-b6819232b411 · outbound

This paper cites A path towards autonomous machine intelligence.OpenReview preprint, 2022.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation A path towards autonomous machine intelligence.OpenReview preprint, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.867340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.010781Z digest=sha256:da42c25cb410c5d3d07a6fb9ec8a4de163f26e02902670ff2595c5ef23083637

Observation df7df0f6-71f8-4b7f-92f3-da69e35974af · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.015774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.015774Z digest=sha256:d2d1d861b532c30b47e656344712454c7337f01d01d2f9bee83fe6bb75438dd5

Observation ed68fb50-386a-41f0-9356-2fe0385ab834 · outbound

This paper cites Swin transformer V2: Scaling up capacity and resolution.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer V2: Scaling up capacity and resolution

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.841633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.020732Z digest=sha256:c4f5f24751d9a29fc5bd1247816f355a7ec976b236ce61689a9bd56c808de711

Observation 23e7dc7b-507c-4dab-86b3-6699f2f59d6f · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.826192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.026096Z digest=sha256:48cccd97b1613a350d03c9589c6ace57996e3fc9c23d14479aa8678d5b402c75

Observation f4fec9ca-7df3-44f8-9754-a72c49232f45 · outbound

This paper cites One-step latent-free image generation with pixel mean flows, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation One-step latent-free image generation with pixel mean flows, 2026

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.809553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.031151Z digest=sha256:e7bbe24e67bbb1a718ec1548dc6419fe83ff675cbbb836e39aa4fc84f02f999d

Observation 2d36acc3-76b4-4d96-a9b7-a0d8875d630b · outbound

This paper cites CMI-Bench: A comprehensive benchmark for evaluating music instruction following.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation CMI-Bench: A comprehensive benchmark for evaluating music instruction following

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.794140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.036013Z digest=sha256:c038cb3ce308c31d9f783c1fbe184db5ad2784b66cf93888616c38859c361797

Observation 75e79708-a7b9-4776-8996-3be58b18a4d3 · outbound

This paper cites Pnp-flow: Plug-and-play image restoration with flow matching.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pnp-flow: Plug-and-play image restoration with flow matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.778168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.040895Z digest=sha256:92842a96f3776911a1904058ec932fd4206f3391f1cfccaa242773221c8ebd69

Observation 85c8f97e-4f54-4634-8f48-c8cedf74184f · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.762821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.045834Z digest=sha256:18012040af688a7cc91862acb353edec59f1e2fd0c783181c3e6033105f8eee2

Observation 134e1ddc-edaa-47a2-b3c8-48fd902bb68c · outbound

This paper cites Polyffusion: A diffusion model for polyphonic score generation with internal and external controls.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Polyffusion: A diffusion model for polyphonic score generation with internal and external controls

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.747648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.050886Z digest=sha256:2f11b74443f5dc3dde6a542b35e05890083db8a56a1c75611100a5b419f64df6

Observation c8146f13-b4f8-4ba4-a32e-9f4b3cf33c60 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.055751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.055751Z digest=sha256:d75f72634558a1c42b07fc5e99111825c73c8cb86af540b7e4dfec83950f6f43

Observation cc2dee26-7156-484b-99a7-ad1273520a00 · outbound

This paper cites Muckley, Ricky T.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muckley, Ricky T

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.060824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.060824Z digest=sha256:d75ecba41778eb531121d4a87f4e401e0a88fa593582b15e7c59086ccd474a9a

Observation 40c9a460-054b-4def-828c-bc6a56b1a33e · outbound

This paper cites PhD thesis, Columbia University, 2016.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation PhD thesis, Columbia University, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.720976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.065674Z digest=sha256:5cbed3d5c1d76de1de89dbf91c2e714c923ade56e47060a1247c4cef14e3e1dc

Observation 223d6f9e-ba15-4769-9dcd-69218455e90f · outbound

This paper cites Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.703918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.070460Z digest=sha256:14ec1060a6a22b73b4cdfbe5786cdc10f7d020a7f5ace2d150876df4713076d6

Observation 0f15c106-a328-41bd-89bd-081b0e7f38c1 · outbound

This paper cites Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.688676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.075034Z digest=sha256:44e369917c1c4b6aab4d1b355c8a4f2f922640bef0bfd876c24dbfc8f456b894

Observation 6bb44a12-d538-46b4-8001-cd46248a1f9f · outbound

This paper cites Muscriptor: An open model for multi-instrument music transcription, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muscriptor: An open model for multi-instrument music transcription, 2026

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.079711Z digest=sha256:dcfc759a306a9cfa699e88b797aa5b5767d67163b0930cd4c9fb238d9b7532c5

Observation c226ce58-bccb-4284-878e-3145fae4044d · outbound

This paper cites Rick Rubin: The 60 minutes interview.60 Minutes, CBS News.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Rick Rubin: The 60 minutes interview.60 Minutes, CBS News

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.655671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.084436Z digest=sha256:6d2ec8d9c5cd8c4f1d113159939a586556bbda7cef35cee9c37bc9ecbc73dd6a

Observation 2c5f8fec-afaa-4ae4-a232-96c8a90cf810 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.640139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.089482Z digest=sha256:775dd238c7387a9d08c19afdece2c77ff10cd0295d34f23f240732092e4ffe8d

Observation 36cb00f1-254e-4088-8729-14943ecbb7c7 · outbound

This paper cites Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.093962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.093962Z digest=sha256:b1eeaa57ae83327f0f19066a6719ebc3a9f38e588139a03dbcf3ab763cf4db38

Observation ab69694a-1cbf-45f4-b5c5-310b296da0c0 · outbound

This paper cites Music-JEPA: Learning a world model of sound from action, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music-JEPA: Learning a world model of sound from action, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.601766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.104016Z digest=sha256:7a1470287a8a20571ee884b4baa854fa7ee2c4acfe5b67e9da0154818ad6aa93

Observation ef97b942-4dd2-4979-ba95-68a7039b397d · outbound

This paper cites Visreg: Variance-invariance-sketching regularization for jepa training, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Visreg: Variance-invariance-sketching regularization for jepa training, 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.585245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.108727Z digest=sha256:c6863f93974b4c71a15d2f25e2b2e8e7e5ae926c4f59051affca33c9e2d16ba9

Observation d4332ef1-0b75-4283-a70f-c45e98a78200 · outbound

This paper cites MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.410262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.113557Z digest=sha256:7f602d7d94c0f56fe03c01960af311bc9e0d128ec39be06f78d2a839f700a585

Observation 8564d54c-c355-4bf7-ade4-163d2526236d · outbound

This paper cites MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-08T18:42:37.387019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.118857Z digest=sha256:b3ab8f05cc3d48afe5870439c6065e8fe9c66ae1e9b09536d0378a252ab031d6

Observation b90d367e-3164-4d36-a123-753077ee38f0 · outbound

This paper cites ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.310363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.123392Z digest=sha256:ba6a3702a9fc7e33eacceaa82dd5352e4b21099851d33791794d3b3a1a914f2a

Observation d31e8fa0-d1c7-49c7-bad9-438cbb904560 · outbound

This paper cites ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.128449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.128449Z digest=sha256:6cebec90fd2abebfec421a52be0e62cb92514c84cf0e42b1a4d6b7492896bf1c

Observation 9c36c57d-20fb-411d-b5cc-6c78bc1b0432 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion Transformers with Representation Autoencoders

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.133267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.133267Z digest=sha256:fee6b2555ad873a7af6e263ce24f8f098a9bc0eb89f8d5d47d3e52c9f758f554

Observation c4c16f69-7feb-45d3-b797-efdfae11cf75 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-08T18:42:37.099100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.099100Z digest=sha256:f6962f045b1a026343d72a779b52d8817a2203ac772926d4bb924111e24139dd

Pith citing papers

No inbound Pith citation observations are available.