Pith. sign in

Paper Citation Record · LEDGER

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04378 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:42:37.133267Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved15
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a8874bf-1f7d-4649-9a84-57fa7402b568 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Self-supervised learning from images with a joint- embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.080911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.912135Z digest=sha256:7d1b9746b872e9a10d15b7bd0cfe1e9fe31d0d5ab5b06782b8ba7313f5d38d2f

Observation 87ab2416-fe04-4adf-a606-c7d2920dc467 · outbound

This paper cites LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.063959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.918353Z digest=sha256:c3990e395c18b0161bd40780d4a47b830b7a8ed42b7080b22e024cde6e7f72cb

Observation f0ce2832-a2bc-4771-a2d8-30745e4a1e7d · outbound

This paper cites Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.048869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.923728Z digest=sha256:9550ee8a1bde8c5f62b7896c0d455464898b2f2a30dd37da9b488bd0fe278618

Observation 2cf73a05-3977-49d2-80f3-fc2571c3f015 · outbound

This paper cites Zapata, and Xavier Serra.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Zapata, and Xavier Serra

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.032439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.929545Z digest=sha256:bb1745c52af02d0cc4be54bf13576bbefa95cc894be06ae96b7fc8c4d5338bba

Observation 8ab0eeaf-e0d4-4b3c-90ec-aca0412ce320 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Emerging properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.017269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.934794Z digest=sha256:4dc9acf6b98021dec926cf39b52b33b3ed6fefb36f172fb34f05a0bce4035f48

Observation 59777f34-56c3-4d60-b615-a86b8be63621 · outbound

This paper cites Codified audio language modeling learns useful representations for music information retrieval.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Codified audio language modeling learns useful representations for music information retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.001701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.939949Z digest=sha256:706bac52bd1f1baceeea5edd21c25a80aa74571f001310938e0707a77d33047c

Observation 5c3baffb-2a94-4a1e-8893-96df08d80ae6 · outbound

This paper cites Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.568953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.945515Z digest=sha256:b0dd708f1254cce0f620d8ee6068b0ac3e74e753422481033cb6717687cecf47

Observation 8b25a6df-49bf-4ad7-9374-fec2b95402df · outbound

This paper cites Pixelflow: Pixel-space generative models with flow, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pixelflow: Pixel-space generative models with flow, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.950691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.950691Z digest=sha256:6178496b0b50fdfebcd468e5bb9818f50666dc2a5253fede36a5ffcb1033e68e

Observation b76b3215-d6ea-4ee0-a166-f468f897a65c · outbound

This paper cites Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.955570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.955570Z digest=sha256:3afd9a6a851634ac75b8a3a976bf91da74494735de956c4edfcdb55f9d12a551

Observation 0872e412-f945-466a-b80e-e54c31e0fc54 · outbound

This paper cites Diffusion models beat GANs on image synthesis.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion models beat GANs on image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.976481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.960678Z digest=sha256:71dc01f32d23094035d0179752ea6ee18c53b0aa79bb1be2d31349796b0c8dc1

Observation 99673cb7-a190-4952-9601-3eaf6166e80b · outbound

This paper cites Generative modelling in latent space, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Generative modelling in latent space, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.962049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.965421Z digest=sha256:ccae5642aa1f78b3485e9d6865c871e1b093fcd591550c93404174d20c8dcdda

Observation ee53d3be-8918-411f-a66a-0973ca62549c · outbound

This paper cites Hawley, and Jordi Pons.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Hawley, and Jordi Pons

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.946586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.970418Z digest=sha256:bca6e27c11ab04616ec71f7518d0dd02f75dd4f318a57a9d651b6588abbc1fda

Observation ec1dd9e5-264d-4417-b06f-3f8a9553fca9 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Learning and Leveraging World Models in Visual Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.975241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.975241Z digest=sha256:60e6af4ac9f0902266478ecfa12dbd224d8effd27b23e3dfa8f119fa4987cd05

Observation 823ffb86-9a19-46ce-8fac-aed9cf3d4be2 · outbound

This paper cites World Models.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.980422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.980422Z digest=sha256:03ef7f921a25702415acebc88d281f7049e250fe9d50cc1c9eb8c7b6cd411e05

Observation 19e4507f-3ca4-4b83-985a-1df8eaaef8b7 · outbound

This paper cites Using a joint-embedding predictive architecture for symbolic music understanding.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Using a joint-embedding predictive architecture for symbolic music understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.929987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.985406Z digest=sha256:e13b10d0463073416da57c8d52503ee35c2b1accc1af255e0522e01e0c7471e0

Observation 2a4bf12f-9026-413b-a7da-0e8174363d51 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.914256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.990376Z digest=sha256:0fbc886284d9ce5f1d6423df08e4652deb4e813089e31d9b2f4b2af90dd39373

Observation dcf967d7-5d7b-4c45-9c3f-0b29962cccbf · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.898382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:36.995429Z digest=sha256:ebdaffc2e37acf9fd845b38072d012bacb7ed47ad27339de4771fdd9bfc60455

Observation 58bff056-ca19-4e5f-94e5-a646e70108bd · outbound

This paper cites EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.882781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.000539Z digest=sha256:d9b00206414e66656c778bc29df5a30401b36b59eae81f962bf316e706b3207a

Observation 776a0ebf-080a-4d0b-bcc9-900b302cffdc · outbound

This paper cites How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.005741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.005741Z digest=sha256:03d5eae0b7b0360db69896e599a876d9236c902a2b96329f5e6a6e994f2bc96e

Observation cb009084-0471-46a6-918f-b6819232b411 · outbound

This paper cites A path towards autonomous machine intelligence.OpenReview preprint, 2022.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation A path towards autonomous machine intelligence.OpenReview preprint, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.867340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.010781Z digest=sha256:9cf431de28430438ee9aa2312bb0ea8f8f2f7e6de548d22fc876aff90143a31d

Observation df7df0f6-71f8-4b7f-92f3-da69e35974af · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.015774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.015774Z digest=sha256:ce126452d8c0f9a5af09431c944b1f1bb25e8d1e947b07dfe916626f50b1d39f

Observation ed68fb50-386a-41f0-9356-2fe0385ab834 · outbound

This paper cites Swin transformer V2: Scaling up capacity and resolution.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer V2: Scaling up capacity and resolution

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.841633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.020732Z digest=sha256:9de08dcb703d1a508fbecca110d8d7259b147adf80a4026db77cf00b0269cd54

Observation 23e7dc7b-507c-4dab-86b3-6699f2f59d6f · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.826192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.026096Z digest=sha256:7d60e3ec593a16021d9ff7c98908ac65c08ebf9c367f2b7e51379ba6f9d631e1

Observation f4fec9ca-7df3-44f8-9754-a72c49232f45 · outbound

This paper cites One-step latent-free image generation with pixel mean flows, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation One-step latent-free image generation with pixel mean flows, 2026

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.809553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.031151Z digest=sha256:582d3f2b21582e6830552c8581add6322e8c930caa3f2f82f2beba3e0646590b

Observation 2d36acc3-76b4-4d96-a9b7-a0d8875d630b · outbound

This paper cites CMI-Bench: A comprehensive benchmark for evaluating music instruction following.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation CMI-Bench: A comprehensive benchmark for evaluating music instruction following

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.794140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.036013Z digest=sha256:6330ef167bdef39b0f9ce72bd5a2d2d51e28c4d39adff55d9655f351926dbb5f

Observation 75e79708-a7b9-4776-8996-3be58b18a4d3 · outbound

This paper cites Pnp-flow: Plug-and-play image restoration with flow matching.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pnp-flow: Plug-and-play image restoration with flow matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.778168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.040895Z digest=sha256:382879df43b1277d017168c564ab7dded2686515bb34eaec35aa095233c1bd60

Observation 85c8f97e-4f54-4634-8f48-c8cedf74184f · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.762821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.045834Z digest=sha256:1f28a11d4ccc2b9e725d3b7e8175d6fd7058856d7750ad971fec66c332f12bde

Observation 134e1ddc-edaa-47a2-b3c8-48fd902bb68c · outbound

This paper cites Polyffusion: A diffusion model for polyphonic score generation with internal and external controls.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Polyffusion: A diffusion model for polyphonic score generation with internal and external controls

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.747648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.050886Z digest=sha256:50a13d450db8d0666dfe89598ee4b0eecd5d69c27474888a479d4bdb89a100de

Observation c8146f13-b4f8-4ba4-a32e-9f4b3cf33c60 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.055751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.055751Z digest=sha256:e59918ce334712967a53706f8642243a0b4f76ad1a18c4a3b25cf1b4207ebcbd

Observation cc2dee26-7156-484b-99a7-ad1273520a00 · outbound

This paper cites Muckley, Ricky T.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muckley, Ricky T

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.060824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.060824Z digest=sha256:ad5d3e87c131eaa8c64ca8025a29b32af01e327027dcdec5e5b8dbecc11adcf1

Observation 40c9a460-054b-4def-828c-bc6a56b1a33e · outbound

This paper cites PhD thesis, Columbia University, 2016.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation PhD thesis, Columbia University, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.720976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.065674Z digest=sha256:3a9b9c1f5d265cb85dee6a6c83731db1fbcff1ba876e475990e25b3c9e2d110a

Observation 223d6f9e-ba15-4769-9dcd-69218455e90f · outbound

This paper cites Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.703918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.070460Z digest=sha256:3ebc80106b1af2184e3adf4e8946d56f1d86b401202bb04663ff603e0253a3f5

Observation 0f15c106-a328-41bd-89bd-081b0e7f38c1 · outbound

This paper cites Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.688676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.075034Z digest=sha256:d9839bffd8d70f4e6a0f1402a93785003f7002abde471eb91af5ca03c1835b07

Observation 6bb44a12-d538-46b4-8001-cd46248a1f9f · outbound

This paper cites Muscriptor: An open model for multi-instrument music transcription, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muscriptor: An open model for multi-instrument music transcription, 2026

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.079711Z digest=sha256:95b092eb3965362c130de55fe7c39956c007f65fafd9d046ea07cc069c83d817

Observation c226ce58-bccb-4284-878e-3145fae4044d · outbound

This paper cites Rick Rubin: The 60 minutes interview.60 Minutes, CBS News.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Rick Rubin: The 60 minutes interview.60 Minutes, CBS News

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.655671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.084436Z digest=sha256:0d6f5029d27b9abb697fa40110e46f12ea0965d925eba4c9036f250f4c8a779c

Observation 2c5f8fec-afaa-4ae4-a232-96c8a90cf810 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.640139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.089482Z digest=sha256:a3b158a75a6ab371c02e7f9e2c0232851b81d91525ad29ad900df64b100bf4fd

Observation 36cb00f1-254e-4088-8729-14943ecbb7c7 · outbound

This paper cites Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.093962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.093962Z digest=sha256:623d023938b4480a7056c6359b3ff1ea66e47b81015a483cca02a5ccc7544808

Observation ab69694a-1cbf-45f4-b5c5-310b296da0c0 · outbound

This paper cites Music-JEPA: Learning a world model of sound from action, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music-JEPA: Learning a world model of sound from action, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.601766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.104016Z digest=sha256:2a95f67c9b52329873b955d292e01dfa509638c681d4fe6c02ac161b623bcb33

Observation ef97b942-4dd2-4979-ba95-68a7039b397d · outbound

This paper cites Visreg: Variance-invariance-sketching regularization for jepa training, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Visreg: Variance-invariance-sketching regularization for jepa training, 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.585245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.108727Z digest=sha256:7546304bd06ec6b86e17a0d7f1520e5d6dadc0f71425ffca2b205596d6e30be5

Observation d4332ef1-0b75-4283-a70f-c45e98a78200 · outbound

This paper cites MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.410262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.113557Z digest=sha256:70aa353988e668ba6fe2a9da31ed5a061bea7dfce9cbb44fd066b03cc7feb380

Observation 8564d54c-c355-4bf7-ade4-163d2526236d · outbound

This paper cites MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-08T18:42:37.387019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.118857Z digest=sha256:cb2c6a70d31bc3bcf04d13cc9af23058c7553dbfb24dfa605d094d14a2a5477c

Observation b90d367e-3164-4d36-a123-753077ee38f0 · outbound

This paper cites ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.310363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T18:42:37.123392Z digest=sha256:9b95c6fef496e87e58a7dda773318ad524c65567c7070be68b185f02829a4205

Observation d31e8fa0-d1c7-49c7-bad9-438cbb904560 · outbound

This paper cites ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.128449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.128449Z digest=sha256:9cb143488c86c16f0f2e40630b352dae6a99d93c1f26320d9c1ed9c287b9777c

Observation 9c36c57d-20fb-411d-b5cc-6c78bc1b0432 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion Transformers with Representation Autoencoders

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.133267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.133267Z digest=sha256:bd396630517a01ae192289ab52f65413cb5dd9608d38f27eb4117a5d5729f116

Observation c4c16f69-7feb-45d3-b797-efdfae11cf75 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-08T18:42:37.099100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.099100Z digest=sha256:667b2df0c8db310a557da35c59556d5feb92c96baa915927a8660f7720799f16

Pith citing papers

No inbound Pith citation observations are available.