Pith. sign in

Paper Citation Record · LEDGER

Foundation Models for Music: A Survey

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2408.14340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.14340 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:16:09.562425Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.406159Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8ca6466e-d401-4f49-b820-ed9d4cceb32e · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Foundation Models for Music: A Survey

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.738757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.738757Z digest=sha256:555a812665f45aef80c8775087e34a7e3b777827f46b3bb9944b94352767da90

Observation 11299f15-9055-4e64-8087-03e957ee6c7a · inbound

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview cites this paper.

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview Foundation Models for Music: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:15:07.222130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:15:07.222130Z digest=sha256:03adada58d32b0c18cb8f970d9b33fe014bf3aeca3abdb6ea4a15c1a54a2e120

Observation c785a807-3bb9-4ea4-a57d-e0b8256c93e0 · inbound

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection cites this paper.

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection Foundation Models for Music: A Survey

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:32:10.244241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:29:43.384422Z digest=sha256:34efa10d8da6832cbcf3286189e332d473d1ad98ba5a0ecde75fa59f88597f00

Observation 1116232f-c27e-4ad3-8a5b-6fe30c146e7d · inbound

Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models cites this paper.

Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models Foundation Models for Music: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:09.562425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:09.562425Z digest=sha256:07dee87bed325f235a51938d7454e44fbc2efd788227da6e102c0cd7f0b70fa4

Observation b8cfe6f9-d930-4ac2-9611-384ebda1c7c3 · inbound

Semantic-Aware Interpretable Multimodal Music Auto-Tagging cites this paper.

Semantic-Aware Interpretable Multimodal Music Auto-Tagging Foundation Models for Music: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:51.795095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:51.795095Z digest=sha256:c1e9d31790757b5c0726c956fc07d25d4b47b7179170cf15501026fa3539099d

Observation 98e6c135-305a-4f67-b27b-19101663bd6b · inbound

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following cites this paper.

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following Foundation Models for Music: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:47.112119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:47.112119Z digest=sha256:cb746b722d9016b1b1c34b1b17ec546af7b6ad3c24dec9caaf90516937667e2c

Observation ab151a50-da97-4c30-99f4-2d050f9148e2 · inbound

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models cites this paper.

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models Foundation Models for Music: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:34.878811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:58:34.878811Z digest=sha256:a4e91b392348b8ad15a2e17fb8e65453fd60a246b505387b700cfc8484dd5c19

Observation 80a30663-649e-451b-a049-5d5ccc8e54fa · inbound

Universal Music Representations? Evaluating Foundation Models on World Music Corpora cites this paper.

Universal Music Representations? Evaluating Foundation Models on World Music Corpora Foundation Models for Music: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.603968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.603968Z digest=sha256:4893259b75d368505f15ed834b07b7a1495836c39aa288c985cff5a6c941513c

Observation 1e95bd5b-1244-4673-969d-1687610ede26 · inbound

CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning cites this paper.

CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning Foundation Models for Music: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:07.347609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:06:07.347609Z digest=sha256:542d0b16f4a52012fb4c38d9e6043e85fee76eca97529cdd6044d2c1bfea2e6f

Observation 865fbb87-2a6a-4007-86df-e97febd206a5 · inbound

Emergent musical properties of a transformer under contrastive self-supervised learning cites this paper.

Emergent musical properties of a transformer under contrastive self-supervised learning Foundation Models for Music: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:35:04.720630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:35:04.720630Z digest=sha256:454f591ca6cf4231a41e331615a19fa9db2fb689ae9f2d17fa4e6bf732207b21

Observation bc65feb1-d30d-446f-8a8a-671b7c708d8a · inbound

Workflow-Based Evaluation of Music Generation Systems cites this paper.

Workflow-Based Evaluation of Music Generation Systems Foundation Models for Music: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:03.070056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:46:03.070056Z digest=sha256:e7825ea1f15f65369f9bbcc54349d4627b05ae40793468e9f608928491e42ff1

Observation 452baee2-4d0c-44ca-9630-7ec2c95cb14b · inbound

Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis cites this paper.

Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis Foundation Models for Music: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:59.084669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:59.084669Z digest=sha256:0b134a726fb47f4d5e8383626915a45cf1ff4f250aa436b607194ba09bb1b8ff

Observation 1371957a-6e91-46cc-b545-60de1bda4fd7 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Foundation Models for Music: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:53.530166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:53.530166Z digest=sha256:0e30965b9e62eae9aff4a6fbf8e3cae2e5ea3cba2f8afac26ef85cd0eb5eb56d

Observation 5ca136b1-a901-4110-838c-4f6a64f914d7 · inbound

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets cites this paper.

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets Foundation Models for Music: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:03.604395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:03.604395Z digest=sha256:2a975d0527f81086eab1b8c1182f1efff0a21ffbc831ffbcc7dab0e789585b02

Observation 3a651c43-f283-4910-a03e-60b151357706 · inbound

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models cites this paper.

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models Foundation Models for Music: A Survey

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.678172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T16:59:25.973854Z digest=sha256:a43f617c3310d5fb6ab42aeec899f9f93361f51f98894f35fc6f638a32127619

Observation 3a86c0b1-3183-4ff6-907f-59d7f68a75db · inbound

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding cites this paper.

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding Foundation Models for Music: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:28.331857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T17:50:26.448655Z digest=sha256:a75bed00a38be7f0434783d72a0c2d2e3464a0e760a188acec69a9127e3c9b5d

Observation 201ae5f6-fe66-4c39-8ed8-5b85629d96b8 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Foundation Models for Music: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.747170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:1bf4f53380882c027c3baa7e053ffd5b5a02ccd8400abcc87102da30d2a8548f

Observation 4763ac22-aa58-4b72-9498-865d0fa61f86 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Foundation Models for Music: A Survey

Reference 255

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.407499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:d4dac1f9e184457016508f4156091982c249cc0e2ff25c2260752596a9278859

Observation 96edaee9-eba1-4775-8e1c-3ee2472fc6eb · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Foundation Models for Music: A Survey

Reference 255

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:1288141ed7c6318216e54bb0bc804c0fa706fa32e64506a2addb5a25eefaad5e

Observation b741c712-5c9b-4f6d-9b32-4042a5cf06a8 · inbound

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling cites this paper.

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling Foundation Models for Music: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:10.657478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:48:10.657478Z digest=sha256:09e247aeee6606b70d27babb9ff9ea31c16c389d603e9398dee2875b6cb19f82

Observation c6505953-5bc5-4576-994e-16060423774c · inbound

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation cites this paper.

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation Foundation Models for Music: A Survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:47:17.683146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:47:17.683146Z digest=sha256:3f9876e5bf431fce6ed6799081cbbe626c655bdc641985edc053c8e83b0e78a1

Observation 8d9bb36b-11e2-43a5-a206-e937dc699946 · inbound

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset cites this paper.

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Foundation Models for Music: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:14.252654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:14.252654Z digest=sha256:cf105d78b08af815e0ee5986c72aca4cb6c62bafadd1759790a14722adf60e3a

Observation 1a5e0262-0caf-43b9-9ee5-6ca05774ce64 · inbound

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset cites this paper.

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Foundation Models for Music: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:44.814420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:44.814420Z digest=sha256:ec9c29662325a794a8a7cc04c0e9f313df834acfeb36d69ef664f1f6bbf2e246

Observation d01292dc-d09a-44ce-8df3-87517abdfd2f · inbound

Where Does AI Innovation Go? Measuring Research Attention Imbalance in AI Music cites this paper.

Where Does AI Innovation Go? Measuring Research Attention Imbalance in AI Music Foundation Models for Music: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:44:12.741007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:44:12.741007Z digest=sha256:3bedc7071b5e99e2b9aec8f2b130768f30803d2b2e33cb3357253d528d401a87