Pith. sign in

Paper Citation Record · LEDGER

Foundation Models for Music: A Survey

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2408.14340.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.14340 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:51.795095Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.406159Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c785a807-3bb9-4ea4-a57d-e0b8256c93e0 · inbound

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection cites this paper.

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection Foundation Models for Music: A Survey

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:32:10.244241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:29:43.384422Z digest=sha256:42a70ddafa6b3a2d396bf947b164952d644728af2b0b205a07b5818c3d98781f

Observation b8cfe6f9-d930-4ac2-9611-384ebda1c7c3 · inbound

Semantic-Aware Interpretable Multimodal Music Auto-Tagging cites this paper.

Semantic-Aware Interpretable Multimodal Music Auto-Tagging Foundation Models for Music: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:51.795095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:51.795095Z digest=sha256:1f93f6c09fc0e6fa7f91f77346b28e9b549854823180c34c5eb583109ca0c4a6

Observation 98e6c135-305a-4f67-b27b-19101663bd6b · inbound

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following cites this paper.

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following Foundation Models for Music: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:47.112119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:47.112119Z digest=sha256:cde4fab2ae011bdb26fdc1e331ca387f1b06d62d0912054a1b6961c0ac4e1b08

Observation ab151a50-da97-4c30-99f4-2d050f9148e2 · inbound

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models cites this paper.

Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models Foundation Models for Music: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:34.878811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:58:34.878811Z digest=sha256:fd213c1f9da94da0bc5c8468d351d4ecefb045c11b443c550c827eddf8596ba2

Observation 80a30663-649e-451b-a049-5d5ccc8e54fa · inbound

Universal Music Representations? Evaluating Foundation Models on World Music Corpora cites this paper.

Universal Music Representations? Evaluating Foundation Models on World Music Corpora Foundation Models for Music: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.603968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.603968Z digest=sha256:2aced2453e88e3e37f0cea6fa50a270eef87c5a7309ba7567f0dd30ed0068982

Observation 865fbb87-2a6a-4007-86df-e97febd206a5 · inbound

Emergent musical properties of a transformer under contrastive self-supervised learning cites this paper.

Emergent musical properties of a transformer under contrastive self-supervised learning Foundation Models for Music: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:35:04.720630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:35:04.720630Z digest=sha256:1ee16b896d59d68d709115022bbca62941e24a638dd41a172d23bb531bd98aa5

Observation bc65feb1-d30d-446f-8a8a-671b7c708d8a · inbound

Workflow-Based Evaluation of Music Generation Systems cites this paper.

Workflow-Based Evaluation of Music Generation Systems Foundation Models for Music: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:03.070056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:46:03.070056Z digest=sha256:a41d840e2c3952ac512b0260d56d08c20fadaaec94a8ef995a75e86e823bf83a

Observation 452baee2-4d0c-44ca-9630-7ec2c95cb14b · inbound

Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis cites this paper.

Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis Foundation Models for Music: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:59.084669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:59.084669Z digest=sha256:8e2c63be76c95292535aabba0a58f93ff31afaa5df42d2b1f2f39a16f039f56b

Observation 1371957a-6e91-46cc-b545-60de1bda4fd7 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Foundation Models for Music: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:53.530166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:53.530166Z digest=sha256:f973680d0ede2bbc425f9e62b2c136f2765da4d21a08e69ebacecf7afc4a2b89

Observation 5ca136b1-a901-4110-838c-4f6a64f914d7 · inbound

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets cites this paper.

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets Foundation Models for Music: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:03.604395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:03.604395Z digest=sha256:05cd8402a321423f0164de5084091a7ef41bf5305cafdfd96ea4364b0c43d721

Observation 3a651c43-f283-4910-a03e-60b151357706 · inbound

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models cites this paper.

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models Foundation Models for Music: A Survey

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.678172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:59:25.973854Z digest=sha256:2bf0f8e944a5d1a5f0c05874aa3ec34864c867cfa360fc58ae8f7e3ceafdd8f8

Observation 3a86c0b1-3183-4ff6-907f-59d7f68a75db · inbound

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding cites this paper.

Can LLMs understand LilyPond? A benchmark for symbolic music generation and understanding Foundation Models for Music: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:28.331857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:50:26.448655Z digest=sha256:3c4ffe039500fc662de248308ae87597ca3838f00c9a4b711bef39a1f45120fa

Observation 201ae5f6-fe66-4c39-8ed8-5b85629d96b8 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Foundation Models for Music: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.747170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:8f1637025634932a7111bf7373623795551bc647204c42dd1434ea8eaf3a0a0d

Observation 4763ac22-aa58-4b72-9498-865d0fa61f86 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Foundation Models for Music: A Survey

Reference 255

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.407499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:9b4dd78eb33c699c82ba0544cdb268f1c4284e0ffeea48b838375d07cb07696a

Observation 96edaee9-eba1-4775-8e1c-3ee2472fc6eb · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Foundation Models for Music: A Survey

Reference 255

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:5133bebaf7a3b7df85390eb92864afc6b5be090c1ad6229cadc9f19623105a3d

Observation b741c712-5c9b-4f6d-9b32-4042a5cf06a8 · inbound

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling cites this paper.

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling Foundation Models for Music: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:48:10.657478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:48:10.657478Z digest=sha256:896c2cafe20f5f9f8250c89bb24dd78261faa33c155d3936737a5b0affaf869d

Observation 8d9bb36b-11e2-43a5-a206-e937dc699946 · inbound

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset cites this paper.

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset Foundation Models for Music: A Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:14.252654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:14.252654Z digest=sha256:79740eface527eed514d35addb94644c857243892be78968c30df772246b5a88