Pith. sign in

Paper Citation Record · LEDGER

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2506.15154.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15154 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:45:57.076238Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T15:40:57.719214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T15:43:32.551173Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10a13a6a-400f-4d3e-80f3-471fe3ae919f · outbound

This paper cites GPT-4 Technical Report.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.837152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.837152Z digest=sha256:a923dc7309c92a01ff05ed9f69690f54c93a103038dc38cb167b6e5da10d2365

Observation 4dbef810-a1f8-4d1b-a73c-bc21669a8075 · outbound

This paper cites MusicLM: Generating Music From Text.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.846792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.846792Z digest=sha256:b7b4478a155edad0f736f8eae9cc944745c1c9ea881ccadcb6ddeb34521b2366

Observation e80cd79b-674b-4b7c-9c95-c90aaf256dec · outbound

This paper cites and Lavie, A.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning and Lavie, A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:58.209579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.853864Z digest=sha256:d8a7b8ba59043be9a6d5e0d55f7b7243dc603095f38158235ebfdf3157d56734

Observation 1e05fde3-fdfc-43fe-9b77-cb09dcd2b9b6 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:58.171723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.861391Z digest=sha256:4119401b31a24e88f5aeedbe1bacb103ff4709dc3d21fccc8de6edf23b7ee261

Observation 01302fff-0fc9-4900-8b4b-b2469ddad949 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.869307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.869307Z digest=sha256:d65c9bacc6ea9fbbd41b849d3829ddf9cdb18c71e271ee1c8ac18809e603f892

Observation 391c0f51-ea8a-4807-b3f6-137a0ce4cbb8 · outbound

This paper cites MIRFLEX: Music Information Retrieval Feature Library for Extraction.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning MIRFLEX: Music Information Retrieval Feature Library for Extraction

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:45:57.477069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.878442Z digest=sha256:60ce54451fa17039ce062eb8bab42d20825d235510d75ae408623525a7467717

Observation 6eec5996-d9f7-4920-8f26-d2eddd11b553 · outbound

This paper cites Qwen2-Audio Technical Report.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Qwen2-Audio Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.886395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.886395Z digest=sha256:77c87ea33419eb6e63d8714070caff6b624a9dd24071158580fe524b169fe6d1

Observation bc7d538c-60ed-4278-bc5e-4d11a1ecc207 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:58.138518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.893047Z digest=sha256:e50b800e49f101386d20f5e7b3163a981db1ee271a4a4535ecdcee6298eeaa00

Observation 86ff009b-f8cf-412e-809f-9d41eddf3a90 · outbound

This paper cites MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.900402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.900402Z digest=sha256:fd9d0aab5b778d5fe8964b1bb5ad542caf6a5a0b7427cd6ca8d38e478cd18a63

Observation 5437857e-0435-4f7e-b07f-2f0e04276508 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:58.096079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.906879Z digest=sha256:ca026aebba4542ec5a26a368d3a9da59d78edbaf41c93f61f4544b0e3de7c913

Observation 52863ebc-80cc-484f-828e-2c7b3f558874 · outbound

This paper cites LLark: A Multimodal Instruction-Following Language Model for Music.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning LLark: A Multimodal Instruction-Following Language Model for Music

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.914108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.914108Z digest=sha256:5e649b4c81809040df362fb05a55fe43d85c295516befdd853cfff2eeaaf1230

Observation 8498e23a-e756-4872-939f-f5ec014c1fa0 · outbound

This paper cites Mistral 7B.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Mistral 7B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.927172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.927172Z digest=sha256:f82232bc481b0532f5cb03b45f0adebee334cde796d2bae38e75dc3b10b56a26

Observation 07a570bb-e9a7-469b-9fb2-f21a5778cf66 · outbound

This paper cites R., and Macha, S.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning R., and Macha, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:58.066868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.935460Z digest=sha256:5da4ccbcc61d9d5f3da9e029418cf3a2f06f96217a6e2ecb2d82e9d482ae4241

Observation 252c1f25-ee90-4d24-8f8b-3bbeb7e18dbf · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:58.022487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.944735Z digest=sha256:f9bd37b1963e081388cdc01d27926dc880f25b3249aea9f4b6564733ec18a730

Observation 99e4e2e5-5d8b-4112-8502-f01b1e1c7a86 · outbound

This paper cites A., Pinkl, C., Perraudin, N., and Wattenhofer, R.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning A., Pinkl, C., Perraudin, N., and Wattenhofer, R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:57.966233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.951431Z digest=sha256:e412cdb83d18747023dbe4ba0322e2ee44cd63c7d475d48555878a6c78741d58

Observation 70bf67ff-e22e-43c2-a744-9441cd1bc5a0 · outbound

This paper cites I., Bay, M., and Downie, J.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning I., Bay, M., and Downie, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:57.934470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.958722Z digest=sha256:4718ade8caf55c76e0c7d770d4bd21fde6ab15c9e3ac432ba9889b6df37a8d58

Observation 7770d045-5339-4e06-8660-f1980f046245 · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.965706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.965706Z digest=sha256:4e33b3ffacecdbad219d4304f6a977f609b3e41c05b5d3315c92b7e0966cfcc6

Observation dfdb8473-7349-4e1d-9212-8c9b67838118 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:56.972481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:56.972481Z digest=sha256:9d8b9f53a0ec11d09331c11922b59da351c8a1ebd2ec7846edbcb0827059f52e

Observation 1146dac3-e389-466b-bd70-ea6c0ed6f4f5 · outbound

This paper cites S., Sun, C., and Shan, Y.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning S., Sun, C., and Shan, Y

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:57.874428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.980761Z digest=sha256:7282be863b18ce3161f029e4f8fe33534bbded09812a22480c92c0d5dabf457b

Observation fc8ab943-664f-4280-856e-677699028217 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:57.847391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.987196Z digest=sha256:fea31de4a02ae5dafab7d15f17e5bd7a36084e51fd5560adfe3f5e135ae0b58b

Observation 10e9c852-e791-4e49-b042-aa0e6beac99c · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:57.817054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:56.996510Z digest=sha256:cd2a04a74c5a3351ba1bd05da92486fc6e8887dba6863aaa59d5488d2fe99f7d

Observation dd87d5b5-8380-412e-9b07-61141754383f · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:57.784413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.011318Z digest=sha256:d653076474a2e86dea28065e7d17f3ae66a0a5c32f4ef9fcec10c88ceabfed31

Observation 934fcba2-be25-4020-b5dc-b0c11459c6a4 · outbound

This paper cites C., Davies, M.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning C., Davies, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:57.747137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.018993Z digest=sha256:13ecaccdde8477639a9f89e258713910f75f3d1881a7d61abcc9b425496c5299

Observation d26b14df-aafa-4c90-b478-6f26a0d16c8c · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:57.026932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:57.026932Z digest=sha256:36514a60cf9b2e07e0eadda6c7f6473806a36cb497966f0238b8ad8f678d4e55

Observation 56fb8458-d3e8-4f10-9cc9-2b8a0ff86c6f · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:45:57.672566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.033183Z digest=sha256:ca8306a5ba84af2ef291e0bdbe198496ba1b54bb541b063a9ede719640799250

Observation 13dc23d5-3091-45e9-908e-22dfd0c91dfe · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:57.647730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.040029Z digest=sha256:6a611a723c1c8bcbb85df19c211ccf8309f92a6210acc95e8dae0d448b9f7dd3

Observation e410f084-a597-47a7-b7b0-3d4ddf1e4669 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:57.046473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:57.046473Z digest=sha256:58c4325beff5629e4b60c062107ccd2a184f5dd30beef1566e997da1f3d190bc

Observation 98814489-b7b2-4623-ba7f-6274e204dc38 · outbound

This paper cites Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:45:57.223689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.053415Z digest=sha256:dd9969cb14fc669dcff6d03b0a5163036ba0b604e90c3e767065107a756df414

Observation 07792add-b5fa-4d37-84ad-274994fab8c0 · outbound

This paper cites an unresolved cited work.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:45:57.621830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T19:45:57.062067Z digest=sha256:1d09a00a865f7af71f8ea8ad31e398a3b9ce5e2797904166610e85398752dee9

Observation 4257b6fe-2671-4320-833a-02fb32ad96d5 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:57.070944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:57.070944Z digest=sha256:5a156943ae3349b2991771af2d65744eae52de27742beab791a805d01bb72a88

Observation 7172946b-9f16-451e-9890-ebc3cc9aa54e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning BERTScore: Evaluating Text Generation with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:57.076238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:57.076238Z digest=sha256:96fb714f37b73bf8f318dc165222ba4a92ceca03b12f475db0310ea0a8bdaae2

Pith citing papers

Observation e9dcf0b2-d6fc-4662-b245-5dde3f7bda8a · inbound

MERIT: Learning Disentangled Music Representations for Audio Similarity cites this paper.

MERIT: Learning Disentangled Music Representations for Audio Similarity SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:43:32.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T15:40:57.719214Z digest=sha256:4e204639db21f153d50a6ac7d0f7b56611b242c4ba81cb7514c9cccf574de358