Pith. sign in

Paper Citation Record · LEDGER

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2408.01337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.01337 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:23:49.282873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T19:35:32.952259Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e14010f1-203f-4f90-8bda-775a670a76eb · inbound

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview cites this paper.

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-12T05:15:07.785245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:15:07.785245Z digest=sha256:49b986befe10e8dfeb37c5c2300a8181bfd390d4935feaefe23911b5790c0022

Observation 93468098-183b-4a30-8a34-31c9cbddcaf2 · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.491889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.491889Z digest=sha256:fd863753cabba81181e9c0fd3b14eb1db22076684c73601506b7be6fb70ae0be

Observation 908d223b-5e37-44d5-898c-7db7a43f6ac7 · inbound

Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning cites this paper.

Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:23:49.282873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:23:49.282873Z digest=sha256:5af512ee0f686bba18bdd644e8628f794013ca1c42f09d9bc9b2c4d48e1b45f5

Observation f254073d-a561-4fe9-9aac-1df1e57fed69 · inbound

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs cites this paper.

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:12:22.024856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T14:11:52.011626Z digest=sha256:c4c47a38b2b2a7669836ea7d97af08b37e316c0c10868ab090a6b320f5e3dd5d

Observation 07e61905-a2e2-4d61-9fec-6569908441fa · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:00.052781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:00.052781Z digest=sha256:84936dd272c3dec7ce6d40782efd696305c904bcc6ec5eedeee27e54bc9b117b

Observation 1f3baa0f-5bc1-405c-9df1-c3e2106b41fa · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:42:45.271465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:2c6b24cc4d4dd4d1b47c50b283900060f99082b55ab5126f76c5e72294d22fa5

Observation d6e8ad95-9497-4d22-af54-95c16feefb10 · inbound

Topology-Aware Layer Pruning for Large Vision-Language Models cites this paper.

Topology-Aware Layer Pruning for Large Vision-Language Models MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:59.448862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:21:04.009413Z digest=sha256:29b840e95a67a0fd82df9ebe540d4730fe71adec981e20bcce839c88a7fc8ad9

Observation 3efb89ea-0c58-4d6a-92bf-91e25f9035ad · inbound

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs cites this paper.

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.338604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:58:51.272971Z digest=sha256:a2612659dbaf0c408cc00571543064b0a10f2c91d7aa0e356ad2f3d8767a72e6

Observation cb3364aa-a6dc-4f9b-aae3-bfbb1ba07bcf · inbound

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models cites this paper.

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:43.020881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T03:31:42.611472Z digest=sha256:cb553f1f7a16e45e6ee2600878f82f7de8276e3958e6da44eea4be744e107431

Observation 147d1cc8-2bf6-49f5-89c7-fe0bd61e0003 · inbound

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music cites this paper.

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:35:32.954819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T19:28:58.523932Z digest=sha256:86d95085388bbce9b0c5cad01f2b03644ffb21ab2b65f9f996101344ce108c61