Pith. sign in

Paper Citation Record · LEDGER

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2506.23009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23009 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:57:08.122917Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T02:12:58.120295Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T02:17:46.305033Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5cc43c0-2f24-47c8-9aa3-c5a96aeb3049 · outbound

This paper cites Acrobat AI Assistant, 2024.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Acrobat AI Assistant, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:14.300132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.029002Z digest=sha256:4824d265b1142cb582d21597d36318a25b388853c90a83a96bf46c8d14e7582f

Observation 940c5739-7079-4665-a70c-b0ff2e4571b7 · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.094544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.094544Z digest=sha256:c0fbe377cab123a114f6ebe87c45f5346e6b4311abe7a3bffb90a5e8465e1c6b

Observation d0e24319-069e-4c41-8abe-6d1fc94fce4e · outbound

This paper cites The basics of reading music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The basics of reading music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:14.081606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.170985Z digest=sha256:36ee37b413192a949f62cb39f9e6e8b49872f4652f12dac8a0afb3538e1e2923

Observation ccb2d247-30ec-4ed1-99b3-f219450ede58 · outbound

This paper cites Reading sheet music facilitates sensorimotor mu- desynchronization in musicians.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Reading sheet music facilitates sensorimotor mu- desynchronization in musicians

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.876029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.243569Z digest=sha256:558a22fa11726a50a8b5a2106f0d77b53ecb8df6924b6284ecb9ff77218cf215

Observation 444834fa-8cde-4875-89f3-9e5004298c0b · outbound

This paper cites Optical Music Recognition: State of the Art and Major Challenges.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical Music Recognition: State of the Art and Major Challenges

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:57:08.894312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.307574Z digest=sha256:04fdebc36cefa5af12fc2694b3e69e97e09847e627f0dcc701810242c2a4878e

Observation b20ed1a5-0a3d-49c0-87c4-c2fcf0ba7def · outbound

This paper cites Understanding optical music recognition.ACM Computing Surveys (CSUR), 53(4):1–35, 2020.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Understanding optical music recognition.ACM Computing Surveys (CSUR), 53(4):1–35, 2020

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.647645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.381957Z digest=sha256:07e599582b21b61cc22953c6c05a448ffa271ce791a504382edfdee50ef9f0d4

Observation 0ad181d0-58c5-4a4d-a5c6-6942d6910ebe · outbound

This paper cites Optical music recognition: state-of-the-art and open issues.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition: state-of-the-art and open issues

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.437024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.441555Z digest=sha256:f40c0ab5123f68e07863e6ba7b8cbe6fb09a4cc8ad76674867830f12e29c1a3c

Observation 3cedc974-803e-4108-9133-15be13d1487a · outbound

This paper cites The challenge of opti- cal music recognition.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The challenge of opti- cal music recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.262422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.530806Z digest=sha256:f176372597574dd3063591fb8b140a7be04d4c6276b3cc7cfc3d87bbf6bebf26

Observation a453cfc9-af66-446f-9777-0952ee80cedc · outbound

This paper cites Optical music recognition using pro- jections.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition using pro- jections

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:13.096993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.612404Z digest=sha256:66ea3313e5ce6c807d26690072d476afa300316150b10e79a275f0d3afd30463

Observation 60d3e49a-7c68-4e09-9e4d-418e773af6dc · outbound

This paper cites Gui agents: A survey.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Gui agents: A survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.704090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.704090Z digest=sha256:040e0a682f75c5eb2b5f411e6ffa9225dccb93c9347f51e56f1c67f29c27c486

Observation ea595109-5701-468c-b1e6-3ea06fd2ff0f · outbound

This paper cites Natural language understand- ing and inference with mllm in visual question answer- ing: A survey.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Natural language understand- ing and inference with mllm in visual question answer- ing: A survey

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.868582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.796792Z digest=sha256:e233b638e1582cd708f02f26bd2b40f8e95b61ced39bc671bb955c07bbad5c71

Observation 5a9369b0-1cf4-4a8c-be2d-cce557455e5b · outbound

This paper cites Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.665173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:03.801270Z digest=sha256:f5ace9d7740edfafd6a1867d93b35ede982fc1ec4ac93fbe331a2e8644f7bd0b

Observation 320ef79e-a083-4b28-85bf-72cc75c4383a · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:03.931871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:03.931871Z digest=sha256:44268f53e55c57b620fe2b096dbdc40f8183799398a50a20a8293b1f08498803

Observation aa2550e8-38f8-4dd0-8235-07e20b3ea904 · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision- language benchmark.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision- language benchmark

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.448383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.036358Z digest=sha256:ccd369e152fc0a101c4826adffdb3c4f2e704f7ebe2c0c3abdc4e8a2ddd07dca

Observation c8f6ce39-1320-4708-a18d-31ee29f69ce0 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PP-OCR: A Practical Ultra Lightweight OCR System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:04.136009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:04.136009Z digest=sha256:4d6c0cc8b248ecd8d7f2dcdd9e3eaed01b26b94dd7ca66e502d6eaf37f6ddc44

Observation 8a612c9c-cc42-4bb0-a4a4-3444743860b3 · outbound

This paper cites Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.295573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.213329Z digest=sha256:9de246c647523430f5997bbca35e270196e77572826b9d76e28112b4bd679e00

Observation 74b8b8f5-9138-4277-8ef6-e870f4d84a8b · outbound

This paper cites CVC-MUSCIMA: A ground-truth of hand- written music score images for writer identification and staff removal.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models CVC-MUSCIMA: A ground-truth of hand- written music score images for writer identification and staff removal

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:12.108776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.318571Z digest=sha256:0c71cc2fcd27d2a456ce3432d0c580d2adf0865ded5de004eccd0888b960aedc

Observation bf1091cd-1469-437c-9c43-c6694acdb18c · outbound

This paper cites Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:04.402673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:04.402673Z digest=sha256:cd4910326577f57e4119c0f19491fd71d16ddf3e9aa258ada1705ec1e555f00b

Observation 25e45c23-f6c6-4e85-8d72-e701d2888fbe · outbound

This paper cites Deepscores-a dataset for segmentation, detection and classification of tiny objects.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Deepscores-a dataset for segmentation, detection and classification of tiny objects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.927933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.520127Z digest=sha256:77d14aad98f4e400a7d89a3d58b0187951bbb48842c7d52508e87516a085003c

Observation c8cf186a-0bc7-471d-a910-eb80ebd255d3 · outbound

This paper cites End-to- end neural optical music recognition of monophonic scores.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to- end neural optical music recognition of monophonic scores

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.748658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.652406Z digest=sha256:b21b1b9c2c973f7aa0c10bc81cce7484dbfd253786bf35fe01ab26aba661c033

Observation dde69ff4-d1c0-4e62-bb9f-2b8729f92b64 · outbound

This paper cites DoReMi: First glance at a universal OMR dataset.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DoReMi: First glance at a universal OMR dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:57:08.584722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.841644Z digest=sha256:16b89aa575b9a4e52d84715c69b753252e81419827b91b2b39cd80e190b5aa28

Observation 58138aea-ac0a-4327-bed4-ac749362497e · outbound

This paper cites A uni- fied representation framework for the evaluation of optical music recognition systems.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models A uni- fied representation framework for the evaluation of optical music recognition systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.509970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:04.953302Z digest=sha256:57d9c4745dd7dd8b848aa9f1460c9ab73f7ec1e60da76ce51ce0cd6b9533718f

Observation b9ccfb9a-77e1-4966-b55f-2c25af092de3 · outbound

This paper cites Practical end-to-end optical music recognition for pianoform music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Practical end-to-end optical music recognition for pianoform music

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.238511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.038931Z digest=sha256:1d81ff85518eea07494aa3f0f270ed0f9ed065f3f3536f91b7bf2d3bffdffec8

Observation 0099d064-ce21-4194-ad6f-d8db0da76193 · outbound

This paper cites Breezewhite/oemer: v0.1.7, October 2023.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Breezewhite/oemer: v0.1.7, October 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:11.038775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.155593Z digest=sha256:a5b7e69c2ad18fe00567fdda5655f17c02390a182b08f8f4c2eb6ec7b1efc83b

Observation acb50a2c-7375-4bd9-9de1-923c41045971 · outbound

This paper cites Optical music recognition in manuscripts from the ricordi archive.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition in manuscripts from the ricordi archive

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.853442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.249120Z digest=sha256:ef92a8ecfc63d458e850a0661fd38fabaa217dc750a582b6b28b10b6bf7fea1b

Observation 5d5ad310-4581-4137-910f-76be0515c404 · outbound

This paper cites Optical music recognition with convolutional sequence-to-sequence models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition with convolutional sequence-to-sequence models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.606216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.356627Z digest=sha256:6d838bd3a801d6a5658ccedfab286abc680c93a4b19c5c99baaad3c01aec0074

Observation 3f4661b1-32d6-469b-a1fd-69aa8172e667 · outbound

This paper cites Tromr:transformer-based polyphonic optical music recognition.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tromr:transformer-based polyphonic optical music recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.417118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.523218Z digest=sha256:3c2cde0ab78ad740886ebd0b2a22bcff93c2d3a0405cc9e6c5a184511dea836a

Observation 1f102df9-df64-4f5e-98b4-c40740b6ed78 · outbound

This paper cites Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.244977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:05.682538Z digest=sha256:9e317f3a023a3bae27f6e04fa1e29239a49d3f9242cb2eda8d1c746a76d8f37c

Observation e6bbf8a2-3f79-4115-9454-0853be89d164 · outbound

This paper cites End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:05.787561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:05.787561Z digest=sha256:8c362efa9dfc291be775039d49a59d8c5d4d70609b8c5412946cf6d493583073

Observation b25e32c3-7c57-4510-a441-9fc3db13dafe · outbound

This paper cites ChatMusician: Understanding and Generating Music Intrinsically with LLM.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models ChatMusician: Understanding and Generating Music Intrinsically with LLM

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:05.952693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:05.952693Z digest=sha256:de7ce74d9cff7296612e3ed7e37cd4f6b57560f52190717773be4c1b601fbd0f

Observation 14ff5879-dfb9-45d9-8db0-5ea7db605ac8 · outbound

This paper cites MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.012244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.012244Z digest=sha256:c16f335a2a2e391b099b6516ecc23df3598e184d1b8f8e7df820e1b7577421dc

Observation d66af441-d8e1-4675-89bf-06959580841b · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.098344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.098344Z digest=sha256:e7babb6173ce083e238064c754404da7487b2aee549125f2919238f21e8b8c03

Observation df89a664-d190-445b-81c5-1049c1ca726c · outbound

This paper cites GPT-4o System Card.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models GPT-4o System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.179397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.179397Z digest=sha256:a75ada3de47a62c8b6a2d401c2da2f0968b903e81dff7103ff80bc7e19e4a6ae

Observation a7908338-7a3c-4c00-b9bc-dd65d225b05f · outbound

This paper cites DeepSeek-V3 Technical Report.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DeepSeek-V3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.325770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.325770Z digest=sha256:6b390303feb179748c706940a87caea7f5949271d291ad319dc5eda29ff4e78a

Observation 369ae79e-0a88-43c2-9875-83f009103400 · outbound

This paper cites Musical scales and the generalized circle of fifths.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Musical scales and the generalized circle of fifths

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:10.090242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:06.425346Z digest=sha256:9eb7b5ad250d95f63764ecfd745a035939c4b83efd7f886877d79909b7778a3a

Observation 0c53bfd0-4e4c-4c08-995b-4f015c15d54b · outbound

This paper cites MusiXTEX.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusiXTEX

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.932628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:06.520548Z digest=sha256:94d0d414cdf4416a8f12132b3f571c8cfa0b65154b0e769e38c5a884d0ad0607

Observation 606c7818-3166-49b8-b16d-772e748a5a56 · outbound

This paper cites Harmonic experience: Tonal harmony from its natural origins to its modern expression.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Harmonic experience: Tonal harmony from its natural origins to its modern expression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.751371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:06.597284Z digest=sha256:ac7ca6f4a768752c1796a2a839fb6c7ace123d8e5ac4e8f3c14c9cb2d829045c

Observation 4676f4b5-d193-46bb-9439-90bd09a41d67 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:06.742697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:06.742697Z digest=sha256:f7a1ea2dc58db3ad2f7c26898346f495affdebcc27a117976a817c2a14f6297f

Observation caa80afb-5dd9-4a46-8e4a-9b2ebfd4c561 · outbound

This paper cites Trins: Towards multimodal language models that can read.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Trins: Towards multimodal language models that can read

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.577261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:06.908892Z digest=sha256:6ee0cca0f79d066875fddd31f904150c6d45829843d8fd3be256bb4904a94d75

Observation 97993c1c-9a6d-4154-9aff-ae7d5edad3c3 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.053168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.053168Z digest=sha256:59633e3e422e3f277166b1e2b5fed338d760f750bfb283fd5702c5994440ebdc

Observation 4e90a859-a7cb-491c-bffa-8b713cd65997 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.157500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.157500Z digest=sha256:70a9f3d0ef946408616b3f46979023b2b5e6509ceead0d734584716583854ede

Observation c63f5167-f5c7-4aaa-978c-3c7e8c33e36c · outbound

This paper cites Music information processing using the humdrum toolkit: Concepts, examples, and lessons.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Music information processing using the humdrum toolkit: Concepts, examples, and lessons

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.407200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:07.301593Z digest=sha256:86fcbfd5efd5a83478705b1bf52c88068bb2ea0dc142048fafa1f573f801ae52

Observation 1bdd0ad0-d2cc-4fb0-9f03-9aee76356e82 · outbound

This paper cites MMR: Evaluating Reading Ability of Large Multimodal Models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MMR: Evaluating Reading Ability of Large Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.432546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.432546Z digest=sha256:c355896f563b11832537c5cdc9bb027fec72eddbb381b9d8b7304a9b03ae19d9

Observation 7e0f840c-504c-4ca6-b0f3-13ac051aa37e · outbound

This paper cites Decoupled Weight Decay Regularization.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Decoupled Weight Decay Regularization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.577643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.577643Z digest=sha256:a0ab343be34f8742749ac59607a7b1116e7ef7a10deb7a8ee1d86c3d57ddfc7b

Observation 8b1be671-60c9-4af2-9463-b46312da531f · outbound

This paper cites Retrieval-augmented generation for knowledge- intensive nlp tasks.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Retrieval-augmented generation for knowledge- intensive nlp tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.694170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.694170Z digest=sha256:2c66bd02a9d2d247fa7e043cb947bb73e1a6b4895500a42999fa34ab6f7e6a30

Observation ffefab5a-ed6e-485d-8181-41742bf42140 · outbound

This paper cites Layoutgpt: Compositional visual planning and generation with large language models.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Layoutgpt: Compositional visual planning and generation with large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.227377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:07.851431Z digest=sha256:11c42c173925e414cab15a6c76ba0e8da7afe1ec4b4a26dbdce51a03c4eb18aa

Observation a1c2e789-ffe0-480d-9692-972da2b8b98d · outbound

This paper cites TextLap: Customizing Language Models for Text-to-Layout Planning.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models TextLap: Customizing Language Models for Text-to-Layout Planning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:07.934980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:07.934980Z digest=sha256:77dd6d2f037052da332c4cfda173b054821054928b04998305b6344ea8e3c14c

Observation 29ef0fcd-2bc9-4672-916c-740cfec8368e · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:08.059391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:08.059391Z digest=sha256:384429cdda734ccc44bcdf0482a2a8ee6a525c04cf4effe4041912754a6b5831

Observation b87514ec-34f8-4519-91d0-5228c6da73f1 · outbound

This paper cites Information not found.

MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Information not found

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:57:09.066912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:57:08.122917Z digest=sha256:67520702d03172daaa8f7128d38c5a7b82d5ac33b5734aeb3da5277343f97369

Pith citing papers

Observation d96761d1-6795-4eab-9920-7f5e5c1453f1 · inbound

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence cites this paper.

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.568533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:03:01.356594Z digest=sha256:1869e94e4b3c78051801785a4958bba77428a30ac0f4ab1a48ec2fdf56df5b8b

Observation d3580fe4-5cda-4565-a885-f9ca87d51454 · inbound

Direct content-based retrieval from music scores images cites this paper.

Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:24:43.335938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:21:33.390581Z digest=sha256:1b63c405cf5157a0e3da6e8b9e5f3ac0e72e464a9bde69234b4dec89ec9f4aa1

Observation eb9799a3-f455-45e9-8103-3613ff422f8a · inbound

Direct content-based retrieval from music scores images cites this paper.

Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.953564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:23:39.285332Z digest=sha256:c8a32ee02246d92daaf42f2c8a69c3a2be60b495b0c435c3628ed1e25aeb21ef

Observation 9c192de3-7c3e-4328-ba8d-4758b25ca7c8 · inbound

LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding cites this paper.

LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T02:17:46.326748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T02:12:58.120295Z digest=sha256:ff66e12943073ac87b7e2eb74bb19483e9d14b98be51fbbc8bab1fe78013b56d

Observation 297aee09-8d2f-435a-a6c9-577e60120dae · inbound

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music cites this paper.

Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T19:35:32.938812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T19:28:58.523932Z digest=sha256:3413e64d2c29355b23aab9d90654f63e89a6af2d4c9c2e71eb123d8ac09a45cd