Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:57:08.122917Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2506.23009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:57:08.122917Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T02:12:58.120295Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T02:17:46.305033Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f5cc43c0-2f24-47c8-9aa3-c5a96aeb3049 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Acrobat AI Assistant, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 940c5739-7079-4665-a70c-b0ff2e4571b7 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e24319-069e-4c41-8abe-6d1fc94fce4e · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The basics of reading music
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccb2d247-30ec-4ed1-99b3-f219450ede58 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Reading sheet music facilitates sensorimotor mu- desynchronization in musicians
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 444834fa-8cde-4875-89f3-9e5004298c0b · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical Music Recognition: State of the Art and Major Challenges
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b20ed1a5-0a3d-49c0-87c4-c2fcf0ba7def · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Understanding optical music recognition.ACM Computing Surveys (CSUR), 53(4):1–35, 2020
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ad181d0-58c5-4a4d-a5c6-6942d6910ebe · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition: state-of-the-art and open issues
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cedc974-803e-4108-9133-15be13d1487a · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models The challenge of opti- cal music recognition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a453cfc9-af66-446f-9777-0952ee80cedc · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition using pro- jections
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60d3e49a-7c68-4e09-9e4d-418e773af6dc · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Gui agents: A survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea595109-5701-468c-b1e6-3ea06fd2ff0f · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Natural language understand- ing and inference with mllm in visual question answer- ing: A survey
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a9369b0-1cf4-4a8c-be2d-cce557455e5b · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 320ef79e-a083-4b28-85bf-72cc75c4383a · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2550e8-38f8-4dd0-8235-07e20b3ea904 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision- language benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8f6ce39-1320-4708-a18d-31ee29f69ce0 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PP-OCR: A Practical Ultra Lightweight OCR System
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a612c9c-cc42-4bb0-a4a4-3444743860b3 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tex- tocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74b8b8f5-9138-4277-8ef6-e870f4d84a8b · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models CVC-MUSCIMA: A ground-truth of hand- written music score images for writer identification and staff removal
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf1091cd-1469-437c-9c43-c6694acdb18c · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Knowledge Discovery in Optical Music Recognition: Enhancing Information Retrieval with Instance Segmentation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e45c23-f6c6-4e85-8d72-e701d2888fbe · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Deepscores-a dataset for segmentation, detection and classification of tiny objects
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8cf186a-0bc7-471d-a910-eb80ebd255d3 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to- end neural optical music recognition of monophonic scores
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dde69ff4-d1c0-4e62-bb9f-2b8729f92b64 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DoReMi: First glance at a universal OMR dataset
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58138aea-ac0a-4327-bed4-ac749362497e · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models A uni- fied representation framework for the evaluation of optical music recognition systems
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9ccfb9a-77e1-4966-b55f-2c25af092de3 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Practical end-to-end optical music recognition for pianoform music
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0099d064-ce21-4194-ad6f-d8db0da76193 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Breezewhite/oemer: v0.1.7, October 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acb50a2c-7375-4bd9-9de1-923c41045971 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition in manuscripts from the ricordi archive
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d5ad310-4581-4137-910f-76be0515c404 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Optical music recognition with convolutional sequence-to-sequence models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f4661b1-32d6-469b-a1fd-69aa8172e667 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Tromr:transformer-based polyphonic optical music recognition
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f102df9-df64-4f5e-98b4-c40740b6ed78 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Sheet music transformer: End-to-end optical music recognition beyond monophonic transcription, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6bbf8a2-3f79-4115-9454-0853be89d164 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25e32c3-7c57-4510-a441-9fc3db13dafe · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models ChatMusician: Understanding and Generating Music Intrinsically with LLM
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ff5879-dfb9-45d9-8db0-5ea7db605ac8 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d66af441-d8e1-4675-89bf-06959580841b · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models PaliGemma 2: A Family of Versatile VLMs for Transfer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df89a664-d190-445b-81c5-1049c1ca726c · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models GPT-4o System Card
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7908338-7a3c-4c00-b9bc-dd65d225b05f · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models DeepSeek-V3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 369ae79e-0a88-43c2-9875-83f009103400 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Musical scales and the generalized circle of fifths
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c53bfd0-4e4c-4c08-995b-4f015c15d54b · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MusiXTEX
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 606c7818-3166-49b8-b16d-772e748a5a56 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Harmonic experience: Tonal harmony from its natural origins to its modern expression
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4676f4b5-d193-46bb-9439-90bd09a41d67 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa80afb-5dd9-4a46-8e4a-9b2ebfd4c561 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Trins: Towards multimodal language models that can read
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97993c1c-9a6d-4154-9aff-ae7d5edad3c3 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e90a859-a7cb-491c-bffa-8b713cd65997 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63f5167-f5c7-4aaa-978c-3c7e8c33e36c · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Music information processing using the humdrum toolkit: Concepts, examples, and lessons
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bdd0ad0-d2cc-4fb0-9f03-9aee76356e82 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models MMR: Evaluating Reading Ability of Large Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0f840c-504c-4ca6-b0f3-13ac051aa37e · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Decoupled Weight Decay Regularization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1be671-60c9-4af2-9463-b46312da531f · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Retrieval-augmented generation for knowledge- intensive nlp tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffefab5a-ed6e-485d-8181-41742bf42140 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Layoutgpt: Compositional visual planning and generation with large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1c2e789-ffe0-480d-9692-972da2b8b98d · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models TextLap: Customizing Language Models for Text-to-Layout Planning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ef0fcd-2bc9-4672-916c-740cfec8368e · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87514ec-34f8-4519-91d0-5228c6da73f1 · outbound
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models Information not found
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d96761d1-6795-4eab-9920-7f5e5c1453f1 · inbound
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3580fe4-5cda-4565-a885-f9ca87d51454 · inbound
Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb9799a3-f455-45e9-8103-3613ff422f8a · inbound
Direct content-based retrieval from music scores images MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c192de3-7c3e-4328-ba8d-4758b25ca7c8 · inbound
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 297aee09-8d2f-435a-a6c9-577e60120dae · inbound
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.