REVIEW 2 major objections 8 minor 24 references
Best music AI transcribes pop songs at only 38% accuracy
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-10 01:42 UTC pith:GZSMIELU
load-bearing objection Solid benchmark filling a real gap, but ground-truth fidelity is the load-bearing question the 2 major comments →
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central finding is that automatic music transcription systems trained primarily on synthetic audio perform poorly on real commercial pop music, with the best model reaching only 38% Onset F1 on a newly constructed benchmark of 572 beat-aligned multitrack MIDI segments. The dataset itself is the core contribution: it is the first to pair fully-produced commercial pop audio with time-aligned multitrack MIDI across diverse genres and decades, filling a gap left by datasets that are either synthetic, single-instrument, or limited in genre diversity.
What carries the argument
The dataset construction pipeline has three stages: (1) metadata matching between the TheoryTab dataset (which provides YouTube audio segments with user annotations) and the Lakh MIDI Dataset (which provides multitrack MIDI), using Levenshtein distance on artist and title fields; (2) beat-based time alignment, where RNN-based beat tracking on the audio is matched to the MIDI beat grid through an anchor beat that links the first audio beat to a specific MIDI beat, with linear interpolation warping all note timings; (3) human annotation, where annotators select the correct anchor beat from up to four algorithmically generated candidates by comparing synthesized MIDI against the original audio.
Load-bearing premise
The dataset construction assumes that the Lakh MIDI files constitute a perfect multitrack transcription of the original commercial recordings. If the MIDI files contain errors, omissions, or arrangement differences relative to the audio, the ground truth labels are noisy, and the reported model scores conflate model failure with label imperfection.
What would settle it
A model that transcribes real pop music correctly but is penalized because the Lakh MIDI ground truth differs from the actual recording arrangement, or conversely, a model that scores well on MulTTiPop by exploiting systematic artifacts in the MIDI-to-audio alignment rather than genuinely transcribing.
If this is right
- The 38% F1 ceiling suggests that training transcription models on synthetic audio (e.g., Slakh2100) does not transfer well to commercial recordings, pointing to a domain gap in timbre and arrangement complexity.
- MulTTiPop can serve as a standard evaluation benchmark for future multitrack AMT systems, enabling direct comparison on real-world pop music.
- The human-verified beat-alignment methodology could be adapted to create larger multitrack transcription datasets by scaling the metadata-matching and anchor-beat pipeline.
Where Pith is reading between the lines
- If the Lakh MIDI files used as ground truth contain arrangement differences or transcription errors relative to the commercial recordings, the 38% F1 score may partly reflect label noise rather than pure model failure, meaning actual model competence on real pop music could be somewhat higher (or lower) than reported.
- The 49.1% annotation success rate suggests the pipeline could be improved with better automatic alignment methods, potentially yielding a substantially larger dataset without proportional increases in human labor.
- The gap between synthetic-training performance and commercial-audio performance implies that data augmentation strategies bridging the timbre gap such as style transfer or more realistic synthesis could be a productive research direction beyond architectural improvements alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MulTTiPop, a dataset of 572 segments (3.5 hours) of commercial pop music audio paired with time-aligned multitrack MIDI transcriptions, intended as an evaluation benchmark for automatic music transcription (AMT) models. The construction pipeline matches audio segments from TheoryTab to multitrack MIDI files from the Lakh MIDI Dataset via metadata matching, beat-based time alignment, and human annotation of anchor beats. The authors evaluate two state-of-the-art AMT models (MT3 and YourMT3+) on the dataset, reporting a best Onset F1 of 38%, and argue this demonstrates substantial room for improvement on real-world multitrack transcription. The dataset fills a genuine gap: existing multitrack AMT datasets use synthesized audio (Slakh2100), while datasets with commercial audio (TheoryTab, McGill-Billboard) lack full multitrack MIDI, and RWC-Pop has limited genre diversity.
Significance. The primary contribution is the dataset itself, which addresses a well-motivated gap in the AMT evaluation landscape. The construction methodology—combining metadata matching, beat-based warping, and human-in-the-loop anchor beat selection—is clearly described and reproducible. The honest reporting of the 49.1% alignment success rate and the breakdown of failure modes (Table 2) is commendable. The evaluation of MT3 and YourMT3+ provides a useful baseline. The dataset's diversity (374 songs, 263 artists, 101 genres) exceeds RWC-Pop despite its smaller size. The recommendation against using the dataset for training (Section 5) is responsible given copyright concerns. The main limitation, discussed below, is the absence of note-level ground-truth verification of the Lakh MIDI files themselves.
major comments (2)
- The benchmark's validity hinges on Lakh MIDI files being note-accurate transcriptions of the commercial audio, but this is only verified perceptually by student annotators, not at the note level. Section 2.4 states annotators were told 'we expect the synthesized MIDI to be a perfect multitrack transcription' and asked to reject samples without a perfect match. However, the Lakh MIDI Dataset consists of user-generated arrangements, not professionally verified transcriptions. A MIDI arrangement can sound perceptually similar to a song while containing incorrect pitches, missing instruments (especially vocals, which are rarely fully transcribed to MIDI), simplified rhythms, or extra notes. If the ground truth itself has, say, 70% note accuracy relative to the true audio content, then the 38% F1 of AMT models reported in Table 4 is confounded: we cannot distinguish model failure from label噪声
- Section 3.1 reports that annotators selected one of the candidate alignments for 49.1% of segments, rejecting 50.9%. Of the rejections, 12.2% were attributed to metadata matching issues, leaving ~38.7% as alignment failures on correct-song matches. The paper does not analyze whether the retained 49.1% is representative of the broader population of pop music segments or whether it is biased toward songs with simpler arrangements, clearer beats, or more faithful MIDI transcriptions. A brief analysis of the rejected segments (e.g., comparing tempo, genre, or instrumentation complexity) would help assess selection bias and clarify the dataset's scope.
minor comments (8)
- Table 1: The 'Size (Hours)' column appears to have a formatting issue where the numeric value and unit run together for some entries (e.g., '34 330' for MusicNet). Consider adding a column separator or reformatting.
- Section 2.2: The notation for beats b_a and b_m uses subscripts that are somewhat hard to parse. Consider using b^{audio} and b^{midi} or explicitly defining the subscript notation.
- Section 3.2: 'Dsepite' should be 'Despite'.
- Appendix B.1: The use of 'Claude Sonnet 4.6' for release year lookup is unusual; consider verifying a sample of these dates manually or noting the potential for hallucinated dates.
- Section 4: The paper evaluates only MT3 and YourMT3+. While these are reasonable choices, including a baseline that is not transformer-based (e.g., a classical DSP method) could provide additional context for the difficulty of the task.
- Section 4.1: The qualitative analysis mentions YourMT3+ uses 'saxophone' for vocal lines. It would help to clarify whether this is a known artifact of the model's training data or a novel observation.
- The paper does not report inter-annotator agreement among the 6 annotators beyond the 70% threshold on the hidden control set. Reporting Cohen's kappa or similar would strengthen the annotation quality assessment.
- Table 3: The genre count for RWC (All Modern) is 43, while MulTTiPop is 101. The methodology for counting genres differs between datasets (Every Noise at Once vs. dataset descriptions), making this comparison somewhat apples-to-oranges. A note acknowledging this would be helpful.
Circularity Check
No circularity: dataset is constructed from external sources and evaluated with external models; ground truth set by human annotators, not by fitted metrics
full rationale
The paper introduces a dataset (MulTTiPop) and evaluates two external AMT models (MT3, YourMT3+) on it. The dataset construction pipeline uses metadata matching, beat-based alignment, and human annotation to select anchor beats. The alignment metrics (chroma similarity, onset correlation, melody matching) are used only to generate candidate anchor beats for human review—they do not define the ground truth labels. Human annotators independently select or reject candidates based on perceptual comparison of synthesized MIDI against original audio. The evaluation models (MT3, YourMT3+) are external, not authored by the paper's authors, and their 38% F1 scores are measured against the human-verified labels. No step in the derivation chain reduces to its own inputs by construction. The concern about Lakh MIDI fidelity (raised in the skeptic attack) is a correctness/label-noise concern, not a circularity issue—the paper does not claim to derive MIDI quality from a metric that was itself fitted to that quality. The paper is self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
free parameters (5)
- Levenshtein distance threshold for metadata matching =
near-identical (unspecified threshold)
- Number of candidate anchor beats =
4
- YouTube timing filter window =
10 seconds
- Melody vs. base metric weighting =
75% / 25%
- Annotator quality control threshold =
70%
axioms (3)
- domain assumption Lakh MIDI files contain accurate multitrack transcriptions of the original commercial recordings.
- domain assumption Beat tracking on commercial pop audio is sufficiently accurate to enable linear-interpolation warping of MIDI.
- domain assumption Onset F1 with 50ms tolerance is an appropriate metric for multitrack transcription quality.
read the original abstract
We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.
Reference graph
Works this paper leans on
-
[1]
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
INTRODUCTION In recent years, automatic music transcription (AMT) systems for converting audio to note-level symbolic representations of music have evolved from transcribing solo piano music [9] to targeting performance on a wide variety of instruments and musical styles [10]. However, current AMT models do not yet meet the task of multitrack transcriptio...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[2]
We first perform metadata matching between audio segments and MIDI files
METHODOLOGY We outline our approach to aligning audio segments sourced from TheoryTab to multitrack MIDI in the Lakh MIDI Dataset. We first perform metadata matching between audio segments and MIDI files. Then, we time align the MIDI and audio by synchronizing the beat grid of the MIDI with beats detected in the audio. Finally, we identify several candida...
-
[3]
DATASET The resulting MulTTiPop dataset contains 572 segments, totaling approximately 3.5 hours of audio. Segments come from 374 unique songs and 263 unique artists. Each seg- ment is 22 seconds on average, and contains an average of 583 MIDI notes. A preview of MulTTiPop including syn- chronized YouTube video and MIDI playback is available at https://gcl...
work page 1939
-
[4]
exact", where instrument labels must match the same MIDI program as defined in the dataset, and
AMT MODEL EV ALUATION We run two high-performing open weights transcription models—MT3 [10] and YourMT3+ [11]—and report the Onset F1 of both models in Table 4. We evaluate these Exact Harmonic-Percussion Model Precision Recall Onset F1 Precision Recall Onset F1 MT3 [10]31.03 28.18 28.42 39.55 37.10 36.83 YourMT3+ [11]29.51 24.10 25.29 43.13 36.65 37.87 T...
-
[5]
RECOMMENDATIONS FOR USE MulTTiPop is strictly designed as anevaluation dataset for automatic music transcription models on multitrack pop music. MulTTiPop is not designed for the training of AMT systems or other machine learning models, especially gener- ative models, both because of the small size of the data and because the labels pertain to copyrighted...
-
[6]
ACKNOWLEDGEMENTS This work was supported by funding from Sony AI, and we thank our Sony AI collaborators for helpful conversations. We are also grateful for the contributions of our data anno- tators at Carnegie Mellon University: Alex Cheng, Woody Li, Arisa Okamura, Sean Xue, Lynn Ye, and Ryan Zhang
-
[7]
Learning Features of Music from Scratch
John Thickstun, Zaid Harchaoui, and Sham Kakade, “Learning features of music from scratch,”arXiv preprint arXiv:1611.09827, 2016
work page internal anchor Pith review Pith/arXiv arXiv 2016
-
[8]
Enabling factorized piano music modeling and generation with the MAESTRO dataset,
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” inInternational Conference on Learning Representations, 2019
work page 2019
-
[9]
Ethan Manilow, Gordon Wichern, Prem Seetharaman, and Jonathan Le Roux, “Cutting music source sep- aration some slakh: A dataset to study the impact of training data quality and quantity,” in2019 IEEE Work- shop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2019, pp. 45–49
work page 2019
-
[10]
An expert ground truth set for audio chord recog- nition and music analysis.,
John Ashley Burgoyne, Jonathan Wild, and Ichiro Fuji- naga, “An expert ground truth set for audio chord recog- nition and music analysis.,” inISMIR, 2011, vol. 11, pp. 633–638
work page 2011
-
[11]
POP909: A Pop-song Dataset for Music Arrangement Generation
Ziyu Wang, Ke Chen, Junyan Jiang, Yiyi Zhang, Mao- ran Xu, Shuqi Dai, Xianbin Gu, and Gus Xia, “Pop909: A pop-song dataset for music arrangement generation,” arXiv preprint arXiv:2008.07142, 2020
work page internal anchor Pith review Pith/arXiv arXiv 2008
-
[12]
Melody transcription via generative pre-training,
Chris Donahue, John Thickstun, and Percy Liang, “Melody transcription via generative pre-training,” in ISMIR, 2022
work page 2022
-
[13]
On the preparation and validation of a large-scale dataset of singing transcription,
Jun-You Wang and Jyh-Shing Roger Jang, “On the preparation and validation of a large-scale dataset of singing transcription,” inICASSP 2021-2021 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 276–280
work page 2021
-
[14]
Rwc music database: Popular, classical and jazz music databases.,
Masataka Goto, Hiroki Hashiguchi, Takuichi Nishimura, and Ryuichi Oka, “Rwc music database: Popular, classical and jazz music databases.,” inIsmir, 2002, vol. 2, pp. 287–288
work page 2002
-
[15]
Onsets and Frames: Dual-Objective Piano Transcription
Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck, “Onsets and frames: Dual-objective piano transcription,”arXiv preprint arXiv:1710.11153, 2017
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[16]
MT3: Multi-Task Multitrack Music Transcription
Josh Gardner, Ian Simon, Ethan Manilow, Cur- tis Hawthorne, and Jesse Engel, “Mt3: Multi- task multitrack music transcription,”arXiv preprint arXiv:2111.03017, 2021
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[17]
Sungkyun Chang, Emmanouil Benetos, Holger Kirch- hoff, and Simon Dixon, “Yourmt3+: Multi-instrument music transcription with enhanced transformer architec- tures and cross-dataset stem augmentation,” in2024 IEEE 34th International Workshop on Machine Learn- ing for Signal Processing (MLSP). IEEE, 2024, pp. 1–6
work page 2024
-
[18]
Colin Raffel,Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching, Columbia University, 2016
work page 2016
-
[19]
Thierry Bertin-Mahieux, Daniel PW Ellis, Brian Whit- man, and Paul Lamere, “The million song dataset,” 2011
work page 2011
-
[20]
Enhanced beat tracking with context-aware neural networks,
Sebastian Böck and Markus Schedl, “Enhanced beat tracking with context-aware neural networks,” inProc. Int. Conf. Digital Audio Effects, 2011, pp. 135–139
work page 2011
-
[21]
madmom: a new Python Audio and Music Signal Processing Library,
Sebastian Böck, Filip Korzeniowski, Jan Schlüter, Flo- rian Krebs, and Gerhard Widmer, “madmom: a new Python Audio and Music Signal Processing Library,” in Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands, 10 2016, pp. 1174–1178
work page 2016
-
[22]
Rwc music database: Music genre database and musical instrument sound database,
Masataka Goto, Hiroki Hashiguchi, Takuichi Nishimura, and Ryuichi Oka, “Rwc music database: Music genre database and musical instrument sound database,” 2003
work page 2003
-
[23]
Evaluation of multiple-f0 estimation and tracking sys- tems.,
Mert Bay, Andreas F Ehmann, and J Stephen Downie, “Evaluation of multiple-f0 estimation and tracking sys- tems.,” inISMIR, 2009, pp. 315–320
work page 2009
-
[24]
Guitarset: A dataset for guitar transcription.,
Qingyang Xi, Rachel M Bittner, Johan Pauwels, Xuzhou Ye, and Juan Pablo Bello, “Guitarset: A dataset for guitar transcription.,” Proceedings of the 19th Interna- tional Society for Music Information Retrieval Confer- ence, 2018. A. ANNOTATION INTERFACE Data annotators are provided a Jupyter notebook (hosted lo- cally or on Google Colab) to perform manual ...
work page 2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.