REVIEW 2 minor 12 references
Pretrained music embeddings outperform from-scratch models on cross-performance jazz standard recognition but remain sensitive to performer identity.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-02 06:11 UTC pith:ZOMXJZMZ
load-bearing objection Pretrained embeddings beat from-scratch models on jazz standard retrieval but still pick up performer identity, with a simple contrastive fix helping only partially.
Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
From-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-k results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. The findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification.
What carries the argument
The lightweight contrastive projection applied to frozen pretrained embeddings to reduce performer identity sensitivity during nearest-neighbor retrieval.
Load-bearing premise
The curated subset of the Jazz Trio Database adequately captures the full range of cross-performance variation without introducing dataset-specific biases that would not generalize to other jazz recordings.
What would settle it
If pretrained embeddings no longer outperform from-scratch models or the contrastive projection fails to reduce performer effects when tested on an independent set of jazz recordings with unseen performers and arrangements, the central claims would not hold.
If this is right
- Pretrained embeddings are more suitable than from-scratch training for cross-performance music retrieval tasks.
- Performer identity sensitivity in embeddings can be addressed with simple projection layers.
- Jazz standard recognition provides a practical benchmark for assessing generalization in music foundation models.
- Retrieval methods can advance toward identifying standards without needing full supervision.
Where Pith is reading between the lines
- Similar performer biases may affect embedding performance in other variable music genres such as classical or folk.
- Training future models with explicit cross-performance contrastive losses could further improve invariance to performer and arrangement.
- The evaluation approach might extend to identifying covers or alternate versions in non-jazz music retrieval.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates pretrained music embeddings for cross-performance jazz standard recognition on a curated subset of the Jazz Trio Database. It compares a from-scratch Harmonic CNN baseline against frozen pretrained representations from music foundation models via supervised probing and nearest-neighbor retrieval, concluding that from-scratch models overfit to training performances while pretrained embeddings yield better top-k results but remain sensitive to performer identity (partially mitigated by a lightweight contrastive projection). The work positions the task as a stress test for music representations.
Significance. If the empirical results hold under scrutiny of the methods and data splits, the findings would usefully demonstrate the generalization advantages of pretrained embeddings over task-specific training on limited jazz data and provide a practical technique (contrastive projection) for reducing performer sensitivity. This contributes an application-driven benchmark that could help evaluate future music foundation models on real-world variation in tempo, key, arrangement, and improvisation.
minor comments (2)
- The abstract references specific models and a 'curated subset' but does not name the exact pretrained embeddings, the size of the subset, or the train/test split criteria; these details are needed to assess reproducibility and the strength of the overfitting claim.
- No quantitative results (e.g., top-k accuracies, statistical significance) appear in the provided abstract, making it impossible to evaluate the magnitude of the reported improvements or the effectiveness of the contrastive projection.
Simulated Author's Rebuttal
We thank the referee for their careful reading and for recognizing the potential value of jazz standard recognition as a stress test for music representations. We address the major comments below.
Circularity Check
No significant circularity
full rationale
This is an empirical comparison study that trains or probes models on a curated dataset and reports observed top-k retrieval metrics. No equations, derivations, or first-principles claims appear; results are presented as experimental observations rather than predictions derived from the model itself. The central claim (pretrained embeddings outperform from-scratch models but remain performer-sensitive) rests on direct measurement, not on any self-definitional mapping, fitted-input renaming, or self-citation chain that reduces the result to its inputs by construction.
Axiom & Free-Parameter Ledger
read the original abstract
Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-$k$ results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: https://github.com/cagries/tipofmyear.
Figures
Reference graph
Works this paper leans on
-
[1]
Bertin-Mahieux, Thierry and Ellis, Daniel P. W. and Whitman, Brian and Lamere, Paul , booktitle =. The
-
[2]
Proceedings of the 4th International Society for Music Information Retrieval Conference , year =
An Industrial-Strength Audio Search Algorithm , author =. Proceedings of the 4th International Society for Music Information Retrieval Conference , year =
-
[3]
Characterization and exploitation of community structure in cover song networks
Characterization and Exploitation of Community Structure in Cover Song Networks , author =. arXiv preprint arXiv:1108.6003 , year =
work page internal anchor Pith review Pith/arXiv arXiv
-
[4]
Xun, Jiahao and Zhang, Shengyu and Yang, Yanting and Zhu, Jieming and Deng, Liqun and Zhao, Zhou and Dong, Zhenhua and Li, Ruiqi and Zhang, Lichao and Wu, Fei , journal =. 2023 , url =
work page 2023
-
[5]
Transactions of the International Society for Music Information Retrieval , volume =
Jazz Trio Database: Automated Annotation of Jazz Piano Trio Recordings Processed Using Audio Source Separation , author =. Transactions of the International Society for Music Information Retrieval , volume =. 2024 , doi =
work page 2024
-
[6]
Li, Yizhi and Yuan, Ruibin and Zhang, Ge and Ma, Yinghao and Chen, Xingran and Yin, Hanzhi and Xiao, Chenghao and Lin, Chenghua and Ragni, Anton and Benetos, Emmanouil and Gyenge, Norbert and Dannenberg, Roger B. and Liu, Ruibo and Chen, Wenhu and Xia, Gus and Shi, Yemin and Huang, Wenhao and Wang, Zili and Guo, Yike and Fu, Jie , journal =. 2024 , url =
work page 2024
-
[7]
Zhu, Haina and Zhou, Yizhi and Chen, Hangting and Yu, Jianwei and Ma, Ziyang and Gu, Rongzhi and Luo, Yi and Tan, Wei and Chen, Xie , journal =. 2025 , url =
work page 2025
-
[8]
Advances in Neural Information Processing Systems 33 , year =
Supervised Contrastive Learning , author =. Advances in Neural Information Processing Systems 33 , year =
-
[9]
arXiv preprint arXiv:2506.17055 , year =
Universal Music Representations? Evaluating Foundation Models on World Music Corpora , author =. arXiv preprint arXiv:2506.17055 , year =
-
[10]
arXiv preprint arXiv:2107.05677 , year =
Codified Audio Language Modeling Learns Useful Representations for Music Information Retrieval , author =. arXiv preprint arXiv:2107.05677 , year =
-
[11]
arXiv preprint arXiv:2210.03799 , year =
Supervised and Unsupervised Learning of Audio Representations for Music Understanding , author =. arXiv preprint arXiv:2210.03799 , year =
-
[12]
Data-Driven Harmonic Filters for Audio Representation Learning , author =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.